如何搜索多个pdf文件的内容?

如何在目录/子目录中搜索PDF文件的内容?我在找一些命令行工具。grep似乎不能搜索PDF文件。

当前回答

Recoll是一个很棒的Unix/Linux全文GUI搜索应用程序，支持几十种不同的格式，包括PDF。它甚至可以将查询的确切页码和搜索词传递给文档查看器，从而允许您直接从它的GUI跳转到结果。

Recoll还提供了一个可行的命令行界面和一个web浏览器界面。

2013-05-29 11:59:04

其他回答

你需要一些工具，如pdf2text，首先将pdf转换为文本文件，然后在文本中搜索。(您可能会错过一些信息或符号)。

如果你正在使用一种编程语言，很可能有专门为此目的编写的pdf库。例如:http://search.cpan.org/dist/CAM-PDF/ for Perl

2011-01-10 03:43:07

谢谢所有的好主意!

我尝试了xargs方法，但正如这里所指出的，xargs将使它不可能(或非常困难)包括打印实际的文件名……

所以我尝试了GNU并行。

parallel "pdftotext -q {} - | grep --with-filename --label='['{}']' --color=always --context=5 'pattern'" ::: *.pdf

This prints not only the pattern, but with --context=5 also 5 lines above and below as well for context. With -q pdftotext won't print any error messages or warnings (quiet). I use brackets [] as labels instead of braces {}. If you wanted braces --label='{'{}'}' will make that happen. Note that {} is replaced by the actual filename by GNU parallel, e.g. 'Example portable document file name with spaces.pdf' ({} is already using single quotes '). By using --label={} only the filename will be printed, which may be the favored way of displaying the filename. I also noticed that the output was without color when I tried it, except when forcing it by adding --color=always with grep. It may be useful to add --ignore-case to the grep command for a case-insensitive keyword search.

如果所有PDF文件都应该递归处理，包括当前目录(.)中的所有子目录，这可以通过find来完成:

find . -type f -iname '*.pdf' -print0 | parallel -0 "pdftotext -q {} - | grep --with-filename --label='['{}']' --color=always --context=5 'pattern'"

With find, -iname '*.pdf' acts case-insensitive. With -name '*.pdf' only lower-case .pdf files will be included (the normal case). Since I sometimes also encountered Windows PDF-files with an upper-case .PDF file extension, I tend to prefer -iname... The above command also works with the -print find option (instead of -print0), so it will be line-based (one file name per line), then -0 (NUL delimiter) must be omitted from the parallel command. Again, including --ignore-case in the grep command will make the search case-insensitive.

作为处理整个命令行的一般建议，parallel -dry-run将打印将要执行的命令。

$ find . -type f -iname '*.pdf' -print0 | parallel --dry-run -0 "pdftotext -q {} - | grep --with-filename --label='['{}']' --color=always --ignore-case --context=5 'pattern'"
pdftotext -q ./test PDF file 1.pdf - | grep --with-filename --label='['./test PDF file 1.pdf']' --color=always --ignore-case --context=5 'pattern'
pdftotext -q ./subdir1/test PDF file 2.pdf - | grep --with-filename --label='['./subdir1/test PDF file 2.pdf']' --color=always --ignore-case --context=5 'pattern'
pdftotext -q ./subdir2/test PDF file 3.pdf - | grep --with-filename --label='['./subdir2/test PDF file 3.pdf']' --color=always --ignore-case --context=5 'pattern'

2022-02-06 15:21:15

你的发行版应该提供一个名为pdftotext的实用程序:

find /path -name '*.pdf' -exec sh -c 'pdftotext "{}" - | grep --with-filename --label="{}" --color "your pattern"' \;

如果要将pdftotext输出到标准输出，而不是输出到文件，则必须使用“-”。 ——with-filename和——label=选项将把文件名放在grep的输出中。可选的——color标志很好，它告诉grep在终端上使用颜色输出。

(在Ubuntu中，pdftotext是由xpdf-utils或poppler-utils包提供的。)

如果您想使用GNU grep中pdfgrep不支持的特性，这种使用pdftotext和grep的方法比pdfgrep更有优势。注意:pdfgrep - 1.3。x支持-C选项打印上下文行。

2011-01-10 03:43:22

Recoll还提供了一个可行的命令行界面和一个web浏览器界面。

2013-05-29 11:59:04

我也遇到了同样的问题，因此我写了一个脚本，搜索指定文件夹中的所有pdf文件的字符串，并打印匹配查询字符串的pdf文件。

也许这对你有帮助。

你可以在这里下载

2012-06-24 14:04:41

如何搜索多个pdf文件的内容?

推荐文章

最新文章

标签