如何搜索多个pdf文件的内容?

如何在目录/子目录中搜索PDF文件的内容?我在找一些命令行工具。grep似乎不能搜索PDF文件。

当前回答

你的发行版应该提供一个名为pdftotext的实用程序:

find /path -name '*.pdf' -exec sh -c 'pdftotext "{}" - | grep --with-filename --label="{}" --color "your pattern"' \;

如果要将pdftotext输出到标准输出，而不是输出到文件，则必须使用“-”。 ——with-filename和——label=选项将把文件名放在grep的输出中。可选的——color标志很好，它告诉grep在终端上使用颜色输出。

(在Ubuntu中，pdftotext是由xpdf-utils或poppler-utils包提供的。)

如果您想使用GNU grep中pdfgrep不支持的特性，这种使用pdftotext和grep的方法比pdfgrep更有优势。注意:pdfgrep - 1.3。x支持-C选项打印上下文行。

2011-01-10 03:43:22

其他回答

还有另一个实用程序叫做ripgrep-all，它是基于ripgrep的。

它不仅可以处理PDF文档，比如Office文档和电影，而且作者声称它比pdfgrep更快。

递归搜索当前目录的命令语法，第二个命令只限制PDF文件:

rga 'pattern' .
rga --type pdf 'pattern' .

2019-07-29 09:06:56

试着在一个简单的脚本中使用'acroread'，就像上面那样

2011-01-10 09:09:49

我喜欢@sjr的答案，但我更喜欢xargs vs -exec。我发现xargs更通用。例如，使用-P，我们可以在必要时利用多个cpu。

find . -name '*.pdf' | xargs -P 5 -I % pdftotext % - | grep --with-filename --label="{}" --color "pattern"

2014-09-26 18:13:38

我写了这个破坏性的小脚本。祝你玩得开心。

function pdfsearch()
{
    find . -iname '*.pdf' | while read filename
    do
        #echo -e "\033[34;1m// === PDF Document:\033[33;1m $filename\033[0m"
        pdftotext -q -enc ASCII7 "$filename" "$filename."; grep -s -H --color=always -i $1 "$filename."
        # remove it!  rm -f "$filename."
    done
}

2011-06-10 15:48:49

Recoll是一个很棒的Unix/Linux全文GUI搜索应用程序，支持几十种不同的格式，包括PDF。它甚至可以将查询的确切页码和搜索词传递给文档查看器，从而允许您直接从它的GUI跳转到结果。

Recoll还提供了一个可行的命令行界面和一个web浏览器界面。

2013-05-29 11:59:04

如何搜索多个pdf文件的内容?

推荐文章

最新文章

标签