使用Python从HTML文件中提取文本

我想使用Python从HTML文件中提取文本。我想从本质上得到相同的输出，如果我从浏览器复制文本，并将其粘贴到记事本。

我想要一些更健壮的东西，而不是使用正则表达式，正则表达式可能会在格式不佳的HTML上失败。我见过很多人推荐Beautiful Soup，但我在使用它时遇到了一些问题。首先，它会抓取不需要的文本，比如JavaScript源代码。此外，它也不解释HTML实体。例如，我会期望'在HTML源代码中转换为文本中的撇号，就像我将浏览器内容粘贴到记事本一样。

更新html2text看起来很有希望。它正确地处理HTML实体，而忽略JavaScript。然而，它并不完全生成纯文本;它产生的降价，然后必须转换成纯文本。它没有示例或文档，但代码看起来很干净。

相关问题:

在python中过滤HTML标签并解析实体在Python中将XML/HTML实体转换为Unicode字符串

当前回答

我发现的最好的一段代码提取文本没有javascript或不想要的东西:

from urllib.request import urlopen
from bs4 import BeautifulSoup

url = "http://news.bbc.co.uk/2/hi/health/2284783.stm"
html = urlopen(url).read()
soup = BeautifulSoup(html, features="html.parser")

# kill all script and style elements
for script in soup(["script", "style"]):
    script.extract()    # rip it out

# get text
text = soup.get_text()

# break into lines and remove leading and trailing space on each
lines = (line.strip() for line in text.splitlines())
# break multi-headlines into a line each
chunks = (phrase.strip() for line in lines for phrase in line.split("  "))
# drop blank lines
text = '\n'.join(chunk for chunk in chunks if chunk)

print(text)

你只需要安装BeautifulSoup:

pip install beautifulsoup4

2014-07-07 19:18:20

其他回答

PyParsing做得很好。PyParsing wiki被杀死了，所以这里有另一个位置，这里有使用PyParsing的示例(示例链接)。花点时间在pyparsing上的一个原因是，他还写了一本非常简短、组织良好的O'Reilly捷径手册，而且价格便宜。

话虽如此，我经常使用BeautifulSoup，处理实体问题并不难，你可以在运行BeautifulSoup之前转换它们。

古德勒克

2008-11-30 15:46:19

Beautiful soup可以转换html实体。考虑到HTML经常有bug并且充满unicode和HTML编码问题，这可能是您最好的选择。这是我用来将html转换为原始文本的代码:

import BeautifulSoup
def getsoup(data, to_unicode=False):
    data = data.replace("&nbsp;", " ")
    # Fixes for bad markup I've seen in the wild.  Remove if not applicable.
    masssage_bad_comments = [
        (re.compile('<!-([^-])'), lambda match: '<!--' + match.group(1)),
        (re.compile('<!WWWAnswer T[=\w\d\s]*>'), lambda match: '<!--' + match.group(0) + '-->'),
    ]
    myNewMassage = copy.copy(BeautifulSoup.BeautifulSoup.MARKUP_MASSAGE)
    myNewMassage.extend(masssage_bad_comments)
    return BeautifulSoup.BeautifulSoup(data, markupMassage=myNewMassage,
        convertEntities=BeautifulSoup.BeautifulSoup.ALL_ENTITIES 
                    if to_unicode else None)

remove_html = lambda c: getsoup(c, to_unicode=True).getText(separator=u' ') if c else ""

2012-11-30 08:23:23

我使用Apache Tika取得了很好的结果。它的目的是从内容中提取元数据和文本，因此底层解析器要进行相应的开箱即用调优。

Tika可以作为服务器运行，在Docker容器中运行/部署很简单，并且可以通过Python绑定访问。

2018-05-07 11:07:18

你也可以在stripogram库中使用html2text方法。

from stripogram import html2text
text = html2text(your_html_string)

需要安装stripogram，请执行sudo easy_install stripogram命令

2009-09-23 03:21:58

下面是我经常使用的代码。

from bs4 import BeautifulSoup
import urllib.request


def processText(webpage):

    # EMPTY LIST TO STORE PROCESSED TEXT
    proc_text = []

    try:
        news_open = urllib.request.urlopen(webpage.group())
        news_soup = BeautifulSoup(news_open, "lxml")
        news_para = news_soup.find_all("p", text = True)

        for item in news_para:
            # SPLIT WORDS, JOIN WORDS TO REMOVE EXTRA SPACES
            para_text = (' ').join((item.text).split())

            # COMBINE LINES/PARAGRAPHS INTO A LIST
            proc_text.append(para_text)

    except urllib.error.HTTPError:
        pass

    return proc_text

我希望这对你有所帮助。

2017-10-25 00:08:25

使用Python从HTML文件中提取文本

推荐文章

最新文章

标签