使用Python从HTML文件中提取文本

我想使用Python从HTML文件中提取文本。我想从本质上得到相同的输出，如果我从浏览器复制文本，并将其粘贴到记事本。

我想要一些更健壮的东西，而不是使用正则表达式，正则表达式可能会在格式不佳的HTML上失败。我见过很多人推荐Beautiful Soup，但我在使用它时遇到了一些问题。首先，它会抓取不需要的文本，比如JavaScript源代码。此外，它也不解释HTML实体。例如，我会期望'在HTML源代码中转换为文本中的撇号，就像我将浏览器内容粘贴到记事本一样。

更新html2text看起来很有希望。它正确地处理HTML实体，而忽略JavaScript。然而，它并不完全生成纯文本;它产生的降价，然后必须转换成纯文本。它没有示例或文档，但代码看起来很干净。

相关问题:

在python中过滤HTML标签并解析实体在Python中将XML/HTML实体转换为Unicode字符串

当前回答

有人尝试过bleach.clean(html,tags=[]，strip=True)与漂白剂吗?这对我很有用。

2017-01-16 14:10:24

其他回答

html2text是一个Python程序，它在这方面做得很好。

2008-11-30 03:23:58

另一种选择是通过基于文本的web浏览器运行html并转储它。例如(使用Lynx):

lynx -dump html_to_convert.html > converted_html.txt

这可以在python脚本中完成，如下所示:

import subprocess

with open('converted_html.txt', 'w') as outputFile:
    subprocess.call(['lynx', '-dump', 'html_to_convert.html'], stdout=testFile)

它不会精确地为您提供HTML文件中的文本，但根据您的用例，它可能比html2text的输出更好。

2014-08-08 02:29:50

PyParsing做得很好。PyParsing wiki被杀死了，所以这里有另一个位置，这里有使用PyParsing的示例(示例链接)。花点时间在pyparsing上的一个原因是，他还写了一本非常简短、组织良好的O'Reilly捷径手册，而且价格便宜。

话虽如此，我经常使用BeautifulSoup，处理实体问题并不难，你可以在运行BeautifulSoup之前转换它们。

古德勒克

2008-11-30 15:46:19

另一个在Python 2.7.9+中使用BeautifulSoup4的例子

包括:

import urllib2
from bs4 import BeautifulSoup

代码:

def read_website_to_text(url):
    page = urllib2.urlopen(url)
    soup = BeautifulSoup(page, 'html.parser')
    for script in soup(["script", "style"]):
        script.extract() 
    text = soup.get_text()
    lines = (line.strip() for line in text.splitlines())
    chunks = (phrase.strip() for line in lines for phrase in line.split("  "))
    text = '\n'.join(chunk for chunk in chunks if chunk)
    return str(text.encode('utf-8'))

解释道:

将url数据读入为html(使用BeautifulSoup)，删除所有脚本和样式元素，并使用.get_text()仅获取文本。分割成行，删除每个标题的开头和结尾空格，然后将多个标题分割成一行，each chunks = (phrase.strip() for line in line for phrase in line。(" "))。然后使用text = '\n'。加入，删除空行，最后返回为批准的utf-8。

注:

一些系统这是运行在https://连接失败，因为SSL问题，你可以关闭验证来解决这个问题。修复示例:http://blog.pengyifan.com/how-to-fix-python-ssl-certificate_verify_failed/ Python < 2.7.9在运行时可能会遇到一些问题 text.encode('utf-8')可能会留下奇怪的编码，可能只需要返回str(text)即可。

2019-08-27 16:52:57

Beautiful soup可以转换html实体。考虑到HTML经常有bug并且充满unicode和HTML编码问题，这可能是您最好的选择。这是我用来将html转换为原始文本的代码:

import BeautifulSoup
def getsoup(data, to_unicode=False):
    data = data.replace("&nbsp;", " ")
    # Fixes for bad markup I've seen in the wild.  Remove if not applicable.
    masssage_bad_comments = [
        (re.compile('<!-([^-])'), lambda match: '<!--' + match.group(1)),
        (re.compile('<!WWWAnswer T[=\w\d\s]*>'), lambda match: '<!--' + match.group(0) + '-->'),
    ]
    myNewMassage = copy.copy(BeautifulSoup.BeautifulSoup.MARKUP_MASSAGE)
    myNewMassage.extend(masssage_bad_comments)
    return BeautifulSoup.BeautifulSoup(data, markupMassage=myNewMassage,
        convertEntities=BeautifulSoup.BeautifulSoup.ALL_ENTITIES 
                    if to_unicode else None)

remove_html = lambda c: getsoup(c, to_unicode=True).getText(separator=u' ') if c else ""

2012-11-30 08:23:23

使用Python从HTML文件中提取文本

推荐文章

最新文章

标签