使用正则表达式解析HTML:为什么不呢?

似乎在stackoverflow上的每个问题中，提问者使用regex从HTML中获取一些信息，将不可避免地有一个“答案”，说不要使用regex解析HTML。

为什么不呢?我知道有一些所谓的“真正的”HTML解析器，比如Beautiful Soup，我相信它们是强大而有用的，但如果您只是在做一些简单、快速或简单的事情，那么当一些正则表达式语句就可以很好地工作时，为什么要麻烦使用如此复杂的东西呢?

此外，是否只是因为我不理解正则表达式的某些基本原理，才使得它们在解析中成为一个糟糕的选择?

当前回答

正则表达式对于HTML这样的语言来说还不够强大。当然，有一些例子可以使用正则表达式。但通常不适合进行解析。

2009-02-26 14:33:51

其他回答

问题是，大多数用户问的问题都与HTML和正则表达式有关，因为他们找不到自己的正则表达式。然后，必须考虑使用DOM或SAX解析器或类似的东西是否会更容易一些。它们是为处理类似xml的文档结构而优化和构造的。

当然，有些问题可以用正则表达式轻松解决。但重点在于容易。

如果您只想找到所有看起来像http://.../的url，那么使用regexp是没问题的。但是如果你想要找到a- element中所有具有'mylink'类的url，你可能最好使用合适的解析器。

2009-02-26 14:30:34

正则表达式并不是为处理嵌套的标记结构而设计的，要处理真正HTML中可能出现的所有边缘情况，往好里说是复杂的(往坏里说是不可能的)。

2009-02-26 14:35:50

请记住，虽然HTML本身不是规则的，但您正在查看的页面的某些部分可能是规则的。

例如，<form>标签被嵌套是一个错误;如果网页正常工作，那么使用正则表达式获取<form>将是完全合理的。

I recently did some web scraping using only Selenium and regular expressions. I got away with it because the data I wanted was put in a <form>, and put in a simple table format (so I could even count on <table>, <tr> and <td> to be non-nested--which is actually highly unusual). In some degree, regular expressions were even almost necessary, because some of the structure I needed to access was delimited by comments. (Beautiful Soup can give you comments, but it would have been difficult to grab  and  blocks using Beautiful Soup.)

但是，如果我不得不担心嵌套表，那么我的方法根本就行不通!我就只能靠《美丽汤》了。但是，即使这样，有时也可以使用正则表达式获取所需的块，然后从那里展开。

2013-02-12 18:34:47

我也试着用正则表达式来做这个。它主要用于查找与下一个HTML标记配对的内容块，它不查找匹配的结束标记，但它将拾取结束标记。用你自己的语言滚动一堆来检查这些。

与“sx”选项一起使用。如果你觉得幸运的话，也可以加上g:

(?P<content>.*?)                # Content up to next tag
(?P<markup>                     # Entire tag
  <!\[CDATA\[(?P<cdata>.+?)]]>| # <![CDATA[ ... ]]>
  <!--(?P<comment>.+?)-->|      # <!-- Comment -->
  </\s*(?P<close_tag>\w+)\s*>|  # </tag>
  <(?P<tag>\w+)                 # <tag ...
    (?P<attributes>
      (?P<attribute>\s+
# <snip>: Use this part to get the attributes out of 'attributes' group.
        (?P<attribute_name>\w+)
        (?:\s*=\s*
          (?P<attribute_value>
            [\w:/.\-]+|         # Unquoted
            (?=(?P<_v>          # Quoted
              (?P<_q>['\"]).*?(?<!\\)(?P=_q)))
            (?P=_v)
          ))?
# </snip>
      )*
    )\s*
  (?P<is_self_closing>/?)   # Self-closing indicator
  >)                        # End of tag

这个是为Python设计的(它可能适用于其他语言，还没有尝试过，它使用了正的反向查找头，负的反向查找头和命名的反向引用)。支持:

打开标签- <div…> 关闭标签- </div> 评论- <!——……--> Cdata - <![CDATA[…]] > 自关闭标签- <div…/> 可选属性值- <input checked> 未加引号/加引号的属性值- <div style='…'> 单引号/双引号- <div style="…" > 转义引号- <a title='John\'s Story'> (这不是真正有效的HTML，但我是一个好人) 等号周围的空格- <a href = '…'> 命名捕获感兴趣的位

它还可以很好地避免在格式错误的标记上触发，比如当您忘记了<或>时。

如果你的regex支持重复命名捕获，那么你是黄金，但Python re不支持(我知道regex支持，但我需要使用香草Python)。以下是你得到的结果:

content - All of the content up to the next tag. You could leave this out. markup - The entire tag with everything in it. comment - If it's a comment, the comment contents. cdata - If it's a <![CDATA[...]]>, the CDATA contents. close_tag - If it's a close tag (</div>), the tag name. tag - If it's an open tag (<div>), the tag name. attributes - All attributes inside the tag. Use this to get all attributes if you don't get repeated groups. attribute - Repeated, each attribute. attribute_name - Repeated, each attribute name. attribute_value - Repeated, each attribute value. This includes the quotes if it was quoted. is_self_closing - This is / if it's a self-closing tag, otherwise nothing. _q and _v - Ignore these; they're used internally for backreferences.

如果您的正则表达式引擎不支持重复的命名捕获，则可以使用一个被调用的部分来获取每个属性。只需在属性组上运行该正则表达式，从中获得每个属性、attribute_name和attribute_value。

演示在这里:https://regex101.com/r/mH8jSu/11

2016-12-28 05:05:01

正则表达式无法解析整个HTML，因为它依赖于匹配开始标记和结束标记，而正则表达式则无法匹配。

正则表达式只能匹配常规语言，但HTML是一种与上下文无关的语言，而不是常规语言(正如@StefanPochmann所指出的，常规语言也是与上下文无关的，因此与上下文无关并不一定意味着不常规)。在HTML上使用regexp唯一能做的事情是启发式，但这并不适用于所有条件。任何正则表达式都可以错误地匹配HTML文件。

2009-02-26 14:32:44

使用正则表达式解析HTML:为什么不呢?

推荐文章

最新文章

标签