如何解析XML并获得特定节点属性的实例?

我在XML中有很多行，我试图获得一个特定节点属性的实例。

<foo>
   <bar>
      <type foobar="1"/>
      <type foobar="2"/>
   </bar>
</foo>

我如何访问属性foobar的值?在这个例子中，我想要“1”和“2”。

当前回答

我很受伤，没有人建议熊猫。Pandas有一个read_xml()函数，它非常适合这种扁平的xml结构。

import pandas as pd

xml = """<foo>
   <bar>
      <type foobar="1"/>
      <type foobar="2"/>
   </bar>
</foo>"""

df = pd.read_xml(xml, xpath=".//type")
print(df)

输出:

   foobar
0       1
1       2

2023-02-03 05:42:19

其他回答

这里有一个使用cElementTree的非常简单但有效的代码。

try:
    import cElementTree as ET
except ImportError:
  try:
    # Python 2.5 need to import a different module
    import xml.etree.cElementTree as ET
  except ImportError:
    exit_err("Failed to import cElementTree from any known place")      

def find_in_tree(tree, node):
    found = tree.find(node)
    if found == None:
        print "No %s in file" % node
        found = []
    return found  

# Parse a xml file (specify the path)
def_file = "xml_file_name.xml"
try:
    dom = ET.parse(open(def_file, "r"))
    root = dom.getroot()
except:
    exit_err("Unable to open and parse input definition file: " + def_file)

# Parse to find the child nodes list of node 'myNode'
fwdefs = find_in_tree(root,"myNode")

这是来自“python xml解析”。

2013-07-08 20:35:26

Python有一个到expat XML解析器的接口。

xml.parsers.expat

它是一个非验证解析器，因此不会捕获糟糕的XML。但如果你知道你的文件是正确的，那么这就很好了，你可能会得到你想要的确切信息，你可以丢弃其余的。

stringofxml = """<foo>
    <bar>
        <type arg="value" />
        <type arg="value" />
        <type arg="value" />
    </bar>
    <bar>
        <type arg="value" />
    </bar>
</foo>"""
count = 0
def start(name, attr):
    global count
    if name == 'type':
        count += 1

p = expat.ParserCreate()
p.StartElementHandler = start
p.Parse(stringofxml)

print count # prints 4

2009-12-16 05:28:00

Minidom是最快速且非常直接的方法。

XML:

<data>
    <items>
        <item name="item1"></item>
        <item name="item2"></item>
        <item name="item3"></item>
        <item name="item4"></item>
    </items>
</data>

Python:

from xml.dom import minidom

dom = minidom.parse('items.xml')
elements = dom.getElementsByTagName('item')

print(f"There are {len(elements)} items:")

for element in elements:
    print(element.attributes['name'].value)

输出:

There are 4 items:
item1
item2
item3
item4

2009-12-16 05:30:15

我推荐ElementTree。同样的API还有其他兼容的实现，比如lxml和Python标准库中的cElementTree;但是，在这种情况下，他们主要增加的是更快的速度——编程的容易程度取决于ElementTree定义的API。

首先从XML中构建一个Element实例根，例如使用XML函数，或者通过解析文件，例如:

import xml.etree.ElementTree as ET
root = ET.parse('thefile.xml').getroot()

或者在ElementTree中显示的许多其他方法中的任何一种。然后这样做:

for type_tag in root.findall('bar/type'):
    value = type_tag.get('foobar')
    print(value)

输出:

1
2

2009-12-16 05:21:55

xml.etree.ElementTree vs. lxml

下面是两个最常用的库的一些优点，在进行选择之前，我应该了解它们。

xml.etree.ElementTree:

来自标准库:不需要安装任何模块

lxml

轻松编写XML声明:例如，您是否需要添加standalone="no"? 漂亮的打印:无需额外代码就可以得到漂亮的缩进XML。 Objectify功能:它允许您像处理普通的Python对象hierarchy.node一样使用XML。 sourceline允许您轻松地获取正在使用的XML元素的行。您还可以使用内置的XSD模式检查器。

2018-11-09 14:42:55

如何解析XML并获得特定节点属性的实例?

推荐文章

最新文章

标签