标记数据错误

我试图使用熊猫操作.csv文件，但我得到这个错误:

pandas.parser.CParserError:标记数据错误。C错误:第3行有2个字段，见12

我试着读过熊猫的文件，但一无所获。

我的代码很简单:

path = 'GOOG Key Ratios.csv'
#print(open(path).read())
data = pd.read_csv(path)

我该如何解决这个问题?我应该使用csv模块还是其他语言?

文件来自晨星公司

当前回答

我相信解决方案，

,engine='python'
, error_bad_lines = False

如果它是虚拟列并且你想要删除它，这将是很好的。在我的例子中，第二行确实有更多的列，我希望这些列被积分，并且有列数= MAX(列)。

请参考下面我无法阅读的解决方案:

try:
    df_data = pd.read_csv(PATH, header = bl_header, sep = str_sep)
except pd.errors.ParserError as err:
    str_find = 'saw '
    int_position = int(str(err).find(str_find)) + len(str_find)
    str_nbCol = str(err)[int_position:]
    l_col = range(int(str_nbCol))
    df_data = pd.read_csv(PATH, header = bl_header, sep = str_sep, names = l_col)

2020-05-13 06:41:07

其他回答

我也有这个问题，但可能是出于不同的原因。我在我的CSV中有一些尾随逗号，添加了熊猫试图读取的额外列。使用以下方法是可行的，但它只是忽略了不好的行:

data = pd.read_csv('file1.csv', error_bad_lines=False)

如果你想让代码行看起来很丑，你可以这样做:

line     = []
expected = []
saw      = []     
cont     = True 

while cont == True:     
    try:
        data = pd.read_csv('file1.csv',skiprows=line)
        cont = False
    except Exception as e:    
        errortype = e.message.split('.')[0].strip()                                
        if errortype == 'Error tokenizing data':                        
           cerror      = e.message.split(':')[1].strip().replace(',','')
           nums        = [n for n in cerror.split(' ') if str.isdigit(n)]
           expected.append(int(nums[0]))
           saw.append(int(nums[2]))
           line.append(int(nums[1])-1)
         else:
           cerror      = 'Unknown'
           print 'Unknown Error - 222'

if line != []:
    # Handle the errors however you want

我接着写了一个脚本，将这些行重新插入到DataFrame中，因为坏的行将由上述代码中的变量“line”给出。这一切都可以通过简单地使用csv阅读器来避免。希望熊猫的开发人员能够在未来更容易地处理这种情况。

2016-02-04 22:16:44

下面的命令序列工作(我丢失了数据的第一行-no header=None present-，但至少它加载):

Df = pd.read_csv(文件名， usecols =范围(0,42)) df。列=[‘年’,‘莫’,‘天’,“人力资源”,“分”,“秒”,“猎狗”, ' error '， ' rectype '， ' lane '， ' speed '， ' class '， ' length ' ' gvw ' ' esal ' ' w1 ' ' s1 ' ' w2 ' ' s2 ' ' w3 ' ' s3 ' ' w4 ' ' s4 ' ' w5 ' ' s5 ' ' w6 ' ' s6 ' ' w7 ' ' s7 ' ' w8 ' ' s8 ' ' w9 ' ' s9 ' ' w10 ' ' s10 ' ' w11 '， ' s11 '， ' w12 '， ' s12 '， ' w13 '， ' s13 '， ' w14 ']

以下不工作:

Df = pd.read_csv(文件名，名称=[‘年’,‘莫’,‘天’,“人力资源”,“分”,“秒”,“猎狗”, ' error '， ' rectype '， ' lane '， ' speed '， ' class '， ' length ' ' gvw ' ' esal ' ' w1 ' ' s1 ' ' w2 ' ' s2 ' ' w3 ' ' s3 ' ' w4 ' ' s4 ' ' w5 ' ' s5 ' ' w6 ' ' s6 ' ' w7 ' ' s7 ' ' w8 ' ' s8 ' ' w9 ' ' s9 ' ' w10 ' ' s10 ' ' w11 '， ' s11 '， ' w12 '， ' s12 '， ' w13 '， ' s13 '， ' w14 ']， usecols =范围(0,42))

CParserError:标记数据错误。C错误:在1605634行中预期有53个字段，看到54 以下不工作:

df = pd read_csv(文件) 标题=郎)

CParserError:标记数据错误。C错误:在1605634行中预期有53个字段，看到54

因此，在你的问题中，你必须传递usecols=range(0,2)

2018-05-23 11:45:25

你可以使用:

pd.read_csv("mycsv.csv", delimiter=";")

熊猫1.4.4

它可以是文件的分隔符，将其作为文本文件打开，查找分隔符。然后，您将拥有可以为空且未命名的列，因为行包含太多分隔符。

因此，您可以使用pandas来处理它们并检查值。对我来说，这比在我的情况下跳过台词要好。

2022-09-23 10:29:20

标记数据错误。C错误:第3行有2个字段，见12

这个错误给出了解决问题“Expected 2 fields in line 3, saw 12”的线索，saw 12表示第二行长度为12，第一行长度为2。

当您有如下所示的数据时，如果您跳过行，那么大部分数据将被跳过

data = """1,2,3
1,2,3,4
1,2,3,4,5
1,2
1,2,3,4"""

如果您不想跳过任何行，请执行以下操作

#First lets find the maximum column for all the rows
with open("file_name.csv", 'r') as temp_f:
    # get No of columns in each line
    col_count = [ len(l.split(",")) for l in temp_f.readlines() ]

### Generate column names  (names will be 0, 1, 2, ..., maximum columns - 1)
column_names = [i for i in range(max(col_count))] 

import pandas as pd
# inside range set the maximum value you can see in "Expected 4 fields in line 2, saw 8"
# here will be 8 
data = pd.read_csv("file_name.csv",header = None,names=column_names )

使用range而不是手动设置名称，因为当您有很多列时，这样做会很麻烦。

此外，如果需要使用均匀的数据长度，可以将NaN值填充为0。如。对于聚类(k-means)

new_data = data.fillna(0)

2020-02-16 09:58:45

在处理类似的解析错误时，我发现另一种方法很有用，它使用CSV模块将数据重新路由到pandas df。例如:

import csv
import pandas as pd
path = 'C:/FileLocation/'
file = 'filename.csv'
f = open(path+file,'rt')
reader = csv.reader(f)

#once contents are available, I then put them in a list
csv_list = []
for l in reader:
    csv_list.append(l)
f.close()
#now pandas has no problem getting into a df
df = pd.DataFrame(csv_list)

我发现CSV模块对于格式不佳的逗号分隔的文件更加健壮，因此已经成功地用这种方法解决了诸如此类的问题。

2018-01-26 20:54:38

标记数据错误

推荐文章

最新文章

标签