使用多个单词边界分隔符将字符串拆分为单词

我想我想做的是一项相当常见的任务，但我在网上找不到任何参考资料。我有带标点符号的文本，我想要一个单词列表。

"Hey, you - what are you doing here!?"

应该是

['hey', 'you', 'what', 'are', 'you', 'doing', 'here']

但Python的str.split（）只对一个参数有效，所以在用空格拆分后，所有单词都带有标点符号。有什么想法吗？

当前回答

创建一个函数，将两个字符串（要拆分的源字符串和分隔符的拆分列表字符串）作为输入，并输出拆分单词列表：

def split_string(source, splitlist):
    output = []  # output list of cleaned words
    atsplit = True
    for char in source:
        if char in splitlist:
            atsplit = True
        else:
            if atsplit:
                output.append(char)  # append new word after split
                atsplit = False
            else: 
                output[-1] = output[-1] + char  # continue copying characters until next split
    return output

2017-05-10 00:58:57

其他回答

我喜欢re，但这是我没有它的解决方案：

from itertools import groupby
sep = ' ,-!?'
s = "Hey, you - what are you doing here!?"
print [''.join(g) for k, g in groupby(s, sep.__contains__) if not k]

sep__contains__是“in”运算符使用的方法。基本上与

lambda ch: ch in sep

但在这里更方便。

groupby获取字符串和函数。它使用该函数将字符串分成多个组：每当函数值发生变化时，就会生成一个新的组。因此，sep__contains__正是我们需要的。

groupby返回一个对序列，其中对[0]是我们函数的结果，对[1]是一个组。使用“if not k”，我们筛选出带有分隔符的组（因为sep.__contains__的结果在分隔符上为True）。好了，就这些了-现在我们有一个组序列，每个组都是一个单词（组实际上是一个可迭代的，所以我们使用join将其转换为字符串）。

这个解决方案非常通用，因为它使用一个函数来分隔字符串（您可以根据需要的任何条件进行拆分）。此外，它不创建中间字符串/列表（您可以删除join，因为每个组都是一个迭代器，所以表达式将变得懒惰）

2013-10-06 17:30:05

首先，我不认为您的意图是在拆分函数中实际使用标点符号作为分隔符。您的描述表明您只是想从生成的字符串中删除标点符号。

我经常遇到这种情况，我通常的解决方案不需要re。

单行lambda函数，带列表理解：

（需要导入字符串）：

split_without_punc = lambda text : [word.strip(string.punctuation) for word in 
    text.split() if word.strip(string.punctuation) != '']

# Call function
split_without_punc("Hey, you -- what are you doing?!")
# returns ['Hey', 'you', 'what', 'are', 'you', 'doing']

功能（传统）

作为传统函数，这仍然只有两行具有列表理解（除了导入字符串）：

def split_without_punctuation2(text):

    # Split by whitespace
    words = text.split()

    # Strip punctuation from each word
    return [word.strip(ignore) for word in words if word.strip(ignore) != '']

split_without_punctuation2("Hey, you -- what are you doing?!")
# returns ['Hey', 'you', 'what', 'are', 'you', 'doing']

它也会自然地保留缩略词和连字符。您可以始终使用text.replace（“-”，“”）在拆分前将连字符转换为空格。

不带Lambda或列表理解的通用函数

对于更一般的解决方案（可以指定要删除的字符），并且不需要列表理解，您可以得到：

def split_without(text: str, ignore: str) -> list:

    # Split by whitespace
    split_string = text.split()

    # Strip any characters in the ignore string, and ignore empty strings
    words = []
    for word in split_string:
        word = word.strip(ignore)
        if word != '':
            words.append(word)

    return words

# Situation-specific call to general function
import string
final_text = split_without("Hey, you - what are you doing?!", string.punctuation)
# returns ['Hey', 'you', 'what', 'are', 'you', 'doing']

当然，您也可以将lambda函数推广到任何指定的字符串。

2014-11-04 19:17:37

专业提示：使用string.translate进行Python最快的字符串操作。

一些证据。。。

首先，缓慢的方式（抱歉pprzemek）：

>>> import timeit
>>> S = 'Hey, you - what are you doing here!?'
>>> def my_split(s, seps):
...     res = [s]
...     for sep in seps:
...         s, res = res, []
...         for seq in s:
...             res += seq.split(sep)
...     return res
... 
>>> timeit.Timer('my_split(S, punctuation)', 'from __main__ import S,my_split; from string import punctuation').timeit()
54.65477919578552

接下来，我们使用re.findall（）（如建议的答案所示）。更快：

>>> timeit.Timer('findall(r"\w+", S)', 'from __main__ import S; from re import findall').timeit()
4.194725036621094

最后，我们使用translate：

>>> from string import translate,maketrans,punctuation 
>>> T = maketrans(punctuation, ' '*len(punctuation))
>>> timeit.Timer('translate(S, T).split()', 'from __main__ import S,T,translate').timeit()
1.2835021018981934

说明：

string.translate是用C实现的，与Python中的许多字符串操作函数不同，string.ttranslate不会生成新字符串。所以它的速度和字符串替换一样快。

不过，这有点尴尬，因为它需要一个翻译表来实现这一魔术。您可以使用maketrans（）方便函数创建转换表。这里的目标是将所有不需要的字符转换为空格。一换一的替代品。同样，不会产生新数据。所以这很快！

接下来，我们使用旧的split（）。默认情况下，split（）将对所有空白字符进行操作，将它们分组以进行拆分。结果将是您想要的单词列表。而且这种方法几乎比re.findall（）快4倍！

2012-08-30 04:05:54

使用maketrans和translate，您可以轻松、整洁地完成

import string
specials = ',.!?:;"()<>[]#$=-/'
trans = string.maketrans(specials, ' '*len(specials))
body = body.translate(trans)
words = body.strip().split()

2018-03-03 23:59:23

我必须想出自己的解决方案，因为我迄今为止测试的所有东西都在某一点上失败了。

>>> import re
>>> def split_words(text):
...     rgx = re.compile(r"((?:(?<!'|\w)(?:\w-?'?)+(?<!-))|(?:(?<='|\w)(?:\w-?'?)+(?=')))")
...     return rgx.findall(text)

至少在下面的例子中，它似乎工作得很好。

>>> split_words("The hill-tops gleam in morning's spring.")
['The', 'hill-tops', 'gleam', 'in', "morning's", 'spring']
>>> split_words("I'd say it's James' 'time'.")
["I'd", 'say', "it's", "James'", 'time']
>>> split_words("tic-tac-toe's tic-tac-toe'll tic-tac'tic-tac we'll--if tic-tac")
["tic-tac-toe's", "tic-tac-toe'll", "tic-tac'tic-tac", "we'll", 'if', 'tic-tac']
>>> split_words("google.com email@google.com split_words")
['google', 'com', 'email', 'google', 'com', 'split_words']
>>> split_words("Kurt Friedrich Gödel (/ˈɡɜːrdəl/;[2] German: [ˈkʊɐ̯t ˈɡøːdl̩] (listen);")
['Kurt', 'Friedrich', 'Gödel', 'ˈɡɜːrdəl', '2', 'German', 'ˈkʊɐ', 't', 'ˈɡøːdl', 'listen']
>>> split_words("April 28, 1906 – January 14, 1978) was an Austro-Hungarian-born Austrian...")
['April', '28', '1906', 'January', '14', '1978', 'was', 'an', 'Austro-Hungarian-born', 'Austrian']

2019-05-23 23:06:56

使用多个单词边界分隔符将字符串拆分为单词

推荐文章

最新文章

标签