从字符串中删除所有特殊字符、标点符号和空格

我需要从字符串中删除所有特殊字符，标点符号和空格，以便我只有字母和数字。

当前回答

import re
my_string = """Strings are amongst the most popular data types in Python. We can create the strings by enclosing characters in quotes. Python treats single quotes the

和双引号一样。”＂＂

# if we need to count the word python that ends with or without ',' or '.' at end

count = 0
for i in text:
    if i.endswith("."):
        text[count] = re.sub("^([a-z]+)(.)?$", r"\1", i)
    count += 1
print("The count of Python : ", text.count("python"))

2018-07-16 11:52:40

其他回答

Python 2 . *

我认为只要filter(str。Isalnum，字符串)工作

In [20]: filter(str.isalnum, 'string with special chars like !,#$% etcs.')
Out[20]: 'stringwithspecialcharslikeetcs'

Python 3。*

在Python3中，filter()函数将返回一个可迭代对象(而不是与上面不同的字符串)。从itertable中获取字符串必须返回连接:

''.join(filter(str.isalnum, string))

或者在连接中传递列表(不确定，但可以快一点)

''.join([*filter(str.isalnum, string)])

注意:unpacking in [*args] valid from Python >= 3.5

2016-04-14 09:32:50

这可以不使用regex完成:

>>> string = "Special $#! characters   spaces 888323"
>>> ''.join(e for e in string if e.isalnum())
'Specialcharactersspaces888323'

你可以使用str.isalnum:

S.isalnum() -> bool 如果S中的所有字符都是字母数字，则返回True 且S中至少有一个字符，否则为假。

如果坚持使用正则表达式，其他解决方案也可以。但是请注意，如果可以在不使用正则表达式的情况下完成，那么这是最好的方法。

2011-04-30 17:47:39

与使用正则表达式的其他人不同，我将尝试排除不是我想要的每个字符，而不是显式地列举我不想要的字符。

例如，如果我只想要字符从'a到z'(大写和小写)和数字，我将排除所有其他:

import re
s = re.sub(r"[^a-zA-Z0-9]","",s)

这意味着“用空字符串替换每个不是数字的字符，或者'a到z'或'a到z'范围内的字符”。

事实上，如果你在正则表达式的第一个位置插入特殊字符^，你将得到否定。

额外提示:如果您还需要将结果小写，您可以使正则表达式更快更简单，只要您现在不会发现任何大写。

import re
s = re.sub(r"[^a-z0-9]","",s.lower())

2018-09-05 10:02:41

function regexFuntion(st) {
  const regx = /[^\w\s]/gi; // allow : [a-zA-Z0-9, space]
  st = st.replace(regx, ''); // remove all data without [a-zA-Z0-9, space]
  st = st.replace(/\s\s+/g, ' '); // remove multiple space

  return st;
}

console.log(regexFuntion('$Hello; # -world--78asdf+-===asdflkj******lkjasdfj67;'));
// Output: Hello world78asdfasdflkjlkjasdfj67

2022-04-06 15:02:44

对于其他语言，如德语，西班牙语，丹麦语，法语等包含特殊字符(如德语“Umlaute”ü， ä， ö)，只需将这些添加到正则表达式搜索字符串:

例如德语:

re.sub('[^A-ZÜÖÄa-z0-9]+', '', mystring)

2020-06-27 10:00:21

从字符串中删除所有特殊字符、标点符号和空格

推荐文章

最新文章

标签