非贪婪(不情愿)正则匹配sed?

我试图使用sed来清理url行来提取域。

所以从:

http://www.suepearson.co.uk/product/174/71/3816/

我想要:

http://www.suepearson.co.uk/

(不管后面有没有斜杠，都没有关系)

我试过:

 sed 's|\(http:\/\/.*?\/\).*|\1|'

And(转义非贪婪量词)

sed 's|\(http:\/\/.*\?\/\).*|\1|'

但我似乎不能让非贪婪量词(?)工作，所以它总是匹配整个字符串。

当前回答

还没有看到这个答案，所以这里是如何用vi或vim做到这一点:

vi -c '%s/\(http:\/\/.\{-}\/\).*/\1/ge | wq' file &>/dev/null

这将全局运行vi:%s替换(后面的g)，如果没有找到模式，则避免引发错误(e)，然后将结果更改保存到磁盘并退出。&>/dev/null可以防止GUI在屏幕上短暂闪烁，这很烦人。

有时候我喜欢用vi来处理超级复杂的正则表达式，因为(1)perl已经奄奄一息了，(2)vim有一个非常先进的正则表达式引擎，(3)在我日常使用的编辑文档中，我已经非常熟悉vi正则表达式了。

2019-04-03 20:38:33

其他回答

sed的| \ (http: \ \ / www \ [a-z.0-9] * \ / \)。|\1|也可以

2013-06-24 15:33:43

使用纯(GNU) sed仍然有希望解决这个问题。尽管这不是一个通用的解决方案，在某些情况下，你可以使用“循环”来消除字符串中所有不必要的部分，就像这样:

sed -r -e ":loop" -e 's|(http://.+)/.*|\1|' -e "t loop"

-r:使用扩展的正则表达式(用于+和未转义的括号) 定义一个名为"loop"的新标签 -e:在sed中添加命令 "t loop":如果有成功的替换，则跳回标记"loop"

这里唯一的问题是它也会切掉最后一个分隔符('/')，但如果你真的需要它，你仍然可以在“循环”结束后简单地把它放回去，只需要在前面的命令行末尾追加这个额外的命令:

-e "s,$,/,"

2016-08-01 12:52:19

下面的解决方案适用于匹配/使用multiply present(链式;串联;复合)HTML或其他标签。例如，我想编辑HTML代码以删除串联出现的<span>标记。

问题:常规sed正则表达式贪婪地匹配从第一个到最后一个的所有标记。

解决方案:非贪婪模式匹配(每个讨论在这个线程的其他地方;例如,https://stackoverflow.com/a/46719361/1904943)。

例子:

echo '<span>Will</span>This <span>remove</span>will <span>this.</span>remain.' | \
sed 's/<span>[^>]*>//g' ; echo

This will remain.

解释:

S /<span>:查找<span> [^>]:后面跟着不是>的任何东西 *>:直到你找到> //g:将任何这样的字符串替换为空。

齿顶高

我试图清理url，但我遇到了困难匹配/排除一个词- href -使用上面的方法。我简要地研究了反向查找(正则表达式来匹配不包含单词的行)，但这种方法似乎过于复杂，并没有提供令人满意的解决方案。

我决定用'(反标记)替换href，做正则表达式替换，然后用href替换'。

示例(为便于阅读，此处格式化):

printf '\n
<a aaa h href="apple">apple</a>
<a bbb "c=ccc" href="banana">banana</a>
<a class="gtm-content-click"
   data-vars-link-text="nope"
   data-vars-click-url="https://blablabla"
   data-vars-event-category="story"
   data-vars-sub-category="story"
   data-vars-item="in_content_link"
   data-vars-link-text
   href="https:example.com">Example.com</a>\n\n' |
sed 's/href/`/g ;
     s/<a[^`]*`/\n<a href/g'

<a href="apple">apple</a> 
<a href="banana">banana</a> 
<a href="https:example.com">Example.com</a>

解释:基本如上。在这里,

S /href/ ':将href替换为'(反引号) s/<a:查找URL开头 [^ ']:后面跟着任何不是'(反撇号) * ':直到你找到一个' /<a href/g:用<a href替换每一个找到的

2021-11-15 01:03:54

您还应该考虑没有匹配界限的情况。你是否想输出这一行。如果不匹配，我这里的示例不会输出任何内容。

你需要前缀到第三个/，所以选择两次字符串的任何长度不包含/和后面的/，然后字符串的任何长度不包含/，然后匹配/后面的任何字符串，然后打印选择。这个想法适用于任何单个的char delims。

echo http://www.suepearson.co.uk/product/174/71/3816/ | \
  sed -nr 's,(([^/]*/){2}[^/]*)/.*,\1,p'

使用sed命令，您可以快速删除前缀或delim选择，如:

echo 'aaa @cee: { "foo":" @cee: " }' | \
  sed -r 't x;s/ @cee: /\n/;D;:x'

这比一次吃焦肉快多了。

如果之前匹配成功，跳转到标签。在第一道线/之前加\n。移除到第一个\n。如果添加了\n，则跳转到结束并打印。

如果有开始和结束delim，很容易删除结束delim，直到你到达你想要的第n -2个元素，然后做D技巧，在结束delim后删除，如果不匹配跳转到删除，在开始delim和打印之前删除。这仅在开始/结束分隔成对出现时有效。

echo 'foobar start block #1 end barfoo start block #2 end bazfoo start block #3 end goo start block #4 end faa' | \
  sed -r 't x;s/end//;s/end/\n/;D;:x;s/(end).*/\1/;T y;s/.*(start)/\1/;p;:y;d'

2021-06-11 14:11:55

在sed中模拟惰性(非贪婪)量词

以及所有其他正则表达式口味!

Finding first occurrence of an expression: POSIX ERE (using -r option) Regex: (EXPRESSION).*|. Sed: sed -r ‍'s/(EXPRESSION).*|./\1/g' # Global `g` modifier should be on Example (finding first sequence of digits) Live demo: $ sed -r 's/([0-9]+).*|./\1/g' <<< 'foo 12 bar 34' 12 How does it work? This regex benefits from an alternation |. At each position engine tries to pick the longest match (this is a POSIX standard which is followed by couple of other engines as well) which means it goes with . until a match is found for ([0-9]+).*. But order is important too. Since global flag is set, engine tries to continue matching character by character up to the end of input string or our target. As soon as the first and only capturing group of left side of alternation is matched (EXPRESSION) rest of line is consumed immediately as well .*. We now hold our value in the first capturing group. POSIX BRE Regex: $\(\(EXPRESSION$.*\)*.\)* Sed: sed 's/$\(\(EXPRESSION$.*\)*.\)*/\3/' Example (finding first sequence of digits): $ sed 's/$\(\([0-9]\{1,\}$.*\)*.\)*/\3/' <<< 'foo 12 bar 34' 12 This one is like ERE version but with no alternation involved. That's all. At each single position engine tries to match a digit. If it is found, other following digits are consumed and captured and the rest of line is matched immediately otherwise since * means more or zero it skips over second capturing group $\([0-9]\{1,\}$.*\)* and arrives at a dot . to match a single character and this process continues. Finding first occurrence of a delimited expression: This approach will match the very first occurrence of a string that is delimited. We can call it a block of string. sed 's/$END-DELIMITER-EXPRESSION$.*/\1/; \ s/$\(START-DELIMITER-EXPRESSION.*$*.\)*/\1/g' Input string: foobar start block #1 end barfoo start block #2 end -EDE: end -SDE: start $ sed 's/$end$.*/\1/; s/$\(start.*$*.\)*/\1/g' Output: start block #1 end First regex $end$.* matches and captures first end delimiter end and substitues all match with recent captured characters which is the end delimiter. At this stage our output is: foobar start block #1 end. Then the result is passed to second regex $\(start.*$*.\)* that is same as POSIX BRE version above. It matches a single character if start delimiter start is not matched otherwise it matches and captures the start delimiter and matches the rest of characters.

直接回答你的问题

使用方法#2(带分隔符的表达式)，你应该选择两个合适的表达式:

艾德:[^]\ / SDE: http:

用法:

$ sed 's/\([^:/]\/\).*/\1/g; s/\(\(http:.*\)*.\)*/\1/' <<< 'http://www.suepearson.co.uk/product/174/71/3816/'

输出:

http://www.suepearson.co.uk/

注意:对于相同的分隔符，这将不起作用。

2016-09-28 16:26:21

非贪婪(不情愿)正则匹配sed?

推荐文章

最新文章

标签