如何迭代字符串的单词？

如何迭代由空格分隔的单词组成的字符串中的单词？

注意，我对C字符串函数或那种字符操作/访问不感兴趣。比起效率，我更喜欢优雅。我当前的解决方案：

#include <iostream>
#include <sstream>
#include <string>

using namespace std;

int main() {
    string s = "Somewhere down the road";
    istringstream iss(s);

    do {
        string subs;
        iss >> subs;
        cout << "Substring: " << subs << endl;
    } while (iss);
}

当前回答

这里有一个拆分函数：

是通用的使用标准C++（无增强）接受多个分隔符忽略空标记（可以轻松更改）模板＜typename T＞矢量<T>拆分（常量T&str，常量T&分隔符）{向量<T>v；typename T:：size_type start=0；自动位置=str.find_first_of（分隔符，开始）；而（pos！=T:：npos）{if（pos！=开始）//忽略空标记v.template_back（str，start，pos-start）；开始=位置+1；pos=str.find_first_of（分隔符，开始）；}if（start＜str.length（））//忽略尾随分隔符v.template_back（str，start，str.length（）-start）；//添加字符串的剩余部分返回v；}

示例用法：

    vector<string> v = split<string>("Hello, there; World", ";,");
    vector<wstring> v = split<wstring>(L"Hello, there; World", L";,");

2012-03-13 00:09:42

其他回答

我有一种与其他解决方案非常不同的方法，它提供了很多其他解决方案所缺乏的价值，但当然也有其缺点。这是一个工作实现，示例是在单词周围放置＜tag＞＜/tag＞。

首先，这个问题可以通过一个循环解决，不需要额外的内存，只需考虑四种逻辑情况。从概念上讲，我们对边界感兴趣。我们的代码应该反映出这一点：让我们遍历字符串，一次查看两个字符，记住字符串的开头和结尾都有特殊情况。

缺点是我们必须编写实现，这有点冗长，但大多是方便的样板。

好处是我们编写了实现，因此很容易根据特定的需要定制它，例如区分左和写单词边界，使用任何一组分隔符，或处理其他情况，例如无边界或错误位置。

using namespace std;

#include <iostream>
#include <string>

#include <cctype>

typedef enum boundary_type_e {
    E_BOUNDARY_TYPE_ERROR = -1,
    E_BOUNDARY_TYPE_NONE,
    E_BOUNDARY_TYPE_LEFT,
    E_BOUNDARY_TYPE_RIGHT,
} boundary_type_t;

typedef struct boundary_s {
    boundary_type_t type;
    int pos;
} boundary_t;

bool is_delim_char(int c) {
    return isspace(c); // also compare against any other chars you want to use as delimiters
}

bool is_word_char(int c) {
    return ' ' <= c && c <= '~' && !is_delim_char(c);
}

boundary_t maybe_word_boundary(string str, int pos) {
    int len = str.length();
    if (pos < 0 || pos >= len) {
        return (boundary_t){.type = E_BOUNDARY_TYPE_ERROR};
    } else {
        if (pos == 0 && is_word_char(str[pos])) {
            // if the first character is word-y, we have a left boundary at the beginning
            return (boundary_t){.type = E_BOUNDARY_TYPE_LEFT, .pos = pos};
        } else if (pos == len - 1 && is_word_char(str[pos])) {
            // if the last character is word-y, we have a right boundary left of the null terminator
            return (boundary_t){.type = E_BOUNDARY_TYPE_RIGHT, .pos = pos + 1};
        } else if (!is_word_char(str[pos]) && is_word_char(str[pos + 1])) {
            // if we have a delimiter followed by a word char, we have a left boundary left of the word char
            return (boundary_t){.type = E_BOUNDARY_TYPE_LEFT, .pos = pos + 1};
        } else if (is_word_char(str[pos]) && !is_word_char(str[pos + 1])) {
            // if we have a word char followed by a delimiter, we have a right boundary right of the word char
            return (boundary_t){.type = E_BOUNDARY_TYPE_RIGHT, .pos = pos + 1};
        }
        return (boundary_t){.type = E_BOUNDARY_TYPE_NONE};
    }
}

int main() {
    string str;
    getline(cin, str);

    int len = str.length();
    for (int i = 0; i < len; i++) {
        boundary_t boundary = maybe_word_boundary(str, i);
        if (boundary.type == E_BOUNDARY_TYPE_LEFT) {
            // whatever
        } else if (boundary.type == E_BOUNDARY_TYPE_RIGHT) {
            // whatever
        }
    }
}

正如您所看到的，代码非常容易理解和微调，代码的实际使用非常简短和简单。使用C++不应阻止我们编写最简单、最容易定制的代码，即使这意味着不使用STL。我认为这是Linus Torvalds所说的“品味”的一个例子，因为我们已经消除了所有不需要的逻辑，而写作风格自然允许在需要处理的时候处理更多的案件。

可以改进此代码的可能是使用enum类，在maybe_word_boundary中接受指向is_word_char的函数指针，而不是直接调用is_word_char，并传递lambda。

2019-01-16 15:14:15

我使用这个simpleton是因为我们得到了字符串类“特殊”（即非标准）：

void splitString(const String &s, const String &delim, std::vector<String> &result) {
    const int l = delim.length();
    int f = 0;
    int i = s.indexOf(delim,f);
    while (i>=0) {
        String token( i-f > 0 ? s.substring(f,i-f) : "");
        result.push_back(token);
        f=i+l;
        i = s.indexOf(delim,f);
    }
    String token = s.substring(f);
    result.push_back(token);
}

2010-09-01 09:25:52

这是我对这个的看法。我必须一个字一个字地处理输入字符串，这可以通过使用空格来计数单词来完成，但我觉得这会很乏味，我应该将单词分割成向量。

#include<iostream>
#include<vector>
#include<string>
#include<stdio.h>
using namespace std;
int main()
{
    char x = '\0';
    string s = "";
    vector<string> q;
    x = getchar();
    while(x != '\n')
    {
        if(x == ' ')
        {
            q.push_back(s);
            s = "";
            x = getchar();
            continue;
        }
        s = s + x;
        x = getchar();
    }
    q.push_back(s);
    for(int i = 0; i<q.size(); i++)
        cout<<q[i]<<" ";
    return 0;
}

不处理多个空间。如果最后一个单词后面没有紧跟换行符，则它包含最后一个词的最后一个字符和换行符之间的空格。

2016-09-10 17:07:56

使用vector作为基类的快速版本，可完全访问其所有运算符：

    // Split string into parts.
    class Split : public std::vector<std::string>
    {
        public:
            Split(const std::string& str, char* delimList)
            {
               size_t lastPos = 0;
               size_t pos = str.find_first_of(delimList);

               while (pos != std::string::npos)
               {
                    if (pos != lastPos)
                        push_back(str.substr(lastPos, pos-lastPos));
                    lastPos = pos + 1;
                    pos = str.find_first_of(delimList, lastPos);
               }
               if (lastPos < str.length())
                   push_back(str.substr(lastPos, pos-lastPos));
            }
    };

用于填充STL集的示例：

std::set<std::string> words;
Split split("Hello,World", ",");
words.insert(split.begin(), split.end());

2012-02-21 21:31:35

这类似于堆栈溢出问题：如何在C++中标记字符串？。需要Boost外部库

#include <iostream>
#include <string>
#include <boost/tokenizer.hpp>

using namespace std;
using namespace boost;

int main(int argc, char** argv)
{
    string text = "token  test\tstring";

    char_separator<char> sep(" \t");
    tokenizer<char_separator<char>> tokens(text, sep);
    for (const string& t : tokens)
    {
        cout << t << "." << endl;
    }
}

2008-10-25 10:58:25

如何迭代字符串的单词？

推荐文章

最新文章

标签