从文件中获取第n行的Bash工具

有没有一种“规范”的方法？我一直在使用head-n|tail-1来完成这个任务，但我一直在想是否有一个Bash工具专门从文件中提取一行（或一系列行）。

我所说的“规范”是指一个主要功能就是这样做的程序。

当前回答

如果有多行由\n分隔（通常为新行）。您也可以使用“cut”：

echo "$data" | cut -f2 -d$'\n'

您将从文件中获得第二行-f3给你第三行。

2016-01-07 16:27:53

其他回答

您也可以使用sed打印并退出：

sed -n '10{p;q;}' file   # print line 10

2011-05-17 11:49:24

这个问题被标记为Bash，下面是Bash（≥4）的方法：使用带有-s（跳过）和-n（计数）选项的mapfile。

如果需要获取文件文件的第42行：

mapfile -s 41 -n 1 ary < file

此时，您将得到一个数组ary，其中的字段包含文件行（包括尾部换行符），我们跳过了前41行（-s 41），并在读取一行（-n 1）后停止。这真的是第42行。要打印出来：

printf '%s' "${ary[0]}"

如果您需要一系列行，请说出范围42–666（含），并说您不想自己计算，然后将它们打印在标准输出上：

mapfile -s $((42-1)) -n $((666-42+1)) ary < file
printf '%s' "${ary[@]}"

如果您也需要处理这些行，那么存储尾随换行符并不太方便。在这种情况下，使用-t选项（trim）：

mapfile -t -s $((42-1)) -n $((666-42+1)) ary < file
# do stuff
printf '%s\n' "${ary[@]}"

你可以让一个函数为你做这件事：

print_file_range() {
    # $1-$2 is the range of file $3 to be printed to stdout
    local ary
    mapfile -s $(($1-1)) -n $(($2-$1+1)) ary < "$3"
    printf '%s' "${ary[@]}"
}

没有外部命令，只有Bash内置命令！

2014-05-17 13:45:29

保存两次按键，不使用括号打印第N行：

sed  -n  Np  <fileName>
      ^   ^
       \   \___ 'p' for printing
        \______ '-n' for not printing by default

例如，要打印第100行：

sed -n 100p foo.txt

2021-05-19 14:18:22

大文件的最快解决方案始终是尾部|头部，前提是两个距离：

从文件开头到开始行。我们称之为S从最后一行到文件结尾的距离。是E吗

是已知的。然后，我们可以使用这个：

mycount="$E"; (( E > S )) && mycount="+$S"
howmany="$(( endline - startline + 1 ))"
tail -n "$mycount"| head -n "$howmany"

多少只是所需的行数。

更多详情请参见https://unix.stackexchange.com/a/216614/79743

2015-07-17 05:34:26

根据我的测试，就性能和可读性而言，我的建议是：

尾部-n+n|头部-1

N是您想要的行号。例如，tail-n+7 input.txt | head-1将打印文件的第7行。

tail-n+n将打印从第n行开始的所有内容，head-1将使其在一行之后停止。

可选的head-N|tail-1可能更可读。例如，这将打印第7行：

head-7 input.txt | tail-1

当谈到性能时，较小的文件大小没有太大的差异，但当文件变大时，尾部|头部（从上方）的性能会优于尾部|头部。

排名靠前的是“NUMq；d’很有意思，但我认为，与头/尾解决方案相比，开箱即用的人更少，而且它也比尾/头慢。

在我的测试中，两个尾部/头部版本都优于sed的NUMq；d’一致。这与发布的其他基准一致。很难找到尾巴/脑袋真的很坏的案例。这也不奇怪，因为这些操作在现代Unix系统中会被大量优化。

为了了解性能差异，以下是我从一个巨大文件（9.3G）中得到的数字：

tail-n+n | head-1:3.7秒头-N|尾-1:4.6秒sed Nq；d： 18.8秒

结果可能有所不同，但总体而言，性能头部|尾部和尾部|头部对于较小的输入来说是可比的，sed总是慢了一个重要因素（大约5倍左右）。

要复制我的基准测试，您可以尝试以下操作，但请注意，它将在当前工作目录中创建一个9.3G文件：

#!/bin/bash
readonly file=tmp-input.txt
readonly size=1000000000
readonly pos=500000000
readonly retries=3

seq 1 $size > $file
echo "*** head -N | tail -1 ***"
for i in $(seq 1 $retries) ; do
    time head "-$pos" $file | tail -1
done
echo "-------------------------"
echo
echo "*** tail -n+N | head -1 ***"
echo

seq 1 $size > $file
ls -alhg $file
for i in $(seq 1 $retries) ; do
    time tail -n+$pos $file | head -1
done
echo "-------------------------"
echo
echo "*** sed Nq;d ***"
echo

seq 1 $size > $file
ls -alhg $file
for i in $(seq 1 $retries) ; do
    time sed $pos'q;d' $file
done
/bin/rm $file

这是在我的机器上运行的输出（ThinkPad X1 Carbon，带有SSD和16G内存）。我假设在最后一次运行中，所有内容都将来自缓存，而不是磁盘：

*** head -N | tail -1 ***
500000000

real    0m9,800s
user    0m7,328s
sys     0m4,081s
500000000

real    0m4,231s
user    0m5,415s
sys     0m2,789s
500000000

real    0m4,636s
user    0m5,935s
sys     0m2,684s
-------------------------

*** tail -n+N | head -1 ***

-rw-r--r-- 1 phil 9,3G Jan 19 19:49 tmp-input.txt
500000000

real    0m6,452s
user    0m3,367s
sys     0m1,498s
500000000

real    0m3,890s
user    0m2,921s
sys     0m0,952s
500000000

real    0m3,763s
user    0m3,004s
sys     0m0,760s
-------------------------

*** sed Nq;d ***

-rw-r--r-- 1 phil 9,3G Jan 19 19:50 tmp-input.txt
500000000

real    0m23,675s
user    0m21,557s
sys     0m1,523s
500000000

real    0m20,328s
user    0m18,971s
sys     0m1,308s
500000000

real    0m19,835s
user    0m18,830s
sys     0m1,004s

2017-07-31 13:10:02

从文件中获取第n行的Bash工具

推荐文章

最新文章

标签