用1mb RAM对100万个8位数进行排序

我有一台有1mb内存的电脑，没有其他本地存储。我必须使用它通过TCP连接接受100万个8位十进制数字，对它们进行排序，然后通过另一个TCP连接发送排序的列表。

数字列表可能包含重复的，我不能丢弃。代码将放在ROM中，所以我不需要从1 MB中减去我的代码的大小。我已经有了驱动以太网端口和处理TCP/IP连接的代码，它需要2 KB的状态数据，包括1 KB的缓冲区，代码将通过它读取和写入数据。这个问题有解决办法吗?

问答来源:

slashdot.org

cleaton.net

当前回答

你用的是哪种电脑?它可能没有任何其他“正常”的本地存储，但它是否有视频RAM，例如?100万像素x每像素32位(比如说)非常接近你所需的数据输入大小。

(我主要是问旧的Acorn RISC PC的内存，如果你选择低分辨率或低颜色深度的屏幕模式，它可以“借用”VRAM来扩展可用的系统RAM !)这在只有几MB普通RAM的机器上非常有用。

2012-10-21 20:15:11

其他回答

在接收流时执行这些步骤。

首先设置一些合理的块大小

伪代码思想:

The first step would be to find all the duplicates and stick them in a dictionary with its count and remove them. The third step would be to place number that exist in sequence of their algorithmic steps and place them in counters special dictionaries with the first number and their step like n, n+1..., n+2, 2n, 2n+1, 2n+2... Begin to compress in chunks some reasonable ranges of number like every 1000 or ever 10000 the remaining numbers that appear less often to repeat. Uncompress that range if a number is found and add it to the range and leave it uncompressed for a while longer. Otherwise just add that number to a byte[chunkSize]

在接收流时继续执行前4步。最后一步是，如果超出内存，则失败，或者在收集完所有数据后开始输出结果，即开始对范围进行排序，并按顺序输出结果，然后按需要解压缩的顺序解压结果，并在得到它们时对它们进行排序。

2014-12-02 21:31:17

我认为解决方案是结合视频编码的技术，即离散余弦变换。在数字视频中，不是将视频的亮度或颜色的变化记录为常规值，如110 112 115 116，而是从最后一个中减去每一个(类似于运行长度编码)。110 112 115 116变成110 2 3 1。这些值，2,3 1比原始值需要更少的比特。

So lets say we create a list of the input values as they arrive on the socket. We are storing in each element, not the value, but the offset of the one before it. We sort as we go, so the offsets are only going to be positive. But the offset could be 8 decimal digits wide which this fits in 3 bytes. Each element can't be 3 bytes, so we need to pack these. We could use the top bit of each byte as a "continue bit", indicating that the next byte is part of the number and the lower 7 bits of each byte need to be combined. zero is valid for duplicates.

当列表填满时，数字之间的距离应该越来越近，这意味着平均只有1个字节用于确定到下一个值的距离。7位值和1位偏移(如果方便的话)，但可能存在一个“继续”值需要少于8位的最佳点。

总之，我做了一些实验。我使用随机数生成器，我可以将100万个排序过的8位十进制数字放入大约1279000字节。每个数字之间的平均间隔始终是99…

public class Test {
    public static void main(String[] args) throws IOException {
        // 1 million values
        int[] values = new int[1000000];

        // create random values up to 8 digits lrong
        Random random = new Random();
        for (int x=0;x<values.length;x++) {
            values[x] = random.nextInt(100000000);
        }
        Arrays.sort(values);

        ByteArrayOutputStream baos = new ByteArrayOutputStream();

        int av = 0;    
        writeCompact(baos, values[0]);     // first value
        for (int x=1;x<values.length;x++) {
            int v = values[x] - values[x-1];  // difference
            av += v;
            System.out.println(values[x] + " diff " + v);
            writeCompact(baos, v);
        }

        System.out.println("Average offset " + (av/values.length));
        System.out.println("Fits in " + baos.toByteArray().length);
    }

    public static void writeCompact(OutputStream os, long value) throws IOException {
        do {
            int b = (int) value & 0x7f;
            value = (value & 0x7fffffffffffffffl) >> 7;
            os.write(value == 0 ? b : (b | 0x80));
        } while (value != 0);
    }
}

2012-10-22 08:33:28

我认为从组合学的角度来思考这个问题:有多少种可能的排序数字的组合?如果我们给出的组合是0,0,0 ....，0代码0，和0,0,0，…，1代码1，和999999999,99999999，…99999999是代码N, N是什么?换句话说，结果空间有多大?

Well, one way to think about this is noticing that this is a bijection of the problem of finding the number of monotonic paths in an N x M grid, where N = 1,000,000 and M = 100,000,000. In other words, if you have a grid that is 1,000,000 wide and 100,000,000 tall, how many shortest paths from the bottom left to the top right are there? Shortest paths of course require you only ever either move right or up (if you were to move down or left you would be undoing previously accomplished progress). To see how this is a bijection of our number sorting problem, observe the following:

您可以将路径中的任何水平支腿想象成排序中的一个数字，其中支腿的Y位置表示值。

所以如果路径只是向右移动一直到最后，然后一直跳到顶部，这相当于顺序为0,0,0，…，0。相反，如果它开始时一直跳到顶部，然后向右移动1,000,000次，这相当于999999999,99999999，……, 99999999。它向右移动一次，然后向上移动一次，然后向右移动一次，然后向上移动一次，等等，直到最后(然后必然会一直跳到顶部)，相当于0,1,2,3，…，999999。

幸运的是，这个问题已经解决了，这样的网格有(N + M)个选择(M)条路径:

(1,000,000 + 100,000,000)选择(100,000,000)~= 2.27 * 10^2436455

N因此等于2.27 * 10^2436455，因此代码0表示0,0,0，…，0和代码2.27 * 10^2436455，一些变化表示999999999,99999999，…, 99999999。

为了存储从0到2.27 * 10^2436455的所有数字，您需要lg2(2.27 * 10^2436455) = 8.0937 * 10^6位。

1兆字节= 8388608比特> 8093700比特

这样看来，我们至少有足够的空间来存储结果!当然，有趣的部分是在数字流进来时进行排序。不确定最好的方法是我们有294908位剩余。我想一个有趣的技巧是在每个点都假设这是整个排序，找到该排序的代码，然后当你收到一个新数字时，返回并更新之前的代码。手，手，手。

2012-10-21 22:46:38

下面是这类问题的一般解决方案:

一般程序

所采取的方法如下。该算法在一个32位字的缓冲区上操作。它在循环中执行以下过程:

We start with a buffer filled with compressed data from the last iteration. The buffer looks like this |compressed sorted|empty| Calculate the maximum amount of numbers that can be stored in this buffer, both compressed and uncompressed. Split the buffer into these two sections, beginning with the space for compressed data, ending with the uncompressed data. The buffer looks like |compressed sorted|empty|empty| Fill the uncompressed section with numbers to be sorted. The buffer looks like |compressed sorted|empty|uncompressed unsorted| Sort the new numbers with an in-place sort. The buffer looks like |compressed sorted|empty|uncompressed sorted| Right-align any already compressed data from the previous iteration in the compressed section. At this point the buffer is partitioned |empty|compressed sorted|uncompressed sorted| Perform a streaming decompression-recompression on the compressed section, merging in the sorted data in the uncompressed section. The old compressed section is consumed as the new compressed section grows. The buffer looks like |compressed sorted|empty|

执行此过程，直到所有数字都已排序。

压缩

当然，这种算法只有在知道实际要压缩什么之前，才有可能计算出新排序缓冲区的最终压缩大小。其次，压缩算法需要足够好来解决实际问题。

所使用的方法使用三个步骤。首先，算法将始终存储排序序列，因此我们可以只存储连续条目之间的差异。每个差值都在[0,99999999]的范围内。

这些差异随后被编码为一元比特流。这个流中的1表示“向累加器添加1,0表示“将累加器作为一个条目发出，并重置”。所以差N由N个1和1个0表示。

所有差异的和将接近算法支持的最大值，所有差异的计数将接近算法中插入的值的数量。这意味着我们期望流在最后包含最大值1和计数0。这允许我们计算流中0和1的期望概率。即，0的概率为count/(count+maxval)， 1的概率为maxval/(count+maxval)。

我们使用这些概率来定义这个比特流上的算术编码模型。这个算术代码将在最佳空间中精确地编码1和0的数量。我们可以计算该模型对于任何中间位流所使用的空间:bits = encoded * log2(1 + amount / maxval) + maxval * log2(1 + maxval / amount)。若要计算算法所需的总空间，请将encoded设置为amount。

为了不需要大量的迭代，可以向缓冲区添加少量开销。这将确保算法将至少对适合这个开销的数量进行操作，因为到目前为止，算法最大的时间成本是每个周期的算术编码压缩和解压缩。

除此之外，在算术编码算法的定点近似中，存储簿记数据和处理轻微的不准确性是需要一些开销的，但总的来说，即使使用可以包含8000个数字的额外缓冲区，该算法也能够容纳1MiB的空间，总共1043916字节的空间。

最优

除了减少算法的开销外，理论上不可能得到更小的结果。为了仅仅包含最终结果的熵，1011717个字节是必要的。如果我们减去为提高效率而增加的额外缓冲区，该算法使用1011916字节来存储最终结果+开销。

2020-04-01 17:23:28

假设这个任务是可能的。在输出之前，内存中会有一个百万个排序数字的表示。有多少种不同的表示法?由于可能有重复的数字，我们不能使用nCr(选择)，但有一种叫做multichoose的操作，它适用于多集。

在0..99,999,999范围内有22e2436455种方法来选择一百万个数字。这需要8,093,730位来表示每个可能的组合，或1,011,717字节。

所以理论上是可能的，如果你能想出一个合理(足够)的数字排序表。例如，一个疯狂的表示可能需要一个10MB的查找表或数千行代码。

但是，如果“1M RAM”意味着100万个字节，那么显然没有足够的空间。事实上，多5%的内存使它在理论上成为可能，这对我来说意味着表示必须非常有效，可能是不理智的。

2012-10-21 20:17:41

用1mb RAM对100万个8位数进行排序

推荐文章

最新文章

标签