多处理vs线程Python

我试图理解多处理相对于线程的优势。我知道多处理绕过了全局解释器锁，但是还有什么其他的优势，线程不能做同样的事情吗?

当前回答

Python文档引用

这个答案的规范版本现在是双重问题:线程模块和多处理模块之间有什么区别?

我已经突出显示了Python文档中关于进程vs线程和GIL的关键引用:什么是CPython中的全局解释器锁(GIL) ?

进程与线程实验

为了更具体地展示差异，我做了一些基准测试。

在基准测试中，我对8超线程CPU上不同数量的线程进行了CPU和IO限制。每个线程提供的功总是相同的，因此线程越多，提供的总功就越多。

结果如下:

图数据。

结论:

对于CPU约束的工作，多处理总是更快，可能是由于GIL IO绑定工作。两者的速度完全相同线程只能扩展到大约4倍，而不是预期的8倍，因为我使用的是8超线程机器。与此相比，C POSIX cpu绑定的工作达到了预期的8倍加速:'real'， 'user'和'sys'在time(1)的输出中是什么意思? 我不知道这是什么原因，肯定有其他Python的低效率因素在起作用。

测试代码:

#!/usr/bin/env python3

import multiprocessing
import threading
import time
import sys

def cpu_func(result, niters):
    '''
    A useless CPU bound function.
    '''
    for i in range(niters):
        result = (result * result * i + 2 * result * i * i + 3) % 10000000
    return result

class CpuThread(threading.Thread):
    def __init__(self, niters):
        super().__init__()
        self.niters = niters
        self.result = 1
    def run(self):
        self.result = cpu_func(self.result, self.niters)

class CpuProcess(multiprocessing.Process):
    def __init__(self, niters):
        super().__init__()
        self.niters = niters
        self.result = 1
    def run(self):
        self.result = cpu_func(self.result, self.niters)

class IoThread(threading.Thread):
    def __init__(self, sleep):
        super().__init__()
        self.sleep = sleep
        self.result = self.sleep
    def run(self):
        time.sleep(self.sleep)

class IoProcess(multiprocessing.Process):
    def __init__(self, sleep):
        super().__init__()
        self.sleep = sleep
        self.result = self.sleep
    def run(self):
        time.sleep(self.sleep)

if __name__ == '__main__':
    cpu_n_iters = int(sys.argv[1])
    sleep = 1
    cpu_count = multiprocessing.cpu_count()
    input_params = [
        (CpuThread, cpu_n_iters),
        (CpuProcess, cpu_n_iters),
        (IoThread, sleep),
        (IoProcess, sleep),
    ]
    header = ['nthreads']
    for thread_class, _ in input_params:
        header.append(thread_class.__name__)
    print(' '.join(header))
    for nthreads in range(1, 2 * cpu_count):
        results = [nthreads]
        for thread_class, work_size in input_params:
            start_time = time.time()
            threads = []
            for i in range(nthreads):
                thread = thread_class(work_size)
                threads.append(thread)
                thread.start()
            for i, thread in enumerate(threads):
                thread.join()
            results.append(time.time() - start_time)
        print(' '.join('{:.6e}'.format(result) for result in results))

相同目录上的GitHub上游+绘图代码。

在Ubuntu 18.10, Python 3.6.7，联想ThinkPad P51笔记本电脑上测试，CPU:英特尔酷睿i7-7820HQ CPU(4核/ 8线程)，RAM: 2倍三星M471A2K43BB1-CRC(2倍16GiB)， SSD:三星MZVLB512HAJQ-000L7 (3000 MB/s)。

可视化给定时间哪些线程正在运行

这篇文章https://rohanvarma.me/GIL/告诉我，你可以运行一个回调每当线程调度与目标=参数的线程。线程和multiprocessing.Process。

这允许我们准确地查看每次运行的线程。当这完成后，我们会看到(我制作了这张特殊的图表):

            +--------------------------------------+
            + Active threads / processes           +
+-----------+--------------------------------------+
|Thread   1 |********     ************             |
|         2 |        *****            *************|
+-----------+--------------------------------------+
|Process  1 |***  ************** ******  ****      |
|         2 |** **** ****** ** ********* **********|
+-----------+--------------------------------------+
            + Time -->                             +
            +--------------------------------------+

这将表明:

线程由GIL完全序列化进程可以并行运行

2019-03-23 23:23:33

其他回答

关键的优势是隔离。进程崩溃不会导致其他进程崩溃，而线程崩溃可能会对其他线程造成严重破坏。

2010-06-15 11:15:29

线程模块使用线程，多处理模块使用进程。不同之处在于线程在相同的内存空间中运行，而进程有单独的内存。这使得在多进程之间共享对象变得有点困难。由于线程使用相同的内存，必须采取预防措施，否则两个线程将同时写入同一内存。这就是全局解释器锁的作用。

生成进程比生成线程要慢一些。

2010-06-15 11:19:25

另一件没有提到的事情是，它取决于你使用的是什么操作系统。在Windows中，进程是昂贵的，所以线程在Windows中会更好，但在unix中，进程比它们的Windows变体更快，所以在unix中使用进程要安全得多，而且生成速度快。

2010-06-15 11:22:41

Threading's job is to enable applications to be responsive. Suppose you have a database connection and you need to respond to user input. Without threading, if the database connection is busy the application will not be able to respond to the user. By splitting off the database connection into a separate thread you can make the application more responsive. Also because both threads are in the same process, they can access the same data structures - good performance, plus a flexible software design.

注意，由于GIL，应用程序实际上并没有同时做两件事，但我们所做的是将数据库上的资源锁放在一个单独的线程中，这样CPU时间就可以在它和用户交互之间切换。CPU时间在线程之间分配。

Multiprocessing is for times when you really do want more than one thing to be done at any given time. Suppose your application needs to connect to 6 databases and perform a complex matrix transformation on each dataset. Putting each job in a separate thread might help a little because when one connection is idle another one could get some CPU time, but the processing would not be done in parallel because the GIL means that you're only ever using the resources of one CPU. By putting each job in a Multiprocessing process, each can run on it's own CPU and run at full efficiency.

2010-06-15 13:38:10