生成具有给定(数值)分布的随机数

我有一个文件，不同的值的一些概率，例如:

我想用这个分布生成随机数。是否存在处理此问题的现有模块?自己编写代码是相当简单的(构建累积密度函数，生成一个随机值[0,1]并选择相应的值)，但这似乎应该是一个常见的问题，可能有人已经为它创建了一个函数/模块。

我需要这个，因为我想生成一个生日列表(它不遵循标准随机模块中的任何分布)。

当前回答

也许有点晚了。但是你可以使用numpy.random.choice()，传递p参数:

val = numpy.random.choice(numpy.arange(1, 7), p=[0.1, 0.05, 0.05, 0.2, 0.4, 0.2])

2013-12-01 00:59:23

其他回答

你可能想看看NumPy随机抽样分布

2010-11-24 11:15:15

另一个答案，可能更快:)

distribution = [(1, 0.2), (2, 0.3), (3, 0.5)]  
# init distribution  
dlist = []  
sumchance = 0  
for value, chance in distribution:  
    sumchance += chance  
    dlist.append((value, sumchance))  
assert sumchance == 1.0 # not good assert because of float equality  

# get random value  
r = random.random()  
# for small distributions use lineair search  
if len(distribution) < 64: # don't know exact speed limit  
    for value, sumchance in dlist:  
        if r < sumchance:  
            return value  
else:  
    # else (not implemented) binary search algorithm

2010-11-24 11:38:00

(好吧，我知道你想要薄膜包装，但也许这些自制的解决方案对你来说不够简洁。: -)

pdf = [(1, 0.1), (2, 0.05), (3, 0.05), (4, 0.2), (5, 0.4), (6, 0.2)]
cdf = [(i, sum(p for j,p in pdf if j < i)) for i,_ in pdf]
R = max(i for r in [random.random()] for i,c in cdf if c <= r)

我伪确认，这是通过目测这个表达式的输出:

sorted(max(i for r in [random.random()] for i,c in cdf if c <= r)
       for _ in range(1000))

2010-11-24 11:32:46

from __future__ import division
import random
from collections import Counter


def num_gen(num_probs):
    # calculate minimum probability to normalize
    min_prob = min(prob for num, prob in num_probs)
    lst = []
    for num, prob in num_probs:
        # keep appending num to lst, proportional to its probability in the distribution
        for _ in range(int(prob/min_prob)):
            lst.append(num)
    # all elems in lst occur proportional to their distribution probablities
    while True:
        # pick a random index from lst
        ind = random.randint(0, len(lst)-1)
        yield lst[ind]

验证:

gen = num_gen([(1, 0.1),
               (2, 0.05),
               (3, 0.05),
               (4, 0.2),
               (5, 0.4),
               (6, 0.2)])
lst = []
times = 10000
for _ in range(times):
    lst.append(next(gen))
# Verify the created distribution:
for item, count in Counter(lst).iteritems():
    print '%d has %f probability' % (item, count/times)

1 has 0.099737 probability
2 has 0.050022 probability
3 has 0.049996 probability 
4 has 0.200154 probability
5 has 0.399791 probability
6 has 0.200300 probability

2015-05-02 00:10:33

基于其他解决方案，您可以生成累积分布(作为整数或浮点数)，然后您可以使用平分使其更快

这是一个简单的例子(我在这里使用整数)

l=[(20, 'foo'), (60, 'banana'), (10, 'monkey'), (10, 'monkey2')]
def get_cdf(l):
    ret=[]
    c=0
    for i in l: c+=i[0]; ret.append((c, i[1]))
    return ret

def get_random_item(cdf):
    return cdf[bisect.bisect_left(cdf, (random.randint(0, cdf[-1][0]),))][1]

cdf=get_cdf(l)
for i in range(100): print get_random_item(cdf),

get_cdf函数会将20、60、10、10转换为20、20+60、20+60+10、20+60+10+10

现在我们随机选择一个20+60+10+10的随机数。然后我们用二分法快速得到实际值

2016-04-26 09:41:45

生成具有给定(数值)分布的随机数

推荐文章

最新文章

标签