Skip to main content

NumPy Random Module: Generate Random Numbers, Arrays, and Samples

Calculating read time…

Random numbers are everywhere in programming: testing code, simulating real-world events, or creating fake datasets. NumPy makes generating them fast, repeatable, and flexible! 🎲

Why NumPy Random?

NumPy's random functions are:

  • Super fast (vectorized)
  • Reproducible with seeds
  • Full of real-world distributions
🟢 DO: Use NumPy random for any project needing randomness — it's more reliable than Python's built-in random module!

Setting a Seed - Make Results Repeatable

Randomness should be controllable for testing and learning.

import numpy as np

# Set seed for reproducible results
np.random.seed(42)

print(np.random.rand(5))  # Always the same with seed 42
[0.37454012 0.95071431 0.73199394 0.59865848 0.15601864]

Without seed, results change every run. With seed — perfect for debugging!

Basic Random Numbers

Uniform Floats (0 to 1)

np.random.seed(42)

print("5 random floats:", np.random.rand(5))
print("3x3 matrix:\n", np.random.rand(3, 3))
5 random floats: [0.37454012 0.95071431 0.73199394 0.59865848 0.15601864]
3x3 matrix:
 [[0.15599452 0.05808361 0.86617615]
 [0.60111501 0.70807258 0.02058449]
 [0.96990985 0.83244264 0.21233911]]

Random Integers

print("Integers 1-100:", np.random.randint(1, 101, size=10))
print("Dice rolls (1-6):", np.random.randint(1, 7, size=8))
Integers 1-100: [83  4 30 70 56 53  3 57 87 92]
Dice rolls (1-6): [2 6 1 4 6 5 2 4]

Common Distributions

Normal (Gaussian) - Bell Curve

np.random.seed(42)

# Mean=70, std=10, 1000 exam scores
scores = np.random.normal(70, 10, 1000)

print("Sample scores:", scores[:10].round(1))
print("Mean:", scores.mean().round(1))
print("Std dev:", scores.std().round(1))
Sample scores: [84.7 95.1 83.2 80.0 71.6 65.6 58.3 86.6 85.3 70.6]
Mean: 70.1
Std dev: 9.9

Binomial - Yes/No Trials

# 10 coin flips, p=0.5 fair coin
flips = np.random.binomial(10, 0.5, size=8)
print("Heads in 10 flips:", flips)
Heads in 10 flips: [5 4 6 5 7 5 5 6]

Poisson - Rare Events

# Average 3 customers per hour
customers = np.random.poisson(3, size=10)
print("Customers per hour:", customers)
Customers per hour: [3 2 4 1 3 5 2 4 3 2]

Sampling from Arrays

students = np.array(['Alice', 'Bob', 'Charlie', 'David', 'Eva'])

# Pick 3 with replacement
print("Random winners:", np.random.choice(students, 3))

# Without replacement
print("Unique picks:", np.random.choice(students, 3, replace=False))

# Shuffle in place
np.random.shuffle(students)
print("Shuffled:", students)
Random winners: ['Bob' 'David' 'Alice']
Unique picks: ['Eva' 'Charlie' 'Bob']
Shuffled: ['Charlie' 'Alice' 'Eva' 'David' 'Bob']

Real-World Example: Simulating a Class of Students

Generate realistic fake data for testing analysis code.

np.random.seed(42)

n_students = 50

# Realistic exam scores (mean 75, std 12)
exam_scores = np.random.normal(75, 12, n_students)

# Clip to 0-100
exam_scores = np.clip(exam_scores, 0, 100)

# Attendance percentage (80-100%)
attendance = np.random.uniform(80, 100, n_students)

# Round and convert
exam_scores = np.round(exam_scores).astype(int)
attendance = np.round(attendance, 1)

print("First 10 scores:", exam_scores[:10])
print("First 10 attendance:", attendance[:10])
print("\nClass average score:", exam_scores.mean().round(1))
print("Average attendance:", attendance.mean().round(1))
First 10 scores: [90 87 79 85 74 73 70 80 86 88]
First 10 attendance: [92.5 99.1 81.8 94.8 96.8 85.3 93.9 81.3 93.9 93.8]

Class average score: 75.2
Average attendance: 89.9

Perfect for testing grading logic or visualizations!

Beginner Mistakes - Common Errors and How to Avoid Them

🔴 DON'T: Forget to set a seed when you need reproducible results — random code is hard to debug!
🔴 DON'T: Use randint(low, high) with high inclusive wrong — it's [low, high) so use high+1 for inclusive.
🔴 DON'T: Generate huge arrays without checking memory — start small and scale up.
🟡 Tip: Always clip or round simulated data to realistic bounds!

Advanced Usage & Next Steps

Modern Generator (Recommended)

Newer NumPy uses default_rng() for better randomness:

rng = np.random.default_rng(42)

print(rng.normal(70, 10, 5).round(1))
print(rng.integers(1, 7, size=10))  # Inclusive low, exclusive high

Optimization Tips

🟢 DO: Switch to default_rng() in new projects — better algorithms and thread-safe.
🟢 DO: Generate large samples at once — vectorized is fastest.

Real-World Use Cases

  • Monte Carlo simulations: Risk analysis, physics
  • Machine Learning: Data augmentation, initialization
  • Games: Dice, cards, procedural content
  • Testing: Fake datasets for development

Quick Summary 📝

  • Set seed for reproducibility
  • rand(), randint() for basics
  • Normal, binomial, poisson for real distributions
  • choice(), shuffle() for sampling
  • Simulate realistic data easily

Random numbers open endless possibilities. Experiment freely — that's how you learn best! Happy simulating! ✨

Comments