Skip to main content

NumPy Mathematical & Statistical Functions: Calculate and Analyze Data Efficiently

Calculating read time…

NumPy comes packed with fast mathematical and statistical functions that work on entire arrays at once. From simple sums to advanced statistics, these tools make data analysis feel effortless! 📊

Why NumPy's Functions Are Special

These functions are vectorized — they operate on whole arrays without loops.

They're also optimized in C, making them much faster than pure Python alternatives.

🟢 DO: Use NumPy's built-in functions whenever possible — they're accurate, fast, and handle edge cases well!

Basic Aggregation Functions

Start with the classics: sum, mean, min, max.

import numpy as np

arr = np.array([10, 20, 30, 40, 50])

print("Sum:", np.sum(arr))
print("Mean:", np.mean(arr))
print("Min:", np.min(arr))
print("Max:", np.max(arr))
Sum: 150
Mean: 30.0
Min: 10
Max: 50

Or use methods directly:

print(arr.sum())
print(arr.mean())

The Power of 'axis' Parameter

For 2D arrays, axis controls direction:

  • axis=0: Down columns
  • axis=1: Across rows
  • No axis: Entire array

Real-World Example: Student Marks Analysis

marks = np.array([
    [85, 88, 92],  # Alice
    [90, 76, 85],  # Bob
    [78, 92, 88],  # Charlie
    [92, 85, 79],  # David
    [88, 90, 94]   # Eva
])

print("Marks:\n", marks)
Marks:
 [[85 88 92]
 [90 76 85]
 [78 92 88]
 [92 85 79]
 [88 90 94]]

Subject-wise Statistics (axis=0)

subject_mean = np.mean(marks, axis=0)
subject_min = np.min(marks, axis=0)
subject_max = np.max(marks, axis=0)

print("Subject means:", subject_mean)
print("Lowest per subject:", subject_min)
print("Highest per subject:", subject_max)
Subject means: [86.6 86.2 87.6]
Lowest per subject: [78 76 79]
Highest per subject: [92 92 94]

Student-wise Statistics (axis=1)

student_total = np.sum(marks, axis=1)
student_avg = np.mean(marks, axis=1)

print("Totals:", student_total)
print("Averages:", student_avg)
Totals: [265 251 258 256 272]
Averages: [88.33333333 83.66666667 86.         85.33333333 90.66666667]

Spread and Variability

# Standard deviation and variance
subject_std = np.std(marks, axis=0)
subject_var = np.var(marks, axis=0)

# Overall spread
overall_std = np.std(marks)

print("Subject std dev:", np.round(subject_std, 2))
print("Subject variance:", np.round(subject_var, 2))
print("Overall std dev:", round(overall_std, 2))
Subject std dev: [4.92 5.69 5.5 ]
Subject variance: [24.24 32.41 30.24]
Overall std dev: 5.57

Science has the highest spread — scores vary most!

Finding Positions: argmin and argmax

< magazin># Who got the highest total?
top_student_idx = np.argmax(student_total)
lowest_student_idx = np.argmin(student_total)

students = ['Alice', 'Bob', 'Charlie', 'David', 'Eva']

print(f"Top performer: {students[top_student_idx]} ({student_total[top_student_idx]} marks)")
print(f"Lowest total: {students[lowest_student_idx]} ({student_total[lowest_student_idx]} marks)")
Top performer: Eva (272 marks)
Lowest total: Bob (251 marks)

Advanced Statistical Functions

Median and Percentiles

overall_median = np.median(marks)
percentile_75 = np.percentile(marks, 75)
percentile_25 = np.percentile(marks, 25)

print("Overall median:", overall_median)
print("75th percentile:", percentile_75)
print("25th percentile:", percentile_25)
Overall median: 88.0
75th percentile: 92.0
25th percentile: 85.0

Correlation

# How subjects relate
correlation_matrix = np.corrcoef(marks.T)  # Transpose for subjects as rows

print("Correlation matrix:\n", np.round(correlation_matrix, 2))
Correlation matrix:
 [[ 1.    0.07 -0.39]
 [ 0.07  1.    0.36]
 [-0.39  0.36  1.  ]]

Weak correlations — performance in one subject doesn't strongly predict others.

Histogram for Distribution

all_marks = marks.flatten()
hist, bins = np.histogram(all_marks, bins=5)

print("Histogram counts:", hist)
print("Bin edges:", bins)
Histogram counts: [2 4 4 3 2]
Bin edges: [76.  80.4 84.8 89.2 93.6 98. ]

Mathematical Functions (Trigonometric, Log, etc.)

angles = np.array([0, 30, 45, 60, 90]) * np.pi / 180  # Degrees to radians

print("Sine:", np.round(np.sin(angles), 2))
print("Cosine:", np.round(np.cos(angles), 2))
print("Log of marks (example):", np.log(marks[0]))  # Natural log
Sine: [0.   0.5  0.71 0.87 1.  ]
Cosine: [1.   0.87 0.71 0.5  0.  ]
Log of marks (example): [4.443 4.477 4.522]

Beginner Mistakes - Common Errors and How to Avoid Them

🔴 DON'T: Forget the axis parameter — it completely changes the result!
🔴 DON'T: Use Python's built-in sum() or min() on large arrays — they're slower than NumPy versions.
🔴 DON'T: Apply np.log to zero or negative values — it returns warnings or NaN.

Optimization Tips

🟢 DO: Use out parameter to avoid temporary arrays: np.sum(arr, out=result)
🟢 DO: For very large data, use np.nanmean, np.nanstd to ignore missing values.

Real-World Use Cases

  • Finance: Portfolio returns, risk (std dev), correlation analysis
  • Science: Signal processing with trig functions
  • Machine Learning: Feature statistics, normalization
  • Business: Sales metrics, percentiles for rankings

Quick Summary 📝

  • Basic: sum, mean, min/max
  • Positions: argmin/argmax
  • Spread: std, var, median, percentile
  • Relations: corrcoef
  • Math: sin, log, exp, etc.

NumPy's math and stats functions turn raw numbers into meaningful insights. Practice daily — soon you'll analyze data effortlessly! ✨

Comments