NumPy Mathematical & Statistical Functions: Calculate and Analyze Data Efficiently
NumPy comes packed with fast mathematical and statistical functions that work on entire arrays at once. From simple sums to advanced statistics, these tools make data analysis feel effortless! 📊
Why NumPy's Functions Are Special
These functions are vectorized — they operate on whole arrays without loops.
They're also optimized in C, making them much faster than pure Python alternatives.
Basic Aggregation Functions
Start with the classics: sum, mean, min, max.
import numpy as np
arr = np.array([10, 20, 30, 40, 50])
print("Sum:", np.sum(arr))
print("Mean:", np.mean(arr))
print("Min:", np.min(arr))
print("Max:", np.max(arr))
Sum: 150
Mean: 30.0
Min: 10
Max: 50
Or use methods directly:
print(arr.sum())
print(arr.mean())
The Power of 'axis' Parameter
For 2D arrays, axis controls direction:
axis=0: Down columnsaxis=1: Across rows- No axis: Entire array
Real-World Example: Student Marks Analysis
marks = np.array([
[85, 88, 92], # Alice
[90, 76, 85], # Bob
[78, 92, 88], # Charlie
[92, 85, 79], # David
[88, 90, 94] # Eva
])
print("Marks:\n", marks)
Marks:
[[85 88 92]
[90 76 85]
[78 92 88]
[92 85 79]
[88 90 94]]
Subject-wise Statistics (axis=0)
subject_mean = np.mean(marks, axis=0)
subject_min = np.min(marks, axis=0)
subject_max = np.max(marks, axis=0)
print("Subject means:", subject_mean)
print("Lowest per subject:", subject_min)
print("Highest per subject:", subject_max)
Subject means: [86.6 86.2 87.6]
Lowest per subject: [78 76 79]
Highest per subject: [92 92 94]
Student-wise Statistics (axis=1)
student_total = np.sum(marks, axis=1)
student_avg = np.mean(marks, axis=1)
print("Totals:", student_total)
print("Averages:", student_avg)
Totals: [265 251 258 256 272]
Averages: [88.33333333 83.66666667 86. 85.33333333 90.66666667]
Spread and Variability
# Standard deviation and variance
subject_std = np.std(marks, axis=0)
subject_var = np.var(marks, axis=0)
# Overall spread
overall_std = np.std(marks)
print("Subject std dev:", np.round(subject_std, 2))
print("Subject variance:", np.round(subject_var, 2))
print("Overall std dev:", round(overall_std, 2))
Subject std dev: [4.92 5.69 5.5 ]
Subject variance: [24.24 32.41 30.24]
Overall std dev: 5.57
Science has the highest spread — scores vary most!
Finding Positions: argmin and argmax
< magazin># Who got the highest total?
top_student_idx = np.argmax(student_total)
lowest_student_idx = np.argmin(student_total)
students = ['Alice', 'Bob', 'Charlie', 'David', 'Eva']
print(f"Top performer: {students[top_student_idx]} ({student_total[top_student_idx]} marks)")
print(f"Lowest total: {students[lowest_student_idx]} ({student_total[lowest_student_idx]} marks)")
Top performer: Eva (272 marks)
Lowest total: Bob (251 marks)
Advanced Statistical Functions
Median and Percentiles
overall_median = np.median(marks)
percentile_75 = np.percentile(marks, 75)
percentile_25 = np.percentile(marks, 25)
print("Overall median:", overall_median)
print("75th percentile:", percentile_75)
print("25th percentile:", percentile_25)
Overall median: 88.0
75th percentile: 92.0
25th percentile: 85.0
Correlation
# How subjects relate
correlation_matrix = np.corrcoef(marks.T) # Transpose for subjects as rows
print("Correlation matrix:\n", np.round(correlation_matrix, 2))
Correlation matrix:
[[ 1. 0.07 -0.39]
[ 0.07 1. 0.36]
[-0.39 0.36 1. ]]
Weak correlations — performance in one subject doesn't strongly predict others.
Histogram for Distribution
all_marks = marks.flatten()
hist, bins = np.histogram(all_marks, bins=5)
print("Histogram counts:", hist)
print("Bin edges:", bins)
Histogram counts: [2 4 4 3 2]
Bin edges: [76. 80.4 84.8 89.2 93.6 98. ]
Mathematical Functions (Trigonometric, Log, etc.)
angles = np.array([0, 30, 45, 60, 90]) * np.pi / 180 # Degrees to radians
print("Sine:", np.round(np.sin(angles), 2))
print("Cosine:", np.round(np.cos(angles), 2))
print("Log of marks (example):", np.log(marks[0])) # Natural log
Sine: [0. 0.5 0.71 0.87 1. ]
Cosine: [1. 0.87 0.71 0.5 0. ]
Log of marks (example): [4.443 4.477 4.522]
Beginner Mistakes - Common Errors and How to Avoid Them
axis parameter — it completely changes the result!
sum() or min() on large arrays — they're slower than NumPy versions.
np.log to zero or negative values — it returns warnings or NaN.
Optimization Tips
out parameter to avoid temporary arrays: np.sum(arr, out=result)
np.nanmean, np.nanstd to ignore missing values.
Real-World Use Cases
- Finance: Portfolio returns, risk (std dev), correlation analysis
- Science: Signal processing with trig functions
- Machine Learning: Feature statistics, normalization
- Business: Sales metrics, percentiles for rankings
Quick Summary 📝
- Basic:
sum,mean,min/max - Positions:
argmin/argmax - Spread:
std,var,median,percentile - Relations:
corrcoef - Math:
sin,log,exp, etc.
NumPy's math and stats functions turn raw numbers into meaningful insights. Practice daily — soon you'll analyze data effortlessly! ✨
Comments
Post a Comment