In NumPy, certain operations produce results that aren't ordinary numbers. These special values — np.nan, np.inf, and np.NINF — represent missing data, infinity, and negative infinity. Understanding them is crucial for clean, reliable data work! 🔍
What Are These Special Values?
NumPy uses IEEE 754 floating-point standards for these:
np.nan: "Not a Number" — for undefined or missing resultsnp.inf: Positive infinitynp.NINF(or-np.inf): Negative infinity
Analogy: np.nan is like a blank answer on a test — something went wrong. Infinity is like dividing by zero — the result grows without bound.
Creating Special Values
import numpy as np
print("NaN:", np.nan)
print("Positive infinity:", np.inf)
print("Negative infinity:", np.NINF) # or -np.inf
print("From operations:")
print("0/0:", np.array([0.0]) / 0) # nan
print("1/0:", np.array([1.0]) / 0) # inf
print("-1/0:", np.array([-1.0]) / 0) # -inf
NaN: nan
Positive infinity: inf
Negative infinity: -inf
From operations:
0/0: [nan]
1/0: [inf]
-1/0: [-inf]
How They Behave in Operations
Special values propagate in predictable ways.
arr = np.array([1, 2, np.nan, 4])
print("Sum with nan:", np.sum(arr)) # nan
print("Any operation with nan → nan:")
print(arr + 10) # [11. 12. nan 14.]
print(np.nan + 5) # nan
print("Infinity:")
print(np.inf - np.inf) # nan
print(np.inf / np.inf) # nan
print(10 * np.inf) # inf
print(-5 * np.inf) # -inf
Sum with nan: nan
Any operation with nan → nan:
[11. 12. nan 14.]
nan
Infinity:
nan
nan
inf
-inf
Detecting Special Values
Use dedicated functions — never compare directly!
data = np.array([1, np.nan, np.inf, -np.inf, 5])
print("isnan:", np.isnan(data))
print("isinf:", np.isinf(data))
print("isfinite:", np.isfinite(data)) # Not nan or inf
print("isnan or isinf:", np.isnan(data) | np.isinf(data))
isnan: [False True False False False]
isinf: [False False True True False]
isfinite: [ True False False False True]
isnan or isinf: [False True True True False]
np.nan == np.nan is False! Always use np.isnan().
Real-World Example: Handling Missing Student Marks
marks = np.array([
[85, 88, np.nan], # Missing English score
[90, 76, 85],
[78, np.nan, 88],
[92, 85, 79],
[88, 90, 94]
])
print("Marks with missing:\n", marks)
# Count missing
missing = np.isnan(marks)
print("Missing values:\n", missing)
print("Total missing:", np.sum(missing))
# Safe statistics (ignore nan)
safe_mean = np.nanmean(marks, axis=0)
safe_sum = np.nansum(marks, axis=1)
print("Subject means (safe):", np.round(safe_mean, 1))
print("Student totals (safe):", safe_sum)
Marks with missing:
[[85. 88. nan]
[90. 76. 85.]
[78. nan 88.]
[92. 85. 79.]
[88. 90. 94.]]
Missing values:
[[False False True]
[False False False]
[False True False]
[False False False]
[False False False]]
Total missing: 2
Subject means (safe): [86.6 84.8 86.5]
Student totals (safe): [173. 251. 166. 256. 272.]
Replacing or Removing NaN
# Replace with 0 (or any value)
filled_zero = np.nan_to_num(marks, nan=0)
print("Filled with 0:\n", filled_zero)
# Replace with mean
col_means = np.nanmean(marks, axis=0)
filled_mean = np.where(np.isnan(marks), col_means, marks)
print("Filled with mean:\n", np.round(filled_mean, 1))
Beginner Mistakes - Common Errors and How to Avoid Them
sum() or mean() on data with nan — results become nan!
== np.nan — always use np.isnan().
Optimization Tips
np.nan_to_num() for quick cleanup before math-heavy operations.
nanmean, nanstd) — they're fast and accurate.
Real-World Use Cases
- Data Cleaning: Handle missing sensor readings or survey responses
- Scientific Computing: Results from overflow or undefined math
- Machine Learning: Mark missing features, avoid breaking models
- Finance: Infinity from invalid rates or divisions
Quick Summary 📝
np.nan: Missing/invalid — propagates, useisnan()np.inf/np.NINF: Infinity from overflow or division by zero- Use nan-functions for safe statistics
- Detect with
isnan,isinf,isfinite
Special values aren't errors — they're information. Handle them wisely, and your analyses will be robust! Happy coding! ✨
Comments
Post a Comment