Easy Guide to dropna() 🧹
Missing values (like NaN) can mess up your data. Pandas has a simple tool called dropna() to delete them. It works on Series (lists) and DataFrames (tables). Let's learn it step by step in easy English!
For a Series (Simple List)
dropna() removes all missing values and keeps only the good ones.
import pandas as pd
import numpy as np
s = pd.Series([1, np.nan, 3, np.nan, 5])
print(s)
Original Series:
0 1.0
1 NaN
2 3.0
3 NaN
4 5.0
dtype: float64
print(s.dropna())
After dropna():
0 1.0
2 3.0
4 5.0
dtype: float64
All NaN are gone! 🎉
For a DataFrame (Table)
In tables, you can control what to delete with extra options.
df = pd.DataFrame({'A': [1, 2, np.nan], 'B': [np.nan, 1, 2]})
print(df)
Original Table:
A B
0 1.0 NaN
1 2.0 1.0
2 NaN 2.0
Common Options
- Default: Delete rows that have any NaN
axis=1: Delete columns instead of rowshow='all': Delete only if all values in row/column are NaNthresh=2: Keep rows that have at least 2 good values
Examples
print(df.dropna()) # Default: drop rows with any NaN
A B
1 2.0 1.0
print(df.dropna(axis=1)) # Drop columns with any NaN
Empty DataFrame
Columns: []
Index: [0, 1, 2]
print(df.dropna(how='all')) # Only drop if ALL are NaN (none here)
A B
0 1.0 NaN
1 2.0 1.0
2 NaN 2.0
print(df.dropna(thresh=2)) # Keep rows with 2+ good values
A B
1 2.0 1.0
That's it! Use dropna() to clean your data quickly. Happy coding! 🚀
Comments
Post a Comment