Skip to main content

Removing Missing Values in Pandas

Calculating read time…

Easy Guide to dropna() 🧹

Missing values (like NaN) can mess up your data. Pandas has a simple tool called dropna() to delete them. It works on Series (lists) and DataFrames (tables). Let's learn it step by step in easy English!

For a Series (Simple List)

dropna() removes all missing values and keeps only the good ones.

import pandas as pd
import numpy as np

s = pd.Series([1, np.nan, 3, np.nan, 5])
print(s)

Original Series:

0    1.0
1    NaN
2    3.0
3    NaN
4    5.0
dtype: float64
print(s.dropna())

After dropna():

0    1.0
2    3.0
4    5.0
dtype: float64

All NaN are gone! 🎉

For a DataFrame (Table)

In tables, you can control what to delete with extra options.

df = pd.DataFrame({'A': [1, 2, np.nan], 'B': [np.nan, 1, 2]})
print(df)

Original Table:

     A    B
0  1.0  NaN
1  2.0  1.0
2  NaN  2.0

Common Options

  • Default: Delete rows that have any NaN
  • axis=1: Delete columns instead of rows
  • how='all': Delete only if all values in row/column are NaN
  • thresh=2: Keep rows that have at least 2 good values

Examples

print(df.dropna())            # Default: drop rows with any NaN
     A    B
1  2.0  1.0
print(df.dropna(axis=1))      # Drop columns with any NaN
Empty DataFrame
Columns: []
Index: [0, 1, 2]
print(df.dropna(how='all'))   # Only drop if ALL are NaN (none here)
     A    B
0  1.0  NaN
1  2.0  1.0
2  NaN  2.0
print(df.dropna(thresh=2))    # Keep rows with 2+ good values
     A    B
1  2.0  1.0

That's it! Use dropna() to clean your data quickly. Happy coding! 🚀

Comments