Outliers are extreme values that are significantly different from other data points. They can skew your analysis, make averages misleading, and lead to incorrect conclusions. This tutorial shows you how to detect and handle outliers using Pandas with simple, practical examples.
Sample Dataset for Practice
Let's create a sample employee salary dataset with some outliers. You can use this exact dataset to practice all the outlier detection techniques!
import pandas as pd
import numpy as np
# Create sample dataset with outliers
data = {
'Employee_ID': [101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115],
'Name': ['Alice', 'Bob', 'Charlie', 'Diana', 'Eve', 'Frank', 'Grace', 'Henry', 'Ivy', 'Jack', 'Kate', 'Leo', 'Mia', 'Noah', 'Olivia'],
'Age': [28, 32, 45, 29, 35, 150, 27, 31, 40, 33, 38, 26, 30, 5, 42],
'Salary': [50000, 52000, 55000, 51000, 54000, 500000, 53000, 49000, 56000, 52500, 10000, 54500, 51500, 53500, 1000000],
'Experience': [3, 5, 15, 4, 8, 2, 3, 5, 12, 6, 1, 2, 4, 7, 20],
'Rating': [4.2, 4.5, 4.8, 4.3, 4.6, 3.9, 4.4, 4.1, 4.7, 4.5, 2.1, 4.3, 4.4, 9.5, 4.9]
}
df = pd.DataFrame(data)
print(df)
print("\nThis dataset contains outliers:")
print("- Extremely high salaries (500000, 1000000)")
print("- Extremely low salary (10000)")
print("- Invalid age (150, 5)")
print("- Invalid rating (9.5)")
print("\nLet's detect and handle these outliers!")
Dataset Preview:
| Employee_ID | Name | Age | Salary | Experience | Rating |
|---|---|---|---|---|---|
| 101 | Alice | 28 | $50,000 | 3 | 4.2 |
| 102 | Bob | 32 | $52,000 | 5 | 4.5 |
| 106 | Frank | 150 ⚠️ | $500,000 ⚠️ | 2 | 3.9 |
| 111 | Kate | 38 | $10,000 ⚠️ | 1 | 2.1 ⚠️ |
| 115 | Olivia | 42 | $1,000,000 ⚠️ | 20 | 4.9 |
Notice: Red marks (⚠️) indicate potential outliers we'll detect using statistical methods!
What Are Outliers?
Outliers are data points that are significantly different from other observations. They can occur due to:
- Data entry errors: Someone typed 1500000 instead of 150000
- Measurement errors: Faulty sensors or incorrect recording
- Natural variation: Genuine extreme values (CEO salary vs regular employee)
- Data processing errors: Mistakes during data cleaning or transformation
Why handle outliers? They can distort averages, affect machine learning models, and lead to wrong business decisions!
1. Detect Outliers Using IQR Method
The IQR (Interquartile Range) method is the most common way to detect outliers. It's based on the concept that outliers are values that fall far outside the normal range of data.
How it works: The IQR is the range between the 25th percentile (Q1) and 75th percentile (Q3). Any value below Q1 - 1.5×IQR or above Q3 + 1.5×IQR is considered an outlier.
# Calculate quartiles and IQR for Salary
Q1 = df['Salary'].quantile(0.25)
Q3 = df['Salary'].quantile(0.75)
IQR = Q3 - Q1
print(f"Q1 (25th percentile): ${Q1:,.0f}")
print(f"Q3 (75th percentile): ${Q3:,.0f}")
print(f"IQR (Q3 - Q1): ${IQR:,.0f}")
# Calculate outlier boundaries
lower_bound = Q1 - 1.5 * IQR
upper_bound = Q3 + 1.5 * IQR
print(f"\nOutlier boundaries:")
print(f"Lower bound: ${lower_bound:,.0f}")
print(f"Upper bound: ${upper_bound:,.0f}")
# Find outliers
outliers = df[(df['Salary'] < lower_bound) | (df['Salary'] > upper_bound)]
print(f"\nOutliers found: {len(outliers)}")
print(outliers[['Name', 'Salary']])
Example Output:
Q1 (25th percentile): $51,000
Q3 (75th percentile): $54,000
IQR (Q3 - Q1): $3,000
Outlier boundaries:
Lower bound: $46,500
Upper bound: $58,500
Outliers found: 3
Name Salary
5 Frank 500000
10 Kate 10000
14 Olivia 1000000
Real-World Use Case: A company analyzing employee salaries to ensure fair compensation and detect data entry errors!
2. Detect Outliers Using Z-Score Method
The Z-score tells you how many standard deviations a value is from the mean. A Z-score above 3 or below -3 is typically considered an outlier.
from scipy import stats
# Calculate Z-scores
df['Salary_Zscore'] = stats.zscore(df['Salary'])
# Find outliers (Z-score > 3 or < -3)
outliers_zscore = df[abs(df['Salary_Zscore']) > 3]
print(f"Outliers using Z-score method: {len(outliers_zscore)}")
print(outliers_zscore[['Name', 'Salary', 'Salary_Zscore']])
Example Output:
Outliers using Z-score method: 2
Name Salary Salary_Zscore
5 Frank 500000 3.45
14 Olivia 1000000 6.89
💡 Pro Tip: Z-score method works best when your data follows a normal distribution (bell curve)!
3. Visualize Outliers with Box Plot
A box plot is a great visual way to see outliers. Points outside the "whiskers" are outliers.
import matplotlib.pyplot as plt
# Create box plot
plt.figure(figsize=(10, 6))
plt.boxplot(df['Salary'], vert=False)
plt.xlabel('Salary ($)')
plt.title('Salary Distribution - Box Plot')
plt.grid(True, alpha=0.3)
plt.show()
# The dots beyond the whiskers are outliers!
What you'll see: The box shows Q1 to Q3 range, the line inside is the median, whiskers extend to 1.5×IQR, and dots beyond whiskers are outliers.
4. Handle Outliers - Method 1: Remove Them
Sometimes the best approach is to simply remove outliers, especially if they're data entry errors.
# Remove outliers using IQR method
df_no_outliers = df[(df['Salary'] >= lower_bound) & (df['Salary'] <= upper_bound)]
print(f"Original dataset: {len(df)} rows")
print(f"After removing outliers: {len(df_no_outliers)} rows")
print(f"Removed: {len(df) - len(df_no_outliers)} rows")
# Compare statistics
print("\nSalary Statistics:")
print(f"Original Mean: ${df['Salary'].mean():,.0f}")
print(f"After Removal Mean: ${df_no_outliers['Salary'].mean():,.0f}")
Example Output:
Original dataset: 15 rows
After removing outliers: 12 rows
Removed: 3 rows
Salary Statistics:
Original Mean: $152,633
After Removal Mean: $52,417
⚠️ Caution: Only remove outliers if you're sure they're errors. Don't remove genuine extreme values!
5. Handle Outliers - Method 2: Cap/Winsorize Them
Instead of removing outliers, you can cap them at the boundary values. This is called winsorization. It keeps all your data but limits extreme values.
# Cap outliers at boundary values
df['Salary_Capped'] = df['Salary'].clip(lower=lower_bound, upper=upper_bound)
print("Before and After Capping:")
print(df[['Name', 'Salary', 'Salary_Capped']])
print("\nSalary Range:")
print(f"Original range: ${df['Salary'].min():,.0f} - ${df['Salary'].max():,.0f}")
print(f"Capped range: ${df['Salary_Capped'].min():,.0f} - ${df['Salary_Capped'].max():,.0f}")
Example Output:
Before and After Capping:
Name Salary Salary_Capped
5 Frank 500000 58500 ← Capped!
10 Kate 10000 46500 ← Capped!
14 Olivia 1000000 58500 ← Capped!
Salary Range:
Original range: $10,000 - $1,000,000
Capped range: $46,500 - $58,500
Real-World Use Case: Machine learning models perform better with capped outliers instead of removed data!
6. Handle Outliers - Method 3: Flag Them
Sometimes you want to keep outliers but mark them for special attention. Create a flag column!
# Create outlier flag
df['Is_Salary_Outlier'] = (df['Salary'] < lower_bound) | (df['Salary'] > upper_bound)
# View flagged records
print("Flagged Outliers:")
print(df[df['Is_Salary_Outlier']][['Name', 'Salary', 'Is_Salary_Outlier']])
# Count outliers
outlier_count = df['Is_Salary_Outlier'].sum()
print(f"\nTotal outliers flagged: {outlier_count}")
print(f"Percentage of outliers: {outlier_count/len(df)*100:.1f}%")
Example Output:
Flagged Outliers:
Name Salary Is_Salary_Outlier
5 Frank 500000 True
10 Kate 10000 True
14 Olivia 1000000 True
Total outliers flagged: 3
Percentage of outliers: 20.0%
Pro Tip: Flagging is great when you need different handling for outliers vs normal data in your analysis!
7. Handle Outliers - Method 4: Transform Data
Apply mathematical transformations like log to reduce the impact of outliers without removing them.
import numpy as np
# Apply log transformation
df['Salary_Log'] = np.log1p(df['Salary']) # log1p handles zeros better
print("Original vs Log-Transformed:")
print(df[['Name', 'Salary', 'Salary_Log']].head(10))
# Compare standard deviations
print(f"\nOriginal Std Dev: ${df['Salary'].std():,.0f}")
print(f"Log-Transformed Std Dev: {df['Salary_Log'].std():.2f}")
print("\n Log transformation reduces the impact of extreme values!")
Real-World Use Case: Financial data and website traffic often use log transformations to handle naturally skewed distributions!
8. Detect Outliers Across Multiple Columns
Check for outliers in all numeric columns at once!
# Function to detect outliers in any column
def detect_outliers_iqr(df, column):
Q1 = df[column].quantile(0.25)
Q3 = df[column].quantile(0.75)
IQR = Q3 - Q1
lower = Q1 - 1.5 * IQR
upper = Q3 + 1.5 * IQR
outliers = df[(df[column] < lower) | (df[column] > upper)]
return outliers, lower, upper
# Check all numeric columns
numeric_columns = ['Age', 'Salary', 'Experience', 'Rating']
for col in numeric_columns:
outliers, lower, upper = detect_outliers_iqr(df, col)
print(f"\n{col}:")
print(f" Bounds: {lower:.2f} - {upper:.2f}")
print(f" Outliers: {len(outliers)} found")
if len(outliers) > 0:
print(f" Values: {outliers[col].tolist()}")
Example Output:
Age:
Bounds: 17.50 - 50.50
Outliers: 2 found
Values: [150, 5]
Salary:
Bounds: 46500.00 - 58500.00
Outliers: 3 found
Values: [500000, 10000, 1000000]
Rating:
Bounds: 3.78 - 4.95
Outliers: 2 found
Values: [2.1, 9.5]
9. Complete Outlier Handling Pipeline
Here's a complete function that detects and handles outliers all at once!
def handle_outliers_complete(df, column, method='cap'):
"""
Complete outlier handling function
Parameters:
- df: DataFrame
- column: Column name to check
- method: 'cap', 'remove', or 'flag'
"""
print(f"\n=== HANDLING OUTLIERS IN {column} ===")
# Calculate IQR
Q1 = df[column].quantile(0.25)
Q3 = df[column].quantile(0.75)
IQR = Q3 - Q1
lower_bound = Q1 - 1.5 * IQR
upper_bound = Q3 + 1.5 * IQR
# Identify outliers
outliers = df[(df[column] < lower_bound) | (df[column] > upper_bound)]
print(f"Outliers found: {len(outliers)}")
if len(outliers) > 0:
print(f"Outlier values: {outliers[column].tolist()}")
# Handle based on method
if method == 'cap':
df[f'{column}_Clean'] = df[column].clip(lower=lower_bound, upper=upper_bound)
print(f" Outliers capped at {lower_bound:.2f} - {upper_bound:.2f}")
elif method == 'remove':
df_clean = df[(df[column] >= lower_bound) & (df[column] <= upper_bound)]
print(f" Removed {len(outliers)} outlier rows")
return df_clean
elif method == 'flag':
df[f'{column}_Outlier'] = (df[column] < lower_bound) | (df[column] > upper_bound)
print(f" Flagged {len(outliers)} outliers")
# Show statistics
print(f"\nOriginal range: {df[column].min():.2f} - {df[column].max():.2f}")
if method == 'cap':
print(f"Cleaned range: {df[f'{column}_Clean'].min():.2f} - {df[f'{column}_Clean'].max():.2f}")
return df
# Use the function
df_result = handle_outliers_complete(df, 'Salary', method='cap')
print(df_result[['Name', 'Salary', 'Salary_Clean']].head(10))
10. Choosing the Right Method
When to use each method:
- Remove Outliers: When they're clearly data entry errors and you have plenty of data
- Cap/Winsorize: For machine learning when you want to keep all records but limit extreme values
- Flag Outliers: When you need to analyze outliers separately (fraud detection, VIP customers)
- Transform Data: When outliers are natural (income, website traffic) and you need to normalize distribution
- Keep Outliers: When they represent important information (high-value customers, rare events)
Quick Reference Summary
# Detect outliers using IQR
Q1 = df['column'].quantile(0.25)
Q3 = df['column'].quantile(0.75)
IQR = Q3 - Q1
lower = Q1 - 1.5 * IQR
upper = Q3 + 1.5 * IQR
outliers = df[(df['column'] < lower) | (df['column'] > upper)]
# Method 1: Remove outliers
df_clean = df[(df['column'] >= lower) & (df['column'] <= upper)]
# Method 2: Cap outliers
df['column_capped'] = df['column'].clip(lower=lower, upper=upper)
# Method 3: Flag outliers
df['is_outlier'] = (df['column'] < lower) | (df['column'] > upper)
# Method 4: Transform data
df['column_log'] = np.log1p(df['column'])
# Detect using Z-score
from scipy import stats
df['zscore'] = stats.zscore(df['column'])
outliers = df[abs(df['zscore']) > 3]
Remember:
- Always visualize your data first (box plots, histograms)
- Understand WHY outliers exist before removing them
- Choose the handling method based on your specific use case
- Document your outlier handling decisions
- Compare statistics before and after handling outliers
Happy Learning !! 🐼
Comments
Post a Comment