Welcome back, data explorer! In our last lesson, we learned about mean (average). Today, we're learning about its cousin: median. Think of median as the "middle value" — and sometimes it tells you MORE truth than the average! Let's discover why using real COVID-19 data.
📚 Source & License
Code examples are adapted from Exploratory Data Analysis with Python Cookbook by Packt Publishing.
GitHub repository: https://github.com/PacktPublishing/Exploratory-Data-Analysis-with-Python-Cookbook
Licensed under the MIT License.
What is Median? (The Middle Value!)
Imagine you and your friends are comparing heights:
- Person 1: 150 cm
- Person 2: 160 cm
- Person 3: 165 cm ← This is the median!
- Person 4: 170 cm
- Person 5: 180 cm
If you line everyone up from shortest to tallest, the middle person is the median. That's it! Simple, right?
Step-by-Step: How to Find Median
Step 1: Arrange numbers in order (smallest to largest)
Example: 10, 50, 100, 200, 500
Step 2: Find the middle number
10, 50, 100, 200, 500 ← Median is 100
What if you have an even number of values?
10, 50, 100, 200, 500, 1000
Take the two middle numbers (100 and 200) and find their average: (100 + 200) ÷ 2 = 150
Mean vs Median: Why Does This Matter?
Here's where it gets interesting! Let's say you have COVID cases in 5 countries:
Country A: 100 cases
Country B: 150 cases
Country C: 200 cases
Country D: 250 cases
Country E: 10,000,000 cases ← Super high!
Mean (average): (100 + 150 + 200 + 250 + 10,000,000) ÷ 5 = 2,000,140 cases
Median (middle): 100, 150, 200, 250, 10,000,000 = 200 cases
See the huge difference?!
- Mean says average is 2 million cases — but that's misleading!
- Median says typical country has 200 cases — much more accurate!
💡 Key Lesson:
Use MEDIAN when you have extreme values (outliers)! Median isn't fooled by super high or super low numbers. It shows you what's "typical" for most of your data.
About Our Real COVID-19 Dataset
Remember our dataset? We have:
- 5,818 rows — Different countries and dates
- 67 columns — Cases, deaths, vaccinations, and more
- Real-world data — With all its messy quirks!
This is perfect for learning median because COVID data has LOTS of extreme values! Some countries had millions of cases while others had just hundreds.
Step 1: Load the COVID-19 Dataset
import pandas as pd
import numpy as np
# Load the COVID-19 dataset
df = pd.read_csv('coviddata.csv')
# Quick peek at the data
print(df.head())
print(f"\nDataset size: {df.shape[0]} rows, {df.shape[1]} columns")
Method 1: Calculate Simple Median 🎯
Let's find the median of total COVID cases. This is as easy as calling .median()!
# Calculate median of total_cases
median_cases = df['total_cases'].median()
print(f"Median total cases: {median_cases:,.0f}")
# Let's compare with mean to see the difference!
mean_cases = df['total_cases'].mean()
print(f"Mean total cases: {mean_cases:,.0f}")
print(f"\nDifference: {abs(mean_cases - median_cases):,.0f}")
Sample Output:
Median total cases: 45,670
Mean total cases: 1,234,567
Difference: 1,188,897
Look at that difference! The mean is WAY higher than the median. This tells us:
- Most countries have around 45,000 cases (median)
- But a few countries with MILLIONS of cases are pulling the average way up!
- Median gives us a better picture of what's "typical"
Method 2: Median for Multiple Columns
Let's calculate median for several important columns at once:
# Calculate median for multiple columns
important_cols = ['total_cases', 'total_deaths', 'new_cases', 'total_vaccinations']
medians = df[important_cols].median()
print("Median values:")
print(medians)
# Let's also compare with means
print("\nMean values:")
means = df[important_cols].mean()
print(means)
# Show the comparison
print("\nMedian vs Mean Comparison:")
comparison = pd.DataFrame({
'Median': medians,
'Mean': means,
'Difference': means - medians
})
print(comparison)
Sample Output:
Median vs Mean Comparison:
Median Mean Difference
total_cases 45670.0 1234567.0 1188897.0
total_deaths 1234.0 23456.0 22222.0
new_cases 125.0 5432.0 5307.0
total_vaccinations 876543.0 5678901.0 4802358.0
What this tells us: For ALL these columns, mean is much higher than median! This confirms our data has many extreme high values (outliers). The median is more trustworthy!
Method 3: Median by Continent
Now let's get really interesting — find the median cases for each continent:
# Calculate median total_cases by continent
continent_median = df.groupby('continent')['total_cases'].median()
print("Median cases by continent:")
print(continent_median.sort_values(ascending=False))
# Compare with mean
continent_mean = df.groupby('continent')['total_cases'].mean()
# Side by side comparison
comparison_df = pd.DataFrame({
'Median': continent_median,
'Mean': continent_mean
}).sort_values('Median', ascending=False)
print("\nMedian vs Mean by Continent:")
print(comparison_df)
Sample Output:
Median vs Mean by Continent:
Median Mean
Europe 789012.0 2345678.0
North America 456789.0 1876543.0
South America 234567.0 1234567.0
Asia 123456.0 987654.0
Oceania 45678.0 234567.0
Africa 23456.0 456789.0
Insights:
- Europe has the highest median cases — this is more reliable than mean!
- For every continent, mean > median, showing we have high outliers
- Africa's median is much lower than mean — suggesting a few countries with very high cases
Method 4: Median by Specific Countries
Let's analyze specific countries and see the median vs mean difference:
# Select countries of interest
countries = ['United States', 'India', 'Brazil', 'United Kingdom', 'Germany', 'France']
# Filter data for these countries
country_data = df[df['location'].isin(countries)]
# Calculate median and mean for each country
country_stats = country_data.groupby('location')['total_cases'].agg(['median', 'mean', 'min', 'max'])
country_stats.columns = ['Median', 'Mean', 'Minimum', 'Maximum']
# Sort by median
country_stats = country_stats.sort_values('Median', ascending=False)
print(country_stats)
Sample Output:
Median Mean Minimum Maximum
United States 45000000.0 50000000.0 0.0 95000000.0
India 32000000.0 35000000.0 0.0 45000000.0
Brazil 18000000.0 25000000.0 0.0 38000000.0
France 12000000.0 18000000.0 0.0 40000000.0
United Kingdom 11000000.0 15000000.0 0.0 24000000.0
Germany 8000000.0 12000000.0 0.0 38000000.0
What we learn:
- The median shows the "middle point" of each country's journey through the pandemic
- All countries started at 0 (minimum) and grew over time
- Mean is higher because later pandemic periods had more cumulative cases
Method 5: Using .describe() for Complete Picture
Remember .describe()? It shows median (as 50%) along with other statistics:
# Get comprehensive statistics
print(df['total_cases'].describe())
Output:
count 5.818000e+03
mean 1.234568e+06 ← Average (can be misleading!)
std 3.456789e+06
min 0.000000e+00
25% 1.234000e+03 ← 25% of data is below this
50% 4.567000e+04 ← MEDIAN! (middle value)
75% 2.345000e+05 ← 75% of data is below this
max 8.900000e+07 ← Highest value (outlier!)
Reading this output:
- 50% (median): Half of all records have fewer than 45,670 cases
- mean: Average is 1.2 million — much higher!
- max: One record has 89 million cases — this is pulling the mean up!
- 25% and 75%: These show the range where most data falls
Pro Tip:
The 50% row in .describe() is your median! You don't always need to call .median() separately.
Method 6: Handling Missing Values (NaN) 🚨
Just like with mean, .median() automatically skips missing values (NaN). Let's verify:
# Check for missing values
total_rows = len(df)
missing_cases = df['total_cases'].isna().sum()
available_cases = df['total_cases'].notna().sum()
print(f"Total rows: {total_rows:,}")
print(f"Missing values: {missing_cases:,}")
print(f"Available values: {available_cases:,}")
print(f"Missing percentage: {(missing_cases/total_rows)*100:.2f}%")
# Calculate median (automatically ignores NaN)
median_value = df['total_cases'].median()
print(f"\nMedian calculated from {available_cases:,} values: {median_value:,.0f}")
Important: Pandas is smart! It ignores NaN values automatically. But you should still check what percentage is missing — if 90% is missing, your median might not be meaningful!
Method 7: Median Over Time (Advanced!)
Let's see how the median changed over time during the pandemic:
# Convert date to datetime
df['date'] = pd.to_datetime(df['date'])
# Extract year and month
df['year_month'] = df['date'].dt.to_period('M')
# Calculate median new_cases by month
monthly_median = df.groupby('year_month')['new_cases'].median()
print("Median new cases by month:")
print(monthly_median.head(10))
# Find which month had highest median
highest_month = monthly_median.idxmax()
highest_value = monthly_median.max()
print(f"\nHighest median new cases: {highest_value:,.0f} in {highest_month}")
# Find which month had lowest median (excluding zeros)
monthly_median_nonzero = monthly_median[monthly_median > 0]
lowest_month = monthly_median_nonzero.idxmin()
lowest_value = monthly_median_nonzero.min()
print(f"Lowest median new cases: {lowest_value:,.0f} in {lowest_month}")
Why this is useful: Tracking median over time shows you the "typical" progression of the pandemic without being skewed by countries with extreme outbreaks!
Real-World Analysis: Complete EDA with Median
Let's combine everything for a comprehensive analysis:
import pandas as pd
# Load data
df = pd.read_csv('coviddata.csv')
print("="*70)
print("COVID-19 MEDIAN ANALYSIS - FINDING THE 'TYPICAL' STORY")
print("="*70)
# 1. Overall Statistics
print("\n1. OVERALL STATISTICS (Total Cases):")
print(f" Median (typical): {df['total_cases'].median():,.0f}")
print(f" Mean (average): {df['total_cases'].mean():,.0f}")
print(f" The mean is {(df['total_cases'].mean() / df['total_cases'].median()):.1f}x higher!")
print(f" → This shows we have extreme outliers!")
# 2. By Continent
print("\n2. MEDIAN CASES BY CONTINENT:")
continent_stats = df.groupby('continent')['total_cases'].agg(['median', 'mean'])
continent_stats['mean/median_ratio'] = continent_stats['mean'] / continent_stats['median']
continent_stats = continent_stats.sort_values('median', ascending=False)
print(continent_stats)
# 3. Deaths Analysis
print("\n3. DEATHS ANALYSIS:")
print(f" Median total deaths: {df['total_deaths'].median():,.0f}")
print(f" Mean total deaths: {df['total_deaths'].mean():,.0f}")
death_by_continent = df.groupby('continent')['total_deaths'].median().sort_values(ascending=False)
print(f" Highest median deaths (continent): {death_by_continent.idxmax()}")
# 4. Vaccination Progress
print("\n4. VACCINATION STATISTICS:")
print(f" Median vaccinations: {df['total_vaccinations'].median():,.0f}")
vacc_continent = df.groupby('continent')['total_vaccinations'].median()
print(f" Best vaccinated (median): {vacc_continent.idxmax()}")
# 5. New Cases Distribution
print("\n5. NEW CASES DISTRIBUTION:")
new_cases_stats = df['new_cases'].describe()
print(f" 25% of days had: ≤ {new_cases_stats['25%']:,.0f} new cases")
print(f" 50% of days had: ≤ {new_cases_stats['50%']:,.0f} new cases (median)")
print(f" 75% of days had: ≤ {new_cases_stats['75%']:,.0f} new cases")
print(f" Worst day had: {new_cases_stats['max']:,.0f} new cases")
# 6. Top Countries by Median
print("\n6. TOP 5 COUNTRIES (by median total cases):")
top_countries_median = df.groupby('location')['total_cases'].median().nlargest(5)
for i, (country, cases) in enumerate(top_countries_median.items(), 1):
print(f" {i}. {country}: {cases:,.0f}")
print("\n" + "="*70)
print("KEY INSIGHT: Median shows the 'typical' experience better than mean!")
print("="*70)
When to Use Median vs Mean?
Use MEDIAN when:
- You have extreme values (outliers) — like income, house prices, COVID cases
- You want to know what's "typical" for most people/countries
- Your data is skewed (not evenly distributed)
- You're dealing with counts that can get very large
Use MEAN when:
- Your data is evenly distributed (no extreme outliers)
- You want to consider ALL values equally
- You're doing further mathematical calculations
- Every data point is equally important
Best Practice: Use BOTH!
Always calculate both median and mean. If they're very different, you have outliers! This difference itself is valuable information.
Quick Reference Cheat Sheet
# Basic median calculation
df['total_cases'].median() # Simple median
# Multiple columns
df[['total_cases', 'total_deaths']].median() # Multiple medians
# Grouped median (most useful!)
df.groupby('continent')['total_cases'].median() # By continent
df.groupby('location')['total_cases'].median() # By country
# Median with other stats
df.groupby('continent')['total_cases'].agg(['median', 'mean', 'max'])
# Get median from describe()
df['total_cases'].describe() # Check the 50% row!
# Compare median vs mean
comparison = pd.DataFrame({
'median': df.groupby('continent')['total_cases'].median(),
'mean': df.groupby('continent')['total_cases'].mean()
})
Common Mistakes to Avoid ⚠️
- Don't use only mean! Always check median too, especially with real-world data
- Don't ignore the difference! If mean >> median, you have outliers — this is important info!
- Don't forget to check sample size! Median of 3 values isn't very meaningful
- Don't compare medians from different sample sizes! Make sure you're comparing apples to apples
- Don't assume median = typical! If you have bimodal data (two peaks), median might fall between both peaks
Practice Exercises 📝
Test your understanding! Try these challenges:
import pandas as pd
df = pd.read_csv('coviddata.csv')
# CHALLENGE 1: Find median population by continent
# Which continent has the highest median population?
continent_pop_median = df.groupby('continent')['population'].median()
print("Challenge 1:", continent_pop_median.idxmax())
# CHALLENGE 2: Calculate the ratio of mean/median for new_deaths
# If ratio > 2, we have strong outliers!
mean_deaths = df['new_deaths'].mean()
median_deaths = df['new_deaths'].median()
ratio = mean_deaths / median_deaths
print(f"Challenge 2: Ratio = {ratio:.2f}")
# CHALLENGE 3: Find countries where median total_cases > 1 million
country_median = df.groupby('location')['total_cases'].median()
high_median_countries = country_median[country_median > 1000000]
print("Challenge 3:", high_median_countries)
# CHALLENGE 4: Compare median of first 1000 rows vs last 1000 rows
# This shows how the pandemic evolved!
first_median = df.head(1000)['new_cases'].median()
last_median = df.tail(1000)['new_cases'].median()
print(f"Challenge 4: First={first_median:.0f}, Last={last_median:.0f}")
print(f"Change: {((last_median - first_median) / first_median * 100):.1f}%")
Key Takeaways 🎓
- Median = middle value when data is sorted. It's not fooled by outliers!
- When mean >> median, you have high outliers pulling the average up
- Median shows what's "typical" better than mean for skewed data
- COVID data is PERFECT for median because countries vary so much
- Always use BOTH median and mean — their difference tells a story!
- Group your data (by continent, country) for meaningful insights
- 50% in .describe() = median — you've been seeing it all along!
Real-World Impact 🌍
Understanding median vs mean has real-world implications:
- Policy decisions: Governments should look at median cases to understand typical countries, not just average
- Resource allocation: Median helps identify where most countries/regions need help
- Trend analysis: Median shows if the "typical" situation is improving or worsening
- Fair comparisons: Comparing medians is fairer than means when countries vary wildly in size
Happy Learning ! 🐼
Comments
Post a Comment