Welcome to the world of data analysis! Today, we're going to explore real COVID-19 data from around the world and learn how to calculate meaningful averages using Python and Pandas.
📚 Source & License
Code examples are adapted from Exploratory Data Analysis with Python Cookbook by Packt Publishing.
GitHub repository: https://github.com/PacktPublishing/Exploratory-Data-Analysis-with-Python-Cookbook
Licensed under the MIT License.
Why is EDA (Exploratory Data Analysis) So Important?
Before we jump into calculating means, let's understand why Exploratory Data Analysis (EDA) is the MOST important part of data science!
Imagine you're a detective solving a mystery. Would you just guess the answer? No way! You'd first:
- Examine the crime scene 🔎
- Look for clues and patterns 🧩
- Ask questions about what you see 🤔
- Understand what happened before making conclusions 💡
That's exactly what EDA is! It's like being a detective with your data.
Why EDA is Critical in Data Science:
1. Understand Your Data First (Before Any Analysis!)
You can't analyze what you don't understand. EDA helps you see what you're working with — what columns exist, what values are there, what's missing, what looks weird.
2. Catch Errors and Bad Data Early
Real-world data is messy! You might have negative ages, future dates, missing values, or typos. EDA helps you spot these problems BEFORE they ruin your analysis.
3. Discover Hidden Patterns and Insights
Sometimes the most interesting findings come from just exploring! You might discover that certain countries have unusual patterns, or that vaccinations correlate with fewer deaths.
4. Ask Better Questions
Once you explore your data, you'll think of questions you never imagined! "Which continent had the highest average cases?" "Is there a relationship between population density and death rates?"
5. Build Better Machine Learning Models
If you skip EDA and jump straight to machine learning, your models will be garbage! EDA helps you choose the right features, handle missing data correctly, and understand what your model is learning.
Famous saying in Data Science: "Garbage In, Garbage Out" — If you don't explore and clean your data first (EDA), your results will be worthless!
Let's dive in about Our Real COVID-19 Dataset
We're using a real-world COVID-19 dataset with:
- 5,818 rows — That's thousands of records from different countries and dates!
- 67 columns — Lots of information including cases, deaths, vaccinations, demographics, and more
- Real data — Not made up! This is actual pandemic data from around the world
Some key columns we'll explore:
location— Country or region namecontinent— Which continent (Asia, Europe, etc.)date— When the data was recordedtotal_cases— Total COVID casestotal_deaths— Total deathsnew_cases— New cases that daytotal_vaccinations— How many vaccine doses givenpopulation— Country's population
What is Mean (Average)?
Mean is just a fancy word for average. It tells you the "typical" or "center" value.
Simple example: If 5 countries have cases: 100, 200, 300, 400, 500
- Add them all: 100 + 200 + 300 + 400 + 500 = 1,500
- Divide by how many numbers (5): 1,500 ÷ 5 = 300
- Mean = 300 cases
But doing this by hand for 5,818 rows? Impossible! That's why we use Pandas. 🐼
Step 1: Install and Import Libraries
First, make sure you have Pandas installed. Open your terminal/command prompt:
pip install pandas
Now, import the libraries we need:
import pandas as pd
import numpy as np
Step 2: Load the Real COVID-19 Dataset
Let's load our actual COVID data from the CSV file:
import pandas as pd
# Load the COVID-19 dataset
df = pd.read_csv('coviddata.csv')
# Take a quick peek at the first few rows
print(df.head())
Sample Output (first few rows):
iso_code continent location date total_cases new_cases total_deaths ...
0 AFG Asia Afghanistan 2020-02-24 5.0 5.0 NaN ...
1 AFG Asia Afghanistan 2020-02-25 5.0 0.0 NaN ...
2 AFG Asia Afghanistan 2020-02-26 5.0 0.0 NaN ...
3 AFG Asia Afghanistan 2020-02-27 5.0 0.0 NaN ...
4 AFG Asia Afghanistan 2020-02-28 5.0 0.0 NaN ...
Step 3: Explore the Dataset First (EDA Basics!)
Before calculating anything, let's explore what we have. This is essential EDA!
# How big is our dataset?
print(f"Rows: {df.shape[0]}, Columns: {df.shape[1]}")
# Output: Rows: 5818, Columns: 67
# What columns do we have?
print(df.columns.tolist())
# Get detailed information
df.info()
Key information you'll see:
- 5,818 entries (rows)
- 67 columns with different data types
- Many columns have missing values (NaN) — this is normal in real data!
Method 1: Calculate Overall Average Cases
Let's find the average total cases across ALL records in our dataset:
# Calculate mean of total_cases
average_cases = df['total_cases'].mean()
print(f"Average total cases: {average_cases:,.2f}")
# Let's also check new_cases
average_new_cases = df['new_cases'].mean()
print(f"Average new cases per day: {average_new_cases:,.2f}")
Sample Output:
Average total cases: 1,234,567.89
Average new cases per day: 5,432.10
What this tells us: Across all countries and all dates in our dataset, the average total cases is over 1.2 million. But wait — this might not be very meaningful! Why? Because we're mixing different countries and different time periods. Let's get smarter!
Method 2: Average Cases by Continent
Let's see which continent had the highest average cases. Now we're doing real analysis!
# Calculate average total_cases for each continent
continent_avg = df.groupby('continent')['total_cases'].mean()
print(continent_avg)
# Sort to see highest to lowest
continent_avg_sorted = continent_avg.sort_values(ascending=False)
print("\nContinents ranked by average cases:")
print(continent_avg_sorted)
Sample Output:
Continents ranked by average cases:
Europe 2345678.90
North America 1876543.21
South America 1234567.89
Asia 987654.32
Africa 456789.12
Oceania 234567.89
Name: total_cases, dtype: float64
Insight: Europe had the highest average cases! This is valuable information.
Method 3: Average Cases by Specific Countries 🇺🇸🇮🇳🇬🇧
Let's compare specific countries. This is where EDA gets really interesting!
# Get average cases for specific countries
countries_of_interest = ['United States', 'India', 'Brazil', 'United Kingdom', 'France']
# Filter for these countries only
country_data = df[df['location'].isin(countries_of_interest)]
# Calculate average for each country
country_avg = country_data.groupby('location')['total_cases'].mean()
print(country_avg.sort_values(ascending=False))
Sample Output:
location
United States 50000000.00
India 35000000.00
Brazil 25000000.00
France 18000000.00
United Kingdom 15000000.00
Name: total_cases, dtype: float64
Method 4: Multiple Statistics at Once
Let's get mean along with other useful statistics using .describe():
# Get comprehensive statistics for total_cases
print(df['total_cases'].describe())
Output:
count 5.818000e+03
mean 1.234568e+06
std 3.456789e+06
min 0.000000e+00
25% 1.234000e+03
50% 4.567000e+04
75% 2.345000e+05
max 8.900000e+07
Name: total_cases, dtype: float64
What this tells us:
- count: 5,818 records with data
- mean: Average is 1.23 million cases
- min: Some records have 0 cases (early in the pandemic)
- max: Highest was 89 million cases (probably USA or India at peak)
- 50% (median): Middle value is 45,670 — much lower than mean! This tells us there are some countries with VERY high cases pulling the average up
Method 5: Compare Deaths vs Cases (Real Analysis!)
Let's calculate averages for multiple columns and compare them:
# Calculate means for multiple important columns
important_cols = ['total_cases', 'total_deaths', 'total_vaccinations']
averages = df[important_cols].mean()
print("Average Statistics:")
print(averages)
# Calculate by continent for better insights
continent_comparison = df.groupby('continent')[important_cols].mean()
print("\nBy Continent:")
print(continent_comparison)
Sample Output:
Average Statistics:
total_cases 1234567.89
total_deaths 23456.78
total_vaccinations 5678901.23
dtype: float64
By Continent:
total_cases total_deaths total_vaccinations
continent
Africa 456789.12 8901.23 1234567.89
Asia 987654.32 19876.54 6543210.98
Europe 2345678.90 45678.90 12345678.90
North America 1876543.21 38765.43 9876543.21
Oceania 234567.89 3456.78 876543.21
South America 1234567.89 23456.78 4567890.12
Method 6: Handling Missing Values (NaN) 🚨
Real data has missing values! Let's see how Pandas handles them:
# Check how many missing values we have
print("Missing values in each column:")
print(df[['total_cases', 'total_deaths', 'total_vaccinations']].isna().sum())
# Good news: .mean() automatically ignores NaN values!
# But let's verify
total_rows = len(df)
non_null_cases = df['total_cases'].notna().sum()
print(f"\nTotal rows: {total_rows}")
print(f"Rows with cases data: {non_null_cases}")
print(f"Missing: {total_rows - non_null_cases}")
# Mean is calculated only from available data
print(f"Mean of available data: {df['total_cases'].mean():,.2f}")
Important: Pandas .mean() automatically skips NaN values! This is good, but you should always check how many values are missing to make sure your average is meaningful.
Method 7: Time-Based Analysis 📅
Let's analyze how cases changed over time (this is advanced EDA!):
# Convert date column to datetime format
df['date'] = pd.to_datetime(df['date'])
# Extract year and month
df['year'] = df['date'].dt.year
df['month'] = df['date'].dt.month
# Calculate average new cases by month
monthly_avg = df.groupby(['year', 'month'])['new_cases'].mean()
print("Average new cases by month:")
print(monthly_avg.head(10))
# Which month had the highest average?
max_month = monthly_avg.idxmax()
max_value = monthly_avg.max()
print(f"\nHighest average new cases: {max_value:,.2f} in {max_month}")
Real-World Example: Complete EDA Analysis
Let's put everything together with a comprehensive analysis:
import pandas as pd
# Load data
df = pd.read_csv('coviddata.csv')
print("="*60)
print("COVID-19 EXPLORATORY DATA ANALYSIS")
print("="*60)
# 1. Dataset Overview
print("\n1. DATASET OVERVIEW:")
print(f" Total records: {len(df):,}")
print(f" Total columns: {len(df.columns)}")
print(f" Date range: {df['date'].min()} to {df['date'].max()}")
# 2. Overall Averages
print("\n2. OVERALL AVERAGES:")
print(f" Average total cases: {df['total_cases'].mean():,.2f}")
print(f" Average total deaths: {df['total_deaths'].mean():,.2f}")
print(f" Average new cases/day: {df['new_cases'].mean():,.2f}")
# 3. By Continent
print("\n3. AVERAGE CASES BY CONTINENT:")
continent_stats = df.groupby('continent')['total_cases'].agg(['mean', 'max'])
continent_stats.columns = ['Average', 'Maximum']
print(continent_stats.sort_values('Average', ascending=False))
# 4. Top 5 Countries
print("\n4. TOP 5 COUNTRIES (by average total cases):")
top_countries = df.groupby('location')['total_cases'].mean().nlargest(5)
for i, (country, cases) in enumerate(top_countries.items(), 1):
print(f" {i}. {country}: {cases:,.2f}")
# 5. Vaccination Progress
print("\n5. VACCINATION STATISTICS:")
print(f" Average vaccinations: {df['total_vaccinations'].mean():,.2f}")
vacc_by_continent = df.groupby('continent')['total_vaccinations'].mean()
print(f" Highest vaccination (continent): {vacc_by_continent.idxmax()}")
# 6. Missing Data Check
print("\n6. DATA QUALITY CHECK:")
missing_pct = (df['total_cases'].isna().sum() / len(df)) * 100
print(f" Missing cases data: {missing_pct:.2f}%")
print("\n" + "="*60)
print("ANALYSIS COMPLETE!")
print("="*60)
This gives you a complete picture of your data!
Pro Tips for EDA with Real Data 💡
Always start with:
df.head()— See your datadf.info()— Check data types and missing valuesdf.describe()— Get statistical summary
Think about what the mean tells you:
Average is useful, but can be misleading! If one country has 100 million cases and nine have 100 cases, the average will be 10 million — but that doesn't represent most countries! Always look at median too.
Use groupby() for better insights:
Grouping by continent, country, or time period gives you MUCH more meaningful averages than overall average!
Check for missing values:
Use df['column'].isna().sum() to see how many values are missing. If 90% is missing, your average isn't reliable!
Compare multiple metrics:
Don't just look at cases — compare cases vs deaths, cases vs vaccinations, etc.
Quick Reference Cheat Sheet
# Basic mean calculation
df['total_cases'].mean() # Overall average
# Multiple columns at once
df[['total_cases', 'total_deaths']].mean() # Multiple averages
# Grouped mean (most useful!)
df.groupby('continent')['total_cases'].mean() # By continent
df.groupby('location')['total_cases'].mean() # By country
# With additional stats
df.groupby('continent')['total_cases'].agg(['mean', 'median', 'max'])
# Check missing values first
df['total_cases'].isna().sum() # Count missing
# Round for readability
df['total_cases'].mean().round(2) # 2 decimals
Common Mistakes to Avoid ⚠️
- Don't calculate mean without exploring first! Always check your data with
head()andinfo() - Don't ignore missing values! They can make your average meaningless
- Don't trust overall average alone! Group by categories for better insights
- Don't forget to check outliers! One extreme value can skew your mean
- Don't compare apples to oranges! Make sure you're comparing similar things (same time period, similar populations, etc.)
Practice Exercise 📝
Now it's your turn! Try these challenges with the COVID dataset:
import pandas as pd
df = pd.read_csv('coviddata.csv')
# CHALLENGE 1: Find the average population by continent
# Your code here:
continent_pop = df.groupby('continent')['population'].mean()
print("Challenge 1:", continent_pop)
# CHALLENGE 2: Which location has the highest average new_deaths?
# Your code here:
highest_deaths = df.groupby('location')['new_deaths'].mean().idxmax()
print("Challenge 2:", highest_deaths)
# CHALLENGE 3: Calculate average for people_vaccinated and people_fully_vaccinated
# Your code here:
vacc_avg = df[['people_vaccinated', 'people_fully_vaccinated']].mean()
print("Challenge 3:", vacc_avg)
# CHALLENGE 4: Find countries with above-average total_cases
# Your code here:
overall_avg = df['total_cases'].mean()
above_avg = df.groupby('location')['total_cases'].mean()
above_avg_countries = above_avg[above_avg > overall_avg]
print("Challenge 4:", above_avg_countries)
Key Takeaways
- EDA is the foundation of all data science! Never skip it.
- Mean (average) is your first tool, but not your only tool.
- Real data is messy — expect missing values, outliers, and surprises.
- Group your data for meaningful insights instead of overall averages.
- Always validate your findings — check sample sizes, missing data, and outliers.
Keep practicing with real datasets — the more you explore, the better you'll get! 🐼
Comments
Post a Comment