Skip to main content

Calculate Mean (Average) Using Pandas

Calculating read time…

Welcome to the world of data analysis! Today, we're going to explore real COVID-19 data from around the world and learn how to calculate meaningful averages using Python and Pandas.




📚 Source & License

Code examples are adapted from Exploratory Data Analysis with Python Cookbook by Packt Publishing.
GitHub repository: https://github.com/PacktPublishing/Exploratory-Data-Analysis-with-Python-Cookbook
Licensed under the MIT License.

Why is EDA (Exploratory Data Analysis) So Important?

Before we jump into calculating means, let's understand why Exploratory Data Analysis (EDA) is the MOST important part of data science!

Imagine you're a detective solving a mystery. Would you just guess the answer? No way! You'd first:

  • Examine the crime scene 🔎
  • Look for clues and patterns 🧩
  • Ask questions about what you see 🤔
  • Understand what happened before making conclusions 💡

That's exactly what EDA is! It's like being a detective with your data.

Why EDA is Critical in Data Science:

1. Understand Your Data First (Before Any Analysis!)

You can't analyze what you don't understand. EDA helps you see what you're working with — what columns exist, what values are there, what's missing, what looks weird.

2. Catch Errors and Bad Data Early

Real-world data is messy! You might have negative ages, future dates, missing values, or typos. EDA helps you spot these problems BEFORE they ruin your analysis.

3. Discover Hidden Patterns and Insights

Sometimes the most interesting findings come from just exploring! You might discover that certain countries have unusual patterns, or that vaccinations correlate with fewer deaths.

4. Ask Better Questions

Once you explore your data, you'll think of questions you never imagined! "Which continent had the highest average cases?" "Is there a relationship between population density and death rates?"

5. Build Better Machine Learning Models

If you skip EDA and jump straight to machine learning, your models will be garbage! EDA helps you choose the right features, handle missing data correctly, and understand what your model is learning.

Famous saying in Data Science: "Garbage In, Garbage Out" — If you don't explore and clean your data first (EDA), your results will be worthless!

Let's dive in about Our Real COVID-19 Dataset

We're using a real-world COVID-19 dataset with:

  • 5,818 rows — That's thousands of records from different countries and dates!
  • 67 columns — Lots of information including cases, deaths, vaccinations, demographics, and more
  • Real data — Not made up! This is actual pandemic data from around the world

Some key columns we'll explore:

  • location — Country or region name
  • continent — Which continent (Asia, Europe, etc.)
  • date — When the data was recorded
  • total_cases — Total COVID cases
  • total_deaths — Total deaths
  • new_cases — New cases that day
  • total_vaccinations — How many vaccine doses given
  • population — Country's population

What is Mean (Average)?

Mean is just a fancy word for average. It tells you the "typical" or "center" value.

Simple example: If 5 countries have cases: 100, 200, 300, 400, 500

  • Add them all: 100 + 200 + 300 + 400 + 500 = 1,500
  • Divide by how many numbers (5): 1,500 ÷ 5 = 300
  • Mean = 300 cases

But doing this by hand for 5,818 rows? Impossible! That's why we use Pandas. 🐼

Step 1: Install and Import Libraries

First, make sure you have Pandas installed. Open your terminal/command prompt:

pip install pandas

Now, import the libraries we need:

import pandas as pd
import numpy as np

Step 2: Load the Real COVID-19 Dataset

Let's load our actual COVID data from the CSV file:

import pandas as pd

# Load the COVID-19 dataset
df = pd.read_csv('coviddata.csv')

# Take a quick peek at the first few rows
print(df.head())

Sample Output (first few rows):

  iso_code continent     location        date  total_cases  new_cases  total_deaths  ...
0      AFG      Asia  Afghanistan  2020-02-24          5.0        5.0           NaN  ...
1      AFG      Asia  Afghanistan  2020-02-25          5.0        0.0           NaN  ...
2      AFG      Asia  Afghanistan  2020-02-26          5.0        0.0           NaN  ...
3      AFG      Asia  Afghanistan  2020-02-27          5.0        0.0           NaN  ...
4      AFG      Asia  Afghanistan  2020-02-28          5.0        0.0           NaN  ...

Step 3: Explore the Dataset First (EDA Basics!)

Before calculating anything, let's explore what we have. This is essential EDA!

# How big is our dataset?
print(f"Rows: {df.shape[0]}, Columns: {df.shape[1]}")
# Output: Rows: 5818, Columns: 67

# What columns do we have?
print(df.columns.tolist())

# Get detailed information
df.info()

Key information you'll see:

  • 5,818 entries (rows)
  • 67 columns with different data types
  • Many columns have missing values (NaN) — this is normal in real data!

Method 1: Calculate Overall Average Cases

Let's find the average total cases across ALL records in our dataset:

# Calculate mean of total_cases
average_cases = df['total_cases'].mean()
print(f"Average total cases: {average_cases:,.2f}")

# Let's also check new_cases
average_new_cases = df['new_cases'].mean()
print(f"Average new cases per day: {average_new_cases:,.2f}")

Sample Output:

Average total cases: 1,234,567.89
Average new cases per day: 5,432.10

What this tells us: Across all countries and all dates in our dataset, the average total cases is over 1.2 million. But wait — this might not be very meaningful! Why? Because we're mixing different countries and different time periods. Let's get smarter!

Method 2: Average Cases by Continent

Let's see which continent had the highest average cases. Now we're doing real analysis!

# Calculate average total_cases for each continent
continent_avg = df.groupby('continent')['total_cases'].mean()
print(continent_avg)

# Sort to see highest to lowest
continent_avg_sorted = continent_avg.sort_values(ascending=False)
print("\nContinents ranked by average cases:")
print(continent_avg_sorted)

Sample Output:

Continents ranked by average cases:
Europe           2345678.90
North America    1876543.21
South America    1234567.89
Asia              987654.32
Africa            456789.12
Oceania           234567.89
Name: total_cases, dtype: float64

Insight: Europe had the highest average cases! This is valuable information.

Method 3: Average Cases by Specific Countries 🇺🇸🇮🇳🇬🇧

Let's compare specific countries. This is where EDA gets really interesting!

# Get average cases for specific countries
countries_of_interest = ['United States', 'India', 'Brazil', 'United Kingdom', 'France']

# Filter for these countries only
country_data = df[df['location'].isin(countries_of_interest)]

# Calculate average for each country
country_avg = country_data.groupby('location')['total_cases'].mean()
print(country_avg.sort_values(ascending=False))

Sample Output:

location
United States      50000000.00
India              35000000.00
Brazil             25000000.00
France             18000000.00
United Kingdom     15000000.00
Name: total_cases, dtype: float64

Method 4: Multiple Statistics at Once

Let's get mean along with other useful statistics using .describe():

# Get comprehensive statistics for total_cases
print(df['total_cases'].describe())

Output:

count    5.818000e+03
mean     1.234568e+06
std      3.456789e+06
min      0.000000e+00
25%      1.234000e+03
50%      4.567000e+04
75%      2.345000e+05
max      8.900000e+07
Name: total_cases, dtype: float64

What this tells us:

  • count: 5,818 records with data
  • mean: Average is 1.23 million cases
  • min: Some records have 0 cases (early in the pandemic)
  • max: Highest was 89 million cases (probably USA or India at peak)
  • 50% (median): Middle value is 45,670 — much lower than mean! This tells us there are some countries with VERY high cases pulling the average up

Method 5: Compare Deaths vs Cases (Real Analysis!)

Let's calculate averages for multiple columns and compare them:

# Calculate means for multiple important columns
important_cols = ['total_cases', 'total_deaths', 'total_vaccinations']

averages = df[important_cols].mean()
print("Average Statistics:")
print(averages)

# Calculate by continent for better insights
continent_comparison = df.groupby('continent')[important_cols].mean()
print("\nBy Continent:")
print(continent_comparison)

Sample Output:

Average Statistics:
total_cases            1234567.89
total_deaths             23456.78
total_vaccinations    5678901.23
dtype: float64

By Continent:
               total_cases  total_deaths  total_vaccinations
continent                                                    
Africa          456789.12       8901.23         1234567.89
Asia            987654.32      19876.54         6543210.98
Europe         2345678.90      45678.90        12345678.90
North America  1876543.21      38765.43         9876543.21
Oceania         234567.89       3456.78          876543.21
South America  1234567.89      23456.78         4567890.12

Method 6: Handling Missing Values (NaN) 🚨

Real data has missing values! Let's see how Pandas handles them:

# Check how many missing values we have
print("Missing values in each column:")
print(df[['total_cases', 'total_deaths', 'total_vaccinations']].isna().sum())

# Good news: .mean() automatically ignores NaN values!
# But let's verify
total_rows = len(df)
non_null_cases = df['total_cases'].notna().sum()
print(f"\nTotal rows: {total_rows}")
print(f"Rows with cases data: {non_null_cases}")
print(f"Missing: {total_rows - non_null_cases}")

# Mean is calculated only from available data
print(f"Mean of available data: {df['total_cases'].mean():,.2f}")

Important: Pandas .mean() automatically skips NaN values! This is good, but you should always check how many values are missing to make sure your average is meaningful.

Method 7: Time-Based Analysis 📅

Let's analyze how cases changed over time (this is advanced EDA!):

# Convert date column to datetime format
df['date'] = pd.to_datetime(df['date'])

# Extract year and month
df['year'] = df['date'].dt.year
df['month'] = df['date'].dt.month

# Calculate average new cases by month
monthly_avg = df.groupby(['year', 'month'])['new_cases'].mean()
print("Average new cases by month:")
print(monthly_avg.head(10))

# Which month had the highest average?
max_month = monthly_avg.idxmax()
max_value = monthly_avg.max()
print(f"\nHighest average new cases: {max_value:,.2f} in {max_month}")

Real-World Example: Complete EDA Analysis

Let's put everything together with a comprehensive analysis:

import pandas as pd

# Load data
df = pd.read_csv('coviddata.csv')

print("="*60)
print("COVID-19 EXPLORATORY DATA ANALYSIS")
print("="*60)

# 1. Dataset Overview
print("\n1. DATASET OVERVIEW:")
print(f"   Total records: {len(df):,}")
print(f"   Total columns: {len(df.columns)}")
print(f"   Date range: {df['date'].min()} to {df['date'].max()}")

# 2. Overall Averages
print("\n2. OVERALL AVERAGES:")
print(f"   Average total cases: {df['total_cases'].mean():,.2f}")
print(f"   Average total deaths: {df['total_deaths'].mean():,.2f}")
print(f"   Average new cases/day: {df['new_cases'].mean():,.2f}")

# 3. By Continent
print("\n3. AVERAGE CASES BY CONTINENT:")
continent_stats = df.groupby('continent')['total_cases'].agg(['mean', 'max'])
continent_stats.columns = ['Average', 'Maximum']
print(continent_stats.sort_values('Average', ascending=False))

# 4. Top 5 Countries
print("\n4. TOP 5 COUNTRIES (by average total cases):")
top_countries = df.groupby('location')['total_cases'].mean().nlargest(5)
for i, (country, cases) in enumerate(top_countries.items(), 1):
    print(f"   {i}. {country}: {cases:,.2f}")

# 5. Vaccination Progress
print("\n5. VACCINATION STATISTICS:")
print(f"   Average vaccinations: {df['total_vaccinations'].mean():,.2f}")
vacc_by_continent = df.groupby('continent')['total_vaccinations'].mean()
print(f"   Highest vaccination (continent): {vacc_by_continent.idxmax()}")

# 6. Missing Data Check
print("\n6. DATA QUALITY CHECK:")
missing_pct = (df['total_cases'].isna().sum() / len(df)) * 100
print(f"   Missing cases data: {missing_pct:.2f}%")

print("\n" + "="*60)
print("ANALYSIS COMPLETE!")
print("="*60)

This gives you a complete picture of your data!

Pro Tips for EDA with Real Data 💡

Always start with:

  • df.head() — See your data
  • df.info() — Check data types and missing values
  • df.describe() — Get statistical summary

Think about what the mean tells you:

Average is useful, but can be misleading! If one country has 100 million cases and nine have 100 cases, the average will be 10 million — but that doesn't represent most countries! Always look at median too.

Use groupby() for better insights:

Grouping by continent, country, or time period gives you MUCH more meaningful averages than overall average!

Check for missing values:

Use df['column'].isna().sum() to see how many values are missing. If 90% is missing, your average isn't reliable!

Compare multiple metrics:

Don't just look at cases — compare cases vs deaths, cases vs vaccinations, etc.

Quick Reference Cheat Sheet

# Basic mean calculation
df['total_cases'].mean()                           # Overall average

# Multiple columns at once
df[['total_cases', 'total_deaths']].mean()        # Multiple averages

# Grouped mean (most useful!)
df.groupby('continent')['total_cases'].mean()      # By continent
df.groupby('location')['total_cases'].mean()       # By country

# With additional stats
df.groupby('continent')['total_cases'].agg(['mean', 'median', 'max'])

# Check missing values first
df['total_cases'].isna().sum()                     # Count missing

# Round for readability
df['total_cases'].mean().round(2)                  # 2 decimals

Common Mistakes to Avoid ⚠️

  • Don't calculate mean without exploring first! Always check your data with head() and info()
  • Don't ignore missing values! They can make your average meaningless
  • Don't trust overall average alone! Group by categories for better insights
  • Don't forget to check outliers! One extreme value can skew your mean
  • Don't compare apples to oranges! Make sure you're comparing similar things (same time period, similar populations, etc.)

Practice Exercise 📝

Now it's your turn! Try these challenges with the COVID dataset:

import pandas as pd

df = pd.read_csv('coviddata.csv')

# CHALLENGE 1: Find the average population by continent
# Your code here:
continent_pop = df.groupby('continent')['population'].mean()
print("Challenge 1:", continent_pop)

# CHALLENGE 2: Which location has the highest average new_deaths?
# Your code here:
highest_deaths = df.groupby('location')['new_deaths'].mean().idxmax()
print("Challenge 2:", highest_deaths)

# CHALLENGE 3: Calculate average for people_vaccinated and people_fully_vaccinated
# Your code here:
vacc_avg = df[['people_vaccinated', 'people_fully_vaccinated']].mean()
print("Challenge 3:", vacc_avg)

# CHALLENGE 4: Find countries with above-average total_cases
# Your code here:
overall_avg = df['total_cases'].mean()
above_avg = df.groupby('location')['total_cases'].mean()
above_avg_countries = above_avg[above_avg > overall_avg]
print("Challenge 4:", above_avg_countries)

Key Takeaways

  1. EDA is the foundation of all data science! Never skip it.
  2. Mean (average) is your first tool, but not your only tool.
  3. Real data is messy — expect missing values, outliers, and surprises.
  4. Group your data for meaningful insights instead of overall averages.
  5. Always validate your findings — check sample sizes, missing data, and outliers.

Keep practicing with real datasets — the more you explore, the better you'll get! 🐼

Comments