We've learned about mean (average) and median (middle value). Now it's time to meet their fun cousin: mode — the "most popular" value! Think of it as finding what happens MOST often in your data. Let's explore this with real COVID-19 data!
📚 Source & License
Code examples are adapted from Exploratory Data Analysis with Python Cookbook by Packt Publishing.
GitHub repository: https://github.com/PacktPublishing/Exploratory-Data-Analysis-with-Python-Cookbook
Licensed under the MIT License.
What is Mode? (The Most Popular Value!)
Imagine you're in a classroom asking everyone their favorite color:
Student 1: Blue
Student 2: Red
Student 3: Blue
Student 4: Green
Student 5: Blue
Student 6: Red
Student 7: Blue
Which color appears MOST often? Blue! It appears 4 times. That's the mode!
📖 Simple Definition:
Mode = The value that appears most frequently in your data.
It's like finding the "winner" — what occurs more than anything else!
Mode vs Mean vs Median: The Complete Picture!
Let's see all three together with a simple example of COVID cases in 10 countries:
Cases: 100, 150, 200, 200, 200, 250, 300, 400, 500, 10000
- Mean (average): (100 + 150 + 200 + 200 + 200 + 250 + 300 + 400 + 500 + 10000) ÷ 10 = 1,230 cases
- Median (middle): Arrange in order, find middle = 225 cases (average of 200 and 250)
- Mode (most common): 200 cases (appears 3 times, more than any other value!)
💡 Why Mode is Special:
- Works with text! You can find the most common continent, country name, or category
- Shows frequency — what value appears most often
- Perfect for categories — like "which test type is most common?"
- Multiple modes possible! Unlike mean and median, you can have 2+ modes
When is Mode Super Useful?
Real-world scenarios where mode shines:
- Categorical data: Most common continent, most frequent test type
- Discrete counts: Most common number of new cases per day
- Customer behavior: Most popular product, most common purchase time
- Medical data: Most common blood type, most frequent symptom
- Survey responses: Most selected answer, most common rating
For COVID data: Mode helps us find things like "Which continent appears most in our data?" or "What's the most common number of daily deaths?"
Important: Mode Can Have Quirks! ⚠️
Before we start coding, understand these special cases:
Case 1: No Mode (All values appear equally)
Data: 10, 20, 30, 40, 50 — each appears once. Technically ALL are modes, but that's not helpful!
Case 2: Multiple Modes (Tie!)
Data: 100, 100, 100, 200, 200, 200, 300 — Both 100 and 200 appear 3 times. This is called bimodal (two modes)!
Case 3: Clear Winner
Data: 100, 200, 200, 200, 300, 400 — 200 appears 3 times, others once. Clear mode = 200!
Pandas handles all these cases for us! Let's see how.
Step 1: Load the COVID-19 Dataset
import pandas as pd
import numpy as np
# Load the COVID-19 dataset
df = pd.read_csv('coviddata.csv')
# Quick overview
print(f"Dataset size: {df.shape[0]} rows, {df.shape[1]} columns")
print(df.head())
Method 1: Calculate Mode - The Basic Way
Let's find the most common (modal) value for new cases:
# Calculate mode of new_cases
mode_value = df['new_cases'].mode()
print("Mode of new_cases:")
print(mode_value)
Output:
0 0.0
dtype: float64
Wait, what?! The mode is 0? Yes! This makes sense because:
- Many days early in the pandemic had ZERO new cases
- Many countries had zero cases on many days
- 0 appears more frequently than any other single number!
💡 Important Note:
.mode() returns a Series, not a single value! If there are multiple modes, it returns all of them. To get just the first mode value, use .mode()[0]
# Get just the first mode value
mode_single = df['new_cases'].mode()[0]
print(f"Most common value for new_cases: {mode_single}")
# Check how many times it appears
mode_count = (df['new_cases'] == mode_single).sum()
print(f"This value appears {mode_count} times")
# What percentage of the data is this?
percentage = (mode_count / len(df)) * 100
print(f"That's {percentage:.2f}% of all records")
Sample Output:
Most common value for new_cases: 0.0
This value appears 1234 times
That's 21.21% of all records
Method 2: Mode for Categorical Data (Text Columns)
This is where mode really shines! Let's find the most common continent in our dataset:
# Find most common continent
most_common_continent = df['continent'].mode()[0]
print(f"Most common continent in dataset: {most_common_continent}")
# How many times does it appear?
continent_count = (df['continent'] == most_common_continent).sum()
print(f"Appears {continent_count} times")
# Show frequency of all continents
print("\nFrequency of all continents:")
print(df['continent'].value_counts())
Sample Output:
Most common continent in dataset: Europe
Appears 1876 times
Frequency of all continents:
Europe 1876
Asia 1543
Africa 1234
North America 987
South America 876
Oceania 302
Name: continent, dtype: int64
Insight: Europe appears most in our dataset! This tells us our data has more European country records than others.
Method 3: Finding Mode for Each Group
Let's find the most common number of new cases for EACH continent:
# Mode of new_cases by continent
# Note: .mode() with groupby can be tricky, so we'll use a custom approach
def get_mode(series):
mode_result = series.mode()
if len(mode_result) > 0:
return mode_result[0]
else:
return None
continent_modes = df.groupby('continent')['new_cases'].apply(get_mode)
print("Most common new_cases value by continent:")
print(continent_modes)
Sample Output:
Most common new_cases value by continent:
Africa 0.0
Asia 0.0
Europe 0.0
North America 0.0
Oceania 0.0
South America 0.0
Name: new_cases, dtype: float64
Why all zeros? Because for every continent, the most frequently occurring value for new_cases is 0 (many days with no new cases).
Method 4: Mode for Non-Zero Values (More Meaningful!) 💡
Let's exclude zeros to find more interesting patterns:
# Filter out zeros
df_nonzero = df[df['new_cases'] > 0]
# Find mode of non-zero new_cases
mode_nonzero = df_nonzero['new_cases'].mode()[0]
print(f"Most common NON-ZERO new_cases: {mode_nonzero}")
# By continent (excluding zeros)
print("\nMost common non-zero new_cases by continent:")
continent_modes_nonzero = df_nonzero.groupby('continent')['new_cases'].apply(get_mode)
print(continent_modes_nonzero)
Sample Output:
Most common NON-ZERO new_cases: 1.0
Most common non-zero new_cases by continent:
Africa 1.0
Asia 1.0
Europe 2.0
North America 3.0
Oceania 1.0
South America 1.0
Name: new_cases, dtype: float64
More interesting! When we exclude zeros, we see small numbers (1-3) are most common — meaning most days with cases had just a few new cases.
Method 5: Finding Most Common Country 🗺
Which country appears most frequently in our dataset?
# Most common location (country)
most_common_country = df['location'].mode()[0]
print(f"Most common country: {most_common_country}")
# Top 10 most frequent countries
print("\nTop 10 countries by frequency:")
top_countries = df['location'].value_counts().head(10)
print(top_countries)
Sample Output:
Most common country: United States
Top 10 countries by frequency:
United States 365
India 365
Brazil 365
United Kingdom 365
Germany 365
France 365
Italy 365
Spain 365
Canada 365
Australia 365
Name: location, dtype: int64
Interesting! All have 365 records — that's one year of daily data! This means our dataset is balanced with equal representation.
Method 6: Detecting Multiple Modes (Bimodal/Multimodal)
Let's check if we have multiple modes (ties for most common):
# Get all modes for new_deaths
all_modes = df['new_deaths'].mode()
print(f"Number of modes: {len(all_modes)}")
print("All mode values:")
print(all_modes)
if len(all_modes) > 1:
print("\nThis data is MULTIMODAL (multiple values appear equally often)")
print("The most common values are:")
for mode in all_modes:
count = (df['new_deaths'] == mode).sum()
print(f" {mode}: appears {count} times")
else:
print(f"\nThis data is UNIMODAL (single mode: {all_modes[0]})")
Sample Output:
Number of modes: 1
All mode values:
0 0.0
dtype: float64
This data is UNIMODAL (single mode: 0.0)
Method 7: Mode vs Value_Counts (Understanding Frequency)
Mode just tells you the "winner." But value_counts() shows you the FULL picture:
# Mode: Just the most common
mode = df['continent'].mode()[0]
print(f"Mode: {mode}")
# Value_counts: See ALL frequencies
print("\nComplete frequency distribution:")
frequencies = df['continent'].value_counts()
print(frequencies)
# Percentage distribution
print("\nAs percentages:")
percentages = df['continent'].value_counts(normalize=True) * 100
print(percentages.round(2))
Sample Output:
Mode: Europe
Complete frequency distribution:
Europe 1876
Asia 1543
Africa 1234
North America 987
South America 876
Oceania 302
Name: continent, dtype: int64
As percentages:
Europe 32.25
Asia 26.53
Africa 21.21
North America 16.97
South America 15.06
Oceania 5.19
Name: continent, dtype: float64
💡 Pro Tip:
Use .mode() when you just need the most common value. Use .value_counts() when you want to see the frequency of ALL values. Often, .value_counts() is more informative!
Method 8: Practical Example - ISO Codes (Real Analysis!)
Let's find the most common ISO code and understand what it means:
# Most common ISO code
most_common_iso = df['iso_code'].mode()[0]
print(f"Most common ISO code: {most_common_iso}")
# Find the country name for this code
country_name = df[df['iso_code'] == most_common_iso]['location'].iloc[0]
print(f"This is: {country_name}")
# How many times does it appear?
iso_count = (df['iso_code'] == most_common_iso).sum()
print(f"Appears {iso_count} times in the dataset")
# Show top 10 ISO codes
print("\nTop 10 most frequent ISO codes:")
print(df['iso_code'].value_counts().head(10))
Method 9: Mode for Binned/Rounded Data
Sometimes finding mode on exact numbers isn't useful. Let's bin the data first:
# Create bins for total_cases
# Let's categorize as: Low, Medium, High, Very High
def categorize_cases(cases):
if pd.isna(cases):
return 'Unknown'
elif cases == 0:
return 'Zero'
elif cases < 1000:
return 'Low'
elif cases < 100000:
return 'Medium'
elif cases < 1000000:
return 'High'
else:
return 'Very High'
df['case_category'] = df['total_cases'].apply(categorize_cases)
# Find most common category
most_common_category = df['case_category'].mode()[0]
print(f"Most common case category: {most_common_category}")
# Show distribution
print("\nDistribution of case categories:")
print(df['case_category'].value_counts())
# As percentages
print("\nAs percentages:")
print(df['case_category'].value_counts(normalize=True).mul(100).round(2))
Sample Output:
Most common case category: Low
Distribution of case categories:
Low 2543
Medium 1876
Very High 876
High 321
Zero 202
Unknown 0
Name: case_category, dtype: int64
As percentages:
Low 43.71
Medium 32.25
Very High 15.06
High 5.52
Zero 3.47
Unknown 0.00
Name: case_category, dtype: float64
Insight: Most records fall in the "Low" category (under 1,000 cases), showing that for most countries on most days, case counts were relatively low!
Real-World Analysis: Complete Mode Analysis
Let's put everything together:
import pandas as pd
# Load data
df = pd.read_csv('coviddata.csv')
print("="*70)
print("COVID-19 MODE ANALYSIS - FINDING WHAT'S MOST COMMON")
print("="*70)
# 1. Most Common Continent
print("\n1. GEOGRAPHIC DISTRIBUTION:")
most_common_continent = df['continent'].mode()[0]
continent_pct = (df['continent'] == most_common_continent).sum() / len(df) * 100
print(f" Most common continent: {most_common_continent} ({continent_pct:.1f}% of data)")
# 2. Most Common Country
print("\n2. COUNTRY REPRESENTATION:")
most_common_country = df['location'].mode()[0]
country_count = (df['location'] == most_common_country).sum()
print(f" Most common country: {most_common_country}")
print(f" Appears {country_count} times")
# 3. Most Common Values (with context)
print("\n3. MOST FREQUENT VALUES:")
print(f" New cases mode: {df['new_cases'].mode()[0]}")
print(f" New deaths mode: {df['new_deaths'].mode()[0]}")
print(f" (Note: 0 is most common because many days had no new cases/deaths)")
# 4. Non-zero analysis
df_active = df[(df['new_cases'] > 0)]
if len(df_active) > 0:
print(f"\n When excluding zeros:")
print(f" Most common new_cases: {df_active['new_cases'].mode()[0]}")
# 5. Check for multimodal data
print("\n4. MULTIMODAL CHECK:")
modes_count = len(df['new_deaths'].mode())
if modes_count > 1:
print(f" new_deaths has {modes_count} modes (multimodal)")
else:
print(f" new_deaths has 1 mode (unimodal)")
# 6. Categorical Analysis
print("\n5. TOP 5 MOST COMMON COUNTRIES:")
top5 = df['location'].value_counts().head(5)
for i, (country, count) in enumerate(top5.items(), 1):
print(f" {i}. {country}: {count} records")
# 7. Test Types (if column exists)
if 'tests_units' in df.columns:
print("\n6. TESTING METHODS:")
most_common_test = df['tests_units'].mode()
if len(most_common_test) > 0:
print(f" Most common test type: {most_common_test[0]}")
print(f"\n All test types:")
print(df['tests_units'].value_counts())
print("\n" + "="*70)
print("KEY INSIGHT: Mode shows us what occurs most frequently!")
print("="*70)
When to Use Mode vs Mean vs Median? 🤔
Use MODE when:
- You have categorical data (countries, continents, categories)
- You want to know the most common value
- You're working with discrete counts that repeat often
- You want to find typical patterns in repeated data
- You're analyzing survey responses or choices
Use MEAN when:
- You have numerical data without extreme outliers
- You want to find the average
- All values should be weighted equally
Use MEDIAN when:
- You have numerical data with outliers
- You want the middle value
- You want to know what's typical (not skewed by extremes)
Quick Reference Cheat Sheet
# Basic mode
df['column'].mode() # Returns Series with all modes
df['column'].mode()[0] # Get first mode value
# Check if multiple modes exist
len(df['column'].mode()) # If > 1, multimodal
# Mode with groupby (custom function needed)
def get_mode(series):
mode = series.mode()
return mode[0] if len(mode) > 0 else None
df.groupby('group')['column'].apply(get_mode)
# Value counts (often more useful than mode!)
df['column'].value_counts() # Frequency of all values
df['column'].value_counts().head(1) # Most common value with count
df['column'].value_counts(normalize=True) # As percentages
# Count how many times mode appears
mode_val = df['column'].mode()[0]
(df['column'] == mode_val).sum() # Count occurrences
# Mode for non-zero values
df[df['column'] > 0]['column'].mode()[0]
Common Mistakes to Avoid ⚠️
- Don't forget [0]!
.mode()returns a Series, not a single value. Use.mode()[0]to get the actual value - Don't use mode for continuous data! For numbers like 123.456789, every value might be unique, making mode useless
- Don't ignore multiple modes! Check
len(df['column'].mode())— you might have ties! - Don't use mode for unique identifiers! IDs, dates, etc. usually don't have meaningful modes
- Don't forget about value_counts()! It often gives more insight than just the mode
- Don't assume mode is meaningful! If mode appears only 2 times out of 10,000 records, it's not very "common"
Comments
Post a Comment