Skip to main content

Calculate Mode Using Pandas

Calculating read time…

We've learned about mean (average) and median (middle value). Now it's time to meet their fun cousin: mode — the "most popular" value! Think of it as finding what happens MOST often in your data. Let's explore this with real COVID-19 data!

📚 Source & License

Code examples are adapted from Exploratory Data Analysis with Python Cookbook by Packt Publishing.
GitHub repository: https://github.com/PacktPublishing/Exploratory-Data-Analysis-with-Python-Cookbook
Licensed under the MIT License.

What is Mode? (The Most Popular Value!)

Imagine you're in a classroom asking everyone their favorite color:

Student 1: Blue
Student 2: Red
Student 3: Blue
Student 4: Green
Student 5: Blue
Student 6: Red
Student 7: Blue

Which color appears MOST often? Blue! It appears 4 times. That's the mode!

📖 Simple Definition:

Mode = The value that appears most frequently in your data.

It's like finding the "winner" — what occurs more than anything else!

Mode vs Mean vs Median: The Complete Picture!

Let's see all three together with a simple example of COVID cases in 10 countries:

Cases: 100, 150, 200, 200, 200, 250, 300, 400, 500, 10000
  • Mean (average): (100 + 150 + 200 + 200 + 200 + 250 + 300 + 400 + 500 + 10000) ÷ 10 = 1,230 cases
  • Median (middle): Arrange in order, find middle = 225 cases (average of 200 and 250)
  • Mode (most common): 200 cases (appears 3 times, more than any other value!)

💡 Why Mode is Special:

  • Works with text! You can find the most common continent, country name, or category
  • Shows frequency — what value appears most often
  • Perfect for categories — like "which test type is most common?"
  • Multiple modes possible! Unlike mean and median, you can have 2+ modes

When is Mode Super Useful?

Real-world scenarios where mode shines:

  • Categorical data: Most common continent, most frequent test type
  • Discrete counts: Most common number of new cases per day
  • Customer behavior: Most popular product, most common purchase time
  • Medical data: Most common blood type, most frequent symptom
  • Survey responses: Most selected answer, most common rating

For COVID data: Mode helps us find things like "Which continent appears most in our data?" or "What's the most common number of daily deaths?"

Important: Mode Can Have Quirks! ⚠️

Before we start coding, understand these special cases:

Case 1: No Mode (All values appear equally)

Data: 10, 20, 30, 40, 50 — each appears once. Technically ALL are modes, but that's not helpful!

Case 2: Multiple Modes (Tie!)

Data: 100, 100, 100, 200, 200, 200, 300 — Both 100 and 200 appear 3 times. This is called bimodal (two modes)!

Case 3: Clear Winner

Data: 100, 200, 200, 200, 300, 400 — 200 appears 3 times, others once. Clear mode = 200!

Pandas handles all these cases for us! Let's see how.

Step 1: Load the COVID-19 Dataset

import pandas as pd
import numpy as np

# Load the COVID-19 dataset
df = pd.read_csv('coviddata.csv')

# Quick overview
print(f"Dataset size: {df.shape[0]} rows, {df.shape[1]} columns")
print(df.head())

Method 1: Calculate Mode - The Basic Way

Let's find the most common (modal) value for new cases:

# Calculate mode of new_cases
mode_value = df['new_cases'].mode()
print("Mode of new_cases:")
print(mode_value)

Output:

0    0.0
dtype: float64

Wait, what?! The mode is 0? Yes! This makes sense because:

  • Many days early in the pandemic had ZERO new cases
  • Many countries had zero cases on many days
  • 0 appears more frequently than any other single number!

💡 Important Note:

.mode() returns a Series, not a single value! If there are multiple modes, it returns all of them. To get just the first mode value, use .mode()[0]

# Get just the first mode value
mode_single = df['new_cases'].mode()[0]
print(f"Most common value for new_cases: {mode_single}")

# Check how many times it appears
mode_count = (df['new_cases'] == mode_single).sum()
print(f"This value appears {mode_count} times")

# What percentage of the data is this?
percentage = (mode_count / len(df)) * 100
print(f"That's {percentage:.2f}% of all records")

Sample Output:

Most common value for new_cases: 0.0
This value appears 1234 times
That's 21.21% of all records

Method 2: Mode for Categorical Data (Text Columns)

This is where mode really shines! Let's find the most common continent in our dataset:

# Find most common continent
most_common_continent = df['continent'].mode()[0]
print(f"Most common continent in dataset: {most_common_continent}")

# How many times does it appear?
continent_count = (df['continent'] == most_common_continent).sum()
print(f"Appears {continent_count} times")

# Show frequency of all continents
print("\nFrequency of all continents:")
print(df['continent'].value_counts())

Sample Output:

Most common continent in dataset: Europe

Appears 1876 times

Frequency of all continents:
Europe           1876
Asia             1543
Africa           1234
North America     987
South America     876
Oceania           302
Name: continent, dtype: int64

Insight: Europe appears most in our dataset! This tells us our data has more European country records than others.

Method 3: Finding Mode for Each Group

Let's find the most common number of new cases for EACH continent:

# Mode of new_cases by continent
# Note: .mode() with groupby can be tricky, so we'll use a custom approach

def get_mode(series):
    mode_result = series.mode()
    if len(mode_result) > 0:
        return mode_result[0]
    else:
        return None

continent_modes = df.groupby('continent')['new_cases'].apply(get_mode)
print("Most common new_cases value by continent:")
print(continent_modes)

Sample Output:

Most common new_cases value by continent:
Africa           0.0
Asia             0.0
Europe           0.0
North America    0.0
Oceania          0.0
South America    0.0
Name: new_cases, dtype: float64

Why all zeros? Because for every continent, the most frequently occurring value for new_cases is 0 (many days with no new cases).

Method 4: Mode for Non-Zero Values (More Meaningful!) 💡

Let's exclude zeros to find more interesting patterns:

# Filter out zeros
df_nonzero = df[df['new_cases'] > 0]

# Find mode of non-zero new_cases
mode_nonzero = df_nonzero['new_cases'].mode()[0]
print(f"Most common NON-ZERO new_cases: {mode_nonzero}")

# By continent (excluding zeros)
print("\nMost common non-zero new_cases by continent:")
continent_modes_nonzero = df_nonzero.groupby('continent')['new_cases'].apply(get_mode)
print(continent_modes_nonzero)

Sample Output:

Most common NON-ZERO new_cases: 1.0

Most common non-zero new_cases by continent:
Africa           1.0
Asia             1.0
Europe           2.0
North America    3.0
Oceania          1.0
South America    1.0
Name: new_cases, dtype: float64

More interesting! When we exclude zeros, we see small numbers (1-3) are most common — meaning most days with cases had just a few new cases.

Method 5: Finding Most Common Country 🗺

Which country appears most frequently in our dataset?

# Most common location (country)
most_common_country = df['location'].mode()[0]
print(f"Most common country: {most_common_country}")

# Top 10 most frequent countries
print("\nTop 10 countries by frequency:")
top_countries = df['location'].value_counts().head(10)
print(top_countries)

Sample Output:

Most common country: United States

Top 10 countries by frequency:
United States    365
India            365
Brazil           365
United Kingdom   365
Germany          365
France           365
Italy            365
Spain            365
Canada           365
Australia        365
Name: location, dtype: int64

Interesting! All have 365 records — that's one year of daily data! This means our dataset is balanced with equal representation.

Method 6: Detecting Multiple Modes (Bimodal/Multimodal)

Let's check if we have multiple modes (ties for most common):

# Get all modes for new_deaths
all_modes = df['new_deaths'].mode()
print(f"Number of modes: {len(all_modes)}")
print("All mode values:")
print(all_modes)

if len(all_modes) > 1:
    print("\nThis data is MULTIMODAL (multiple values appear equally often)")
    print("The most common values are:")
    for mode in all_modes:
        count = (df['new_deaths'] == mode).sum()
        print(f"  {mode}: appears {count} times")
else:
    print(f"\nThis data is UNIMODAL (single mode: {all_modes[0]})")

Sample Output:

Number of modes: 1
All mode values:
0    0.0
dtype: float64

This data is UNIMODAL (single mode: 0.0)

Method 7: Mode vs Value_Counts (Understanding Frequency)

Mode just tells you the "winner." But value_counts() shows you the FULL picture:

# Mode: Just the most common
mode = df['continent'].mode()[0]
print(f"Mode: {mode}")

# Value_counts: See ALL frequencies
print("\nComplete frequency distribution:")
frequencies = df['continent'].value_counts()
print(frequencies)

# Percentage distribution
print("\nAs percentages:")
percentages = df['continent'].value_counts(normalize=True) * 100
print(percentages.round(2))

Sample Output:

Mode: Europe

Complete frequency distribution:
Europe           1876
Asia             1543
Africa           1234
North America     987
South America     876
Oceania           302
Name: continent, dtype: int64

As percentages:
Europe           32.25
Asia             26.53
Africa           21.21
North America    16.97
South America    15.06
Oceania           5.19
Name: continent, dtype: float64

💡 Pro Tip:

Use .mode() when you just need the most common value. Use .value_counts() when you want to see the frequency of ALL values. Often, .value_counts() is more informative!

Method 8: Practical Example - ISO Codes (Real Analysis!)

Let's find the most common ISO code and understand what it means:

# Most common ISO code
most_common_iso = df['iso_code'].mode()[0]
print(f"Most common ISO code: {most_common_iso}")

# Find the country name for this code
country_name = df[df['iso_code'] == most_common_iso]['location'].iloc[0]
print(f"This is: {country_name}")

# How many times does it appear?
iso_count = (df['iso_code'] == most_common_iso).sum()
print(f"Appears {iso_count} times in the dataset")

# Show top 10 ISO codes
print("\nTop 10 most frequent ISO codes:")
print(df['iso_code'].value_counts().head(10))

Method 9: Mode for Binned/Rounded Data

Sometimes finding mode on exact numbers isn't useful. Let's bin the data first:

# Create bins for total_cases
# Let's categorize as: Low, Medium, High, Very High

def categorize_cases(cases):
    if pd.isna(cases):
        return 'Unknown'
    elif cases == 0:
        return 'Zero'
    elif cases < 1000:
        return 'Low'
    elif cases < 100000:
        return 'Medium'
    elif cases < 1000000:
        return 'High'
    else:
        return 'Very High'

df['case_category'] = df['total_cases'].apply(categorize_cases)

# Find most common category
most_common_category = df['case_category'].mode()[0]
print(f"Most common case category: {most_common_category}")

# Show distribution
print("\nDistribution of case categories:")
print(df['case_category'].value_counts())

# As percentages
print("\nAs percentages:")
print(df['case_category'].value_counts(normalize=True).mul(100).round(2))

Sample Output:

Most common case category: Low

Distribution of case categories:
Low          2543
Medium       1876
Very High     876
High          321
Zero          202
Unknown         0
Name: case_category, dtype: int64

As percentages:
Low          43.71
Medium       32.25
Very High    15.06
High          5.52
Zero          3.47
Unknown       0.00
Name: case_category, dtype: float64

Insight: Most records fall in the "Low" category (under 1,000 cases), showing that for most countries on most days, case counts were relatively low!

Real-World Analysis: Complete Mode Analysis

Let's put everything together:

import pandas as pd

# Load data
df = pd.read_csv('coviddata.csv')

print("="*70)
print("COVID-19 MODE ANALYSIS - FINDING WHAT'S MOST COMMON")
print("="*70)

# 1. Most Common Continent
print("\n1. GEOGRAPHIC DISTRIBUTION:")
most_common_continent = df['continent'].mode()[0]
continent_pct = (df['continent'] == most_common_continent).sum() / len(df) * 100
print(f"   Most common continent: {most_common_continent} ({continent_pct:.1f}% of data)")

# 2. Most Common Country
print("\n2. COUNTRY REPRESENTATION:")
most_common_country = df['location'].mode()[0]
country_count = (df['location'] == most_common_country).sum()
print(f"   Most common country: {most_common_country}")
print(f"   Appears {country_count} times")

# 3. Most Common Values (with context)
print("\n3. MOST FREQUENT VALUES:")
print(f"   New cases mode: {df['new_cases'].mode()[0]}")
print(f"   New deaths mode: {df['new_deaths'].mode()[0]}")
print(f"   (Note: 0 is most common because many days had no new cases/deaths)")

# 4. Non-zero analysis
df_active = df[(df['new_cases'] > 0)]
if len(df_active) > 0:
    print(f"\n   When excluding zeros:")
    print(f"   Most common new_cases: {df_active['new_cases'].mode()[0]}")

# 5. Check for multimodal data
print("\n4. MULTIMODAL CHECK:")
modes_count = len(df['new_deaths'].mode())
if modes_count > 1:
    print(f"   new_deaths has {modes_count} modes (multimodal)")
else:
    print(f"   new_deaths has 1 mode (unimodal)")

# 6. Categorical Analysis
print("\n5. TOP 5 MOST COMMON COUNTRIES:")
top5 = df['location'].value_counts().head(5)
for i, (country, count) in enumerate(top5.items(), 1):
    print(f"   {i}. {country}: {count} records")

# 7. Test Types (if column exists)
if 'tests_units' in df.columns:
    print("\n6. TESTING METHODS:")
    most_common_test = df['tests_units'].mode()
    if len(most_common_test) > 0:
        print(f"   Most common test type: {most_common_test[0]}")
        print(f"\n   All test types:")
        print(df['tests_units'].value_counts())

print("\n" + "="*70)
print("KEY INSIGHT: Mode shows us what occurs most frequently!")
print("="*70)

When to Use Mode vs Mean vs Median? 🤔

Use MODE when:

  • You have categorical data (countries, continents, categories)
  • You want to know the most common value
  • You're working with discrete counts that repeat often
  • You want to find typical patterns in repeated data
  • You're analyzing survey responses or choices

Use MEAN when:

  • You have numerical data without extreme outliers
  • You want to find the average
  • All values should be weighted equally

Use MEDIAN when:

  • You have numerical data with outliers
  • You want the middle value
  • You want to know what's typical (not skewed by extremes)

Quick Reference Cheat Sheet

# Basic mode
df['column'].mode()              # Returns Series with all modes
df['column'].mode()[0]           # Get first mode value

# Check if multiple modes exist
len(df['column'].mode())         # If > 1, multimodal

# Mode with groupby (custom function needed)
def get_mode(series):
    mode = series.mode()
    return mode[0] if len(mode) > 0 else None

df.groupby('group')['column'].apply(get_mode)

# Value counts (often more useful than mode!)
df['column'].value_counts()                    # Frequency of all values
df['column'].value_counts().head(1)            # Most common value with count
df['column'].value_counts(normalize=True)      # As percentages

# Count how many times mode appears
mode_val = df['column'].mode()[0]
(df['column'] == mode_val).sum()               # Count occurrences

# Mode for non-zero values
df[df['column'] > 0]['column'].mode()[0]

Common Mistakes to Avoid ⚠️

  • Don't forget [0]! .mode() returns a Series, not a single value. Use .mode()[0] to get the actual value
  • Don't use mode for continuous data! For numbers like 123.456789, every value might be unique, making mode useless
  • Don't ignore multiple modes! Check len(df['column'].mode()) — you might have ties!
  • Don't use mode for unique identifiers! IDs, dates, etc. usually don't have meaningful modes
  • Don't forget about value_counts()! It often gives more insight than just the mode
  • Don't assume mode is meaningful! If mode appears only 2 times out of 10,000 records, it's not very "common"

Comments