📖 Simple Definition:
Variance = A measure of how spread out numbers are from their average.
Low variance = Numbers are close together (consistent, predictable)
High variance = Numbers are far apart (inconsistent, unpredictable)
📚 Source & License
Code examples are adapted from Exploratory Data Analysis with Python Cookbook by Packt Publishing.
GitHub repository: https://github.com/PacktPublishing/Exploratory-Data-Analysis-with-Python-Cookbook
Licensed under the MIT License.
What is Variance?
Imagine you're comparing daily temperatures in two cities for a week:
City A: 20°C, 21°C, 20°C, 19°C, 20°C, 21°C, 19°C
City B: 10°C, 30°C, 5°C, 35°C, 15°C, 25°C, 20°C
Both cities have the same average temperature (20°C), but they feel completely different! City A is stable and predictable. City B is all over the place — wildly unpredictable!
Variance is the number that captures this difference. It tells you how much the data "varies" or "spreads out" from the average.
The Math Behind Variance
Here's how variance is calculated step-by-step:
Example: Data = [10, 12, 14, 16, 18]
Step 1: Find the mean (average)
Mean = (10 + 12 + 14 + 16 + 18) ÷ 5 = 70 ÷ 5 = 14
Step 2: Subtract the mean from each value (this shows how far each point is from the average)
10 - 14 = -4
12 - 14 = -2
14 - 14 = 0
16 - 14 = 2
18 - 14 = 4
Step 3: Square each difference (this makes all numbers positive and emphasizes larger differences)
(-4)² = 16
(-2)² = 4
(0)² = 0
(2)² = 4
(4)² = 16
Step 4: Find the average of these squared differences
Variance = (16 + 4 + 0 + 4 + 16) ÷ 5 = 40 ÷ 5 = 8
💡 Key Insight:
Variance is always a positive number (because we square the differences). The bigger the variance, the more "spread out" your data is!
Sample Variance vs Population Variance
There are actually TWO types of variance, and this confuses many beginners:
🔵 Population Variance (ddof=0)
Use when you have ALL the data for the entire group.
Formula: Sum of squared differences ÷ N (total number of values)
Example: You have test scores for ALL students in the class
🟠 Sample Variance (ddof=1) — DEFAULT in Pandas!
Use when you have only a SAMPLE of the data.
Formula: Sum of squared differences ÷ (N - 1)
Example: You surveyed 100 people out of a city of 1 million
Why N-1? This corrects for bias when estimating the variance of the full population from a sample (it's called "Bessel's correction")
Important: Pandas uses ddof=1 by default (sample variance), which is what you want 99% of the time when analyzing real-world data!
Step 1: Load the COVID-19 Dataset
import pandas as pd
import numpy as np
# Load the COVID-19 dataset
df = pd.read_csv('coviddata.csv')
# Quick overview
print(f"Dataset size: {df.shape[0]} rows, {df.shape[1]} columns")
print(df.head())
Method 1: Calculate Variance - The Basic Way
Let's find the variance of new COVID cases:
# Calculate variance of new_cases
variance = df['new_cases'].var()
print(f"Variance of new_cases: {variance}")
# Let's also see the mean for context
mean = df['new_cases'].mean()
print(f"Mean of new_cases: {mean}")
Sample Output:
Variance of new_cases: 4523891.45
Mean of new_cases: 1234.56
What does this mean? The variance is HUGE (over 4.5 million) compared to the mean (1,234). This tells us that COVID case numbers varied wildly — some days had very few cases, other days had massive spikes. The data is very "spread out"!
⚠️ Important Note:
Variance has squared units. If new_cases is measured in "number of cases," then variance is in "cases squared" — which is hard to interpret! That's why we often use standard deviation (the square root of variance) for easier interpretation.
Method 2: Population Variance vs Sample Variance
Let's compare both types of variance:
# Sample variance (default, ddof=1)
sample_var = df['new_cases'].var()
print(f"Sample variance (ddof=1): {sample_var}")
# Population variance (ddof=0)
population_var = df['new_cases'].var(ddof=0)
print(f"Population variance (ddof=0): {population_var}")
# The difference
difference = sample_var - population_var
print(f"Difference: {difference}")
print(f"Sample variance is slightly larger (more conservative estimate)")
Sample Output:
Sample variance (ddof=1): 4523891.45
Population variance (ddof=0): 4522134.78
Difference: 1756.67
Sample variance is slightly larger (more conservative estimate)
Notice: Sample variance is always slightly larger! This is intentional — it gives a more conservative (safer) estimate when you're working with sample data.
Method 3: Variance for Multiple Columns
Calculate variance for several columns at once:
# Variance for multiple columns
variance_multiple = df[['new_cases', 'new_deaths', 'total_cases']].var()
print("Variance for multiple columns:")
print(variance_multiple)
print()
# Make it more readable
print("Formatted output:")
for col in ['new_cases', 'new_deaths', 'total_cases']:
var = df[col].var()
print(f"{col}: {var:,.2f}")
Sample Output:
Variance for multiple columns:
new_cases 4523891.45
new_deaths 12345.67
total_cases 8900000000.00
dtype: float64
Formatted output:
new_cases: 4,523,891.45
new_deaths: 12,345.67
total_cases: 8,900,000,000.00
Insight: total_cases has ENORMOUS variance (billions!) because it's a cumulative number that keeps growing. new_deaths has lower variance than new_cases, meaning death counts were more consistent (less variable) than case counts.
Method 4: Variance by Group (Continent)
Let's see how variance differs across continents:
# Variance of new_cases by continent
variance_by_continent = df.groupby('continent')['new_cases'].var()
print("Variance of new_cases by continent:")
print(variance_by_continent.sort_values(ascending=False))
print()
# With mean for context
print("Mean and Variance by continent:")
stats = df.groupby('continent')['new_cases'].agg(['mean', 'var'])
stats['std'] = stats['var'] ** 0.5 # Standard deviation for easier interpretation
print(stats.sort_values('var', ascending=False))
Sample Output:
Variance of new_cases by continent:
Europe 8934567.89
North America 7234567.12
Asia 6543210.45
South America 4321098.76
Africa 2109876.54
Oceania 987654.32
Name: new_cases, dtype: float64
Mean and Variance by continent:
mean var std
Europe 2345.67 8934567.89 2989.08
North America 1987.34 7234567.12 2689.72
Asia 1654.23 6543210.45 2558.17
South America 1432.10 4321098.76 2078.72
Africa 876.54 2109876.54 1452.54
Oceania 234.56 987654.32 993.81
Insight: Europe has the highest variance in new cases — meaning case numbers fluctuated wildly in Europe! Oceania has the lowest variance, indicating more stable (consistent) case numbers over time.
Method 5: Understanding Variance with Standard Deviation
Variance is hard to interpret because of squared units. Let's compare it with standard deviation:
# Calculate both variance and standard deviation
variance = df['new_cases'].var()
std_dev = df['new_cases'].std()
mean = df['new_cases'].mean()
print(f"Mean: {mean:.2f} cases")
print(f"Variance: {variance:,.2f} cases²")
print(f"Standard Deviation: {std_dev:.2f} cases")
print()
print("Notice: Standard Deviation = √Variance")
print(f"Check: {variance**0.5:.2f} = {std_dev:.2f}")
print()
print("Interpretation:")
print(f" - On average, new_cases is about {mean:.0f}")
print(f" - Typically, values deviate by ±{std_dev:.0f} from the mean")
print(f" - That means most days have between {mean-std_dev:.0f} and {mean+std_dev:.0f} new cases")
Sample Output:
Mean: 1234.56 cases
Variance: 4,523,891.45 cases²
Standard Deviation: 2127.04 cases
Notice: Standard Deviation = √Variance
Check: 2127.04 = 2127.04
Interpretation:
- On average, new_cases is about 1235
- Typically, values deviate by ±2127 from the mean
- That means most days have between -892 and 3362 new cases
💡 Pro Tip:
Always look at variance AND standard deviation together! Variance is useful for mathematical calculations, but standard deviation is easier to interpret because it's in the same units as your original data.
Method 6: Variance for Non-Null Values Only
Handle missing data properly when calculating variance:
# Check for missing values
print(f"Missing values in new_cases: {df['new_cases'].isna().sum()}")
print(f"Total values: {len(df)}")
print()
# Variance automatically skips NaN values
variance_with_nulls = df['new_cases'].var()
print(f"Variance (auto-skips NaN): {variance_with_nulls:,.2f}")
print()
# Explicitly drop NaN first (same result)
variance_no_nulls = df['new_cases'].dropna().var()
print(f"Variance (explicit dropna): {variance_no_nulls:,.2f}")
print()
# Count non-null values used
non_null_count = df['new_cases'].notna().sum()
print(f"Non-null values used in calculation: {non_null_count}")
Method 7: Comparing Variance Across Countries
Which countries had the most volatile (unstable) COVID case numbers?
# Variance by country
country_variance = df.groupby('location')['new_cases'].var()
# Top 10 countries with highest variance (most volatile)
print("Top 10 countries with HIGHEST variance (most volatile):")
top_volatile = country_variance.nlargest(10)
for country, var in top_volatile.items():
mean = df[df['location'] == country]['new_cases'].mean()
std = var ** 0.5
print(f" {country}: Variance={var:,.0f}, Mean={mean:.0f}, StdDev={std:.0f}")
print()
# Top 10 countries with lowest variance (most stable)
print("Top 10 countries with LOWEST variance (most stable):")
top_stable = country_variance.nsmallest(10)
for country, var in top_stable.items():
mean = df[df['location'] == country]['new_cases'].mean()
std = var ** 0.5
print(f" {country}: Variance={var:,.0f}, Mean={mean:.0f}, StdDev={std:.0f}")
Sample Output:
Top 10 countries with HIGHEST variance (most volatile):
United States: Variance=89,234,567, Mean=12,345, StdDev=9,446
India: Variance=78,456,123, Mean=10,234, StdDev=8,857
Brazil: Variance=67,890,234, Mean=8,765, StdDev=8,239
...
Top 10 countries with LOWEST variance (most stable):
Small Island A: Variance=12, Mean=5, StdDev=3
Small Island B: Variance=23, Mean=8, StdDev=5
...
Insight: Large countries with big populations tend to have higher variance (more unpredictable case numbers). Small countries with smaller populations have lower variance (more stable, predictable case numbers).
Method 8: Coefficient of Variation (Normalized Variance)
Sometimes we want to compare variability between different scales. The coefficient of variation helps:
# Coefficient of Variation (CV) = (Standard Deviation / Mean) × 100%
# This tells you the variability as a percentage of the mean
def coefficient_of_variation(series):
"""Calculate coefficient of variation as a percentage"""
mean = series.mean()
std = series.std()
if mean == 0:
return 0
cv = (std / mean) * 100
return cv
# Calculate CV for new_cases
cv_cases = coefficient_of_variation(df['new_cases'])
print(f"Coefficient of Variation for new_cases: {cv_cases:.2f}%")
print()
# By continent
print("Coefficient of Variation by continent:")
cv_by_continent = df.groupby('continent')['new_cases'].apply(coefficient_of_variation)
print(cv_by_continent.sort_values(ascending=False))
print()
print("Interpretation:")
print(" - High CV (>100%): Very inconsistent data")
print(" - Medium CV (50-100%): Moderate variability")
print(" - Low CV (<50 code="" consistent="" data="" relatively="">50>
Sample Output:
Coefficient of Variation for new_cases: 172.34%
Coefficient of Variation by continent:
Oceania 423.89
Africa 165.87
South America 145.23
Asia 154.67
North America 135.34
Europe 127.45
Name: new_cases, dtype: float64
Interpretation:
- High CV (>100%): Very inconsistent data
- Medium CV (50-100%): Moderate variability
- Low CV (<50 code="" consistent="" data="" relatively="">50>
Insight: All continents have CV > 100%, meaning COVID case numbers were highly variable everywhere! Oceania had the highest relative variability (over 400%), meaning their case numbers were extremely inconsistent relative to their average.
Method 9: Rolling Variance (Variance Over Time)
See how variance changes over time using a rolling window:
# For a specific country, calculate 7-day rolling variance
country_data = df[df['location'] == 'United States'].copy()
country_data = country_data.sort_values('date')
# Calculate 7-day rolling variance
country_data['rolling_variance'] = country_data['new_cases'].rolling(window=7).var()
country_data['rolling_std'] = country_data['new_cases'].rolling(window=7).std()
print("7-day Rolling Variance Sample:")
print(country_data[['date', 'new_cases', 'rolling_variance', 'rolling_std']].tail(10))
print()
# Find periods of highest and lowest variance
max_var_date = country_data.loc[country_data['rolling_variance'].idxmax(), 'date']
min_var_date = country_data.loc[country_data['rolling_variance'].idxmin(), 'date']
print(f"Most volatile 7-day period: {max_var_date}")
print(f"Most stable 7-day period: {min_var_date}")
Insight: Rolling variance helps identify when case numbers were most unpredictable (high variance) vs most stable (low variance) over time.
Method 10: Complete Variance Analysis
Let's put everything together in a comprehensive analysis:
import pandas as pd
# Load data
df = pd.read_csv('coviddata.csv')
print("="*70)
print("COVID-19 VARIANCE ANALYSIS - UNDERSTANDING DATA SPREAD")
print("="*70)
# 1. Overall Variance
print("\n1. OVERALL STATISTICS:")
for col in ['new_cases', 'new_deaths']:
mean = df[col].mean()
variance = df[col].var()
std = df[col].std()
cv = (std / mean * 100) if mean != 0 else 0
print(f"\n{col}:")
print(f" Mean: {mean:,.2f}")
print(f" Variance: {variance:,.2f}")
print(f" Std Dev: {std:,.2f}")
print(f" Coefficient of Variation: {cv:.2f}%")
# 2. Variance by Continent
print("\n2. VARIANCE BY CONTINENT (new_cases):")
continent_stats = df.groupby('continent')['new_cases'].agg(['mean', 'var', 'std'])
continent_stats['cv'] = (continent_stats['std'] / continent_stats['mean'] * 100)
continent_stats = continent_stats.sort_values('var', ascending=False)
print(continent_stats)
# 3. Sample vs Population Variance
print("\n3. SAMPLE vs POPULATION VARIANCE:")
sample_var = df['new_cases'].var(ddof=1)
pop_var = df['new_cases'].var(ddof=0)
print(f"Sample variance (ddof=1): {sample_var:,.2f}")
print(f"Population variance (ddof=0): {pop_var:,.2f}")
print(f"Difference: {sample_var - pop_var:,.2f}")
# 4. Most and Least Volatile Countries
print("\n4. MOST vs LEAST VOLATILE COUNTRIES:")
country_var = df.groupby('location')['new_cases'].var()
print("\nTop 5 Most Volatile:")
for country, var in country_var.nlargest(5).items():
print(f" {country}: {var:,.0f}")
print("\nTop 5 Least Volatile:")
for country, var in country_var.nsmallest(5).items():
print(f" {country}: {var:,.0f}")
# 5. Interpretation Guide
print("\n" + "="*70)
print("KEY INSIGHTS:")
print("="*70)
print("High variance = Data is spread out (inconsistent, unpredictable)")
print("Low variance = Data is clustered (consistent, predictable)")
print("Use standard deviation for easier interpretation (same units as data)")
print("Use coefficient of variation to compare variability across different scales")
print("="*70)
When to Use Variance? 🤔
Use VARIANCE when:
- You need to measure how spread out your data is
- You're doing statistical calculations (variance is mathematically convenient)
- You're comparing variability between different groups
- You need to identify inconsistent or volatile patterns
- You're performing hypothesis testing or ANOVA
Use STANDARD DEVIATION when:
- You want to interpret the spread in the same units as your data
- You're explaining results to non-technical audiences
- You want to identify outliers (values beyond 2-3 standard deviations)
Quick Reference Cheat Sheet
# Basic variance
df['column'].var() # Sample variance (ddof=1, default)
df['column'].var(ddof=0) # Population variance
df['column'].std() # Standard deviation (√variance)
# Variance for multiple columns
df[['col1', 'col2', 'col3']].var()
# Variance by group
df.groupby('group')['column'].var()
# Multiple statistics
df.groupby('group')['column'].agg(['mean', 'var', 'std'])
# Coefficient of Variation
(df['column'].std() / df['column'].mean()) * 100
# Rolling variance (7-day window)
df['column'].rolling(window=7).var()
# Variance with conditions
df[df['column'] > 0]['column'].var() # Variance of positive values only
Common Mistakes to Avoid ⚠️
- Don't confuse variance with standard deviation! They measure the same thing, but variance is squared. Always check which one you're using.
- Don't forget about units! Variance is in squared units (cases²), which is hard to interpret. Use standard deviation for better interpretation.
- Don't ignore missing values! Pandas automatically skips NaN, but you should know how many values you're working with.
- Don't use variance for categorical data! Variance only works with numerical data. You can't calculate variance of country names!
- Don't assume high variance is bad! Sometimes high variance is expected (like stock prices). Context matters!
- Don't compare raw variances across different scales! Use coefficient of variation instead when comparing data with different units or scales.
Comments
Post a Comment