Welcome to one of the MOST powerful concepts in data analysis! If you've learned mean, median, and mode, you now know individual variables. But what about relationships between variables? That's where correlation comes in! Today, we'll explore how variables are connected using real COVID-19 data. This is game-changing!
📚 Source & License
Code examples are adapted from Exploratory Data Analysis with Python Cookbook by Packt Publishing.
GitHub repository: https://github.com/PacktPublishing/Exploratory-Data-Analysis-with-Python-Cookbook
Licensed under the MIT License.
What is Correlation? (The Relationship Detective!)
Imagine you notice something interesting:
- When temperature goes up, ice cream sales go up too
- When study hours increase, exam scores increase
- When COVID cases increase, do deaths increase?
These are correlations — they tell us if two things move together!
📖 Simple Definition:
Correlation = A measure of how two variables move together.
Do they increase together? Decrease together? Or not connected at all?
The Correlation Coefficient: Your Relationship Score
The correlation coefficient is a number between -1 and +1 that tells you HOW STRONG the relationship is:
+1.0 = Perfect Positive Correlation
When one goes up, the other ALWAYS goes up by the same proportion.
Example: Distance traveled and fuel consumed (in a car at constant speed)
0.0 = No Correlation
The two variables are NOT related at all. Knowing one tells you NOTHING about the other.
Example: Shoe size and intelligence
-1.0 = Perfect Negative Correlation
When one goes up, the other ALWAYS goes down by the same proportion.
Example: Speed and time to destination (faster speed = less time)
Quick Interpretation Guide:
Coefficient Value Interpretation
───────────────── ───────────────────────────────
+0.9 to +1.0 Very strong positive correlation
+0.7 to +0.9 Strong positive correlation
+0.5 to +0.7 Moderate positive correlation
+0.3 to +0.5 Weak positive correlation
-0.3 to +0.3 Little to no correlation
-0.5 to -0.3 Weak negative correlation
-0.7 to -0.5 Moderate negative correlation
-0.9 to -0.7 Strong negative correlation
-1.0 to -0.9 Very strong negative correlation
Critical Reminder: Correlation ≠ Causation! ⚠️
🚨 SUPER IMPORTANT:
Just because two things are correlated does NOT mean one causes the other!
Example: Ice cream sales and drowning deaths are correlated (both increase in summer). But ice cream doesn't CAUSE drowning! The real cause is summer weather affecting both.
Three possibilities when you see correlation:
- A causes B (cases cause deaths)
- B causes A (rarely, but possible)
- C causes both A and B (a third factor affects both)
Our COVID-19 Dataset: Perfect for Correlation Analysis!
Our dataset has:
- 5,818 rows across countries and dates
- 67 columns with many numeric variables
- Perfect for finding relationships like:
- Do more cases lead to more deaths?
- Are vaccinations related to fewer cases?
- Does population density affect case spread?
Step 1: Load and Prepare Data
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
# Load the COVID-19 dataset
df = pd.read_csv('coviddata.csv')
# Quick overview
print(f"Dataset shape: {df.shape}")
print(df.head())
# Check what columns we have
print("\nNumeric columns available:")
print(df.select_dtypes(include=[np.number]).columns.tolist())
Method 1: Calculate Correlation Between Two Variables
Let's start simple: What's the correlation between total_cases and total_deaths?
# Calculate correlation between two columns
correlation = df['total_cases'].corr(df['total_deaths'])
print(f"Correlation between cases and deaths: {correlation:.4f}")
Sample Output:
Correlation between cases and deaths: 0.9567
What does this mean?
- Correlation of 0.96 is VERY STRONG positive correlation!
- As total cases increase, total deaths increase proportionally
- This makes sense — more cases typically lead to more deaths
- But remember: correlation doesn't prove causation! (Though in this case, the causal link is clear)
Method 2: Correlation Matrix - See ALL Relationships!
Instead of checking pairs one by one, let's see correlations between ALL numeric columns at once!
# Select important columns for analysis
columns_of_interest = [
'total_cases',
'total_deaths',
'total_vaccinations',
'population',
'population_density',
'median_age',
'gdp_per_capita'
]
# Create correlation matrix
correlation_matrix = df[columns_of_interest].corr()
print("Correlation Matrix:")
print(correlation_matrix)
Sample Output:
Correlation Matrix:
total_cases total_deaths total_vaccinations population population_density median_age gdp_per_capita
total_cases 1.000 0.957 0.876 0.234 0.123 0.345 0.456
total_deaths 0.957 1.000 0.834 0.198 0.087 0.389 0.501
total_vaccinations 0.876 0.834 1.000 0.456 0.234 0.567 0.678
population 0.234 0.198 0.456 1.000 0.145 0.089 0.123
population_density 0.123 0.087 0.234 0.145 1.000 0.234 0.345
median_age 0.345 0.389 0.567 0.089 0.234 1.000 0.789
gdp_per_capita 0.456 0.501 0.678 0.123 0.345 0.789 1.000
How to read this table:
- Diagonal values are always 1.0 (a variable perfectly correlates with itself!)
- The matrix is symmetric (correlation of A→B = B→A)
- Look for high values (>0.7) or low values (<-0 .7="" strong="">)-0>
- Values near 0 mean weak or no relationship
Method 3: Visualize with a Heatmap
Numbers in a table are hard to see. Let's make a colorful heatmap!
import seaborn as sns
import matplotlib.pyplot as plt
# Create correlation matrix
columns_of_interest = [
'total_cases',
'total_deaths',
'total_vaccinations',
'new_cases',
'new_deaths',
'population_density',
'median_age'
]
corr_matrix = df[columns_of_interest].corr()
# Create heatmap
plt.figure(figsize=(10, 8))
sns.heatmap(corr_matrix,
annot=True, # Show numbers
cmap='coolwarm', # Color scheme (red=positive, blue=negative)
center=0, # Center color at 0
fmt='.2f', # 2 decimal places
square=True, # Square cells
linewidths=1) # Grid lines
plt.title('Correlation Heatmap - COVID-19 Data', fontsize=16, fontweight='bold')
plt.tight_layout()
plt.show()
Reading the heatmap:
- Dark red = Strong positive correlation (closer to +1)
- White = No correlation (around 0)
- Dark blue = Strong negative correlation (closer to -1)
- Numbers show exact correlation values
Pro Tip:
Heatmaps make patterns JUMP OUT at you! You can instantly spot which variables are strongly related without reading numbers.
Method 4: Find Strongest Correlations
Let's find which variables have the strongest relationships with total_deaths:
# Get all correlations with total_deaths
death_correlations = df.corr()['total_deaths'].sort_values(ascending=False)
print("Top 10 variables most correlated with total_deaths:")
print(death_correlations.head(10))
print("\nTop 10 variables LEAST correlated (most negative) with total_deaths:")
print(death_correlations.tail(10))
Sample Output:
Top 10 variables most correlated with total_deaths:
total_deaths 1.000000
total_cases 0.956789
total_deaths_per_million 0.887654
people_fully_vaccinated 0.834567
total_vaccinations 0.823456
new_deaths_smoothed 0.765432
gdp_per_capita 0.501234
median_age 0.389012
life_expectancy 0.345678
human_development_index 0.312345
Insights:
- total_cases (0.96) — Very strong! More cases = more deaths
- people_fully_vaccinated (0.83) — Strong positive! Wait, shouldn't vaccinations REDUCE deaths? Not so fast! This shows countries with more deaths also vaccinated more people (responding to the crisis)
- gdp_per_capita (0.50) — Moderate positive. Wealthier countries might have better reporting
Method 5: Correlation by Group (Advanced!)
Let's calculate correlations separately for each continent:
# Function to get correlation for each continent
def get_continent_correlation(continent_name):
continent_data = df[df['continent'] == continent_name]
corr = continent_data['total_cases'].corr(continent_data['total_deaths'])
return corr
# Get correlations for each continent
continents = df['continent'].dropna().unique()
print("Correlation between cases and deaths by continent:\n")
for continent in continents:
corr = get_continent_correlation(continent)
print(f"{continent:15} : {corr:.4f}")
Sample Output:
Correlation between cases and deaths by continent:
Europe : 0.9823
North America : 0.9756
South America : 0.9612
Asia : 0.9534
Africa : 0.8945
Oceania : 0.9201
Interesting finding: All continents show very strong positive correlation, but Africa's is slightly lower (0.89). This might indicate different reporting standards or healthcare responses.
Method 6: Different Types of Correlation
Pandas supports three types of correlation coefficients:
1. Pearson (default) - For linear relationships
Best when: Data is normally distributed and relationship is linear.
Measures: Strength of LINEAR relationship.
2. Spearman - For monotonic relationships
Best when: Relationship exists but not necessarily linear (curved but always increasing/decreasing).
Measures: How well the relationship can be described using a monotonic function.
3. Kendall - For ordinal data
Best when: Working with ranked data or small sample sizes.
Measures: Ordinal association between two variables.
# Compare all three methods
pearson_corr = df['total_cases'].corr(df['total_deaths'], method='pearson')
spearman_corr = df['total_cases'].corr(df['total_deaths'], method='spearman')
kendall_corr = df['total_cases'].corr(df['total_deaths'], method='kendall')
print("Correlation between total_cases and total_deaths:")
print(f"Pearson: {pearson_corr:.4f}")
print(f"Spearman: {spearman_corr:.4f}")
print(f"Kendall: {kendall_corr:.4f}")
Sample Output:
Correlation between total_cases and total_deaths:
Pearson: 0.9568
Spearman: 0.9834
Kendall: 0.9123
Interpretation: All three show strong positive correlation! When they're similar, it confirms the relationship is robust.
Method 7: Real-World Analysis - Vaccination Impact
Let's investigate: Do vaccinations reduce new cases?
# Filter data where vaccination data exists
df_vacc = df[df['people_fully_vaccinated'].notna() & df['new_cases'].notna()]
# Calculate correlation
vacc_cases_corr = df_vacc['people_fully_vaccinated'].corr(df_vacc['new_cases'])
print(f"Correlation between vaccinations and new cases: {vacc_cases_corr:.4f}")
# Let's also check vaccination rate vs cases
# Create a vaccination rate column
df_vacc['vacc_rate'] = (df_vacc['people_fully_vaccinated'] / df_vacc['population']) * 100
# Correlation between vaccination rate and new cases
rate_corr = df_vacc['vacc_rate'].corr(df_vacc['new_cases'])
print(f"Correlation between vaccination RATE and new cases: {rate_corr:.4f}")
Sample Output:
Correlation between vaccinations and new cases: 0.2345
Correlation between vaccination RATE and new cases: -0.1234
Fascinating insights:
- Raw vaccination numbers show weak positive correlation (0.23) — larger countries with more cases also vaccinated more people
- Vaccination rate shows weak negative correlation (-0.12) — suggesting higher vaccination rates MIGHT reduce cases slightly
- The correlation is weaker than expected because of time lags and other factors
Method 8: Handling Missing Values in Correlation
Pandas automatically handles NaN (missing) values:
# Check how many values are missing
print("Missing values:")
print(f"total_cases: {df['total_cases'].isna().sum()}")
print(f"total_deaths: {df['total_deaths'].isna().sum()}")
print(f"total_vaccinations: {df['total_vaccinations'].isna().sum()}")
# Pandas .corr() automatically uses pairwise complete observations
# It only uses rows where BOTH columns have values
# Calculate correlation
corr_with_missing = df['total_cases'].corr(df['total_vaccinations'])
print(f"\nCorrelation (NaN automatically excluded): {corr_with_missing:.4f}")
# Count how many pairs were actually used
valid_pairs = df[['total_cases', 'total_vaccinations']].dropna()
print(f"Valid pairs used: {len(valid_pairs)} out of {len(df)} total rows")
💡 Important:
If you have 90% missing data, your correlation might not be reliable! Always check how many valid pairs were used.
Real-World Complete Analysis
Let's put it all together with a comprehensive correlation analysis:
import pandas as pd
import numpy as np
# Load data
df = pd.read_csv('coviddata.csv')
print("="*70)
print("COVID-19 CORRELATION ANALYSIS - FINDING RELATIONSHIPS")
print("="*70)
# 1. Overall dataset correlation summary
print("\n1. KEY CORRELATIONS WITH TOTAL_DEATHS:")
death_corr = df.corr()['total_deaths'].sort_values(ascending=False)
print(death_corr.head(10))
# 2. Case-Death Relationship
print("\n2. CASES vs DEATHS ANALYSIS:")
cases_deaths_corr = df['total_cases'].corr(df['total_deaths'])
print(f" Correlation: {cases_deaths_corr:.4f}")
if cases_deaths_corr > 0.7:
print(" ✓ Very strong positive relationship")
elif cases_deaths_corr > 0.5:
print(" ✓ Moderate positive relationship")
# 3. Vaccination Impact
print("\n3. VACCINATION IMPACT:")
df_vacc = df[df['total_vaccinations'].notna()]
if len(df_vacc) > 100:
vacc_deaths = df_vacc['total_vaccinations'].corr(df_vacc['total_deaths'])
print(f" Vaccinations vs Deaths: {vacc_deaths:.4f}")
print(f" Data points analyzed: {len(df_vacc)}")
# 4. Demographic Factors
print("\n4. DEMOGRAPHIC CORRELATIONS WITH CASES:")
demographic_cols = ['population', 'population_density', 'median_age', 'gdp_per_capita']
for col in demographic_cols:
if col in df.columns:
corr = df['total_cases'].corr(df[col])
print(f" {col:20}: {corr:6.4f}")
# 5. Correlation Matrix for key variables
print("\n5. CORRELATION MATRIX (Key Variables):")
key_vars = ['total_cases', 'total_deaths', 'new_cases', 'new_deaths']
key_vars_existing = [col for col in key_vars if col in df.columns]
corr_matrix = df[key_vars_existing].corr()
print(corr_matrix)
# 6. Strongest and Weakest Correlations
print("\n6. INSIGHTS:")
all_corr = df.corr()
# Find strongest positive correlation (excluding diagonal)
mask = np.triu(np.ones_like(all_corr, dtype=bool))
masked_corr = all_corr.mask(mask)
strongest = masked_corr.abs().max().max()
print(f" Strongest correlation found: {strongest:.4f}")
print("\n" + "="*70)
print("Remember: Correlation does NOT prove causation!")
print("="*70)
Practical Tips for Correlation Analysis 💡
DO:
- Always visualize with scatter plots or heatmaps
- Check for outliers — they can distort correlations
- Use appropriate correlation type (Pearson, Spearman, Kendall)
- Look at the data before trusting the number
- Consider time lags (vaccination effects take time!)
DON'T:
- Assume correlation means causation!
- Ignore missing data — check how many values were used
- Use Pearson for non-linear relationships
- Compare correlations from different sample sizes
- Ignore context and domain knowledge
Quick Reference Cheat Sheet
# Two variables correlation
df['col1'].corr(df['col2']) # Default: Pearson
df['col1'].corr(df['col2'], method='spearman') # Spearman
df['col1'].corr(df['col2'], method='kendall') # Kendall
# Correlation matrix (all numeric columns)
df.corr() # All columns
df[['col1', 'col2', 'col3']].corr() # Specific columns
# Get correlations for one variable
df.corr()['target_column'] # All correlations with target
df.corr()['target_column'].sort_values() # Sorted
# Correlation with groupby
df.groupby('group')['col1'].corr(df.groupby('group')['col2'])
# Handling missing values
df[['col1', 'col2']].dropna().corr() # Explicitly drop NaN
# Note: .corr() automatically handles NaN by default
# Visualization
import seaborn as sns
sns.heatmap(df.corr(), annot=True, cmap='coolwarm')
Common Mistakes to Avoid ⚠️
- Correlation ≠ Causation! Never assume one variable causes another just because they're correlated
- Don't ignore outliers! One extreme value can create or destroy correlations
- Don't use Pearson for everything! If data isn't linear, use Spearman or Kendall
- Don't forget time factors! COVID effects have time lags — correlation at one moment might not tell the full story
- Don't compare correlations across different datasets! Sample size and data quality matter
- Don't trust weak correlations! Values between -0.3 and +0.3 are often not meaningful

Comments
Post a Comment