Skip to main content

Understanding Correlation Coefficient

Calculating read time…

Welcome to one of the MOST powerful concepts in data analysis! If you've learned mean, median, and mode, you now know individual variables. But what about relationships between variables? That's where correlation comes in! Today, we'll explore how variables are connected using real COVID-19 data. This is game-changing!



📚 Source & License

Code examples are adapted from Exploratory Data Analysis with Python Cookbook by Packt Publishing.
GitHub repository: https://github.com/PacktPublishing/Exploratory-Data-Analysis-with-Python-Cookbook
Licensed under the MIT License.

What is Correlation? (The Relationship Detective!)

Imagine you notice something interesting:

  • When temperature goes up, ice cream sales go up too
  • When study hours increase, exam scores increase
  • When COVID cases increase, do deaths increase?

These are correlations — they tell us if two things move together!

📖 Simple Definition:

Correlation = A measure of how two variables move together.

Do they increase together? Decrease together? Or not connected at all?

The Correlation Coefficient: Your Relationship Score

The correlation coefficient is a number between -1 and +1 that tells you HOW STRONG the relationship is:

+1.0 = Perfect Positive Correlation

When one goes up, the other ALWAYS goes up by the same proportion.
Example: Distance traveled and fuel consumed (in a car at constant speed)

0.0 = No Correlation

The two variables are NOT related at all. Knowing one tells you NOTHING about the other.
Example: Shoe size and intelligence

-1.0 = Perfect Negative Correlation

When one goes up, the other ALWAYS goes down by the same proportion.
Example: Speed and time to destination (faster speed = less time)

Quick Interpretation Guide:

Coefficient Value     Interpretation
─────────────────     ───────────────────────────────
 +0.9 to +1.0         Very strong positive correlation
 +0.7 to +0.9         Strong positive correlation
 +0.5 to +0.7         Moderate positive correlation
 +0.3 to +0.5         Weak positive correlation
 -0.3 to +0.3         Little to no correlation
 -0.5 to -0.3         Weak negative correlation
 -0.7 to -0.5         Moderate negative correlation
 -0.9 to -0.7         Strong negative correlation
 -1.0 to -0.9         Very strong negative correlation

Critical Reminder: Correlation ≠ Causation! ⚠️

🚨 SUPER IMPORTANT:

Just because two things are correlated does NOT mean one causes the other!

Example: Ice cream sales and drowning deaths are correlated (both increase in summer). But ice cream doesn't CAUSE drowning! The real cause is summer weather affecting both.

Three possibilities when you see correlation:

  1. A causes B (cases cause deaths)
  2. B causes A (rarely, but possible)
  3. C causes both A and B (a third factor affects both)

Our COVID-19 Dataset: Perfect for Correlation Analysis!

Our dataset has:

  • 5,818 rows across countries and dates
  • 67 columns with many numeric variables
  • Perfect for finding relationships like:
    • Do more cases lead to more deaths?
    • Are vaccinations related to fewer cases?
    • Does population density affect case spread?

Step 1: Load and Prepare Data

import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns

# Load the COVID-19 dataset
df = pd.read_csv('coviddata.csv')

# Quick overview
print(f"Dataset shape: {df.shape}")
print(df.head())

# Check what columns we have
print("\nNumeric columns available:")
print(df.select_dtypes(include=[np.number]).columns.tolist())

Method 1: Calculate Correlation Between Two Variables

Let's start simple: What's the correlation between total_cases and total_deaths?

# Calculate correlation between two columns
correlation = df['total_cases'].corr(df['total_deaths'])
print(f"Correlation between cases and deaths: {correlation:.4f}")

Sample Output:

Correlation between cases and deaths: 0.9567

What does this mean?

  • Correlation of 0.96 is VERY STRONG positive correlation!
  • As total cases increase, total deaths increase proportionally
  • This makes sense — more cases typically lead to more deaths
  • But remember: correlation doesn't prove causation! (Though in this case, the causal link is clear)

Method 2: Correlation Matrix - See ALL Relationships!

Instead of checking pairs one by one, let's see correlations between ALL numeric columns at once!

# Select important columns for analysis
columns_of_interest = [
    'total_cases', 
    'total_deaths', 
    'total_vaccinations',
    'population',
    'population_density',
    'median_age',
    'gdp_per_capita'
]

# Create correlation matrix
correlation_matrix = df[columns_of_interest].corr()
print("Correlation Matrix:")
print(correlation_matrix)

Sample Output:

Correlation Matrix:
                      total_cases  total_deaths  total_vaccinations  population  population_density  median_age  gdp_per_capita
total_cases                 1.000         0.957               0.876       0.234               0.123       0.345           0.456
total_deaths                0.957         1.000               0.834       0.198               0.087       0.389           0.501
total_vaccinations          0.876         0.834               1.000       0.456               0.234       0.567           0.678
population                  0.234         0.198               0.456       1.000               0.145       0.089           0.123
population_density          0.123         0.087               0.234       0.145               1.000       0.234           0.345
median_age                  0.345         0.389               0.567       0.089               0.234       1.000           0.789
gdp_per_capita              0.456         0.501               0.678       0.123               0.345       0.789           1.000

How to read this table:

  • Diagonal values are always 1.0 (a variable perfectly correlates with itself!)
  • The matrix is symmetric (correlation of A→B = B→A)
  • Look for high values (>0.7) or low values (<-0 .7="" strong="">)
  • Values near 0 mean weak or no relationship

Method 3: Visualize with a Heatmap

Numbers in a table are hard to see. Let's make a colorful heatmap!

import seaborn as sns
import matplotlib.pyplot as plt

# Create correlation matrix
columns_of_interest = [
    'total_cases', 
    'total_deaths', 
    'total_vaccinations',
    'new_cases',
    'new_deaths',
    'population_density',
    'median_age'
]

corr_matrix = df[columns_of_interest].corr()

# Create heatmap
plt.figure(figsize=(10, 8))
sns.heatmap(corr_matrix, 
            annot=True,           # Show numbers
            cmap='coolwarm',      # Color scheme (red=positive, blue=negative)
            center=0,             # Center color at 0
            fmt='.2f',            # 2 decimal places
            square=True,          # Square cells
            linewidths=1)         # Grid lines

plt.title('Correlation Heatmap - COVID-19 Data', fontsize=16, fontweight='bold')
plt.tight_layout()
plt.show()

Reading the heatmap:

  • Dark red = Strong positive correlation (closer to +1)
  • White = No correlation (around 0)
  • Dark blue = Strong negative correlation (closer to -1)
  • Numbers show exact correlation values

Pro Tip:

Heatmaps make patterns JUMP OUT at you! You can instantly spot which variables are strongly related without reading numbers.

Method 4: Find Strongest Correlations

Let's find which variables have the strongest relationships with total_deaths:

# Get all correlations with total_deaths
death_correlations = df.corr()['total_deaths'].sort_values(ascending=False)

print("Top 10 variables most correlated with total_deaths:")
print(death_correlations.head(10))

print("\nTop 10 variables LEAST correlated (most negative) with total_deaths:")
print(death_correlations.tail(10))

Sample Output:

Top 10 variables most correlated with total_deaths:
total_deaths                        1.000000
total_cases                         0.956789
total_deaths_per_million            0.887654
people_fully_vaccinated             0.834567
total_vaccinations                  0.823456
new_deaths_smoothed                 0.765432
gdp_per_capita                      0.501234
median_age                          0.389012
life_expectancy                     0.345678
human_development_index             0.312345

Insights:

  • total_cases (0.96) — Very strong! More cases = more deaths
  • people_fully_vaccinated (0.83) — Strong positive! Wait, shouldn't vaccinations REDUCE deaths? Not so fast! This shows countries with more deaths also vaccinated more people (responding to the crisis)
  • gdp_per_capita (0.50) — Moderate positive. Wealthier countries might have better reporting

Method 5: Correlation by Group (Advanced!)

Let's calculate correlations separately for each continent:

# Function to get correlation for each continent
def get_continent_correlation(continent_name):
    continent_data = df[df['continent'] == continent_name]
    corr = continent_data['total_cases'].corr(continent_data['total_deaths'])
    return corr

# Get correlations for each continent
continents = df['continent'].dropna().unique()
print("Correlation between cases and deaths by continent:\n")

for continent in continents:
    corr = get_continent_correlation(continent)
    print(f"{continent:15} : {corr:.4f}")

Sample Output:

Correlation between cases and deaths by continent:

Europe          : 0.9823
North America   : 0.9756
South America   : 0.9612
Asia            : 0.9534
Africa          : 0.8945
Oceania         : 0.9201

Interesting finding: All continents show very strong positive correlation, but Africa's is slightly lower (0.89). This might indicate different reporting standards or healthcare responses.

Method 6: Different Types of Correlation

Pandas supports three types of correlation coefficients:

1. Pearson (default) - For linear relationships

Best when: Data is normally distributed and relationship is linear.
Measures: Strength of LINEAR relationship.

2. Spearman - For monotonic relationships

Best when: Relationship exists but not necessarily linear (curved but always increasing/decreasing).
Measures: How well the relationship can be described using a monotonic function.

3. Kendall - For ordinal data

Best when: Working with ranked data or small sample sizes.
Measures: Ordinal association between two variables.

# Compare all three methods
pearson_corr = df['total_cases'].corr(df['total_deaths'], method='pearson')
spearman_corr = df['total_cases'].corr(df['total_deaths'], method='spearman')
kendall_corr = df['total_cases'].corr(df['total_deaths'], method='kendall')

print("Correlation between total_cases and total_deaths:")
print(f"Pearson:  {pearson_corr:.4f}")
print(f"Spearman: {spearman_corr:.4f}")
print(f"Kendall:  {kendall_corr:.4f}")

Sample Output:

Correlation between total_cases and total_deaths:
Pearson:  0.9568
Spearman: 0.9834
Kendall:  0.9123

Interpretation: All three show strong positive correlation! When they're similar, it confirms the relationship is robust.

Method 7: Real-World Analysis - Vaccination Impact

Let's investigate: Do vaccinations reduce new cases?

# Filter data where vaccination data exists
df_vacc = df[df['people_fully_vaccinated'].notna() & df['new_cases'].notna()]

# Calculate correlation
vacc_cases_corr = df_vacc['people_fully_vaccinated'].corr(df_vacc['new_cases'])
print(f"Correlation between vaccinations and new cases: {vacc_cases_corr:.4f}")

# Let's also check vaccination rate vs cases
# Create a vaccination rate column
df_vacc['vacc_rate'] = (df_vacc['people_fully_vaccinated'] / df_vacc['population']) * 100

# Correlation between vaccination rate and new cases
rate_corr = df_vacc['vacc_rate'].corr(df_vacc['new_cases'])
print(f"Correlation between vaccination RATE and new cases: {rate_corr:.4f}")

Sample Output:

Correlation between vaccinations and new cases: 0.2345
Correlation between vaccination RATE and new cases: -0.1234

Fascinating insights:

  • Raw vaccination numbers show weak positive correlation (0.23) — larger countries with more cases also vaccinated more people
  • Vaccination rate shows weak negative correlation (-0.12) — suggesting higher vaccination rates MIGHT reduce cases slightly
  • The correlation is weaker than expected because of time lags and other factors

Method 8: Handling Missing Values in Correlation

Pandas automatically handles NaN (missing) values:

# Check how many values are missing
print("Missing values:")
print(f"total_cases: {df['total_cases'].isna().sum()}")
print(f"total_deaths: {df['total_deaths'].isna().sum()}")
print(f"total_vaccinations: {df['total_vaccinations'].isna().sum()}")

# Pandas .corr() automatically uses pairwise complete observations
# It only uses rows where BOTH columns have values

# Calculate correlation
corr_with_missing = df['total_cases'].corr(df['total_vaccinations'])
print(f"\nCorrelation (NaN automatically excluded): {corr_with_missing:.4f}")

# Count how many pairs were actually used
valid_pairs = df[['total_cases', 'total_vaccinations']].dropna()
print(f"Valid pairs used: {len(valid_pairs)} out of {len(df)} total rows")

💡 Important:

If you have 90% missing data, your correlation might not be reliable! Always check how many valid pairs were used.

Real-World Complete Analysis

Let's put it all together with a comprehensive correlation analysis:

import pandas as pd
import numpy as np

# Load data
df = pd.read_csv('coviddata.csv')

print("="*70)
print("COVID-19 CORRELATION ANALYSIS - FINDING RELATIONSHIPS")
print("="*70)

# 1. Overall dataset correlation summary
print("\n1. KEY CORRELATIONS WITH TOTAL_DEATHS:")
death_corr = df.corr()['total_deaths'].sort_values(ascending=False)
print(death_corr.head(10))

# 2. Case-Death Relationship
print("\n2. CASES vs DEATHS ANALYSIS:")
cases_deaths_corr = df['total_cases'].corr(df['total_deaths'])
print(f"   Correlation: {cases_deaths_corr:.4f}")
if cases_deaths_corr > 0.7:
    print("   ✓ Very strong positive relationship")
elif cases_deaths_corr > 0.5:
    print("   ✓ Moderate positive relationship")

# 3. Vaccination Impact
print("\n3. VACCINATION IMPACT:")
df_vacc = df[df['total_vaccinations'].notna()]
if len(df_vacc) > 100:
    vacc_deaths = df_vacc['total_vaccinations'].corr(df_vacc['total_deaths'])
    print(f"   Vaccinations vs Deaths: {vacc_deaths:.4f}")
    print(f"   Data points analyzed: {len(df_vacc)}")

# 4. Demographic Factors
print("\n4. DEMOGRAPHIC CORRELATIONS WITH CASES:")
demographic_cols = ['population', 'population_density', 'median_age', 'gdp_per_capita']
for col in demographic_cols:
    if col in df.columns:
        corr = df['total_cases'].corr(df[col])
        print(f"   {col:20}: {corr:6.4f}")

# 5. Correlation Matrix for key variables
print("\n5. CORRELATION MATRIX (Key Variables):")
key_vars = ['total_cases', 'total_deaths', 'new_cases', 'new_deaths']
key_vars_existing = [col for col in key_vars if col in df.columns]
corr_matrix = df[key_vars_existing].corr()
print(corr_matrix)

# 6. Strongest and Weakest Correlations
print("\n6. INSIGHTS:")
all_corr = df.corr()
# Find strongest positive correlation (excluding diagonal)
mask = np.triu(np.ones_like(all_corr, dtype=bool))
masked_corr = all_corr.mask(mask)
strongest = masked_corr.abs().max().max()
print(f"   Strongest correlation found: {strongest:.4f}")

print("\n" + "="*70)
print("Remember: Correlation does NOT prove causation!")
print("="*70)

Practical Tips for Correlation Analysis 💡

DO:

  • Always visualize with scatter plots or heatmaps
  • Check for outliers — they can distort correlations
  • Use appropriate correlation type (Pearson, Spearman, Kendall)
  • Look at the data before trusting the number
  • Consider time lags (vaccination effects take time!)

DON'T:

  • Assume correlation means causation!
  • Ignore missing data — check how many values were used
  • Use Pearson for non-linear relationships
  • Compare correlations from different sample sizes
  • Ignore context and domain knowledge

Quick Reference Cheat Sheet

# Two variables correlation
df['col1'].corr(df['col2'])                      # Default: Pearson
df['col1'].corr(df['col2'], method='spearman')   # Spearman
df['col1'].corr(df['col2'], method='kendall')    # Kendall

# Correlation matrix (all numeric columns)
df.corr()                                        # All columns
df[['col1', 'col2', 'col3']].corr()             # Specific columns

# Get correlations for one variable
df.corr()['target_column']                       # All correlations with target
df.corr()['target_column'].sort_values()         # Sorted

# Correlation with groupby
df.groupby('group')['col1'].corr(df.groupby('group')['col2'])

# Handling missing values
df[['col1', 'col2']].dropna().corr()            # Explicitly drop NaN
# Note: .corr() automatically handles NaN by default

# Visualization
import seaborn as sns
sns.heatmap(df.corr(), annot=True, cmap='coolwarm')

Common Mistakes to Avoid ⚠️

  • Correlation ≠ Causation! Never assume one variable causes another just because they're correlated
  • Don't ignore outliers! One extreme value can create or destroy correlations
  • Don't use Pearson for everything! If data isn't linear, use Spearman or Kendall
  • Don't forget time factors! COVID effects have time lags — correlation at one moment might not tell the full story
  • Don't compare correlations across different datasets! Sample size and data quality matter
  • Don't trust weak correlations! Values between -0.3 and +0.3 are often not meaningful

Comments