Skip to main content

Correlation in Pandas DataFrames

Calculating read time…

C orrelation in Pandas DataFrames

If you're just starting with pandas and data analysis, correlation is one of the first things you'll want to understand. It helps you answer questions like:

  • Do taller people tend to weigh more?
  • Does more study time lead to higher exam scores?
  • Are temperature and ice cream sales related?

Correlation measures how strongly two numeric columns are related and in which direction.

Our Sample DataFrame

Let's create a simple DataFrame with numeric data that has obvious relationships:

import pandas as pd

data = {
    'Height_cm': [150, 160, 170, 180, 190],
    'Weight_kg': [50,  60,  70,  80,  90],
    'Shoe_Size': [5,   6,   7,   8,   9],
    'Study_Hours': [1, 3, 2, 5, 4],
    'Exam_Score': [55, 75, 65, 90, 85]
}

df = pd.DataFrame(data)
print(df)

This will display:

Height_cm Weight_kg Shoe_Size Study_Hours Exam_Score
0 150 50 5 1 55
1 160 60 6 3 75
2 170 70 7 2 65
3 180 80 8 5 90
4 190 90 9 4 85

What Does Correlation Mean?

The correlation coefficient is a number between -1 and +1:

  • +1 — Perfect positive relationship: as one increases, the other increases exactly together.
  • Close to +1 — Strong positive relationship.
  • 0 — No linear relationship.
  • Close to -1 — Strong negative relationship: as one increases, the other decreases.
  • -1 — Perfect negative relationship.

Important: Correlation only measures linear relationships. It does not mean causation!

How to Calculate Correlation in Pandas

Use the df.corr() method. It computes correlation between all pairs of numeric columns and returns a correlation matrix.

# Default: Pearson correlation
correlation_matrix = df.corr()
print(correlation_matrix.round(2))

This will display (rounded to 2 decimals):

Height_cm Weight_kg Shoe_Size Study_Hours Exam_Score
Height_cm 1.00 1.00 1.00 0.79 0.79
Weight_kg 1.00 1.00 1.00 0.79 0.79
Shoe_Size 1.00 1.00 1.00 0.79 0.79
Study_Hours 0.79 0.79 0.79 1.00 0.99
Exam_Score 0.79 0.79 0.79 0.99 1.00

What we see:

  • Height, Weight, and Shoe_Size are perfectly correlated (1.00) with each other.
  • Study_Hours and Exam_Score are very strongly correlated (0.99).
  • The diagonal is always 1.00 (a variable is perfectly correlated with itself).

Different Types of Correlation

df.corr() has a method parameter:

  • 'pearson' (default): Standard method for linear relationships.
  • 'spearman': Rank-based, good for monotonic relationships (even non-linear).
  • 'kendall': Another rank-based method, useful for small datasets.
# Spearman example
df.corr(method='spearman')

Key Tips for Beginners

  • Only numeric columns are included — text columns are ignored.
  • Check for missing values first with df.isna().sum().
  • High correlation ≠ causation (example: ice cream sales and shark attacks both rise in summer).
  • Use df.corr().round(2) for cleaner output.

Happy Coding ! 🐼

Comments