C orrelation in Pandas DataFrames
If you're just starting with pandas and data analysis, correlation is one of the first things you'll want to understand. It helps you answer questions like:
- Do taller people tend to weigh more?
- Does more study time lead to higher exam scores?
- Are temperature and ice cream sales related?
Correlation measures how strongly two numeric columns are related and in which direction.
Our Sample DataFrame
Let's create a simple DataFrame with numeric data that has obvious relationships:
import pandas as pd
data = {
'Height_cm': [150, 160, 170, 180, 190],
'Weight_kg': [50, 60, 70, 80, 90],
'Shoe_Size': [5, 6, 7, 8, 9],
'Study_Hours': [1, 3, 2, 5, 4],
'Exam_Score': [55, 75, 65, 90, 85]
}
df = pd.DataFrame(data)
print(df)
This will display:
| Height_cm | Weight_kg | Shoe_Size | Study_Hours | Exam_Score | |
|---|---|---|---|---|---|
| 0 | 150 | 50 | 5 | 1 | 55 |
| 1 | 160 | 60 | 6 | 3 | 75 |
| 2 | 170 | 70 | 7 | 2 | 65 |
| 3 | 180 | 80 | 8 | 5 | 90 |
| 4 | 190 | 90 | 9 | 4 | 85 |
What Does Correlation Mean?
The correlation coefficient is a number between -1 and +1:
- +1 — Perfect positive relationship: as one increases, the other increases exactly together.
- Close to +1 — Strong positive relationship.
- 0 — No linear relationship.
- Close to -1 — Strong negative relationship: as one increases, the other decreases.
- -1 — Perfect negative relationship.
Important: Correlation only measures linear relationships. It does not mean causation!
How to Calculate Correlation in Pandas
Use the df.corr() method. It computes correlation between all pairs of numeric columns and returns a correlation matrix.
# Default: Pearson correlation correlation_matrix = df.corr() print(correlation_matrix.round(2))
This will display (rounded to 2 decimals):
| Height_cm | Weight_kg | Shoe_Size | Study_Hours | Exam_Score | |
|---|---|---|---|---|---|
| Height_cm | 1.00 | 1.00 | 1.00 | 0.79 | 0.79 |
| Weight_kg | 1.00 | 1.00 | 1.00 | 0.79 | 0.79 |
| Shoe_Size | 1.00 | 1.00 | 1.00 | 0.79 | 0.79 |
| Study_Hours | 0.79 | 0.79 | 0.79 | 1.00 | 0.99 |
| Exam_Score | 0.79 | 0.79 | 0.79 | 0.99 | 1.00 |
What we see:
- Height, Weight, and Shoe_Size are perfectly correlated (1.00) with each other.
- Study_Hours and Exam_Score are very strongly correlated (0.99).
- The diagonal is always 1.00 (a variable is perfectly correlated with itself).
Different Types of Correlation
df.corr() has a method parameter:
- 'pearson' (default): Standard method for linear relationships.
- 'spearman': Rank-based, good for monotonic relationships (even non-linear).
- 'kendall': Another rank-based method, useful for small datasets.
# Spearman example df.corr(method='spearman')
Key Tips for Beginners
- Only numeric columns are included — text columns are ignored.
- Check for missing values first with
df.isna().sum(). - High correlation ≠ causation (example: ice cream sales and shark attacks both rise in summer).
- Use
df.corr().round(2)for cleaner output.
Happy Coding ! 🐼
Comments
Post a Comment