Master Histograms: The Ultimate Guide to Understanding Data Distribution with Amsterdam House Prices
📌 Dataset Source: To follow along with this tutorial, download the Amsterdam House Prices Dataset from Kaggle. You can access it directly here:
https://www.kaggle.com/datasets/thomasnibb/amsterdam-house-price-prediction
"The histogram is the Swiss Army knife of data visualization - simple yet incredibly powerful for revealing what's really happening in your data."
What is a Histogram? (The Distribution Detective!)
A histogram is a special type of bar chart that shows the distribution of numerical data. It answers questions like:
- 📊 Where do most values cluster? (What's the typical house price?)
- 📈 How spread out are the values? (Are prices similar or vary widely?)
- 🔍 Are there unusual patterns? (Multiple price peaks? Gaps?)
- 📏 What's the shape of the data? (Symmetric? Skewed?)
📖 Simple Analogy:
Imagine sorting Amsterdam houses into price "buckets":
• Bucket 1: €100,000 - €200,000 (10 houses)
• Bucket 2: €200,000 - €300,000 (25 houses)
• Bucket 3: €300,000 - €400,000 (50 houses)
A histogram is simply a bar chart of these bucket counts!
Essential Matplotlib Terms - Explained Simply
Before we dive into code, let's understand the key terms you'll encounter. These are the building blocks of every matplotlib visualization:
1. plt.figure() - Your Drawing Canvas
What it is: Creates a new figure (window/canvas) for your plot.
Analogy: Like getting a blank sheet of paper before you start drawing.
Why it matters: Without a figure, you have nowhere to draw your plot!
# Creating a figure - your drawing canvas plt.figure(figsize=(10, 6)) # 10 inches wide, 6 inches tall # Now you have a blank canvas to work on!
2. figsize - Setting Your Canvas Size
What it is: Controls the width and height of your figure in inches.
Format: figsize=(width, height)
Common sizes:
• (10, 6) - Standard report size
• (12, 8) - Large presentation size
• (8, 8) - Square plot
• (16, 4) - Wide timeline
plt.figure(figsize=(12, 8)) # 12 inches wide, 8 inches tall # This creates a larger canvas for more detailed plots
3. plt.hist() - The Histogram Function
What it is: The main function that creates histograms.
Basic syntax: plt.hist(data, bins=10)
Key parameters:
• data: Your numerical data (list, array, or pandas Series)
• bins: Number of buckets/bars (MOST IMPORTANT!)
• color: Bar color ('blue', 'red', '#FF5733')
• edgecolor: Color of bar edges ('black', 'white')
• alpha: Transparency (0=invisible, 1=solid)
• density: If True, shows probability instead of count
# Basic histogram plt.hist(house_prices, bins=20, color='blue', edgecolor='black') # This creates 20 blue bars with black edges
4. bins - The Secret Ingredient
What it is: Number of bars/buckets in your histogram.
Why it matters: Bins control how detailed or smooth your histogram looks!
Different ways to specify bins:
• bins=10 - Exactly 10 bars
• bins='auto' - Let matplotlib decide
• bins=30 - Good default for most data
• bins=[0, 100, 200, 300] - Custom bin edges
• bins=np.arange(0, 1000, 50) - Bins every 50 units
🎯 Bin Selection Rule:
• Too few bins (5): Lose detail, oversimplify
• Good bins (30): Show patterns clearly
• Too many bins (100): Too noisy, hard to interpret
Start with 30 bins, then adjust!
5. plt.subplots() - Multiple Plots in One Figure
What it is: Creates a grid of multiple plots (subplots).
Format: fig, axes = plt.subplots(rows, columns)
Example: plt.subplots(2, 3) creates 2 rows × 3 columns = 6 plots
Returns:
• fig: The overall figure
• axes: Array of individual plot areas (axes)
Use when: Comparing multiple datasets or views
# Create 2 rows, 2 columns of plots (4 total) fig, axes = plt.subplots(2, 2, figsize=(12, 10)) # Access each subplot individually axes[0, 0].hist(data1, bins=20) # Top-left plot axes[0, 1].hist(data2, bins=20) # Top-right plot axes[1, 0].hist(data3, bins=20) # Bottom-left plot axes[1, 1].hist(data4, bins=20) # Bottom-right plot
6. plt.suptitle() - Main Title for Multiple Plots
What it is: Adds a main title to a figure with multiple subplots.
Difference from plt.title():
• plt.title() - Titles an individual plot
• plt.suptitle() - Titles the ENTIRE figure (all subplots)
Position: Appears at the top of the entire figure
Use with: plt.subplots() when you have multiple plots
fig, axes = plt.subplots(2, 2, figsize=(12, 10))
fig.suptitle('Amsterdam House Price Analysis by Neighborhood',
fontsize=16, fontweight='bold')
# This title appears above ALL 4 subplots
7. plt.xlabel() / plt.ylabel() - Axis Labels
What they are: Add descriptive labels to X and Y axes.
Why they matter: Without labels, no one knows what they're looking at!
Key parameters:
• label: The text to display
• fontsize: Text size (12 is standard)
• fontweight: 'normal', 'bold'
• labelpad: Space between label and axis
plt.xlabel('House Price (Euros)', fontsize=12, fontweight='bold')
plt.ylabel('Number of Houses', fontsize=12, fontweight='bold')
# Now viewers know what the axes represent!
8. plt.axvline() / plt.axhline() - Reference Lines
What they are: Add vertical or horizontal reference lines.
Common uses:
• Show mean/median values
• Indicate thresholds
• Highlight specific values
Key parameters:
• x or y: Position of the line
• color: Line color
• linestyle: '--' (dashed), '-.' (dash-dot), ':' (dotted)
• linewidth: Thickness of line
# Add mean line (red dashed)
mean_price = df['Price'].mean()
plt.axvline(mean_price, color='red', linestyle='--', linewidth=2,
label=f'Mean: €{mean_price:,.0f}')
# Add median line (green dash-dot)
median_price = df['Price'].median()
plt.axvline(median_price, color='green', linestyle='-.', linewidth=2,
label=f'Median: €{median_price:,.0f}')
9. plt.legend() - The Plot Legend
What it is: Creates a legend explaining colors/lines/symbols.
When to use: When you have multiple data series or reference lines.
Key parameters:
• loc: Position ('upper right', 'lower left', 'best')
• fontsize: Text size in legend
• frameon: Show box around legend (True/False)
Requires: label= parameter in your plot commands
plt.hist(data1, bins=20, color='blue', label='Apartments')
plt.hist(data2, bins=20, color='red', label='Houses')
plt.axvline(mean_price, color='green', linestyle='--', linewidth=2,
label='Average Price')
plt.legend(loc='upper right', fontsize=11)
# Now viewers know what each color/line represents!
10. plt.tight_layout() - The Space Manager
What it is: Automatically adjusts spacing between plots.
Why it matters: Prevents labels from overlapping or getting cut off.
When to use: ALWAYS use with multiple subplots!
What it fixes:
• Overlapping titles/labels
• Plots too close together
• Labels cut off at edges
Simple rule: Call plt.tight_layout() before plt.show()
fig, axes = plt.subplots(2, 2, figsize=(12, 10)) # ... create your plots ... plt.tight_layout() # Fixes all spacing issues! plt.show()
11. plt.show() - Display Your Masterpiece
What it is: Displays the final plot.
When to use: ALWAYS at the end of your plotting code!
What happens: Opens a window with your visualization
In Jupyter: Shows the plot directly in the notebook
Important: Nothing appears until you call plt.show()!
# All your plotting code goes here...
plt.figure(figsize=(10, 6))
plt.hist(data, bins=30)
plt.xlabel('Price')
plt.ylabel('Count')
plt.title('My Histogram')
plt.show() # THIS MAKES IT APPEAR!
The Complete Plotting Workflow
Here's the standard sequence for creating ANY matplotlib plot:
# STEP 1: Create figure (your canvas)
plt.figure(figsize=(12, 8))
# STEP 2: Create the plot (histogram, scatter, etc.)
plt.hist(data, bins=30, color='blue', edgecolor='black')
# STEP 3: Add labels and title
plt.xlabel('X-axis Label', fontsize=12)
plt.ylabel('Y-axis Label', fontsize=12)
plt.title('Descriptive Title', fontsize=14, fontweight='bold')
# STEP 4: Add reference lines (optional)
plt.axvline(mean_value, color='red', linestyle='--', label='Mean')
# STEP 5: Add legend if needed
plt.legend()
# STEP 6: Adjust layout
plt.tight_layout()
# STEP 7: Show the plot
plt.show()
Loading and Exploring Amsterdam House Data
Now let's load our dataset and understand what we're working with:
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
# Load the Amsterdam house prices dataset
df = pd.read_csv('HousingPricesAmsterdam.csv')
print("🏠 AMSTERDAM HOUSE PRICES DATASET")
print("="*50)
print(f"Dataset shape: {df.shape}") # (rows, columns)
print(f"\nFirst 3 houses:")
print(df.head(3))
print(f"\n📊 Column Information:")
print(df.info())
print(f"\n💰 Price Statistics (in Euros):")
print(f"Minimum price: €{df['Price'].min():,.0f}")
print(f"Maximum price: €{df['Price'].max():,.0f}")
print(f"Average price: €{df['Price'].mean():,.0f}")
print(f"Median price: €{df['Price'].median():,.0f}")
# Check for missing values
print(f"\n🔍 Missing Values Check:")
print(df.isnull().sum())
💡 First Insight:
Notice something interesting? The average price (€583,574) is higher than the median price (€495,000). This suggests our data might be right-skewed (a few very expensive houses pulling the average up). Let's confirm this with our first histogram!
Method 1: Your First Histogram (The Absolute Basics)
Let's create the simplest possible histogram to see house price distribution:
# METHOD 1: Basic histogram with 20 bins
plt.figure(figsize=(10, 6)) # Step 1: Create canvas (10x6 inches)
# Step 2: Create histogram
plt.hist(df['Price'], bins=20) # 20 bars/buckets
# Step 3: Add labels and title
plt.xlabel('House Price (Euros)', fontsize=12)
plt.ylabel('Number of Houses', fontsize=12)
plt.title('Amsterdam House Price Distribution', fontsize=14, fontweight='bold')
# Step 4: Add grid for readability
plt.grid(True, alpha=0.3) # alpha controls transparency (0=invisible, 1=solid)
# Step 5: Format x-axis to show Euro symbols
plt.gca().xaxis.set_major_formatter(
plt.FuncFormatter(lambda x, p: f'€{x:,.0f}'))
# Step 6: Show the plot
plt.tight_layout() # Adjusts spacing
plt.show()
# Let's also calculate what the histogram shows us
print("📈 WHAT THIS HISTOGRAM TELLS US:")
print("="*40)
# Calculate basic statistics
mean_price = df['Price'].mean()
median_price = df['Price'].median()
std_price = df['Price'].std()
print(f"1. Average Price: €{mean_price:,.0f}")
print(f"2. Median Price: €{median_price:,.0f}")
print(f"3. Price Range: €{df['Price'].min():,.0f} to €{df['Price'].max():,.0f}")
print(f"4. Standard Deviation: €{std_price:,.0f} (measure of spread)")
print(f"5. Shape: The histogram is RIGHT-SKEWED (tail extends to the right)")
print(f"6. Most Common Price Range: Around €300,000 - €500,000")
✅ What You Just Learned:
plt.figure(figsize=(10, 6))- Creates canvasplt.hist(data, bins=20)- Creates histogram with 20 barsplt.xlabel()/plt.ylabel()- Adds axis labelsplt.title()- Adds plot titleplt.grid()- Adds grid linesplt.tight_layout()- Adjusts spacingplt.show()- Displays the plot
Method 2: Understanding Bins - The Most Important Parameter!
Bins are the SECRET to great histograms! Let's see how different bin choices change the story:
# METHOD 2: Comparing different bin sizes
fig, axes = plt.subplots(2, 2, figsize=(14, 10)) # 2 rows, 2 columns of plots
fig.suptitle('How Bin Size Changes Your Histogram Story',
fontsize=16, fontweight='bold')
# Plot 1: Too few bins (5) - Top-left
axes[0, 0].hist(df['Price'], bins=5, color='lightblue', edgecolor='black')
axes[0, 0].set_title('Too Few Bins (5 bins)', fontsize=12)
axes[0, 0].set_xlabel('Price (Euros)')
axes[0, 0].set_ylabel('Count')
axes[0, 0].grid(True, alpha=0.3)
axes[0, 0].xaxis.set_major_formatter(
plt.FuncFormatter(lambda x, p: f'€{x:,.0f}'))
# Plot 2: Default bins ('auto') - Top-right
axes[0, 1].hist(df['Price'], bins='auto', color='lightgreen', edgecolor='black')
axes[0, 1].set_title("Auto Bins (Matplotlib's Choice)", fontsize=12)
axes[0, 1].set_xlabel('Price (Euros)')
axes[0, 1].set_ylabel('Count')
axes[0, 1].grid(True, alpha=0.3)
axes[0, 1].xaxis.set_major_formatter(
plt.FuncFormatter(lambda x, p: f'€{x:,.0f}'))
# Plot 3: Good number of bins (30) - Bottom-left
axes[1, 0].hist(df['Price'], bins=30, color='lightcoral', edgecolor='black')
axes[1, 0].set_title('Good Detail (30 bins)', fontsize=12)
axes[1, 0].set_xlabel('Price (Euros)')
axes[1, 0].set_ylabel('Count')
axes[1, 0].grid(True, alpha=0.3)
axes[1, 0].xaxis.set_major_formatter(
plt.FuncFormatter(lambda x, p: f'€{x:,.0f}'))
# Plot 4: Too many bins (100) - Bottom-right
axes[1, 1].hist(df['Price'], bins=100, color='lightyellow', edgecolor='black')
axes[1, 1].set_title('Too Many Bins (100 bins - noisy)', fontsize=12)
axes[1, 1].set_xlabel('Price (Euros)')
axes[1, 1].set_ylabel('Count')
axes[1, 1].grid(True, alpha=0.3)
axes[1, 1].xaxis.set_major_formatter(
plt.FuncFormatter(lambda x, p: f'€{x:,.0f}'))
plt.tight_layout() # Fix spacing
plt.show() # Display all 4 plots
print("🧮 BIN SELECTION RULES (Different Methods):")
print("="*45)
# Rule 1: Square root rule
sqrt_bins = int(np.sqrt(len(df['Price'])))
print(f"1. Square Root Rule: {sqrt_bins} bins (√n = √{len(df['Price'])})")
# Rule 2: Sturges' formula
sturges_bins = int(np.ceil(np.log2(len(df['Price'])) + 1))
print(f"2. Sturges' Formula: {sturges_bins} bins")
# Rule 3: Freedman-Diaconis Rule (more robust)
q75, q25 = np.percentile(df['Price'], [75, 25])
iqr = q75 - q25
bin_width = 2 * iqr / (len(df['Price']) ** (1/3))
fd_bins = int(np.ceil((df['Price'].max() - df['Price'].min()) / bin_width))
print(f"3. Freedman-Diaconis: {fd_bins} bins")
print(f"\n💡 Recommendation: Start with 30 bins for this dataset, then adjust!")
🎯 Bin Selection Rule of Thumb:
Too few bins (5): Lose detail, oversimplify
Good bins (30): Shows patterns clearly
Too many bins (100): Too noisy, hard to see patterns
Start with √n (square root of data points) or 30 bins, then adjust!
Method 3: Professional Customizations
Let's add professional touches to make our histogram publication-ready:
# METHOD 3: Professional histogram with customizations
plt.figure(figsize=(14, 8)) # Larger canvas for detailed plot
# Create histogram with professional settings
n, bins, patches = plt.hist(df['Price'], bins=30, # 30 bars
color='#2E86AB', # Professional blue color
alpha=0.85, # 85% opacity (slight transparency)
edgecolor='black', # Black edges on bars
linewidth=1.2, # Edge thickness
density=False) # Show count, not probability
# Add mean and median lines
mean_price = df['Price'].mean()
median_price = df['Price'].median()
plt.axvline(mean_price, color='#E63946', linestyle='--', linewidth=2.5,
label=f'Mean: €{mean_price:,.0f}')
plt.axvline(median_price, color='#F4A261', linestyle='-.', linewidth=2.5,
label=f'Median: €{median_price:,.0f}')
# Add labels with professional formatting
plt.xlabel('House Price (Euros)', fontsize=13, fontweight='bold', labelpad=10)
plt.ylabel('Number of Houses', fontsize=13, fontweight='bold', labelpad=10)
plt.title('Amsterdam House Price Distribution 2023',
fontsize=16, fontweight='bold', pad=20)
# Add text box with statistics
stats_text = f"""Dataset Statistics:
• Total Houses: {len(df):,}
• Price Range: €{df['Price'].min():,.0f} - €{df['Price'].max():,.0f}
• Average Price: €{mean_price:,.0f}
• Median Price: €{median_price:,.0f}
• Standard Deviation: €{df['Price'].std():,.0f}
• Most Common Range: €300K - €500K"""
plt.text(0.02, 0.98, stats_text, transform=plt.gca().transAxes,
fontsize=10, verticalalignment='top',
bbox=dict(boxstyle='round', facecolor='wheat', alpha=0.8))
# Format axes professionally (show €K for thousands)
plt.gca().xaxis.set_major_formatter(
plt.FuncFormatter(lambda x, p: f'€{x/1000:.0f}K'))
# Add grid (subtle)
plt.grid(True, alpha=0.2, linestyle='--')
# Add legend
plt.legend(loc='upper right', fontsize=11)
# Adjust layout
plt.tight_layout()
plt.show()
print("📊 PROFESSIONAL FEATURES ADDED:")
print("="*45)
print("1. Professional color scheme (#2E86AB = nice blue)")
print("2. alpha=0.85 for slight transparency")
print("3. edgecolor='black' for bar separation")
print("4. Mean/median lines with different linestyles")
print("5. Text box with key statistics")
print("6. Formatted axes (€K instead of €)")
print("7. Subtle grid with alpha=0.2")
print("8. Bold labels with labelpad for spacing")
Method 4: Density Histogram (Probability View)
Sometimes you want to see probabilities instead of counts. This is especially useful when comparing datasets of different sizes:
# METHOD 4: Density histogram (area = 1)
plt.figure(figsize=(14, 8))
# Density histogram - area under curve = 1
plt.hist(df['Price'], bins=30, density=True, # density=True is key!
color='#6A0572', alpha=0.7, edgecolor='black')
# Add KDE (Kernel Density Estimate) for smooth curve
sns.kdeplot(df['Price'], color='#FF6B6B', linewidth=3,
label='KDE Smooth Curve')
# Calculate and plot normal distribution for comparison
from scipy import stats
mean, std = df['Price'].mean(), df['Price'].std()
x = np.linspace(df['Price'].min(), df['Price'].max(), 1000)
pdf = stats.norm.pdf(x, mean, std)
plt.plot(x, pdf, 'g--', linewidth=2.5, alpha=0.8,
label='Normal Distribution')
# Labels and title
plt.xlabel('House Price (Euros)', fontsize=12, fontweight='bold')
plt.ylabel('Probability Density', fontsize=12, fontweight='bold')
plt.title('Probability Density of Amsterdam House Prices',
fontsize=14, fontweight='bold')
plt.gca().xaxis.set_major_formatter(
plt.FuncFormatter(lambda x, p: f'€{x/1000:.0f}K'))
# Add area percentages
plt.text(0.02, 0.95, 'Price Distribution Analysis:',
transform=plt.gca().transAxes,
fontsize=11, fontweight='bold', verticalalignment='top')
# Calculate what percentage of houses fall in different price ranges
price_500k = len(df[df['Price'] <= 500000]) / len(df) * 100
price_1m = len(df[df['Price'] <= 1000000]) / len(df) * 100
price_2m = len(df[df['Price'] <= 2000000]) / len(df) * 100
plt.text(0.02, 0.88, f'• ≤ €500K: {price_500k:.1f}% of houses',
transform=plt.gca().transAxes, fontsize=10, verticalalignment='top')
plt.text(0.02, 0.83, f'• ≤ €1M: {price_1m:.1f}% of houses',
transform=plt.gca().transAxes, fontsize=10, verticalalignment='top')
plt.text(0.02, 0.78, f'• ≤ €2M: {price_2m:.1f}% of houses',
transform=plt.gca().transAxes, fontsize=10, verticalalignment='top')
plt.grid(True, alpha=0.2)
plt.legend(fontsize=11)
plt.tight_layout()
plt.show()
print("📈 DENSITY HISTOGRAM EXPLAINED:")
print("="*45)
print(f"1. {price_500k:.1f}% of Amsterdam houses cost ≤ €500,000")
print(f"2. {price_1m:.1f}% of Amsterdam houses cost ≤ €1,000,000")
print(f"3. {price_2m:.1f}% of Amsterdam houses cost ≤ €2,000,000")
print(f"4. Only {100 - price_2m:.1f}% of houses cost > €2,000,000")
print("\n💡 Key Differences from Regular Histogram:")
print("• Y-axis shows PROBABILITY DENSITY, not count")
print("• Area under histogram = 1 (100% probability)")
print("• KDE curve shows SMOOTHED probability distribution")
print("• Green dashed line shows what a NORMAL distribution would look like")
print("• Our data is NOT normal - it's right-skewed with a long tail!")
Method 5: Cumulative Histogram (Wealth Distribution)
Cumulative histograms answer questions like: "What percentage of houses cost less than X?" Perfect for understanding wealth distribution!
# METHOD 5: Cumulative histogram
fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(16, 6))
# Regular histogram for comparison
ax1.hist(df['Price'], bins=30, color='#1E88E5', edgecolor='black')
ax1.set_xlabel('House Price (Euros)', fontsize=11)
ax1.set_ylabel('Number of Houses', fontsize=11)
ax1.set_title('Regular Histogram (Counts)', fontsize=13, fontweight='bold')
ax1.grid(True, alpha=0.3)
ax1.xaxis.set_major_formatter(
plt.FuncFormatter(lambda x, p: f'€{x/1000:.0f}K'))
# Add some threshold lines
thresholds = [300000, 500000, 1000000]
colors = ['green', 'orange', 'red']
labels = ['€300K', '€500K', '€1M']
for threshold, color, label in zip(thresholds, colors, labels):
ax1.axvline(threshold, color=color, linestyle='--', alpha=0.7, linewidth=1.5)
ax1.text(threshold, ax1.get_ylim()[1]*0.9, label,
color=color, fontsize=10, ha='center')
# CUMULATIVE histogram (cumulative=True is key!)
ax2.hist(df['Price'], bins=30, cumulative=True, # cumulative=True!
color='#D81B60', edgecolor='black', linewidth=1.5)
ax2.set_xlabel('House Price (Euros)', fontsize=11)
ax2.set_ylabel('Cumulative Number of Houses', fontsize=11)
ax2.set_title('Cumulative Histogram', fontsize=13, fontweight='bold')
ax2.grid(True, alpha=0.3)
ax2.xaxis.set_major_formatter(
plt.FuncFormatter(lambda x, p: f'€{x/1000:.0f}K'))
# Add percentage scale on right side
ax2_percent = ax2.twinx()
ax2_percent.set_ylabel('Cumulative Percentage', fontsize=11, color='#004D40')
ax2_percent.set_ylim(0, len(df['Price']))
ax2_percent.set_yticks(np.arange(0, len(df['Price'])+1, len(df['Price'])/5))
ax2_percent.set_yticklabels([f'{p}%' for p in range(0, 101, 20)])
ax2_percent.tick_params(axis='y', labelcolor='#004D40')
# Add threshold lines and annotations
for threshold, color, label in zip(thresholds, colors, labels):
ax2.axvline(threshold, color=color, linestyle='--', alpha=0.7, linewidth=1.5)
# Calculate what percentage of houses are below this threshold
pct_below = (df['Price'] <= threshold).sum() / len(df) * 100
ax2.text(threshold, (df['Price'] <= threshold).sum(), f'{pct_below:.1f}%',
color=color, fontsize=10, ha='center', va='bottom', fontweight='bold')
plt.tight_layout()
plt.show()
print("💰 CUMULATIVE HISTOGRAM - WEALTH DISTRIBUTION ANALYSIS:")
print("="*55)
print("This answers: 'What percentage of houses cost LESS than X?'")
print("\nKey Thresholds in Amsterdam:")
print("-"*40)
for threshold in thresholds:
pct_below = (df['Price'] <= threshold).sum() / len(df) * 100
count_below = (df['Price'] <= threshold).sum()
print(f"€{threshold/1000:.0f}K Threshold:")
print(f" • {pct_below:.1f}% of houses ({count_below:,} houses) cost ≤ €{threshold/1000:.0f}K")
print(f" • {100-pct_below:.1f}% of houses ({len(df)-count_below:,} houses) cost > €{threshold/1000:.0f}K")
print()
print("💡 Insight: The cumulative histogram shows that house prices in Amsterdam")
print("follow a Pareto-like distribution - most houses are relatively affordable,")
print("but a small percentage are extremely expensive!")
Method 6: Comparing Distributions (By Number of Bedrooms)
Let's see how house prices differ based on number of bedrooms:
# METHOD 6: Comparing multiple distributions
fig, axes = plt.subplots(2, 3, figsize=(16, 10))
fig.suptitle('House Price Distribution by Number of Bedrooms',
fontsize=16, fontweight='bold')
# Get unique bedroom counts (sorted)
bedroom_counts = sorted(df['Bedroom'].dropna().unique())
# Colors for different bedroom counts
colors = ['#FF9999', '#66B2FF', '#99FF99', '#FFB366', '#FF99FF', '#FFD700']
for idx, bedroom_count in enumerate(bedroom_counts):
if idx >= 6: # Only show first 6
break
row = idx // 3 # Calculate row position
col = idx % 3 # Calculate column position
# Filter data for this bedroom count
bedroom_data = df[df['Bedroom'] == bedroom_count]['Price']
# Create histogram in the appropriate subplot
axes[row, col].hist(bedroom_data, bins=20,
color=colors[idx], edgecolor='black', alpha=0.8)
# Add mean line
mean_price = bedroom_data.mean()
axes[row, col].axvline(mean_price, color='red', linestyle='--', linewidth=2,
alpha=0.7, label=f'Mean: €{mean_price/1000:.0f}K')
# Add median line
median_price = bedroom_data.median()
axes[row, col].axvline(median_price, color='green', linestyle='-.', linewidth=2,
alpha=0.7, label=f'Median: €{median_price/1000:.0f}K')
# Customize subplot
axes[row, col].set_title(f'{bedroom_count} Bedroom{"s" if bedroom_count != 1 else ""}',
fontsize=12, fontweight='bold')
axes[row, col].set_xlabel('Price (Euros)')
axes[row, col].set_ylabel('Count')
axes[row, col].grid(True, alpha=0.3)
axes[row, col].legend(fontsize=9)
# Format x-axis
axes[row, col].xaxis.set_major_formatter(
plt.FuncFormatter(lambda x, p: f'€{x/1000:.0f}K'))
# Add statistics text
stats_text = f"""N = {len(bedroom_data):,}
Mean: €{mean_price/1000:.0f}K
Median: €{median_price/1000:.0f}K
Std: €{bedroom_data.std()/1000:.0f}K"""
axes[row, col].text(0.02, 0.98, stats_text,
transform=axes[row, col].transAxes,
fontsize=9, verticalalignment='top',
bbox=dict(boxstyle='round', facecolor='white', alpha=0.8))
# Remove empty subplots if any
for idx in range(len(bedroom_counts), 6):
row = idx // 3
col = idx % 3
fig.delaxes(axes[row, col])
plt.tight_layout()
plt.show()
print("🛏️ BEDROOM ANALYSIS - KEY FINDINGS:")
print("="*45)
# Create summary table
summary_data = []
for bedroom_count in bedroom_counts:
bedroom_data = df[df['Bedroom'] == bedroom_count]['Price']
if len(bedroom_data) > 0:
summary_data.append({
'Bedrooms': bedroom_count,
'Count': len(bedroom_data),
'Mean Price': f'€{bedroom_data.mean()/1000:.0f}K',
'Median Price': f'€{bedroom_data.median()/1000:.0f}K',
'Price Range': f'€{bedroom_data.min()/1000:.0f}K-€{bedroom_data.max()/1000:.0f}K'
})
# Convert to DataFrame for nice display
summary_df = pd.DataFrame(summary_data)
print(summary_df.to_string(index=False))
print("\n💡 Insights:")
print("1. More bedrooms generally mean higher prices (but not always linearly)")
print("2. 2-3 bedroom houses are most common in Amsterdam")
print("3. Price variability increases with bedroom count")
print("4. Single bedroom apartments have the narrowest price range")
print("5. Luxury 5+ bedroom houses show extreme price variation")
Histogram Cheat Sheet (Quick Reference)
# 1. BASIC SETUP
plt.figure(figsize=(10, 6)) # Create canvas (width, height in inches)
plt.hist(data, bins=30) # Basic histogram with 30 bins
plt.xlabel('X Label') # X-axis label
plt.ylabel('Y Label') # Y-axis label
plt.title('Plot Title') # Plot title
plt.grid(True, alpha=0.3) # Add grid with transparency
plt.tight_layout() # Fix spacing issues
plt.show() # Display plot
# 2. CUSTOMIZING HISTOGRAM
plt.hist(data, bins=30, color='blue', alpha=0.7, edgecolor='black')
# 3. DIFFERENT BIN OPTIONS
bins=30 # Fixed number of bins
bins='auto' # Let matplotlib decide
bins=[0, 100, 200, 300] # Custom bin edges
bins=np.arange(0, 1000, 50) # Bins from 0 to 1000, step 50
# 4. SPECIAL HISTOGRAMS
plt.hist(data, bins=30, density=True) # Density histogram
plt.hist(data, bins=30, cumulative=True) # Cumulative histogram
# 5. MULTIPLE PLOTS
fig, axes = plt.subplots(2, 2) # 2x2 grid of plots
axes[0, 0].hist(data1, bins=20) # Top-left plot
axes[0, 1].hist(data2, bins=20) # Top-right plot
fig.suptitle('Overall Title') # Title for all subplots
# 6. REFERENCE LINES
plt.axvline(mean_value, color='red', linestyle='--', label='Mean')
plt.axhline(threshold, color='green', linestyle='-', label='Threshold')
# 7. FORMATTING AXES
plt.gca().xaxis.set_major_formatter(plt.FuncFormatter(lambda x, p: f'€{x:,.0f}'))
plt.gca().xaxis.set_major_formatter(plt.FuncFormatter(lambda x, p: f'€{x/1000:.0f}K'))
# 8. ADDING TEXT
plt.text(x_position, y_position, 'Text here', fontsize=10)
# 9. LEGEND
plt.legend(loc='upper right', fontsize=10)
# 10. SAVING
plt.savefig('histogram.png', dpi=300, bbox_inches='tight')
Common Mistakes to Avoid ⚠️
❌ COMMON MISTAKES:
- Wrong bin size: Too few (hides patterns) or too many (creates noise)
- Forgetting plt.show(): Plot won't display without this!
- No labels/title: Viewers won't know what they're looking at
- Overlapping text: Forgetting plt.tight_layout() with subplots
- Using default colors: Hard to read, not publication-quality
- Wrong axis formatting: Large numbers without thousands separators
- No grid: Hard to read exact values from bars
- Small figure size: Everything cramped and hard to read
✅ BEST PRACTICES:
- Always start with 30 bins: Then adjust based on your data
- Use descriptive labels: Include units (Euros, meters, etc.)
- Add reference lines: Mean, median, important thresholds
- Choose good colors: Colorblind-friendly, publication-ready
- Format large numbers: Use €K for thousands, €M for millions
- Add a legend: When you have multiple data series
- Use plt.tight_layout(): Especially with multiple subplots
- Save high-quality images: dpi=300 for publications
Your Histogram Mastery Checklist
🎓 WHAT YOU'VE LEARNED:
✅ Can create basic histograms with plt.hist()
✅ Understand how bins affect histogram appearance
✅ Can customize colors, transparency, and edges
✅ Know how to create multiple plots with plt.subplots()
✅ Can add mean/median lines with plt.axvline()
✅ Understand density vs cumulative histograms
✅ Can format axes with Euro symbols and K/M abbreviations
✅ Know to always use plt.tight_layout() and plt.show()
Remember: Histograms are your window into understanding data distributions. They reveal patterns that summary statistics alone cannot show. With the Amsterdam house price data, you've seen not just average prices, but the entire market structure - from affordable apartments to luxury mansions.
Happy histogramming! 🎯
Comments
Post a Comment