Skip to main content

Mastering Boxplot (EDA)

Calculating read time…

📌 Dataset Source: To follow along with this tutorial, download the Amsterdam House Prices Dataset from Kaggle. You can access it directly here:
https://www.kaggle.com/datasets/thomasnibb/amsterdam-house-price-data

"Boxplots are like X-ray vision for your data - they let you see inside the distribution and spot outliers instantly!"

What is a Boxplot? (The Data Detective's X-Ray!)

A boxplot (or box-and-whisker plot) is a standardized way of displaying the distribution of data based on a five-number summary: minimum, first quartile (Q1), median, third quartile (Q3), and maximum. It answers questions like:

  • 📊 What's the spread of the data? (How variable are house prices?)
  • 🔍 Are there outliers? (Extremely cheap or expensive houses?)
  • 📏 Is the data symmetric or skewed?
  • 📈 How do different groups compare? (Prices by neighborhood?)

📖 Simple Analogy:

Imagine sorting all Amsterdam houses by price, then dividing them into 4 equal groups. The boxplot shows:
• Bottom whisker: Cheapest 25% of houses
• Box (Q1 to Q3): Middle 50% of houses (the "normal" range)
• Line in box: Median price (middle value)
• Top whisker: Most expensive 25% of houses
• Dots: Outliers (unusually cheap/expensive houses)

Essential Boxplot Terms - Explained Simply

Before we dive into code, let's understand the key components of a boxplot:

1. Median (Q2) - The Middle Value

What it is: The value separating the higher half from the lower half of the data.
Calculation: Middle value when data is sorted
Why it matters: Less affected by outliers than the mean
In boxplot: The line inside the box

# Calculating median
median_price = df['Price'].median()
print(f"Median House Price: €{median_price:,.0f}")
# This is the line you see in the middle of the boxplot!

2. Quartiles (Q1 & Q3) - The Box Boundaries

Q1 (First Quartile): 25th percentile - 25% of data is below this
Q3 (Third Quartile): 75th percentile - 75% of data is below this
IQR (Interquartile Range): Q3 - Q1 (the box height)
Why it matters: Shows where the middle 50% of data lies

# Calculating quartiles
Q1 = df['Price'].quantile(0.25)  # 25th percentile
Q3 = df['Price'].quantile(0.75)  # 75th percentile
IQR = Q3 - Q1  # Interquartile Range

print(f"Q1 (25th percentile): €{Q1:,.0f}")
print(f"Q3 (75th percentile): €{Q3:,.0f}")
print(f"IQR (Middle 50% range): €{IQR:,.0f}")
# The box in boxplot spans from Q1 to Q3

3. Whiskers - The Normal Range

What they are: Lines extending from the box
Default calculation: 1.5 × IQR from quartiles
Lower whisker: Q1 - (1.5 × IQR) or minimum value (whichever is higher)
Upper whisker: Q3 + (1.5 × IQR) or maximum value (whichever is lower)
Why it matters: Shows the expected range of "normal" values

# Calculating whiskers
lower_whisker = max(df['Price'].min(), Q1 - 1.5 * IQR)
upper_whisker = min(df['Price'].max(), Q3 + 1.5 * IQR)

print(f"Lower whisker: €{lower_whisker:,.0f}")
print(f"Upper whisker: €{upper_whisker:,.0f}")
print(f"Normal price range: €{lower_whisker:,.0f} to €{upper_whisker:,.0f}")

4. Outliers - The Special Cases

What they are: Data points outside the whiskers
Calculation: Values < Q1 - 1.5×IQR or > Q3 + 1.5×IQR
In boxplot: Shown as individual dots
Why they matter: Could be errors, special cases, or extreme values

# Identifying outliers
outliers = df[(df['Price'] < Q1 - 1.5 * IQR) | 
              (df['Price'] > Q3 + 1.5 * IQR)]

print(f"Number of outliers: {len(outliers)}")
print(f"Percentage of outliers: {len(outliers)/len(df)*100:.1f}%")
print(f"\nOutlier price range: €{outliers['Price'].min():,.0f} to €{outliers['Price'].max():,.0f}")
# These are the dots you see beyond the whiskers!

Method 1: Your First Boxplot (The Absolute Basics)

Let's create the simplest possible boxplot to understand house price distribution:

# METHOD 1: Basic boxplot
plt.figure(figsize=(10, 6))  # Create canvas

# Create basic boxplot
plt.boxplot(df['Price'])

# Add labels and title
plt.ylabel('House Price (Euros)', fontsize=12)
plt.title('Amsterdam House Price Distribution', fontsize=14, fontweight='bold')

# Format y-axis to show Euro symbols
plt.gca().yaxis.set_major_formatter(
    plt.FuncFormatter(lambda x, p: f'€{x:,.0f}'))

# Add grid for readability
plt.grid(True, alpha=0.3, axis='y')

# Show the plot
plt.tight_layout()
plt.show()

# Calculate and display boxplot statistics
print("📊 BOXPLOT STATISTICS:")
print("="*45)

price_data = df['Price']
Q1 = price_data.quantile(0.25)
median = price_data.median()
Q3 = price_data.quantile(0.75)
IQR = Q3 - Q1

print(f"Minimum: €{price_data.min():,.0f}")
print(f"Q1 (25th percentile): €{Q1:,.0f}")
print(f"Median: €{median:,.0f}")
print(f"Q3 (75th percentile): €{Q3:,.0f}")
print(f"Maximum: €{price_data.max():,.0f}")
print(f"IQR (Middle 50% range): €{IQR:,.0f}")
print(f"Lower whisker (approx): €{max(price_data.min(), Q1 - 1.5 * IQR):,.0f}")
print(f"Upper whisker (approx): €{min(price_data.max(), Q3 + 1.5 * IQR):,.0f}")

# Count outliers
outliers = price_data[(price_data < Q1 - 1.5 * IQR) | 
                      (price_data > Q3 + 1.5 * IQR)]
print(f"Number of outliers: {len(outliers)}")
print(f"Outlier percentage: {len(outliers)/len(price_data)*100:.1f}%")

✅ What You Just Learned:

  • plt.boxplot(data) - Creates basic boxplot
  • plt.ylabel() - Adds y-axis label
  • plt.grid(axis='y') - Adds horizontal grid lines
  • How to interpret quartiles, median, and outliers
  • How to calculate boxplot statistics manually

Method 2: Horizontal Boxplot (Better for Price Data)

Horizontal boxplots are often more readable for price data:

# METHOD 2: Horizontal boxplot
plt.figure(figsize=(12, 6))

# Create horizontal boxplot (vert=False)
plt.boxplot(df['Price'], vert=False)

# Add labels and title
plt.xlabel('House Price (Euros)', fontsize=12, fontweight='bold')
plt.title('Amsterdam House Price Distribution (Horizontal View)', 
          fontsize=14, fontweight='bold')

# Format x-axis to show Euro symbols with K for thousands
plt.gca().xaxis.set_major_formatter(
    plt.FuncFormatter(lambda x, p: f'€{x/1000:.0f}K'))

# Add vertical grid lines
plt.grid(True, alpha=0.3, axis='x')

# Add mean line for comparison
mean_price = df['Price'].mean()
plt.axvline(mean_price, color='red', linestyle='--', linewidth=2,
            alpha=0.7, label=f'Mean: €{mean_price/1000:.0f}K')

# Add median line (already in boxplot, but for emphasis)
plt.axvline(df['Price'].median(), color='green', linestyle='-.', 
            linewidth=1.5, alpha=0.5, label='Median')

# Add legend
plt.legend(fontsize=10)

# Add text with key statistics
stats_text = f"""Key Statistics:
• Median: €{df['Price'].median()/1000:.0f}K
• Mean: €{mean_price/1000:.0f}K
• IQR: €{(df['Price'].quantile(0.75) - df['Price'].quantile(0.25))/1000:.0f}K
• Outliers: {len(outliers)} houses"""
plt.text(0.02, 0.98, stats_text, transform=plt.gca().transAxes,
         fontsize=10, verticalalignment='top',
         bbox=dict(boxstyle='round', facecolor='wheat', alpha=0.8))

plt.tight_layout()
plt.show()

print("📈 HORIZONTAL BOXPLOT ADVANTAGES:")
print("="*45)
print("1. Easier to read price labels on x-axis")
print("2. Natural left-to-right reading (cheap to expensive)")
print("3. More space for price formatting (€K, €M)")
print("4. Better for comparing with other horizontal charts")

Method 3: Multiple Boxplots for Comparison

Compare house prices across different categories:

# METHOD 3: Multiple boxplots by number of bedrooms
plt.figure(figsize=(14, 8))

# Prepare data for boxplot (list of arrays)
bedroom_data = []
bedroom_labels = []

# Get unique bedroom counts (sorted)
bedroom_counts = sorted(df['Bedroom'].dropna().unique())

for bedrooms in bedroom_counts:
    if bedrooms <= 5:  # Limit to 0-5 bedrooms for clarity
        bedroom_prices = df[df['Bedroom'] == bedrooms]['Price']
        if len(bedroom_prices) > 5:  # Only include if enough data
            bedroom_data.append(bedroom_prices)
            bedroom_labels.append(f'{bedrooms} Bedroom{"s" if bedrooms != 1 else ""}')

# Create multiple boxplots
bp = plt.boxplot(bedroom_data, labels=bedroom_labels, patch_artist=True)

# Customize box colors
colors = ['#FF9999', '#66B2FF', '#99FF99', '#FFB366', '#FF99FF', '#FFD700']
for patch, color in zip(bp['boxes'], colors[:len(bedroom_data)]):
    patch.set_facecolor(color)
    patch.set_alpha(0.7)

# Customize median lines
for median in bp['medians']:
    median.set(color='black', linewidth=2)

# Add labels and title
plt.ylabel('House Price (Euros)', fontsize=12, fontweight='bold')
plt.xlabel('Number of Bedrooms', fontsize=12, fontweight='bold')
plt.title('House Price Distribution by Number of Bedrooms', 
          fontsize=16, fontweight='bold')

# Format y-axis
plt.gca().yaxis.set_major_formatter(
    plt.FuncFormatter(lambda x, p: f'€{x/1000:.0f}K'))

# Add grid
plt.grid(True, alpha=0.3, axis='y')

# Add mean markers for each group
for i, data in enumerate(bedroom_data):
    mean_price = data.mean()
    plt.plot(i + 1, mean_price, 'ro', markersize=8, 
             label='Mean' if i == 0 else "")

# Add legend for mean markers
plt.legend(['Median (black line)', 'Mean (red dot)'], fontsize=10)

plt.tight_layout()
plt.show()

print("🛏️ BEDROOM COMPARISON ANALYSIS:")
print("="*45)

# Create comparison table
comparison_data = []
for i, bedrooms in enumerate(bedroom_counts):
    if bedrooms <= 5:
        bedroom_prices = df[df['Bedroom'] == bedrooms]['Price']
        if len(bedroom_prices) > 5:
            comparison_data.append({
                'Bedrooms': bedrooms,
                'Count': len(bedroom_prices),
                'Median Price': f'€{bedroom_prices.median()/1000:.0f}K',
                'Mean Price': f'€{bedroom_prices.mean()/1000:.0f}K',
                'Price Range': f'€{bedroom_prices.min()/1000:.0f}K-€{bedroom_prices.max()/1000:.0f}K'
            })

# Convert to DataFrame for nice display
comparison_df = pd.DataFrame(comparison_data)
print(comparison_df.to_string(index=False))

print("\n💡 Key Insights:")
print("1. Price generally increases with number of bedrooms")
print("2. But not always linear - check the median progression")
print("3. Variation increases with bedroom count (taller boxes)")
print("4. Notice any outliers in each category?")

Method 4: Grouped Boxplots (Multiple Comparisons)

Compare prices by multiple categories simultaneously:

# METHOD 4: Grouped boxplots by neighborhood and house type
plt.figure(figsize=(16, 10))

# Select top neighborhoods for clarity
top_neighborhoods = df['Neighborhood'].value_counts().head(5).index.tolist()
df_top = df[df['Neighborhood'].isin(top_neighborhoods)]

# Get unique house types
house_types = sorted(df_top['Type'].dropna().unique())

# Prepare data structure for grouped boxplot
data_to_plot = []
positions = []
group_labels = []
pos = 1

# Create grouped data
for i, neighborhood in enumerate(top_neighborhoods):
    neighborhood_data = []
    for j, house_type in enumerate(house_types):
        subset = df_top[(df_top['Neighborhood'] == neighborhood) & 
                        (df_top['Type'] == house_type)]
        if len(subset) > 3:  # Only include if enough data points
            neighborhood_data.append(subset['Price'].values)
            positions.append(pos)
            group_labels.append(f'{neighborhood}\n{house_type}')
            pos += 1
        else:
            positions.append(pos)
            group_labels.append('')
            pos += 1
    data_to_plot.extend(neighborhood_data)
    pos += 1  # Add space between neighborhoods

# Create the boxplot
bp = plt.boxplot(data_to_plot, positions=positions, 
                 patch_artist=True, widths=0.6)

# Color by neighborhood
neighborhood_colors = ['#FF6B6B', '#4ECDC4', '#45B7D1', '#96CEB4', '#FFEAA7']
for i, (box, position) in enumerate(zip(bp['boxes'], positions)):
    neighborhood_idx = (position - 1) // (len(house_types) + 1)
    box.set_facecolor(neighborhood_colors[neighborhood_idx % len(neighborhood_colors)])
    box.set_alpha(0.7)

# Customize median lines
for median in bp['medians']:
    median.set(color='black', linewidth=1.5)

# Customize whiskers and caps
for whisker in bp['whiskers']:
    whisker.set(color='gray', linewidth=1, linestyle='--')
    
for cap in bp['caps']:
    cap.set(color='gray', linewidth=1)

# Set x-ticks and labels
plt.xticks(positions, group_labels, rotation=45, ha='right', fontsize=9)

# Add labels and title
plt.ylabel('House Price (Euros)', fontsize=12, fontweight='bold')
plt.title('House Price Distribution by Neighborhood and House Type', 
          fontsize=16, fontweight='bold', pad=20)

# Format y-axis
plt.gca().yaxis.set_major_formatter(
    plt.FuncFormatter(lambda x, p: f'€{x/1000:.0f}K'))

# Add grid
plt.grid(True, alpha=0.2, axis='y')

# Add legend for neighborhoods
from matplotlib.patches import Patch
legend_elements = [Patch(facecolor=neighborhood_colors[i], alpha=0.7, 
                         label=top_neighborhoods[i]) 
                   for i in range(len(top_neighborhoods))]
plt.legend(handles=legend_elements, title='Neighborhoods', 
           fontsize=10, title_fontsize=11, loc='upper left')

plt.tight_layout()
plt.show()

print("🏙️ NEIGHBORHOOD & HOUSE TYPE ANALYSIS:")
print("="*55)

# Analyze price differences
print("\nMost Expensive Neighborhoods (Median Prices):")
print("-"*45)

neighborhood_stats = []
for neighborhood in top_neighborhoods:
    neighborhood_prices = df_top[df_top['Neighborhood'] == neighborhood]['Price']
    neighborhood_stats.append({
        'Neighborhood': neighborhood,
        'Median Price': f'€{neighborhood_prices.median()/1000:.0f}K',
        'Average Price': f'€{neighborhood_prices.mean()/1000:.0f}K',
        'Price Range': f'€{neighborhood_prices.min()/1000:.0f}K-€{neighborhood_prices.max()/1000:.0f}K',
        'Number of Houses': len(neighborhood_prices)
    })

neighborhood_df = pd.DataFrame(neighborhood_stats)
print(neighborhood_df.to_string(index=False))

print("\n💡 Key Insights from Grouped Boxplots:")
print("1. Some neighborhoods consistently more expensive (taller boxes)")
print("2. Certain house types may be outliers in specific neighborhoods")
print("3. Price variation differs by neighborhood (box height differences)")
print("4. Some neighborhoods have more outliers than others")

Method 5: Violin Plot Alternative (Density + Boxplot)

Violin plots combine boxplots with kernel density estimation:

# METHOD 5: Violin plots (enhanced boxplots)
fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(16, 8))

# Plot 1: Regular boxplot for comparison
ax1.boxplot(df['Price'], vert=False)
ax1.set_xlabel('House Price (Euros)', fontsize=11)
ax1.set_title('Traditional Boxplot', fontsize=13, fontweight='bold')
ax1.grid(True, alpha=0.3, axis='x')
ax1.xaxis.set_major_formatter(
    plt.FuncFormatter(lambda x, p: f'€{x/1000:.0f}K'))

# Add statistics to boxplot
Q1 = df['Price'].quantile(0.25)
median = df['Price'].median()
Q3 = df['Price'].quantile(0.75)

ax1.axvline(median, color='green', linestyle='-', alpha=0.5, linewidth=1)
ax1.text(median, 1.05, f'Median: €{median/1000:.0f}K', 
         color='green', fontsize=9, ha='center')

# Plot 2: Violin plot
violin_parts = ax2.violinplot(df['Price'], vert=False, showmeans=True, 
                              showmedians=True, showextrema=True)

# Customize violin plot colors
for pc in violin_parts['bodies']:
    pc.set_facecolor('#FF6B6B')
    pc.set_alpha(0.7)
    pc.set_edgecolor('black')

# Customize median and mean lines
violin_parts['cmedians'].set_color('black')
violin_parts['cmedians'].set_linewidth(2)
violin_parts['cmeans'].set_color('blue')
violin_parts['cmeans'].set_linewidth(2)

ax2.set_xlabel('House Price (Euros)', fontsize=11)
ax2.set_title('Violin Plot (Boxplot + Density)', fontsize=13, fontweight='bold')
ax2.grid(True, alpha=0.3, axis='x')
ax2.xaxis.set_major_formatter(
    plt.FuncFormatter(lambda x, p: f'€{x/1000:.0f}K'))

# Add legend
ax2.legend([violin_parts['bodies'][0], violin_parts['cmedians'], violin_parts['cmeans']],
           ['Density', 'Median', 'Mean'], loc='upper right')

plt.tight_layout()
plt.show()

print("🎻 VIOLIN PLOT VS BOXPLOT:")
print("="*45)

print("\n📦 Traditional Boxplot:")
print("• Shows: Quartiles, median, outliers")
print("• Strengths: Clear outliers, simple statistics")
print("• Limitations: Doesn't show density/multimodality")

print("\n🎻 Violin Plot:")
print("• Shows: Everything in boxplot + density shape")
print("• Strengths: Reveals bimodal/multimodal distributions")
print("• Can show: Density peaks, gaps in data, skewness shape")

print("\n💡 When to Use Which:")
print("1. Use Boxplot when: Focus on outliers/comparison")
print("2. Use Violin Plot when: Understanding distribution shape")
print("3. Use Both when: Comprehensive EDA needed")

# Analyze distribution shape from violin plot
from scipy import stats

print("\n📊 Distribution Shape Analysis:")
print("-"*35)

skewness = stats.skew(df['Price'].dropna())
kurtosis = stats.kurtosis(df['Price'].dropna())

print(f"Skewness: {skewness:.2f}")
if skewness > 1:
    print("  → Strongly right-skewed (long tail to high prices)")
elif skewness > 0.5:
    print("  → Moderately right-skewed")
elif skewness > -0.5:
    print("  → Approximately symmetric")
else:
    print("  → Left-skewed")

print(f"\nKurtosis: {kurtosis:.2f}")
if kurtosis > 3:
    print("  → Leptokurtic (heavy tails, more outliers)")
elif kurtosis < 3:
    print("  → Platykurtic (light tails, fewer outliers)")
else:
    print("  → Mesokurtic (normal distribution tails)")

Method 6: Notched Boxplots (Confidence Intervals)

Notched boxplots show confidence intervals around the median:

# METHOD 6: Notched boxplots with confidence intervals
fig, axes = plt.subplots(2, 2, figsize=(14, 10))
fig.suptitle('Notched Boxplots: Comparing Price Distributions', 
             fontsize=16, fontweight='bold')

# Data preparation: Compare different price segments
price_segments = [
    df[df['Price'] <= 500000]['Price'],
    df[(df['Price'] > 500000) & (df['Price'] <= 1000000)]['Price'],
    df[(df['Price'] > 1000000) & (df['Price'] <= 2000000)]['Price'],
    df[df['Price'] > 2000000]['Price']
]

segment_labels = ['≤ €500K', '€500K-€1M', '€1M-€2M', '> €2M']
colors = ['#FF9999', '#66B2FF', '#99FF99', '#FFB366']

# Plot 1: Basic notched boxplot
axes[0, 0].boxplot(price_segments, labels=segment_labels, notch=True)
axes[0, 0].set_title('Notched Boxplot (Basic)', fontsize=12)
axes[0, 0].set_ylabel('Price (Euros)', fontsize=11)
axes[0, 0].grid(True, alpha=0.3, axis='y')
axes[0, 0].yaxis.set_major_formatter(
    plt.FuncFormatter(lambda x, p: f'€{x/1000:.0f}K'))

# Plot 2: Notched boxplot with custom confidence interval
axes[0, 1].boxplot(price_segments, labels=segment_labels, 
                   notch=True, bootstrap=10000)  # More bootstrap samples
axes[0, 1].set_title('Notched Boxplot (10K Bootstrap Samples)', fontsize=12)
axes[0, 1].grid(True, alpha=0.3, axis='y')
axes[0, 1].yaxis.set_major_formatter(
    plt.FuncFormatter(lambda x, p: f'€{x/1000:.0f}K'))

# Plot 3: Notched horizontal boxplot
axes[1, 0].boxplot(price_segments, labels=segment_labels, 
                   notch=True, vert=False, patch_artist=True)

# Color the boxes
for i, box in enumerate(axes[1, 0].patches):
    box.set_facecolor(colors[i % len(colors)])
    box.set_alpha(0.7)

axes[1, 0].set_title('Horizontal Notched Boxplot', fontsize=12)
axes[1, 0].set_xlabel('Price (Euros)', fontsize=11)
axes[1, 0].grid(True, alpha=0.3, axis='x')
axes[1, 0].xaxis.set_major_formatter(
    plt.FuncFormatter(lambda x, p: f'€{x/1000:.0f}K'))

# Plot 4: Notched vs Regular comparison
bp1 = axes[1, 1].boxplot([price_segments[0], price_segments[1]], 
                         positions=[1, 2], widths=0.35,
                         labels=['Regular', 'Notched'], notch=False)
bp2 = axes[1, 1].boxplot([price_segments[0], price_segments[1]], 
                         positions=[1.4, 2.4], widths=0.35,
                         labels=['', ''], notch=True, patch_artist=True)

# Color the notched boxes
for box in bp2['boxes']:
    box.set_facecolor('#FF9999')
    box.set_alpha(0.7)

axes[1, 1].set_title('Regular vs Notched Comparison', fontsize=12)
axes[1, 1].set_ylabel('Price (€500K Segment)', fontsize=11)
axes[1, 1].set_xticks([1.2, 2.2])
axes[1, 1].set_xticklabels(['≤ €500K', '€500K-€1M'])
axes[1, 1].grid(True, alpha=0.3, axis='y')
axes[1, 1].yaxis.set_major_formatter(
    plt.FuncFormatter(lambda x, p: f'€{x/1000:.0f}K'))

# Add legend
axes[1, 1].legend([bp1['boxes'][0], bp2['boxes'][0]], 
                  ['Regular Boxplot', 'Notched Boxplot'], 
                  fontsize=9)

plt.tight_layout()
plt.show()

print("🎯 NOTCHED BOXPLOTS EXPLAINED:")
print("="*45)

print("\nWhat are the notches?")
print("• Confidence interval around the median")
print("• Typically 95% confidence interval")
print("• Calculated as: median ± 1.57 × IQR / √n")

print("\nHow to interpret notches:")
print("1. If notches DON'T overlap: Medians are significantly different")
print("2. If notches DO overlap: Medians might not be significantly different")
print("3. Width of notch: Indicates confidence level (narrow = more confidence)")

print("\n📊 Statistical Significance Check:")
print("-"*35)

for i, segment in enumerate(price_segments):
    if len(segment) > 0:
        median = segment.median()
        n = len(segment)
        IQR = segment.quantile(0.75) - segment.quantile(0.25)
        notch_width = 1.57 * IQR / np.sqrt(n)
        
        print(f"\n{segment_labels[i]}:")
        print(f"  • Median: €{median/1000:.0f}K")
        print(f"  • Sample size: {n}")
        print(f"  • 95% CI: €{(median - notch_width)/1000:.0f}K to €{(median + notch_width)/1000:.0f}K")
        print(f"  • Notch width: €{notch_width/1000:.0f}K")

Method 7: Advanced Customization & Styling

Create publication-ready boxplots with advanced styling:

# METHOD 7: Professional publication-ready boxplot
plt.figure(figsize=(16, 8))

# Create more detailed data segmentation
conditions = [
    (df['Price'] <= 300000),
    ((df['Price'] > 300000) & (df['Price'] <= 500000)),
    ((df['Price'] > 500000) & (df['Price'] <= 750000)),
    ((df['Price'] > 750000) & (df['Price'] <= 1000000)),
    ((df['Price'] > 1000000) & (df['Price'] <= 1500000)),
    (df['Price'] > 1500000)
]

segments = []
segment_labels = ['≤ €300K', '€300-500K', '€500-750K', 
                  '€750K-€1M', '€1-1.5M', '> €1.5M']

for condition in conditions:
    segments.append(df[condition]['Price'])

# Create the boxplot with extensive customization
bp = plt.boxplot(segments, labels=segment_labels, patch_artist=True,
                 notch=True,  # Add notches for confidence intervals
                 bootstrap=5000,  # Bootstrap for notch calculation
                 widths=0.7,  # Box width
                 showmeans=True,  # Show mean markers
                 meanline=True,  # Show mean as line
                 showfliers=True,  # Show outliers
                 flierprops=dict(marker='o', markerfacecolor='red', 
                                markersize=6, alpha=0.6, 
                                markeredgecolor='none'),  # Outlier style
                 medianprops=dict(color='black', linewidth=2.5),  # Median line
                 meanprops=dict(color='blue', linewidth=2.5, 
                               linestyle='--'),  # Mean line
                 boxprops=dict(linewidth=1.5),  # Box outline
                 whiskerprops=dict(linewidth=1.5, linestyle='-'),  # Whiskers
                 capprops=dict(linewidth=1.5))  # Whisker caps

# Custom box colors using a gradient
colors = plt.cm.viridis(np.linspace(0.2, 0.8, len(segments)))
for patch, color in zip(bp['boxes'], colors):
    patch.set_facecolor(color)
    patch.set_alpha(0.8)
    patch.set_edgecolor('black')

# Add labels and title with professional formatting
plt.ylabel('House Price (Euros)', fontsize=14, fontweight='bold', labelpad=15)
plt.xlabel('Price Segment', fontsize=14, fontweight='bold', labelpad=15)
plt.title('Amsterdam House Price Distribution by Segment\nProfessional Analysis', 
          fontsize=18, fontweight='bold', pad=20)

# Format y-axis professionally
plt.gca().yaxis.set_major_formatter(
    plt.FuncFormatter(lambda x, p: f'€{x/1000:.0f}K'))
plt.gca().yaxis.set_tick_params(labelsize=11)

# Rotate x-tick labels for better readability
plt.xticks(rotation=30, ha='right', fontsize=11)

# Add subtle grid
plt.grid(True, alpha=0.15, axis='y', linestyle='--')

# Add text annotations for key insights
insights = [
    "Most houses in €300-500K range",
    "Right-skewed distribution",
    "Significant luxury segment (>€1.5M)",
    "Widening IQR with price increase"
]

for i, insight in enumerate(insights):
    plt.text(0.02, 0.95 - i*0.05, f"• {insight}", 
             transform=plt.gca().transAxes,
             fontsize=10, verticalalignment='top',
             bbox=dict(boxstyle='round', facecolor='white', 
                      alpha=0.8, edgecolor='gray'))

# Add statistics table
import matplotlib.patches as mpatches

# Create custom legend
median_patch = mpatches.Patch(color='black', label='Median', alpha=0.7)
mean_patch = mpatches.Patch(color='blue', label='Mean', alpha=0.7)
outlier_patch = mpatches.Patch(color='red', label='Outliers', alpha=0.7)

plt.legend(handles=[median_patch, mean_patch, outlier_patch],
           loc='upper right', fontsize=10, framealpha=0.9)

# Add overall statistics text
total_houses = len(df)
outliers_total = sum([len(segment[segment > segment.quantile(0.75) + 1.5 * 
                         (segment.quantile(0.75) - segment.quantile(0.25))]) 
                     for segment in segments if len(segment) > 0])

stats_text = f"""Overall Statistics:
• Total Houses: {total_houses:,}
• Overall Median: €{df['Price'].median()/1000:.0f}K
• Overall Mean: €{df['Price'].mean()/1000:.0f}K
• Price Range: €{df['Price'].min()/1000:.0f}K - €{df['Price'].max()/1000:.0f}K
• Total Outliers: {outliers_total:,} ({outliers_total/total_houses*100:.1f}%)"""

plt.text(0.98, 0.02, stats_text, transform=plt.gca().transAxes,
         fontsize=10, verticalalignment='bottom', horizontalalignment='right',
         bbox=dict(boxstyle='round', facecolor='lightyellow', 
                  alpha=0.9, edgecolor='orange'))

plt.tight_layout()
plt.show()

print("🎨 PROFESSIONAL BOXPLOT FEATURES:")
print("="*45)
print("1. Gradient color scheme (viridis colormap)")
print("2. Notched boxes with confidence intervals")
print("3. Both mean and median lines shown")
print("4. Custom outlier styling (red circles)")
print("5. Professional typography and spacing")
print("6. Statistical annotations")
print("7. Comprehensive legend")
print("8. Overall statistics box")
print("9. Subtle grid for readability")
print("10. Publication-ready formatting")

Boxplot Cheat Sheet (Quick Reference)

# 1. BASIC BOXPLOT
plt.boxplot(data)                     # Vertical boxplot
plt.boxplot(data, vert=False)         # Horizontal boxplot

# 2. MULTIPLE BOXPLOTS
plt.boxplot([data1, data2, data3])    # Multiple groups
plt.boxplot([data1, data2], labels=['Group A', 'Group B'])

# 3. CUSTOMIZATION OPTIONS
plt.boxplot(data, patch_artist=True)  # Enable color filling
plt.boxplot(data, notch=True)         # Add confidence intervals
plt.boxplot(data, showmeans=True)     # Show mean markers
plt.boxplot(data, showfliers=False)   # Hide outliers

# 4. STYLING PROPERTIES
boxprops = dict(linewidth=2, color='blue')      # Box style
medianprops = dict(linewidth=2.5, color='red')  # Median line
whiskerprops = dict(linewidth=1.5)              # Whisker style
capprops = dict(linewidth=1.5)                  # Whisker cap style
flierprops = dict(marker='o', color='red', alpha=0.5)  # Outlier style
meanprops = dict(marker='D', markeredgecolor='black')  # Mean style

# 5. COMPLETE EXAMPLE
plt.boxplot(data, patch_artist=True, notch=True,
            boxprops=dict(facecolor='lightblue', color='blue'),
            medianprops=dict(color='red'),
            whiskerprops=dict(color='blue'),
            capprops=dict(color='blue'),
            flierprops=dict(marker='o', color='red', alpha=0.5))

# 6. HORIZONTAL FORMATTING
plt.boxplot(data, vert=False)
plt.xlabel('Values')
plt.gca().xaxis.set_major_formatter(plt.FuncFormatter(lambda x, p: f'€{x:,.0f}'))

# 7. GROUPED BOXPLOTS (Advanced)
positions = [1, 2, 4, 5, 7, 8]  # Custom positions
labels = ['A', 'B', 'C', 'D', 'E', 'F']
plt.boxplot([data1, data2, data3, data4, data5, data6], 
            positions=positions, labels=labels)

# 8. VIOLIN PLOT ALTERNATIVE
plt.violinplot(data, showmeans=True, showmedians=True)

# 9. STATISTICS CALCULATION
Q1 = np.percentile(data, 25)
median = np.median(data)
Q3 = np.percentile(data, 75)
IQR = Q3 - Q1
outliers = data[(data < Q1 - 1.5*IQR) | (data > Q3 + 1.5*IQR)]

# 10. SAVING
plt.savefig('boxplot.png', dpi=300, bbox_inches='tight', facecolor='white')

Common Boxplot Mistakes to Avoid ⚠️

❌ COMMON MISTAKES:

  • Ignoring outliers: Not investigating why they exist
  • Wrong scale: Using linear scale for highly skewed data
  • Overcrowding: Too many categories in one plot
  • No labels: Forgetting to label groups or axes
  • Color misuse: Using similar colors for different groups
  • Misinterpreting notches: Treating overlap as proof of no difference
  • Small sample sizes: Creating boxplots with < 5 data points
  • Forgetting units: Not showing currency or measurement units

✅ BEST PRACTICES:

  • Always investigate outliers: Are they errors or meaningful?
  • Use log scale when needed: For highly skewed data
  • Limit categories: Max 6-8 groups per plot for readability
  • Use color effectively: Different colors for different groups
  • Add mean markers: Especially when data is skewed
  • Consider violin plots: When you need to see density
  • Add sample sizes: Show n for each group
  • Use consistent formatting: Same scale for comparison plots

Your Boxplot Mastery Checklist

🎓 WHAT YOU'VE LEARNED:

✅ Can create basic boxplots with plt.boxplot()
✅ Understand quartiles, median, IQR, and outliers
✅ Can create horizontal boxplots for better readability
✅ Know how to compare multiple groups
✅ Can create notched boxplots for confidence intervals
✅ Understand violin plots as an alternative
✅ Can create publication-ready boxplots
✅ Know common mistakes and best practices

Remember: Boxplots are powerful tools for understanding data distribution at a glance. They reveal central tendency, spread, skewness, and outliers all in one visualization. With the Amsterdam house price data, boxplots help you see not just average prices, but the entire market structure and variability.

Happy boxplotting! 🎯

Comments