Skip to main content

Mastering BAR Chart (EDA)

Calculating read time…

📌 Dataset Source: To follow along with this tutorial, download the Amsterdam House Prices Dataset from Kaggle. You can access it directly here:
https://www.kaggle.com/datasets/thomasnibb/amsterdam-house-price-data

"Bar charts are the workhorses of data visualization - simple, intuitive, and incredibly powerful for comparing categorical data and showing proportions."

What is a Bar Chart? (The Comparison Champion!)

A bar chart displays categorical data with rectangular bars where the length/height is proportional to the values they represent. It's ideal for comparing different groups or tracking changes over time. It answers questions like:

  • 📊 Which category is largest/smallest? (Most common house type?)
  • 📈 How do different groups compare? (Prices by neighborhood?)
  • 📅 How have things changed over time? (Yearly price trends?)
  • 🧮 What are the proportions between categories? (Market share?)

📖 Simple Analogy:

Imagine counting different types of houses in Amsterdam:
• 🏢 Apartments: 150 bars tall
• 🏘️ Houses: 100 bars tall
• 🏰 Maisons: 50 bars tall
A bar chart simply shows these counts as bars of different heights - instantly revealing which type is most common!

Bar Chart vs Histogram - Know the Difference!

🎯 CRITICAL DISTINCTION:

📊 BAR CHART:

  • For: Categorical data (names, labels)
  • Bars: Separate, have gaps between them
  • X-axis: Categories (neighborhoods, types)
  • Order: Can be sorted by value or category
  • Example: Count of houses by type

📈 HISTOGRAM:

  • For: Numerical data (prices, sizes)
  • Bars: Touching, no gaps
  • X-axis: Numerical ranges (price bins)
  • Order: Fixed numerical order
  • Example: Distribution of house prices

💡 Quick Rule: If you can change the order of bars without losing meaning → Bar Chart. If order matters → Histogram.

Essential Bar Chart Types - Explained Simply

Different bar charts serve different purposes. Choose wisely:

1. Vertical Bar Chart (Column Chart)

What it is: Bars extend vertically from x-axis
Best for: Comparing different categories
When to use: Category names are short, 4-8 categories
Example: House counts by neighborhood

# Vertical bar chart (most common)
categories = ['Apartment', 'House', 'Maison']
counts = [150, 100, 50]

plt.figure(figsize=(10, 6))
plt.bar(categories, counts, color=['#FF6B6B', '#4ECDC4', '#FFD166'])
plt.title('House Types in Amsterdam', fontsize=14, fontweight='bold')
plt.xlabel('House Type', fontsize=12)
plt.ylabel('Number of Houses', fontsize=12)
plt.show()

2. Horizontal Bar Chart

What it is: Bars extend horizontally from y-axis
Best for: Long category names, many categories
When to use: More than 8 categories, ranking comparisons
Example: Average price by neighborhood (sorted)

# Horizontal bar chart (better for many/long labels)
neighborhoods = ['Jordaan', 'De Pijp', 'Oud-West', 'Centrum', 'Zuid']
avg_prices = [650000, 620000, 580000, 720000, 850000]

plt.figure(figsize=(10, 6))
plt.barh(neighborhoods, avg_prices, color='skyblue')
plt.title('Average House Price by Neighborhood', fontsize=14, fontweight='bold')
plt.xlabel('Average Price (Euros)', fontsize=12)
plt.ylabel('Neighborhood', fontsize=12)
plt.gca().xaxis.set_major_formatter(
    plt.FuncFormatter(lambda x, p: f'€{x/1000:.0f}K'))
plt.show()

3. Grouped Bar Chart

What it is: Multiple bars per category (side-by-side)
Best for: Comparing sub-categories across main categories
When to use: 2-3 sub-categories per category
Example: Old vs New house prices by neighborhood

# Grouped bar chart
import numpy as np

categories = ['Apartment', 'House', 'Maison']
old_prices = [450000, 520000, 680000]
new_prices = [480000, 550000, 720000]

x = np.arange(len(categories))
width = 0.35

plt.figure(figsize=(10, 6))
plt.bar(x - width/2, old_prices, width, label='Old Houses', color='#FF6B6B')
plt.bar(x + width/2, new_prices, width, label='New Houses', color='#4ECDC4')

plt.xlabel('House Type', fontsize=12)
plt.ylabel('Average Price (Euros)', fontsize=12)
plt.title('Average Price: Old vs New Houses', fontsize=14, fontweight='bold')
plt.xticks(x, categories)
plt.legend()
plt.gca().yaxis.set_major_formatter(
    plt.FuncFormatter(lambda x, p: f'€{x/1000:.0f}K'))
plt.show()

4. Stacked Bar Chart

What it is: Bars stacked on top of each other
Best for: Showing part-to-whole relationships
When to use: Showing composition/proportions
Example: Price segments within each neighborhood

# Stacked bar chart
neighborhoods = ['Jordaan', 'De Pijp', 'Oud-West']
budget = [30, 40, 50]      # ≤ €400K
midrange = [40, 35, 30]    # €400K-€800K
luxury = [30, 25, 20]      # > €800K

plt.figure(figsize=(10, 6))
plt.bar(neighborhoods, budget, label='Budget (≤ €400K)', color='#FF9999')
plt.bar(neighborhoods, midrange, bottom=budget, 
        label='Mid-Range (€400-800K)', color='#66B2FF')
plt.bar(neighborhoods, luxury, bottom=np.array(budget) + np.array(midrange),
        label='Luxury (> €800K)', color='#99FF99')

plt.ylabel('Percentage of Houses', fontsize=12)
plt.title('Price Segment Distribution by Neighborhood', 
          fontsize=14, fontweight='bold')
plt.legend(loc='upper right')
plt.ylim(0, 100)
plt.show()

Method 1: Your First Bar Chart (Count Analysis)

Let's start with the most basic bar chart - counting houses by type:

# METHOD 1: Basic bar chart - Count houses by type
plt.figure(figsize=(12, 7))

# Get house type counts
type_counts = df['Type'].value_counts()

# Create bar chart
bars = plt.bar(type_counts.index, type_counts.values, 
               color=['#FF6B6B', '#4ECDC4', '#FFD166'],  # Custom colors
               edgecolor='black', linewidth=1.2,          # Bar borders
               alpha=0.8)                                 # Transparency

# Add labels and title
plt.xlabel('House Type', fontsize=12, fontweight='bold')
plt.ylabel('Number of Houses', fontsize=12, fontweight='bold')
plt.title('House Type Distribution in Amsterdam', 
          fontsize=16, fontweight='bold')

# Add value labels on top of bars
for bar in bars:
    height = bar.get_height()
    plt.text(bar.get_x() + bar.get_width()/2., height + 10,
             f'{int(height):,}', ha='center', va='bottom',
             fontsize=10, fontweight='bold')

# Add percentage labels
total_houses = len(df)
for i, (type_name, count) in enumerate(type_counts.items()):
    percentage = (count / total_houses) * 100
    plt.text(i, count/2, f'{percentage:.1f}%', 
             ha='center', va='center', fontsize=11,
             fontweight='bold', color='white')

# Add grid for better readability
plt.grid(True, alpha=0.3, axis='y')

# Add total count annotation
plt.text(0.02, 0.98, f'Total Houses: {total_houses:,}',
         transform=plt.gca().transAxes, fontsize=11,
         verticalalignment='top',
         bbox=dict(boxstyle='round', facecolor='wheat', alpha=0.8))

plt.tight_layout()
plt.show()

# Print detailed statistics
print("🏠 HOUSE TYPE ANALYSIS:")
print("="*45)

for type_name, count in type_counts.items():
    percentage = (count / total_houses) * 100
    avg_price = df[df['Type'] == type_name]['Price'].mean()
    median_price = df[df['Type'] == type_name]['Price'].median()
    
    print(f"\n{type_name}:")
    print(f"  • Count: {count:,} houses")
    print(f"  • Percentage: {percentage:.1f}% of market")
    print(f"  • Average Price: €{avg_price:,.0f}")
    print(f"  • Median Price: €{median_price:,.0f}")
    
    if type_name == type_counts.index[0]:  # Most common
        print(f"  • 🏆 MOST COMMON TYPE")
    if avg_price == max([df[df['Type'] == t]['Price'].mean() for t in type_counts.index]):
        print(f"  • 💰 MOST EXPENSIVE TYPE")

print(f"\n💡 Market Insight: {type_counts.index[0]} dominates with "
      f"{type_counts.iloc[0]/total_houses*100:.1f}% market share")

✅ What You Just Learned:

  • plt.bar(x, height) - Creates vertical bar chart
  • value_counts() - Counts categorical values
  • Adding value labels with plt.text()
  • Customizing bar colors and edges
  • Adding percentage calculations
  • Using grid for better readability

Method 2: Horizontal Bar Chart (Ranking Analysis)

Horizontal bar charts are perfect for ranking and comparing many categories:

# METHOD 2: Horizontal bar chart - Top neighborhoods by average price
plt.figure(figsize=(14, 8))

# Get top 10 neighborhoods by house count
neighborhood_counts = df['Neighborhood'].value_counts().head(10)
top_neighborhoods = neighborhood_counts.index.tolist()

# Calculate average prices for top neighborhoods
avg_prices = []
median_prices = []
house_counts = []

for neighborhood in top_neighborhoods:
    nb_data = df[df['Neighborhood'] == neighborhood]
    avg_prices.append(nb_data['Price'].mean())
    median_prices.append(nb_data['Price'].median())
    house_counts.append(len(nb_data))

# Sort by average price (descending) for ranking
sorted_indices = np.argsort(avg_prices)[::-1]  # Descending order
top_neighborhoods = [top_neighborhoods[i] for i in sorted_indices]
avg_prices = [avg_prices[i] for i in sorted_indices]
median_prices = [median_prices[i] for i in sorted_indices]
house_counts = [house_counts[i] for i in sorted_indices]

# Create horizontal bar chart
y_pos = np.arange(len(top_neighborhoods))
colors = plt.cm.viridis(np.linspace(0.2, 0.8, len(top_neighborhoods)))

bars = plt.barh(y_pos, avg_prices, color=colors, 
                edgecolor='black', linewidth=1, alpha=0.8)

# Add labels and title
plt.xlabel('Average House Price (Euros)', fontsize=12, fontweight='bold')
plt.ylabel('Neighborhood', fontsize=12, fontweight='bold')
plt.title('Top 10 Neighborhoods by Average House Price', 
          fontsize=16, fontweight='bold')
plt.yticks(y_pos, top_neighborhoods)

# Format x-axis with Euro symbols
plt.gca().xaxis.set_major_formatter(
    plt.FuncFormatter(lambda x, p: f'€{x/1000:.0f}K'))

# Add value labels on bars
for i, (bar, price, median, count) in enumerate(zip(bars, avg_prices, median_prices, house_counts)):
    width = bar.get_width()
    
    # Price label
    plt.text(width + 5000, bar.get_y() + bar.get_height()/2,
             f'€{price/1000:.0f}K', ha='left', va='center',
             fontsize=10, fontweight='bold')
    
    # Median indicator
    plt.scatter([median], [bar.get_y() + bar.get_height()/2],
                color='red', s=50, zorder=5, label='Median' if i == 0 else "")
    
    # Count annotation
    plt.text(10000, bar.get_y() + bar.get_height()/2,
             f'{count:,} houses', ha='left', va='center',
             fontsize=9, color='gray', style='italic')

# Add grid
plt.grid(True, alpha=0.2, axis='x', linestyle='--')

# Add legend
plt.legend(['Average Price (bar)', 'Median Price (dot)'], 
           loc='lower right', fontsize=10)

# Add statistical summary
total_avg = df['Price'].mean()
summary_text = f"""Market Summary:
• Overall Average: €{total_avg/1000:.0f}K
• Top vs Bottom: {avg_prices[0]/avg_prices[-1]:.1f}x difference
• Price Range: €{avg_prices[-1]/1000:.0f}K - €{avg_prices[0]/1000:.0f}K
• Top 10 represent {sum(house_counts):,} houses"""

plt.text(0.02, 0.98, summary_text, transform=plt.gca().transAxes,
         fontsize=10, verticalalignment='top',
         bbox=dict(boxstyle='round', facecolor='lightyellow', alpha=0.9))

plt.tight_layout()
plt.show()

print("🏙️ NEIGHBORHOOD RANKING ANALYSIS:")
print("="*50)

# Create ranking table
ranking_data = []
for i, (nb, avg, med, count) in enumerate(zip(top_neighborhoods, avg_prices, 
                                               median_prices, house_counts)):
    ranking_data.append({
        'Rank': i + 1,
        'Neighborhood': nb,
        'Avg Price': f'€{avg/1000:.0f}K',
        'Median Price': f'€{med/1000:.0f}K',
        'House Count': f'{count:,}',
        'Price Premium': f'+{(avg - total_avg)/total_avg*100:.0f}%'
    })

ranking_df = pd.DataFrame(ranking_data)
print(ranking_df.to_string(index=False))

print(f"\n💡 Market Insights:")
print(f"1. Most expensive neighborhood: {top_neighborhoods[0]} "
      f"(€{avg_prices[0]/1000:.0f}K average)")
print(f"2. Most affordable in top 10: {top_neighborhoods[-1]} "
      f"(€{avg_prices[-1]/1000:.0f}K average)")
print(f"3. Price premium in {top_neighborhoods[0]}: "
      f"+{(avg_prices[0] - total_avg)/total_avg*100:.0f}% above market")
print(f"4. Total houses in top 10: {sum(house_counts):,} "
      f"({sum(house_counts)/len(df)*100:.1f}% of market)")

Method 3: Grouped Bar Chart (Comparison Analysis)

Compare multiple metrics across categories side-by-side:

# METHOD 3: Grouped bar chart - Compare price metrics by house type
plt.figure(figsize=(14, 8))

# Prepare data
house_types = df['Type'].unique()
price_metrics = {}

for h_type in house_types:
    type_data = df[df['Type'] == h_type]['Price']
    price_metrics[h_type] = {
        'Average': type_data.mean(),
        'Median': type_data.median(),
        'Minimum': type_data.min(),
        'Maximum': type_data.max(),
        '25th Percentile': type_data.quantile(0.25),
        '75th Percentile': type_data.quantile(0.75)
    }

# Set up positions for grouped bars
x = np.arange(len(house_types))
width = 0.15  # Width of each bar
metrics_to_plot = ['Average', 'Median', '25th Percentile', '75th Percentile']
colors = ['#FF6B6B', '#4ECDC4', '#FFD166', '#06D6A0']

# Create grouped bars
for i, (metric, color) in enumerate(zip(metrics_to_plot, colors)):
    values = [price_metrics[h_type][metric] for h_type in house_types]
    offset = (i - len(metrics_to_plot)/2) * width + width/2
    bars = plt.bar(x + offset, values, width, 
                   label=metric, color=color, alpha=0.8,
                   edgecolor='black', linewidth=0.5)

    # Add value labels for average and median only (to avoid clutter)
    if metric in ['Average', 'Median']:
        for bar, value in zip(bars, values):
            height = bar.get_height()
            plt.text(bar.get_x() + bar.get_width()/2, height + 5000,
                     f'€{value/1000:.0f}K', ha='center', va='bottom',
                     fontsize=9, rotation=0)

# Add labels and title
plt.xlabel('House Type', fontsize=12, fontweight='bold')
plt.ylabel('Price (Euros)', fontsize=12, fontweight='bold')
plt.title('Price Statistics Comparison by House Type', 
          fontsize=16, fontweight='bold')
plt.xticks(x, house_types)

# Format y-axis
plt.gca().yaxis.set_major_formatter(
    plt.FuncFormatter(lambda x, p: f'€{x/1000:.0f}K'))

# Add grid
plt.grid(True, alpha=0.2, axis='y', linestyle=':')

# Add legend
plt.legend(title='Price Metrics', title_fontsize=11,
           fontsize=10, loc='upper left')

# Add range annotations (min to max)
for i, h_type in enumerate(house_types):
    min_price = price_metrics[h_type]['Minimum']
    max_price = price_metrics[h_type]['Maximum']
    
    # Add min-max range line
    plt.plot([x[i] - width*2, x[i] + width*2], 
             [min_price, min_price], 'k--', alpha=0.5, linewidth=1)
    plt.plot([x[i] - width*2, x[i] + width*2], 
             [max_price, max_price], 'k--', alpha=0.5, linewidth=1)
    
    # Add range label
    range_text = f'Range:\n€{min_price/1000:.0f}K-€{max_price/1000:.0f}K'
    plt.text(x[i], max_price + 20000, range_text,
             ha='center', va='bottom', fontsize=8,
             bbox=dict(boxstyle='round', facecolor='white', alpha=0.7))

# Add statistical insights
insights = [
    "• Apartments have narrowest price range",
    "• Maisons have highest average but also highest variability",
    "• Median < Average indicates right-skew in all types",
    "• 25th-75th percentiles show middle 50% price range"
]

for i, insight in enumerate(insights):
    plt.text(0.02, 0.95 - i*0.05, insight, transform=plt.gca().transAxes,
             fontsize=9, verticalalignment='top',
             bbox=dict(boxstyle='round', facecolor='lightblue', alpha=0.3))

plt.tight_layout()
plt.show()

print("📊 PRICE METRICS COMPARISON:")
print("="*45)

# Create detailed comparison table
comparison_data = []
for h_type in house_types:
    type_data = df[df['Type'] == h_type]['Price']
    comparison_data.append({
        'Type': h_type,
        'Count': len(type_data),
        'Average': f'€{type_data.mean()/1000:.0f}K',
        'Median': f'€{type_data.median()/1000:.0f}K',
        'Min': f'€{type_data.min()/1000:.0f}K',
        'Max': f'€{type_data.max()/1000:.0f}K',
        'Range': f'€{type_data.max()/type_data.min():.1f}x',
        'Std Dev': f'€{type_data.std()/1000:.0f}K',
        '25th %ile': f'€{type_data.quantile(0.25)/1000:.0f}K',
        '75th %ile': f'€{type_data.quantile(0.75)/1000:.0f}K'
    })

comparison_df = pd.DataFrame(comparison_data)
print(comparison_df.to_string(index=False))

print("\n💡 Key Business Insights:")
print("1. Apartments: Most common, most predictable pricing")
print("2. Houses: Balanced across all metrics")
print("3. Maisons: Premium segment with high variability")
print("4. All types show right-skew (Median < Average)")
print("5. Price ranges increase with property type luxury")

Method 4: Stacked Bar Chart (Composition Analysis)

Show how different segments contribute to the whole:

# METHOD 4: Stacked bar chart - Price segments by neighborhood
plt.figure(figsize=(16, 9))

# Define price segments
def get_price_segment(price):
    if price <= 400000:
        return 'Budget (≤ €400K)'
    elif price <= 700000:
        return 'Mid-Range (€400-700K)'
    elif price <= 1000000:
        return 'Premium (€700K-€1M)'
    else:
        return 'Luxury (> €1M)'

# Apply segmentation
df['Price_Segment'] = df['Price'].apply(get_price_segment)

# Get top 8 neighborhoods
top_neighborhoods = df['Neighborhood'].value_counts().head(8).index.tolist()
df_top = df[df['Neighborhood'].isin(top_neighborhoods)]

# Prepare data for stacking
segments = ['Budget (≤ €400K)', 'Mid-Range (€400-700K)', 
            'Premium (€700K-€1M)', 'Luxury (> €1M)']
segment_colors = ['#FF9999', '#66B2FF', '#99FF99', '#FFB366']

# Calculate counts for each segment in each neighborhood
segment_counts = {}
for segment in segments:
    segment_counts[segment] = []
    for nb in top_neighborhoods:
        count = len(df_top[(df_top['Neighborhood'] == nb) & 
                          (df_top['Price_Segment'] == segment)])
        segment_counts[segment].append(count)

# Create stacked bars
bottom_values = np.zeros(len(top_neighborhoods))

for i, (segment, color) in enumerate(zip(segments, segment_colors)):
    counts = segment_counts[segment]
    bars = plt.bar(top_neighborhoods, counts, bottom=bottom_values,
                   color=color, edgecolor='white', linewidth=0.5,
                   label=segment, alpha=0.8)
    
    # Add segment percentage labels (only if segment > 10% of total)
    for j, (bar, count, bottom) in enumerate(zip(bars, counts, bottom_values)):
        total = sum(segment_counts[s][j] for s in segments)
        percentage = (count / total) * 100 if total > 0 else 0
        
        if percentage >= 15:  # Only show significant segments
            y_pos = bottom + count / 2
            plt.text(bar.get_x() + bar.get_width()/2, y_pos,
                     f'{percentage:.0f}%', ha='center', va='center',
                     fontsize=9, fontweight='bold', color='black')
    
    bottom_values += counts

# Add labels and title
plt.xlabel('Neighborhood', fontsize=12, fontweight='bold')
plt.ylabel('Number of Houses', fontsize=12, fontweight='bold')
plt.title('Price Segment Composition by Neighborhood (Stacked Bar Chart)', 
          fontsize=16, fontweight='bold')

# Rotate x-tick labels for better readability
plt.xticks(rotation=30, ha='right')

# Add legend with custom positioning
plt.legend(title='Price Segments', title_fontsize=11,
           fontsize=10, loc='upper left', bbox_to_anchor=(1, 1))

# Add total count labels on top of each stack
for i, nb in enumerate(top_neighborhoods):
    total = sum(segment_counts[s][i] for s in segments)
    plt.text(i, total + 5, f'{total:,}', ha='center', va='bottom',
             fontsize=10, fontweight='bold', color='darkblue')

# Add grid
plt.grid(True, alpha=0.2, axis='y', linestyle='--')

# Add market composition analysis
market_totals = {}
for segment in segments:
    market_totals[segment] = len(df[df['Price_Segment'] == segment])

total_houses = len(df)
composition_text = "Overall Market Composition:\n"
for segment in segments:
    percentage = (market_totals[segment] / total_houses) * 100
    composition_text += f"• {segment}: {percentage:.1f}%\n"

plt.text(0.02, 0.98, composition_text, transform=plt.gca().transAxes,
         fontsize=10, verticalalignment='top',
         bbox=dict(boxstyle='round', facecolor='lightyellow', alpha=0.9))

plt.tight_layout()
plt.show()

print("🎯 MARKET SEGMENT ANALYSIS:")
print("="*45)

# Create neighborhood segment analysis
print("\nNeighborhood Segment Profiles:")
print("-"*40)

for nb in top_neighborhoods:
    nb_data = df_top[df_top['Neighborhood'] == nb]
    total = len(nb_data)
    
    print(f"\n{nb} ({total:,} houses):")
    
    for segment in segments:
        count = len(nb_data[nb_data['Price_Segment'] == segment])
        percentage = (count / total) * 100
        
        if percentage >= 10:  # Only show significant segments
            bar = '█' * int(percentage / 5)  # Visual bar
            print(f"  {segment[:15]:<20} {bar} {percentage:5.1f}% ({count:,})")
    
    # Identify dominant segment
    dominant_segment = max(segments, 
                          key=lambda s: len(nb_data[nb_data['Price_Segment'] == s]))
    dominant_percentage = len(nb_data[nb_data['Price_Segment'] == dominant_segment]) / total * 100
    
    if dominant_percentage > 40:
        print(f"  🏆 Dominant: {dominant_segment} ({dominant_percentage:.1f}%)")

print("\n💡 Market Positioning Insights:")
print("1. Some neighborhoods are clearly budget-focused")
print("2. Others have balanced segment distribution")
print("3. Luxury concentration varies significantly")
print("4. Market segmentation helps target marketing")
print("5. Can identify underserved segments per neighborhood")

Method 5: Percentage Bar Chart (Normalized Comparison)

Compare proportions rather than absolute counts:

# METHOD 5: Percentage bar chart - Room distribution by house type
plt.figure(figsize=(14, 8))

# Prepare data: Room distribution by house type
house_types = df['Type'].unique()
room_counts = sorted(df['Room'].dropna().unique())
room_counts = [r for r in room_counts if r <= 6]  # Limit to 1-6 rooms

# Create matrix: house_type × room_count
distribution_matrix = np.zeros((len(house_types), len(room_counts)))

for i, h_type in enumerate(house_types):
    type_data = df[df['Type'] == h_type]
    for j, rooms in enumerate(room_counts):
        count = len(type_data[type_data['Room'] == rooms])
        distribution_matrix[i, j] = count

# Convert to percentages
percent_matrix = distribution_matrix / distribution_matrix.sum(axis=1, keepdims=True) * 100

# Set up positions
x = np.arange(len(house_types))
width = 0.8  # Total width for all room segments
segment_width = width / len(room_counts)

# Create color palette
colors = plt.cm.Set3(np.linspace(0, 1, len(room_counts)))

# Create stacked percentage bars
bottom_values = np.zeros(len(house_types))

for i, (rooms, color) in enumerate(zip(room_counts, colors)):
    percentages = percent_matrix[:, i]
    bars = plt.bar(x, percentages, segment_width, 
                   bottom=bottom_values, color=color,
                   edgecolor='white', linewidth=0.5,
                   label=f'{rooms} Room{"s" if rooms != 1 else ""}')
    
    # Add room count labels (only if percentage > 5%)
    for j, (bar, percentage, bottom) in enumerate(zip(bars, percentages, bottom_values)):
        if percentage >= 8:  # Only label significant segments
            y_pos = bottom + percentage / 2
            plt.text(bar.get_x() + bar.get_width()/2, y_pos,
                     f'{rooms}', ha='center', va='center',
                     fontsize=9, fontweight='bold', color='black')
    
    bottom_values += percentages

# Add labels and title
plt.xlabel('House Type', fontsize=12, fontweight='bold')
plt.ylabel('Percentage of Houses', fontsize=12, fontweight='bold')
plt.title('Room Distribution by House Type (Percentage Bar Chart)', 
          fontsize=16, fontweight='bold')
plt.xticks(x, house_types)
plt.ylim(0, 100)

# Add legend with room counts
plt.legend(title='Number of Rooms', title_fontsize=11,
           fontsize=10, loc='upper left', bbox_to_anchor=(1, 1))

# Add total count annotation for each house type
for i, h_type in enumerate(house_types):
    total_count = int(distribution_matrix[i].sum())
    plt.text(i, 102, f'n={total_count:,}', ha='center', va='bottom',
             fontsize=10, fontweight='bold', color='darkred')

# Add grid
plt.grid(True, alpha=0.2, axis='y', linestyle=':')

# Add statistical insights
insights = [
    "• Apartments: Mostly 1-3 rooms",
    "• Houses: Concentrated in 3-5 rooms",
    "• Maisons: More rooms on average",
    "• Room count correlates with property type"
]

for i, insight in enumerate(insights):
    plt.text(0.02, 0.95 - i*0.05, insight, transform=plt.gca().transAxes,
             fontsize=9, verticalalignment='top',
             bbox=dict(boxstyle='round', facecolor='lightblue', alpha=0.3))

plt.tight_layout()
plt.show()

print("🚪 ROOM DISTRIBUTION ANALYSIS:")
print("="*45)

# Create detailed distribution table
print("\nRoom Distribution by House Type (%):")
print("-"*50)

# Header
header = "Type       " + "".join([f"{rooms:>8} " for rooms in room_counts]) + "  Total"
print(header)
print("-" * len(header))

# Data rows
for i, h_type in enumerate(house_types):
    row = f"{h_type:<10}"
    for j, rooms in enumerate(room_counts):
        percentage = percent_matrix[i, j]
        if percentage >= 1:
            row += f"{percentage:>7.1f}% "
        else:
            row += "        "
    
    total_count = int(distribution_matrix[i].sum())
    row += f"  {total_count:>6,}"
    print(row)

# Summary statistics
print("\n📊 Room Count Statistics:")
print("-"*30)

room_stats = []
for h_type in house_types:
    type_data = df[df['Type'] == h_type]['Room'].dropna()
    room_stats.append({
        'Type': h_type,
        'Avg Rooms': f'{type_data.mean():.1f}',
        'Median Rooms': f'{type_data.median():.0f}',
        'Min Rooms': f'{type_data.min():.0f}',
        'Max Rooms': f'{type_data.max():.0f}',
        'Most Common': f'{type_data.mode().iloc[0]:.0f} rooms'
    })

room_stats_df = pd.DataFrame(room_stats)
print(room_stats_df.to_string(index=False))

print("\n💡 Property Type Insights:")
print("1. Apartments: Compact living (1-3 rooms typical)")
print("2. Houses: Family-oriented (3-5 rooms typical)")
print("3. Maisons: Spacious living (more rooms)")
print("4. Room count strongly indicates property type")
print("5. Market gaps: Few 6+ room apartments")

Method 6: Advanced Bar Chart with Error Bars

Add statistical confidence to your comparisons:

# METHOD 6: Bar chart with error bars - Price statistics with confidence
plt.figure(figsize=(16, 9))

# Analyze price by number of bedrooms
bedroom_counts = sorted(df['Bedroom'].dropna().unique())
bedroom_counts = [b for b in bedroom_counts if 1 <= b <= 5]  # Limit to 1-5

# Calculate statistics for each bedroom count
avg_prices = []
median_prices = []
std_errors = []
conf_intervals = []
sample_sizes = []

for bedrooms in bedroom_counts:
    price_data = df[df['Bedroom'] == bedrooms]['Price'].values
    
    if len(price_data) >= 10:  # Only include with sufficient data
        avg = np.mean(price_data)
        median = np.median(price_data)
        std = np.std(price_data)
        n = len(price_data)
        
        avg_prices.append(avg)
        median_prices.append(median)
        
        # Standard error
        se = std / np.sqrt(n)
        std_errors.append(se)
        
        # 95% confidence interval (1.96 * SE)
        ci = 1.96 * se
        conf_intervals.append(ci)
        sample_sizes.append(n)

# Create positions
x = np.arange(len(bedroom_counts))
width = 0.35

# Plot average prices with error bars
fig, (ax1, ax2) = plt.subplots(2, 1, figsize=(14, 12), 
                               gridspec_kw={'height_ratios': [2, 1]})

# Plot 1: Average prices with confidence intervals
bars1 = ax1.bar(x - width/2, avg_prices, width, 
                yerr=conf_intervals, capsize=8,
                label='Average Price (±95% CI)',
                color='#4ECDC4', alpha=0.8,
                error_kw=dict(ecolor='red', elinewidth=2, capthick=2))

# Plot median prices
bars2 = ax1.bar(x + width/2, median_prices, width,
                label='Median Price',
                color='#FF6B6B', alpha=0.8)

# Customize plot 1
ax1.set_xlabel('Number of Bedrooms', fontsize=12, fontweight='bold')
ax1.set_ylabel('Price (Euros)', fontsize=12, fontweight='bold')
ax1.set_title('House Prices by Number of Bedrooms\nWith Confidence Intervals', 
              fontsize=16, fontweight='bold')
ax1.set_xticks(x)
ax1.set_xticklabels([f'{b} Bedroom{"s" if b != 1 else ""}' for b in bedroom_counts])
ax1.yaxis.set_major_formatter(
    plt.FuncFormatter(lambda y, p: f'€{y/1000:.0f}K'))
ax1.legend(fontsize=10)
ax1.grid(True, alpha=0.2, axis='y')

# Add value labels
for bars in [bars1, bars2]:
    for bar in bars:
        height = bar.get_height()
        ax1.text(bar.get_x() + bar.get_width()/2, height + 10000,
                 f'€{height/1000:.0f}K', ha='center', va='bottom',
                 fontsize=9, fontweight='bold')

# Add sample size annotations
for i, n in enumerate(sample_sizes):
    ax1.text(i, ax1.get_ylim()[0] + 50000, f'n={n:,}',
             ha='center', va='bottom', fontsize=9,
             bbox=dict(boxstyle='round', facecolor='white', alpha=0.7))

# Add statistical insights
price_growth = []
for i in range(1, len(avg_prices)):
    growth = (avg_prices[i] - avg_prices[i-1]) / avg_prices[i-1] * 100
    price_growth.append(growth)

growth_text = "Price Increase per Additional Bedroom:\n"
for i, growth in enumerate(price_growth):
    growth_text += f"• {bedroom_counts[i]}→{bedroom_counts[i+1]}: +{growth:.1f}%\n"

ax1.text(0.02, 0.98, growth_text, transform=ax1.transAxes,
         fontsize=9, verticalalignment='top',
         bbox=dict(boxstyle='round', facecolor='lightyellow', alpha=0.8))

# Plot 2: Standard errors and sample sizes
bar_colors = plt.cm.viridis(np.linspace(0.2, 0.8, len(bedroom_counts)))

# Create bar for standard errors
bars_se = ax2.bar(x, std_errors, color=bar_colors, alpha=0.7,
                  edgecolor='black', linewidth=1)

ax2.set_xlabel('Number of Bedrooms', fontsize=12, fontweight='bold')
ax2.set_ylabel('Standard Error (Euros)', fontsize=12, fontweight='bold')
ax2.set_title('Price Variability by Bedroom Count', 
              fontsize=14, fontweight='bold')
ax2.set_xticks(x)
ax2.set_xticklabels(bedroom_counts)
ax2.yaxis.set_major_formatter(
    plt.FuncFormatter(lambda y, p: f'€{y/1000:.1f}K'))
ax2.grid(True, alpha=0.2, axis='y')

# Add sample size labels on secondary y-axis
ax2_twin = ax2.twinx()
ax2_twin.plot(x, sample_sizes, 'ro-', linewidth=2, markersize=8,
              label='Sample Size')
ax2_twin.set_ylabel('Sample Size (Number of Houses)', 
                    fontsize=12, fontweight='bold', color='red')
ax2_twin.tick_params(axis='y', labelcolor='red')

# Combine legends
lines1, labels1 = ax2.get_legend_handles_labels()
lines2, labels2 = ax2_twin.get_legend_handles_labels()
ax2.legend(lines1 + lines2, labels1 + labels2, loc='upper left')

plt.tight_layout()
plt.show()

print("📈 STATISTICAL BAR CHART ANALYSIS:")
print("="*50)

# Create detailed statistical table
print("\nPrice Statistics by Number of Bedrooms:")
print("-"*55)

stat_table = []
for i, bedrooms in enumerate(bedroom_counts):
    stat_table.append({
        'Bedrooms': bedrooms,
        'Sample Size': f'{sample_sizes[i]:,}',
        'Average Price': f'€{avg_prices[i]/1000:.1f}K',
        'Median Price': f'€{median_prices[i]/1000:.1f}K',
        'Std Error': f'€{std_errors[i]/1000:.1f}K',
        '95% CI (±)': f'€{conf_intervals[i]/1000:.1f}K',
        'Price/BR': f'€{avg_prices[i]/bedrooms/1000:.1f}K'
    })

stat_df = pd.DataFrame(stat_table)
print(stat_df.to_string(index=False))

print("\n🔍 Statistical Insights:")
print("1. Confidence intervals show estimate precision")
print("2. Larger samples = smaller confidence intervals")
print("3. Median vs Average shows skewness in each category")
print("4. Standard error indicates variability within category")
print("5. Price per bedroom decreases with more bedrooms")

# Calculate statistical significance
print("\n🎯 Statistical Significance Check:")
print("-"*35)

for i in range(len(bedroom_counts)-1):
    # Simple overlap check for confidence intervals
    ci1_lower = avg_prices[i] - conf_intervals[i]
    ci1_upper = avg_prices[i] + conf_intervals[i]
    ci2_lower = avg_prices[i+1] - conf_intervals[i+1]
    ci2_upper = avg_prices[i+1] + conf_intervals[i+1]
    
    if ci1_upper < ci2_lower or ci2_upper < ci1_lower:
        significance = "Significantly different"
    else:
        significance = "Not significantly different"
    
    print(f"{bedroom_counts[i]} vs {bedroom_counts[i+1]} bedrooms: {significance}")

Bar Chart Cheat Sheet (Quick Reference)

# 1. BASIC BAR CHARTS
plt.bar(x, height)                    # Vertical bar chart
plt.barh(y, width)                    # Horizontal bar chart

# 2. BAR CUSTOMIZATION
plt.bar(x, height, color='blue')                 # Single color
plt.bar(x, height, color=['red', 'green', 'blue']) # Different colors
plt.bar(x, height, edgecolor='black', linewidth=2) # Border
plt.bar(x, height, alpha=0.7)                    # Transparency
plt.bar(x, height, hatch='//')                   # Pattern

# 3. GROUPED BAR CHARTS
width = 0.35
x = np.arange(len(categories))
plt.bar(x - width/2, data1, width, label='Group 1')
plt.bar(x + width/2, data2, width, label='Group 2')
plt.xticks(x, categories)

# 4. STACKED BAR CHARTS
plt.bar(categories, bottom_data, label='Bottom')
plt.bar(categories, top_data, bottom=bottom_data, label='Top')

# 5. ERROR BARS
plt.bar(x, y, yerr=error_values, capsize=5)     # Vertical error
plt.barh(y, width, xerr=error_values, capsize=5) # Horizontal error

# 6. VALUE LABELS
for bar in bars:
    height = bar.get_height()
    plt.text(bar.get_x() + bar.get_width()/2., height,
             f'{height}', ha='center', va='bottom')

# 7. BAR WIDTH AND SPACING
plt.bar(x, height, width=0.8)          # Bar width (default 0.8)
plt.bar(x, height, align='center')     # Bar alignment
plt.bar(x, height, align='edge')       # Alternative alignment

# 8. PANDAS INTEGRATION
df['column'].value_counts().plot(kind='bar')          # Count plot
df.groupby('category')['value'].mean().plot(kind='bar') # Mean plot
df.plot(kind='bar', stacked=True)                     # Stacked plot

# 9. SEABORN BAR CHARTS
import seaborn as sns
sns.barplot(x='category', y='value', data=df)          # Basic
sns.barplot(x='category', y='value', data=df, hue='group') # Grouped
sns.barplot(x='category', y='value', data=df, ci=95)   # Confidence

# 10. PROFESSIONAL FORMATTING
plt.bar(x, height, color=plt.cm.viridis(np.linspace(0, 1, len(x))))
plt.gca().yaxis.set_major_formatter(plt.FuncFormatter(lambda x, p: f'€{x:,.0f}'))
plt.xticks(rotation=45, ha='right')
plt.grid(True, alpha=0.3, axis='y')

Common Bar Chart Mistakes to Avoid ⚠️

❌ COMMON MISTAKES:

  • Truncated y-axis: Starting above zero distorts comparisons
  • Too many categories: More than 10-12 bars becomes unreadable
  • Random ordering: Not sorting by value when appropriate
  • Unreadable labels: Long category names overlapping
  • 3D effects: Distorts perception of bar heights
  • Inconsistent colors: Same color for unrelated categories
  • Missing labels: No value labels when exact numbers matter
  • Wrong chart type: Using bar chart for time series (use line)

✅ BEST PRACTICES:

  • Always start at zero: Maintain proportional accuracy
  • Sort meaningfully: By value (ranking) or category (alphabetical)
  • Use horizontal bars: For long category labels or many categories
  • Add value labels: Especially for important comparisons
  • Limit color palette: Use colors consistently across charts
  • Include sample sizes: Show n for each category
  • Consider log scale: When data spans multiple orders of magnitude
  • Add context: Benchmark lines, averages, or targets

When to Use Different Bar Chart Types

🎯 CHOOSING THE RIGHT BAR CHART:

📊 Vertical Bar Chart:

  • 4-8 categories
  • Short category names
  • Time-based data (months)
  • When height comparison is natural

↔️ Horizontal Bar Chart:

  • 8+ categories
  • Long category names
  • Ranking comparisons
  • When reading left-to-right is intuitive

👥 Grouped Bar Chart:

  • Comparing sub-categories
  • 2-3 groups per category
  • Side-by-side comparison needed
  • When totals aren't the focus

🗃️ Stacked Bar Chart:

  • Showing part-to-whole
  • Comparing compositions
  • When totals matter
  • 4-6 segments per bar max

💡 Pro Tip: Use horizontal bars for ranking, vertical for time series, grouped for comparison, stacked for composition, and percentage bars for proportional analysis.

Your Bar Chart Mastery Checklist

🎓 WHAT YOU'VE LEARNED:

✅ Can create basic vertical/horizontal bar charts
✅ Understand bar vs histogram differences
✅ Can create grouped and stacked bar charts
✅ Know how to add value labels and annotations
✅ Can create percentage/100% stacked charts
✅ Understand how to add error bars and confidence intervals
✅ Can create publication-ready bar charts
✅ Know when to use each bar chart type

Remember: Bar charts are one of the most versatile and widely understood visualization types. They excel at comparing categorical data, showing rankings, and displaying proportions. With the Amsterdam house price data, bar charts help you quickly identify market leaders, understand market segmentation, and compare different neighborhoods or property types at a glance.

Happy bar charting! 📊🎯

Comments