📌 Dataset Source: To follow along with this tutorial, download the Amsterdam House Prices Dataset from Kaggle. You can access it directly here:
https://www.kaggle.com/datasets/thomasnibb/amsterdam-house-price-data
"Violin plots are the symphony of data visualization - combining the precision of boxplots with the elegance of density plots to reveal the full story of your data."
What is a Violin Plot? (The Data Orchestra!)
A violin plot combines a boxplot and kernel density plot into a single visualization. It shows the full distribution of the data, revealing patterns that boxplots alone cannot. It answers questions like:
- 🎻 What's the data density at different values? (Where are houses most concentrated?)
- 📊 Is the distribution bimodal or multimodal? (Multiple price peaks?)
- 📏 How does the distribution shape compare across groups?
- 🔍 Are there gaps or unusual patterns in the data?
📖 Simple Analogy:
Imagine a boxplot got married to a density plot, and their child was a violin plot:
• Boxplot part: Shows median, quartiles, and outliers (the skeleton)
• Density part: Shows how data points are distributed (the flesh)
• Width at any point: Indicates density of data at that value
• Symmetry: Mirror image shows smoothed distribution
Why Violin Plots? The Superpowers Over Boxplots
🎯 VIOLIN PLOT ADVANTAGES:
📦 Boxplot Limitations:
- Hides distribution shape
- Can't show multimodality
- Loses density information
- Shows only 5-number summary
🎻 Violin Plot Advantages:
- Shows full distribution shape
- Reveals multiple peaks
- Preserves density information
- Combines summary + detail
Essential Violin Plot Terms - Explained Simply
Before we dive into code, let's understand the key components of a violin plot:
1. Kernel Density Estimation (KDE) - The Magic Behind
What it is: A non-parametric way to estimate probability density
Analogy: Like smoothing a histogram with a moving window
Bandwidth: Controls smoothness (larger = smoother)
Why it matters: Creates the smooth "violin" shape from raw data
# Understanding KDE
import seaborn as sns
import numpy as np
# Generate sample data
sample_data = df['Price'].sample(1000, random_state=42)
# Plot KDE curve
plt.figure(figsize=(10, 4))
sns.kdeplot(sample_data, bw_adjust=0.5, label='Bandwidth = 0.5')
sns.kdeplot(sample_data, bw_adjust=1.0, label='Bandwidth = 1.0')
sns.kdeplot(sample_data, bw_adjust=2.0, label='Bandwidth = 2.0')
plt.title('KDE with Different Bandwidths')
plt.xlabel('Price (Euros)')
plt.ylabel('Density')
plt.legend()
plt.show()
print("💡 Bandwidth controls smoothness:")
print("• Small bandwidth: Follows data closely (can be noisy)")
print("• Large bandwidth: Very smooth (can oversmooth)")
print("• Default: Usually optimal for most datasets")
2. Violin Body - The Distribution Shape
What it is: The smoothed density plot mirrored around center
Width: Proportional to data density at that value
Shape reveals: Modality, skewness, kurtosis, gaps
Key patterns: Unimodal, bimodal, uniform, skewed
# Different distribution shapes in violin plots
fig, axes = plt.subplots(2, 3, figsize=(15, 8))
# Generate different distribution shapes
np.random.seed(42)
data_shapes = {
'Normal': np.random.normal(0, 1, 1000),
'Bimodal': np.concatenate([np.random.normal(-2, 1, 500),
np.random.normal(2, 1, 500)]),
'Right-Skewed': np.random.exponential(1, 1000),
'Uniform': np.random.uniform(-3, 3, 1000),
'Multimodal': np.concatenate([np.random.normal(-3, 0.5, 250),
np.random.normal(0, 0.5, 500),
np.random.normal(3, 0.5, 250)]),
'Outlier-Heavy': np.concatenate([np.random.normal(0, 1, 950),
np.random.normal(10, 0.5, 50)])
}
# Plot each distribution
for ax, (name, data) in zip(axes.flat, data_shapes.items()):
ax.violinplot(data, showmeans=True, showmedians=True)
ax.set_title(name, fontweight='bold')
ax.set_xticks([])
ax.grid(True, alpha=0.3)
plt.tight_layout()
plt.show()
print("🎻 VIOLIN SHAPE INTERPRETATION:")
print("="*45)
print("1. Narrow middle, fat ends: Bimodal distribution")
print("2. Fat right tail: Right-skewed data")
print("3. Symmetrical: Normal distribution")
print("4. Flat and wide: Uniform distribution")
print("5. Multiple peaks: Multimodal distribution")
print("6. Long thin tail with bulge: Outlier cluster")
3. Inner Elements - Boxplot Inside Violin
What they are: Statistical markers inside the violin
Common options: Median, mean, quartiles, min/max
Box inside: Can show IQR box within violin
Why it matters: Combines summary statistics with density
# Different inner plot options
fig, axes = plt.subplots(2, 2, figsize=(12, 10))
# Data for demonstration
sample_prices = df['Price'].sample(500, random_state=42).values
# Plot 1: Default (box)
axes[0, 0].violinplot(sample_prices, showmeans=False, showmedians=True)
axes[0, 0].set_title('Default (Median only)', fontweight='bold')
axes[0, 0].set_ylabel('Price')
# Plot 2: Show mean and median
axes[0, 1].violinplot(sample_prices, showmeans=True, showmedians=True)
axes[0, 1].set_title('Mean + Median', fontweight='bold')
# Plot 3: Show quartiles
axes[1, 0].violinplot(sample_prices, showmeans=True, showmedians=True,
showextrema=True, quantiles=[[0.25, 0.75]])
axes[1, 0].set_title('Quartiles (25th & 75th)', fontweight='bold')
axes[1, 0].set_ylabel('Price')
axes[1, 0].set_xlabel('Different inner displays')
# Plot 4: Minimal (just violin)
axes[1, 1].violinplot(sample_prices, showmeans=False, showmedians=False,
showextrema=False)
axes[1, 1].set_title('Violin Only (No markers)', fontweight='bold')
axes[1, 1].set_xlabel('Different inner displays')
plt.tight_layout()
plt.show()
print("📊 INNER ELEMENT OPTIONS:")
print("="*45)
print("showmeans=True - Shows mean as a point/line")
print("showmedians=True - Shows median line")
print("showextrema=True - Shows min/max whiskers")
print("quantiles=[...] - Shows custom percentiles")
print("points=100 - Shows individual data points")
Method 1: Your First Violin Plot (Basic)
Let's create a basic violin plot to see house price distribution:
# METHOD 1: Basic violin plot
plt.figure(figsize=(12, 7))
# Create basic violin plot
vp = plt.violinplot(df['Price'].values, showmeans=True, showmedians=True)
# Customize violin color
for pc in vp['bodies']:
pc.set_facecolor('#6A0572')
pc.set_alpha(0.7)
pc.set_edgecolor('black')
# Customize median and mean lines
vp['cmedians'].set_color('yellow')
vp['cmedians'].set_linewidth(2)
vp['cmeans'].set_color('red')
vp['cmeans'].set_linewidth(2)
# Add labels and title
plt.ylabel('House Price (Euros)', fontsize=12, fontweight='bold')
plt.title('Amsterdam House Price Distribution (Violin Plot)',
fontsize=14, fontweight='bold')
# Format y-axis with Euro symbols
plt.gca().yaxis.set_major_formatter(
plt.FuncFormatter(lambda x, p: f'€{x/1000:.0f}K'))
# Add grid for readability
plt.grid(True, alpha=0.3, axis='y')
# Add legend
plt.legend([vp['bodies'][0], vp['cmedians'], vp['cmeans']],
['Density', 'Median', 'Mean'],
loc='upper right', fontsize=10)
# Add text with key insights
price_data = df['Price']
Q1 = price_data.quantile(0.25)
median = price_data.median()
Q3 = price_data.quantile(0.75)
skewness = price_data.skew()
insight_text = f"""Key Insights:
• Median: €{median/1000:.0f}K
• IQR: €{Q1/1000:.0f}K - €{Q3/1000:.0f}K
• Skewness: {skewness:.2f} (right-skewed)
• Note wide tail at high prices"""
plt.text(0.02, 0.98, insight_text, transform=plt.gca().transAxes,
fontsize=10, verticalalignment='top',
bbox=dict(boxstyle='round', facecolor='wheat', alpha=0.8))
plt.tight_layout()
plt.show()
print("🎻 FIRST VIOLIN PLOT ANALYSIS:")
print("="*45)
print("What we can see that a boxplot wouldn't show:")
print("1. Density concentration around €300-500K")
print("2. Smooth tapering to very high prices (long right tail)")
print("3. No secondary peaks (unimodal distribution)")
print("4. Most houses cluster in lower price ranges")
print("5. Gradual decrease in density as price increases")
print("\n💡 Key Observations:")
print(f"• {len(df[df['Price'] <= 500000])/len(df)*100:.1f}% of houses ≤ €500K")
print(f"• {len(df[df['Price'] > 1000000])/len(df)*100:.1f}% of houses > €1M")
print("• Distribution is unimodal but highly right-skewed")
print("• Luxury market (>€1M) forms a long thin tail")
✅ What You Just Learned:
plt.violinplot(data)- Creates basic violin plotshowmeans=True/showmedians=True- Adds statistical markersvp['bodies']- Access violin bodies for customization- How to interpret violin width as data density
- How violin plots reveal distribution shape better than boxplots
Method 2: Multiple Violin Plots for Comparison
Compare distributions across different categories:
# METHOD 2: Multiple violin plots by number of rooms
plt.figure(figsize=(14, 8))
# Prepare data for different room counts
room_data = []
room_labels = []
# Get room counts with sufficient data
room_counts = sorted(df['Room'].dropna().unique())
for rooms in room_counts:
if rooms <= 6: # Limit to reasonable room counts
room_prices = df[df['Room'] == rooms]['Price']
if len(room_prices) >= 20: # Only include if enough data
room_data.append(room_prices.values)
room_labels.append(f'{rooms} Room{"s" if rooms != 1 else ""}')
# Create positions for violins
positions = np.arange(1, len(room_data) + 1)
# Create violin plot
vp = plt.violinplot(room_data, positions=positions,
showmeans=True, showmedians=True)
# Customize colors using a colormap
colors = plt.cm.Set3(np.linspace(0, 1, len(room_data)))
for i, pc in enumerate(vp['bodies']):
pc.set_facecolor(colors[i])
pc.set_alpha(0.8)
pc.set_edgecolor('black')
pc.set_linewidth(1)
# Customize statistical lines
vp['cmedians'].set_color('black')
vp['cmedians'].set_linewidth(2)
vp['cmeans'].set_color('darkred')
vp['cmeans'].set_linewidth(1.5)
vp['cmeans'].set_linestyle('--')
# Add labels and title
plt.ylabel('House Price (Euros)', fontsize=12, fontweight='bold')
plt.xlabel('Number of Rooms', fontsize=12, fontweight='bold')
plt.title('House Price Distribution by Number of Rooms',
fontsize=16, fontweight='bold')
plt.xticks(positions, room_labels)
# Format y-axis
plt.gca().yaxis.set_major_formatter(
plt.FuncFormatter(lambda x, p: f'€{x/1000:.0f}K'))
# Add grid
plt.grid(True, alpha=0.2, axis='y', linestyle=':')
# Add mean price annotations
for i, (pos, data) in enumerate(zip(positions, room_data)):
mean_price = np.mean(data)
median_price = np.median(data)
# Annotate mean
plt.text(pos, mean_price + 50000, f'€{mean_price/1000:.0f}K',
ha='center', va='bottom', fontsize=9,
bbox=dict(boxstyle='round', facecolor='white', alpha=0.7))
# Annotate median
plt.text(pos, median_price - 50000, f'€{median_price/1000:.0f}K',
ha='center', va='top', fontsize=8, color='darkgreen',
bbox=dict(boxstyle='round', facecolor='lightgreen', alpha=0.5))
# Add legend
from matplotlib.patches import Patch
legend_elements = [
Patch(facecolor='gray', alpha=0.5, label='Density Distribution'),
Patch(facecolor='none', edgecolor='black', label='Median (solid)'),
Patch(facecolor='none', edgecolor='darkred', linestyle='--',
label='Mean (dashed)')
]
plt.legend(handles=legend_elements, loc='upper left', fontsize=10)
plt.tight_layout()
plt.show()
print("🏠 ROOM COUNT ANALYSIS:")
print("="*45)
# Create analysis table
analysis_data = []
for i, (rooms, data) in enumerate(zip(room_counts, room_data)):
if rooms <= 6 and len(df[df['Room'] == rooms]) >= 20:
room_df = df[df['Room'] == rooms]
analysis_data.append({
'Rooms': rooms,
'Count': len(room_df),
'Median Price': f'€{room_df["Price"].median()/1000:.0f}K',
'Mean Price': f'€{room_df["Price"].mean()/1000:.0f}K',
'Price Range': f'€{room_df["Price"].min()/1000:.0f}K-€{room_df["Price"].max()/1000:.0f}K',
'Skewness': f'{room_df["Price"].skew():.2f}'
})
analysis_df = pd.DataFrame(analysis_data)
print(analysis_df.to_string(index=False))
print("\n💡 Distribution Shape Insights:")
print("1. 1-2 rooms: Right-skewed (many cheap apartments)")
print("2. 3-4 rooms: More symmetric (family homes)")
print("3. 5+ rooms: Wide distribution (luxury variation)")
print("4. Notice how shape changes with room count")
print("5. Width indicates price variability in each category")
Method 3: Split Violin Plots (Compare Groups)
Split violin plots compare two groups within each category:
# METHOD 3: Split violin plots using Seaborn
plt.figure(figsize=(15, 8))
# Create a sample for demonstration (full dataset might be too large)
sample_df = df.sample(2000, random_state=42).copy()
# Create binary category for demonstration (e.g., Old vs New)
sample_df['Is_Old'] = sample_df['Year'] < 1950
sample_df['Era'] = sample_df['Is_Old'].map({True: 'Pre-1950', False: 'Post-1950'})
# Get top neighborhoods for comparison
top_neighborhoods = sample_df['Neighborhood'].value_counts().head(6).index.tolist()
sample_filtered = sample_df[sample_df['Neighborhood'].isin(top_neighborhoods)]
# Create split violin plot using Seaborn
ax = sns.violinplot(x='Neighborhood', y='Price', hue='Era',
data=sample_filtered, split=True,
palette={'Pre-1950': '#FF6B6B', 'Post-1950': '#4ECDC4'},
inner='quartile', # Show quartiles inside
scale='width', # Scale violins by data count
bw=0.6) # Bandwidth adjustment
# Customize plot
plt.title('House Price Distribution: Old vs New by Neighborhood\n(Split Violin Plot)',
fontsize=16, fontweight='bold', pad=20)
plt.xlabel('Neighborhood', fontsize=12, fontweight='bold')
plt.ylabel('House Price (Euros)', fontsize=12, fontweight='bold')
# Format y-axis
ax.yaxis.set_major_formatter(
plt.FuncFormatter(lambda x, p: f'€{x/1000:.0f}K'))
# Rotate x-tick labels for better readability
plt.xticks(rotation=30, ha='right')
# Add grid
plt.grid(True, alpha=0.2, axis='y', linestyle='--')
# Add legend with custom position
plt.legend(title='Construction Era', title_fontsize=11,
fontsize=10, loc='upper left')
# Add statistics annotations
for i, neighborhood in enumerate(top_neighborhoods):
neighborhood_data = sample_filtered[sample_filtered['Neighborhood'] == neighborhood]
if len(neighborhood_data) > 0:
# Calculate price difference between eras
old_median = neighborhood_data[neighborhood_data['Era'] == 'Pre-1950']['Price'].median()
new_median = neighborhood_data[neighborhood_data['Era'] == 'Post-1950']['Price'].median()
if not pd.isna(old_median) and not pd.isna(new_median):
price_diff = new_median - old_median
diff_percent = (price_diff / old_median) * 100
# Add annotation if difference is significant
if abs(diff_percent) > 10:
color = 'green' if price_diff > 0 else 'red'
symbol = '↑' if price_diff > 0 else '↓'
plt.text(i, ax.get_ylim()[1] * 0.95,
f'{symbol} {abs(diff_percent):.0f}%',
ha='center', va='bottom', fontsize=9,
color=color, fontweight='bold',
bbox=dict(boxstyle='round', facecolor='white', alpha=0.8))
plt.tight_layout()
plt.show()
print("🏛️ SPLIT VIOLIN PLOT ANALYSIS:")
print("="*50)
print("\nWhat Split Violin Plots Reveal:")
print("1. Distribution shape differences between groups")
print("2. Density comparisons side by side")
print("3. Median/mean shifts between categories")
print("4. Overlap or separation in distributions")
print("\n🔍 Key Questions Answered:")
print("• Which neighborhoods have biggest old/new price gaps?")
print("• Are old houses consistently cheaper?")
print("• Which neighborhoods preserve historical value?")
print("• Where is the luxury market concentrated (old vs new)?")
# Calculate detailed statistics
print("\n📊 Price Premium Analysis (New vs Old):")
print("-"*40)
premium_data = []
for neighborhood in top_neighborhoods:
nb_data = sample_filtered[sample_filtered['Neighborhood'] == neighborhood]
old_houses = nb_data[nb_data['Era'] == 'Pre-1950']
new_houses = nb_data[nb_data['Era'] == 'Post-1950']
if len(old_houses) > 5 and len(new_houses) > 5:
old_median = old_houses['Price'].median()
new_median = new_houses['Price'].median()
premium_data.append({
'Neighborhood': neighborhood,
'Old Median': f'€{old_median/1000:.0f}K',
'New Median': f'€{new_median/1000:.0f}K',
'Price Diff': f'€{(new_median - old_median)/1000:.0f}K',
'Premium %': f'{(new_median - old_median)/old_median*100:.0f}%',
'Old Count': len(old_houses),
'New Count': len(new_houses)
})
premium_df = pd.DataFrame(premium_data)
print(premium_df.to_string(index=False))
Method 4: Half Violin Plots (Raincloud Alternative)
Half violin plots (raincloud plots) combine multiple visualization types:
# METHOD 4: Half-violin plots (Raincloud style)
fig, (ax1, ax2, ax3) = plt.subplots(3, 1, figsize=(14, 12))
# Sample data for different house types
house_types = ['Apartment', 'House', 'Maison']
sample_data = {}
for h_type in house_types:
type_data = df[df['Type'] == h_type]['Price']
if len(type_data) > 100:
sample_data[h_type] = type_data.sample(min(500, len(type_data)),
random_state=42).values
# Plot 1: Traditional violin plot (for comparison)
positions = np.arange(1, len(sample_data) + 1)
vp1 = ax1.violinplot(list(sample_data.values()), positions=positions,
showmeans=True, showmedians=True)
# Customize
colors = ['#FF9999', '#66B2FF', '#99FF99']
for i, pc in enumerate(vp1['bodies']):
pc.set_facecolor(colors[i])
pc.set_alpha(0.7)
ax1.set_title('Traditional Violin Plot', fontsize=13, fontweight='bold')
ax1.set_xticks(positions)
ax1.set_xticklabels(list(sample_data.keys()))
ax1.set_ylabel('Price (Euros)')
ax1.yaxis.set_major_formatter(
plt.FuncFormatter(lambda x, p: f'€{x/1000:.0f}K'))
ax1.grid(True, alpha=0.2, axis='y')
# Plot 2: Half-violin plot (left side only)
from matplotlib.collections import PolyCollection
# Create half violins manually
for i, (h_type, data) in enumerate(sample_data.items()):
# Calculate KDE
from scipy import stats
kde = stats.gaussian_kde(data)
x = np.linspace(min(data), max(data), 100)
y = kde(x)
# Normalize y to desired width
y = y / y.max() * 0.4
# Create left half only
vertices = np.vstack([np.zeros_like(x) + i + 0.6, x]).T
vertices = np.vstack([vertices, [(i + 0.6, x[-1]), (i + 0.6, x[0])]])
poly = PolyCollection([vertices], facecolors=colors[i],
alpha=0.7, edgecolors='black')
ax2.add_collection(poly)
# Add median line
median = np.median(data)
ax2.plot([i + 0.6, i + 0.6 + 0.4], [median, median],
'k-', linewidth=2)
# Add boxplot on top
bp = ax2.boxplot(list(sample_data.values()), positions=positions,
widths=0.15, patch_artist=True)
for i, box in enumerate(bp['boxes']):
box.set_facecolor('white')
box.set_alpha(0.8)
ax2.set_title('Half-Violin + Boxplot (Raincloud Style)',
fontsize=13, fontweight='bold')
ax2.set_xticks(positions)
ax2.set_xticklabels(list(sample_data.keys()))
ax2.set_ylabel('Price (Euros)')
ax2.yaxis.set_major_formatter(
plt.FuncFormatter(lambda x, p: f'€{x/1000:.0f}K'))
ax2.grid(True, alpha=0.2, axis='y')
# Plot 3: Full raincloud plot (half-violin + jitter)
ax3.set_title('Raincloud Plot (Half-Violin + Jittered Points)',
fontsize=13, fontweight='bold')
# Create half violins
for i, (h_type, data) in enumerate(sample_data.items()):
# Half violin (same as above)
kde = stats.gaussian_kde(data)
x = np.linspace(min(data), max(data), 100)
y = kde(x)
y = y / y.max() * 0.4
vertices = np.vstack([np.zeros_like(x) + i + 0.6, x]).T
vertices = np.vstack([vertices, [(i + 0.6, x[-1]), (i + 0.6, x[0])]])
poly = PolyCollection([vertices], facecolors=colors[i],
alpha=0.5, edgecolors='none')
ax3.add_collection(poly)
# Add jittered individual points
jitter = np.random.normal(i + 1, 0.05, len(data))
ax3.scatter(jitter, data, alpha=0.3, s=20,
color=colors[i], edgecolors='none')
# Add median line
median = np.median(data)
ax3.plot([i + 0.8, i + 1.2], [median, median],
'k-', linewidth=2, zorder=5)
ax3.set_xticks(positions)
ax3.set_xticklabels(list(sample_data.keys()))
ax3.set_ylabel('Price (Euros)')
ax3.yaxis.set_major_formatter(
plt.FuncFormatter(lambda x, p: f'€{x/1000:.0f}K'))
ax3.grid(True, alpha=0.2, axis='y')
plt.tight_layout()
plt.show()
print("☁️ RAINCLOUD PLOTS EXPLAINED:")
print("="*45)
print("\nRaincloud Plot Components:")
print("1. ☁️ Half-violin plot (left): Density distribution")
print("2. 📦 Boxplot (center): Summary statistics")
print("3. 🌧️ Jittered points (right): Individual data points")
print("4. 📏 Median line: Central tendency")
print("\n🎯 Advantages Over Traditional Violin Plots:")
print("• Shows individual data points (not just density)")
print("• More intuitive (points + density + statistics)")
print("• Less ink used (half-violin is sufficient)")
print("• Better for small to medium datasets")
print("\n💡 When to Use Raincloud Plots:")
print("• Small to medium sample sizes (< 1000 points)")
print("• When you want to show individual observations")
print("• For publication-quality visualizations")
print("• When comparing distributions across groups")
Method 5: Violin Plot with Swarm/Strip Overlay
Combine violin plots with individual data points:
# METHOD 5: Violin plot with swarm/strip overlay
fig, axes = plt.subplots(1, 3, figsize=(16, 8))
# Sample data for demonstration
np.random.seed(42)
sample_size = 200
sample_df = df.sample(sample_size, random_state=42).copy()
# Create price segments for comparison
def get_price_segment(price):
if price <= 400000:
return 'Budget (≤ €400K)'
elif price <= 800000:
return 'Mid-Range (€400-800K)'
else:
return 'Luxury (> €800K)'
sample_df['Segment'] = sample_df['Price'].apply(get_price_segment)
segments = sample_df['Segment'].unique()
# Plot 1: Violin plot only
vp1 = axes[0].violinplot([sample_df[sample_df['Segment'] == seg]['Price'].values
for seg in segments],
positions=range(len(segments)),
showmeans=True, showmedians=True)
# Color violins
colors = ['#FF9999', '#66B2FF', '#99FF99']
for i, pc in enumerate(vp1['bodies']):
pc.set_facecolor(colors[i])
pc.set_alpha(0.6)
axes[0].set_title('Violin Plot Only', fontsize=12, fontweight='bold')
axes[0].set_xticks(range(len(segments)))
axes[0].set_xticklabels(segments, rotation=15)
axes[0].set_ylabel('Price (Euros)')
axes[0].yaxis.set_major_formatter(
plt.FuncFormatter(lambda x, p: f'€{x/1000:.0f}K'))
axes[0].grid(True, alpha=0.2, axis='y')
# Plot 2: Violin + Strip plot
# First create violin
vp2 = axes[1].violinplot([sample_df[sample_df['Segment'] == seg]['Price'].values
for seg in segments],
positions=range(len(segments)),
showmeans=False, showmedians=False)
# Color violins with more transparency
for i, pc in enumerate(vp2['bodies']):
pc.set_facecolor(colors[i])
pc.set_alpha(0.3)
# Add strip plot (jittered points)
for i, seg in enumerate(segments):
seg_data = sample_df[sample_df['Segment'] == seg]['Price']
# Add jitter to x-position
jitter = np.random.normal(i, 0.1, len(seg_data))
axes[1].scatter(jitter, seg_data, alpha=0.5, s=30,
color=colors[i], edgecolors='white', linewidth=0.5)
axes[1].set_title('Violin + Strip Plot', fontsize=12, fontweight='bold')
axes[1].set_xticks(range(len(segments)))
axes[1].set_xticklabels(segments, rotation=15)
axes[1].set_ylabel('Price (Euros)')
axes[1].yaxis.set_major_formatter(
plt.FuncFormatter(lambda x, p: f'€{x/1000:.0f}K'))
axes[1].grid(True, alpha=0.2, axis='y')
# Plot 3: Violin + Swarm plot (using seaborn)
import seaborn as sns
# Create violin with seaborn for easier swarm overlay
sns.violinplot(x='Segment', y='Price', data=sample_df,
ax=axes[2], inner=None, palette=colors, alpha=0.3)
# Overlay swarm plot
sns.swarmplot(x='Segment', y='Price', data=sample_df,
ax=axes[2], color='black', alpha=0.6, size=3)
axes[2].set_title('Violin + Swarm Plot (Seaborn)', fontsize=12, fontweight='bold')
axes[2].set_xlabel('')
axes[2].set_ylabel('Price (Euros)')
axes[2].yaxis.set_major_formatter(
plt.FuncFormatter(lambda x, p: f'€{x/1000:.0f}K'))
axes[2].grid(True, alpha=0.2, axis='y')
plt.tight_layout()
plt.show()
print("🐝 SWARM/STRIP PLOT OVERLAYS:")
print("="*45)
print("\nWhy Overlay Individual Points?")
print("1. Shows exact data points (not just density)")
print("2. Reveals clustering and gaps")
print("3. Helps identify outliers visually")
print("4. Shows sample size directly")
print("5. More transparent visualization")
print("\n🎯 Different Overlay Types:")
print("Strip Plot: Points with jitter (simple)")
print("Swarm Plot: Points avoid overlap (better)")
print("Beeswarm Plot: Optimized arrangement (best)")
print("\n💡 Best Practices for Point Overlays:")
print("• Use with small to medium datasets (< 500 points)")
print("• Adjust point size and transparency")
print("• Consider color coding by additional variables")
print("• Use swarm over strip when points overlap")
print("• Combine with semi-transparent violins")
# Analyze clustering patterns
print("\n🔍 Clustering Pattern Analysis:")
print("-"*35)
for seg in segments:
seg_data = sample_df[sample_df['Segment'] == seg]['Price']
# Check for gaps in distribution
sorted_prices = np.sort(seg_data)
gaps = np.diff(sorted_prices)
large_gaps = gaps[gaps > np.percentile(gaps, 90)]
print(f"\n{seg}:")
print(f" • Sample size: {len(seg_data)}")
print(f" • Price range: €{seg_data.min()/1000:.0f}K - €{seg_data.max()/1000:.0f}K")
if len(large_gaps) > 0:
print(f" • Has {len(large_gaps)} significant gaps in distribution")
print(f" • Largest gap: €{large_gaps.max()/1000:.1f}K")
else:
print(f" • Relatively continuous distribution")
Method 6: Advanced Customization & Professional Styling
Create publication-ready violin plots with advanced features:
# METHOD 6: Professional publication-ready violin plot
plt.figure(figsize=(16, 10))
# Prepare data: Price by neighborhood (top 5)
top_neighborhoods = df['Neighborhood'].value_counts().head(5).index.tolist()
df_top = df[df['Neighborhood'].isin(top_neighborhoods)].copy()
# Sort neighborhoods by median price for better visualization
neighborhood_medians = {}
for nb in top_neighborhoods:
neighborhood_medians[nb] = df_top[df_top['Neighborhood'] == nb]['Price'].median()
# Sort neighborhoods by median price
sorted_neighborhoods = sorted(top_neighborhoods,
key=lambda x: neighborhood_medians[x])
# Prepare data in sorted order
violin_data = []
for nb in sorted_neighborhoods:
violin_data.append(df_top[df_top['Neighborhood'] == nb]['Price'].values)
# Create positions
positions = np.arange(1, len(sorted_neighborhoods) + 1)
# Create the violin plot with extensive customization
vp = plt.violinplot(violin_data, positions=positions,
showmeans=True, showmedians=True,
showextrema=True,
widths=0.8, # Width of violins
bw_method='scott') # Bandwidth method
# Customize violin bodies with gradient colors
cmap = plt.cm.viridis
norm = plt.Normalize(vmin=0, vmax=len(violin_data)-1)
for i, pc in enumerate(vp['bodies']):
color = cmap(norm(i))
pc.set_facecolor(color)
pc.set_alpha(0.7)
pc.set_edgecolor('black')
pc.set_linewidth(1)
# Add gradient fill within each violin
from matplotlib.colors import LinearSegmentedColormap
gradient = np.linspace(0, 1, 256).reshape(1, -1)
gradient = np.vstack((gradient, gradient))
# Create gradient mask
from matplotlib.path import Path
from matplotlib.patches import PathPatch
# Get violin path vertices
vertices = pc.get_paths()[0].vertices
path = Path(vertices)
patch = PathPatch(path, facecolor='none', edgecolor='none')
plt.gca().add_patch(patch)
# Add subtle gradient (commented for clarity)
# im = plt.imshow(gradient, aspect='auto', cmap='Blues',
# extent=[i+0.6, i+1.4, min(violin_data[i]), max(violin_data[i])],
# alpha=0.2, clip_path=patch, clip_on=True)
# Customize statistical elements
vp['cmedians'].set_color('gold')
vp['cmedians'].set_linewidth(2.5)
vp['cmeans'].set_color('red')
vp['cmeans'].set_linewidth(2)
vp['cmeans'].set_linestyle='--'
vp['cmins'].set_color('gray')
vp['cmaxes'].set_color('gray')
vp['cbars'].set_color('gray')
vp['cbars'].set_linewidth(1)
# Add labels and title with professional formatting
plt.ylabel('House Price (Euros)', fontsize=14, fontweight='bold', labelpad=15)
plt.xlabel('Neighborhood', fontsize=14, fontweight='bold', labelpad=15)
plt.title('Amsterdam House Price Distribution by Neighborhood\nProfessional Violin Plot Analysis',
fontsize=18, fontweight='bold', pad=25)
# Set x-ticks with rotated labels
plt.xticks(positions, sorted_neighborhoods, rotation=30,
ha='right', fontsize=11, fontweight='bold')
# Format y-axis professionally
plt.gca().yaxis.set_major_formatter(
plt.FuncFormatter(lambda x, p: f'€{x/1000:.0f}K'))
plt.gca().yaxis.set_tick_params(labelsize=11)
# Add subtle grid
plt.grid(True, alpha=0.15, axis='y', linestyle='--', linewidth=0.5)
# Add mean and median annotations
for i, (pos, data) in enumerate(zip(positions, violin_data)):
mean_val = np.mean(data)
median_val = np.median(data)
# Annotate mean
plt.annotate(f'Mean: €{mean_val/1000:.0f}K',
xy=(pos, mean_val), xytext=(pos, mean_val + 150000),
ha='center', va='bottom', fontsize=9,
arrowprops=dict(arrowstyle='->', color='red', alpha=0.7),
bbox=dict(boxstyle='round', facecolor='white', alpha=0.8))
# Annotate median
plt.annotate(f'Median: €{median_val/1000:.0f}K',
xy=(pos, median_val), xytext=(pos, median_val - 150000),
ha='center', va='top', fontsize=9,
arrowprops=dict(arrowstyle='->', color='gold', alpha=0.7),
bbox=dict(boxstyle='round', facecolor='white', alpha=0.8))
# Add distribution shape indicators
shape_indicators = []
for i, data in enumerate(violin_data):
from scipy import stats
skewness = stats.skew(data)
if abs(skewness) > 0.5:
shape = 'Right-skewed' if skewness > 0 else 'Left-skewed'
color = '#FF6B6B' if skewness > 0 else '#4ECDC4'
shape_indicators.append(f'{sorted_neighborhoods[i]}: {shape}')
# Add shape info box
if shape_indicators:
shape_text = "Distribution Shapes:\n" + "\n".join(shape_indicators)
plt.text(0.98, 0.02, shape_text, transform=plt.gca().transAxes,
fontsize=9, verticalalignment='bottom',
horizontalalignment='right',
bbox=dict(boxstyle='round', facecolor='lightyellow',
alpha=0.9, edgecolor='orange'))
# Add colorbar for gradient (if used)
# sm = plt.cm.ScalarMappable(cmap=cmap, norm=norm)
# sm.set_array([])
# cbar = plt.colorbar(sm, ax=plt.gca(), pad=0.02)
# cbar.set_label('Neighborhood Index', fontsize=10)
# Add comprehensive legend
from matplotlib.patches import Patch
legend_elements = [
Patch(facecolor='gray', alpha=0.7, label='Density Distribution'),
Patch(facecolor='none', edgecolor='gold', linewidth=2, label='Median'),
Patch(facecolor='none', edgecolor='red', linestyle='--',
linewidth=2, label='Mean'),
Patch(facecolor=cmap(0.2), alpha=0.7, label='Lower Priced Areas'),
Patch(facecolor=cmap(0.8), alpha=0.7, label='Higher Priced Areas')
]
plt.legend(handles=legend_elements, loc='upper left',
fontsize=10, framealpha=0.9)
# Add overall statistics
total_houses = sum(len(data) for data in violin_data)
overall_median = df_top['Price'].median()
overall_mean = df_top['Price'].mean()
stats_text = f"""Overall Statistics (Top {len(sorted_neighborhoods)} Areas):
• Total Houses: {total_houses:,}
• Overall Median: €{overall_median/1000:.0f}K
• Overall Mean: €{overall_mean/1000:.0f}K
• Price Range: €{df_top['Price'].min()/1000:.0f}K - €{df_top['Price'].max()/1000:.0f}K
• Coefficient of Variation: {(df_top['Price'].std() / overall_mean)*100:.1f}%"""
plt.text(0.02, 0.98, stats_text, transform=plt.gca().transAxes,
fontsize=10, verticalalignment='top',
bbox=dict(boxstyle='round', facecolor='white',
alpha=0.9, edgecolor='blue'))
plt.tight_layout()
plt.show()
print("🎨 PROFESSIONAL VIOLIN PLOT FEATURES:")
print("="*45)
print("1. Gradient color scheme (viridis colormap)")
print("2. Multiple statistical markers (mean, median, min/max)")
print("3. Annotated values with arrows")
print("4. Distribution shape analysis")
print("5. Comprehensive legend")
print("6. Overall statistics box")
print("7. Professional typography and spacing")
print("8. Subtle grid for readability")
print("9. Sorted display for better comparison")
print("10. Publication-ready formatting")
print("\n📊 NEIGHBORHOOD DISTRIBUTION INSIGHTS:")
print("-"*40)
for i, nb in enumerate(sorted_neighborhoods):
data = df_top[df_top['Neighborhood'] == nb]['Price']
skew = data.skew()
kurt = data.kurtosis()
print(f"\n{nb}:")
print(f" • Median: €{data.median()/1000:.0f}K")
print(f" • IQR: €{data.quantile(0.25)/1000:.0f}K - €{data.quantile(0.75)/1000:.0f}K")
print(f" • Skewness: {skew:.2f} ({'Right-skewed' if skew > 0 else 'Left-skewed'})")
print(f" • Kurtosis: {kurt:.2f} ({'Heavy tails' if kurt > 3 else 'Light tails'})")
print(f" • Sample size: {len(data):,}")
Violin Plot Cheat Sheet (Quick Reference)
# 1. BASIC VIOLIN PLOT
plt.violinplot(data) # Basic violin
plt.violinplot(data, showmeans=True) # Show mean
plt.violinplot(data, showmedians=True) # Show median
plt.violinplot(data, showextrema=True) # Show min/max
# 2. MULTIPLE VIOLIN PLOTS
plt.violinplot([data1, data2, data3]) # Multiple groups
plt.violinplot([data1, data2], positions=[1, 3]) # Custom positions
# 3. CUSTOMIZATION
# Color and style
for pc in vp['bodies']:
pc.set_facecolor('lightblue')
pc.set_alpha(0.7)
pc.set_edgecolor('black')
pc.set_linewidth(1)
# Statistical lines
vp['cmedians'].set_color('red')
vp['cmedians'].set_linewidth(2)
vp['cmeans'].set_color('green')
vp['cmeans'].set_linewidth(1.5)
# 4. BANDWIDTH CONTROL
plt.violinplot(data, bw_method='scott') # Default
plt.violinplot(data, bw_method='silverman') # Alternative
plt.violinplot(data, bw_method=0.5) # Fixed bandwidth
# 5. QUANTILES (CUSTOM PERCENTILES)
plt.violinplot(data, quantiles=[[0.25, 0.75]]) # Show quartiles
plt.violinplot(data, quantiles=[[0.1, 0.9]]) # Show 10th & 90th
# 6. SEABORN VIOLIN PLOTS (EASIER)
import seaborn as sns
sns.violinplot(x='category', y='value', data=df) # Basic
sns.violinplot(x='category', y='value', data=df, split=True) # Split
sns.violinplot(x='category', y='value', data=df, inner='box') # Box inside
sns.violinplot(x='category', y='value', data=df, hue='group') # Grouped
# 7. HALF VIOLIN PLOTS (MANUAL)
from matplotlib.collections import PolyCollection
vertices = [...] # Calculate vertices for half shape
poly = PolyCollection([vertices], facecolors='blue', alpha=0.7)
ax.add_collection(poly)
# 8. RAINCLOUD PLOTS (COMBINATION)
# Half violin + boxplot + jittered points
# Use libraries like ptitprince or create manually
# 9. COMMON PARAMETERS
widths=0.8 # Width of violins
positions=[...] # X positions
vert=True # Vertical (False for horizontal)
points=100 # Points to evaluate KDE
# 10. SAVING
plt.savefig('violin_plot.png', dpi=300, bbox_inches='tight',
facecolor='white', transparent=False)
Common Violin Plot Mistakes to Avoid ⚠️
❌ COMMON MISTAKES:
- Wrong bandwidth: Too small (noisy) or too large (oversmoothed)
- Overplotting: Too many violins in one plot
- Misinterpreting width: Width shows density, not frequency
- Ignoring multimodality: Not investigating multiple peaks
- Small sample sizes: Violin plots need sufficient data
- Color misuse: Using colors that don't represent density well
- Missing scale: Not showing what width represents
- Overcomplicating: Too many inner elements confusing the plot
✅ BEST PRACTICES:
- Check bandwidth: Test different bw_method values
- Limit groups: Max 6-8 violins for readability
- Use split violins: For comparing two groups per category
- Add individual points: For small to medium datasets
- Show sample sizes: Annotate with n for each group
- Use consistent scaling: Same y-axis for comparison
- Consider alternatives: Boxplots for large datasets
- Document bandwidth: Note bandwidth method in caption
When to Use Violin Plots vs Alternatives
🎯 CHOOSING THE RIGHT PLOT:
🎻 Use Violin Plots When:
- Need to see distribution shape
- Looking for multimodality
- Comparing groups with similar spreads
- Small to medium datasets (n < 10,000)
- Publication-quality visuals needed
📦 Use Boxplots When:
- Focus on outliers and extremes
- Large datasets (n > 10,000)
- Need simple, fast visualization
- Comparing many groups (> 8)
- Non-technical audience
☁️ Use Raincloud Plots When:
- Small datasets (n < 500)
- Want to show individual points
- Need maximum transparency
- Statistical detail + raw data
- Research publications
📊 Use Histograms When:
- Need exact frequency counts
- Comparing absolute frequencies
- Data has natural binning
- Simple distribution overview
- Teaching basic statistics
Your Violin Plot Mastery Checklist
🎓 WHAT YOU'VE LEARNED:
✅ Can create basic violin plots with plt.violinplot()
✅ Understand KDE and bandwidth control
✅ Can create multiple violin plots for comparison
✅ Know how to make split violin plots
✅ Can create half-violin and raincloud plots
✅ Understand how to overlay individual points
✅ Can create publication-ready violin plots
✅ Know when to use violin vs box vs raincloud plots
Remember: Violin plots are powerful tools that reveal the full story of your data distribution. They combine the statistical rigor of boxplots with the visual richness of density plots. With the Amsterdam house price data, violin plots help you see not just where prices cluster, but how they distribute, where gaps exist, and how different neighborhoods compare in their entire price spectrum.
Happy violin plotting! 🎻🎯
Comments
Post a Comment