Real-world data is almost always messy. People enter information differently, systems export in weird formats, and small errors creep in over time. If you jump straight into analysis without checking for problems, your results can be completely wrong.
Pandas gives you simple, powerful tools to spot these issues early. In this tutorial, we'll walk through the most common data problems, show real examples with a deliberately messy dataset, and give you a step-by-step checklist you can use on every project.
Our Messy Example Dataset
Let's create a small but realistic customer dataset full of typical problems (duplicate customers, inconsistent capitalization, wrong data types, missing values, weird date formats, numbers stored as strings, etc.).
Customer_ID
Name
Age
Email
Phone
Join_Date
Salary
Rating
0
C001
john doe
28
john@email.com
(123) 456-7890
2020-01-15
75,000
4.5
1
C002
JANE SMITH
'32'
jane@email.com
9876543210
05/20/2019
85000
3.8
2
C003
bob johnson
25
bob@email.com
555-1234
March 10, 2021
65,000
'good'
Common Data Problems & How to Spot Them
1. Missing Values
df.isna().sum() or df.info()
2. Duplicates
df.duplicated().sum() or df.duplicated(subset=['Customer_ID'])
3. Inconsistent Formatting
Names mixed case, dates/phone in different styles
4. Wrong/Mixed Data Types
df.dtypes and check per column with .apply(type)
5. Outliers/Invalid Values
df.describe()
A Simpler, More Powerful Data Health Check Function
- Added overall summary (shape, duplicates, missing overview) — you see big issues immediately
- Kept per-column details but made them cleaner and more focused
- Raised unique value limit to 15 (better for small categories)
- Used clearer formatting and emojis for quick scanning
- Still very easy to read and copy-paste
0:
print(f" Missing values:\n{missing[missing > 0]}")
else:
print(" No missing values")
print("\n" + "—" * 50)
# Per column details
for col in df.columns:
print(f"\n {col} (type: {df[col].dtype})")
print(f" • Non-null count: {df[col].notna().sum()}/{len(df)}")
# Show unique values if not too many
if df[col].nunique() < 15:
print(f" • Unique values ({df[col].nunique()}): {df[col].unique()}")
# Detect mixed types
types = df[col].apply(lambda x: type(x).__name__).value_counts()
if len(types) > 1:
print(f" Mixed types detected: {dict(types)}")
# Sample values
print(f" • Sample (first 3): {df[col].head(3).tolist()}")
# Run it
quick_data_check(df)
Your 5-Minute Data Check Checklist
df.head()+df.sample(5)— Visual feeldf.info()— Types & missingdf.describe(include='all')— Statsdf.duplicated().sum()— Duplicates
Happy Learning !! 🐼
Comments
Post a Comment