Skip to main content

Spotting Problematic Data in Pandas

Calculating read time…

Real-world data is almost always messy. People enter information differently, systems export in weird formats, and small errors creep in over time. If you jump straight into analysis without checking for problems, your results can be completely wrong.

Pandas gives you simple, powerful tools to spot these issues early. In this tutorial, we'll walk through the most common data problems, show real examples with a deliberately messy dataset, and give you a step-by-step checklist you can use on every project.

Our Messy Example Dataset

Let's create a small but realistic customer dataset full of typical problems (duplicate customers, inconsistent capitalization, wrong data types, missing values, weird date formats, numbers stored as strings, etc.).

    
        
            
            Customer_ID
            Name
            Age
            Email
            Phone
            Join_Date
            Salary
            Rating
        
    
    
        
            0
            C001
            john doe
            28
            john@email.com
            (123) 456-7890
            2020-01-15
            75,000
            4.5
        
        
            1
            C002
            JANE SMITH
            '32'
            jane@email.com
            9876543210
            05/20/2019
            85000
            3.8
        
        
            2
            C003
            bob johnson
            25
            bob@email.com
            555-1234
            March 10, 2021
            65,000
            'good'
        
        
    


Common Data Problems & How to Spot Them

1. Missing Values

df.isna().sum() or df.info()

2. Duplicates

df.duplicated().sum() or df.duplicated(subset=['Customer_ID'])

3. Inconsistent Formatting

Names mixed case, dates/phone in different styles

4. Wrong/Mixed Data Types

df.dtypes and check per column with .apply(type)

5. Outliers/Invalid Values

df.describe()

A Simpler, More Powerful Data Health Check Function

  • Added overall summary (shape, duplicates, missing overview) — you see big issues immediately
  • Kept per-column details but made them cleaner and more focused
  • Raised unique value limit to 15 (better for small categories)
  • Used clearer formatting and emojis for quick scanning
  • Still very easy to read and copy-paste
 0:
        print(f" Missing values:\n{missing[missing > 0]}")
    else:
        print(" No missing values")
    
    print("\n" + "—" * 50)
    
    # Per column details
    for col in df.columns:
        print(f"\n {col} (type: {df[col].dtype})")
        print(f"   • Non-null count: {df[col].notna().sum()}/{len(df)}")
        
        # Show unique values if not too many
        if df[col].nunique() < 15:
            print(f"   • Unique values ({df[col].nunique()}): {df[col].unique()}")
        
        # Detect mixed types
        types = df[col].apply(lambda x: type(x).__name__).value_counts()
        if len(types) > 1:
            print(f"  Mixed types detected: {dict(types)}")
        
        # Sample values
        print(f"   • Sample (first 3): {df[col].head(3).tolist()}")

# Run it
quick_data_check(df)

Your 5-Minute Data Check Checklist

  1. df.head() + df.sample(5) — Visual feel
  2. df.info() — Types & missing
  3. df.describe(include='all') — Stats
  4. df.duplicated().sum() — Duplicates

Happy Learning !! 🐼

Comments