Skip to main content

Posts

Showing posts with the label Pandas

Manage Raw JSON Lines - Pandas

Working with nested JSON data can be tricky! JSON often contains lists inside objects, objects inside lists, and multiple levels of nesting. This comprehensive tutorial shows you how to handle complex JSON structures in Pandas using lines=True , explode() , json_normalize() , and max_level parameter. What is JSON Lines Format? Before we dive in, let's understand JSON Lines (JSONL) format. Unlike regular JSON where everything is in one array, JSON Lines has one JSON object per line . This format is extremely popular for: Streaming data (logs, events, API responses) Large datasets (easier to process line-by-line) Database exports Machine learning training data Regular JSON: [ {"name": "Alice", "age": 25}, {"name": "Bob", "age": 30} ] JSON Lines Format: {"name": "Alice", "age": 25} {"name": "Bob", "age": 30} Sample Dataset - E-commer...

Pandas - Fix Wrong Data Formats

Data rarely comes in perfect format. Dates might be stored as text, numbers might have currency symbols, or ages might be stored as strings. This tutorial shows you how to fix wrong data formats in Pandas and convert everything to the correct type for analysis. We'll cover dates, numbers, currency, percentages, and mixed formats with plenty of practical examples! Sample Dataset for Practice Let's create a messy dataset with various format issues. This is what real-world data often looks like! import pandas as pd import numpy as np # Create dataset with format issues data = { 'Employee_ID': ['E001', 'E002', 'E003', 'E004', 'E005', 'E006', 'E007', 'E008', 'E009', 'E010'], 'Name': ['Alice Smith', 'Bob Jones', 'Charlie Brown', 'Diana Prince', 'Eve Wilson', 'Frank Miller', 'Grace Lee', 'Henry Ford', 'Ivy...

Pandas - Outliers Detection

Outliers are extreme values that are significantly different from other data points. They can skew your analysis, make averages misleading, and lead to incorrect conclusions. This tutorial shows you how to detect and handle outliers using Pandas with simple, practical examples. Sample Dataset for Practice Let's create a sample employee salary dataset with some outliers. You can use this exact dataset to practice all the outlier detection techniques! import pandas as pd import numpy as np # Create sample dataset with outliers data = { 'Employee_ID': [101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115], 'Name': ['Alice', 'Bob', 'Charlie', 'Diana', 'Eve', 'Frank', 'Grace', 'Henry', 'Ivy', 'Jack', 'Kate', 'Leo', 'Mia', 'Noah', 'Olivia'], 'Age': [28, 32, 45, 29, 35, 150, 27, 31, 40, 33, 38, 26, 30, 5, 4...

Pandas - Data validation

This Titorial shows you how to create validation rules to catch errors and ensure data quality using Pandas. We'll cover everything from basic null checks to complex business rules with practical, beginner-friendly examples. Sample Dataset for Practice Let's create a sample customer dataset that contains various data quality issues. You can use this exact dataset to practice all the validation techniques in this guide! import pandas as pd import numpy as np # Create sample dataset with intentional errors data = { 'Customer_ID': [101, 102, 103, 104, None, 106, 107, 108, 109, 110], 'Name': ['Alice Smith', 'bob jones', 'CHARLIE BROWN', 'Diana Prince', 'Eve Wilson', 'frank miller', 'Grace Lee', 'Henry Ford', 'Ivy Chen', 'Jack Ryan'], 'Age': [25, 150, 30, -5, 45, 28, 200, 35, 22, 40], 'Email': ['alice@email.com', 'bobemailcom', ...

Pandas - Data Cleaning

Pandas makes data cleaning super easy with powerful built-in functions. This guide explains 10 essential cleaning patterns step-by-step with plenty of simple examples, so even complete beginners can follow along and start using them right away. We'll use practical examples throughout to make it easy to understand and apply immediately. What is Data Cleaning? Data cleaning is the process of fixing messy, inconsistent, or incorrect data before analysis. Real-world data is often messy with extra spaces, wrong formats, duplicates, missing values, and inconsistencies. These patterns will help you handle all of these issues efficiently. 1. Remove Whitespace Extra spaces at the beginning or end of text can cause problems when matching or sorting data. The strip() method removes them instantly. # Remove leading and trailing whitespace df['city'] = df['city'].str.strip() # Before: " New York " # After: "New York" Real-World Use Case: Y...

Spotting Problematic Data in Pandas

Real-world data is almost always messy . People enter information differently, systems export in weird formats, and small errors creep in over time. If you jump straight into analysis without checking for problems, your results can be completely wrong. Pandas gives you simple, powerful tools to spot these issues early. In this tutorial, we'll walk through the most common data problems, show real examples with a deliberately messy dataset, and give you a step-by-step checklist you can use on every project. Our Messy Example Dataset Let's create a small but realistic customer dataset full of typical problems (duplicate customers, inconsistent capitalization, wrong data types, missing values, weird date formats, numbers stored as strings, etc.). import pandas as pd messy_data = { 'Customer_ID': ['C001', 'C002', 'C003', 'C004', 'C005', 'C001', 'C006', 'C007', 'C008', None], 'Name...