A Beginner's Guide 🐼
When you work with data in Pandas, you need to explore and understand it first. Think of it like getting a new puzzle — before solving it, you need to see how many pieces you have, what colors there are, and which pieces are missing! These 8 methods are like your toolbox for data exploration. Let's learn them step by step, super easy!
Our Example DataFrame
Let's create a simple dataset about people to practice with:
import pandas as pd
data = {
'Name': ['John', 'Anna', 'Peter', 'Linda', 'James', 'Sarah', 'Mike', 'Emma'],
'Age': [25, 30, 35, 28, 32, 27, 33, 29],
'City': ['New York', 'London', 'Paris', 'Tokyo', 'Berlin', 'Madrid', 'Toronto', 'Dubai'],
'Salary': [50000, 60000, 75000, 55000, 70000, 52000, 68000, 62000]
}
df = pd.DataFrame(data)
Now we have a table with 8 people and 4 columns. Let's explore it! 🔍
1. df.head() - See the Top Rows
Imagine you have a huge book with 1000 pages. Instead of reading everything at once, you just peek at the first few pages to get an idea. That's exactly what df.head() does!
# Show first 5 rows (default)
print(df.head())
Output:
Name Age City Salary
0 John 25 New York 50000
1 Anna 30 London 60000
2 Peter 35 Paris 75000
3 Linda 28 Tokyo 55000
4 James 32 Berlin 70000
Want to see more or less?
print(df.head(3)) # First 3 rows only
print(df.head(10)) # First 10 rows (will show all 8 since we only have 8)
When to use this:
- When you first load a CSV file and want to quickly check if it loaded correctly
- To verify column names are what you expected
- To see actual data examples without waiting for the whole dataset to print
- Perfect for huge datasets with millions of rows — you don't want to print everything!
2. df.tail() - See the Bottom Rows
This is like reading the last pages of your book. Sometimes the end tells you important things! Maybe your data was cut off? Maybe there's a pattern at the end?
# Show last 3 rows
print(df.tail(3))
Output:
Name Age City Salary
5 Sarah 27 Madrid 52000
6 Mike 33 Toronto 68000
7 Emma 29 Dubai 62000
Why is this useful?
- Sometimes data gets cut off or corrupted at the end when loading from files
- In time-series data (like stock prices or temperature readings), the most recent data is at the bottom
- Check if your data sorting worked correctly
- Verify the last entries were added properly
print(df.tail()) # Default: last 5 rows
print(df.tail(1)) # Just the very last row
Real-world example: If you're tracking daily sales and load yesterday's data, use tail() to verify yesterday's date is actually there!
3. df.sample() - Pick Random Rows
Imagine picking names from a hat — totally random! This gives you a surprise view of your data without any bias.
# Get 2 random rows
print(df.sample(2))
Output (will be different each time!):
Name Age City Salary
3 Linda 28 Tokyo 55000
6 Mike 33 Toronto 68000
Every time you run it, you get different rows! This is super powerful for:
- Getting an unbiased peek at your data (not always looking at the same top rows)
- Testing your code on random examples to make sure it works everywhere
- Spotting unusual patterns you might miss if you only look at the beginning
- Creating training/testing datasets in machine learning
print(df.sample(5)) # 5 random rows
print(df.sample(frac=0.5)) # Random 50% of all rows (4 rows here)
print(df.sample(1)) # Just 1 random row
Pro tip: If you want the same random rows every time (for testing), use a seed:
df.sample(3, random_state=42) # Same 3 random rows every time!
4. df.shape - Know Your Size
How many rows and columns do you have? df.shape tells you instantly — like checking how big your puzzle is before starting!
print(df.shape)
Output:
(8, 4)
This means: 8 rows (people) and 4 columns (Name, Age, City, Salary)
The format is always: (number_of_rows, number_of_columns)
You can use this in your code!
num_rows = df.shape[0] # Get number of rows: 8
num_columns = df.shape[1] # Get number of columns: 4
print(f"We have {num_rows} people in our dataset")
# Output: We have 8 people in our dataset
print(f"Each person has {num_columns} attributes")
# Output: Each person has 4 attributes
When to use:
- Before processing: if you expect 1000 rows but see 10, something went wrong!
- After filtering: check how many rows match your criteria
- After merging datasets: verify the merge worked as expected
- In loops: know how many iterations you'll need
5. df.columns - List All Column Names
What information do you have? This shows all the column labels — like reading the headers in a spreadsheet.
print(df.columns)
Output:
Index(['Name', 'Age', 'City', 'Salary'], dtype='object')
You can turn it into a regular Python list:
column_list = df.columns.tolist()
print(column_list)
# ['Name', 'Age', 'City', 'Salary']
Why is this super useful?
- Check if a column exists before using it (avoid errors!)
- Loop through all columns automatically
- Find typos in column names (like "Agee" instead of "Age")
- Get column count:
len(df.columns)gives you 4
Practical examples:
# Check if a column exists
if 'Age' in df.columns:
print("Age column found!")
else:
print("Age column missing!")
# Loop through all columns
for col in df.columns:
print(f"Column: {col}")
# Output:
# Column: Name
# Column: Age
# Column: City
# Column: Salary
6. df.dtypes - Check Data Types
Are your numbers really numbers? Or did Pandas think they're text? df.dtypes tells you how Pandas is interpreting each column. This is SUPER important!
print(df.dtypes)
Output:
Name object
Age int64
City object
Salary int64
dtype: object
What do these mean?
object= Text (strings) like "John" or "Tokyo" or mixed typesint64= Whole numbers like 25, 30, 35 (no decimals)float64= Decimal numbers like 3.14, 99.99, 25.0datetime64= Dates and times like "2024-01-15"bool= True/False values
Why this matters — REALLY important:
If numbers are stored as "object" (text), you cannot do math on them! Imagine trying to add "25" + "30" as text — you get "2530" not 55!
# Example: What if Age was stored as text?
# You'd need to convert it:
df['Age'] = df['Age'].astype(int)
# Check specific column type
print(df['Age'].dtype) # int64
# Check if a column is numeric
print(df['Salary'].dtype in ['int64', 'float64']) # True
Common problems this helps catch:
- Numbers with commas like "1,000" are read as text (object)
- Missing values in number columns turn them into float64
- Dates stored as text instead of datetime64
7. df.info() - Get the Full Report Card
This is like a complete health checkup for your data. It tells you EVERYTHING important in one command! This is the FIRST thing you should run on new data.
df.info()
Output:
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 8 entries, 0 to 7
Data columns (total 4 columns):
# Column Non-Null Count Dtype
--- ------ -------------- -----
0 Name 8 non-null object
1 Age 8 non-null int64
2 City 8 non-null object
3 Salary 8 non-null int64
dtypes: int64(2), object(2)
memory usage: 384.0+ bytes
Breaking it down — what you learn:
- RangeIndex: 8 entries, 0 to 7 = You have 8 rows, numbered from 0 to 7
- Data columns (total 4 columns) = You have 4 columns
- Non-Null Count = How many values are NOT missing
- All show "8 non-null" which means perfect — no missing data!
- If you saw "5 non-null" out of 8 rows, that means 3 values are missing (NaN)
- Dtype = Data type of each column (we learned this above)
- dtypes: int64(2), object(2) = Summary: 2 integer columns, 2 text columns
- memory usage: 384.0+ bytes = How much RAM this takes (tiny!)
This is PERFECT for finding missing data:
# Example with missing values:
# If you see:
# 0 Name 8 non-null object
# 1 Age 5 non-null int64 <-- Only 5 out of 8!
# 2 City 8 non-null object
# 3 Salary 7 non-null int64 <-- Only 7 out of 8!
# This immediately tells you:
# - Age column is missing 3 values (8-5=3)
# - Salary column is missing 1 value (8-7=1)
When to use: ALWAYS run this FIRST when you get new data. It's your data's complete resume!
8. df.describe() - Statistics at a Glance
Want to know the average age? Minimum salary? Maximum? df.describe() gives you instant math summary for all number columns! It's like a calculator on steroids.
print(df.describe())
Output:
Age Salary
count 8.000000 8.000000
mean 29.875000 61500.000000
std 3.227486 8602.325267
min 25.000000 50000.000000
25% 27.750000 53500.000000
50% 29.500000 61000.000000
75% 32.250000 68500.000000
max 35.000000 75000.000000
Breaking it down — what each number means:
- count = 8.0 → We counted 8 people (all rows have values)
- mean = 29.875 → Average age is about 30 years, average salary is $61,500
- std = 3.227 → Standard deviation (how spread out the numbers are)
- Low std = everyone is similar (ages are close to 30)
- High std = big differences (salaries range from 50k to 75k)
- min = 25 → Youngest person is 25, lowest salary is $50,000
- 25% = 27.75 → 25% of people are younger than 27.75 years (quartile)
- 50% = 29.5 → Middle value (median) — half are younger, half are older
- 75% = 32.25 → 75% of people are younger than 32.25 years
- max = 35 → Oldest person is 35, highest salary is $75,000
What about text columns? By default they're not shown, but you can include them:
print(df.describe(include='all'))
Output (showing everything):
Name Age City Salary
count 8 8.0 8 8.0
unique 8 NaN 8 NaN
top John NaN New York NaN
freq 1 NaN 1 NaN
mean NaN 29.9 NaN 61500.0
std NaN 3.2 NaN 8602.3
min NaN 25.0 NaN 50000.0
25% NaN 27.8 NaN 53500.0
50% NaN 29.5 NaN 61000.0
75% NaN 32.3 NaN 68500.0
max NaN 35.0 NaN 75000.0
Now for text columns it shows:
- unique = 8 unique names (all different people)
- top = Most common value (here "John" and "New York" each appear once)
- freq = Frequency of the top value (1 means it appears once)
This helps you spot problems quickly:
- If min age is -5, you have bad data!
- If max salary is 9,999,999, that's probably a typo
- If mean is way different from median (50%), you have outliers
Quick Cheat Sheet for Beginners 🚀
df.head() # First 5 rows - quick peek at the top
df.tail(3) # Last 3 rows - check the end
df.sample(2) # 2 random rows - unbiased view
df.shape # (rows, columns) - how big is my data?
df.columns # Column names - what info do I have?
df.dtypes # Data types - are numbers really numbers?
df.info() # Full report - missing values? memory usage?
df.describe() # Statistics - min, max, average, etc.
Pro Tips for Beginners
- Always start with:
df.shape,df.info(), anddf.head() - Use
df.sample()instead of alwayshead()to avoid bias - Check
df.dtypesif calculations aren't working → probably wrong type! df.describe()quickly spots outliers (values too high or too low)- Missing values in
df.info()shows where you need to clean data - Combine methods: After filtering, check shape to see how many rows remain
Common Mistakes to Avoid ⚠️
- Don't print entire huge datasets — use
head()instead! - Don't ignore data types — "25" as text ≠ 25 as number
- Don't skip
df.info()— it saves hours of debugging - Don't only look at
head()— important patterns might be at the end
Keep practicing with Pandas — it gets easier every time! 🐼✨
Comments
Post a Comment