Skip to main content

Essential Pandas DataFrame Methods

Calculating read time…

A Beginner's Guide 🐼

When you work with data in Pandas, you need to explore and understand it first. Think of it like getting a new puzzle — before solving it, you need to see how many pieces you have, what colors there are, and which pieces are missing! These 8 methods are like your toolbox for data exploration. Let's learn them step by step, super easy!

Our Example DataFrame

Let's create a simple dataset about people to practice with:

import pandas as pd

data = {
    'Name': ['John', 'Anna', 'Peter', 'Linda', 'James', 'Sarah', 'Mike', 'Emma'],
    'Age': [25, 30, 35, 28, 32, 27, 33, 29],
    'City': ['New York', 'London', 'Paris', 'Tokyo', 'Berlin', 'Madrid', 'Toronto', 'Dubai'],
    'Salary': [50000, 60000, 75000, 55000, 70000, 52000, 68000, 62000]
}

df = pd.DataFrame(data)

Now we have a table with 8 people and 4 columns. Let's explore it! 🔍

1. df.head() - See the Top Rows

Imagine you have a huge book with 1000 pages. Instead of reading everything at once, you just peek at the first few pages to get an idea. That's exactly what df.head() does!

# Show first 5 rows (default)
print(df.head())

Output:

    Name  Age      City  Salary
0   John   25  New York   50000
1   Anna   30    London   60000
2  Peter   35     Paris   75000
3  Linda   28     Tokyo   55000
4  James   32    Berlin   70000

Want to see more or less?

print(df.head(3))  # First 3 rows only
print(df.head(10)) # First 10 rows (will show all 8 since we only have 8)

When to use this:

  • When you first load a CSV file and want to quickly check if it loaded correctly
  • To verify column names are what you expected
  • To see actual data examples without waiting for the whole dataset to print
  • Perfect for huge datasets with millions of rows — you don't want to print everything!

2. df.tail() - See the Bottom Rows

This is like reading the last pages of your book. Sometimes the end tells you important things! Maybe your data was cut off? Maybe there's a pattern at the end?

# Show last 3 rows
print(df.tail(3))

Output:

   Name  Age     City  Salary
5  Sarah   27   Madrid   52000
6   Mike   33  Toronto   68000
7   Emma   29    Dubai   62000

Why is this useful?

  • Sometimes data gets cut off or corrupted at the end when loading from files
  • In time-series data (like stock prices or temperature readings), the most recent data is at the bottom
  • Check if your data sorting worked correctly
  • Verify the last entries were added properly
print(df.tail())   # Default: last 5 rows
print(df.tail(1))  # Just the very last row

Real-world example: If you're tracking daily sales and load yesterday's data, use tail() to verify yesterday's date is actually there!

3. df.sample() - Pick Random Rows

Imagine picking names from a hat — totally random! This gives you a surprise view of your data without any bias.

# Get 2 random rows
print(df.sample(2))

Output (will be different each time!):

    Name  Age     City  Salary
3  Linda   28    Tokyo   55000
6   Mike   33  Toronto   68000

Every time you run it, you get different rows! This is super powerful for:

  • Getting an unbiased peek at your data (not always looking at the same top rows)
  • Testing your code on random examples to make sure it works everywhere
  • Spotting unusual patterns you might miss if you only look at the beginning
  • Creating training/testing datasets in machine learning
print(df.sample(5))        # 5 random rows
print(df.sample(frac=0.5)) # Random 50% of all rows (4 rows here)
print(df.sample(1))        # Just 1 random row

Pro tip: If you want the same random rows every time (for testing), use a seed:

df.sample(3, random_state=42)  # Same 3 random rows every time!

4. df.shape - Know Your Size

How many rows and columns do you have? df.shape tells you instantly — like checking how big your puzzle is before starting!

print(df.shape)

Output:

(8, 4)

This means: 8 rows (people) and 4 columns (Name, Age, City, Salary)

The format is always: (number_of_rows, number_of_columns)

You can use this in your code!

num_rows = df.shape[0]      # Get number of rows: 8
num_columns = df.shape[1]   # Get number of columns: 4

print(f"We have {num_rows} people in our dataset")
# Output: We have 8 people in our dataset

print(f"Each person has {num_columns} attributes")
# Output: Each person has 4 attributes

When to use:

  • Before processing: if you expect 1000 rows but see 10, something went wrong!
  • After filtering: check how many rows match your criteria
  • After merging datasets: verify the merge worked as expected
  • In loops: know how many iterations you'll need

5. df.columns - List All Column Names

What information do you have? This shows all the column labels — like reading the headers in a spreadsheet.

print(df.columns)

Output:

Index(['Name', 'Age', 'City', 'Salary'], dtype='object')

You can turn it into a regular Python list:

column_list = df.columns.tolist()
print(column_list)
# ['Name', 'Age', 'City', 'Salary']

Why is this super useful?

  • Check if a column exists before using it (avoid errors!)
  • Loop through all columns automatically
  • Find typos in column names (like "Agee" instead of "Age")
  • Get column count: len(df.columns) gives you 4

Practical examples:

# Check if a column exists
if 'Age' in df.columns:
    print("Age column found!")
else:
    print("Age column missing!")

# Loop through all columns
for col in df.columns:
    print(f"Column: {col}")

# Output:
# Column: Name
# Column: Age
# Column: City
# Column: Salary

6. df.dtypes - Check Data Types

Are your numbers really numbers? Or did Pandas think they're text? df.dtypes tells you how Pandas is interpreting each column. This is SUPER important!

print(df.dtypes)

Output:

Name       object
Age         int64
City       object
Salary      int64
dtype: object

What do these mean?

  • object = Text (strings) like "John" or "Tokyo" or mixed types
  • int64 = Whole numbers like 25, 30, 35 (no decimals)
  • float64 = Decimal numbers like 3.14, 99.99, 25.0
  • datetime64 = Dates and times like "2024-01-15"
  • bool = True/False values

Why this matters — REALLY important:

If numbers are stored as "object" (text), you cannot do math on them! Imagine trying to add "25" + "30" as text — you get "2530" not 55!

# Example: What if Age was stored as text?
# You'd need to convert it:
df['Age'] = df['Age'].astype(int)

# Check specific column type
print(df['Age'].dtype)  # int64

# Check if a column is numeric
print(df['Salary'].dtype in ['int64', 'float64'])  # True

Common problems this helps catch:

  • Numbers with commas like "1,000" are read as text (object)
  • Missing values in number columns turn them into float64
  • Dates stored as text instead of datetime64

7. df.info() - Get the Full Report Card

This is like a complete health checkup for your data. It tells you EVERYTHING important in one command! This is the FIRST thing you should run on new data.

df.info()

Output:

<class 'pandas.core.frame.DataFrame'>
RangeIndex: 8 entries, 0 to 7
Data columns (total 4 columns):
 #   Column  Non-Null Count  Dtype 
---  ------  --------------  ----- 
 0   Name    8 non-null      object
 1   Age     8 non-null      int64 
 2   City    8 non-null      object
 3   Salary  8 non-null      int64 
dtypes: int64(2), object(2)
memory usage: 384.0+ bytes

Breaking it down — what you learn:

  • RangeIndex: 8 entries, 0 to 7 = You have 8 rows, numbered from 0 to 7
  • Data columns (total 4 columns) = You have 4 columns
  • Non-Null Count = How many values are NOT missing
    • All show "8 non-null" which means perfect — no missing data!
    • If you saw "5 non-null" out of 8 rows, that means 3 values are missing (NaN)
  • Dtype = Data type of each column (we learned this above)
  • dtypes: int64(2), object(2) = Summary: 2 integer columns, 2 text columns
  • memory usage: 384.0+ bytes = How much RAM this takes (tiny!)

This is PERFECT for finding missing data:

# Example with missing values:
# If you see:
# 0   Name    8 non-null    object
# 1   Age     5 non-null    int64   <-- Only 5 out of 8!
# 2   City    8 non-null    object
# 3   Salary  7 non-null    int64   <-- Only 7 out of 8!

# This immediately tells you:
# - Age column is missing 3 values (8-5=3)
# - Salary column is missing 1 value (8-7=1)

When to use: ALWAYS run this FIRST when you get new data. It's your data's complete resume!

8. df.describe() - Statistics at a Glance

Want to know the average age? Minimum salary? Maximum? df.describe() gives you instant math summary for all number columns! It's like a calculator on steroids.

print(df.describe())

Output:

             Age        Salary
count   8.000000      8.000000
mean   29.875000  61500.000000
std     3.227486   8602.325267
min    25.000000  50000.000000
25%    27.750000  53500.000000
50%    29.500000  61000.000000
75%    32.250000  68500.000000
max    35.000000  75000.000000

Breaking it down — what each number means:

  • count = 8.0 → We counted 8 people (all rows have values)
  • mean = 29.875 → Average age is about 30 years, average salary is $61,500
  • std = 3.227 → Standard deviation (how spread out the numbers are)
    • Low std = everyone is similar (ages are close to 30)
    • High std = big differences (salaries range from 50k to 75k)
  • min = 25 → Youngest person is 25, lowest salary is $50,000
  • 25% = 27.75 → 25% of people are younger than 27.75 years (quartile)
  • 50% = 29.5 → Middle value (median) — half are younger, half are older
  • 75% = 32.25 → 75% of people are younger than 32.25 years
  • max = 35 → Oldest person is 35, highest salary is $75,000

What about text columns? By default they're not shown, but you can include them:

print(df.describe(include='all'))

Output (showing everything):

        Name   Age        City        Salary
count      8   8.0           8           8.0
unique     8   NaN           8           NaN
top     John   NaN    New York           NaN
freq       1   NaN           1           NaN
mean     NaN  29.9         NaN       61500.0
std      NaN   3.2         NaN        8602.3
min      NaN  25.0         NaN       50000.0
25%      NaN  27.8         NaN       53500.0
50%      NaN  29.5         NaN       61000.0
75%      NaN  32.3         NaN       68500.0
max      NaN  35.0         NaN       75000.0

Now for text columns it shows:

  • unique = 8 unique names (all different people)
  • top = Most common value (here "John" and "New York" each appear once)
  • freq = Frequency of the top value (1 means it appears once)

This helps you spot problems quickly:

  • If min age is -5, you have bad data!
  • If max salary is 9,999,999, that's probably a typo
  • If mean is way different from median (50%), you have outliers

Quick Cheat Sheet for Beginners 🚀

df.head()       # First 5 rows - quick peek at the top
df.tail(3)      # Last 3 rows - check the end
df.sample(2)    # 2 random rows - unbiased view
df.shape        # (rows, columns) - how big is my data?
df.columns      # Column names - what info do I have?
df.dtypes       # Data types - are numbers really numbers?
df.info()       # Full report - missing values? memory usage?
df.describe()   # Statistics - min, max, average, etc.

Pro Tips for Beginners

  • Always start with: df.shape, df.info(), and df.head()
  • Use df.sample() instead of always head() to avoid bias
  • Check df.dtypes if calculations aren't working → probably wrong type!
  • df.describe() quickly spots outliers (values too high or too low)
  • Missing values in df.info() shows where you need to clean data
  • Combine methods: After filtering, check shape to see how many rows remain

Common Mistakes to Avoid ⚠️

  • Don't print entire huge datasets — use head() instead!
  • Don't ignore data types — "25" as text ≠ 25 as number
  • Don't skip df.info() — it saves hours of debugging
  • Don't only look at head() — important patterns might be at the end

Keep practicing with Pandas — it gets easier every time! 🐼✨

Comments