Skip to main content

Pandas Data Selection

Calculating read time…

A Beginner-Friendly Guide : loc[] vs iloc[]

If you're learning Pandas, selecting data from a DataFrame can feel confusing. Pandas gives us two main tools for this: loc[] (label-based) and iloc[] (position-based).

We’ll use the below sample DataFrame throughout all examples so everything is easy to follow.

Our Sample DataFrame

First, create a simple DataFrame:

import pandas as pd

data = {
    'Name': ['Alice', 'Bob', 'Charlie', 'David', 'Eve'],
    'Age': [25, 30, 35, 28, 42],
    'Salary': [50000, 60000, 70000, 55000, 80000]
}

df = pd.DataFrame(data)
print(df)

This will display:

Name Age Salary
0 Alice 25 50000
1 Bob 30 60000
2 Charlie 35 70000
3 David 28 55000
4 Eve 42 80000

The numbers 0, 1, 2, 3, 4 on the left are the row index labels (by default, Pandas uses 0-based integers).

loc[] – Label-Based Selection

loc[] uses the actual labels of rows and columns (not their positions). It’s perfect when you want to select data by meaningful names or index values.

Key points for beginners:

  • Slicing with loc includes both start and end labels.
  • You can use column names directly (strings).
  • It works great with conditions (boolean indexing).
# 1. Select a single row by its index label
df.loc[0]                    # Returns the entire row for Alice

# 2. Select multiple specific rows
df.loc[[0, 2, 4]]            # Rows with labels 0, 2, 4 (Alice, Charlie, Eve)

# 3. Select a slice of rows (inclusive!)
df.loc[1:3]                   # Rows 1 to 3 → Bob, Charlie, David

# 4. Select specific rows and specific columns
df.loc[0:2, ['Name', 'Salary']]   # Rows 0-2, only Name and Salary columns

# 5. Select a single cell (row label, column name)
df.loc[2, 'Age']              # Returns 35 (Charlie's age)

# 6. Boolean indexing – select rows based on a condition
df.loc[df['Age'] > 30]         # All people older than 30

# 7. Multiple conditions
df.loc[(df['Age'] > 30) & (df['Salary'] > 65000)]   # Age > 30 AND Salary > 65000

iloc[] – Position-Based Selection

iloc[] uses integer positions (like normal Python lists). It ignores the actual index labels and column names.

Key points for beginners:

  • It counts from 0 (first row is position 0, first column is position 0).
  • Slicing excludes the end position (just like Python lists).
  • You always use numbers for both rows and columns.
# 1. Select the first row (position 0)
df.iloc[0]                   # Alice's row

# 2. Select first 3 rows (positions 0, 1, 2)
df.iloc[0:3]                 # Rows 0 to 2 (excludes position 3)

# 3. Select specific rows by position
df.iloc[[0, 2, 4]]           # Same rows as loc[[0,2,4]]

# 4. First 3 rows, first 2 columns (positions 0 and 1)
df.iloc[0:3, 0:2]            # Name and Age for first 3 people

# 5. Select a single cell by position
df.iloc[2, 1]                # Row position 2, column position 1 → Age of Charlie (35)

# 6. Last row (using negative indexing, like Python lists)
df.iloc[-1]                  # Eve's row

Key Differences at a Glance

Feature loc[] iloc[]
Uses Labels (index values & column names) Integer positions (0, 1, 2...)
Slicing Inclusive (loc[1:3] → rows 1,2,3) Exclusive (iloc[1:3] → rows 1,2 only)
Best for When you know names/labels or use conditions When you want "first N rows" or positional access

Happy coding 🐼!

Comments