Skip to main content

Indexing and Slicing

Calculating read time…

NumPy arrays power everything from simple calculations to machine learning models. Mastering indexing and slicing is your key to unlocking their full potential — it's how you grab exactly the data you need, fast and efficiently. 🔍

We'll start simple and gradually build to advanced, real-world techniques. By the end, you'll handle complex multi-dimensional data like a pro! 🐍✨

Creating NumPy Arrays - Your Foundation

Always start with:

import numpy as np

1D, 2D, and 3D Arrays

# 1D: Like a list or row
arr_1d = np.array([10, 20, 30, 40, 50])

# 2D: Like a table
arr_2d = np.array([
    [1, 2, 3],
    [4, 5, 6],
    [7, 8, 9]
])

# 3D: Like a stack of tables (e.g., RGB images)
arr_3d = np.array([
    [[101, 102, 103], [104, 105, 106]],  # Layer 1
    [[201, 202, 203], [204, 205, 206]],  # Layer 2
    [[301, 302, 303], [304, 305, 306]]   # Layer 3
])

print("3D shape:", arr_3d.shape)  # (3, 2, 3)
3D shape: (3, 2, 3)
🟢 DO: Use higher dimensions early — real data (images, videos, sensor readings) is rarely just 2D!

Basic Indexing - Single Elements

Zero-based indexing, just like Python lists.

1D and 2D

print(arr_1d[0])      # 10
print(arr_1d[-1])     # 50 (last)

print(arr_2d[1, 2])   # Row 1, Col 2 → 6

3D Indexing (layer, row, column)

print(arr_3d[2, 0, 1])  # Layer 2, Row 0, Col 1 → 302
🟡 Tip: Always check array.shape first — it tells you the exact order of dimensions!

Basic Slicing - Sub-Arrays

Format: start:end:step (end exclusive)

# 1D examples
print(arr_1d[1:4])     # [20 30 40]
print(arr_1d[::2])     # [10 30 50] (every second)
print(arr_1d[::-1])    # [50 40 30 20 10] (reverse)

# 2D examples
print(arr_2d[0:2, 1:])  # Rows 0-1, columns 1 to end
[[2 3]
 [5 6]]

Advanced Slicing with Steps and Negative Indices

# Reverse rows but keep columns normal
print(arr_2d[::-1, :])

# Every second element in both dimensions
print(arr_2d[::2, ::2])
[[7 8 9]
 [4 5 6]
 [1 2 3]]

[[1 3]
 [7 9]]

Real-World Example: Student Performance Dashboard

Let's build a complete analysis using only NumPy — no Pandas yet!

students = np.array(['Alice', 'Bob', 'Charlie', 'David', 'Eva'])

# Shape: (5 students, 3 subjects, 2 metrics: marks + attendance)
data = np.array([
    [[85, 95], [88, 90], [92, 94]],  # Alice: Math, Science, English
    [[90, 87], [76, 85], [85, 88]],
    [[78, 92], [92, 90], [88, 85]],
    [[92, 88], [85, 86], [79, 90]],
    [[88, 96], [90, 92], [94, 95]]
])

marks = data[:, :, 0]      # All students, all subjects, marks only
attendance = data[:, :, 1] # Attendance

print("Marks:\n", marks)
print("Attendance:\n", attendance)
Marks:
 [[85 88 92]
 [90 76 85]
 [78 92 88]
 [92 85 79]
 [88 90 94]]
Attendance:
 [[95 90 94]
 [87 85 88]
 [92 90 85]
 [88 86 90]
 [96 92 95]]

Complex Calculations

# Total and average marks per student
total_marks = marks.sum(axis=1)
avg_marks = marks.mean(axis=1)

# Average attendance per subject
avg_attendance_per_subject = attendance.mean(axis=0)

print("Total marks:", total_marks)
print("Average marks:", avg_marks)
print("Avg attendance per subject:", avg_attendance_per_subject)
Total marks: [265 251 258 256 272]
Average marks: [88.33333333 83.66666667 86.         85.33333333 90.66666667]
Avg attendance per subject: [91.4 88.6 90.4]

Conditional Grading (np.where)

grades = np.where(avg_marks >= 90, 'A',
              np.where(avg_marks >= 80, 'B', 'C'))

print("Grades:", grades)
Grades: ['B' 'C' 'B' 'B' 'A']

Ranking Students

ranks = (-total_marks).argsort().argsort() + 1  # Descending rank
print("Ranks:", ranks)
Ranks: [2 5 3 4 1]
🟢 DO: Use argsort() twice for stable ranking — it's a pro trick!

Boolean Indexing - Powerful Filtering

# Students with average >= 88
top_students = students[avg_marks >= 88]
print("Top performers:", top_students)

# Marks of students with low attendance in any subject (< 90)
low_att = students[np.any(attendance < 90, axis=1)]
print("Need attendance warning:", low_att)
Top performers: ['Alice' 'Eva']
Need attendance warning: ['Bob' 'Charlie' 'David']

Fancy Indexing - Irregular Selections

# Marks of Alice, Charlie, Eva (indices 0,2,4)
selected = marks[[0, 2, 4]]
print("Selected marks:\n", selected)

# Swap two students' data
marks[[0, 1]] = marks[[1, 0]]  # Swap Alice and Bob
print("After swap:\n", marks[0])  # Now shows Bob's original

Views vs Copies - Critical for Large Data

sub = marks[:2]      # View!
sub[0, 0] = 999
print("Original changed:", marks[0, 0])  # 999

safe = marks[:2].copy()  # Independent copy
safe[0, 0] = 111
print("Original safe:", marks[0, 0])     # Still 999
🔴 DON'T: Assume slices are copies — they are views! Always .copy() when you want independence.

Advanced Real-World: Image Cropping Simulation

# Simulate a 100x100 RGB image
img = np.random.randint(0, 256, (100, 100, 3), dtype=np.uint8)

# Crop center 50x50 region
center_crop = img[25:75, 25:75]

# Flip horizontally
flipped = img[:, ::-1]

# Extract only red channel
red_channel = img[:, :, 0]

print("Center crop shape:", center_crop.shape)
Center crop shape: (50, 50, 3)

Beginner Mistakes to Avoid

🔴 DON'T: Use single index on multi-dim arrays when you need a slice — it reduces dimensions unexpectedly.
🔴 DON'T: Forget copy() in data pipelines — bugs from shared memory are hard to trace.

Comments