NumPy arrays power everything from simple calculations to machine learning models. Mastering indexing and slicing is your key to unlocking their full potential — it's how you grab exactly the data you need, fast and efficiently. 🔍
We'll start simple and gradually build to advanced, real-world techniques. By the end, you'll handle complex multi-dimensional data like a pro! 🐍✨
Creating NumPy Arrays - Your Foundation
Always start with:
import numpy as np
1D, 2D, and 3D Arrays
# 1D: Like a list or row
arr_1d = np.array([10, 20, 30, 40, 50])
# 2D: Like a table
arr_2d = np.array([
[1, 2, 3],
[4, 5, 6],
[7, 8, 9]
])
# 3D: Like a stack of tables (e.g., RGB images)
arr_3d = np.array([
[[101, 102, 103], [104, 105, 106]], # Layer 1
[[201, 202, 203], [204, 205, 206]], # Layer 2
[[301, 302, 303], [304, 305, 306]] # Layer 3
])
print("3D shape:", arr_3d.shape) # (3, 2, 3)
3D shape: (3, 2, 3)
Basic Indexing - Single Elements
Zero-based indexing, just like Python lists.
1D and 2D
print(arr_1d[0]) # 10
print(arr_1d[-1]) # 50 (last)
print(arr_2d[1, 2]) # Row 1, Col 2 → 6
3D Indexing (layer, row, column)
print(arr_3d[2, 0, 1]) # Layer 2, Row 0, Col 1 → 302
array.shape first — it tells you the exact order of dimensions!
Basic Slicing - Sub-Arrays
Format: start:end:step (end exclusive)
# 1D examples
print(arr_1d[1:4]) # [20 30 40]
print(arr_1d[::2]) # [10 30 50] (every second)
print(arr_1d[::-1]) # [50 40 30 20 10] (reverse)
# 2D examples
print(arr_2d[0:2, 1:]) # Rows 0-1, columns 1 to end
[[2 3]
[5 6]]
Advanced Slicing with Steps and Negative Indices
# Reverse rows but keep columns normal
print(arr_2d[::-1, :])
# Every second element in both dimensions
print(arr_2d[::2, ::2])
[[7 8 9]
[4 5 6]
[1 2 3]]
[[1 3]
[7 9]]
Real-World Example: Student Performance Dashboard
Let's build a complete analysis using only NumPy — no Pandas yet!
students = np.array(['Alice', 'Bob', 'Charlie', 'David', 'Eva'])
# Shape: (5 students, 3 subjects, 2 metrics: marks + attendance)
data = np.array([
[[85, 95], [88, 90], [92, 94]], # Alice: Math, Science, English
[[90, 87], [76, 85], [85, 88]],
[[78, 92], [92, 90], [88, 85]],
[[92, 88], [85, 86], [79, 90]],
[[88, 96], [90, 92], [94, 95]]
])
marks = data[:, :, 0] # All students, all subjects, marks only
attendance = data[:, :, 1] # Attendance
print("Marks:\n", marks)
print("Attendance:\n", attendance)
Marks:
[[85 88 92]
[90 76 85]
[78 92 88]
[92 85 79]
[88 90 94]]
Attendance:
[[95 90 94]
[87 85 88]
[92 90 85]
[88 86 90]
[96 92 95]]
Complex Calculations
# Total and average marks per student
total_marks = marks.sum(axis=1)
avg_marks = marks.mean(axis=1)
# Average attendance per subject
avg_attendance_per_subject = attendance.mean(axis=0)
print("Total marks:", total_marks)
print("Average marks:", avg_marks)
print("Avg attendance per subject:", avg_attendance_per_subject)
Total marks: [265 251 258 256 272]
Average marks: [88.33333333 83.66666667 86. 85.33333333 90.66666667]
Avg attendance per subject: [91.4 88.6 90.4]
Conditional Grading (np.where)
grades = np.where(avg_marks >= 90, 'A',
np.where(avg_marks >= 80, 'B', 'C'))
print("Grades:", grades)
Grades: ['B' 'C' 'B' 'B' 'A']
Ranking Students
ranks = (-total_marks).argsort().argsort() + 1 # Descending rank
print("Ranks:", ranks)
Ranks: [2 5 3 4 1]
argsort() twice for stable ranking — it's a pro trick!
Boolean Indexing - Powerful Filtering
# Students with average >= 88
top_students = students[avg_marks >= 88]
print("Top performers:", top_students)
# Marks of students with low attendance in any subject (< 90)
low_att = students[np.any(attendance < 90, axis=1)]
print("Need attendance warning:", low_att)
Top performers: ['Alice' 'Eva']
Need attendance warning: ['Bob' 'Charlie' 'David']
Fancy Indexing - Irregular Selections
# Marks of Alice, Charlie, Eva (indices 0,2,4)
selected = marks[[0, 2, 4]]
print("Selected marks:\n", selected)
# Swap two students' data
marks[[0, 1]] = marks[[1, 0]] # Swap Alice and Bob
print("After swap:\n", marks[0]) # Now shows Bob's original
Views vs Copies - Critical for Large Data
sub = marks[:2] # View!
sub[0, 0] = 999
print("Original changed:", marks[0, 0]) # 999
safe = marks[:2].copy() # Independent copy
safe[0, 0] = 111
print("Original safe:", marks[0, 0]) # Still 999
.copy() when you want independence.
Advanced Real-World: Image Cropping Simulation
# Simulate a 100x100 RGB image
img = np.random.randint(0, 256, (100, 100, 3), dtype=np.uint8)
# Crop center 50x50 region
center_crop = img[25:75, 25:75]
# Flip horizontally
flipped = img[:, ::-1]
# Extract only red channel
red_channel = img[:, :, 0]
print("Center crop shape:", center_crop.shape)
Center crop shape: (50, 50, 3)
Beginner Mistakes to Avoid
copy() in data pipelines — bugs from shared memory are hard to trace.
Comments
Post a Comment