Think of a NumPy array like a container with a very specific structure. Just like you need to know if a box is small or large, flat or deep, and what kind of items it holds, you need to understand your array's shape, dimensions, and data type. These three concepts are the foundation of working with NumPy!
Why Do Shape and Dimensions Matter?
Imagine you're trying to add two arrays together, but one is a list of 5 numbers and the other is a table with 3 rows and 4 columns. What should happen? NumPy will throw an error because the shapes don't match!
import numpy as np
a = np.array([1, 2, 3, 4, 5]) # Shape: (5,)
b = np.array([[1, 2, 3, 4],
[5, 6, 7, 8],
[9, 10, 11, 12]]) # Shape: (3, 4)
# result = a + b # ERROR! Shapes don't match
Understanding shapes prevents these errors and helps you manipulate data correctly. Let's break it down step by step! 🎯
What is a Dimension? (ndim)
The dimension (or rank) tells you how many axes an array has. Think of it as "how many index numbers do you need to access a single element?"
0D Array - A Single Number (Scalar)
Zero dimensions means just one value - no axes at all.
import numpy as np
scalar = np.array(42)
print(scalar)
# Output: 42
print(f"Dimensions: {scalar.ndim}")
# Output: Dimensions: 0
print(f"Shape: {scalar.shape}")
# Output: Shape: ()
# Access the value
print(scalar[()]) # Empty tuple for zero dimensions
# Output: 42
Real-world analogy: A single temperature reading: 23°C
1D Array - A List of Numbers (Vector)
One dimension means one axis - like a single row or column. You need one index to access an element.
import numpy as np
vector = np.array([10, 20, 30, 40, 50])
print(vector)
# Output: [10 20 30 40 50]
print(f"Dimensions: {vector.ndim}")
# Output: Dimensions: 1
print(f"Shape: {vector.shape}")
# Output: Shape: (5,)
# Access an element with ONE index
print(vector[2])
# Output: 30
Real-world analogy: Temperatures for 5 consecutive days: [23, 25, 22, 24, 26]
For a 1D array with 5 elements, the shape is (5,) - note the comma!
This distinguishes it from a scalar or higher dimensions.
The comma means "this is a tuple with one element."
2D Array - A Table (Matrix)
Two dimensions means two axes - rows and columns. You need two indices to access an element: [row, column]
import numpy as np
matrix = np.array([[1, 2, 3, 4],
[5, 6, 7, 8],
[9, 10, 11, 12]])
print(matrix)
# Output:
# [[ 1 2 3 4]
# [ 5 6 7 8]
# [ 9 10 11 12]]
print(f"Dimensions: {matrix.ndim}")
# Output: Dimensions: 2
print(f"Shape: {matrix.shape}")
# Output: Shape: (3, 4)
# Access an element with TWO indices [row, column]
print(matrix[1, 2]) # Row 1, Column 2
# Output: 7
Real-world analogy: A spreadsheet with 3 rows (days) and 4 columns (cities), showing temperatures for each city on each day.
3D Array - A Stack of Tables (Tensor)
Three dimensions means three axes - like multiple tables stacked together. You need three indices: [depth, row, column]
import numpy as np
tensor = np.array([
[[1, 2], [3, 4]], # First "table"
[[5, 6], [7, 8]], # Second "table"
[[9, 10], [11, 12]] # Third "table"
])
print(tensor)
# Output:
# [[[ 1 2]
# [ 3 4]]
#
# [[ 5 6]
# [ 7 8]]
#
# [[ 9 10]
# [11 12]]]
print(f"Dimensions: {tensor.ndim}")
# Output: Dimensions: 3
print(f"Shape: {tensor.shape}")
# Output: Shape: (3, 2, 2)
# Access an element with THREE indices [depth, row, column]
print(tensor[1, 0, 1]) # Second table, first row, second column
# Output: 6
Real-world analogy: A video is a 3D array! Each frame is a 2D image, and stacking frames creates the third dimension (time). For colored images, you'd even have 4D: [time, height, width, color_channels]
Understanding Shape in Detail
The shape is a tuple that tells you the size along each dimension. It's one of the most important attributes you'll use!
Reading Shape Values
import numpy as np
# 1D array
arr1d = np.array([1, 2, 3, 4, 5])
print(f"Shape: {arr1d.shape}")
# Output: Shape: (5,)
# Meaning: 5 elements in the first (and only) dimension
# 2D array
arr2d = np.array([[1, 2, 3],
[4, 5, 6]])
print(f"Shape: {arr2d.shape}")
# Output: Shape: (2, 3)
# Meaning: 2 rows, 3 columns
# 3D array
arr3d = np.array([
[[1, 2], [3, 4], [5, 6]],
[[7, 8], [9, 10], [11, 12]]
])
print(f"Shape: {arr3d.shape}")
# Output: Shape: (2, 3, 2)
# Meaning: 2 "blocks", each with 3 rows and 2 columns
1D: (n,) → n elements
2D: (rows, columns) → rows × columns grid
3D: (depth, rows, columns) → depth × rows × columns cube
Shape and Size Relationship
The size attribute tells you the total number of elements. It's always the product of all numbers in the shape!
import numpy as np
arr = np.array([[1, 2, 3, 4],
[5, 6, 7, 8],
[9, 10, 11, 12]])
print(f"Shape: {arr.shape}")
# Output: Shape: (3, 4)
print(f"Size: {arr.size}")
# Output: Size: 12
# Verify: 3 rows × 4 columns = 12 total elements ✓
Visualizing Different Shapes
Shape (5,) - 1D Array:
[■][■][■][■][■]
Shape (2, 3) - 2D Array:
[■][■][■]
[■][■][■]
Shape (3, 4) - 2D Array:
[■][■][■][■]
[■][■][■][■]
[■][■][■][■]
Shape (2, 2, 3) - 3D Array:
Layer 1: Layer 2:
[■][■][■] [■][■][■]
[■][■][■] [■][■][■]
Data Types (dtype) - What Kind of Numbers?
Every NumPy array holds elements of the same data type. The dtype tells you what kind of data you're working with.
Common Data Types
import numpy as np
# Integer types
int_array = np.array([1, 2, 3, 4, 5])
print(f"Dtype: {int_array.dtype}")
# Output: Dtype: int64 (64-bit integers)
# Float types
float_array = np.array([1.5, 2.7, 3.2])
print(f"Dtype: {float_array.dtype}")
# Output: Dtype: float64 (64-bit floating point)
# Boolean type
bool_array = np.array([True, False, True, True])
print(f"Dtype: {bool_array.dtype}")
# Output: Dtype: bool
# String type
string_array = np.array(['apple', 'banana', 'cherry'])
print(f"Dtype: {string_array.dtype}")
# Output: Dtype:
int32, int64: Whole numbers (32-bit or 64-bit)
float32, float64: Decimal numbers (32-bit or 64-bit precision)
bool: True or False values
complex64, complex128: Complex numbers (a + bi)
Why Data Type Matters
Different data types use different amounts of memory and have different precision.
import numpy as np
# int32 uses 4 bytes per element
arr_int32 = np.array([1, 2, 3], dtype=np.int32)
print(f"Memory per element: {arr_int32.itemsize} bytes")
# Output: Memory per element: 4 bytes
# int64 uses 8 bytes per element (more memory, larger range)
arr_int64 = np.array([1, 2, 3], dtype=np.int64)
print(f"Memory per element: {arr_int64.itemsize} bytes")
# Output: Memory per element: 8 bytes
# For 1 million elements:
# int32: 1,000,000 × 4 = 4 MB
# int64: 1,000,000 × 8 = 8 MB
Explicit Type Specification
You can force a specific data type when creating arrays:
import numpy as np
# Force float even though inputs are integers
arr = np.array([1, 2, 3, 4], dtype=float)
print(arr)
# Output: [1. 2. 3. 4.]
print(f"Dtype: {arr.dtype}")
# Output: Dtype: float64
# Force int32 instead of default int64
arr = np.array([1, 2, 3, 4], dtype=np.int32)
print(f"Dtype: {arr.dtype}")
# Output: Dtype: int32
# Create array of zeros with specific type
zeros = np.zeros(5, dtype=int)
print(zeros)
# Output: [0 0 0 0 0]
print(f"Dtype: {zeros.dtype}")
# Output: Dtype: int64
Type Conversion (Casting)
Convert an existing array to a different data type:
import numpy as np
# Start with floats
float_arr = np.array([1.7, 2.3, 3.9, 4.1])
print(f"Original: {float_arr}, dtype: {float_arr.dtype}")
# Output: Original: [1.7 2.3 3.9 4.1], dtype: float64
# Convert to integers (decimals are truncated, not rounded!)
int_arr = float_arr.astype(int)
print(f"As int: {int_arr}, dtype: {int_arr.dtype}")
# Output: As int: [1 2 3 4], dtype: int64
# Convert to string
str_arr = float_arr.astype(str)
print(f"As str: {str_arr}, dtype: {str_arr.dtype}")
# Output: As str: ['1.7' '2.3' '3.9' '4.1'], dtype:
Converting from float to int truncates decimals (doesn't round)! 1.7 becomes 1, not 2. 4.9 becomes 4, not 5. Always be careful when converting types - you might lose precision or data.
Real-World Example: Image Data
Let's understand shapes and dtypes with a practical image example.
Grayscale Image (2D)
import numpy as np
# A tiny 5x5 grayscale image (values 0-255)
# 0 = black, 255 = white
grayscale_image = np.array([
[0, 50, 100, 150, 200],
[50, 100, 150, 200, 255],
[100, 150, 200, 255, 200],
[150, 200, 255, 200, 150],
[200, 255, 200, 150, 100]
], dtype=np.uint8) # Unsigned 8-bit integer (0-255)
print(f"Shape: {grayscale_image.shape}")
# Output: Shape: (5, 5)
print(f"Dimensions: {grayscale_image.ndim}")
# Output: Dimensions: 2
print(f"Data type: {grayscale_image.dtype}")
# Output: Data type: uint8
print(f"Total pixels: {grayscale_image.size}")
# Output: Total pixels: 25
print(f"Memory used: {grayscale_image.nbytes} bytes")
# Output: Memory used: 25 bytes (1 byte per pixel)
Color Image (3D)
import numpy as np
# A 3x3 color image (RGB - Red, Green, Blue channels)
# Shape: (height, width, channels)
color_image = np.array([
[[255, 0, 0], [0, 255, 0], [0, 0, 255]], # Row 1: Red, Green, Blue
[[255, 255, 0], [255, 0, 255], [0, 255, 255]], # Row 2: Yellow, Magenta, Cyan
[[255, 255, 255], [128, 128, 128], [0, 0, 0]] # Row 3: White, Gray, Black
], dtype=np.uint8)
print(f"Shape: {color_image.shape}")
# Output: Shape: (3, 3, 3)
# Meaning: 3 rows, 3 columns, 3 color channels (RGB)
print(f"Dimensions: {color_image.ndim}")
# Output: Dimensions: 3
print(f"Total values: {color_image.size}")
# Output: Total values: 27 (3×3×3)
# Access a specific pixel's RGB values
print(f"Top-left pixel (should be red): {color_image[0, 0]}")
# Output: Top-left pixel (should be red): [255 0 0]
# Access just the Red channel of the entire image
red_channel = color_image[:, :, 0]
print(f"Red channel shape: {red_channel.shape}")
# Output: Red channel shape: (3, 3)
Grayscale Photo (1920×1080): Shape (1080, 1920), 2D
Color Photo (1920×1080): Shape (1080, 1920, 3), 3D
Video (10 seconds, 30fps): Shape (300, 1080, 1920, 3), 4D
Reshaping Arrays - Changing Structure
Sometimes you need to reorganize your data into a different shape. The total number of elements must stay the same!
Basic Reshaping
import numpy as np
# Start with 1D array of 12 elements
arr = np.arange(12)
print(f"Original shape: {arr.shape}")
print(arr)
# Output: [0 1 2 3 4 5 6 7 8 9 10 11]
# Reshape to 3 rows × 4 columns
arr_2d = arr.reshape(3, 4)
print(f"\nReshaped to (3, 4):")
print(arr_2d)
# Output:
# [[ 0 1 2 3]
# [ 4 5 6 7]
# [ 8 9 10 11]]
# Reshape to 4 rows × 3 columns
arr_2d_alt = arr.reshape(4, 3)
print(f"\nReshaped to (4, 3):")
print(arr_2d_alt)
# Output:
# [[ 0 1 2]
# [ 3 4 5]
# [ 6 7 8]
# [ 9 10 11]]
# Reshape to 3D
arr_3d = arr.reshape(2, 3, 2)
print(f"\nReshaped to (2, 3, 2):")
print(arr_3d)
print(f"Shape: {arr_3d.shape}")
# Output: Shape: (2, 3, 2)
Auto-Calculating Dimensions with -1
Use -1 to let NumPy automatically figure out one dimension:
import numpy as np
arr = np.arange(24)
# "I want 4 rows, you figure out columns"
reshaped = arr.reshape(4, -1)
print(f"Shape: {reshaped.shape}")
# Output: Shape: (4, 6)
# NumPy calculated: 24 elements ÷ 4 rows = 6 columns
# "I want 3 columns, you figure out rows"
reshaped = arr.reshape(-1, 3)
print(f"Shape: {reshaped.shape}")
# Output: Shape: (8, 3)
# NumPy calculated: 24 elements ÷ 3 columns = 8 rows
# Works with 3D too!
reshaped = arr.reshape(2, -1, 3)
print(f"Shape: {reshaped.shape}")
# Output: Shape: (2, 4, 3)
# NumPy calculated: 24 ÷ (2 × 3) = 4
arr = np.arange(12) # 12 elements
# This will ERROR!
# arr.reshape(3, 5) # 3×5 = 15 elements (doesn't match 12)
# This will ERROR!
# arr.reshape(-1, -1) # Can't auto-calculate both dimensions!
The new shape must have the same total number of elements. You can only use -1 for ONE dimension.
Flattening to 1D
import numpy as np
arr_2d = np.array([[1, 2, 3],
[4, 5, 6],
[7, 8, 9]])
# Method 1: flatten() - creates a copy
flat1 = arr_2d.flatten()
print(flat1)
# Output: [1 2 3 4 5 6 7 8 9]
# Method 2: ravel() - returns a view (faster, shares memory)
flat2 = arr_2d.ravel()
print(flat2)
# Output: [1 2 3 4 5 6 7 8 9]
# Method 3: reshape to (-1)
flat3 = arr_2d.reshape(-1)
print(flat3)
# Output: [1 2 3 4 5 6 7 8 9]
Adding and Removing Dimensions
Adding Dimensions with np.newaxis
import numpy as np
# 1D array
arr = np.array([1, 2, 3, 4, 5])
print(f"Original shape: {arr.shape}")
# Output: Original shape: (5,)
# Add dimension to make it a column vector
arr_column = arr[:, np.newaxis]
print(f"Column shape: {arr_column.shape}")
# Output: Column shape: (5, 1)
print(arr_column)
# Output:
# [[1]
# [2]
# [3]
# [4]
# [5]]
# Add dimension to make it a row vector
arr_row = arr[np.newaxis, :]
print(f"Row shape: {arr_row.shape}")
# Output: Row shape: (1, 5)
print(arr_row)
# Output: [[1 2 3 4 5]]
Removing Dimensions with squeeze()
import numpy as np
# Array with unnecessary dimensions
arr = np.array([[[1, 2, 3]]])
print(f"Original shape: {arr.shape}")
# Output: Original shape: (1, 1, 3)
# Remove all dimensions of size 1
squeezed = arr.squeeze()
print(f"Squeezed shape: {squeezed.shape}")
# Output: Squeezed shape: (3,)
print(squeezed)
# Output: [1 2 3]
Transposing - Swapping Dimensions
import numpy as np
# 2D array
arr = np.array([[1, 2, 3],
[4, 5, 6]])
print(f"Original shape: {arr.shape}")
# Output: Original shape: (2, 3)
print(arr)
# Output:
# [[1 2 3]
# [4 5 6]]
# Transpose: swap rows and columns
transposed = arr.T
print(f"Transposed shape: {transposed.shape}")
# Output: Transposed shape: (3, 2)
print(transposed)
# Output:
# [[1 4]
# [2 5]
# [3 6]]
# Alternative method
transposed2 = np.transpose(arr)
print(f"Same result: {np.array_equal(transposed, transposed2)}")
# Output: Same result: True
Transposing Color Images
import numpy as np
# Color image: (height, width, channels)
image = np.random.randint(0, 256, size=(480, 640, 3), dtype=np.uint8)
print(f"Original image shape: {image.shape}")
# Output: Original image shape: (480, 640, 3)
# 480 pixels tall, 640 pixels wide, 3 color channels
# Some libraries expect (channels, height, width)
# Use transpose to rearrange
image_transposed = np.transpose(image, (2, 0, 1))
print(f"Transposed shape: {image_transposed.shape}")
# Output: Transposed shape: (3, 480, 640)
# 3 channels, 480 height, 640 width
Common Beginner Mistakes
Mistake 1: Confusing Shape and Size
arr = np.array([[1, 2, 3],
[4, 5, 6]])
# Thinking size is the same as shape
# size is NOT (2, 3)!
arr = np.array([[1, 2, 3],
[4, 5, 6]])
print(f"Shape: {arr.shape}") # Output: (2, 3) - tuple
print(f"Size: {arr.size}") # Output: 6 - total elements
print(f"ndim: {arr.ndim}") # Output: 2 - number of dimensions
# Remember:
# shape = dimensions of the array (tuple)
# size = total number of elements (integer)
# ndim = number of axes (integer)
Mistake 2: Forgetting Shape is Immutable with reshape()
arr = np.arange(12)
arr.reshape(3, 4) # Reshape but don't save result!
print(arr.shape)
# Output: (12,) - Still 1D! reshape() didn't modify arr
arr = np.arange(12)
arr = arr.reshape(3, 4) # Save the result!
print(arr.shape)
# Output: (3, 4) - Now it's 2D!
# Or use inplace with resize (modifies original)
arr = np.arange(12)
arr.resize(3, 4) # Modifies arr directly
print(arr.shape)
# Output: (3, 4)
Mistake 3: Type Conversion Surprises
arr = np.array([1.1, 2.5, 3.9, 4.3])
int_arr = arr.astype(int)
print(int_arr)
# Output: [1 2 3 4] - Decimals truncated, not rounded!
# 1.1 → 1 (not rounded to 1)
# 2.5 → 2 (not rounded to 3)
# 3.9 → 3 (not rounded to 4)
arr = np.array([1.1, 2.5, 3.9, 4.3])
# Round FIRST, then convert to int
int_arr = np.round(arr).astype(int)
print(int_arr)
# Output: [1 2 4 4] - Properly rounded!
# Or use np.rint() for round-to-nearest-even
int_arr = np.rint(arr).astype(int)
print(int_arr)
# Output: [1 2 4 4]
Mistake 4: Mixing Up Axis Numbers
grades = np.array([[85, 90, 78, 92],
[88, 76, 92, 85],
[92, 88, 85, 90]])
# Want: average per student (each row)
avg = np.mean(grades, axis=0) # WRONG! This averages columns
print(avg)
# Output: [88.33... 84.66... 85. 89.] - Class avg per subject
grades = np.array([[85, 90, 78, 92],
[88, 76, 92, 85],
[92, 88, 85, 90]])
# Average per student: axis=1 (across columns)
avg = np.mean(grades, axis=1)
print(avg)
# Output: [86.25 85.25 88.75] - Each student's average
# Remember:
# axis=0: operate DOWN the rows (result: one value per column)
# axis=1: operate ACROSS the columns (result: one value per row)
Advanced Tips and Tricks
Tip 1: Check Shape Before Operations
import numpy as np
def safe_array_operation(arr1, arr2):
"""Always check shapes before combining arrays"""
print(f"Array 1 shape: {arr1.shape}")
print(f"Array 2 shape: {arr2.shape}")
if arr1.shape != arr2.shape:
print("Warning: Shapes don't match!")
print("Broadcasting may occur or operation may fail.")
result = arr1 + arr2
print(f"Result shape: {result.shape}")
return result
# Test it
a = np.array([[1, 2], [3, 4]])
b = np.array([[5, 6], [7, 8]])
safe_array_operation(a, b)
Tip 2: Use dtype to Save Memory
import numpy as np
# Default: int64 (8 bytes per element)
arr_default = np.arange(1000000)
print(f"Memory (default int64): {arr_default.nbytes / 1024 / 1024:.2f} MB")
# Output: Memory (default int64): 7.63 MB
# Use int32 if values fit (4 bytes per element)
arr_int32 = np.arange(1000000, dtype=np.int32)
print(f"Memory (int32): {arr_int32.nbytes / 1024 / 1024:.2f} MB")
# Output: Memory (int32): 3.81 MB
# Use int16 for small values (2 bytes per element)
arr_int16 = np.arange(1000000, dtype=np.int16)
print(f"Memory (int16): {arr_int16.nbytes / 1024 / 1024:.2f} MB")
# Output: Memory (int16): 1.91 MB
# Save 75% memory by using int16 instead of int64!
Smaller data types have smaller ranges! int16 can only hold -32,768 to 32,767. If you overflow, values will wrap around (1000000 as int16 becomes 16960). Always ensure your data fits the type you choose!
Tip 3: Use Broadcasting Effectively
import numpy as np
# Normalize each column (subtract mean, divide by std)
data = np.array([[1, 2, 3],
[4, 5, 6],
[7, 8, 9]], dtype=float)
# Calculate mean and std for each column
col_means = np.mean(data, axis=0) # Shape: (3,)
col_stds = np.std(data, axis=0) # Shape: (3,)
print(f"Column means: {col_means}")
# Output: Column means: [4. 5. 6.]
print(f"Column stds: {col_stds}")
# Output: Column stds: [2.44948974 2.44948974 2.44948974]
# Broadcasting: (3,3) array with (3,) arrays
normalized = (data - col_means) / col_stds
print(f"Normalized data:\n{normalized}")
# Each column now has mean≈0, std≈1
Tip 4: Understand Memory Layout (C vs Fortran Order)
import numpy as np
# C-order (row-major, default): rows are contiguous in memory
arr_c = np.array([[1, 2, 3],
[4, 5, 6]], order='C')
# Fortran-order (column-major): columns are contiguous in memory
arr_f = np.array([[1, 2, 3],
[4, 5, 6]], order='F')
print(f"C-order flags: {arr_c.flags['C_CONTIGUOUS']}")
# Output: C-order flags: True
print(f"F-order flags: {arr_f.flags['F_CONTIGUOUS']}")
# Output: F-order flags: True
# Performance tip: operate along the contiguous dimension
# C-order: iterate rows faster
# F-order: iterate columns faster
Real-World Use Cases
Use Case 1: Preparing Data for Machine Learning
import numpy as np
# Simulate dataset: 1000 samples, 5 features
# Shape: (samples, features)
X = np.random.randn(1000, 5)
print(f"Dataset shape: {X.shape}")
# Output: Dataset shape: (1000, 5)
print(f"Data type: {X.dtype}")
# Output: Data type: float64
# Add bias term (column of 1s) for linear regression
bias = np.ones((1000, 1))
X_with_bias = np.hstack([bias, X])
print(f"With bias shape: {X_with_bias.shape}")
# Output: With bias shape: (1000, 6)
# Reshape labels for compatibility
# From (1000,) to (1000, 1)
y = np.random.randint(0, 2, size=1000)
y_reshaped = y.reshape(-1, 1)
print(f"Labels shape: {y_reshaped.shape}")
# Output: Labels shape: (1000, 1)
Use Case 2: Processing Video Frames
import numpy as np
# Simulate video: 100 frames, 480×640 pixels, RGB
# Shape: (frames, height, width, channels)
video = np.random.randint(0, 256, size=(100, 480, 640, 3), dtype=np.uint8)
print(f"Video shape: {video.shape}")
# Output: Video shape: (100, 480, 640, 3)
print(f"Video size: {video.nbytes / 1024 / 1024:.2f} MB")
# Output: Video size: 88.47 MB
# Extract first frame
first_frame = video[0]
print(f"Frame shape: {first_frame.shape}")
# Output: Frame shape: (480, 640, 3)
# Convert to grayscale (average RGB channels)
grayscale_video = np.mean(video, axis=3).astype(np.uint8)
print(f"Grayscale shape: {grayscale_video.shape}")
# Output: Grayscale shape: (100, 480, 640)
# Reduce size by 67%!
print(f"Grayscale size: {grayscale_video.nbytes / 1024 / 1024:.2f} MB")
# Output: Grayscale size: 29.49 MB
Use Case 3: Financial Time Series
import numpy as np
# Stock prices: 30 days, 5 stocks
# Shape: (days, stocks)
prices = np.random.uniform(100, 200, size=(30, 5))
print(f"Prices shape: {prices.shape}")
# Output: Prices shape: (30, 5)
# Calculate daily returns for each stock
# Returns = (Price_today - Price_yesterday) / Price_yesterday
returns = (prices[1:] - prices[:-1]) / prices[:-1]
print(f"Returns shape: {returns.shape}")
# Output: Returns shape: (29, 5)
# One less day because we need previous day for comparison
# Average return per stock
avg_returns = np.mean(returns, axis=0)
print(f"Avg returns per stock: {avg_returns}")
# Volatility per stock (std of returns)
volatility = np.std(returns, axis=0)
print(f"Volatility per stock: {volatility}")
Quick Summary 📝
ndim: Number of dimensions (axes) → 1D, 2D, 3D, etc.
shape: Size along each dimension → (rows, columns) for 2D
size: Total number of elements → product of all shape values
dtype: Data type → int32, float64, bool, etc.
reshape(): Change shape → must keep same total elements
Keep practicing, and remember: understanding shape, dimensions, and data types is the key to mastering NumPy! 🎉
Comments
Post a Comment