Skip to main content

Posts

Showing posts with the label NumPy

NumPy Data Cleaning: Handling NaN, Null, and Infinite Values

In NumPy, certain operations produce results that aren't ordinary numbers. These special values — np.nan , np.inf , and np.NINF — represent missing data, infinity, and negative infinity. Understanding them is crucial for clean, reliable data work! 🔍 What Are These Special Values? NumPy uses IEEE 754 floating-point standards for these: np.nan : "Not a Number" — for undefined or missing results np.inf : Positive infinity np.NINF (or -np.inf ): Negative infinity Analogy: np.nan is like a blank answer on a test — something went wrong. Infinity is like dividing by zero — the result grows without bound. 🟢 DO: Treat these as signals — they tell you something important about your data or calculation! Creating Special Values import numpy as np print("NaN:", np.nan) print("Positive infinity:", np.inf) print("Negative infinity:", np.NINF) # or -np.inf print("From operations:") print("0/0:", np.array([...

NumPy Universal Functions (ufuncs): Fast Element-Wise Operations on Arrays

Universal Functions, or ufuncs , are the secret behind NumPy's incredible speed. They perform element-wise operations on arrays — fast, clean, and without any Python loops. 🧩 Every time you write arr + 10 or np.sin(arr) , you're using a ufunc. What Makes ufuncs Special? ufuncs are functions that: Operate element-by-element Support broadcasting Are implemented in fast C code Preserve array shape (usually) Analogy: Think of a ufunc as a factory machine that processes every item on a conveyor belt (your array) simultaneously — no manual handling needed! 🟢 DO: Use ufuncs for any element-wise calculation — they're the most efficient way in NumPy! Basic Arithmetic ufuncs Common operators are actually ufuncs. import numpy as np arr = np.array([1, 2, 3, 4]) print("Add:", np.add(arr, 10)) # or arr + 10 print("Subtract:", np.subtract(arr, 1)) # or arr - 1 print("Multiply:", np.multiply(arr, 2)) # or arr * 2 ...

Python Data Type Conversion: Convert Between Strings, Integers, Floats, Lists, and Other Types

NumPy arrays have a fixed data type (dtype) for every element. This is what makes them fast and memory-efficient — but it also means you sometimes need to convert types to fit your needs. 🔄 What Are Data Types in NumPy? Think of dtype like choosing the right box size for your items: int8 : Tiny box (1 byte) — holds -128 to 127 int32 : Medium box (4 bytes) — much larger range float64 : Default for decimals — precise but uses more space Smaller dtypes save memory and speed up calculations. 🟢 DO: Choose the smallest dtype that fits your data — it saves memory and boosts performance! Checking and Setting Dtypes import numpy as np arr = np.array([1, 2, 3]) print("Default dtype:", arr.dtype) # Usually int64 or int32 # Specify dtype small_int = np.array([1, 2, 3], dtype=np.int8) print("int8 dtype:", small_int.dtype) print("Memory used:", small_int.nbytes, "bytes") # 3 bytes Default dtype: int64 int8 dtype: int8 Memory u...

NumPy vs Python Lists: Performance, Memory Usage, and Speed Compared

One of NumPy's biggest advantages is speed. When handling large numerical data, NumPy arrays can be 10-100 times faster than regular Python lists — and use less memory too! ⚡ Why NumPy Arrays Are Faster Think of Python lists like a chain of boxes — each box can hold anything (numbers, strings, even other boxes), but they're scattered around memory. NumPy arrays are like a tight grid of identical crates — all the same type, packed continuously. This lets NumPy use optimized C code and hardware tricks. Fixed data type → No type checking overhead Contiguous memory → Faster access and cache-friendly Vectorized operations → No slow Python loops 🟢 DO: Switch to NumPy arrays for numerical data with more than a few hundred elements — the speedup is worth it! Speed Comparison: Element-wise Operations Let's add two large collections of numbers. import numpy as np import time size = 1_000_000 # 1 million elements # Python lists list1 = list(range(size))...

NumPy Random Module: Generate Random Numbers, Arrays, and Samples

Random numbers are everywhere in programming: testing code, simulating real-world events, or creating fake datasets. NumPy makes generating them fast, repeatable, and flexible! 🎲 Why NumPy Random? NumPy's random functions are: Super fast (vectorized) Reproducible with seeds Full of real-world distributions 🟢 DO: Use NumPy random for any project needing randomness — it's more reliable than Python's built-in random module! Setting a Seed - Make Results Repeatable Randomness should be controllable for testing and learning. import numpy as np # Set seed for reproducible results np.random.seed(42) print(np.random.rand(5)) # Always the same with seed 42 [0.37454012 0.95071431 0.73199394 0.59865848 0.15601864] Without seed, results change every run. With seed — perfect for debugging! Basic Random Numbers Uniform Floats (0 to 1) np.random.seed(42) print("5 random floats:", np.random.rand(5)) print("3x3 matrix:\n", np.ra...

NumPy Reshaping and Combining Arrays: Reshape, Stack, and Concatenate

Reshaping lets you change the structure of your data without copying it — like rearranging furniture in a room to fit better. Combining arrays merges multiple datasets into one — perfect for building bigger pictures from smaller pieces! 🧩 Why Reshape Arrays? Real data often comes in the "wrong" shape for your needs. Reshaping fixes that quickly and efficiently. Most reshaping functions return a view — no data is copied, so it's fast and memory-friendly. 🟢 DO: Reshape early and often — it makes broadcasting, plotting, and analysis much easier! Basic Reshaping with reshape() reshape() changes the array's dimensions while keeping the data. import numpy as np arr = np.array([1, 2, 3, 4, 5, 6]) # Reshape to 2 rows, 3 columns matrix = arr.reshape(2, 3) print("2x3 matrix:\n", matrix) # Use -1 to let NumPy figure out one dimension column = arr.reshape(-1, 1) # Auto-calculates rows print("Column vector shape:", column.shape) 2x3 ma...

NumPy Mathematical & Statistical Functions: Calculate and Analyze Data Efficiently

NumPy comes packed with fast mathematical and statistical functions that work on entire arrays at once. From simple sums to advanced statistics, these tools make data analysis feel effortless! 📊 Why NumPy's Functions Are Special These functions are vectorized — they operate on whole arrays without loops. They're also optimized in C, making them much faster than pure Python alternatives. 🟢 DO: Use NumPy's built-in functions whenever possible — they're accurate, fast, and handle edge cases well! Basic Aggregation Functions Start with the classics: sum, mean, min, max. import numpy as np arr = np.array([10, 20, 30, 40, 50]) print("Sum:", np.sum(arr)) print("Mean:", np.mean(arr)) print("Min:", np.min(arr)) print("Max:", np.max(arr)) Sum: 150 Mean: 30.0 Min: 10 Max: 50 Or use methods directly: print(arr.sum()) print(arr.mean()) The Power of 'axis' Parameter For 2D arrays, axis controls direction:...

NumPy Broadcasting Explained: Efficient Array Operations Without Loops

Broadcasting is one of NumPy's most powerful features. It lets you perform operations on arrays with different shapes — without writing loops or making unnecessary copies. It's like NumPy automatically stretches smaller arrays to match larger ones! ✨ What is Broadcasting? A Simple Analogy Imagine you have a big grid of numbers (a 2D array) and want to add 10 to every value. In regular Python, you'd loop through everything. In NumPy, you just write grid + 10 — the scalar 10 is "broadcast" to match the grid's shape. It's efficient, fast, and saves memory because no extra copies are made. 🟢 DO: Embrace broadcasting — it's the key to clean, fast NumPy code! Basic Broadcasting - Start Simple Import NumPy first: import numpy as np Scalar Broadcasting arr = np.array([1, 2, 3, 4, 5]) print(arr + 10) print(arr * 2.5) [11 12 13 14 15] [ 2.5 5. 7.5 10. 12.5] The scalar stretches to match the array. 1D with 2D - Adding a Vector ...

Vectorized Operations in NumPy: Fast and Efficient Array Computation

NumPy's real superpower is vectorization : doing operations on entire arrays at once, without writing slow Python loops. It's like upgrading from walking to flying! 🚀 We'll start with simple math on arrays and build up to powerful, real-world data analysis. No prior NumPy knowledge needed — just follow along! 🐍 What Are Vectorized Operations? Think of regular Python lists: to add two lists element-wise, you need a loop. NumPy arrays let you skip the loop entirely — operations apply to every element automatically. This is fast because it's written in optimized C code under the hood. 🟢 DO: Always prefer vectorized operations over Python loops when working with NumPy arrays — speed gains can be 10-100x! Basic Arithmetic - No Loops Needed First, import NumPy: import numpy as np arr = np.array([1, 2, 3, 4, 5]) # Add 10 to every element print(arr + 10) # Multiply by 2 print(arr * 2) # Element-wise operations between arrays arr2 = np.array([10, 20, 30...

Indexing and Slicing

NumPy arrays power everything from simple calculations to machine learning models. Mastering indexing and slicing is your key to unlocking their full potential — it's how you grab exactly the data you need, fast and efficiently. 🔍 We'll start simple and gradually build to advanced, real-world techniques. By the end, you'll handle complex multi-dimensional data like a pro! 🐍✨ Creating NumPy Arrays - Your Foundation Always start with: import numpy as np 1D, 2D, and 3D Arrays # 1D: Like a list or row arr_1d = np.array([10, 20, 30, 40, 50]) # 2D: Like a table arr_2d = np.array([ [1, 2, 3], [4, 5, 6], [7, 8, 9] ]) # 3D: Like a stack of tables (e.g., RGB images) arr_3d = np.array([ [[101, 102, 103], [104, 105, 106]], # Layer 1 [[201, 202, 203], [204, 205, 206]], # Layer 2 [[301, 302, 303], [304, 305, 306]] # Layer 3 ]) print("3D shape:", arr_3d.shape) # (3, 2, 3) 3D shape: (3, 2, 3) 🟢 DO: Use higher dimensions early —...

NumPy Indexing and Slicing: Access and Manipulate Array Elements

Think of a NumPy array like a container with a very specific structure . Just like you need to know if a box is small or large, flat or deep, and what kind of items it holds, you need to understand your array's shape, dimensions, and data type. These three concepts are the foundation of working with NumPy! Why Do Shape and Dimensions Matter? Imagine you're trying to add two arrays together, but one is a list of 5 numbers and the other is a table with 3 rows and 4 columns. What should happen? NumPy will throw an error because the shapes don't match! import numpy as np a = np.array([1, 2, 3, 4, 5]) # Shape: (5,) b = np.array([[1, 2, 3, 4], [5, 6, 7, 8], [9, 10, 11, 12]]) # Shape: (3, 4) # result = a + b # ERROR! Shapes don't match Understanding shapes prevents these errors and helps you manipulate data correctly. Let's break it down step by step! 🎯 What is a Dimension? (ndim) The dimension (or rank) tells you how man...

NumPy Arrays Explained: Create, Index, and Manipulate Arrays in Python

Think of NumPy arrays as supercharged lists that can do math at lightning speed. While Python lists are great for general tasks, NumPy arrays are built specifically for numbers and calculations. They're the backbone of data science, machine learning, and scientific computing in Python! Why NumPy? The Problem with Regular Lists Let's say you have a list of temperatures in Celsius and want to convert them all to Fahrenheit. # Using regular Python lists celsius = [0, 10, 20, 30, 40] # You have to use a loop fahrenheit = [] for temp in celsius: fahrenheit.append((temp * 9/5) + 32) print(fahrenheit) # [32.0, 50.0, 68.0, 86.0, 104.0] This works, but it's slow and requires writing loops for every calculation. Now imagine you have 1 million temperatures to convert! With NumPy, the same task becomes super simple: import numpy as np # Using NumPy arrays celsius = np.array([0, 10, 20, 30, 40]) # Just do the math directly - no loops! fahrenheit = (celsius * 9...