Python Data Type Conversion: Convert Between Strings, Integers, Floats, Lists, and Other Types
NumPy arrays have a fixed data type (dtype) for every element. This is what makes them fast and memory-efficient — but it also means you sometimes need to convert types to fit your needs. 🔄
What Are Data Types in NumPy?
Think of dtype like choosing the right box size for your items:
int8: Tiny box (1 byte) — holds -128 to 127int32: Medium box (4 bytes) — much larger rangefloat64: Default for decimals — precise but uses more space
Smaller dtypes save memory and speed up calculations.
Checking and Setting Dtypes
import numpy as np
arr = np.array([1, 2, 3])
print("Default dtype:", arr.dtype) # Usually int64 or int32
# Specify dtype
small_int = np.array([1, 2, 3], dtype=np.int8)
print("int8 dtype:", small_int.dtype)
print("Memory used:", small_int.nbytes, "bytes") # 3 bytes
Default dtype: int64
int8 dtype: int8
Memory used: 3 bytes
Converting with astype()
astype() creates a new array with the desired type.
# Float to int (truncates decimals)
float_arr = np.array([1.9, 2.5, 3.7])
int_arr = float_arr.astype(np.int32)
print("Converted to int:", int_arr) # [1 2 3]
# Int to float
back_to_float = int_arr.astype(np.float64)
print("Back to float:", back_to_float)
Converted to int: [1 2 3]
Back to float: [1. 2. 3.]
astype() always returns a copy — safe but uses extra memory.
Real-World Example: Student Marks Dataset
Marks are whole numbers (0-100), attendance percentages have decimals.
marks = np.array([
[85, 88, 92],
[90, 76, 85],
[78, 92, 88],
[92, 85, 79],
[88, 90, 94]
])
print("Original dtype:", marks.dtype) # Likely int64
print("Memory:", marks.nbytes, "bytes") # 15 elements * 8 bytes = 120
Original dtype: int64
Memory: 120 bytes
Downcast to Smaller Integer
# Safe downcast to uint8 (unsigned 0-255)
marks_uint8 = marks.astype(np.uint8)
print("New dtype:", marks_uint8.dtype)
print("Memory now:", marks_uint8.nbytes, "bytes") # 15 bytes!
print("Values unchanged:", np.array_equal(marks, marks_uint8))
New dtype: uint8
Memory now: 15 bytes
Values unchanged: True
8x memory savings — huge for large datasets!
Adding Decimal Data (Attendance)
attendance = np.array([
[95.5, 90.0, 94.2],
[87.3, 85.1, 88.8],
[92.0, 90.5, 85.0],
[88.8, 86.2, 90.1],
[96.0, 92.3, 95.5]
])
# Combine — NumPy upcasts to float64
combined = np.hstack((marks_uint8, attendance.astype(np.float32)))
print("Combined dtype:", combined.dtype)
Combined dtype: float64
Unsafe Conversions and Overflow
# Overflow example
large_nums = np.array([200, 300], dtype=np.uint8)
print("Overflow wraps around:", large_nums) # 200 and 44 (300-256)
# Float to int truncation
decimals = np.array([1.9, -2.3])
negative_int = decimals.astype(np.int8)
print("Truncated:", negative_int) # [1 -2]
Overflow wraps around: [200 44]
Truncated: [ 1 -2]
Viewing as Different Type (Advanced)
Use view() to reinterpret bytes — no copy, but risky!
float_arr = np.array([1.0, 2.0], dtype=np.float64)
int_view = float_arr.view(np.int64)
print("As integers:", int_view) # Raw byte representation
Beginner Mistakes - Common Errors and How to Avoid Them
view() casually — it's for experts and can cause hard-to-debug issues.
Optimization Tips
float32 instead of float64 for ML/models — halves memory with minimal precision loss.
arr.min() and arr.max() before downcasting.
Real-World Use Cases
- Image Processing: uint8 for pixels (0-255)
- Machine Learning: float32 for weights
- IoT/Sensors: Small integers for limited memory
- Big Data: Downcasting saves gigabytes
Quick Summary 📝
- Dtypes control precision and memory
astype()for safe conversions- Downcast when possible, watch for overflow
- Upcasting happens automatically when needed
Mastering dtypes makes your code faster and leaner ✨
Comments
Post a Comment