Skip to main content

Python Data Type Conversion: Convert Between Strings, Integers, Floats, Lists, and Other Types

Calculating read time…

NumPy arrays have a fixed data type (dtype) for every element. This is what makes them fast and memory-efficient — but it also means you sometimes need to convert types to fit your needs. 🔄

What Are Data Types in NumPy?

Think of dtype like choosing the right box size for your items:

  • int8: Tiny box (1 byte) — holds -128 to 127
  • int32: Medium box (4 bytes) — much larger range
  • float64: Default for decimals — precise but uses more space

Smaller dtypes save memory and speed up calculations.

🟢 DO: Choose the smallest dtype that fits your data — it saves memory and boosts performance!

Checking and Setting Dtypes

import numpy as np

arr = np.array([1, 2, 3])
print("Default dtype:", arr.dtype)  # Usually int64 or int32

# Specify dtype
small_int = np.array([1, 2, 3], dtype=np.int8)
print("int8 dtype:", small_int.dtype)
print("Memory used:", small_int.nbytes, "bytes")  # 3 bytes
Default dtype: int64
int8 dtype: int8
Memory used: 3 bytes

Converting with astype()

astype() creates a new array with the desired type.

# Float to int (truncates decimals)
float_arr = np.array([1.9, 2.5, 3.7])
int_arr = float_arr.astype(np.int32)
print("Converted to int:", int_arr)  # [1 2 3]

# Int to float
back_to_float = int_arr.astype(np.float64)
print("Back to float:", back_to_float)
Converted to int: [1 2 3]
Back to float: [1. 2. 3.]
🟡 Tip: astype() always returns a copy — safe but uses extra memory.

Real-World Example: Student Marks Dataset

Marks are whole numbers (0-100), attendance percentages have decimals.

marks = np.array([
    [85, 88, 92],
    [90, 76, 85],
    [78, 92, 88],
    [92, 85, 79],
    [88, 90, 94]
])

print("Original dtype:", marks.dtype)  # Likely int64
print("Memory:", marks.nbytes, "bytes")  # 15 elements * 8 bytes = 120
Original dtype: int64
Memory: 120 bytes

Downcast to Smaller Integer

# Safe downcast to uint8 (unsigned 0-255)
marks_uint8 = marks.astype(np.uint8)
print("New dtype:", marks_uint8.dtype)
print("Memory now:", marks_uint8.nbytes, "bytes")  # 15 bytes!
print("Values unchanged:", np.array_equal(marks, marks_uint8))
New dtype: uint8
Memory now: 15 bytes
Values unchanged: True

8x memory savings — huge for large datasets!

Adding Decimal Data (Attendance)

attendance = np.array([
    [95.5, 90.0, 94.2],
    [87.3, 85.1, 88.8],
    [92.0, 90.5, 85.0],
    [88.8, 86.2, 90.1],
    [96.0, 92.3, 95.5]
])

# Combine — NumPy upcasts to float64
combined = np.hstack((marks_uint8, attendance.astype(np.float32)))
print("Combined dtype:", combined.dtype)
Combined dtype: float64

Unsafe Conversions and Overflow

# Overflow example
large_nums = np.array([200, 300], dtype=np.uint8)
print("Overflow wraps around:", large_nums)  # 200 and 44 (300-256)

# Float to int truncation
decimals = np.array([1.9, -2.3])
negative_int = decimals.astype(np.int8)
print("Truncated:", negative_int)  # [1 -2]
Overflow wraps around: [200  44]
Truncated: [ 1 -2]

Viewing as Different Type (Advanced)

Use view() to reinterpret bytes — no copy, but risky!

float_arr = np.array([1.0, 2.0], dtype=np.float64)
int_view = float_arr.view(np.int64)
print("As integers:", int_view)  # Raw byte representation

Beginner Mistakes - Common Errors and How to Avoid Them

🔴 DON'T: Downcast without checking range — overflow corrupts data silently!
🔴 DON'T: Convert float to int expecting rounding — it truncates toward zero.
🔴 DON'T: Use view() casually — it's for experts and can cause hard-to-debug issues.

Optimization Tips

🟢 DO: Use float32 instead of float64 for ML/models — halves memory with minimal precision loss.
🟢 DO: Check arr.min() and arr.max() before downcasting.

Real-World Use Cases

  • Image Processing: uint8 for pixels (0-255)
  • Machine Learning: float32 for weights
  • IoT/Sensors: Small integers for limited memory
  • Big Data: Downcasting saves gigabytes

Quick Summary 📝

  • Dtypes control precision and memory
  • astype() for safe conversions
  • Downcast when possible, watch for overflow
  • Upcasting happens automatically when needed

Mastering dtypes makes your code faster and leaner ✨

Comments