What is a Pandas Series?
A Series is a one-dimensional labeled array in Pandas that can hold data of any type (integers, floats, strings, Python objects, etc.). Think of it as:
- A single column in an Excel sheet
- A NumPy array with labels (index) attached to each value
- A labeled vector
The two main components are:
- Values: the actual data
- Index: labels for each value (by default, 0, 1, 2, …)
Series are the building blocks of the more powerful DataFrame (which is essentially a collection of Series).
Importing Pandas
Always start with:
import pandas as pd
import numpy as np # Optional, but useful for some examples
1. Creating a Basic Series from a List
The simplest way:
data = [10, 20, 30, 40, 50]
series = pd.Series(data)
print(series)
Output:
0 10
1 20
2 30
3 40
4 50
dtype: int64
Here the index is automatically created as integers starting from 0.
2. Creating a Series with Custom Index
We can provide our own index labels:
data = [10, 20, 30, 40, 50]
custom_index = [1, 2, 3, 4, 5]
series_with_index = pd.Series(data, index=custom_index)
print(series_with_index)
Output:
1 10
2 20
3 30
4 40
5 50
dtype: int64
Accessing Elements by Label
print(series_with_index[1]) # Access by index label
Output: 10
Slicing (Important Note on Label vs Position)
print(series_with_index[1:3]) # This slices by LABEL (1 and 2)
Output:
1 10
2 20
dtype: int64
Tip: If you want to slice by position (like regular Python lists), use .iloc:
print(series_with_index.iloc[1:3]) # Positions 1 and 2 → values 20 and 30
Output:
2 20
3 30
dtype: int64
3. Creating a Series from a Dictionary
Dictionaries are perfect for labeled data — keys become the index:
data_dict = {'a': 100, 'b': 200, 'c': 300}
series_from_dict = pd.Series(data_dict)
print(series_from_dict)
Output:
a 100
b 200
c 300
dtype: int64
Accessing:
print(series_from_dict['a']) # 100
Bonus: Other Useful Ways and Operations
- From a scalar value (broadcasts to all indices):
pd.Series(5, index=['x', 'y', 'z'])
Output: x 5, y 5, z 5
- Arithmetic operations align by index automatically (great for real-world data):
s1 = pd.Series([1, 2, 3], index=['a', 'b', 'c'])
s2 = pd.Series([4, 5, 6], index=['b', 'c', 'd'])
print(s1 + s2)
Missing values become NaN
- Common methods:
.head(),.tail(),.describe(),.unique(),.value_counts()
df.head(2) # Shows first 2 rows
df.tail(2) # Shows last 2 rows
df.info() # Tells you about columns and data types
df.describe() # Quick stats (works best on numbers)
Conclusion
The Pandas Series is a powerful and flexible way to handle one-dimensional labeled data. The custom indexing makes it much more useful than plain lists or NumPy arrays when working with real-world datasets.
Happy coding!
— Ritesh
Comments
Post a Comment