Pandas makes it super easy to read JSON with the pd.read_json() function. This guide explains everything step-by-step with plenty of simple examples, so even complete beginners can follow along and start using it right away.
We’ll use the same sample datasets throughout to make it easy to compare.
What Does JSON Look Like?
JSON data is made of key-value pairs, arrays, and nested objects. Here are a few simple examples we'll use:
Sample 1: Flat List of Records (Most Common)
[
{"Name": "Alice", "Age": 25, "City": "New York"},
{"Name": "Bob", "Age": 30, "City": "London"},
{"Name": "Charlie", "Age": 35, "City": "Tokyo"},
{"Name": "David", "Age": 28, "City": "Paris"}
]
Sample 2: Nested JSON (Real-World Example)
[
{
"Name": "Alice",
"Age": 25,
"Address": {"Street": "123 Main St", "City": "New York", "Zip": "10001"}
},
{
"Name": "Bob",
"Age": 30,
"Address": {"Street": "456 Elm St", "City": "London", "Zip": "SW1A"}
}
]
Sample 3: JSON Lines Format (One JSON Object Per Line)
{"Name": "Alice", "Age": 25, "Salary": 50000}
{"Name": "Bob", "Age": 30, "Salary": 60000}
{"Name": "Charlie", "Age": 35, "Salary": 70000}
1. Reading JSON from a String
If your JSON data is already in Python as a string (e.g., from an API response), you can read it directly.
import pandas as pd
import json # Optional, for pretty printing
json_string = '''
[
{"Name": "Alice", "Age": 25, "City": "New York"},
{"Name": "Bob", "Age": 30, "City": "London"},
{"Name": "Charlie", "Age": 35, "City": "Tokyo"}
]
'''
df = pd.read_json(json_string)
print(df)
Output:
| Name | Age | City |
|---|---|---|
| Alice | 25 | New York |
| Bob | 30 | London |
| Charlie | 35 | Tokyo |
2. Reading JSON from a File
Save the JSON data to a file (e.g., data.json) and read it.
df = pd.read_json('data.json')
print(df)
Gives the same DataFrame as above.
3. Different JSON Orientations
JSON can be structured in many ways. Use the orient parameter to tell Pandas how your data is organized.
orient='records' (Default for list of objects)
df = pd.read_json(json_string, orient='records')
orient='index'
json_index = '''
{
"0": {"Name": "Alice", "Age": 25},
"1": {"Name": "Bob", "Age": 30}
}
'''
df = pd.read_json(json_index, orient='index')
print(df)
Rows become the index.
orient='columns'
json_columns = '''
{
"Name": ["Alice", "Bob"],
"Age": [25, 30]
}
'''
df = pd.read_json(json_columns, orient='columns')
print(df)
4. Reading Nested JSON (Using json_normalize)
Flat JSON is easy, but real JSON is often nested. Use pd.json_normalize() to flatten it.
import json
nested_json = '''
[
{
"Name": "Alice",
"Age": 25,
"Address": {"Street": "123 Main St", "City": "New York", "Zip": "10001"}
},
{
"Name": "Bob",
"Age": 30,
"Address": {"Street": "456 Elm St", "City": "London", "Zip": "SW1A"}
}
]
'''
data = json.loads(nested_json)
df_nested = pd.json_normalize(data)
print(df_nested)
Output:
| Name | Age | Address.Street | Address.City | Address.Zip |
|---|---|---|---|---|
| Alice | 25 | 123 Main St | New York | 10001 |
| Bob | 30 | 456 Elm St | London | SW1A |
5. Reading JSON Lines (JSONL) Files
Large datasets are often stored as JSON Lines (one JSON object per line).
df = pd.read_json('data.jsonl', lines=True)
print(df)
Use lines=True – it's much faster and memory-efficient for big files.
Common Parameters You’ll Use
orient: 'records', 'index', 'columns', 'values', 'split'lines=True: For JSONL filesdtype: Specify data types (e.g.,dtype={'Age': int})convert_dates: Automatically parse date columns
Happy Learning !! 🐼
Comments
Post a Comment