Skip to main content

Pandas: Reading JSON Data

Calculating read time…

Pandas makes it super easy to read JSON with the pd.read_json() function. This guide explains everything step-by-step with plenty of simple examples, so even complete beginners can follow along and start using it right away.

We’ll use the same sample datasets throughout to make it easy to compare.

What Does JSON Look Like?

JSON data is made of key-value pairs, arrays, and nested objects. Here are a few simple examples we'll use:

Sample 1: Flat List of Records (Most Common)

[
    {"Name": "Alice", "Age": 25, "City": "New York"},
    {"Name": "Bob", "Age": 30, "City": "London"},
    {"Name": "Charlie", "Age": 35, "City": "Tokyo"},
    {"Name": "David", "Age": 28, "City": "Paris"}
]

Sample 2: Nested JSON (Real-World Example)

[
    {
        "Name": "Alice",
        "Age": 25,
        "Address": {"Street": "123 Main St", "City": "New York", "Zip": "10001"}
    },
    {
        "Name": "Bob",
        "Age": 30,
        "Address": {"Street": "456 Elm St", "City": "London", "Zip": "SW1A"}
    }
]

Sample 3: JSON Lines Format (One JSON Object Per Line)

{"Name": "Alice", "Age": 25, "Salary": 50000}
{"Name": "Bob", "Age": 30, "Salary": 60000}
{"Name": "Charlie", "Age": 35, "Salary": 70000}

1. Reading JSON from a String

If your JSON data is already in Python as a string (e.g., from an API response), you can read it directly.

import pandas as pd
import json   # Optional, for pretty printing

json_string = '''
[
    {"Name": "Alice", "Age": 25, "City": "New York"},
    {"Name": "Bob", "Age": 30, "City": "London"},
    {"Name": "Charlie", "Age": 35, "City": "Tokyo"}
]
'''

df = pd.read_json(json_string)
print(df)

Output:

Name Age City
Alice 25 New York
Bob 30 London
Charlie 35 Tokyo

2. Reading JSON from a File

Save the JSON data to a file (e.g., data.json) and read it.

df = pd.read_json('data.json')
print(df)

Gives the same DataFrame as above.

3. Different JSON Orientations

JSON can be structured in many ways. Use the orient parameter to tell Pandas how your data is organized.

orient='records' (Default for list of objects)

df = pd.read_json(json_string, orient='records')

orient='index'

json_index = '''
{
    "0": {"Name": "Alice", "Age": 25},
    "1": {"Name": "Bob", "Age": 30}
}
'''

df = pd.read_json(json_index, orient='index')
print(df)

Rows become the index.

orient='columns'

json_columns = '''
{
    "Name": ["Alice", "Bob"],
    "Age": [25, 30]
}
'''

df = pd.read_json(json_columns, orient='columns')
print(df)

4. Reading Nested JSON (Using json_normalize)

Flat JSON is easy, but real JSON is often nested. Use pd.json_normalize() to flatten it.

import json

nested_json = '''
[
    {
        "Name": "Alice",
        "Age": 25,
        "Address": {"Street": "123 Main St", "City": "New York", "Zip": "10001"}
    },
    {
        "Name": "Bob",
        "Age": 30,
        "Address": {"Street": "456 Elm St", "City": "London", "Zip": "SW1A"}
    }
]
'''

data = json.loads(nested_json)
df_nested = pd.json_normalize(data)
print(df_nested)

Output:

Name Age Address.Street Address.City Address.Zip
Alice 25 123 Main St New York 10001
Bob 30 456 Elm St London SW1A

5. Reading JSON Lines (JSONL) Files

Large datasets are often stored as JSON Lines (one JSON object per line).

df = pd.read_json('data.jsonl', lines=True)
print(df)

Use lines=True – it's much faster and memory-efficient for big files.

Common Parameters You’ll Use

  • orient: 'records', 'index', 'columns', 'values', 'split'
  • lines=True: For JSONL files
  • dtype: Specify data types (e.g., dtype={'Age': int})
  • convert_dates: Automatically parse date columns

Happy Learning !! 🐼

Comments