Think of Hugging Face Datasets like the world's biggest, smartest library of data — but instead of books, it stores text, images, audio, and more, all ready for AI training. The best part? You can load any dataset with just one line of code.
What is the Hugging Face Datasets Library?
Before an AI model can learn anything, it needs data — thousands or millions of examples to study. Finding, downloading, and preparing that data used to take days of painful work.
Hugging Face Datasets solves all of that. It gives you access to over 100,000 datasets from research labs, universities, and companies — all in one place, all free, all loadable in one line.
💡 Think of it like: A magical school library where every textbook ever written is already organised, labelled, and handed to you the moment you ask for it. No searching, no downloading ZIP files, no formatting headaches! 📖
What can you do with it?
- Load → Download any dataset instantly with
load_dataset() - Explore → Inspect examples, check column names, see data types
- Filter → Keep only the rows you need
- Map → Apply transformations to every row at once
- Split → Divide into train, validation, and test sets
- Stream → Work with datasets too large to fit in memory
- Push → Upload your own datasets and share them with the world
Installing the Library
📌 What this code does: Installs the Hugging Face Datasets library on your computer or Colab notebook. Think of it like downloading the app before you can use it — you only need to do this once!
pip install datasets
For working with audio or image datasets, install the extras too:
pip install datasets[audio,vision]
Loading Your First Dataset
Let's load the famous IMDB movie reviews dataset — 50,000 movie reviews labelled as Positive or Negative. This is the "Hello World" of NLP datasets!
Step 1: Load the Dataset
📌 What this code does: Downloads the IMDB dataset from the Hugging Face Hub and loads it into memory. It automatically splits it into a training set and a test set for you. Think of it as the librarian handing you two folders — one for studying, one for the final exam!
from datasets import load_dataset
dataset = load_dataset("imdb")
print(dataset)
Output:
DatasetDict({
train: Dataset({
features: ['text', 'label'],
num_rows: 25000
})
test: Dataset({
features: ['text', 'label'],
num_rows: 25000
})
})
25,000 examples for training and 25,000 for testing — all downloaded automatically! 🎉
Step 2: Explore the Dataset
📌 What this code does: Lets you peek inside the dataset — like opening the folder and reading a few pages of one example. This helps you understand what the data looks like before you start working with it.
# Look at the training split
train_data = dataset["train"]
# See the first example
print(train_data[0])
Output:
{
'text': 'I rented I AM CURIOUS-YELLOW from my video store because of all the
controversy that surrounded it when it was first released in 1967...',
'label': 0
}
What does label 0 or 1 mean?
label = 0→ Negative review 👎label = 1→ Positive review 👍
Step 3: Check Dataset Features
📌 What this code does: Shows you the column names and data types of the dataset — like looking at the header row of a spreadsheet. This tells you exactly what information each example contains.
print(train_data.features)
print(f"Number of rows: {len(train_data)}")
print(f"Column names: {train_data.column_names}")
Output:
{'text': Value(dtype='string', id=None),
'label': ClassLabel(names=['neg', 'pos'], id=None)}
Number of rows: 25000
Column names: ['text', 'label']
See? Two columns — text holds the review, label holds the sentiment. Simple and clean! 🎯
Browsing Multiple Examples at Once
📌 What this code does: Fetches several examples at the same time using a slice (like grabbing rows 0 to 4 from a spreadsheet). Useful for getting a quick feel for what the data looks like across many examples.
# Get the first 3 examples
examples = train_data[:3]
for i in range(3):
label_name = "Positive 👍" if examples['label'][i] == 1 else "Negative 👎"
text_preview = examples['text'][i][:80] # first 80 characters only
print(f"Example {i+1} | {label_name}")
print(f"Text: {text_preview}...")
print()
Output:
Example 1 | Negative 👎
Text: I rented I AM CURIOUS-YELLOW from my video store because of all the controversy...
Example 2 | Negative 👎
Text: "I Am Curious: Yellow" is a risque film mainly for the sex and nudity...
Example 3 | Positive 👍
Text: If only to avoid making this type of film in the future. This film is...
Now you can see real examples from the dataset! 📊
Understanding DatasetDict — The Container
When you call load_dataset(), you get back a DatasetDict. Think of it like a filing cabinet 🗄️:
- The cabinet itself =
DatasetDict(holds everything) - Each drawer = a split like
train,test,validation - Files inside each drawer = the actual rows of data
Most datasets come with these standard splits:
train→ The examples used to teach the modelvalidation→ Used to check progress during training (not all datasets have this)test→ Used only at the very end to measure final performance
Loading Only Part of a Dataset
Sometimes datasets are massive and you just want a small piece to experiment with. Here are three handy tricks:
Trick 1: Load Only One Split
📌 What this code does: Downloads only the training portion of the dataset instead of everything. Saves time and memory when you don't need the test set yet.
train_only = load_dataset("imdb", split="train")
print(type(train_only))
print(len(train_only))
Output:
<class 'datasets.arrow_dataset.Dataset'>
25000
Notice: when you request a specific split, you get a Dataset object directly — not a DatasetDict! 💡
Trick 2: Load a Percentage of the Data
📌 What this code does: Loads only the first 10% of the training data. Perfect when you want to quickly prototype something without waiting to download millions of rows.
small_train = load_dataset("imdb", split="train[:10%]")
print(f"Rows in small sample: {len(small_train)}")
Output:
Rows in small sample: 2500
You can also do split="train[1000:2000]" to grab rows 1000 to 2000 — just like slicing a Python list! 🍕
Trick 3: Load a Specific Number of Rows
📌 What this code does: Loads the dataset in streaming mode and takes only the first 100 rows. This is perfect for huge datasets that are too large to download fully — you stream them like a Netflix video instead of downloading the whole movie first! 🎬
streamed = load_dataset("imdb", split="train", streaming=True)
# Take just the first 100 examples
tiny_sample = list(streamed.take(100))
print(f"Got {len(tiny_sample)} examples")
print(tiny_sample[0]['label'])
Output:
Got 100 examples
0
Filtering Data — Keeping Only What You Need
Just like filtering a spreadsheet, you can keep only the rows that match a condition.
📌 What this code does: Goes through every row in the dataset and keeps only the positive reviews (label = 1). Think of it as a teacher who keeps only the A-grade papers in a separate pile!
# Keep only positive reviews
positive_only = train_data.filter(lambda example: example['label'] == 1)
print(f"Total training rows: {len(train_data)}")
print(f"Positive reviews only: {len(positive_only)}")
Output:
Total training rows: 25000
Positive reviews only: 12500
Exactly half — because IMDB is balanced with 12,500 positive and 12,500 negative reviews! ⚖️
Let's also filter by text length — keep only reviews with more than 200 characters:
📌 What this code does: Removes short reviews from the dataset by checking the length of each review text. Useful when you know very short texts are too noisy or uninformative for training.
long_reviews = train_data.filter(lambda example: len(example['text']) > 200)
print(f"Original size: {len(train_data)}")
print(f"Long reviews only: {len(long_reviews)}")
Output:
Original size: 25000
Long reviews only: 24892
The map() Function — Transform Every Row
The map() function is one of the most powerful tools in the Datasets library. It applies a function to every single row in the dataset — automatically, efficiently, and in parallel.
💡 Think of it like: A factory conveyor belt 🏭. Each row goes in, your function does something to it, and the changed row comes out the other end. You never write a for-loop!
Example 1: Add a New Column — Text Length
📌 What this code does: Adds a brand new column called text_length to every row, containing the number of characters in the review. This is useful for analysis — for example, checking whether longer reviews tend to be more positive or negative.
def add_text_length(example):
example['text_length'] = len(example['text'])
return example
dataset_with_length = train_data.map(add_text_length)
print(dataset_with_length.column_names)
print(f"First review length: {dataset_with_length[0]['text_length']} characters")
Output:
['text', 'label', 'text_length']
First review length: 1967 characters
A new column appeared! And we never wrote a single for-loop. 🎉
Example 2: Rename Labels — Make Them Human-Friendly
📌 What this code does: Converts the numeric labels (0 and 1) into human-readable text ("negative" and "positive"). This makes it much easier to understand what you are looking at when exploring the data.
def label_to_text(example):
example['sentiment'] = "positive" if example['label'] == 1 else "negative"
return example
labelled = train_data.map(label_to_text)
# Check the first few examples
for i in range(3):
print(f"Label: {labelled[i]['label']} | Sentiment: {labelled[i]['sentiment']}")
Output:
Label: 0 | Sentiment: negative
Label: 0 | Sentiment: negative
Label: 1 | Sentiment: positive
Example 3: Tokenize the Whole Dataset in One Go
📌 What this code does: Runs every review through a tokenizer — converting raw text into numbers the AI model can understand. This is the most common real-world use of map(). The batched=True option makes it run much faster by processing many rows at once instead of one at a time.
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")
def tokenize(examples):
return tokenizer(
examples["text"],
max_length=128,
truncation=True,
padding="max_length"
)
tokenized = train_data.map(tokenize, batched=True)
print(tokenized.column_names)
print(f"Input IDs for first example: {tokenized[0]['input_ids'][:10]}...")
Output:
['text', 'label', 'input_ids', 'attention_mask', 'token_type_ids']
Input IDs for first example: [101, 1045, 7422, 1045, 2572, 8025, 1011, 3756, 2013, 2026]...
Three new columns appeared: input_ids, attention_mask, and token_type_ids — exactly what BERT needs to learn! 🤖
Selecting and Removing Columns
After tokenizing, the raw text column is no longer needed. Let's clean up:
📌 What this code does: Removes the original text column and keeps only the columns the model actually needs for training. Like throwing away the rough draft once you've typed up the final version!
# Remove columns the model doesn't need
clean = tokenized.remove_columns(["text"])
print(clean.column_names)
Output:
['label', 'input_ids', 'attention_mask', 'token_type_ids']
Or if you want to keep only specific columns:
slim = tokenized.select_columns(["input_ids", "attention_mask", "label"])
print(slim.column_names)
Output:
['input_ids', 'attention_mask', 'label']
Renaming Columns
Some models expect the label column to be called labels (plural) instead of label. Easy fix:
📌 What this code does: Renames the label column to labels. This matters because the Hugging Face Trainer looks for a column specifically named labels when calculating the loss during training.
renamed = slim.rename_column("label", "labels")
print(renamed.column_names)
Output:
['input_ids', 'attention_mask', 'labels']
Shuffling and Splitting Your Own Dataset
Shuffling — Mix Up the Order
📌 What this code does: Randomly shuffles the order of all rows in the dataset. This is important because if all the negative reviews are grouped at the top and positives at the bottom, the model will see a very biased order during training and learn poorly.
shuffled = train_data.shuffle(seed=42)
# Confirm labels are now mixed up
for i in range(5):
print(f"Row {i}: label = {shuffled[i]['label']}")
Output:
Row 0: label = 1
Row 1: label = 0
Row 2: label = 1
Row 3: label = 1
Row 4: label = 0
A nice mix of 0s and 1s! The seed=42 makes sure you get the same shuffle every time you run it — great for reproducibility. 🔀
Splitting — Create Your Own Train / Validation Split
📌 What this code does: Takes one big dataset and splits it into two parts — 90% for training and 10% for validation. Like dividing your study notes into "notes I'll study from" and "notes I'll use to test myself" before the exam!
split = shuffled.train_test_split(test_size=0.1, seed=42)
print(split)
print(f"Training examples: {len(split['train'])}")
print(f"Validation examples: {len(split['test'])}")
Output:
DatasetDict({
train: Dataset({
features: ['text', 'label'],
num_rows: 22500
})
test: Dataset({
features: ['text', 'label'],
num_rows: 2500
})
})
Training examples: 22500
Validation examples: 2500
Selecting Specific Rows
📌 What this code does: Picks out specific row numbers from the dataset by their index — just like choosing specific pages from a book. You could use this to create a tiny test sample or grab specific examples you want to examine.
# Select rows 0, 5, 10, 15, and 20
sample = train_data.select([0, 5, 10, 15, 20])
print(f"Selected {len(sample)} rows")
for i in range(len(sample)):
print(f"Row: label={sample[i]['label']}, text starts: {sample[i]['text'][:50]}...")
Output:
Selected 5 rows
Row: label=0, text starts: I rented I AM CURIOUS-YELLOW from my video store...
Row: label=0, text starts: This film is a good film. I enjoyed it. It is...
Row: label=1, text starts: Encouraged by the positive comments about this film...
Row: label=0, text starts: If you like original gut wrenching laughter then...
Row: label=1, text starts: Phil the Alien is one of those quirky films where...
Sorting the Dataset
📌 What this code does: Sorts all rows in the dataset by the value of a column — just like sorting a spreadsheet. Here we sort by label so all the negative reviews come first, then all the positive ones.
sorted_data = train_data.sort("label")
print("First 3 labels (should all be 0 = negative):")
print([sorted_data[i]['label'] for i in range(3)])
print("Last 3 labels (should all be 1 = positive):")
print([sorted_data[-(i+1)]['label'] for i in range(3)])
Output:
First 3 labels (should all be 0 = negative):
[0, 0, 0]
Last 3 labels (should all be 1 = positive):
[1, 1, 1]
Streaming — For Huge Datasets 🌊
Some datasets are so large they cannot fit in your computer's memory. The Hugging Face Datasets library handles this with streaming mode.
💡 Think of it like: A water tap 🚰. You don't fill the entire ocean into your house — you just turn on the tap and use water as you need it!
📌 What this code does: Loads a very large dataset (The Pile — used to train GPT-style models) without downloading it all. It streams examples one at a time, so you can start working with the data in seconds even if the full dataset is hundreds of gigabytes.
from datasets import load_dataset
# Stream a massive dataset — doesn't download everything at once!
streamed_dataset = load_dataset(
"EleutherAI/pile",
split="train",
streaming=True
)
# Iterate over the first 5 examples only
for i, example in enumerate(streamed_dataset):
print(f"Example {i+1}: {example['text'][:80]}...")
if i == 4:
break
How streaming compares to normal loading:
- Normal loading → Downloads everything first, then you can work with it. Great for small datasets
- Streaming → Downloads one batch at a time as you need it. Essential for giant datasets like The Pile, Common Crawl, or LAION
You can still use filter() and map() on streamed datasets — they just work lazily, processing data as it arrives! 🎯
Saving and Loading Datasets Locally
Once you have cleaned, tokenized, and prepared your dataset, save it locally so you don't have to redo all that work next time!
Save to Disk
📌 What this code does: Saves the fully processed dataset to a folder on your hard drive. Next time you run your code, you load it directly from disk instead of re-downloading and re-processing everything. Huge time saver! ⏱️
# Save the tokenized dataset to disk
tokenized.save_to_disk("./my_imdb_tokenized")
print("Dataset saved! ✅")
Output:
Dataset saved! ✅
Load from Disk
📌 What this code does: Loads the dataset back from the folder you saved it to. This is instant — no downloading, no tokenizing, no waiting. Just load and go!
from datasets import load_from_disk
loaded = load_from_disk("./my_imdb_tokenized")
print(loaded)
print(f"Loaded {len(loaded)} rows from disk!")
Output:
Dataset({
features: ['text', 'label', 'input_ids', 'attention_mask', 'token_type_ids'],
num_rows: 25000
})
Loaded 25000 rows from disk!
Converting to Pandas — For Exploration
Sometimes it's easier to explore data in a Pandas DataFrame. Hugging Face Datasets makes this a one-liner:
📌 What this code does: Converts the Hugging Face Dataset into a familiar Pandas DataFrame so you can use all the usual Pandas tools — describe(), value_counts(), groupby(), and more!
import pandas as pd
df = train_data.to_pandas()
print(df.shape)
print(df.head())
print(df['label'].value_counts())
Output:
(25000, 2)
text label
0 I rented I AM CURIOUS-YELLOW from my video st... 0
1 "I Am Curious: Yellow" is a risque film mainly... 0
2 If only to avoid making this type of film in t... 0
3 This film was probably inspired by Godard's Ma... 0
4 Oh, brother...after hearing about this ridicul... 0
label
0 12500
1 12500
Name: label, dtype: int64
Perfectly balanced — 12,500 negative and 12,500 positive reviews! And going back from Pandas to a Dataset is just as easy:
from datasets import Dataset
back_to_dataset = Dataset.from_pandas(df)
print(back_to_dataset)
Loading Different Types of Datasets
The Datasets library handles much more than just text. Here are the most common formats:
Loading from a CSV File
📌 What this code does: Loads data directly from a CSV file on your computer — just like pd.read_csv() in Pandas, but the result is a Hugging Face Dataset ready for AI training.
csv_dataset = load_dataset("csv", data_files="my_data.csv")
print(csv_dataset)
Loading from a JSON File
📌 What this code does: Loads data from a JSON or JSONL (JSON Lines) file. JSONL is very common for NLP datasets where each line is one example. Perfect for loading your own custom datasets!
json_dataset = load_dataset("json", data_files="my_data.jsonl")
print(json_dataset)
Loading from Multiple Files at Once
📌 What this code does: Loads different files as different splits — train from one file, test from another. Useful when your data is already split into separate files on disk.
split_dataset = load_dataset(
"csv",
data_files={
"train": "train_data.csv",
"test": "test_data.csv"
}
)
print(split_dataset)
Output:
DatasetDict({
train: Dataset({features: [...], num_rows: ...})
test: Dataset({features: [...], num_rows: ...})
})
Creating Your Own Dataset from Scratch
You don't always need to download a dataset — sometimes you want to create one from data you already have in Python.
📌 What this code does: Creates a brand new Hugging Face Dataset from a plain Python dictionary — just like creating a Pandas DataFrame from a dictionary. Perfect for wrapping your own data so it works with all the Hugging Face training tools.
from datasets import Dataset
# Your own data as a Python dictionary
my_data = {
"question": [
"What is the capital of France?",
"Who invented the telephone?",
"What is 2 + 2?",
"Which planet is closest to the Sun?"
],
"answer": [
"Paris",
"Alexander Graham Bell",
"4",
"Mercury"
],
"category": [
"geography",
"history",
"math",
"science"
]
}
my_dataset = Dataset.from_dict(my_data)
print(my_dataset)
print(my_dataset[0])
Output:
Dataset({
features: ['question', 'answer', 'category'],
num_rows: 4
})
{'question': 'What is the capital of France?',
'answer': 'Paris',
'category': 'geography'}
Now your custom data has all the powers of the Datasets library — map(), filter(), save_to_disk(), everything! 🎉
Searching and Finding Datasets on the Hub
There are over 100,000 datasets on the Hugging Face Hub. You can search and discover them directly from code:
📌 What this code does: Searches the Hugging Face Hub for datasets related to a keyword — here "sentiment". Returns a list of matching datasets you can then load with load_dataset().
from huggingface_hub import list_datasets
# Search for sentiment datasets
results = list(list_datasets(search="sentiment", limit=5))
for dataset_info in results:
print(f"Name: {dataset_info.id}")
print(f"Downloads last month: {dataset_info.downloads}")
print()
Output:
Name: stanfordnlp/sst2
Downloads last month: 142301
Name: cardiffnlp/tweet_sentiment_multilingual
Downloads last month: 89234
Name: mteb/amazon_reviews_multi
Downloads last month: 76112
...
Pushing Your Own Dataset to the Hub
Once you have created or cleaned a dataset, you can share it with the world on the Hugging Face Hub — completely free!
Step 1: Log In
📌 What this code does: Logs you into your Hugging Face account. You need to create a free account at huggingface.co and generate an access token from your profile settings.
from huggingface_hub import login
login() # will prompt you to paste your access token
Step 2: Push the Dataset
📌 What this code does: Uploads your dataset to the Hugging Face Hub under your username. After this runs, anyone in the world can load your dataset with load_dataset("your-username/your-dataset-name") — just like you did with IMDB!
my_dataset.push_to_hub("your-username/my-quiz-dataset")
print("Dataset published! 🚀")
print("Anyone can now load it with:")
print('load_dataset("your-username/my-quiz-dataset")')
A Real End-to-End Example — From Load to Ready-for-Training
Let's put everything together! Here is a complete pipeline that loads, inspects, tokenizes, cleans, and saves a dataset — ready to be fed into a Hugging Face Trainer.
Step 1: Load
📌 What this code does: Loads the full IMDB dataset — both train and test splits — in one go.
from datasets import load_dataset
dataset = load_dataset("imdb")
print(f"Train: {len(dataset['train'])} rows")
print(f"Test: {len(dataset['test'])} rows")
Output:
Train: 25000 rows
Test: 25000 rows
Step 2: Shuffle and Create a Validation Split
📌 What this code does: Shuffles the training data and carves off 10% of it to use as a validation set. Now we have three clean splits: train, validation, and test.
shuffled_train = dataset["train"].shuffle(seed=42)
train_val = shuffled_train.train_test_split(test_size=0.1, seed=42)
print(f"Train: {len(train_val['train'])} rows")
print(f"Validation: {len(train_val['test'])} rows")
print(f"Test: {len(dataset['test'])} rows")
Output:
Train: 22500 rows
Validation: 2500 rows
Test: 25000 rows
Step 3: Tokenize Everything
📌 What this code does: Runs all three splits through the tokenizer at once using map() with batched=True. The batched=True flag processes 1,000 rows at a time instead of one at a time — making it roughly 10× faster!
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")
def tokenize_batch(examples):
return tokenizer(
examples["text"],
max_length=128,
truncation=True,
padding="max_length"
)
tokenized_train = train_val["train"].map(tokenize_batch, batched=True)
tokenized_val = train_val["test"].map(tokenize_batch, batched=True)
tokenized_test = dataset["test"].map(tokenize_batch, batched=True)
print("All splits tokenized! ✅")
Step 4: Clean Up Columns and Rename Labels
📌 What this code does: Removes the raw text column (no longer needed after tokenizing) and renames label to labels so the Hugging Face Trainer can find it automatically during training.
def clean_split(ds):
ds = ds.remove_columns(["text"])
ds = ds.rename_column("label", "labels")
ds.set_format("torch") # convert to PyTorch tensors automatically
return ds
train_ready = clean_split(tokenized_train)
val_ready = clean_split(tokenized_val)
test_ready = clean_split(tokenized_test)
print(train_ready.column_names)
print(f"Ready to train on {len(train_ready)} examples!")
Output:
['labels', 'input_ids', 'attention_mask', 'token_type_ids']
Ready to train on 22500 examples!
Step 5: Save Everything to Disk
📌 What this code does: Saves all three prepared splits to disk so you never have to repeat all that tokenizing work again. Next time, just call load_from_disk() and you are immediately ready to train!
train_ready.save_to_disk("./prepared/train")
val_ready.save_to_disk("./prepared/val")
test_ready.save_to_disk("./prepared/test")
print("All splits saved! ✅")
print("Next time, just load from disk — no re-processing needed!")
The DataCollator — The Final Piece
When training, your model receives data in batches. A DataCollator is a helper that takes a list of individual examples and neatly stacks them into one batch.
💡 Think of it like: A waiter in a restaurant 🍽️. Each customer orders separately, but the waiter picks up all the dishes at once from the kitchen and delivers them together on a tray. The DataCollator is that tray!
📌 What this code does: Creates a DataCollatorWithPadding object that automatically pads shorter sequences in a batch to match the longest one. This is more memory-efficient than padding everything to the maximum length upfront.
from transformers import DataCollatorWithPadding
data_collator = DataCollatorWithPadding(tokenizer=tokenizer)
# Pass it to the Trainer
from transformers import Trainer, TrainingArguments
trainer = Trainer(
model=model,
args=TrainingArguments(output_dir="./output", num_train_epochs=3),
train_dataset=train_ready,
eval_dataset=val_ready,
data_collator=data_collator,
)
print("Trainer is ready to go! 🚀")
Popular Datasets to Know
Here are the most widely used datasets you will encounter in real AI projects:
Text Classification
- imdb → Movie review sentiment (Positive / Negative)
- stanfordnlp/sst2 → Short phrase sentiment from movie reviews
- ag_news → News article topic classification (4 categories)
Question Answering
- rajpurkar/squad → Reading comprehension — find the answer in a paragraph
- google-research-datasets/natural_questions → Real Google search questions with answers
Language Modelling / LLM Training
- HuggingFaceFW/fineweb → High-quality web text, used to train modern LLMs (2024–2026)
- allenai/dolma → Open LLM pre-training dataset from Allen AI
- open-thoughts/OpenThoughts-114k → Chain-of-thought reasoning data (popular in 2025–2026)
Instruction Tuning / Chat
- HuggingFaceH4/ultrachat_200k → 200,000 multi-turn chat conversations
- teknium/OpenHermes-2.5 → High-quality instruction-following dataset
- argilla/magpie-ultra-v0.1 → Synthetic instruction data (2025–2026 trend)
Multimodal (Text + Images)
- HuggingFaceM4/OBELICS → Interleaved image and text for vision-language models
- laion/laion-400m → 400 million image-caption pairs
Common Mistakes — What NOT to Do 🚫
🚫 Don't forget to shuffle before splitting. If you split without shuffling, your validation set might contain all the examples from one category — completely useless for evaluation!
🚫 Don't tokenize without batched=True on large datasets. Without batching, map() processes one row at a time. With batched=True, it processes 1,000 at once — dramatically faster.
🚫 Don't load massive datasets without streaming. If a dataset is 500 GB and you try to load it normally, your computer will crash or freeze. Always use streaming=True for datasets larger than your RAM.
🚫 Don't skip saving processed datasets to disk. Tokenizing 50,000 examples takes time. If you don't save, you have to redo it every single run. Use save_to_disk() after processing!
🚫 Don't forget to call set_format("torch") before training. Without this, your Dataset returns Python lists. PyTorch needs tensors. One line saves you a confusing error!
Good Practices — What TO Do ✅
✅ Always explore before processing. Print a few examples, check features, check value_counts in Pandas. Understand your data before throwing it at a model!
✅ Use batched=True in map(). It is almost always faster. For tokenization especially, it can be 10–30× quicker than row-by-row processing.
✅ Save processed datasets to disk. Always call save_to_disk() after the full preparation pipeline. Load with load_from_disk() in future runs.
✅ Use seed=42 in shuffle() and train_test_split(). This makes your data splits reproducible — you and your teammates always get the exact same train/validation split.
✅ Check for class imbalance with value_counts(). If 90% of your labels are class 0 and 10% are class 1, you have an imbalanced dataset. Handle this before training or your model will just predict class 0 for everything!
Quick Summary 📝
What we learned :
- Loading datasets →
load_dataset("name")downloads any of 100,000+ datasets instantly - Exploring → Check
features,column_names, slice withdataset[0]ordataset[:5] - Filtering → Keep rows that match a condition with
dataset.filter(lambda x: ...) - Transforming → Apply changes to every row with
dataset.map(function, batched=True) - Splitting → Create train / validation splits with
train_test_split(test_size=0.1) - Streaming → Work with giant datasets without downloading them using
streaming=True - Saving → Save processed data with
save_to_disk(), reload withload_from_disk() - Own data → Wrap any Python dict with
Dataset.from_dict()or load CSVs and JSONs directly - Sharing → Publish your dataset to the world with
push_to_hub()
Happy learning! 🤗✨
Comments
Post a Comment