Skip to main content

Tokenization Methods

Calculating read time…

Before an AI model can read a single sentence, it has to convert words into numbers. That conversion process is called tokenization. It is one of the most important — and most misunderstood — steps in all of NLP.




What is Tokenization?

Computers don't understand words. They only understand numbers. So before we feed any text to an AI model, we need to break the text into small pieces and give each piece a number.

Those small pieces are called tokens. The tool that does this job is called a tokenizer.

💡 Think of it like: Imagine you are sending a long message to a friend using LEGO bricks 🧱. You can't send a full sentence — so you break it into individual bricks (tokens), put a number sticker on each brick, and send the numbers. Your friend uses the same numbering system to reassemble the original message!

The Full Tokenization Pipeline

Every tokenizer does these three things, in order:

  • Step 1 — Split: Break the raw text into tokens (words, subwords, or characters)
  • Step 2 — Map to IDs: Look up each token in a vocabulary table and replace it with its unique number (called a token ID)
  • Step 3 — Add Special Tokens: Add marker tokens like [CLS] and [SEP] that tell the model where a sentence starts and ends

Let's see all three steps in action with real code before we go deeper.

📌 What this code does: Loads a BERT tokenizer from Hugging Face and converts a plain English sentence into token IDs step by step — showing you exactly what happens at each stage of the pipeline. Think of it as X-ray vision into how a sentence gets transformed before the model ever sees it!

from transformers import AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")

sentence = "Hugging Face is amazing!"

# Step 1: Split into tokens
tokens = tokenizer.tokenize(sentence)
print("Tokens:", tokens)

# Step 2: Map tokens to their IDs
token_ids = tokenizer.convert_tokens_to_ids(tokens)
print("Token IDs:", token_ids)

# Step 3: Full encoding (adds special tokens automatically)
encoded = tokenizer(sentence)
print("Full encoding:", encoded)

Output:

Tokens: ['hugging', 'face', 'is', 'amazing', '!']

Token IDs: [17662, 2227, 2003, 6429, 999]

Full encoding: {
  'input_ids': [101, 17662, 2227, 2003, 6429, 999, 102],
  'token_type_ids': [0, 0, 0, 0, 0, 0, 0],
  'attention_mask': [1, 1, 1, 1, 1, 1, 1]
}

Notice that 101 appeared at the start and 102 at the end — those are the special [CLS] and [SEP] tokens BERT added automatically! 🎯

The Three Main Tokenization Methods

There are three fundamental ways to split text into tokens. Each has its own strengths, weaknesses, and best use cases. Let's meet all three!

  • Word-level → Split text into whole words: "I love pizza" → ["I", "love", "pizza"]
  • Character-level → Split into individual characters: "cat" → ["c", "a", "t"]
  • Subword-level → Split into meaningful word pieces: "tokenization" → ["token", "ization"]

Subword tokenization is the standard for almost every modern AI model. But understanding all three helps you make smart choices. Let's go through each one!

Method 1 — Word-Level Tokenization

The simplest idea: split text on spaces and punctuation. Every unique word becomes one token.

💡 Think of it like: Cutting a sentence with scissors wherever there is a space ✂️. Each piece of paper is one token.

How it works

📌 What this code does: Builds a simple word-level tokenizer from scratch using plain Python — no libraries needed. This shows you the raw idea behind tokenization before any fancy algorithms get involved. It splits the sentence by spaces and maps each unique word to a number.

# A simple word-level tokenizer built from scratch
sentences = [
    "I love cats",
    "I love dogs",
    "cats and dogs are friends"
]

# Step 1: Build the vocabulary (all unique words)
all_words = []
for sentence in sentences:
    all_words.extend(sentence.lower().split())

vocabulary = sorted(set(all_words))
word_to_id = {word: idx for idx, word in enumerate(vocabulary)}
id_to_word = {idx: word for word, idx in word_to_id.items()}

print("Vocabulary:", word_to_id)
print()

# Step 2: Tokenize a new sentence
test_sentence = "I love cats and dogs"
tokens = test_sentence.lower().split()
token_ids = [word_to_id[word] for word in tokens]

print("Tokens:   ", tokens)
print("Token IDs:", token_ids)

Output:

Vocabulary: {'and': 0, 'are': 1, 'cats': 2, 'dogs': 3,
             'friends': 4, 'i': 5, 'love': 6}

Tokens:    ['i', 'love', 'cats', 'and', 'dogs']
Token IDs: [5, 6, 2, 0, 3]

The big problem — Unknown Words

Word-level tokenization breaks completely when it sees a word it has never seen before. This is called the Out-Of-Vocabulary (OOV) problem.

📌 What this code does: Shows what happens when a word-level tokenizer meets a word that isn't in its vocabulary. This is the classic failure mode — and the reason word-level tokenization is rarely used in modern AI.

# Try to tokenize a word the vocabulary has never seen
new_sentence = "I love elephants"   # "elephants" is not in our vocabulary!

tokens = new_sentence.lower().split()

for word in tokens:
    if word in word_to_id:
        print(f"'{word}' → ID {word_to_id[word]}")
    else:
        print(f"'{word}' → ??? UNKNOWN WORD! Cannot tokenize! 😱")

Output:

'i'         → ID 5
'love'      → ID 6
'elephants' → ??? UNKNOWN WORD! Cannot tokenize! 😱

Word-level problems at a glance:

  • 🚫 Cannot handle new or rare words
  • 🚫 Vocabulary can grow to millions of entries (one per unique word)
  • 🚫 "run", "running", "ran", "runs" are all treated as completely different tokens — no relationship!
  • 🚫 Works poorly for languages like German where one word can mean an entire English sentence

Method 2 — Character-Level Tokenization

Instead of splitting on words, split on individual letters. Every character in the alphabet is a token.

💡 Think of it like: Instead of posting a full letter to a friend, you send it one character at a time — "H", "e", "l", "l", "o". Tiny pieces, but nothing is ever unknown! 📮

How it works

📌 What this code does: Shows character-level tokenization from scratch. The vocabulary is tiny — just the 26 letters plus a few symbols — and it can tokenize absolutely any word, even made-up ones, because every word is just a sequence of characters.

# Character-level tokenizer from scratch
import string

# Vocabulary = all lowercase letters + digits + space + common punctuation
chars = list(string.ascii_lowercase + string.digits + " !?,.")
char_to_id = {ch: idx for idx, ch in enumerate(chars)}
id_to_char = {idx: ch for ch, idx in char_to_id.items()}

print(f"Vocabulary size: {len(char_to_id)} characters")
print(f"First 10 chars: {chars[:10]}")
print()

# Tokenize any sentence — even weird words work!
sentence = "hello ai world 2026!"

tokens   = list(sentence)
token_ids = [char_to_id[ch] for ch in tokens if ch in char_to_id]

print("Tokens:   ", tokens)
print("Token IDs:", token_ids)

Output:

Vocabulary size: 64 characters
First 10 chars: ['a', 'b', 'c', 'd', 'e', 'f', 'g', 'h', 'i', 'j']

Tokens:    ['h', 'e', 'l', 'l', 'o', ' ', 'a', 'i', ' ', 'w', 'o', 'r', 'l', 'd', ' ', '2', '0', '2', '6', '!']
Token IDs: [7, 4, 11, 11, 14, 52, 0, 8, 52, 22, 14, 17, 11, 3, 52, 27, 26, 27, 32, 54]

Where character-level works brilliantly

  • ✅ Chinese, Japanese, Korean — these languages use characters as meaningful units (each character can be a whole word or concept)
  • ✅ Spelling correction — working at the character level helps detect and fix typos
  • ✅ Code generation — useful when every character matters, like in programming
  • ✅ Rare languages — tiny vocabulary covers every word, even ones never seen before

The big problem — Sequences Get Very Long

The word "unbelievable" is 13 characters long. At character level, that is 13 tokens. A whole paragraph becomes thousands of tokens. Transformers struggle with very long sequences because memory usage grows quadratically with length — very expensive! 💸

Character-level problems at a glance:

  • 🚫 Sequences become very long — expensive for Transformers
  • 🚫 Individual characters carry very little meaning on their own
  • 🚫 The model has to learn to assemble meaning from letters — much harder than from meaningful word pieces

Method 3 — Subword-Level Tokenization

This is the clever middle ground between word-level and character-level. It keeps common words whole and breaks rare or long words into meaningful pieces.

💡 Think of it like: A smart chef who buys whole apples (common words) and dices rare exotic fruits into smaller pieces (rare words into subwords) 🍎🔪. You always have pieces you can work with — nothing is ever completely unknown!

Instead of splitting on spaces or characters, subword methods learn from data which combinations of characters appear most frequently and build a vocabulary of the most useful pieces.

Example — What Subword Looks Like

Word-level:      "tokenization"  →  ["tokenization"]        (1 token)
Character-level: "tokenization"  →  ["t","o","k","e","n","i","z","a","t","i","o","n"]  (12 tokens)
Subword-level:   "tokenization"  →  ["token", "ization"]    (2 tokens)  ✅

The subword approach splits at a meaningful boundary — token is a real English word, and ization is a common suffix that appears in hundreds of words (realization, modernization, organization). The model learns that pattern once and applies it everywhere! 🧠

The Three Main Subword Algorithms

There are three algorithms used to decide how to split words into subwords. Each is used by different popular models:

  • BPE (Byte-Pair Encoding) → Used by GPT-2, GPT-4, RoBERTa, LLaMA, Mistral. Merges the most frequent pairs of characters/subwords repeatedly until the vocabulary reaches the target size
  • WordPiece → Used by BERT and DistilBERT. Similar to BPE but chooses merges that maximise the probability of the training data. Adds ## prefix to continuation pieces
  • SentencePiece / Unigram → Used by T5, Gemma, LLaMA 3. Works directly on raw text without pre-tokenizing on spaces first — great for multilingual models and languages without spaces between words

Deep Dive — BPE (Byte-Pair Encoding)

BPE is the most widely used subword algorithm. Understanding it will help you understand how GPT-4, LLaMA 3, and Mistral all handle text under the hood.

How BPE Learns a Vocabulary

BPE starts with individual characters and merges the most frequent pair over and over until it reaches the target vocabulary size. Here is the intuition:

Start:  "l o w", "l o w e r", "n e w e s t", "w i d e s t"

Step 1: Most frequent pair → ('e', 's')  →  merge to 'es'
        "l o w", "l o w e r", "n e w es t", "w i d es t"

Step 2: Most frequent pair → ('es', 't')  →  merge to 'est'
        "l o w", "l o w e r", "n e w est", "w i d est"

Step 3: Most frequent pair → ('l', 'o')  →  merge to 'lo'
        "lo w", "lo w e r", "n e w est", "w i d est"

...keeps going until vocabulary size is reached!

After enough merges, frequently occurring word fragments get their own vocabulary entry. Rare words stay as individual characters — which means nothing is ever truly unknown! 🎉

BPE in Action — Using GPT-2's Tokenizer

📌 What this code does: Loads the GPT-2 tokenizer (which uses BPE) and tokenizes several sentences — including tricky ones with rare words and made-up words. This shows you exactly how BPE handles things that would break a word-level tokenizer completely.

from transformers import AutoTokenizer

# GPT-2 uses BPE tokenization
tokenizer = AutoTokenizer.from_pretrained("gpt2")

sentences = [
    "Hello world!",
    "tokenization",            # common word — probably one token
    "antidisestablishmentarianism",  # very rare long word
    "ChatGPT4turbo",           # made-up compound — never in training data
    "2026年のAI"               # mixed language
]

for sentence in sentences:
    tokens = tokenizer.tokenize(sentence)
    print(f"Input:  '{sentence}'")
    print(f"Tokens: {tokens}")
    print(f"Count:  {len(tokens)} tokens")
    print()

Output:

Input:  'Hello world!'
Tokens: ['Hello', 'Ġworld', '!']
Count:  3 tokens

Input:  'tokenization'
Tokens: ['token', 'ization']
Count:  2 tokens

Input:  'antidisestablishmentarianism'
Tokens: ['ant', 'idis', 'establish', 'ment', 'arian', 'ism']
Count:  6 tokens

Input:  'ChatGPT4turbo'
Tokens: ['Chat', 'G', 'PT', '4', 't', 'urbo']
Count:  6 tokens

Input:  '2026年のAI'
Tokens: ['2026', 'å¹´', 'ã', 'Ī', 'AI']
Count:  5 tokens

Notice the Ġ character in front of world — GPT-2's BPE uses this to indicate a space before a word. And antidisestablishmentarianism — one of the longest words in English — got broken into 6 meaningful pieces. Nothing crashed! ✅

Deep Dive — WordPiece (BERT's Tokenizer)

WordPiece is what BERT and its family of models use. It is similar to BPE but with one visible difference — it adds ## at the start of any subword that is not the beginning of a word.

💡 Think of it like: When you split a word, the first piece has no marker, but every continuation piece gets labelled "this is a continuation — not a new word" with the ## prefix. Like putting a hyphen on the second half of a split word 🔗

WordPiece in Action — BERT's Tokenizer

📌 What this code does: Loads the BERT tokenizer and tokenizes a variety of words — common, rare, made-up, and misspelled. This clearly shows the ## continuation marker that distinguishes WordPiece from BPE, and demonstrates that BERT can handle any word gracefully.

from transformers import AutoTokenizer

# BERT uses WordPiece tokenization
tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")

words = [
    "playing",           # common word with suffix
    "tokenization",      # longer word
    "unbelievable",      # un + believable
    "ChatGPT",           # modern term
    "covid19",           # compound word
    "xylophone",         # rare word
]

for word in words:
    tokens = tokenizer.tokenize(word)
    print(f"'{word}'  →  {tokens}")

Output:

'playing'       →  ['playing']
'tokenization'  →  ['token', '##ization']
'unbelievable'  →  ['un', '##believable']
'ChatGPT'       →  ['chat', '##g', '##pt']
'covid19'       →  ['covid', '##19']
'xylophone'     →  ['xy', '##lo', '##phone']

See how ##ization, ##believable, and ##phone all have the ## marker? Those are continuation pieces — they never appear at the start of a word on their own. This helps the model understand word structure! 🧩

BERT's Special Tokens

BERT adds two special tokens around every input:

  • [CLS] (ID 101) → Always the very first token. Stands for "classification". The model's overall understanding of the whole sentence gets stored here
  • [SEP] (ID 102) → Marks the end of a sentence. When feeding two sentences to BERT at once, [SEP] separates them
  • [PAD] (ID 0) → Padding token. Used to fill up shorter sequences in a batch to the same length
  • [UNK] (ID 100) → Unknown token. Used for characters the vocabulary cannot handle at all
  • [MASK] (ID 103) → Used during BERT's training — a token is hidden and the model has to predict what it was

📌 What this code does: Shows BERT's full tokenization of a two-sentence input — including how [CLS], [SEP], and padding work together. This is exactly the format you need to feed text into BERT for any classification task.

from transformers import AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")

# Two sentences — BERT can process pairs for tasks like sentence similarity
sentence_a = "The cat sat on the mat."
sentence_b = "The dog lay on the rug."

encoded = tokenizer(
    sentence_a,
    sentence_b,
    max_length=30,
    padding="max_length",
    truncation=True
)

# Decode the token IDs back to words so we can see what happened
decoded_tokens = tokenizer.convert_ids_to_tokens(encoded["input_ids"])

print("Token IDs:   ", encoded["input_ids"])
print()
print("Tokens:      ", decoded_tokens)
print()
print("Type IDs:    ", encoded["token_type_ids"])
print("(0 = first sentence, 1 = second sentence)")
print()
print("Attn Mask:   ", encoded["attention_mask"])
print("(1 = real token, 0 = padding)")

Output:

Token IDs:   [101, 1996, 4937, 2938, 2006, 1996, 13523, 1012, 102,
              1996, 3899, 3477, 2006, 1996, 11724, 1012, 102,
              0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]

Tokens:      ['[CLS]', 'the', 'cat', 'sat', 'on', 'the', 'mat', '.',
              '[SEP]', 'the', 'dog', 'lay', 'on', 'the', 'rug', '.',
              '[SEP]', '[PAD]', '[PAD]', '[PAD]', '[PAD]', '[PAD]',
              '[PAD]', '[PAD]', '[PAD]', '[PAD]', '[PAD]', '[PAD]', '[PAD]']

Type IDs:    [0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 1, 1, 1, 1, 1,
              0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]
(0 = first sentence, 1 = second sentence)

Attn Mask:   [1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1,
              0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]
(1 = real token, 0 = padding)

Three important things to notice here:

  • The first sentence tokens all have type ID 0, the second have type ID 1 — BERT uses this to tell the two sentences apart
  • Padding tokens at the end have attention mask 0 — this tells the model "ignore these, they are just filler"
  • There are two [SEP] tokens — one at the end of sentence A, one at the end of sentence B

Deep Dive — SentencePiece (T5, LLaMA, Gemma)

SentencePiece is the tokenizer used by modern large language models like T5, LLaMA 2 & 3, Gemma, and Mistral. It has one big design difference from BPE and WordPiece:

SentencePiece treats spaces as part of the token. It does not split text on spaces first. Instead, it adds a special underscore ▁ at the start of tokens that follow a space. This makes it language-agnostic — it works equally well for English, Chinese, Arabic, Hindi, or any mix of languages.

💡 Think of it like: Most tokenizers first cut the text at spaces, then process each word separately. SentencePiece processes the raw text as one continuous stream — like reading a sentence without lifting your pen ✍️. This makes it perfect for multilingual models.

📌 What this code does: Loads the T5 tokenizer (which uses SentencePiece) and tokenizes several sentences, including a multilingual example. Notice the ▁ character that marks the beginning of words — the key visual difference from BERT's ## system.

from transformers import AutoTokenizer

# T5 uses SentencePiece tokenization
tokenizer = AutoTokenizer.from_pretrained("t5-small")

sentences = [
    "Hello world",
    "tokenization is fun",
    "Bonjour le monde",           # French
    "Hola mundo",                 # Spanish
    "नमस्ते दुनिया",              # Hindi
]

for sentence in sentences:
    tokens = tokenizer.tokenize(sentence)
    print(f"Input:  '{sentence}'")
    print(f"Tokens: {tokens}")
    print()

Output:

Input:  'Hello world'
Tokens: ['▁Hello', '▁world']

Input:  'tokenization is fun'
Tokens: ['▁token', 'ization', '▁is', '▁fun']

Input:  'Bonjour le monde'
Tokens: ['▁Bon', 'jour', '▁le', '▁monde']

Input:  'Hola mundo'
Tokens: ['▁Hol', 'a', '▁mundo']

Input:  'नमस्ते दुनिया'
Tokens: ['▁', 'न', 'म', 'स', '्', 'त', 'े', '▁द', 'ु', 'न', 'ि', 'य', 'ा']

The ▁ underscore marks "this token starts after a space". Without it, the model could not tell whether ization is the start of a new word or a continuation of the previous one! 🔍

Byte-Level BPE — How GPT-4 and LLaMA 3 Tokenize

The very latest LLMs in 2025–2026 (GPT-4, LLaMA 3, Mistral, Gemma 2) use a special version called Byte-Level BPE. Instead of starting with characters, it starts with the 256 possible byte values.

This means it can tokenize absolutely any text in any language without ever producing an unknown token — because every possible string of characters can be represented as a sequence of bytes. No [UNK] token needed! ✅

📌 What this code does: Uses the LLaMA 3 tokenizer (byte-level BPE via tiktoken or SentencePiece) to tokenize text in several different languages including emoji and code — demonstrating that it handles everything without any unknown tokens, no matter how unusual the input.

from transformers import AutoTokenizer

# LLaMA 3 uses byte-level BPE
tokenizer = AutoTokenizer.from_pretrained("meta-llama/Meta-Llama-3-8B")

inputs = [
    "Hello! How are you?",
    "日本語のテキスト",         # Japanese text
    "print('hello world')",    # Python code
    "2 + 2 = 4 🎉",           # math with emoji
    "café résumé naïve",       # French accents
]

for text in inputs:
    tokens = tokenizer.tokenize(text)
    print(f"Input:  '{text}'")
    print(f"Tokens: {tokens}")
    print(f"Count:  {len(tokens)}")
    print()

No matter what you throw at it — emoji, Japanese, Python code, accented letters — byte-level BPE always produces a valid tokenization. There are no unknown words because every possible byte value is in the vocabulary from the start! 🌍

Comparing All Three Methods Side by Side

Let's tokenize the exact same sentence with BERT (WordPiece), GPT-2 (BPE), and T5 (SentencePiece) to see how differently they each handle text:

📌 What this code does: Runs the same sentence through three different tokenizers and compares the output side by side. This is the clearest way to see the practical difference between WordPiece, BPE, and SentencePiece. Think of it as putting the same text through three different meat grinders — same input, three very different outputs!

from transformers import AutoTokenizer

sentence = "The tokenization of unbelievable words is fascinating!"

models = {
    "BERT (WordPiece)":         "bert-base-uncased",
    "GPT-2 (BPE)":              "gpt2",
    "T5 (SentencePiece)":       "t5-small",
}

for model_label, model_name in models.items():
    tokenizer = AutoTokenizer.from_pretrained(model_name)
    tokens = tokenizer.tokenize(sentence)
    print(f"--- {model_label} ---")
    print(f"Tokens: {tokens}")
    print(f"Count:  {len(tokens)} tokens")
    print()

Output:

--- BERT (WordPiece) ---
Tokens: ['the', 'token', '##ization', 'of', 'un', '##believable',
         'words', 'is', 'fascinating', '!']
Count:  10 tokens

--- GPT-2 (BPE) ---
Tokens: ['The', 'Ġtoken', 'ization', 'Ġof', 'Ġun', 'belie', 'vable',
         'Ġwords', 'Ġis', 'Ġfascinating', '!']
Count:  11 tokens

--- T5 (SentencePiece) ---
Tokens: ['▁The', '▁token', 'ization', '▁of', '▁un', 'believable',
         '▁words', '▁is', '▁fascinating', '!']
Count:  10 tokens

What to notice:

  • All three split tokenization into token + continuation — they all learned this is a common split!
  • BERT uses ## prefix, GPT-2 uses Ġ for space, T5 uses ▁ for space — three different conventions, same idea
  • The token count is similar across all three — subword methods are fairly consistent in efficiency

Vocabulary Size — Why It Matters

Every tokenizer has a fixed vocabulary — a lookup table of all the tokens it knows. The size of this vocabulary is a critical design choice:

  • BERT: 30,522 tokens (WordPiece, English-focused)
  • GPT-2: 50,257 tokens (BPE, English-focused)
  • LLaMA 3: 128,256 tokens (Byte-level BPE, multilingual)
  • Gemma 2: 256,000 tokens (SentencePiece, multilingual)

💡 Why does size matter?

  • Larger vocabulary → common words stay as single tokens → shorter sequences → faster training and inference
  • Smaller vocabulary → more words get split into pieces → longer sequences → slower but better at rare words and languages
  • Modern multilingual LLMs (2025–2026) use very large vocabularies (100k–300k) to handle dozens of languages efficiently

The Attention Mask — Telling the Model What to Ignore

When you batch multiple sequences together, shorter ones need to be padded to the same length. But you don't want the model to pay attention to those padding tokens — they carry no meaning!

The attention mask is a list of 1s and 0s that tells the model exactly which tokens to look at and which to ignore.

📌 What this code does: Tokenizes three sentences of different lengths in one batch. The tokenizer automatically pads the shorter ones and creates attention masks showing which positions are real tokens versus padding. This is what happens every single time you train a model on a batch of text!

from transformers import AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")

# Three sentences of very different lengths
sentences = [
    "Hi!",
    "The weather is nice today.",
    "Artificial intelligence is transforming the way we build software in 2026."
]

encoded = tokenizer(
    sentences,
    padding=True,        # pad all to the length of the longest
    truncation=True,
    max_length=20,
    return_tensors="pt"  # return PyTorch tensors
)

print("Input IDs shape:", encoded["input_ids"].shape)
print()

for i, sentence in enumerate(sentences):
    print(f"Sentence: '{sentence}'")
    print(f"  input_ids:      {encoded['input_ids'][i].tolist()}")
    print(f"  attention_mask: {encoded['attention_mask'][i].tolist()}")
    real_tokens = encoded['attention_mask'][i].sum().item()
    print(f"  Real tokens: {real_tokens} | Padding tokens: {20 - real_tokens}")
    print()

Output:

Input IDs shape: torch.Size([3, 20])

Sentence: 'Hi!'
  input_ids:      [101, 7632, 999, 102, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]
  attention_mask: [1, 1, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]
  Real tokens: 4 | Padding tokens: 16

Sentence: 'The weather is nice today.'
  input_ids:      [101, 1996, 4633, 2003, 3835, 2651, 1012, 102, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]
  attention_mask: [1, 1, 1, 1, 1, 1, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]
  Real tokens: 8 | Padding tokens: 12

Sentence: 'Artificial intelligence is transforming the way we build software in 2026.'
  input_ids:      [101, 9262, 4454, 2003, 19361, 1996, 2126, 2057, 3857, 4007, 1999, 25682, 1012, 102, 0, 0, 0, 0, 0, 0]
  attention_mask: [1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 0, 0, 0, 0, 0, 0]
  Real tokens: 14 | Padding tokens: 6

"Hi!" only needs 4 real tokens — but it gets padded to 20 to match the others. The attention mask of all-zeros tells the model "those 16 zeros are not real — ignore them completely!" 👻

Decoding — Converting Numbers Back to Text

Tokenization goes both ways! When a model generates output, it produces token IDs which need to be converted back into human-readable text. This is called decoding.

📌 What this code does: Takes a list of token IDs and converts them back into readable text. This is what happens every time a chatbot or language model shows you its response — the model outputs numbers, and the tokenizer's decode function turns them back into words.

from transformers import AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")

# Some token IDs — let's decode them back to text
token_ids = [101, 17662, 2227, 2003, 6429, 999, 102]

# Decode without skipping special tokens
raw_decode = tokenizer.decode(token_ids)
print("With special tokens:   ", raw_decode)

# Decode while skipping special tokens ([CLS], [SEP], [PAD])
clean_decode = tokenizer.decode(token_ids, skip_special_tokens=True)
print("Without special tokens:", clean_decode)

Output:

With special tokens:    [CLS] hugging face is amazing! [SEP]
Without special tokens: hugging face is amazing!

The round trip works perfectly — text → IDs → text again! 🔄

Tokenizer — Inspecting the Vocabulary

Want to look inside a tokenizer's vocabulary? You can see exactly which tokens exist and what IDs they map to:

📌 What this code does: Opens up the tokenizer's vocabulary like a dictionary and lets you look up individual tokens by name or by ID. This is useful for debugging tokenization issues, checking if a specific word is in the vocabulary, or understanding what the tokenizer knows.

from transformers import AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")

# Total vocabulary size
print(f"Vocabulary size: {tokenizer.vocab_size:,} tokens")

# Look up a specific token's ID
print(f"ID of 'hello':     {tokenizer.convert_tokens_to_ids('hello')}")
print(f"ID of '[CLS]':     {tokenizer.convert_tokens_to_ids('[CLS]')}")
print(f"ID of '[SEP]':     {tokenizer.convert_tokens_to_ids('[SEP]')}")
print(f"ID of '[PAD]':     {tokenizer.convert_tokens_to_ids('[PAD]')}")
print(f"ID of '[MASK]':    {tokenizer.convert_tokens_to_ids('[MASK]')}")

# Look up what token ID 1000 is
print(f"Token at ID 1000:  {tokenizer.convert_ids_to_tokens(1000)}")

# Check the first 10 tokens in the vocabulary
vocab = tokenizer.get_vocab()
first_10 = sorted(vocab.items(), key=lambda x: x[1])[:10]
print(f"\nFirst 10 tokens: {first_10}")

Output:

Vocabulary size: 30,522 tokens

ID of 'hello':    7592
ID of '[CLS]':    101
ID of '[SEP]':    102
ID of '[PAD]':    0
ID of '[MASK]':   103

Token at ID 1000: '##s'

First 10 tokens: [('[PAD]', 0), ('[unused0]', 1), ('[unused1]', 2),
                  ('[unused2]', 3), ('[unused3]', 4), ('[unused4]', 5),
                  ('[unused5]', 6), ('[unused6]', 7), ('[unused7]', 8),
                  ('[unused8]', 9)]

Saving and Loading a Tokenizer

Once you download a tokenizer, save it locally so you never need to re-download it. This is especially useful when working offline or when deploying your model:

📌 What this code does: Saves the tokenizer files to a local folder and then loads them back. After saving, you never need an internet connection to use that tokenizer again. This is the standard practice when shipping models to production!

from transformers import AutoTokenizer

# Download and save
tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")
tokenizer.save_pretrained("./my_tokenizer")

print("Tokenizer saved! Files created:")
import os
for f in os.listdir("./my_tokenizer"):
    print(f"  {f}")

# Load from local folder — no internet needed!
local_tokenizer = AutoTokenizer.from_pretrained("./my_tokenizer")

# Test it works
test = local_tokenizer("Hello from local disk!")
print("\nLoaded from disk successfully!")
print("Token IDs:", test["input_ids"])

Output:

Tokenizer saved! Files created:
  tokenizer_config.json
  vocab.txt
  tokenizer.json
  special_tokens_map.json

Loaded from disk successfully!
Token IDs: [101, 7592, 2013, 2334, 4202, 999, 102]

Training Your Own Tokenizer

If you are working with a specialised domain — medical text, legal documents, programming code in a rare language, or a low-resource language — you might want to train your own tokenizer from scratch on your own data.

💡 Think of it like: Standard tokenizers learned English from the internet. If your text is full of medical terms like "pneumonoultramicroscopicsilicovolcanoconiosis", a standard tokenizer will chop it into dozens of tiny pieces. A tokenizer trained on medical text would know that word and handle it efficiently! 🏥

📌 What this code does: Trains a brand new BPE tokenizer from scratch on a small custom corpus. This teaches the tokenizer the most common word pieces in your specific domain, resulting in much more efficient tokenization for your use case compared to a general-purpose tokenizer.

from tokenizers import ByteLevelBPETokenizer

# Your custom training data — replace this with your actual text files
training_corpus = [
    "Machine learning models process text as tokens.",
    "Neural networks learn patterns from large datasets.",
    "Transformers use attention mechanisms to understand context.",
    "Fine-tuning adapts pre-trained models to new tasks.",
    "Tokenization is the first step in any NLP pipeline.",
    # ... in real use, you would have thousands of sentences
]

# Save corpus to a file (tokenizer trainer reads from files)
with open("corpus.txt", "w") as f:
    f.write("\n".join(training_corpus))

# Train a Byte-Level BPE tokenizer on your corpus
tokenizer = ByteLevelBPETokenizer()

tokenizer.train(
    files=["corpus.txt"],
    vocab_size=500,          # small vocab for this demo — use 32000+ in real projects
    min_frequency=2,         # a pair must appear at least twice to be merged
    special_tokens=["", "", "", "", ""]
)

tokenizer.save_model("./custom_tokenizer")
print("Custom tokenizer trained and saved! ✅")

# Test it
output = tokenizer.encode("Transformers use attention mechanisms.")
print("Tokens:", output.tokens)
print("IDs:   ", output.ids)

Output:

Custom tokenizer trained and saved! ✅
Tokens: ['ĠTrans', 'formers', 'Ġuse', 'Ġat', 'ten', 'tion', 'Ġmechan', 'isms', '.']
IDs:    [312, 198, 431, 89, 204, 271, 389, 276, 13]

Your tokenizer learned the specific vocabulary of your domain! The word "Transformers" now gets split based on what appears most frequently in your corpus — not what appeared most often on the general internet. 🎯

Fast vs Slow Tokenizers

Hugging Face offers two versions of most tokenizers. This is important to understand when working at scale:

  • Slow tokenizer → Written in Python. Easier to understand and modify. Good for small experiments
  • Fast tokenizer → Written in Rust (much faster language). Parallelised, up to 100× faster on large datasets. Always use this in production!

📌 What this code does: Loads both versions of the BERT tokenizer and times them on a large batch of text to show the dramatic speed difference. In real training pipelines with millions of examples, this difference matters enormously.

from transformers import BertTokenizer, BertTokenizerFast
import time

# Generate a large batch of text to tokenize
texts = ["The quick brown fox jumps over the lazy dog. " * 10] * 1000

# Time the slow (Python) tokenizer
slow_tokenizer = BertTokenizer.from_pretrained("bert-base-uncased")
start = time.time()
for text in texts:
    slow_tokenizer(text, max_length=128, truncation=True)
slow_time = time.time() - start

# Time the fast (Rust) tokenizer
fast_tokenizer = BertTokenizerFast.from_pretrained("bert-base-uncased")
start = time.time()
for text in texts:
    fast_tokenizer(text, max_length=128, truncation=True)
fast_time = time.time() - start

print(f"Slow tokenizer: {slow_time:.2f} seconds")
print(f"Fast tokenizer: {fast_time:.2f} seconds")
print(f"Speed-up:       {slow_time / fast_time:.1f}x faster!")

Output (approximate):

Slow tokenizer: 12.40 seconds
Fast tokenizer:  0.38 seconds
Speed-up:       32.6x faster!

32× faster on the same text! When you call AutoTokenizer.from_pretrained(), Hugging Face automatically loads the fast version if it's available — so you get this speed boost for free without thinking about it! ✅

Which Tokenizer Does Each Model Use?

As a quick reference, here is a summary of which tokenization method each popular model family uses:

  • BERT, DistilBERT, ALBERT → WordPiece, vocabulary ~30k, ## prefix for continuations
  • GPT-2, GPT-Neo, GPT-J → BPE, vocabulary ~50k, Ġ for spaces
  • RoBERTa, DeBERTa → Byte-level BPE, vocabulary ~50k
  • T5, mT5, Flan-T5 → SentencePiece (Unigram), vocabulary ~32k, ▁ for spaces
  • LLaMA 2 → SentencePiece (BPE), vocabulary 32k
  • LLaMA 3, Mistral, Mixtral → Tiktoken (Byte-level BPE), vocabulary 128k
  • Gemma, Gemma 2 → SentencePiece, vocabulary 256k
  • GPT-4, GPT-4o → Tiktoken cl100k_base, vocabulary 100k
  • Qwen 2.5 → Tiktoken, vocabulary 151k

Common Mistakes — What NOT to Do 🚫

🚫 Don't mix tokenizers between models. Using BERT's tokenizer with GPT-2's model will produce garbage results. Always use the tokenizer that was trained with the model — AutoTokenizer.from_pretrained("model-name") gets the right one automatically.

🚫 Don't forget to set truncation=True. If you tokenize without truncation and your text is longer than the model's maximum length, you will get an error during training — or worse, silent incorrect results.

🚫 Don't ignore the attention mask. If you do padding without passing the attention mask to the model, the model will treat padding tokens as real text and learn from garbage. Always pass attention_mask!

🚫 Don't assume one word = one token. "unbelievable" could be 3 tokens. "ChatGPT4" could be 5 tokens. Always check your actual token counts — they directly affect memory usage and the model's effective context window.

🚫 Don't use the slow tokenizer in production. Always use the fast (Rust) tokenizer for training and inference. It is 10–100× faster and already the default when you use AutoTokenizer.

Good Practices — What TO Do ✅

✅ Always use AutoTokenizer.from_pretrained(). It automatically picks the correct tokenizer and the fast version if available. Never manually instantiate BertTokenizer or GPT2Tokenizer unless you have a specific reason.

✅ Check your token counts before training. Run len(tokenizer(text)["input_ids"]) on a sample of your data. If most texts are much shorter than max_seq_length, you are wasting memory. If many are longer, you are silently truncating important content.

✅ Save your tokenizer with your model. Always call tokenizer.save_pretrained() alongside model.save_pretrained(). Whoever loads your model later needs the exact same tokenizer — including any custom settings you made.

✅ Use batched=True when tokenizing large datasets. When using the Datasets library, always pass batched=True to dataset.map(tokenize_function). The fast tokenizer processes thousands of sequences in parallel — many times faster than one-at-a-time.

✅ Use return_tensors="pt" when feeding to PyTorch. Adding return_tensors="pt" to your tokenizer call returns PyTorch tensors directly — no manual conversion needed. Use "tf" for TensorFlow.

Quick Summary 📝

What we learned today:

  • Word-level → Simple but fails on unknown words. Vocabulary grows huge. Rarely used.
  • Character-level → No unknown words, but sequences get very long. Used for some languages and tasks
  • Subword-level → The standard. Best of both worlds — keeps common words whole, splits rare ones into meaningful pieces
  • BPE → Used by GPT family and LLaMA. Merges most frequent character pairs iteratively
  • WordPiece → Used by BERT. Adds ## to continuation pieces. Chooses merges by likelihood
  • SentencePiece → Used by T5, Gemma, LLaMA. Processes raw text, uses ▁ for spaces. Best for multilingual
  • Special tokens → [CLS], [SEP], [PAD], [MASK] — each has a specific purpose the model learns during training
  • Attention mask → Tells the model which tokens are real and which are padding. Always pass it!
  • Fast tokenizers → Written in Rust, up to 100× faster. The default when using AutoTokenizer

Open Google Colab right now, run AutoTokenizer.from_pretrained("bert-base-uncased"), and start tokenizing your own sentences! You will be surprised how much you learn just by experimenting. Happy tokenizing! 🤗✨

Comments