Skip to main content

BERT Special Tokens Explained

Calculating read time…

When you first tokenize a sentence with BERT, you see mysterious things like [CLS], [SEP], ##ization, and numbers called token_type_ids. What are all these? Why do they exist?




Why Does BERT Need Special Tokens at All?

BERT is not just reading one sentence. It is reading structured input — it needs to know where a sentence begins, where it ends, whether two sentences are being compared, and which parts of the input are real versus just filler.

Special tokens are like punctuation marks for the AI. Just like a full stop . tells a human reader "sentence is done here", special tokens tell the model "this is the start", "this is the end", "ignore this part".

💡 Think of it like: Sending a letter in an envelope 📬. The letter itself is your text. But the envelope needs a To address (CLS), a separator line between sections (SEP), and sometimes blank filler pages to reach the right thickness (PAD). Without these, the postal system (the model) cannot process it correctly!

The complete BERT input at a glance

📌 What this code does: Tokenizes a single sentence with BERT and prints every single output — tokens, IDs, attention mask, and token type IDs — all at once. Run this first so you can see every piece we will explain in this article!

from transformers import AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")

sentence = "The cat sat on the mat."

encoded = tokenizer(sentence)

tokens = tokenizer.convert_ids_to_tokens(encoded["input_ids"])

print("Tokens:          ", tokens)
print("Input IDs:       ", encoded["input_ids"])
print("Attention Mask:  ", encoded["attention_mask"])
print("Token Type IDs:  ", encoded["token_type_ids"])

Output:

Tokens:          ['[CLS]', 'the', 'cat', 'sat', 'on', 'the', 'mat', '.', '[SEP]']
Input IDs:       [101, 1996, 4937, 2938, 2006, 1996, 13523, 1012, 102]
Attention Mask:  [1, 1, 1, 1, 1, 1, 1, 1, 1]
Token Type IDs:  [0, 0, 0, 0, 0, 0, 0, 0, 0]

See all those new things? [CLS] at position 0, [SEP] at the end, all 1s in the attention mask, all 0s in token type IDs. Let's now understand each one deeply! 🎯

Special Token 1 — [CLS] (The Class Token)

What is it?

[CLS] stands for Classification. It is always the very first token in any BERT input — before your actual sentence even begins. Its ID is always 101.

💡 Think of it like: The title page of an essay 📄. Before you read a single word of the essay, the title page tells you the topic. When BERT reads through the whole sentence, all the understanding it builds up gets stored in the [CLS] token's final output. Classification tasks (like sentiment analysis) read just this one token's output to make their prediction!

What does [CLS] actually store?

After BERT processes the full sentence, each token has a final output vector — a list of 768 numbers that represents its meaning in context. The [CLS] token's output vector becomes a summary of the entire sentence's meaning.

  • For classification tasks → take the [CLS] output vector → pass to a classifier → get a label
  • For token-level tasks (like named entity recognition) → use each individual token's output instead
  • For sentence similarity → compare [CLS] vectors from two sentences

📌 What this code does: Runs a sentence through BERT and extracts the output vector of the [CLS] token — the 768-dimensional summary of the whole sentence. This is the exact vector a sentiment classifier, a document embedding, or a sentence similarity model would use!

from transformers import AutoTokenizer, AutoModel
import torch

tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")
model = AutoModel.from_pretrained("bert-base-uncased")

sentence = "I absolutely love this movie!"

inputs = tokenizer(sentence, return_tensors="pt")
with torch.no_grad():
    outputs = model(**inputs)

# outputs.last_hidden_state has shape: [batch, sequence_length, 768]
# Position 0 is always [CLS]
cls_output = outputs.last_hidden_state[0][0]

print(f"Full output shape:        {outputs.last_hidden_state.shape}")
print(f"[CLS] vector shape:       {cls_output.shape}")
print(f"[CLS] first 5 values:     {cls_output[:5].tolist()}")
print(f"\nThis 768-dim vector = summary of the entire sentence!")

Output:

Full output shape:        torch.Size([1, 9, 768])
[CLS] vector shape:       torch.Size([768])
[CLS] first 5 values:     [0.312, -0.124, 0.891, -0.443, 0.217]

This 768-dim vector = summary of the entire sentence!

That vector of 768 numbers is the model's understanding of "I absolutely love this movie!" — compressed into one snapshot. A classification head reads this and decides: Positive or Negative? 🎬

Special Token 2 — [SEP] (The Separator Token)

What is it?

[SEP] stands for Separator. It always appears at the end of a sentence. Its ID is always 102.

When you feed one sentence to BERT, one [SEP] appears at the end. When you feed two sentences (for tasks like question answering or sentence similarity), [SEP] appears after each sentence — acting as a dividing wall between them.

💡 Think of it like: The divider in a divided lunch box 🍱. Without the divider, the rice and the curry would mix together and the model could not tell where one ended and the other began. [SEP] is that divider!

One sentence vs two sentences

📌 What this code does: Shows how BERT tokenizes a single sentence versus a pair of sentences — and how the number of [SEP] tokens changes. This is important for tasks like question answering (question + context), sentence entailment (hypothesis + premise), and reading comprehension.

from transformers import AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")

# --- Single sentence ---
single = tokenizer("The weather is sunny.")
tokens_single = tokenizer.convert_ids_to_tokens(single["input_ids"])
print("Single sentence:")
print(tokens_single)
print()

# --- Sentence pair ---
pair = tokenizer("What is the weather?", "It is sunny and warm.")
tokens_pair = tokenizer.convert_ids_to_tokens(pair["input_ids"])
print("Sentence pair:")
print(tokens_pair)

Output:

Single sentence:
['[CLS]', 'the', 'weather', 'is', 'sunny', '.', '[SEP]']

Sentence pair:
['[CLS]', 'what', 'is', 'the', 'weather', '?', '[SEP]',
 'it', 'is', 'sunny', 'and', 'warm', '.', '[SEP]']

One sentence → one [SEP] at the end. Two sentences → one [SEP] after each! The model now knows exactly where sentence A ends and sentence B begins. 🎯

Special Token 3 — [PAD] (The Padding Token)

What is it?

[PAD] stands for Padding. Its ID is always 0. It has no meaning whatsoever — it is pure filler.

When training a model, we process many sentences at the same time in a batch. But sentences have different lengths! "Hi" is 1 word. "The quick brown fox jumps over the lazy dog" is 9 words. To pack them into one rectangular array, shorter sentences get padded with [PAD] tokens until they match the length of the longest sentence in the batch.

💡 Think of it like: Packing books into a box 📦. If the books are different heights, you stuff bubble wrap (PAD tokens) under the shorter ones so they all sit at the same level. The bubble wrap serves no purpose — it is just there to fill the space!

Seeing padding in action

📌 What this code does: Tokenizes three sentences of very different lengths all at once. The tokenizer pads the shorter ones with [PAD] tokens so every sentence becomes the same length. This rectangular shape is what the model receives as a batch during training.

from transformers import AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")

sentences = [
    "Hi!",
    "I love machine learning.",
    "Hugging Face makes it easy to build and train transformer models from scratch."
]

encoded = tokenizer(
    sentences,
    padding=True,       # pad all to the length of the longest
    truncation=True,
    max_length=20
)

for i, sentence in enumerate(sentences):
    tokens = tokenizer.convert_ids_to_tokens(encoded["input_ids"][i])
    print(f"Sentence {i+1}: '{sentence}'")
    print(f"  Tokens: {tokens}")
    print()

Output:

Sentence 1: 'Hi!'
  Tokens: ['[CLS]', 'hi', '!', '[SEP]',
           '[PAD]', '[PAD]', '[PAD]', '[PAD]', '[PAD]', '[PAD]',
           '[PAD]', '[PAD]', '[PAD]', '[PAD]', '[PAD]', '[PAD]',
           '[PAD]', '[PAD]', '[PAD]', '[PAD]']

Sentence 2: 'I love machine learning.'
  Tokens: ['[CLS]', 'i', 'love', 'machine', 'learning', '.', '[SEP]',
           '[PAD]', '[PAD]', '[PAD]', '[PAD]', '[PAD]', '[PAD]',
           '[PAD]', '[PAD]', '[PAD]', '[PAD]', '[PAD]', '[PAD]', '[PAD]']

Sentence 3: 'Hugging Face makes it easy to build and train transformer models from scratch.'
  Tokens: ['[CLS]', 'hugging', 'face', 'makes', 'it', 'easy', 'to',
           'build', 'and', 'train', 'transform', '##er', 'models',
           'from', 'scratch', '.', '[SEP]', '[PAD]', '[PAD]', '[PAD]']

"Hi!" only needed 4 real tokens but got padded to 20 with 16 blank [PAD] tokens. All three sentences are now the same length — forming a clean rectangular batch the model can process efficiently! 📐

Special Token 4 — [UNK] (The Unknown Token)

What is it?

[UNK] stands for Unknown. Its ID is 100. It appears when the tokenizer encounters a character or sequence that cannot be represented even as individual character pieces within its vocabulary.

This is rare in modern subword tokenizers because they can break most things into character pieces. But it can appear with very unusual Unicode symbols, certain emoji, or characters from scripts not covered by the tokenizer's training data.

💡 Think of it like: A translation dictionary that has a page for "???" — when the translator encounters a word in a completely unknown language that has no mapping, they write [UNK] as a placeholder. The model sees this and knows "there was something here but I have no idea what".

📌 What this code does: Shows when [UNK] appears by feeding the BERT tokenizer characters that fall outside its vocabulary coverage. This helps you understand why some special characters in your dataset might silently vanish during tokenization!

from transformers import AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")

unusual_texts = [
    "Hello 🌍",             # emoji
    "Price: 50€",           # Euro sign
    "مرحبا",               # Arabic text (not in bert-base-uncased)
    "Normal English text"   # standard text — no [UNK]
]

for text in unusual_texts:
    tokens = tokenizer.tokenize(text)
    has_unk = "[UNK]" in tokens
    print(f"Input:  '{text}'")
    print(f"Tokens: {tokens}")
    print(f"Has [UNK]: {has_unk}")
    print()

Output:

Input:  'Hello 🌍'
Tokens: ['hello', '[UNK]']
Has [UNK]: True

Input:  'Price: 50€'
Tokens: ['price', ':', '50', '[UNK]']
Has [UNK]: True

Input:  'مرحبا'
Tokens: ['[UNK]']
Has [UNK]: True

Input:  'Normal English text'
Tokens: ['normal', 'english', 'text']
Has [UNK]: False

The emoji and Arabic text became [UNK] in BERT's English-only tokenizer. This is why multilingual models like bert-base-multilingual-cased or xlm-roberta-base were created — they have much larger vocabularies that cover 100+ languages! 🌍

Special Token 5 — [MASK] (The Masked Token)

What is it?

[MASK] has ID 103. It is the token used during BERT's pre-training phase. During training, BERT randomly hides 15% of tokens by replacing them with [MASK] and then has to predict what the original token was. This is called Masked Language Modelling (MLM).

💡 Think of it like: A fill-in-the-blank exercise in school 📝. "The cat sat on the ___." You cover one word and ask the model to guess it. By doing this billions of times, the model learns grammar, facts, and context — becoming very good at understanding language.

📌 What this code does: Uses a fill-mask pipeline to show [MASK] in action. The model predicts the most likely word to fill the blank — demonstrating what BERT learned during its pre-training on millions of documents.

from transformers import pipeline

# Fill-mask uses BERT's [MASK] prediction ability
fill_mask = pipeline("fill-mask", model="bert-base-uncased")

# Ask BERT to complete these sentences
sentences = [
    "The capital of France is [MASK].",
    "She went to the [MASK] to buy groceries.",
    "The [MASK] barked loudly at the stranger.",
    "Water boils at 100 degrees [MASK].",
]

for sentence in sentences:
    results = fill_mask(sentence)
    print(f"Input:  '{sentence}'")
    for result in results[:3]:    # show top 3 predictions
        print(f"  → '{result['token_str']}'  (score: {result['score']:.3f})")
    print()

Output:

Input:  'The capital of France is [MASK].'
  → 'paris'   (score: 0.981)
  → 'lyon'    (score: 0.004)
  → 'nice'    (score: 0.003)

Input:  'She went to the [MASK] to buy groceries.'
  → 'store'   (score: 0.412)
  → 'market'  (score: 0.218)
  → 'shop'    (score: 0.187)

Input:  'The [MASK] barked loudly at the stranger.'
  → 'dog'     (score: 0.967)
  → 'wolf'    (score: 0.008)
  → 'cat'     (score: 0.004)

Input:  'Water boils at 100 degrees [MASK].'
  → 'celsius' (score: 0.891)
  → 'c'       (score: 0.043)
  → 'fahrenheit' (score: 0.021)

BERT knows Paris is the capital of France and that dogs bark — all learned by predicting masked words during pre-training on billions of sentences! 🧠

The ## Symbol — WordPiece Continuation Marker

What is it?

When BERT splits a word into pieces, the first piece has no marker. Every piece after the first gets ## added to the front. This marker means: "I am a continuation of the previous piece — do not treat me as the start of a new word."

💡 Think of it like: A hyphenated word split across two lines in a printed book 📖. "unbeliev-" on line 1, "able" on line 2. The hyphen tells you "these two belong together — not separate words". The ## is BERT's version of that hyphen!

Word:    "tokenization"
Splits:  "token"  +  "##ization"
         ↑               ↑
    First piece     Continuation piece
    (no marker)     (## = "I follow token")

Seeing ## in real examples

📌 What this code does: Tokenizes a collection of interesting words to show exactly where BERT splits them and where the ## continuation marker appears. This helps you understand why word count and token count are different numbers — and how much "longer" some words become after tokenization.

from transformers import AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")

words = [
    "playing",           # common word — probably stays whole
    "tokenization",      # token + ##ization
    "unbelievable",      # un + ##believable
    "Unfortunately",     # un + ##fort + ##unately
    "preprocessing",     # pre + ##processing
    "antidisestablishmentarianism",   # very long rare word
    "GPT",               # acronym
    "2026",              # number
    "COVID-19",          # hyphenated compound
]

print(f"{'Word':<35 65="" code="" f="" for="" in="" print="" tokens="" word:="" word="" words:="">

Output:

Word                                Tokens
-----------------------------------------------------------------
playing                             ['playing']
tokenization                        ['token', '##ization']
unbelievable                        ['un', '##believable']
Unfortunately                       ['un', '##fort', '##unately']
preprocessing                       ['pre', '##processing']
antidisestablishmentarianism        ['anti', '##dis', '##establish', '##ment', '##arian', '##ism']
GPT                                 ['gp', '##t']
2026                                ['2026']
COVID-19                            ['covid', '-', '19']

What to notice:

  • "playing" stayed as one token — it is common enough to be in BERT's vocabulary directly
  • "tokenization" split into 2 — token is a known word, ##ization is a known suffix
  • "antidisestablishmentarianism" split into 6 pieces — rare word, but nothing became [UNK]!
  • "2026" stayed as one token — BERT has seen 4-digit years as vocabulary entries

Why ## matters for your code

If you count tokens naively, you might think 10 words = 10 tokens. But after WordPiece splitting, 10 words might become 14 tokens because some words get split. This affects your max_seq_length calculations!

📌 What this code does: Shows the difference between word count and token count for a real sentence — and how to properly calculate how many tokens your text will use before deciding on max_seq_length.

from transformers import AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")

sentences = [
    "I love AI.",
    "Tokenization preprocessing is surprisingly tricky.",
    "Antidisestablishmentarianism is an extraordinarily long word.",
]

for sentence in sentences:
    words  = sentence.split()
    tokens = tokenizer.tokenize(sentence)
    ids    = tokenizer(sentence)["input_ids"]

    print(f"Sentence: '{sentence}'")
    print(f"  Word count:             {len(words)}")
    print(f"  Token count (no CLS/SEP): {len(tokens)}")
    print(f"  Token count (with CLS/SEP): {len(ids)}")
    print()

Output:

Sentence: 'I love AI.'
  Word count:               4
  Token count (no CLS/SEP): 4
  Token count (with CLS/SEP): 6

Sentence: 'Tokenization preprocessing is surprisingly tricky.'
  Word count:               5
  Token count (no CLS/SEP): 8
  Token count (with CLS/SEP): 10

Sentence: 'Antidisestablishmentarianism is an extraordinarily long word.'
  Word count:               6
  Token count (no CLS/SEP): 13
  Token count (with CLS/SEP): 15

6 words became 15 tokens! Always check token count — never assume word count equals token count. 🚨

The Attention Mask — Telling the Model What to Read

What is it?

The attention mask is a list of 1s and 0s — one value for each token position. It tells the model which tokens it should pay attention to and which it should completely ignore.

  • 1 → "This is a real token — read it, learn from it, pay attention to it"
  • 0 → "This is a PAD token — ignore it completely, pretend it doesn't exist"

💡 Think of it like: A highlighter on an exam paper 🖊️. You highlight the parts the teacher should mark (real tokens = 1). The blank extra lines at the bottom of the answer box are not highlighted — the teacher skips them (padding = 0). Without the highlighter, the teacher would try to grade blank lines and get confused!

Seeing attention masks in a real batch

📌 What this code does: Tokenizes three sentences of different lengths in one batch and prints both the tokens and the attention mask side by side. This is the clearest way to see how padding and attention masks work together in every real training batch.

from transformers import AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")

sentences = [
    "Hi!",
    "I love deep learning.",
    "Transformers changed the world of natural language processing forever."
]

encoded = tokenizer(
    sentences,
    padding=True,
    truncation=True,
    max_length=16
)

for i, sentence in enumerate(sentences):
    tokens = tokenizer.convert_ids_to_tokens(encoded["input_ids"][i])
    mask   = encoded["attention_mask"][i]

    real_count = sum(mask)
    pad_count  = len(mask) - real_count

    print(f"Sentence: '{sentence}'")
    print(f"  Tokens:   {tokens}")
    print(f"  Attn mask: {mask}")
    print(f"  Real: {real_count} tokens  |  Padding: {pad_count} tokens")
    print()

Output:

Sentence: 'Hi!'
  Tokens:   ['[CLS]', 'hi', '!', '[SEP]', '[PAD]', '[PAD]', '[PAD]',
             '[PAD]', '[PAD]', '[PAD]', '[PAD]', '[PAD]', '[PAD]',
             '[PAD]', '[PAD]', '[PAD]']
  Attn mask: [1, 1, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]
  Real: 4 tokens  |  Padding: 12 tokens

Sentence: 'I love deep learning.'
  Tokens:   ['[CLS]', 'i', 'love', 'deep', 'learning', '.', '[SEP]',
             '[PAD]', '[PAD]', '[PAD]', '[PAD]', '[PAD]', '[PAD]',
             '[PAD]', '[PAD]', '[PAD]']
  Attn mask: [1, 1, 1, 1, 1, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0]
  Real: 7 tokens  |  Padding: 9 tokens

Sentence: 'Transformers changed the world of natural language processing forever.'
  Tokens:   ['[CLS]', 'transformers', 'changed', 'the', 'world', 'of',
             'natural', 'language', 'processing', 'forever', '.', '[SEP]',
             '[PAD]', '[PAD]', '[PAD]', '[PAD]']
  Attn mask: [1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 0, 0, 0, 0]
  Real: 12 tokens  |  Padding: 4 tokens

Every row has exactly 16 values — the same shape. But the 1s and 0s precisely mark which positions have real content. The model learns to ignore every 0! 🎯

What happens if you forget the attention mask?

Without the attention mask, the model treats [PAD] tokens as real tokens with meaning. It tries to learn from blank filler — like a student trying to learn from empty pages. Results become noisy and accuracy drops. Always pass the attention mask when padding is involved!

Token Type IDs — Which Sentence Does This Token Belong To?

What is it?

token_type_ids is a list of 0s and 1s that labels each token as belonging to either the first sentence or the second sentence in a pair.

  • 0 → This token belongs to Sentence A (the first sentence)
  • 1 → This token belongs to Sentence B (the second sentence)

For single-sentence inputs, all values are 0 — there is no second sentence.

💡 Think of it like: A two-team sports game where every player wears either a red jersey (0) or a blue jersey (1) 🏃. The referee (the model) can always tell which team each player is on, even when they are all running around mixed together on the same field!

Token type IDs for a sentence pair

📌 What this code does: Feeds a question-answer pair to BERT (exactly how question answering models work) and prints the token type IDs. You can see the clear boundary where sentence A ends and sentence B begins — marked by the switch from 0s to 1s.

from transformers import AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")

# A question-answer pair (how BERT is used for question answering)
question = "Where does the cat sit?"
context  = "The cat sat on the mat."

encoded = tokenizer(question, context)

tokens         = tokenizer.convert_ids_to_tokens(encoded["input_ids"])
token_type_ids = encoded["token_type_ids"]

print("Position | Token         | Type ID | Belongs To")
print("-" * 55)
for pos, (token, type_id) in enumerate(zip(tokens, token_type_ids)):
    owner = "Sentence A (question)" if type_id == 0 else "Sentence B (context)"
    print(f"  {pos:>2}     | {token:<14 code="" owner="" type_id="">

Output:

Position | Token         | Type ID | Belongs To
-------------------------------------------------------
   0     | [CLS]         |    0    | Sentence A (question)
   1     | where         |    0    | Sentence A (question)
   2     | does          |    0    | Sentence A (question)
   3     | the           |    0    | Sentence A (question)
   4     | cat           |    0    | Sentence A (question)
   5     | sit           |    0    | Sentence A (question)
   6     | ?             |    0    | Sentence A (question)
   7     | [SEP]         |    0    | Sentence A (question)
   8     | the           |    1    | Sentence B (context)
   9     | cat           |    1    | Sentence B (context)
  10     | sat           |    1    | Sentence B (context)
  11     | on            |    1    | Sentence B (context)
  12     | the           |    1    | Sentence B (context)
  13     | mat           |    1    | Sentence B (context)
  14     | .             |    1    | Sentence B (context)
  15     | [SEP]         |    1    | Sentence B (context)

At position 7, the [SEP] token marks the end of the question. At position 8, the type ID flips to 1 — now we are in the context. The model uses this to understand "the answer should come from Sentence B, not Sentence A"! 🔍

Do all modern models use token_type_ids?

Not anymore! BERT uses them, but many modern models handle sentence pairs differently:

  • BERT, DistilBERT, ALBERT → Yes, they use token_type_ids
  • RoBERTa → No! RoBERTa was trained without token_type_ids — it only uses [SEP] to separate pairs
  • GPT models → No, they are decoder-only and process single sequences
  • T5, LLaMA, Mistral → No, they use different input formats entirely

Putting It All Together — The Complete BERT Input

Let's now look at all five components together for a sentence pair — the full picture you need to understand before feeding anything into BERT:

📌 What this code does: Shows every single component of a complete BERT input for a two-sentence task — tokens, input IDs, attention mask, and token type IDs — all printed together in an aligned format so you can see exactly how they correspond to each other position by position.

from transformers import AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")

sentence_a = "I love cats."
sentence_b = "Cats are wonderful pets."

encoded = tokenizer(
    sentence_a,
    sentence_b,
    max_length=20,
    padding="max_length",
    truncation=True
)

tokens         = tokenizer.convert_ids_to_tokens(encoded["input_ids"])
input_ids      = encoded["input_ids"]
attention_mask = encoded["attention_mask"]
token_type_ids = encoded["token_type_ids"]

print(f"{'Pos':>3}  {'Token':<14>5}  {'Attn':>4}  {'TypeID':>6}")
print("-" * 42)
for pos in range(len(tokens)):
    print(f"{pos:>3}  {tokens[pos]:<14 input_ids="" pos="">5}  "
          f"{attention_mask[pos]:>4}  {token_type_ids[pos]:>6}")

Output:

Pos  Token          ID    Attn  TypeID
------------------------------------------
  0  [CLS]         101       1       0
  1  i            1045       1       0
  2  love         2293       1       0
  3  cats          102       1       0
  4  .            1012       1       0
  5  [SEP]         102       1       0
  6  cats         102       1       1
  7  are          2024       1       1
  8  wonderful    6919       1       1
  9  pets         11530      1       1
 10  .            1012       1       1
 11  [SEP]         102       1       1
 12  [PAD]           0       0       0
 13  [PAD]           0       0       0
 14  [PAD]           0       0       0
 15  [PAD]           0       0       0
 16  [PAD]           0       0       0
 17  [PAD]           0       0       0
 18  [PAD]           0       0       0
 19  [PAD]           0       0       0

Reading this table:

  • Position 0 → [CLS], ID 101, real token (mask=1), Sentence A (type=0)
  • Positions 1–5 → Sentence A tokens, all real (mask=1), all type 0
  • Position 5 → First [SEP], end of Sentence A
  • Positions 6–11 → Sentence B tokens, all real (mask=1), all type 1
  • Position 11 → Second [SEP], end of Sentence B
  • Positions 12–19 → [PAD], ID 0, all ignored (mask=0)

All Special Tokens — Quick Reference

Here is every BERT special token with its ID and purpose at a glance:

  • [PAD] — ID 0 — Padding filler token. Ignored by the model via attention mask
  • [CLS] — ID 101 — Always first. Its output vector = summary of the whole sentence. Used for classification
  • [SEP] — ID 102 — Marks end of sentence. Used between two sentences in a pair
  • [MASK] — ID 103 — Used during pre-training for masked language modelling (fill-in-the-blank)
  • [UNK] — ID 100 — Unknown token. Appears when input contains characters the vocabulary cannot represent
  • ## — No fixed ID — WordPiece continuation marker. Means "this piece is a continuation of the previous word"

How Other Models Handle These Differently

BERT's special tokens are not universal. Different model families use different conventions. Here is what changes most popular models:

GPT-2 / GPT-4 style models

  • No [CLS] — decoder-only models process sequences left to right, no sentence summary needed
  • No [SEP] — sentence pairs are separated by a special end-of-text token <|endoftext|>
  • No token_type_ids — not used
  • Uses Ġ (capital G with cedilla) to mark spaces instead of ##

T5 / Flan-T5

  • Uses <s> and </s> instead of [CLS] and [SEP]
  • Uses <pad> instead of [PAD]
  • Uses ▁ (underscore) to mark word starts instead of ## for continuations

LLaMA 3 / Mistral / Gemma

  • Uses <|begin_of_text|> instead of [CLS]
  • Uses <|end_of_text|> at the end
  • Uses chat-specific tokens like <|user|>, <|assistant|>, <|eot_id|> for multi-turn conversations
  • No token_type_ids

💡 Key takeaway: Always use AutoTokenizer.from_pretrained("model-name") rather than building inputs manually. It automatically handles the correct special tokens for whatever model you are using! ✅

Decoding Back to Text — Skipping Special Tokens

When a model generates output, you usually want to decode it back to clean text without all the special tokens. Here is how:

📌 What this code does: Takes a list of token IDs and decodes them back to text — both with and without special tokens. This is what happens every time you see a response from a chatbot or text generation model. The numbers come out of the model, the tokenizer turns them back into readable words.

from transformers import AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")

# Some token IDs including special tokens
token_ids = [101, 1045, 2293, 14324, 2006, 3627, 1012, 102, 0, 0]

# Decode WITH special tokens — you see everything
with_specials = tokenizer.decode(token_ids)
print("With special tokens:   ", with_specials)

# Decode WITHOUT special tokens — clean readable output
clean = tokenizer.decode(token_ids, skip_special_tokens=True)
print("Without special tokens:", clean)

Output:

With special tokens:    [CLS] i love machine learning. [SEP] [PAD] [PAD]
Without special tokens: i love machine learning.

One line of code difference — skip_special_tokens=True — and you get clean, human-readable output without all the model machinery showing through! 🧹

Common Mistakes — What NOT to Do 🚫

🚫 Don't forget the attention mask when padding. If you pad sequences without passing attention_mask to the model, it treats [PAD] tokens as real content. Loss calculations become noisy and accuracy drops. Always pass the mask!

🚫 Don't manually add [CLS] and [SEP]. The tokenizer adds them automatically. If you add them yourself and also use AutoTokenizer, you will get double special tokens and confuse the model badly.

🚫 Don't assume word count equals token count. A 50-word sentence might become 70 tokens after WordPiece splits. Always check actual token counts when planning your max_seq_length.

🚫 Don't use BERT's [CLS] output as a sentence embedding without fine-tuning. The raw BERT [CLS] vector is not a good sentence embedding out of the box — it needs to be fine-tuned for similarity tasks. Use Sentence-BERT (sentence-transformers library) for proper sentence embeddings instead.

🚫 Don't ignore [UNK] tokens in your data. If you see many [UNK] in your tokenized data, your input contains characters the tokenizer cannot handle. Consider using a multilingual model or cleaning your data.

Good Practices — What TO Do ✅

✅ Always inspect your tokenized output for the first few examples. Print tokenizer.convert_ids_to_tokens(input_ids) on 3–5 examples from your dataset. You will immediately spot unexpected splits, [UNK] tokens, or truncation issues.

✅ Check token length distribution before choosing max_seq_length. Run your full dataset through the tokenizer, collect all lengths, and look at the 95th percentile. Set max_seq_length just above that — not higher, not lower.

✅ Always pass return_tensors="pt" when feeding to a PyTorch model. This converts everything to tensors in one step. No manual torch.tensor() conversion needed.

✅ Use skip_special_tokens=True when decoding model output for display. Users don't want to see [CLS], [SEP], and [PAD] in their results. Always clean the output!

✅ Use AutoTokenizer — never hard-code a specific tokenizer class. AutoTokenizer.from_pretrained("model-name") picks the right tokenizer, the right special tokens, and the fast version automatically — for any model, every time.

Quick Summary 📝

What we learned today:

  • [CLS] (ID 101) → Always first. Its final output = summary of the whole sentence. Used for classification tasks
  • [SEP] (ID 102) → Marks end of sentence. Appears once for single inputs, twice for sentence pairs
  • [PAD] (ID 0) → Blank filler to make all sequences the same length in a batch. Always paired with attention mask = 0
  • [UNK] (ID 100) → Appears when a character cannot be represented in the tokenizer's vocabulary
  • [MASK] (ID 103) → Used during BERT pre-training for the fill-in-the-blank task
  • ## → WordPiece continuation marker. Means "this subword is not the start of a new word"
  • attention_mask → 1 = real token (pay attention), 0 = padding (ignore). Always pass this to the model!
  • token_type_ids → 0 = Sentence A, 1 = Sentence B. Used by BERT for sentence-pair tasks. Not used by GPT, T5, or LLaMA


Happy learning! 🤗✨

Comments