Deep learning models can't read text directly — every model first has to convert words into numbers, and how it does that conversion shapes everything downstream. Tokenization is the step that breaks raw text into discrete units a model can count and index; embeddings are the step that turns each of those discrete units into a dense vector of numbers a network can actually compute with. Together they're the translation layer between human language and the matrix multiplications covered elsewhere in this series. 🔤
This matters in production because tokenization and embedding choices directly determine a model's vocabulary size, its ability to handle rare or unseen words, and its memory footprint — get them wrong and you'll see garbled outputs, exploding vocabulary tables, or a model that silently fails on text unlike anything it saw in training. ⚙️
Figure 1. Original diagram: a token ID is just a row index into a table of learned vectors.
📑 In This Post
- Foundations: why models need tokens and vectors, not raw text
- Tokenization: splitting text into countable units
- Subword tokenization: solving the rare-word problem
- Embeddings: turning IDs into meaningful vectors
- Real example: word2vec's analogy result
- Implementation walkthrough with short code examples
- Enterprise rollout: versioning vocabularies and embeddings
- Common mistakes and why they hurt production systems
- FAQ
- References & further reading
- Summary
🔀 Quick Comparison: Word-Level vs. Subword vs. Character-Level Tokenization
| Property | Word-level | Subword (e.g. BPE) | Character-level |
|---|---|---|---|
| Vocabulary size | Very large | Fixed, moderate | Very small |
| Handles unseen words | Poorly (out-of-vocabulary) | Well, by falling back to smaller pieces | Always, but loses word-level structure |
| Sequence length | Shortest | Moderate | Longest |
| Documented in | Classic NLP pipelines | Sennrich, Haddow & Birch, ACL 2016 | Classic NLP pipelines |
1. Foundations: Why Models Need Tokens and Vectors, Not Raw Text
🧠 Child-friendly analogy first: imagine trying to do arithmetic with the actual English words "three" and "four" instead of the numbers 3 and 4 — you can't multiply letters. Before any math can happen, the words first need to become numbers, and then those numbers need to carry some meaning beyond just being an arbitrary label. Tokenization is the first conversion (words to countable pieces), and embeddings are the second (pieces to meaningful numbers).
A neural network, as covered throughout this series, only ever computes with tensors of numbers — weighted sums, activation functions, matrix multiplications. Text arrives as a string of characters, which has no numeric structure a network can act on directly. Two separate problems need solving before text reaches the first layer of a model: deciding what the countable "units" of text even are (tokenization), and deciding what number, or more precisely what vector of numbers, each unit should be represented by (embedding).
🎯 Use this when explaining to a teammate why "just feed the model the text" is never literally what happens under the hood.
2. Tokenization: Splitting Text into Countable Units
🧠 Analogy: think of cutting a sentence into individual LEGO bricks so you can sort and count them. If your bricks are whole words, "unhappiness" is one giant, unique brick you may never have seen before. If your bricks are individual letters, you have very few unique brick shapes, but you need dozens of them to rebuild even a short word. Tokenization is choosing the size of brick you cut everything into.
What it does: a tokenizer takes a raw string and splits it into a sequence of discrete units — words, subword pieces, or individual characters, depending on the scheme — and then maps each unit to an integer ID from a fixed vocabulary.
Why it's needed: a model can only look up an embedding for a token it has an ID for. Simple whole-word tokenization runs into a hard problem: any word not seen during vocabulary construction has no ID at all, which is typically called the out-of-vocabulary problem.
What fails without it: a fixed word-level vocabulary either has to grow enormous to cover a language's full range of words (including rare names, typos, and technical terms), or it has to give up and map every unseen word to one generic "unknown" token, throwing away whatever information that specific word carried.
🎯 Use this when reasoning about why a model's vocabulary size and choice of tokenization scheme were decided the way they were.
3. Subword Tokenization: Solving the Rare-Word Problem
🧠 Analogy: imagine a language where you don't need a separate word for every possible concept, because you can build new ones out of a moderate set of reusable syllables — "un-", "happi-", "-ness" combine into "unhappiness" without "unhappiness" ever needing to be memorized as its own separate unit. Subword tokenization gives a model exactly this ability: a moderate, fixed set of reusable pieces that can still spell out words it has never seen whole.
What it does: Sennrich, Haddow, and Birch's 2016 paper introduces byte pair encoding (BPE) as a word segmentation method for neural machine translation, encoding rare and unknown words as sequences of subword units built from a compression algorithm rather than requiring every whole word to appear in the vocabulary. (Sennrich, Haddow & Birch, "Neural Machine Translation of Rare Words with Subword Units," ACL 2016)
Why it's needed: the same paper's abstract explains the underlying intuition directly: various word classes are translatable via smaller units than whole words — names via character copying or transliteration, compound words via compositional translation, and cognates or loanwords via phonological and morphological transformations. Building the vocabulary out of these smaller, reusable units lets a fixed-size vocabulary still represent an effectively open-ended set of whole words.
✅ Worked example: the same paper's abstract reports a specific, measured result: subword models improved over a back-off dictionary baseline on the WMT 2015 translation tasks by up to 1.1 BLEU for English-to-German and up to 1.3 BLEU for English-to-Russian. (Sennrich, Haddow & Birch, ACL 2016)
What fails without it: a purely word-level vocabulary has no principled way to handle a word it has never seen, while a purely character-level vocabulary avoids that problem but produces much longer token sequences for the same text, since every character needs its own step through the model. Subword tokenization is a middle ground between these two failure modes.
🎯 Use this when deciding whether a new model needs a subword tokenizer instead of a simpler whole-word vocabulary.
4. Embeddings: Turning IDs into Meaningful Vectors
🧠 Analogy: imagine a giant coat-check counter where every possible token has its own numbered hook, and hanging on that hook is a whole outfit's worth of descriptive tags — not just an arbitrary number, but a rich profile capturing how that token tends to behave. Looking up a token's embedding is exactly like handing over its numbered ticket and getting back that rich profile instead of just an empty peg number.
What it does: PyTorch's own documentation describes its embedding module directly as a simple lookup table that stores embeddings of a fixed dictionary and size, taking a list of integer indices as input and returning the corresponding dense vectors as output. (PyTorch Embedding documentation) Under the hood, this lookup table is simply a matrix with one row per vocabulary entry and one column per embedding dimension; retrieving a token's embedding is just selecting that token's row.
Why it's needed: a raw integer token ID carries no meaning beyond being distinct from other IDs — token 812 isn't "closer" to token 813 in any linguistically useful sense. An embedding vector, by contrast, is a set of learned coordinates that training can shape so that tokens used in similar contexts end up with similar vectors, giving the model a genuinely useful geometric notion of similarity to work with.
How it works, step by step:
- Build a vocabulary of tokens and assign each one a unique integer ID.
- Create an embedding matrix with one row per vocabulary entry and a chosen number of columns (the embedding dimension).
- Initialize this matrix, typically with small random values before any training.
- For each token ID in an input sequence, retrieve the corresponding row from the matrix, exactly as PyTorch's Embedding module does.
- Feed the retrieved vectors into the rest of the network, and let backpropagation update the embedding matrix's rows along with every other trainable parameter.
What fails without it: PyTorch's documentation also notes a specific, easy-to-miss detail: if a padding_idx is specified, that row's embedding does not contribute to the gradient and is never updated during training, so it remains fixed as intended for representing padding rather than a real token. Forgetting to specify this for an actual padding token means the model wastes capacity learning a "meaning" for what should be an empty placeholder.
🎯 Use this when setting up a new embedding layer, especially when deciding how to handle padding tokens in variable-length sequences.
5. Real Example: Word2vec's Analogy Result
Figure 2. Original diagram: the same directional offset captures a relationship, regardless of which word pair you start from.
One of the most cited demonstrations that trained embeddings capture genuine relationships, not just arbitrary numbers, comes from Tomas Mikolov, Wen-tau Yih, and Geoffrey Zweig's 2013 paper "Linguistic Regularities in Continuous Space Word Representations." Its abstract states directly that these vector-space word representations are surprisingly good at capturing syntactic and semantic regularities, with each relationship characterized by a relation-specific vector offset that allows vector-oriented reasoning; the paper's specific example is that the male/female relationship is automatically learned, such that computing "King minus Man plus Woman" results in a vector very close to the vector for "Queen." (Mikolov, Yih & Zweig, "Linguistic Regularities in Continuous Space Word Representations," NAACL 2013)
This result depends on the specific model and training data used to produce those embeddings, and later research has further examined the strength and generality of this kind of analogy behavior across different embedding methods; it's cited here as a documented, historically significant demonstration rather than a claim that every embedding space exhibits this property equally well.
🎯 Use this when explaining why embeddings are described as capturing "meaning" rather than just serving as arbitrary lookup indices.
6. Implementation Walkthrough
The short, illustrative examples below show the concepts above as code. They are simplified teaching snippets, not production-ready modules.
# Illustrative example: a tiny embedding lookup, from scratch
vocab_size = 10
embedding_dim = 4
embedding_table = random_matrix(shape=(vocab_size, embedding_dim))
def look_up(token_ids, table):
return [table[token_id] for token_id in token_ids]
token_ids = [4, 2, 9] # produced by a tokenizer elsewhere
vectors = look_up(token_ids, embedding_table)
# Illustrative example: an embedding layer with a reserved padding index
embedding_layer = create_embedding_layer(
num_embeddings=30000,
embedding_dim=256,
padding_idx=0, # index 0 is reserved for padding, never updated
)
batch = create_tensor([[15, 302, 0, 0], [88, 4021, 12, 0]]) # padded to equal length
vectors = embedding_layer.forward(batch) # -> shape (2, 4, 256)
Best practices: always set padding_idx explicitly when your sequences are padded to a common length, and log the actual vocabulary size a tokenizer produces against the embedding table's configured size, since a mismatch between the two raises an index-out-of-range error the moment an unexpected token ID appears.
🎯 Use this when wiring up a new text pipeline's tokenizer and embedding layer together for the first time.
7. Enterprise Rollout: Versioning Vocabularies and Embeddings
A tokenizer's vocabulary and an embedding table's learned weights are both load-bearing artifacts that must stay in sync with the model that was trained alongside them.
Ownership and governance: version the tokenizer's vocabulary file alongside the model checkpoint it was trained with, since a tokenizer update that changes how even a few words are split shifts every downstream token ID, silently invalidating a model trained against the older vocabulary.
CI gates: add a test that tokenizes a fixed set of known sentences and asserts the exact resulting token IDs, so any change to the tokenizer's vocabulary or splitting rules is caught immediately, before it silently changes how production text is processed.
Dataset and checkpoint versioning: store the tokenizer version and vocabulary size as explicit metadata on every model checkpoint, since loading a checkpoint's embedding table with the wrong tokenizer produces token IDs that map to the wrong rows entirely, with no error raised if the vocabulary sizes happen to match.
Dashboards and alerts: track the rate of unknown or fallback tokens seen in production traffic, since a rising rate can signal a shift in the kind of text a model is now receiving, such as new terminology or a new language the tokenizer wasn't designed for.
Incident response: when a text model produces unexpectedly poor output, capture the exact tokenized sequence, not just the raw input text, as part of the incident record, since a tokenization mismatch upstream of the model is often invisible until you look at the actual token IDs the model received.
🎯 Use this when a text-processing pipeline is being handed off to a team that will retrain or update the tokenizer over time.
8. Common Mistakes
Using a different tokenizer version at inference time than the model was trained with. Because token IDs are only meaningful relative to the specific vocabulary that assigned them, even a small vocabulary update shifts what ID a given piece of text maps to. The production impact is a model receiving token IDs that don't correspond to what its embedding table was trained to expect, producing degraded or nonsensical output with no explicit error.
Forgetting to set padding_idx for padded sequences. Without it, PyTorch's Embedding layer treats the padding token like any other token, learning and updating a real embedding vector for what should be a meaningless placeholder. The production impact is wasted model capacity and, in some architectures, subtle interference from the padding token's now-nonzero embedding leaking into downstream computations.
Choosing a word-level vocabulary for a domain with many rare or novel terms. As Sennrich, Haddow, and Birch's paper's own motivation describes, whole-word vocabularies handle the open-vocabulary nature of real text poorly. The production impact is a high rate of unknown-token fallbacks on exactly the specialized or rare terms that may matter most for a given application.
Assuming embedding similarity always reflects true semantic similarity in every direction. Word2vec's documented analogy result is a genuine, historically important finding, but it describes a specific type of learned relationship in a specific embedding space; treating every dimension of similarity in every embedding space as automatically meaningful without evaluation is an overgeneralization the original narrow finding does not support.
Letting vocabulary size and embedding dimension grow without considering memory cost. The embedding table's total parameter count is the vocabulary size multiplied by the embedding dimension, and for large vocabularies this can become one of the largest components of a model's total parameter count, an easily overlooked cost when tuning vocabulary size in isolation from the rest of the architecture.
❓ FAQ
What's the actual difference between a token ID and an embedding?
A token ID is just an arbitrary integer index assigned by the tokenizer's vocabulary, carrying no meaning on its own. An embedding is the dense vector of numbers retrieved from that ID's row in a lookup table, and it's this vector, not the raw ID, that the rest of the network actually computes with.
Why did byte pair encoding become popular for tokenization?
Sennrich, Haddow, and Birch's 2016 paper demonstrated it as a simpler and more effective way to handle open-vocabulary translation than backing off to a dictionary for rare words, with a measured improvement of up to 1.1 to 1.3 BLEU on their tested translation tasks.
Does every model use the same tokenizer?
No. Different models are trained with different tokenizers and vocabularies, which is exactly why a model's checkpoint and its tokenizer must be versioned and distributed together; token IDs from one tokenizer are meaningless to a model trained with a different one.
Are embedding vectors fixed once training finishes, or can they be updated later?
They're ordinary learnable parameters, updated by backpropagation just like any other weight during training, and they stop changing once training ends unless the model is fine-tuned further, following the same rules covered in this series' pretrained-models article.
Is the "king minus man plus woman equals queen" result still considered accurate today?
It remains a genuine, well-documented historical finding from the cited 2013 paper's specific experiments. Whether it holds with equal cleanliness across every modern embedding method and every analogy is a separate empirical question this article doesn't claim to resolve; it's presented here as one specific, citable result rather than a universal guarantee.
🔗 References & Further Reading
- PyTorch — torch.nn.Embedding (official documentation)
- Sennrich, Haddow & Birch — "Neural Machine Translation of Rare Words with Subword Units," ACL 2016 (original paper)
- Mikolov, Yih & Zweig — "Linguistic Regularities in Continuous Space Word Representations," NAACL 2013 (original paper)
PyTorch is a trademark of the PyTorch Foundation. Both academic papers are credited to their original authors and linked above.
📝 Summary
- Tokenization splits raw text into discrete units and assigns each one an integer ID from a fixed vocabulary.
- Subword schemes like byte pair encoding solve the rare-word problem by building words from smaller, reusable pieces.
- An embedding layer is a simple lookup table, turning an arbitrary token ID into a dense, trainable vector of numbers.
- Word2vec's 2013 analogy result showed that trained embeddings can capture genuine relational structure, not just arbitrary labels.
- A model's tokenizer and its embedding table are inseparable, versioned artifacts; mismatching them silently breaks the model.
- Treat vocabulary version, padding index handling, and embedding dimension as one governed, tested configuration.
From LEGO-brick word pieces to a coat-check ticket for meaning — that's how text becomes something a network can actually compute with. Happy building! 🚀
Comments
Post a Comment