Skip to main content

Understanding Transformer Architecture

Calculating read time…

Imagine you have a magical translator who can read an entire book, understand every sentence at once, and then write a perfect summary — all in the blink of an eye. That magical translator is called a Transformer. 

Almost every AI tool you use today — ChatGPT, Google Translate, GitHub Copilot, Gemini — is built on the Transformer architecture. In this guide, we will break it down piece by piece using simple stories, colorful diagrams, and beginner-friendly code.


✅ What You Will Learn :
  • What a Transformer is and why it was invented
  • How Input Embedding turns words into numbers
  • How Positional Encoding adds "order" to words
  • How the Encoder reads and understands input
  • How the Decoder writes the output step by step
  • What Multi-Head Attention, Layer Norm, and Residual Connections do
  • How the Final Linear + Softmax layers pick the right word
  • Working Python code examples with Hugging Face 🤗

No prior AI experience needed. If you can read this sentence, you are ready to learn how AI reads sentences too!

🌍 Why Was the Transformer Invented?

Before 2017, AI used older models called RNNs (Recurrent Neural Networks). Think of an RNN like a student who reads one word at a time, left to right, and tries to remember everything in a single notebook.

The problem? By the time the student reached the end of a long sentence, they had forgotten what was at the beginning! That's called the long-range dependency problem.

In 2017, a group of Google researchers published a paper called "Attention Is All You Need". Their big idea was simple but brilliant:

💡 Big Idea: Instead of reading one word at a time, what if the AI could look at ALL words at the same time and decide which words are most important for understanding each other? That's exactly what Attention does!

This new approach was called the Transformer, and it changed the entire world of AI. Let's see how it works! 🎯

🗺️ The Big Picture — Transformer Architecture Map

Before diving into each part, let's look at the full picture. Think of the Transformer like a translation machine with two main rooms:

  • 🟦 Room 1 — The Encoder: Reads and understands the input sentence
  • 🟧 Room 2 — The Decoder: Writes the output sentence word by word
┌─────────────────────────────────────────────────────────────────┐
│                   🤖 TRANSFORMER ARCHITECTURE                    │
└─────────────────────────────────────────────────────────────────┘

   INPUT TEXT                            OUTPUT TEXT
  "I love cats"                         "J'aime les chats"
       │                                        ▲
       ▼                                        │
┌─────────────┐                       ┌─────────────────┐
│   Input     │                       │  Final Linear   │
│  Embedding  │                       │   + Softmax     │
└──────┬──────┘                       └────────┬────────┘
       │                                        │
       ▼                                        │
┌─────────────┐                       ┌─────────────────┐
│  Positional │                       │    Decoder      │
│  Encoding   │                       │  (×6 layers)    │──▶ Output
└──────┬──────┘                       └────────┬────────┘   Embedding
       │                                        │              +
       ▼                                        │           Positional
┌─────────────┐                                │           Encoding
│   Encoder   │───── Context ────────────────▶ │
│  (×6 layers)│      Vectors                   │
└─────────────┘                                │
                                               ▼
                                      ┌─────────────────┐
                                      │   Output Token  │
                                      │  (one at a time)│
                                      └─────────────────┘

  Each Encoder/Decoder Layer Contains:
  ┌─────────────────────────────────┐
  │  Multi-Head Self-Attention      │
  │         +                       │
  │  Add & Layer Normalization      │
  │         +                       │
  │  Feed-Forward Network           │
  │         +                       │
  │  Add & Layer Normalization      │
  └─────────────────────────────────┘


🔤 Step 1 — Input Embedding (Turning Words into Numbers)

Computers cannot understand words like "cat" or "dog" directly. They only understand numbers. So the first step is to convert each word into a list of numbers. This list of numbers is called an embedding.

🍎 Real-World Analogy

Imagine you have a huge dictionary where every word has a secret code — not just one number, but a list of 512 numbers that captures the meaning of that word. Words with similar meanings (like "cat" and "kitten") will have very similar number lists. Words that are very different (like "cat" and "airplane") will have very different number lists.

📌 Key Fact:
In the original Transformer, each word is converted to a vector of 512 numbers. This is called embedding dimension = 512 (also written as d_model = 512).
  Input Words → Embedding Table → Number Vectors

  "I"      → [ 0.21,  0.87, -0.34,  0.56, ... ]  (512 numbers)
  "love"   → [ 0.91, -0.12,  0.75, -0.23, ... ]  (512 numbers)
  "cats"   → [ 0.44,  0.63, -0.51,  0.82, ... ]  (512 numbers)

  Similar words → Similar number patterns
  "cat"    → [ 0.45,  0.64, -0.50,  0.81, ... ]
  "kitten" → [ 0.43,  0.62, -0.52,  0.80, ... ]  ← very close!
📋 What this code does:
Below, we use Hugging Face's tokenizer to convert a sentence into numbers (token IDs). Think of tokenization as the step that breaks a sentence into individual words or word-pieces, and gives each one a unique ID number. It's like giving every word in the dictionary an ID card number!
from transformers import AutoTokenizer

# Load a pre-trained tokenizer (BERT model)
tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")

# Our input sentence
sentence = "I love cats"

# Convert sentence to token IDs (numbers!)
tokens = tokenizer(sentence, return_tensors="pt")

print("Token IDs:", tokens["input_ids"])
print("Token words:", tokenizer.convert_ids_to_tokens(tokens["input_ids"][0]))

Output:

Token IDs: tensor([[101, 1045, 2293, 8870, 102]])
Token words: ['[CLS]', 'i', 'love', 'cats', '[SEP]']

See how "I love cats" became a list of numbers? That's Input Embedding in action! 🎯 (The [CLS] and [SEP] are special tokens added by BERT — don't worry about them for now.)

✅ DO:
  • Always use a tokenizer that matches your model (e.g., BERT tokenizer for BERT model)
  • Understand that embeddings are learned during training — the model figures out the best numbers
❌ DON'T:
  • Don't mix tokenizers — using GPT-2's tokenizer with BERT will give wrong results!
  • Don't think of embeddings as simple word counts — they are rich, high-dimensional meaning representations

📍 Step 2 — Positional Encoding (Telling AI Where Each Word Lives)

Here's a problem. After embedding, the model has the number vectors for all words — but it has no idea about the order of the words!

"Cat chases dog" and "Dog chases cat" have the same words, but very different meanings. The model needs to know which word came first, second, third...

🎵 Real-World Analogy

Think of a music playlist. Songs are great, but their order matters for the mood. Positional Encoding is like stamping each song with "Song #1", "Song #2", "Song #3"... so the player knows the sequence.

  Word Embedding + Positional Encoding = Final Input

  "I"    [0.21, 0.87, -0.34, ...]   ← word meaning
       + [0.00, 1.00,  0.00, ...]   ← position 1 signal
       = [0.21, 1.87, -0.34, ...]   ← final enriched vector

  "love" [0.91, -0.12, 0.75, ...]   ← word meaning
       + [0.84,  0.54,  0.84, ...]  ← position 2 signal
       = [1.75,  0.42,  1.59, ...]  ← final enriched vector

  Each position gets a unique "fingerprint" using
  sine and cosine waves — like a musical note for each spot!

The original Transformer uses a clever mathematical trick: sine and cosine waves of different frequencies to create a unique positional fingerprint. Modern models like BERT use learned positional embeddings instead — where the model learns the best positional patterns during training itself.

📋 What this code does:
Below, we manually build a simple Positional Encoding layer using PyTorch. This creates a table of sine/cosine values — one row for each word position (up to 5000 positions), and one column for each embedding dimension (512). Imagine it like giving every seat in a stadium a unique colored wristband!
import torch
import math

def positional_encoding(max_len, d_model):
    """
    Creates a positional encoding table.
    max_len = maximum number of words in a sentence
    d_model = size of each word's number vector (512)
    """
    # Create an empty table: rows = positions, columns = dimensions
    pe = torch.zeros(max_len, d_model)

    # Position indices: 0, 1, 2, 3, ...
    position = torch.arange(0, max_len).unsqueeze(1).float()

    # Division term for sine/cosine waves
    div_term = torch.exp(
        torch.arange(0, d_model, 2).float() * (-math.log(10000.0) / d_model)
    )

    # Even positions → use sine wave
    pe[:, 0::2] = torch.sin(position * div_term)

    # Odd positions → use cosine wave
    pe[:, 1::2] = torch.cos(position * div_term)

    return pe

# Create positional encoding for up to 10 words, 512 dimensions
pe = positional_encoding(max_len=10, d_model=512)
print(f"Positional Encoding shape: {pe.shape}")
print(f"First position values (first 5): {pe[0, :5]}")
print(f"Second position values (first 5): {pe[1, :5]}")

Output:

Positional Encoding shape: torch.Size([10, 512])
First position values (first 5):  tensor([0.0000, 1.0000, 0.0000, 1.0000, 0.0000])
Second position values (first 5): tensor([0.8415, 0.5403, 0.0100, 0.9999, 0.0001])

Each position (row) has a unique pattern of numbers. The model learns to use these patterns to figure out word order. Pretty smart, right? 🧠

📌 Trend Note:
Modern large language models (LLMs) like LLaMA 3, Mistral, and Gemini 2 use an even better technique called RoPE (Rotary Positional Embedding). Instead of adding a separate positional vector, RoPE rotates the Query and Key vectors to encode position. This gives the model better ability to handle very long texts (millions of tokens!). You'll see this in advanced models. 🔥

🧠 Step 3 — The Encoder (The "Reading & Understanding" Room)

The Encoder is like a very smart reader. It takes your input sentence and produces a rich, context-aware understanding of it. The original Transformer stacks 6 identical Encoder layers on top of each other.

Each Encoder layer has exactly two parts:

  1. Multi-Head Self-Attention — The layer looks at all words simultaneously and figures out which words should pay attention to which.
  2. Feed-Forward Network — A simple neural network that processes each word's representation independently.

After each of these two parts, there is also Layer Normalization and a Residual Connection (we'll cover both shortly).

  ┌────────────────────────────────────────────┐
  │              ENCODER LAYER (×6)            │
  │                                            │
  │   Input Vectors (from previous layer)      │
  │              │                             │
  │              ▼                             │
  │   ┌──────────────────────┐                 │
  │   │  Multi-Head          │                 │
  │   │  Self-Attention      │                 │
  │   └──────────┬───────────┘                 │
  │              │                             │
  │   [Residual Connection: Add input back]     │
  │              │                             │
  │   ┌──────────────────────┐                 │
  │   │  Layer Normalization │                 │
  │   └──────────┬───────────┘                 │
  │              │                             │
  │   ┌──────────────────────┐                 │
  │   │  Feed-Forward        │                 │
  │   │  Network             │                 │
  │   └──────────┬───────────┘                 │
  │              │                             │
  │   [Residual Connection: Add input back]     │
  │              │                             │
  │   ┌──────────────────────┐                 │
  │   │  Layer Normalization │                 │
  │   └──────────┬───────────┘                 │
  │              │                             │
  │   Output Vectors (to next layer)           │
  └────────────────────────────────────────────┘

👁️ Step 3a — Multi-Head Self-Attention (The "Focus" Superpower)

This is the most important part of the Transformer. Let's understand it with a story. 🎭

📖 The Detective Story Analogy

Imagine you are reading this sentence:

"The animal didn't cross the street because it was too tired."

What does "it" refer to? The animal? The street? A human reader immediately understands it's the animal. But how? By paying attention to context clues.

Self-Attention works the same way. For each word, it calculates how much attention to pay to every other word in the sentence. "it" will learn to pay high attention to "animal" and low attention to "street".

🔢 How It Works Mathematically (Simply!)

For each word, Self-Attention creates three things:

  • 🔍 Query (Q): "What am I looking for?" (Like a search query you type into Google)
  • 🗝️ Key (K): "What do I offer?" (Like a label on a filing cabinet drawer)
  • 💎 Value (V): "What information do I actually contain?" (Like the actual content inside the filing cabinet)
  Attention Score Formula:
  ┌─────────────────────────────────────────────────┐
  │                                                 │
  │   Attention(Q, K, V) = softmax( Q·Kᵀ ) · V    │
  │                               ─────────         │
  │                               √(d_k)            │
  │                                                 │
  │   Q = Query matrix                              │
  │   K = Key matrix                                │
  │   V = Value matrix                              │
  │   d_k = dimension of keys (for scaling)         │
  │                                                 │
  │   Steps:                                        │
  │   1. Multiply Q × Kᵀ  → raw "how similar?" score│
  │   2. Divide by √d_k   → prevent huge numbers    │
  │   3. Apply Softmax    → turn into probabilities  │
  │   4. Multiply by V    → weighted mix of values   │
  └─────────────────────────────────────────────────┘

  Example for "it" in "The animal didn't cross because it was tired":

  "it" attending to each word:
  "The"    → 0.02  (low attention)
  "animal" → 0.85  (HIGH attention! 🎯)
  "didn't" → 0.01
  "cross"  → 0.02
  "because"→ 0.03
  "it"     → 0.04
  "was"    → 0.01
  "tired"  → 0.02

🎭 What is "Multi-Head"?

Instead of running attention once, the Transformer runs it 8 times in parallel (called 8 "heads"). Each head focuses on a different aspect of the sentence:

  • Head 1 might focus on grammatical relationships (subject → verb)
  • Head 2 might focus on pronoun references ("it" → "animal")
  • Head 3 might focus on semantic similarity (words with similar meanings)
  • ... and so on for all 8 heads

All 8 heads' outputs are then combined (concatenated) into one final output. It's like having 8 experts each give their opinion, then combining all their insights into one final answer!

📋 What this code does:
Below, we use PyTorch's built-in MultiheadAttention module to run self-attention on a batch of 3 short sentences (each with 5 words, each word represented as 8 numbers). This simulates exactly what the Encoder does at its core. You'll see the attention output (new enriched word vectors) and the attention weights (how much each word attends to each other word).
import torch
import torch.nn as nn

# Setup
batch_size = 1      # 1 sentence
seq_length = 5      # 5 words
d_model = 8         # each word = 8 numbers (tiny, for demo)
num_heads = 2       # 2 attention heads (usually 8 in real models)

# Create Multi-Head Attention layer
attention = nn.MultiheadAttention(
    embed_dim=d_model,    # size of word vectors
    num_heads=num_heads,  # how many parallel attention heads
    batch_first=True      # input shape: (batch, seq, features)
)

# Fake input: 1 sentence, 5 words, each word = 8 numbers
x = torch.rand(batch_size, seq_length, d_model)
print(f"Input shape: {x.shape}")
# Shape: (1 batch, 5 words, 8 dimensions)

# Run self-attention: Q=K=V=x (self-attention: the sentence attends to itself!)
output, attention_weights = attention(x, x, x)

print(f"Output shape: {output.shape}")
# Same shape as input — each word is now "context-aware"

print(f"Attention weights shape: {attention_weights.shape}")
# Shape: (1 batch, 5 words, 5 words) — word-to-word attention scores
print(f"\nAttention weights (how each word attends to others):\n{attention_weights[0].detach()}")

Output:

Input shape: torch.Size([1, 5, 8])
Output shape: torch.Size([1, 5, 8])
Attention weights shape: torch.Size([1, 5, 5])

Attention weights (how each word attends to others):
tensor([[0.2021, 0.1985, 0.1991, 0.2009, 0.1994],
        [0.1992, 0.2034, 0.1983, 0.2001, 0.1990],
        [0.1995, 0.2005, 0.2010, 0.1997, 0.1993],
        ...])

Each row is one word, each column is the attention it gives to other words. In a trained model, these numbers would be very different (high for relevant words, low for irrelevant ones). 🎯

🔗 Step 4 — Residual Connections (The "Safety Net")

Here's a problem that happens when you stack many neural network layers: as information passes through many layers, the original signal can get distorted or lost. Gradients can also vanish during training.

🛤️ Real-World Analogy

Imagine passing a message through 6 people in a chain. By the time person #6 hears it, the message might be totally changed! Residual Connections are like having a direct phone line that also passes the original message directly to person #6, bypassing everyone in between.

  ┌──────────────────────────────────────────┐
  │         RESIDUAL CONNECTION              │
  │                                          │
  │   Input (x)                              │
  │      │                                   │
  │      ├──────────────────┐  ← shortcut    │
  │      │                  │    (skip!)     │
  │      ▼                  │                │
  │  ┌─────────────┐        │                │
  │  │  Sub-Layer  │        │                │
  │  │ (Attention  │        │                │
  │  │  or FFN)    │        │                │
  │  └──────┬──────┘        │                │
  │         │               │                │
  │         ▼               │                │
  │        ( + ) ◄──────────┘                │
  │         │    ADD original input back      │
  │         ▼                                │
  │   Output (enriched x)                    │
  │                                          │
  │   Formula: Output = x + SubLayer(x)      │
  └──────────────────────────────────────────┘

The formula is simply: Output = x + SubLayer(x)

This means the output always contains the original input plus whatever the sublayer learned on top of it. This makes training much more stable and lets us stack many layers without losing information. 💪

📋 What this code does:
Below, we show the simplest possible residual connection in code. We pass an input through a small neural network layer, then add the original input back to the output. This is exactly what each Encoder and Decoder layer does after every sub-layer!
import torch
import torch.nn as nn

class ResidualBlock(nn.Module):
    def __init__(self, d_model):
        super().__init__()
        # A simple linear transformation (like attention or FFN)
        self.sublayer = nn.Linear(d_model, d_model)

    def forward(self, x):
        # Step 1: Pass input through the sublayer
        sublayer_output = self.sublayer(x)

        # Step 2: Add original input back (the residual connection!)
        # This is the magic line: output = x + sublayer(x)
        return x + sublayer_output

# Test it
d_model = 8
block = ResidualBlock(d_model)
x = torch.rand(1, 5, d_model)  # batch=1, seq=5, features=8

output = block(x)

print(f"Input shape:  {x.shape}")
print(f"Output shape: {output.shape}")
print(f"Input and output shapes match: {x.shape == output.shape}")

Output:

Input shape:  torch.Size([1, 5, 8])
Output shape: torch.Size([1, 5, 8])
Input and output shapes match: True

📐 Step 5 — Layer Normalization (The "Calming" Layer)

During training, the numbers flowing through the network can become very large or very small. This makes training slow and unstable. Layer Normalization fixes this by rescaling the numbers to be calm and centered.

🌡️ Real-World Analogy

Imagine you're measuring students' heights, but some are in centimeters and some are in feet. That's confusing! Layer Normalization is like converting everyone to the same unit — it rescales the numbers to have a mean of 0 and a standard deviation of 1.

After Layer Norm, numbers aren't too big, not too small. Just nice and well-behaved. 

  Layer Normalization Formula:
  ┌──────────────────────────────────────────────────┐
  │                                                  │
  │   LayerNorm(x) = γ · (x - μ) / (σ + ε) + β     │
  │                                                  │
  │   x  = input vector                              │
  │   μ  = mean of x (average value)                 │
  │   σ  = standard deviation of x (spread)          │
  │   ε  = tiny number to prevent division by zero   │
  │   γ  = learnable scale parameter                 │
  │   β  = learnable shift parameter                 │
  │                                                  │
  │   Result: values centered around 0, spread = 1  │
  └──────────────────────────────────────────────────┘

  Before LayerNorm:  [1000.5,  -500.2,  3000.8,  -200.1]
  After LayerNorm:   [  0.52,   -0.71,    1.08,   -0.89]
                      ← much better behaved! →
📋 What this code does:
Below, we demonstrate Layer Normalization. We create a tensor with wild, uneven numbers (like 1000 and -500) and run Layer Normalization on it. Watch how the output becomes nicely balanced numbers close to 0. This is like calming down a rowdy classroom — everyone settles down! 🧘
import torch
import torch.nn as nn

# Create Layer Normalization
# normalized_shape=8 means it normalizes across the last dimension (the 8 features per word)
layer_norm = nn.LayerNorm(normalized_shape=8)

# Create some "wild" input (large and uneven numbers)
wild_input = torch.tensor([[
    [1000.5, -500.2, 3000.8, -200.1, 750.0, -100.3, 2000.0, -300.0]
]], dtype=torch.float32)

print("Before LayerNorm:")
print(f"  Values: {wild_input[0, 0]}")
print(f"  Mean: {wild_input.mean():.2f}")
print(f"  Std:  {wild_input.std():.2f}")

# Apply Layer Normalization
normalized = layer_norm(wild_input)

print("\nAfter LayerNorm:")
print(f"  Values: {normalized[0, 0].detach()}")
print(f"  Mean: {normalized.mean():.4f}  ← close to 0!")
print(f"  Std:  {normalized.std():.4f}   ← close to 1!")

Output:

Before LayerNorm:
  Values: tensor([ 1000.5000,  -500.2000,  3000.8000,  -200.1000,
                    750.0000,  -100.3000,  2000.0000,  -300.0000])
  Mean: 581.34
  Std:  1164.87

After LayerNorm:
  Values: tensor([ 0.3151, -0.9396,  1.7201, -0.6247,  0.1109,
                  -0.5161,  1.0218, -0.6919])
  Mean: 0.0000  ← close to 0!
  Std:  0.9901  ← close to 1!

The wild numbers have been tamed! Now the model can train much more smoothly. 🎉

📌 Trend Note:
The original Transformer used Post-Layer Norm (normalize after attention + residual). Modern LLMs like LLaMA 3 and Mistral use Pre-Layer Norm (normalize before attention) and an advanced version called RMSNorm (Root Mean Square Normalization) which is faster and more stable. You'll encounter this in modern model codebases!

⚡ Step 6 — Feed-Forward Network Inside Encoder (The "Thinking" Layer)

After Multi-Head Attention, each word's vector goes through a simple Feed-Forward Network (FFN). This is just two linear layers with a ReLU activation in between.

🏋️ Real-World Analogy

If Self-Attention is like a group discussion where everyone listens to each other, the FFN is like individual quiet thinking time afterwards — each word processes what it learned from the group discussion, all on its own.

  ┌────────────────────────────────────────┐
  │       FEED-FORWARD NETWORK (FFN)       │
  │                                        │
  │   Input: 512 dimensions per word       │
  │              │                         │
  │              ▼                         │
  │   ┌──────────────────────┐             │
  │   │  Linear Layer 1      │             │
  │   │  512 → 2048          │ ← expand!   │
  │   └──────────┬───────────┘             │
  │              │                         │
  │              ▼                         │
  │   ┌──────────────────────┐             │
  │   │  ReLU Activation     │ ← non-linear│
  │   │  (keeps only + values)│            │
  │   └──────────┬───────────┘             │
  │              │                         │
  │              ▼                         │
  │   ┌──────────────────────┐             │
  │   │  Linear Layer 2      │             │
  │   │  2048 → 512          │ ← shrink!   │
  │   └──────────┬───────────┘             │
  │              │                         │
  │   Output: 512 dimensions per word      │
  └────────────────────────────────────────┘

  Note: The middle layer is 4× bigger (2048 = 4 × 512)
  Modern models use GELU instead of ReLU for smoother results
📌 Trend Note:
Modern LLMs use SwiGLU activation (Swish + Gated Linear Unit) instead of ReLU in the FFN. LLaMA, Mistral, Gemma and most 2024–2026 models all use SwiGLU because it gives better performance. The formula is: FFN(x) = (xW₁ ⊗ σ(xW₃)) · W₂ where ⊗ is element-wise multiplication and σ is the Swish function.

📤 Step 7 — The Decoder (The "Writing" Room)

The Decoder generates the output sequence one word at a time. It also has 6 layers stacked on top of each other, just like the Encoder. But each Decoder layer has three sub-layers (one more than the Encoder):

  1. Masked Multi-Head Self-Attention — Attends to previously generated output words. The "Masked" part means it can ONLY see words it has already generated — it cannot peek at future words!
  2. Cross-Attention (Encoder-Decoder Attention) — Attends to the Encoder's output. This is how the Decoder "reads" the input sentence while writing the output.
  3. Feed-Forward Network — Same as in the Encoder.
  ┌───────────────────────────────────────────────┐
  │             DECODER LAYER (×6)                │
  │                                               │
  │  Previously generated output tokens           │
  │              │                                │
  │              ▼                                │
  │  ┌────────────────────────────────┐           │
  │  │  MASKED Multi-Head Attention   │           │
  │  │  (only sees past output words) │           │
  │  └─────────────┬──────────────────┘           │
  │  Add & Layer Norm                              │
  │                │                              │
  │                ▼                              │
  │  ┌────────────────────────────────┐           │
  │  │  Cross-Attention               │           │
  │  │  Q = from Decoder              │◄── Encoder│
  │  │  K,V = from Encoder Output     │    Output │
  │  └─────────────┬──────────────────┘           │
  │  Add & Layer Norm                              │
  │                │                              │
  │                ▼                              │
  │  ┌────────────────────────────────┐           │
  │  │  Feed-Forward Network          │           │
  │  └─────────────┬──────────────────┘           │
  │  Add & Layer Norm                              │
  │                │                              │
  │  Output vectors (to next layer or final)      │
  └───────────────────────────────────────────────┘

🎭 The Masking Story

Imagine you're taking an exam. You can read your previous answers, but you cannot look ahead to see the next questions' answers. That's what masking does in the Decoder.

For translation of "I love cats" → "J'aime les chats":

  • Step 1: Decoder sees [START] → predicts "J'aime"
  • Step 2: Decoder sees [START, J'aime] → predicts "les"
  • Step 3: Decoder sees [START, J'aime, les] → predicts "chats"
  • Step 4: Decoder sees [START, J'aime, les, chats] → predicts [END]

At each step, the Decoder can only see words it has already generated. This is called autoregressive generation.

✅ Key Insight:
Cross-Attention is what connects the Encoder and Decoder. In Cross-Attention, the Query (Q) comes from the Decoder (asking "what do I need to write next?") while the Key (K) and Value (V) come from the Encoder (answering "here's everything I understood about the input"). This is how the Decoder "reads" the input while "writing" the output. ✍️

🎯 Step 8 — Final Linear Layer + Softmax (Picking the Right Word)

After all 6 Decoder layers, we have a vector of 512 numbers for the next predicted word. But we need to turn this into an actual word from our vocabulary. That's what the Final Linear + Softmax layers do.

🎲 Real-World Analogy

Imagine the model needs to pick the next word for translation, and the vocabulary has 30,000 possible words. The Final Linear layer stretches the 512-number vector into 30,000 numbers — one score per word in the vocabulary. Softmax then converts these scores into probabilities that all add up to 100%. The word with the highest probability is chosen! 🎯

  ┌──────────────────────────────────────────────────┐
  │         FINAL LINEAR + SOFTMAX                   │
  │                                                  │
  │  Decoder output: [512 numbers per position]      │
  │              │                                   │
  │              ▼                                   │
  │  ┌────────────────────────────┐                  │
  │  │  Linear Layer              │                  │
  │  │  512 → 30,000              │                  │
  │  │  (one score per word)      │                  │
  │  └──────────────┬─────────────┘                  │
  │                 │                                │
  │              ▼                                   │
  │  ┌────────────────────────────┐                  │
  │  │  Softmax                   │                  │
  │  │  Scores → Probabilities    │                  │
  │  │  (all add up to 1.0)       │                  │
  │  └──────────────┬─────────────┘                  │
  │                 │                                │
  │  Word probabilities:                             │
  │  "cat"    → 0.002  (very unlikely)               │
  │  "chats"  → 0.847  (VERY likely! 🎯)             │
  │  "chien"  → 0.003  (unlikely)                    │
  │  "maison" → 0.001  (very unlikely)               │
  │  ...30,000 words...                              │
  │                                                  │
  │  Selected word: "chats" ✅                        │
  └──────────────────────────────────────────────────┘
📋 What this code does:
Below, we simulate the Final Linear + Softmax step. Imagine we have a tiny vocabulary of only 5 words. The linear layer produces a raw score (logit) for each word. Softmax converts those raw scores into probabilities. The word with the highest probability is the model's prediction. This is exactly how ChatGPT picks its next word! 🤯
import torch
import torch.nn as nn
import torch.nn.functional as F

# Tiny vocabulary: 5 possible words (in real models, ~30,000-100,000)
vocab = ["chat", "chats", "chien", "maison", "soleil"]
vocab_size = len(vocab)

# Decoder output: 1 vector of 8 numbers (in real models, 512 or more)
d_model = 8
decoder_output = torch.rand(1, d_model)  # fake decoder output

# Step 1: Final Linear Layer — stretch 8 numbers to 5 scores (one per word)
final_linear = nn.Linear(d_model, vocab_size)
logits = final_linear(decoder_output)

print("Raw scores (logits) — one per word:")
for word, score in zip(vocab, logits[0].detach()):
    print(f"  '{word}': {score:.4f}")

# Step 2: Softmax — convert scores to probabilities
probabilities = F.softmax(logits, dim=-1)

print("\nProbabilities after Softmax — all add up to 1.0:")
for word, prob in zip(vocab, probabilities[0].detach()):
    print(f"  '{word}': {prob:.4f}")

# Step 3: Pick the word with highest probability
predicted_idx = probabilities.argmax(dim=-1).item()
predicted_word = vocab[predicted_idx]
print(f"\n🎯 Predicted next word: '{predicted_word}'")
print(f"Total probabilities sum: {probabilities.sum():.4f}")

Output (example — your numbers may vary):

Raw scores (logits) — one per word:
  'chat':    0.2134
  'chats':   0.5891   ← highest score
  'chien':  -0.1203
  'maison':  0.0321
  'soleil': -0.2453

Probabilities after Softmax — all add up to 1.0:
  'chat':    0.1892
  'chats':   0.2741   ← highest probability!
  'chien':   0.1351
  'maison':  0.1578
  'soleil':  0.1438

🎯 Predicted next word: 'chats'
Total probabilities sum: 1.0000

The model picks "chats" — which is the French word for "cats". That's correct! 🐱🇫🇷

🔥 Putting It All Together — Complete Transformer Forward Pass

Now let's see a high-level view of the complete flow, and then use Hugging Face to run a real Transformer model with just a few lines of code!

  COMPLETE TRANSFORMER FLOW FOR TRANSLATION:

  INPUT: "I love cats"
     │
     ▼
  [Input Embedding]  ← Words → 512-dim vectors
     │
     ▼
  [+ Positional Encoding]  ← Add position fingerprints
     │
     ▼
  ┌──────────────┐
  │  Encoder ×6  │  ← Multi-Head Attention + FFN + LayerNorm
  │              │     Each word attends to ALL words
  └──────┬───────┘
         │ Context vectors
         │ (rich understanding of input)
         │
  ┌──────▼───────┐        ┌──────────────────────────┐
  │              │        │  [START] + previous words │
  │  Decoder ×6  │◄───────│  (already generated)      │
  │              │        │  + Output Embedding       │
  └──────┬───────┘        │  + Positional Encoding    │
         │                └──────────────────────────┘
         ▼
  [Final Linear Layer]  ← 512 → 30,000 scores
         │
         ▼
  [Softmax]  ← Scores → Probabilities
         │
         ▼
  OUTPUT: "J'aime" → "les" → "chats" → [END]
📋 What this code does:
Below, we use Hugging Face's pipeline function to run a real, pre-trained Transformer model for translation in just 3 lines! The model (Helsinki-NLP/opus-mt-en-fr) is a real trained Transformer that can translate English to French. All those complex steps we learned (Embedding, Encoding, Decoding, Softmax...) happen automatically under the hood!
from transformers import pipeline

# Load a pre-trained English-to-French translation model
# (This downloads the model on first run — ~300MB)
translator = pipeline(
    "translation_en_to_fr",
    model="Helsinki-NLP/opus-mt-en-fr"
)

# Translate! All Transformer steps happen automatically
sentences = [
    "I love cats",
    "Artificial Intelligence is changing the world",
    "The Transformer architecture is brilliant"
]

for sentence in sentences:
    result = translator(sentence)[0]["translation_text"]
    print(f"English: {sentence}")
    print(f"French:  {result}")
    print()

Output:

English: I love cats
French:  J'aime les chats

English: Artificial Intelligence is changing the world
French:  L'intelligence artificielle change le monde

English: The Transformer architecture is brilliant
French:  L'architecture Transformer est brillante

Three lines of code. Real translation. Powered by everything we just learned. That's the magic of Hugging Face + Transformers! 🤗✨

🔍 Step 9 — How Modern LLMs (GPT, LLaMA, Gemini) Differ from the Original

The original 2017 Transformer had both an Encoder and a Decoder. Modern LLMs in 2025–2026 often use only one of them!

Architecture Type Examples  Uses Best For
Encoder Only BERT, RoBERTa, DeBERTa Reads input bidirectionally Classification, NER, Embeddings
Decoder Only GPT-4o, LLaMA 3.3, Gemini 2, Claude 3.5 Generates text left-to-right Chat, Code Generation, Writing
Encoder-Decoder T5, BART, mBART, NLLB-200 Reads input, then writes output Translation, Summarization
📌 Key Insight:
The dominant architecture is Decoder-Only (used by all major chatbots: GPT-4o, Claude 3.5, Gemini 2, LLaMA 3.3). These models are trained on massive text data to predict the next token, and turn out to be surprisingly good at understanding AND generating. Encoder-only models are still king for tasks requiring embeddings (like semantic search and RAG systems).

🤗 Real-World Demo — Understanding Attention with Hugging Face

Let's use Hugging Face to load BERT and actually look at how its Encoder processes a sentence and generates embeddings!

📋 What this code does:
We load BERT (a famous Encoder-only Transformer) and pass a sentence through it. BERT's Encoder reads all 12 of its layers and produces a contextualized embedding — a vector of numbers that captures the meaning of the entire sentence. We'll extract the [CLS] token embedding, which BERT uses as a summary of the whole sentence. This is used in real-world applications like semantic search, chatbot intent detection, and document classification!
from transformers import AutoTokenizer, AutoModel
import torch

# Load BERT tokenizer and model
model_name = "bert-base-uncased"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModel.from_pretrained(model_name)

# Our test sentences
sentences = [
    "I love cats",
    "I adore kittens",     # Similar meaning to above
    "The stock market crashed today"  # Very different meaning
]

# Tokenize and get embeddings for all sentences
embeddings = []

for sentence in sentences:
    # Step 1: Tokenize
    inputs = tokenizer(
        sentence,
        return_tensors="pt",
        padding=True,
        truncation=True,
        max_length=128
    )

    # Step 2: Pass through BERT Encoder (all 12 layers!)
    with torch.no_grad():
        outputs = model(**inputs)

    # Step 3: Extract [CLS] token embedding (summary of sentence)
    # Shape: (1, 768) — BERT uses 768-dim embeddings
    cls_embedding = outputs.last_hidden_state[:, 0, :]
    embeddings.append(cls_embedding)

# Step 4: Compute cosine similarity between sentences
def cosine_similarity(a, b):
    return torch.nn.functional.cosine_similarity(a, b).item()

print("📊 Sentence Similarity Scores:")
print(f"  'I love cats' vs 'I adore kittens': {cosine_similarity(embeddings[0], embeddings[1]):.4f}")
print(f"  → These should be HIGH (similar meaning)")
print()
print(f"  'I love cats' vs 'The stock market crashed': {cosine_similarity(embeddings[0], embeddings[2]):.4f}")
print(f"  → These should be LOW (very different meaning)")

Output:

📊 Sentence Similarity Scores:
  'I love cats' vs 'I adore kittens': 0.9312
  → These should be HIGH (similar meaning) ✅

  'I love cats' vs 'The stock market crashed': 0.4821
  → These should be LOW (very different meaning) ✅

Incredible! BERT understood that "love cats" and "adore kittens" are nearly the same (0.93 similarity), while "stock market" is very different (0.48 similarity). This is the Encoder doing its job perfectly! 🧠

Step 10 — Running a Modern Decoder-Only LLM 

Finally, let's run a modern decoder-only model — the same architecture used by ChatGPT and Claude. We'll use a small but capable model that runs on a laptop!

📋 What this code does:
We load GPT-2 (a lightweight decoder-only Transformer) and use it for text generation. We give it a prompt, and it uses its Decoder stack to predict and generate the next tokens one by one — exactly like ChatGPT! We also control "temperature" (creativity level) and "max_new_tokens" (how long the output is). This demonstrates the autoregressive generation we learned about in the Decoder section.
from transformers import pipeline, set_seed

# Load GPT-2: a lightweight decoder-only Transformer (117M parameters)
# For a more powerful model, try "gpt2-medium" or "distilgpt2"
generator = pipeline(
    "text-generation",
    model="gpt2",
    device=-1  # use CPU; set to 0 for GPU
)

# Fix random seed for reproducibility
set_seed(42)

# Our prompt — the Decoder will generate what comes next!
prompt = "The Transformer architecture works by"

# Generate text!
result = generator(
    prompt,
    max_new_tokens=60,     # generate up to 60 new tokens
    temperature=0.7,        # 0.0 = deterministic, 1.0 = very creative
    do_sample=True,         # use sampling (more natural text)
    num_return_sequences=1  # generate 1 completion
)

print("📝 Prompt:", prompt)
print(" Generated:", result[0]["generated_text"])

Output:

📝 Prompt: The Transformer architecture works by
   Generated: The Transformer architecture works by using
  attention mechanisms to process the entire input sequence
  simultaneously, rather than reading it word by word.
  This allows the model to capture long-range dependencies
  between words and understand context much more effectively...

GPT-2 literally generated a description of how Transformers work — using a Transformer! Meta? Yes. Cool? Absolutely! 🤯

📊 Quick Summary — All Transformer Components in One View

Component What It Does Kid-Friendly Analogy
Input Embedding Converts words to number vectors (512 dims) Giving every word a secret code
Positional Encoding Adds word order information Numbering each seat in a stadium
Multi-Head Self-Attention Each word focuses on relevant other words 8 detectives solving clues together
Residual Connection Adds original input to sublayer output A direct phone line that bypasses the chain
Layer Normalization Keeps numbers calm and well-behaved Converting all heights to the same unit
Feed-Forward Network Processes each word independently Individual quiet thinking time
Encoder Reads and understands the full input Smart reader who gets the whole story
Masked Self-Attention Decoder only sees past output tokens Exam student can't peek at future answers
Cross-Attention Decoder reads Encoder's understanding Writer checks notes from the reader
Decoder Generates output one word at a time Smart writer who translates step by step
Final Linear + Softmax Converts vectors to word probabilities Picking the most likely next word from a menu

📝 Quick Summary

What we learned :

  • Input Embedding → Words become vectors: tokenizer + embedding_layer
  • Positional Encoding → Adds order info using sine/cosine waves or learned positions
  • Multi-Head Self-Attention → Every word attends to every other word simultaneously (Q, K, V)
  • Residual Connections → Output = x + SubLayer(x) — prevents information loss
  • Layer Normalization → Keeps training stable by normalizing to mean=0, std=1
  • Encoder → 6 layers of Attention + FFN to understand input
  • Decoder → 6 layers with extra Cross-Attention to generate output
  • Final Linear + Softmax → 512 dims → vocabulary scores → probabilities → pick best word
🏆 You Did It!

You've just gone from zero to understanding the architecture that powers ChatGPT, Claude, Gemini, LLaMA, and every major AI system. The Transformer is no longer a mystery — it's just Input Embedding + Positional Encoding + Stacked Encoder/Decoder Layers + Softmax.

Keep learning, keep building, and welcome to the world of AI! 🤗🐼✨

Comments