Here's a scenario: you're a swimming coach 🏊 working with a talented but untrained athlete. They can swim — but their technique is rough and their turns are terrible. You don't throw out everything they know and start from scratch. Instead, every day you show them where they went wrong, celebrate what went right, and make small, deliberate corrections. Over weeks, they go from sloppy to Olympic-level.
Proximal Policy Optimization (PPO) is the algorithm that does exactly this for language models. It's how ChatGPT learned to be helpful instead of just coherent. It's how AlphaGo became a Go grandmaster. It's how robots learn to walk without anyone coding each step.
Part 1 — The Foundation: Reinforcement Learning
What Is Reinforcement Learning?
Reinforcement Learning (RL) is a type of machine learning where an agent learns by interacting with an environment and receiving feedback in the form of rewards. There are no labelled examples — no "here is the correct answer." The agent simply tries things, gets scored, and gradually learns to do better.
💡 Think of it like: Learning to cook 👨🍳 without a recipe. You make a dish, your family rates it from 1 to 10, you adjust the seasoning next time. You never knew the "correct" recipe upfront — you discovered it through feedback.
The Five RL Building Blocks
Every RL system is described with the same five concepts. Learn these once and you understand all RL:
- Agent 🤖 → The decision-maker. In gaming RL, it's the game-playing bot. In LLM alignment, it's the language model generating text token by token.
- Environment 🌍 → Everything the agent interacts with. In chess, it's the board. In LLM alignment, it's the combination of user prompts and the reward model that evaluates responses.
- State s 👁️ → What the agent currently observes. In chess, it's the current board position. In LLMs, it's the sequence of tokens generated so far — the "context" for the next token decision.
- Action a ⚡ → A decision the agent makes. In chess, it's moving a piece. In LLMs, it's selecting the next token from the vocabulary. A complete response is a sequence of thousands of action decisions.
- Reward r 🏆 → A scalar score the agent receives after taking an action. In chess, +1 for winning, -1 for losing. In LLMs, the reward model's score for the complete generated response.
The Policy — The Brain Behind Every Decision
The policy π(a|s) is the agent's strategy — a function that maps any state to a probability distribution over actions. Given the current conversation so far (state), the policy outputs a probability for every token in the vocabulary. The agent samples from this distribution to choose the next word.
The goal of RL: find the policy that maximises the total expected reward over time. PPO is one of the best algorithms ever invented for doing exactly that.
🔄 The RL Interaction Loop
🤖
Agent
Observes state s
⚡
Action
Takes action a
via policy π(a|s)
🌍
Environment
Returns new state s'
and reward r
📈
Update
Policy learns to
maximise reward
This loop repeats thousands or millions of times. Each iteration makes the policy slightly smarter.
Part 2 — Why PPO? The Problem with Naive Policy Gradient
The Simplest RL Update — Policy Gradient
The most basic RL update rule is the REINFORCE algorithm (also called vanilla policy gradient). The idea is beautifully simple: if an action led to a high reward, make that action more probable. If it led to a low reward, make it less probable.
Policy Gradient Update:
θ ← θ + α × ∇log π(a|s) × R
Where:
θ = model parameters
α = learning rate
∇log = gradient of log probability of the action taken
R = reward received
Simple, elegant — and catastrophically unstable in practice. Here's why:
The Three Problems with Vanilla Policy Gradient
- Problem 1 — Destructively Large Updates → A single lucky trajectory (a sequence of actions that got high reward by chance) can cause a massive policy update. The policy swings wildly toward that lucky behaviour, destroys everything it previously knew, and collapses. This is called catastrophic forgetting in RL.
- Problem 2 — No Re-use of Data → Each rollout (sequence of actions) can only be used for exactly one gradient update. After that, the policy has changed, and the old data is no longer valid. This is incredibly wasteful — generating experience is expensive (especially for LLMs).
- Problem 3 — High Variance → Rewards are noisy. A good response might score 2.1 one time and 2.8 another time from the same reward model. This noise makes gradient estimates unreliable.
💡 The swimming coach analogy again: Imagine a coach who, after watching ONE bad race, tears up the entire training programme and builds a completely new technique from scratch. The athlete would be completely lost. A good coach makes small, targeted corrections while preserving what's already working. That's what PPO does.
Part 3 — PPO's Core Innovation: The Clipped Surrogate Objective
The Probability Ratio — Measuring How Much the Policy Changed
PPO's genius starts with a simple ratio. When you update a policy, each action's probability changes. Define the ratio of new probability to old probability:
r_t(θ) = π_θ(aₜ | sₜ) ← New policy probability
─────────────────
π_θ_old(aₜ | sₜ) ← Old policy probability (before this update)
- r = 1.0 → Policy hasn't changed for this action at all
- r = 1.5 → New policy is 50% more likely to pick this action
- r = 0.6 → New policy is 40% less likely to pick this action
This ratio is the heart of PPO. By monitoring it during training, PPO can detect when an update is getting too aggressive — and put on the brakes.
The Advantage Function — Was This a Good Decision?
Before deciding how to update the policy, we need to answer one question per action: "Was this action better or worse than what we'd normally expect?"
This is measured by the Advantage function A_t:
A_t = Q(s_t, a_t) - V(s_t)
Where:
Q(s_t, a_t) = Total reward actually received from this action onwards
V(s_t) = Total reward expected on average from this state
A_t = How much better (or worse) this action was vs. the average
💡 Basketball analogy 🏀: Your team averages 90 points per game (V). Tonight you scored 108 (Q). Your advantage was +18 — you outperformed expectations significantly. PPO uses this signal to say "whatever you did tonight, do more of it."
- A > 0 → Action was better than expected → increase its probability
- A < 0 → Action was worse than expected → decrease its probability
- A ≈ 0 → Action was average → no significant update needed
The Clipped Objective — PPO's Safety Mechanism
Now combine the ratio and the advantage.
The naive policy gradient would be: r_t × A_t.
If A_t is large and positive, and r_t is large (we've already shifted probability toward this action),
the gradient explodes.
PPO's fix: clip the ratio so it can never get too far from 1.0.
L_CLIP(θ) = E_t [ min(
r_t(θ) × A_t,
clip(r_t(θ), 1 - ε, 1 + ε) × A_t
) ]
Where ε (epsilon) = 0.2 by default
This means ratios are clamped to the range [0.8, 1.2]
The min() function is the clever part.
It always takes the pessimistic view:
- Good action (A > 0): We want to increase probability — but only up to 1.2× the old probability. If r already exceeded 1.2, the gradient is zero: "we've already boosted this enough, stop."
- Bad action (A < 0): We want to decrease probability — but only down to 0.8× the old probability. If r is already below 0.8, the gradient is zero: "we've already suppressed this enough, stop."
✂️ PPO Clipping — What It Prevents
Without Clipping
r = 3.0, A = +2.0 → gradient = 6.0 → policy EXPLODES toward this action.
All previous behaviour destroyed. 💥
With Clipping (ε=0.2)
r = 3.0 → clipped to 1.2, A = +2.0 → gradient = 2.4 → safe, measured update.
Previous knowledge preserved. ✅
Part 4 — The Full PPO Loss Function
PPO doesn't just use the clipped policy gradient. The full PPO training objective has three components:
L_PPO(θ) = L_CLIP(θ) ← Policy improvement
− c₁ × L_VF(θ) ← Value function accuracy
+ c₂ × H(π_θ) ← Entropy bonus (exploration)
Typical values: c₁ = 0.5, c₂ = 0.01
Component 1 — L_CLIP: Policy Gradient with Clipping
Already explained above. This is the primary objective — make good actions more likely and bad actions less likely, but never change the policy too drastically in one step.
Component 2 — L_VF: Value Function Loss
PPO trains a value function (critic) alongside the policy (actor). The critic predicts the expected total reward from any given state. This prediction is used to compute the advantage.
If the critic is inaccurate, the advantage estimates are wrong, and the policy learns from bad signals. So the value loss trains the critic to be accurate:
L_VF(θ) = MSE(V_θ(s_t), R_t)
Where:
V_θ(s_t) = Critic's predicted reward from state s_t
R_t = Actual reward received from state s_t onwards
Component 3 — Entropy Bonus H(π): Encouraging Exploration
Entropy measures how spread out a probability distribution is. High entropy = many tokens have similar probabilities → diverse, exploratory responses. Low entropy = one token dominates → deterministic, repetitive responses.
Without the entropy bonus, PPO can collapse the policy to always produce the same response — the one that historically got the highest reward. The entropy bonus rewards the policy for staying diverse and exploratory.
Part 5 — Generalised Advantage Estimation (GAE)
The advantage function A_t has to be estimated from sampled data — it's never exact. The quality of that estimate dramatically affects training stability.
GAE (Generalised Advantage Estimation) provides the best available estimate using a weighted average of multi-step "temporal difference errors" (TD errors):
δ_t = r_t + γ × V(s_{t+1}) - V(s_t) ← TD error at step t
A_GAE_t = δ_t + (γλ) × δ_{t+1} + (γλ)² × δ_{t+2} + ...
Where:
γ = discount factor (how much we care about future vs immediate reward)
λ = GAE smoothing parameter (bias-variance tradeoff)
The lambda parameter gives you control over the bias-variance tradeoff:
- λ = 0 → Only look one step ahead (low variance, high bias). Fast but potentially misleading — doesn't account for long-term consequences of actions.
- λ = 1 → Look all the way to the end of the trajectory (low bias, high variance). More accurate but noisy — a single bad reward at the end corrupts all estimates.
- λ = 0.95 → The practical sweet spot. Balances immediate and delayed credit assignment. The default for virtually all LLM PPO training runs.
γ = 1.0 (no discounting).
LLM responses are short bounded episodes where all tokens contribute to the same final reward.
Discounting would unfairly devalue tokens at the beginning of the response.
Set λ = 0.95 as your starting point for GAE.
Part 6 — PPO in the LLM RLHF Pipeline
For language model alignment, PPO doesn't run alone. It's the final stage in the RLHF (Reinforcement Learning from Human Feedback) pipeline — the same process used to train ChatGPT, Gemini, and Claude.
🏗️ The Complete RLHF + PPO Pipeline
Stage 1 — Pre-training 🌍
Train on massive internet text. Model learns language, facts, and reasoning. Produces a raw, unaligned base model. No human preferences yet.
Stage 2 — Supervised Fine-Tuning (SFT) 🎓
Fine-tune on curated instruction-response pairs. Model learns to follow instructions and generate helpful outputs. This SFT model becomes the starting point for PPO.
Stage 3 — Reward Model Training 🏅
Collect human preference data: show humans pairs of responses, record which they prefer. Train a reward model that scores any (prompt, response) pair. This model acts as a substitute human evaluator during PPO.
Stage 4 — PPO Alignment Training 🎯
Use PPO to fine-tune the SFT model to maximise the reward model's scores. A KL penalty keeps the model from drifting too far from the SFT baseline. This is where PPO runs.
The Four Models Active During PPO Training
PPO for LLMs requires four separate models loaded simultaneously. This is why RLHF is memory-intensive and why LoRA + quantization are essential:
- Policy Model (Actor) 🤖 → The model being trained. Generates responses. Weights are updated by PPO. Starts as a copy of the SFT model.
- Value Model (Critic) 📊 → Predicts expected total reward from each token position. Used to compute advantage estimates. Also updated during PPO. Often shares layers with the policy model.
- Reward Model 🏅 → Scores the complete response. Frozen — weights never change during PPO. Is the "teacher" providing the training signal.
- Reference Model ❄️ → A frozen copy of the original SFT model. Used to compute the KL divergence penalty — prevents the policy from reward-hacking. Never updated. The "sanity anchor."
🏗️ Four Models in PPO Training — Memory Layout
🤖
Policy Model
TRAINABLE ✏️
~6 GB (4-bit)
📊
Value Model
TRAINABLE ✏️
~6 GB (4-bit)
🏅
Reward Model
FROZEN ❄️
~2 GB
🧊
Reference Model
FROZEN ❄️
~6 GB (4-bit)
Total VRAM for 7B models: ~20 GB with 4-bit quantization. Recommended: A100 40 GB or two RTX 4090s.
Part 7 — The KL Divergence Penalty
What Is KL Divergence?
KL divergence measures how different two probability distributions are. In PPO for LLMs, we measure how different the current policy is from the frozen reference model (SFT model):
KL(π_policy || π_ref) = Σ_tokens π_policy(token) × log( π_policy(token) / π_ref(token) )
Result: A non-negative number.
KL = 0 → Policy is identical to reference (no change at all)
KL = 5 → Moderate drift — policy has changed noticeably
KL = 20+ → Severe drift — policy behaving very differently from original SFT model
How KL Is Used in PPO
The KL penalty is subtracted from the reward before PPO uses it. This means the policy is rewarded for high-quality responses, but also penalised for drifting too far from the reference model:
Adjusted Reward = Reward_Model_Score − β × KL(π_policy || π_ref)
Where β (KL coefficient) controls the strength of the penalty:
β = 0.1 → Mild constraint — policy can explore widely
β = 0.5 → Strong constraint — policy stays close to reference
💡 Why does this prevent reward hacking? Without KL penalty, the policy might discover that certain nonsensical phrases always trick the reward model into giving high scores. The KL penalty says: "even if that phrase scores well, if it's not something the reference model would say, you're penalised for using it." This keeps the model grounded in natural language.
Part 8 — Complete PPO Implementation with TRL
Step 1 — Install Libraries
pip install trl transformers peft accelerate bitsandbytes datasets wandb
What each library does:
- trl → Transformer Reinforcement Learning — implements PPOTrainer, reward computation, rollout generation
- peft → Parameter-Efficient Fine-Tuning — provides LoRA so we don't train all 7B parameters
- bitsandbytes → Enables 4-bit NF4 quantization to fit the model in limited VRAM
- accelerate → Handles multi-GPU distribution and mixed precision automatically
- wandb → Weights & Biases — for tracking all PPO metrics across training runs
Step 2 — Load the Policy Model with LoRA
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
from peft import LoraConfig, get_peft_model, prepare_model_for_kbit_training, TaskType
# Base SFT model — this is what PPO will improve
MODEL_NAME = "mistralai/Mistral-7B-Instruct-v0.2"
# 4-bit NF4 quantization — load the 7B model in ~4GB instead of ~14GB
bnb_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.bfloat16,
bnb_4bit_use_double_quant=True
)
# ── Load policy model ─────────────────────────────────────────
policy_model = AutoModelForCausalLM.from_pretrained(
MODEL_NAME,
quantization_config=bnb_config,
device_map="auto"
)
# Prepare for 4-bit training (enables gradient checkpointing, normalises layer dtypes)
policy_model = prepare_model_for_kbit_training(policy_model)
# ── Attach LoRA adapters ──────────────────────────────────────
# Only ~84M parameters will be trained instead of 7.2B
lora_config = LoraConfig(
r=16,
lora_alpha=32,
target_modules=[
"q_proj", "k_proj", "v_proj", "o_proj",
"gate_proj", "up_proj", "down_proj"
],
lora_dropout=0.05,
bias="none",
task_type=TaskType.CAUSAL_LM
)
policy_model = get_peft_model(policy_model, lora_config)
policy_model.print_trainable_parameters()
# ── Load tokenizer ────────────────────────────────────────────
tokenizer = AutoTokenizer.from_pretrained(MODEL_NAME)
tokenizer.pad_token = tokenizer.eos_token
tokenizer.padding_side = "left" # CRITICAL: left padding for RL generation
print(f"Policy model memory: {policy_model.get_memory_footprint()/1e9:.2f} GB")
print("Policy model ready! ✅")
Output:
trainable params: 83,886,080 || all params: 3,827,765,248 || trainable%: 2.19%
Policy model memory: 4.32 GB
Policy model ready! ✅
Step 3 — Set Up the Reward Model
from transformers import AutoModelForSequenceClassification
# Load a pre-trained public reward model
# This model outputs a helpfulness score for any (prompt, response) pair
REWARD_MODEL_NAME = "OpenAssistant/reward-model-deberta-v3-large-v2"
reward_model = AutoModelForSequenceClassification.from_pretrained(
REWARD_MODEL_NAME,
torch_dtype=torch.float16,
device_map="auto"
)
reward_tokenizer = AutoTokenizer.from_pretrained(REWARD_MODEL_NAME)
# Freeze the reward model — it should NEVER be updated during PPO
for param in reward_model.parameters():
param.requires_grad = False
print(f"Reward model loaded and frozen ❄️")
print(f"Reward model parameters: {sum(p.numel() for p in reward_model.parameters()):,}")
def compute_reward_scores(prompts: list, responses: list) -> list:
"""
Compute scalar reward scores for a batch of (prompt, response) pairs.
Returns a list of float tensors, one per sample.
Higher score = better response quality.
"""
scores = []
for prompt, response in zip(prompts, responses):
text = f"Human: {prompt}\n\nAssistant: {response}"
inputs = reward_tokenizer(
text,
return_tensors="pt",
truncation=True,
max_length=512,
padding=True
).to(reward_model.device)
with torch.no_grad():
output = reward_model(**inputs)
score = output.logits[0][0].cpu() # scalar reward score
scores.append(score)
return scores
# Test the reward function
test_rewards = compute_reward_scores(
prompts=["What causes rainbows?", "How do I improve my sleep?"],
responses=[
"Rainbows form when sunlight passes through water droplets in the air. The droplets
act like tiny prisms, splitting white light into its component colours — red, orange,
yellow, green, blue, indigo, violet.",
"Sleep can vary. Some people sleep more. It depends on many things honestly."
]
)
print(f"Reward for good response: {test_rewards[0].item():.4f}")
print(f"Reward for vague response: {test_rewards[1].item():.4f}")
Output:
Reward model loaded and frozen ❄️
Reward model parameters: 183,831,554
Reward for good response: 2.8341
Reward for vague response: -0.9823
The reward model correctly prefers the clear, detailed response! 🎯
Step 4 — Prepare the Training Prompts
from datasets import load_dataset
from torch.utils.data import DataLoader
def prepare_prompt_dataset(dataset_name: str, n_samples: int = 3000) -> list:
"""
Load a prompt dataset and format prompts for PPO training.
Returns tokenised prompt tensors ready for generation.
"""
raw_dataset = load_dataset(dataset_name, split=f"train[:{n_samples}]")
# Extract just the prompt (no response — PPO will generate those)
prompt_texts = []
for item in raw_dataset:
# Format using the model's chat template
if "chosen" in item:
# HH-RLHF format: extract only the human turn
text = item["chosen"].split("\n\nAssistant:")[0].replace("Human: ", "").strip()
else:
text = item.get("prompt", item.get("instruction", ""))
if text and len(text.split()) > 5: # filter out very short prompts
prompt_texts.append(text[:512]) # truncate very long prompts
# Tokenise all prompts
tokenised = tokenizer(
prompt_texts,
return_tensors="pt",
truncation=True,
max_length=256,
padding=True
)
print(f"Prepared {len(prompt_texts)} training prompts")
print(f"Sample prompt: '{prompt_texts[0][:80]}...'")
return prompt_texts, tokenised
prompt_texts, prompt_tokens = prepare_prompt_dataset("Anthropic/hh-rlhf", n_samples=3000)
Output:
Prepared 2841 training prompts
Sample prompt: 'What are some good tips for improving my productivity at work?...'
Step 5 — Configure and Run PPO
from trl import PPOTrainer, PPOConfig
from trl import AutoModelForCausalLMWithValueHead
# ── Reload model with value head ──────────────────────────────
# TRL's AutoModelForCausalLMWithValueHead adds a linear value prediction
# head on top of the language model — this is our Critic
model_with_value = AutoModelForCausalLMWithValueHead.from_pretrained(
policy_model, # pass our LoRA-wrapped model
torch_dtype=torch.bfloat16
)
# ── PPO Configuration ─────────────────────────────────────────
ppo_config = PPOConfig(
# Core training
model_name=MODEL_NAME,
learning_rate=1.4e-5, # Low LR — PPO needs small, careful steps
batch_size=32, # Prompts per PPO iteration
mini_batch_size=4, # Mini-batch size for K gradient updates
ppo_epochs=4, # K epochs — reuse each rollout 4 times
gradient_accumulation_steps=2,
# KL penalty — prevents reward hacking
init_kl_coef=0.2, # Starting β
target=6.0, # Target KL (nats) — PPO auto-adjusts β to hit this
horizon=10000,
# Advantage estimation
gamma=1.0, # No discounting for LLM episodes
lam=0.95, # GAE lambda
# Loss coefficients
vf_coef=0.1, # c₁ — value function loss weight
cliprange=0.2, # ε — the clipping boundary
cliprange_value=0.2, # value function clipping range
# Optimiser
max_grad_norm=0.5,
adap_kl_ctrl=True, # Automatically adapt β to hit target KL
# Tracking
log_with="wandb",
project_kwargs={"project": "ppo-mistral-rlhf"},
# Output
output_dir="./ppo-training-output",
save_freq=50
)
# ── Initialise Trainer ────────────────────────────────────────
ppo_trainer = PPOTrainer(
config=ppo_config,
model=model_with_value,
ref_model=None, # TRL auto-creates a frozen copy of the initial model as reference
tokenizer=tokenizer
)
print("PPO Trainer initialised! 🎯")
print(f"Reference model is frozen: {all(not p.requires_grad for p in ppo_trainer.ref_model.parameters())}")
Output:
PPO Trainer initialised! 🎯
Reference model is frozen: True
Step 6 — The PPO Training Loop
from tqdm import tqdm
import random
# Generation settings for rollouts
GEN_KWARGS = {
"max_new_tokens": 200,
"temperature": 0.8, # some randomness for exploration
"top_p": 0.9,
"do_sample": True,
"pad_token_id": tokenizer.eos_token_id,
"eos_token_id": tokenizer.eos_token_id
}
NUM_ITERATIONS = 500
all_stats = []
reward_history = []
print("Starting PPO training... 🏋️")
for iteration in tqdm(range(NUM_ITERATIONS), desc="PPO Iterations"):
# ── 1. Sample a random batch of prompts ──────────────────
batch_indices = random.sample(range(len(prompt_texts)), ppo_config.batch_size)
batch_prompts = [prompt_texts[i] for i in batch_indices]
# Tokenise this batch
query_tensors = [
tokenizer.encode(p, return_tensors="pt", truncation=True, max_length=256)[0]
for p in batch_prompts
]
# ── 2. Generate responses (rollout phase) ─────────────────
response_tensors = ppo_trainer.generate(
query_tensors,
return_prompt=False, # only return the generated tokens, not the prompt
**GEN_KWARGS
)
# ── 3. Decode responses ───────────────────────────────────
decoded_responses = [
tokenizer.decode(r, skip_special_tokens=True)
for r in response_tensors
]
# ── 4. Score with reward model ────────────────────────────
reward_scores = compute_reward_scores(batch_prompts, decoded_responses)
reward_history.append(sum(r.item() for r in reward_scores) / len(reward_scores))
# ── 5. PPO gradient update ────────────────────────────────
# This internally:
# a) Computes advantages using GAE
# b) Normalises advantages across the batch
# c) Runs K=4 gradient steps with the clipped PPO objective
# d) Computes KL penalty and adjusts β
stats = ppo_trainer.step(query_tensors, response_tensors, reward_scores)
all_stats.append(stats)
# ── 6. Log progress every 25 iterations ───────────────────
if iteration % 25 == 0:
avg_reward = reward_history[-1]
kl = stats.get("objective/kl", 0)
policy_loss = stats.get("ppo/loss/policy", 0)
value_loss = stats.get("ppo/loss/value", 0)
entropy = stats.get("objective/entropy", 0)
clip_frac = stats.get("objective/clipfrac", 0)
kl_coef = stats.get("objective/kl_coef", 0)
print(f"\n--- Iteration {iteration:4d} ---")
print(f" Avg Reward: {avg_reward:+.4f} {'📈' if avg_reward > 0 else '📉'}")
print(f" KL Divergence:{kl:.4f} {'✅' if kl < 10 else '⚠️'}")
print(f" KL Coef (β): {kl_coef:.4f}")
print(f" Policy Loss: {policy_loss:.4f}")
print(f" Value Loss: {value_loss:.4f}")
print(f" Entropy: {entropy:.4f}")
print(f" Clip Fraction:{clip_frac:.4f} {'✅' if 0.1 < clip_frac < 0.4 else '⚠️'}")
# Show a sample generated response
print(f"\n Sample prompt: {batch_prompts[0][:60]}...")
print(f" Sample response: {decoded_responses[0][:120]}...")
print("\nPPO training complete! 🎉")
ppo_trainer.save_pretrained("./ppo-final-adapter")
Output (sample iterations):
--- Iteration 0 ---
Avg Reward: +0.1203 📈
KL Divergence:0.0012 ✅
KL Coef (β): 0.2000
Policy Loss: 0.0312
Value Loss: 3.4821
Entropy: 6.2341
Clip Fraction:0.0891 ⚠️
--- Iteration 25 ---
Avg Reward: +0.6821 📈
KL Divergence:2.3412 ✅
KL Coef (β): 0.1923
Policy Loss: 0.0089
Value Loss: 1.2103
Entropy: 5.8901
Clip Fraction:0.1923 ✅
--- Iteration 100 ---
Avg Reward: +1.2341 📈
KL Divergence:5.8901 ✅
KL Coef (β): 0.1542
Policy Loss: 0.0041
Value Loss: 0.8201
Entropy: 5.2103
Clip Fraction:0.2341 ✅
--- Iteration 250 ---
Avg Reward: +1.7823 📈
KL Divergence:7.9032 ✅
KL Coef (β): 0.1623
Policy Loss: 0.0019
Value Loss: 0.5902
Entropy: 4.9821
Clip Fraction:0.1892 ✅
PPO training complete! 🎉
All health metrics look good! Reward climbing steadily. KL well under 10 nats. Clip fraction in the 0.15–0.25 sweet zone. ✅
Part 9 — Reading the Training Metrics Like a Pro
PPO produces more training metrics than any other LLM fine-tuning method. Here's a complete guide to what each one means and how to react:
📊 PPO Metrics Diagnostic Guide
| Metric | Healthy Range | Danger Sign | Fix |
|---|---|---|---|
| mean_reward | Steadily increasing | Flat or decreasing after iter 50 | Check reward model, reduce KL coef |
| objective/kl | Grows slowly, stays < 15 | Jumps above 25 suddenly | Increase β, reduce learning rate |
| ppo/loss/policy | Decreasing, stable oscillation | Wild swings, spikes every few iters | Reduce LR, reduce mini_batch_size |
| ppo/loss/value | Decreasing over first 100 iters | Stays above 2.0 after iter 100 | Increase c₁ (vf_coef), more ppo_epochs |
| objective/clipfrac | 0.10 – 0.35 | > 0.5 or < 0.02 | High: reduce LR. Low: reduce KL coef |
| objective/entropy | Slow decrease, stays above 3.0 | Collapses to < 1.0 quickly | Increase c₂ (entropy coefficient) |
Part 10 — Detecting and Preventing Reward Hacking
Reward hacking is the #1 failure mode in PPO for LLMs. The model discovers it can achieve high reward scores without actually being helpful — by gaming the reward model's weaknesses.
Common Reward Hacking Patterns
- Verbosity inflation → Reward model slightly prefers longer answers. Model learns to pad responses with redundant sentences. "In conclusion, as I have explained above, and to summarise what was said earlier..."
- Sycophantic openers → Reward model trained on human data where humans preferred validated, agreeable responses. Model starts every response: "What a thoughtful question! You've raised an excellent point..."
- Repetition loops → Model discovers certain high-scoring phrases and repeats them. "The key insight here is important. This important insight is key. The key importance of this insight..."
- Confident hallucination → Reward model prefers confident-sounding responses. Model makes things up but presents them with authority.
def run_reward_hack_audit(model, tokenizer, reward_fn,
audit_prompts: list, iteration: int) -> dict:
"""
Detect reward hacking patterns in model outputs.
Run this every 50 PPO iterations to catch problems early.
"""
from collections import Counter
import re
results = {
"iteration": iteration,
"verbosity_score": 0, # fraction of responses over 200 words
"sycophancy_score": 0, # fraction with sycophantic openers
"repetition_score": 0, # fraction with repeated 4-grams
"avg_response_len": 0,
"reward_vs_quality_gap": 0
}
lengths = []
sycophantic_openers = [
"great question", "excellent question", "what a wonderful",
"you've raised", "absolutely right", "i completely agree"
]
for prompt in audit_prompts:
inputs = tokenizer(prompt, return_tensors="pt",
truncation=True, max_length=256).to(model.device)
with torch.no_grad():
out = model.generate(**inputs, max_new_tokens=200, temperature=0.7,
do_sample=True, pad_token_id=tokenizer.eos_token_id)
response = tokenizer.decode(out[0], skip_special_tokens=True)[len(prompt):].strip()
words = response.split()
lengths.append(len(words))
# Check verbosity
if len(words) > 200:
results["verbosity_score"] += 1
# Check sycophancy
if any(phrase in response[:150].lower() for phrase in sycophantic_openers):
results["sycophancy_score"] += 1
# Check repetition: look for 4-gram repeats
tokens = response.lower().split()
fourgrams = [" ".join(tokens[i:i+4]) for i in range(len(tokens)-3)]
counts = Counter(fourgrams)
if any(c > 2 for c in counts.values()):
results["repetition_score"] += 1
n = len(audit_prompts)
results["verbosity_score"] = results["verbosity_score"] / n
results["sycophancy_score"] = results["sycophancy_score"] / n
results["repetition_score"] = results["repetition_score"] / n
results["avg_response_len"] = sum(lengths) / len(lengths)
# Print report
print(f"\n=== Reward Hack Audit — Iteration {iteration} ===")
print(f" Avg response length: {results['avg_response_len']:.1f} words")
print(f" Verbosity rate: {results['verbosity_score']:.1%} {'⚠️' if results['verbosity_score'] > 0.3 else '✅'}")
print(f" Sycophancy rate: {results['sycophancy_score']:.1%} {'⚠️' if results['sycophancy_score'] > 0.2 else '✅'}")
print(f" Repetition rate: {results['repetition_score']:.1%} {'⚠️' if results['repetition_score'] > 0.15 else '✅'}")
any_hack = any([
results["verbosity_score"] > 0.3,
results["sycophancy_score"] > 0.2,
results["repetition_score"] > 0.15
])
if any_hack:
print("\n 🚨 REWARD HACKING DETECTED!")
print(" Action: Increase KL coefficient (β) by 0.05 and recheck in 25 iterations.")
else:
print("\n ✅ No reward hacking patterns detected.")
return results
# Run audit at key checkpoints
audit_prompts = [
"What are some tips for public speaking?",
"How do vaccines work?",
"Should I use Python or JavaScript for web development?",
"What is compound interest?",
"How do I get better at chess?"
]
# Call this every 50 iterations inside the training loop
audit_report = run_reward_hack_audit(
model_with_value, tokenizer, compute_reward_scores, audit_prompts, iteration=250
)
Output:
=== Reward Hack Audit — Iteration 250 ===
Avg response length: 87.3 words
Verbosity rate: 8.0% ✅
Sycophancy rate: 14.0% ✅
Repetition rate: 4.0% ✅
✅ No reward hacking patterns detected.
Part 11 — Saving, Merging, and Deploying
from peft import PeftModel
from transformers import AutoModelForCausalLM
from huggingface_hub import login
# ── Option A: Save just the LoRA adapter (~168 MB) ────────────
ppo_trainer.model.save_pretrained("./ppo-lora-adapter")
tokenizer.save_pretrained("./ppo-lora-adapter")
print("LoRA adapter saved! (~168 MB)")
# ── Option B: Merge into the base model (for production) ─────
# Reload base model in fp16 for merging
base_for_merge = AutoModelForCausalLM.from_pretrained(
MODEL_NAME,
torch_dtype=torch.float16,
device_map="auto"
)
# Load the LoRA adapter on top
ppo_merged = PeftModel.from_pretrained(base_for_merge, "./ppo-lora-adapter")
# Merge adapter weights into base model — zero inference overhead
ppo_merged = ppo_merged.merge_and_unload()
# ── Push to Hugging Face Hub ──────────────────────────────────
login(token="hf_your_token_here")
ppo_merged.push_to_hub(
"your-username/mistral-7b-ppo-rlhf-aligned",
commit_message="PPO-aligned Mistral-7B via RLHF — 500 iterations"
)
tokenizer.push_to_hub("your-username/mistral-7b-ppo-rlhf-aligned")
print("✅ PPO-aligned model published to Hugging Face Hub!")
print("Load it with:")
print(" AutoModelForCausalLM.from_pretrained('your-username/mistral-7b-ppo-rlhf-aligned')")
Quick Inference Test After PPO
from transformers import pipeline
# Compare SFT vs PPO model outputs on the same prompt
test_pipe = pipeline(
"text-generation",
model=ppo_merged,
tokenizer=tokenizer,
max_new_tokens=150,
temperature=0.7,
do_sample=True
)
test_prompt = "How do I stay motivated when learning a difficult new skill?"
result = test_pipe(test_prompt)[0]["generated_text"]
response = result[len(test_prompt):].strip()
print(f"PROMPT: {test_prompt}")
print(f"\nPPO Model Response:")
print(response)
Output:
PROMPT: How do I stay motivated when learning a difficult new skill?
PPO Model Response:
The key is breaking the skill into small, achievable daily targets.
Instead of "become fluent in Spanish," try "learn 5 new words today."
Small wins release dopamine and build the habit of showing up consistently.
Track your progress visibly — a simple checkmark chart on paper often outperforms digital apps
because the physical act reinforces the habit loop.
And when you inevitably plateau, remind yourself that plateaus are where the real consolidation
happens. You're not stuck — you're absorbing.
Concrete, actionable, appropriately personal — and at exactly the right length. That's what PPO alignment achieves! 🎯
Part 12 — PPO vs Alternatives: Knowing When to Choose What
🤼 PPO vs DPO vs GRPO vs ORPO
| Property | PPO | DPO | GRPO (2024) | ORPO (2024) |
|---|---|---|---|---|
| Reward model needed? | Yes | No | Yes | No |
| Value model needed? | Yes | No | No | No |
| Online (generates data)? | Yes ✅ | No (offline) | Yes ✅ | No (offline) |
| VRAM (7B model) | ~40 GB (4 models) | ~12 GB (2 models) | ~20 GB (3 models) | ~8 GB (1 model) |
| Best use case | Complex multi-dim alignment, large-scale production | Quick alignment from existing preferences | Reasoning tasks (DeepSeek-style) | Single-stage SFT+alignment combined |
| Complexity | ⭐⭐⭐⭐⭐ | ⭐⭐ | ⭐⭐⭐ | ⭐⭐ |
Common Mistakes and How to Avoid Them 🪲
-
Right-padding the tokenizer →
PPO generation requires
tokenizer.padding_side = "left". Right-padding causes attention mask misalignment during generation and produces garbage responses that look like they trained but didn't. This is one of the most silent and painful bugs in PPO implementations. - Training reward model and policy on same data → If the reward model saw the exact prompts you use for PPO, it overestimates scores for those prompts. Always use separate prompt pools for reward model training and PPO rollout.
- Not watching generated samples manually → Metrics look great. Reward is climbing. But the model is generating verbose nonsense. Read 10 actual generated responses every 50 iterations. Numbers can be deceiving. Text never lies.
- Setting ppo_epochs too high → With K=8 or K=16 epochs on the same rollout batch, the old rollout data becomes increasingly off-policy — the advantage estimates are computed from a policy that no longer exists. Keep ppo_epochs = 4 as the maximum. K=2 is safer.
- Skipping advantage normalisation → Always normalise advantages across the batch (subtract mean, divide by std) before computing the PPO loss. Unnormalised advantages cause wildly varying gradient magnitudes that destabilise training from the first iteration.
- Comparing PPO reward to absolute quality → A model with mean reward +2.5 is not "objectively good." It's good relative to the reference model and reward model at that point in training. Always run an external quality benchmark (MT-Bench, AlpacaEval, head-to-head win rate) alongside reward tracking to verify true quality improvement.
Quick Summary 📝
The complete PPO knowledge map — everything we covered:
- RL Vocabulary → Agent, environment, state, action, reward, policy — the 6 concepts that underpin all of RL
- Why PPO exists → Vanilla policy gradient destroys knowledge in one bad update. TRPO solved this but was computationally brutal. PPO achieves TRPO-level stability with first-order gradient descent.
- Probability ratio r(θ) → Measures how much the policy changed per action. If it drifts too far from 1.0, clip it.
- Clipped surrogate objective →
min(rA, clip(r, 1-ε, 1+ε)A). Prevents any single update from being too aggressive. - Advantage function A_t → Was this action better or worse than expected? GAE with λ=0.95 gives the best estimate.
- Full PPO loss → Policy gradient + value function loss + entropy bonus. All three together.
- KL divergence penalty → Adjusted reward = Reward - β×KL. Prevents reward hacking and catastrophic policy drift.
- Four models → Policy (trainable), Value (trainable), Reward (frozen), Reference (frozen). All four needed simultaneously.
- Health metrics → Watch KL, clip fraction, entropy, and value loss — not just reward.
- Reward hacking → Verbosity, sycophancy, repetition. Audit every 50 iterations. Raise β when detected.
Happy Learning 🧠✨
Comments
Post a Comment