Skip to main content

Reward Models and RLHF Explained: How LLMs Learn from Human Feedback

Calculating read time…

Have you ever wondered why ChatGPT sounds helpful and polite instead of rambling on like a broken search engine? The secret weapon behind that behaviour is called RLHF — Reinforcement Learning from Human Feedback.

And at the very heart of RLHF sits a special component called the Reward Model. It's the judge, the scorekeeper, the compass that points the LLM toward responses that real humans actually prefer.


Why Do LLMs Need RLHF at All?

Imagine you train a student entirely by having them read millions of books. They become incredibly knowledgeable — but they never had a teacher correct their behaviour.

When you ask this student a question, they might give you a technically accurate but brutally unhelpful answer. They might write three pages when you needed three sentences. They might even confidently say something wrong.

Pre-training an LLM on raw internet text has the exact same problem. The model learns to predict the next word — not to be helpful to humans. These are very different goals. 🎯

💡 Think of it like this:

  • Pre-training = teaching a parrot to mimic every sound it has ever heard
  • RLHF = teaching that parrot to say things its owner actually wants to hear

RLHF is the bridge between "a model that can generate text" and "a model that generates text humans find genuinely useful and safe."

🗺️ The Full RLHF Pipeline — A Bird's Eye View

Before diving deep, let's see the big picture. RLHF has three distinct phases that happen in sequence:

📐 The RLHF Pipeline

Phase 1
📚
Supervised
Fine-Tuning (SFT)

Teach the model
basic instruction following

→
Phase 2
⚖️
Train the
Reward Model (RM)

Learn what humans
prefer

→
Phase 3
🤖
RL Fine-Tuning
with PPO

Use reward signal to
improve the LLM

↩ Human feedback continuously improves the reward model over multiple cycles

We'll explore each phase in detail. But the star of the show — the Reward Model — is what makes or breaks the whole pipeline.

Phase 1 — 📚 Supervised Fine-Tuning (SFT)

Before we can do any reinforcement learning, we need a base model that already knows how to follow instructions — at least roughly.

SFT is simply training the pre-trained LLM on a curated dataset of (instruction → ideal response) pairs, written or verified by humans.

What SFT Produces

  • A model that understands the format of instructions and responses
  • A starting point with sensible, generally helpful behaviour
  • A much better base for RLHF than the raw pre-trained model

Think of SFT as giving the model its first job training before performance reviews begin. It knows what's expected — now we refine how well it does it.

🟢 DO: Use a high-quality, diverse instruction dataset for SFT. Datasets like OpenHermes 2.5, SlimOrca, or Dolly 15k are excellent starting points. The reward model can only improve what the SFT model already knows — it cannot teach brand-new skills.

Phase 2 — ⚖️ The Reward Model — The Heart of RLHF

This is the most important and most misunderstood part of RLHF. Let's take it apart carefully.

What Is a Reward Model?

A reward model is a neural network that takes a (prompt, response) pair as input and outputs a single scalar score — a number that represents how good that response is.

Higher score = humans would prefer this response. Lower score = humans would find this response less useful, harmful, or low-quality.

💡 Real-world analogy:

Imagine you're a restaurant critic 🍽️. You visit a restaurant, taste a dish, and give it a rating from 1 to 10. You don't write the recipe — you just judge how good the dish is.

The reward model is that critic. It doesn't generate responses — it judges them. And the LLM is the chef, trying to cook dishes the critic will rate highly.

How Is a Reward Model Built? — The Preference Data

The reward model is trained on human preference data. Here is exactly how this data is collected:

  1. Take a prompt (a real user question)
  2. Generate two or more different responses using the SFT model
  3. Show both responses to a human annotator
  4. Ask the annotator: "Which response do you prefer, and why?"
  5. Record the result as a preference pair: (prompt, chosen response, rejected response)

Do this thousands of times across diverse prompts, and you have a preference dataset that captures what humans actually value.

A Real Preference Data Example

Prompt:

User: Explain what a neural network is. I'm a complete beginner.

Response A (Chosen ✅):

A neural network is inspired by how the human brain works.
Think of it as a system of small decision-makers called "neurons."
Each neuron takes some input, does a simple calculation,
and passes the result to the next neuron.

By layering thousands of these neurons together,
the network learns to recognise patterns — like whether an image
contains a cat, or whether an email is spam.

It learns from examples, just like a child learns to recognise
a dog after seeing many dogs. The more examples it sees,
the better it gets.

Response B (Rejected ❌):

A neural network is a computational graph with weighted edges
connecting nodes across multiple layers.
Forward propagation computes activations using non-linear
functions like ReLU or sigmoid,
while backpropagation updates weights via gradient descent
to minimise a loss function.

Response A wins because it uses an analogy, explains step by step, and is written for a beginner — exactly what the prompt asked for. Response B is technically correct but completely ignores the audience.

The reward model learns to associate this kind of preference signal with specific qualities: clarity, relevance, appropriate complexity, helpfulness.

Reward Model Architecture — Under the Hood

A reward model is typically built by taking the same LLM used for SFT and replacing its final language-modelling head with a regression head — a single linear layer that outputs one number.

🏗️ Reward Model Architecture

Input: [Prompt + Response] → Tokenised
↓
Transformer Backbone (same as SFT model)
All hidden layers — frozen or partially fine-tuned
↓
Final Token Hidden State → Linear Layer
↓
Output: Single Scalar Score (e.g., 7.4)

The model is trained using a ranking loss. For every preference pair (chosen, rejected), we want the reward model to assign a higher score to the chosen response than to the rejected one.

The training objective in plain English:

Loss = - log( sigmoid( reward(chosen) - reward(rejected) ) )

This pushes the score gap between chosen and rejected responses to grow larger over training. The wider the gap, the more confident and accurate the reward model becomes.

Step-by-Step: Training a Simple Reward Model in Python

Let's build a minimal reward model trainer using Hugging Face and PyTorch:

Step 1: Install Dependencies

pip install transformers datasets trl torch accelerate

Step 2: Load a Preference Dataset

from datasets import load_dataset # Load the Anthropic HH-RLHF preference dataset dataset = load_dataset("Anthropic/hh-rlhf", split="train") # Each sample has: # - 'chosen' → the full conversation with the preferred response # - 'rejected' → the full conversation with the less preferred response print(dataset[0]['chosen'][:300]) print("---") print(dataset[0]['rejected'][:300])

Output (truncated):

Human: What are some ways to make friends? Assistant: Making friends starts with showing genuine interest in others. Try joining clubs, attending community events, or taking classes where you meet people with shared interests... --- Human: What are some ways to make friends? Assistant: Just talk to people. Say hi. It's not that complicated.

Step 3: Build the Reward Model

import torch from transformers import AutoModelForSequenceClassification, AutoTokenizer model_name = "distilbert-base-uncased" # Lightweight for demonstration # Load model with a regression head (num_labels=1 = single score output) reward_model = AutoModelForSequenceClassification.from_pretrained( model_name, num_labels=1 ) tokenizer = AutoTokenizer.from_pretrained(model_name) print("Reward model loaded!") print(f"Output head: {reward_model.classifier}")

Output:

Reward model loaded! Output head: Linear(in_features=768, out_features=1, bias=True)

Step 4: Define the Ranking Loss and Training Loop

import torch.nn as nn from torch.optim import AdamW optimizer = AdamW(reward_model.parameters(), lr=1e-5) def compute_ranking_loss(chosen_score, rejected_score): # We want: chosen_score > rejected_score # Loss is low when the gap is large (chosen clearly beats rejected) return -torch.log(torch.sigmoid(chosen_score - rejected_score)).mean() def get_score(text): tokens = tokenizer( text, return_tensors="pt", truncation=True, max_length=512, padding=True ) output = reward_model(**tokens) return output.logits.squeeze() # Single scalar score # ── One training step ────────────────────────────────────── reward_model.train() chosen_text = "Making friends starts with showing genuine interest in others..." rejected_text = "Just talk to people. Say hi. It's not that complicated." chosen_score = get_score(chosen_text) rejected_score = get_score(rejected_text) loss = compute_ranking_loss(chosen_score, rejected_score) optimizer.zero_grad() loss.backward() optimizer.step() print(f"Chosen score: {chosen_score.item():.4f}") print(f"Rejected score: {rejected_score.item():.4f}") print(f"Training loss: {loss.item():.4f}")

Output (after some training):

Chosen score:   2.3471
Rejected score: -1.0823
Training loss:   0.1247

The chosen response is scoring much higher than the rejected one. The reward model is learning what "better" means! 🎯

Step 5: Using TRL's RewardTrainer (Production Way)

In practice, you don't write the training loop from scratch. Hugging Face's TRL (Transformer Reinforcement Learning) library handles all of this cleanly:

from trl import RewardTrainer, RewardConfig from transformers import AutoModelForSequenceClassification, AutoTokenizer from datasets import load_dataset model_name = "meta-llama/Llama-3.2-1B" model = AutoModelForSequenceClassification.from_pretrained( model_name, num_labels=1 ) tokenizer = AutoTokenizer.from_pretrained(model_name) tokenizer.pad_token = tokenizer.eos_token dataset = load_dataset("Anthropic/hh-rlhf", split="train[:5000]") training_args = RewardConfig( output_dir="./reward_model_output", num_train_epochs=1, per_device_train_batch_size=4, learning_rate=1e-5, logging_steps=50, save_steps=500, remove_unused_columns=False, ) trainer = RewardTrainer( model=model, args=training_args, train_dataset=dataset, tokenizer=tokenizer, ) trainer.train() trainer.save_model("./my_reward_model") print("Reward model saved! ✅")
🟢 DO: Always use the same tokenizer and architecture family for your reward model as for your SFT model. Mismatched tokenizers cause silent scoring errors that are very difficult to debug later.
🔴 DON'T: Underestimate the importance of preference data quality. If your annotators disagree with each other more than 30% of the time (low inter-annotator agreement), your reward model will learn a blurry, confused signal — and your LLM fine-tuning will make the model worse, not better.

Phase 3 — 🤖 RL Fine-Tuning with PPO

Now we have a trained reward model that can score any response. Phase 3 uses this signal to actually improve the LLM through reinforcement learning.

The Core Idea — Reward-Guided Improvement

At every training step, the LLM generates a response to a prompt. The reward model scores that response. The RL algorithm uses that score to update the LLM's weights — nudging it toward responses that earn higher rewards.

💡 Think of it like this:

The LLM is a contestant on a cooking show 👨‍🍳. Every dish it prepares gets a score from the judge (the reward model). The contestant studies which dishes scored well and adjusts their technique. Over hundreds of rounds, they become significantly better at pleasing the judge.

What is PPO?

PPO — Proximal Policy Optimisation — is the RL algorithm used in the original InstructGPT (ChatGPT's predecessor) paper. It's designed to update the model's behaviour without making changes so drastic that the model forgets everything it learned during SFT.

PPO balances two competing goals during training:

  • Maximise reward — generate responses the reward model scores highly
  • Stay close to the SFT model — don't drift so far that the model loses its language quality

The second goal is enforced by a term called the KL divergence penalty. It measures how different the current model's outputs are from the original SFT model's outputs. If the model drifts too far, the KL penalty increases and pulls it back.

🔄 PPO Training Loop (One Step)

Sample a batch of prompts from the dataset
↓
LLM generates a response for each prompt
↓
Reward Model scores each (prompt, response) pair
↓
Compute KL penalty (how far did we drift from SFT model?)
↓
Final reward = reward score − β × KL penalty
↓
Update LLM weights using PPO gradient step
↓ Repeat for thousands of steps

PPO Fine-Tuning in Code — Using TRL

from trl import PPOTrainer, PPOConfig, AutoModelForCausalLMWithValueHead from transformers import AutoTokenizer from datasets import load_dataset import torch # ── Load the SFT model with a value head ──────────────────── # The value head predicts expected future reward — needed for PPO model = AutoModelForCausalLMWithValueHead.from_pretrained( "path/to/your/sft_model" ) # ── Load the reward model ──────────────────────────────────── from transformers import AutoModelForSequenceClassification reward_model = AutoModelForSequenceClassification.from_pretrained( "./my_reward_model" ) reward_model.eval() tokenizer = AutoTokenizer.from_pretrained("path/to/your/sft_model") tokenizer.pad_token = tokenizer.eos_token # ── PPO configuration ──────────────────────────────────────── ppo_config = PPOConfig( model_name="my_rlhf_model", learning_rate=1.41e-5, batch_size=16, mini_batch_size=4, gradient_accumulation_steps=4, ppo_epochs=4, kl_penalty="kl", # KL divergence to prevent reward hacking target_kl=6.0, # Stop updating if KL exceeds this threshold init_kl_coef=0.2, # Initial KL penalty coefficient (β) ) ppo_trainer = PPOTrainer( config=ppo_config, model=model, tokenizer=tokenizer, ) # ── One PPO training step ──────────────────────────────────── prompts = ["Explain what gravity is.", "How do I make pasta?"] for prompt_text in prompts: # Tokenise prompt query_tensor = tokenizer.encode(prompt_text, return_tensors="pt") # LLM generates a response response_tensor = ppo_trainer.generate( query_tensor, max_new_tokens=200, do_sample=True, temperature=0.7 ) response_text = tokenizer.decode(response_tensor[0], skip_special_tokens=True) # Reward model scores the response reward_input = tokenizer( response_text, return_tensors="pt", truncation=True, max_length=512 ) with torch.no_grad(): reward_output = reward_model(**reward_input) reward_score = reward_output.logits.squeeze() print(f"Prompt: {prompt_text}") print(f"Response: {response_text[:100]}...") print(f"Reward: {reward_score.item():.4f}") print("---")

Output (example):

Prompt: Explain what gravity is. Response: Gravity is a natural force that pulls objects toward each other. The more mass an object has, the stronger its gravitational pull... Reward: 3.8724 --- Prompt: How do I make pasta? Response: Boil salted water, add pasta, cook for 8-10 minutes until al dente, drain and serve with your favourite sauce... Reward: 4.1293 ---
🟡 TIP — Reward Hacking: If the KL penalty is set too low, the LLM will find shortcuts to get high reward scores without actually being helpful. For example, it might learn to always add emojis or always say "Great question!" — because those patterns happened to score well in training. This is called reward hacking and is one of the key challenges in RLHF. Always monitor KL divergence throughout training.

⚡ Modern Alternative — DPO (Direct Preference Optimisation)

PPO-based RLHF is powerful but complex. It requires training and maintaining two separate models simultaneously (the LLM and the reward model), careful KL tuning, and a lot of GPU memory.

In 2023, researchers introduced DPO — Direct Preference Optimisation. DPO eliminates the need for a separate reward model entirely.

How DPO Works

DPO takes the same preference data (chosen, rejected pairs) and fine-tunes the LLM directly to prefer chosen responses over rejected ones. Mathematically, it bakes the reward model into the LLM's loss function.

The result: same alignment quality, much simpler training.

  • No separate reward model to train or maintain
  • No PPO sampling loop
  • No KL penalty tuning
  • Uses standard supervised fine-tuning infrastructure
  • Significantly less GPU memory required

DPO Training in Code

from trl import DPOTrainer, DPOConfig from transformers import AutoModelForCausalLM, AutoTokenizer from datasets import load_dataset model_name = "meta-llama/Llama-3.2-1B-Instruct" # Load the SFT model (the one to be aligned) model = AutoModelForCausalLM.from_pretrained(model_name) # DPO also needs a frozen reference model (the original SFT weights) ref_model = AutoModelForCausalLM.from_pretrained(model_name) tokenizer = AutoTokenizer.from_pretrained(model_name) tokenizer.pad_token = tokenizer.eos_token # Dataset must have: 'prompt', 'chosen', 'rejected' columns dataset = load_dataset("Anthropic/hh-rlhf", split="train[:3000]") dpo_config = DPOConfig( output_dir="./dpo_aligned_model", num_train_epochs=1, per_device_train_batch_size=2, learning_rate=5e-7, # Very small LR — we're making subtle adjustments beta=0.1, # DPO's equivalent of the KL penalty strength logging_steps=25, save_steps=500, ) trainer = DPOTrainer( model=model, ref_model=ref_model, # Reference model kept frozen throughout args=dpo_config, train_dataset=dataset, tokenizer=tokenizer, ) trainer.train() trainer.save_model("./my_dpo_model") print("DPO alignment complete! ✅")
🟢 DO: For most practical use cases and smaller teams, start with DPO rather than full PPO-RLHF. DPO is easier to implement, requires less compute, and achieves comparable results on most alignment tasks. Switch to PPO only if DPO doesn't meet your quality requirements or if you need very fine-grained reward control.

RLHF (PPO) vs DPO — Quick Comparison

  • RLHF with PPO → Separate reward model + RL loop. More flexible. More complex. More compute. Used by OpenAI (InstructGPT), Anthropic (early Claude versions).
  • DPO → No separate reward model. Direct fine-tuning on preferences. Simpler. Faster. Used in Llama 2 Chat, Zephyr, many open-source models.
  • RLAIF → Reward signal from another AI (not humans). Scales cheaply. Used in Claude's Constitutional AI approach.
  • GRPO → Group Relative Policy Optimisation, used in DeepSeek-R1. Removes the value head entirely, further simplifying RL training.

🧪 Evaluating Your Reward Model

Before using your reward model in PPO training, you must verify it actually learned to distinguish good from bad responses.

Method 1 — Accuracy on Held-Out Preference Pairs

from transformers import AutoModelForSequenceClassification, AutoTokenizer import torch reward_model = AutoModelForSequenceClassification.from_pretrained("./my_reward_model") tokenizer = AutoTokenizer.from_pretrained("./my_reward_model") reward_model.eval() def score(text): tokens = tokenizer(text, return_tensors="pt", truncation=True, max_length=512) with torch.no_grad(): return reward_model(**tokens).logits.item() # Test on held-out preference pairs test_pairs = [ { "chosen": "Paris is the capital of France, known for the Eiffel Tower.", "rejected": "Paris is a city somewhere in Europe I think." }, { "chosen": "To centre a div in CSS: use display:flex; justify-content:center; align-items:center;", "rejected": "Just add margin:auto and it should work somehow." }, ] correct = 0 for pair in test_pairs: score_chosen = score(pair["chosen"]) score_rejected = score(pair["rejected"]) is_correct = score_chosen > score_rejected correct += int(is_correct) print(f"Chosen score: {score_chosen:.3f}") print(f"Rejected score: {score_rejected:.3f}") print(f"Correct: {'✅' if is_correct else '❌'}") print("---") print(f"Accuracy: {correct}/{len(test_pairs)} = {correct/len(test_pairs)*100:.1f}%")

Output:

Chosen score:   3.847
Rejected score: -0.923
Correct: ✅
---
Chosen score:   4.102
Rejected score: 0.312
Correct: ✅
---
Accuracy: 2/2 = 100.0%

Method 2 — Reward Score Distribution Analysis

Plot the distribution of reward scores for chosen vs. rejected responses across your entire validation set.

A well-trained reward model produces clearly separated distributions: chosen responses cluster at high scores, rejected ones cluster at low scores. Overlapping distributions mean the model is confused.

🟡 TIP — Target Accuracy: A reward model achieving 65–70% accuracy on held-out preference pairs is reasonable for training. Above 75% is excellent. Below 60% means your preference data is too noisy or your model needs more training. Don't proceed to PPO with an under-trained reward model.

⚠️ Key Challenges in RLHF — What Can Go Wrong

Challenge 1 — Reward Hacking

The LLM finds patterns that fool the reward model into giving high scores without the response actually being helpful. Common examples: excessive hedging, sycophantic phrases like "Excellent question!", or padding responses with repetitive filler to appear thorough.

🔴 DON'T: Train PPO for too many steps without monitoring reward hacking. Always track response quality (not just reward scores) using human evaluation or a separate, independent evaluator model at regular checkpoints.

Challenge 2 — Annotation Inconsistency

Human annotators don't always agree. What one annotator marks as "preferred" another might reject. Low inter-annotator agreement (below 70%) produces noisy training signal.

Solutions include: clear annotation guidelines, annotator calibration sessions, majority-vote aggregation across multiple annotators per pair, and using model-based consistency checks to flag high-disagreement samples.

Challenge 3 — Distribution Shift

The reward model is trained on responses from the SFT model. As PPO training progresses, the LLM generates increasingly different responses — and the reward model may not score them accurately because they're outside its training distribution.

The solution is iterative RLHF: periodically collect new preference data from the latest version of the LLM and retrain the reward model with this fresh data.

Challenge 4 — Over-Optimisation

Running PPO for too long maximises the reward model score but eventually degrades actual response quality. The model "over-fits" to the reward model's preferences, which are only an imperfect approximation of true human values.

Monitor real human preference win rates — not just reward scores — to know when to stop PPO training.

🌍 Real-World RLHF Implementations

InstructGPT / ChatGPT (OpenAI)

The original paper that popularised RLHF for LLMs. Used PPO with a 6B parameter reward model trained on ~33,000 human preference comparisons. Showed that the 1.3B InstructGPT model was preferred by humans over the 175B GPT-3 on most tasks — proving alignment quality matters more than raw size.

Claude (Anthropic)

Anthropic extended RLHF with Constitutional AI (CAI) — a technique where an AI model critiques and revises its own outputs against a written "constitution" of principles. This reduces dependence on human annotation for safety-related preferences.

Llama 2 Chat (Meta)

Used two separate reward models — one for helpfulness and one for safety — then combined their scores. This two-reward architecture prevents the helpfulness reward from overriding safety signals.

Zephyr / Mistral Instruct

Used DPO instead of PPO, with AI-generated preference data (RLAIF). Achieved competitive alignment quality at a fraction of the annotation cost.

📊 Monitoring RLHF in Production (MLOps Perspective)

From an MLOps standpoint, RLHF is not a one-time event. It's an ongoing process that needs continuous monitoring and iteration.

Key Metrics to Track

  • Mean reward score — Average reward model score over batches. Should increase during training, but flag if it grows too fast (reward hacking signal).
  • KL divergence from SFT baseline — Should stay within the target range. Spikes indicate policy instability.
  • Human win rate — What percentage of the time do real humans prefer the RLHF model's output over the SFT baseline? This is the ground truth metric.
  • Helpfulness score — Measured by a separate evaluator on standardised prompts.
  • Safety violation rate — Percentage of outputs flagged by a safety classifier.
  • Response length trend — A sudden increase in average length often indicates verbosity hacking.
🟢 DO: Set up experiment tracking (Weights & Biases or MLflow) from the very start of RLHF training. Reward hacking and policy instability can emerge suddenly after hundreds of steps of apparently normal training. You need the full training history to diagnose what happened.

🗂️ Complete RLHF Workflow — End to End Summary

🗺️ Full RLHF Workflow at a Glance

  1. Collect instruction data → Write or curate (prompt, ideal response) pairs
  2. SFT fine-tuning → Train the base LLM to follow instructions
  3. Generate response pairs → For each prompt, produce 2+ responses using SFT model
  4. Human preference annotation → Annotators rank which response they prefer
  5. Train reward model → Fit a regression head on the preference data using ranking loss
  6. Evaluate reward model → Confirm 65%+ accuracy on held-out pairs
  7. PPO or DPO training → Use the reward signal to fine-tune the LLM
  8. Monitor training → Track KL divergence, reward scores, response quality
  9. Human evaluation → Measure real win rate against the SFT baseline
  10. Iterate → Collect new preferences, retrain reward model, repeat

Quick Summary 📝

What we learned today:

  • Why RLHF? → Pre-training alone produces capable but misaligned models. RLHF teaches them to be genuinely helpful.
  • SFT (Phase 1) → Fine-tune on human-written (instruction, response) pairs to create a strong base model.
  • Reward Model (Phase 2) → A regression-head LLM trained on preference pairs to score responses. Heart of the RLHF pipeline.
  • PPO (Phase 3) → RL algorithm that uses reward scores to iteratively improve the LLM, with KL penalty to prevent reward hacking.
  • DPO → Modern simpler alternative to PPO. No separate reward model. Same preference data. Easier to implement.
  • Key risks → Reward hacking, annotation inconsistency, distribution shift, over-optimisation.
  • MLOps → Track mean reward, KL divergence, response length, and human win rate throughout training.

RLHF is the difference between a raw language model and a model you can actually trust and deploy. Happy training! 🤖✨

Comments