Skip to main content

Preference Datasets & DPO: Aligning Large Language Models with Human Preferences

Calculating read time…

You've trained a language model. It can write text, answer questions, and generate code. But there's a problem — it answers every question the same way it learned from training data. Sometimes it's too verbose. Sometimes it's rude. Sometimes it confidently makes things up.

The model knows how to generate text. But it doesn't know which text a human would actually prefer. That gap — between what a model can say and what it should say — is exactly what preference alignment solves.

💡 Think of it like this: Imagine you hire a talented new writer who reads everything on the internet. They can write in any style, on any topic. But left to their own devices, they write in a style that's technically accurate but cold, confusing, or inappropriate. You need to sit with them, show them examples of good vs. bad writing, and help them develop a sense of what you actually want. That training process — learning from human preferences — is alignment. 📝

The Alignment Problem — Why Models Need Preferences

When you pre-train a language model on internet text, it learns to predict the most statistically likely next word. The internet contains brilliant writing — and terrible writing. The model doesn't know which is which.

After supervised fine-tuning (SFT) on instruction datasets, the model learns to follow instructions. But following an instruction and following it well are different things.

Consider this prompt: "What is the capital of France?"

  • Response A: "Paris."
  • Response B: "The capital of France is Paris, a city that has served as the country's political and cultural centre for centuries, nestled along the Seine river in the north of the country, with a population of approximately 2.1 million within city limits and over 12 million in the greater metropolitan area, established as..."

Both responses are factually correct. Both follow the instruction. But response A is clearly better for this question. Response B shows the model doesn't understand that concision is a virtue here.

To teach a model these judgements, you need a special kind of training data: preference data.

What Is a Preference Dataset?

A preference dataset is a collection of examples, where each example contains a prompt and two or more responses — one labelled as preferred (chosen) and one as less preferred (rejected).

Instead of teaching the model "here is the correct answer," preference data teaches the model "between these two options, humans prefer this one — and here's what that reveals about quality."

The Core Structure

Every row in a preference dataset has exactly three components:

  • prompt → The question or instruction given to the model
  • chosen → The response humans judged as better (the preferred answer)
  • rejected → The response humans judged as worse (the less preferred answer)

Here is what a raw preference dataset entry looks like in JSON:

{
  "prompt": "Explain what a neural network is.",

  "chosen": "A neural network is a system loosely inspired by the human brain.
             It consists of layers of simple units called neurons,
             each of which receives inputs, applies a transformation,
             and passes results forward to the next layer.
             Through training on examples, the network learns which
             transformations lead to correct outputs.",

  "rejected": "Neural networks are complex mathematical structures invented
               in the mid-20th century that use matrix multiplications
               and non-linear activation functions organised in directed
               acyclic computation graphs to approximate arbitrary functions
               via stochastic gradient descent over parametrised weight spaces."
}

Both responses are accurate. The chosen response is preferred because it uses plain language and builds from the familiar (brain) to the unfamiliar (layers, training). The rejected response is technically correct but uses jargon that a beginner finds inaccessible.

📐 Anatomy of a Preference Dataset Sample

💬 PROMPT

The question or task given to the model.

→

✅ CHOSEN (preferred)

The response humans judged as better quality.

❌ REJECTED (less preferred)

The response humans judged as lower quality.

→

🧠 MODEL LEARNS

"I should produce responses more like chosen and less like rejected."

What Makes One Response Better Than Another?

Before you can collect preferences, you need to understand what dimensions of quality humans care about. The most widely used framework comes from alignment research and covers seven dimensions:

  • Helpfulness → Does the response actually solve the user's problem? A response that's technically correct but doesn't help the user accomplish their goal is not helpful.
  • Harmlessness → Does the response avoid causing harm? This includes physical harm, psychological harm, misinformation, and harmful stereotypes.
  • Honesty → Does the response accurately represent what the model knows and doesn't know? Confident hallucinations are worse than honest uncertainty.
  • Conciseness → Does the response use the right amount of words? Neither padding answers with unnecessary filler nor omitting critical details.
  • Coherence → Does the response flow logically? Does it stay on topic without contradicting itself?
  • Instruction-following → Does the response exactly follow the format, length, or constraints specified in the prompt?
  • Safety → Does the response appropriately decline harmful requests rather than complying with them?
🟡 TIP: When creating your own preference dataset, be explicit about which dimensions your annotators should prioritise. If your use case is a medical chatbot, harmlessness and honesty outweigh conciseness. If your use case is a coding assistant, accuracy and instruction-following come first. Different applications need different preference signals.

Famous Public Preference Datasets

You don't always need to create preference data from scratch. Several high-quality public datasets are freely available:

  • Anthropic HH-RLHF → The dataset released by Anthropic from their Constitutional AI research. Contains human-labelled conversations, each with a chosen and rejected response. Focuses on helpfulness and harmlessness. Around 170,000 examples.
  • OpenAI WebGPT Comparisons → Comparison data from training WebGPT, OpenAI's web-browsing model. Humans compare two model-generated answers and indicate which is more accurate.
  • UltraFeedback → 64,000 instructions each answered by four different LLMs, then scored by GPT-4 on helpfulness, honesty, and instruction-following. One of the highest-quality synthetic preference datasets currently available.
  • OpenHermes Preferences / Capybara → Community-curated preference datasets covering diverse instruction types. Widely used for DPO fine-tuning open-source models.
  • Orca DPO Pairs → Preference pairs generated from Orca's reasoning dataset using GPT-4 for quality judgement. Strong for improving chain-of-thought reasoning quality.

How to Create Your Own Preference Dataset

Building a custom preference dataset is the most impactful thing you can do to align a model for your specific use case. There are three main strategies — each with different costs, quality levels, and scalability.

Strategy 1 — Human Annotation (Highest Quality)

The gold standard. Real humans read pairs of model responses and indicate which they prefer. This is how the original RLHF datasets at OpenAI and Anthropic were built.

The process step by step:

  • Collect a diverse set of prompts covering your target use cases
  • Generate two or more responses to each prompt using your model (or multiple models)
  • Present each prompt + response pair to human annotators
  • Annotators select the better response and optionally write a brief reason
  • Use majority vote (3–5 annotators per pair) to reduce individual bias
import json
from datetime import datetime

def create_annotation_task(prompt: str, response_a: str, response_b: str) -> dict:
    """
    Creates a structured annotation task for human labellers.
    Returns a dict ready to be sent to an annotation platform.
    """
    return {
        "task_id": f"pref_{hash(prompt) % 100000:05d}",
        "created_at": datetime.now().isoformat(),
        "prompt": prompt,
        "response_a": response_a,
        "response_b": response_b,
        "annotation_guidelines": {
            "primary_criteria": "helpfulness",
            "secondary_criteria": ["accuracy", "conciseness", "tone"],
            "options": ["A is better", "B is better", "roughly equal", "both are bad"]
        }
    }

# Example task
task = create_annotation_task(
    prompt="How should I deal with a coworker who takes credit for my work?",
    response_a="Document your contributions carefully and have a direct, 
    professional conversation with your coworker. If it continues, involve your manager with specific documented examples.",
    response_b="This is a difficult interpersonal situation that many professionals encounter. 
    There are multiple approaches you could consider. First, you might want to think about documenting your work..."
)

print(json.dumps(task, indent=2))

Output:

{
  "task_id": "pref_42891",
  "created_at": "2025-01-15T09:23:41",
  "prompt": "How should I deal with a coworker who takes credit for my work?",
  "response_a": "Document your contributions carefully and have a direct, professional conversation...",
  "response_b": "This is a difficult interpersonal situation that many professionals encounter...",
  "annotation_guidelines": {
    "primary_criteria": "helpfulness",
    "secondary_criteria": ["accuracy", "conciseness", "tone"],
    "options": ["A is better", "B is better", "roughly equal", "both are bad"]
  }
}
🟢 DO: Write clear, detailed annotation guidelines before collecting human labels. Annotators without guidance default to personal preferences — one annotator prizes brevity while another prizes thoroughness. The result is contradictory signal that confuses your model rather than aligning it. Agree on your quality rubric first, then annotate.

Strategy 2 — LLM-as-a-Judge (Scalable & Practical)

Use a powerful LLM like GPT-4 or Claude to generate preference labels automatically. This produces thousands of high-quality preference pairs per hour, at a fraction of the cost of human annotation.

The quality of synthetic preference data is surprisingly close to human labels — research has shown that GPT-4's preferences correlate strongly with human judgements on most tasks, with the exception of highly subjective or culturally nuanced content.

Step 1 — Generate Multiple Responses to the Same Prompt

from openai import OpenAI

client = OpenAI()

def generate_response_pair(prompt: str, model_a: str, model_b: str) -> dict:
    """
    Generate two responses to the same prompt using different
    models or different sampling temperatures.
    """
    # Response A — lower temperature (more focused)
    resp_a = client.chat.completions.create(
        model=model_a,
        messages=[{"role": "user", "content": prompt}],
        temperature=0.3,
        max_tokens=512
    )

    # Response B — higher temperature (more diverse, sometimes worse)
    resp_b = client.chat.completions.create(
        model=model_b,
        messages=[{"role": "user", "content": prompt}],
        temperature=1.0,
        max_tokens=512
    )

    return {
        "prompt": prompt,
        "response_a": resp_a.choices[0].message.content,
        "response_b": resp_b.choices[0].message.content
    }

pair = generate_response_pair(
    prompt="Explain how a vaccine works to someone with no science background.",
    model_a="gpt-4o-mini",
    model_b="gpt-4o-mini"
)

Step 2 — Score Both Responses with a Judge LLM

import json

JUDGE_SYSTEM_PROMPT = """You are an expert evaluator of AI assistant responses.
Your job is to compare two responses to the same user prompt and determine which is better.

Evaluation criteria (in order of priority):
1. Accuracy — is the information factually correct?
2. Helpfulness — does it actually address what the user needs?
3. Clarity — is it easy to understand for the intended audience?
4. Appropriate length — neither too brief nor unnecessarily padded?

Return ONLY a valid JSON object with this exact structure:
{
  "winner": "A" or "B" or "tie",
  "chosen": the full text of the winning response,
  "rejected": the full text of the losing response,
  "reasoning": "one concise sentence explaining the choice",
  "scores": {"A": 1-5, "B": 1-5}
}"""

def judge_response_pair(prompt: str, response_a: str, response_b: str) -> dict:
    """
    Use GPT-4 as a judge to compare two responses and determine preference.
    Returns structured preference data ready for DPO training.
    """
    user_message = f"""PROMPT: {prompt}

RESPONSE A:
{response_a}

RESPONSE B:
{response_b}

Which response is better? Apply the evaluation criteria strictly."""

    result = client.chat.completions.create(
        model="gpt-4o",
        messages=[
            {"role": "system", "content": JUDGE_SYSTEM_PROMPT},
            {"role": "user", "content": user_message}
        ],
        temperature=0,        # deterministic judgement
        response_format={"type": "json_object"}
    )

    judgement = json.loads(result.choices[0].message.content)
    return {
        "prompt": prompt,
        "chosen": judgement["chosen"],
        "rejected": judgement["rejected"],
        "winner": judgement["winner"],
        "reasoning": judgement["reasoning"],
        "scores": judgement["scores"]
    }

# Judge our pair
preference_sample = judge_response_pair(
    pair["prompt"],
    pair["response_a"],
    pair["response_b"]
)

print(json.dumps(preference_sample, indent=2))

Output:

{
  "prompt": "Explain how a vaccine works to someone with no science background.",
  "chosen": "Think of your immune system as an army that protects your body.
             A vaccine is like a training drill — it shows your army a
             harmless version of an enemy (a virus) so they learn to
             recognise and fight it. When the real virus arrives later,
             your immune army already knows exactly how to defeat it.",
  "rejected": "Vaccines work through the principle of immunological memory.
               They introduce antigens — typically attenuated or inactivated
               pathogens, or their subunits — which trigger an adaptive
               immune response including T-cell activation and B-cell
               differentiation into memory cells...",
  "winner": "A",
  "reasoning": "Response A uses a simple, memorable analogy perfect for a non-scientist, 
  while Response B uses heavy medical jargon that contradicts the stated audience.",
  "scores": {"A": 5, "B": 2}
}

Step 3 — Build the Dataset at Scale

from datasets import Dataset
import pandas as pd
from tqdm import tqdm

def build_preference_dataset(prompts: list, output_path: str) -> Dataset:
    """
    Generate a full preference dataset from a list of prompts.
    Saves to disk and returns a Hugging Face Dataset.
    """
    preference_data = []
    failed = 0

    for prompt in tqdm(prompts, desc="Generating preference pairs"):
        try:
            # Generate two responses
            pair = generate_response_pair(prompt, "gpt-4o-mini", "gpt-4o-mini")

            # Skip if both responses are identical or very similar
            if pair["response_a"].strip() == pair["response_b"].strip():
                continue

            # Judge the pair
            pref = judge_response_pair(
                pair["prompt"],
                pair["response_a"],
                pair["response_b"]
            )

            # Only keep clear winners — skip ties
            if pref["winner"] != "tie":
                preference_data.append({
                    "prompt": pref["prompt"],
                    "chosen": pref["chosen"],
                    "rejected": pref["rejected"]
                })

        except Exception as e:
            failed += 1
            continue

    print(f"Generated: {len(preference_data)} preference pairs")
    print(f"Failed/skipped: {failed}")

    # Save as both JSON and Hugging Face Dataset
    df = pd.DataFrame(preference_data)
    df.to_json(f"{output_path}.json", orient="records", indent=2)

    dataset = Dataset.from_pandas(df)
    dataset.save_to_disk(output_path)

    return dataset


# Example list of seed prompts
seed_prompts = [
    "What is the best way to learn programming as a complete beginner?",
    "How do I politely decline a meeting invitation at work?",
    "Explain the difference between a virus and a bacterium.",
    "What should I do if I find a lost dog?",
    "How can I improve my focus while studying?",
    # ... add hundreds more
]

dataset = build_preference_dataset(seed_prompts, "./my_preference_dataset")
print(dataset)

Output:

Generated: 487 preference pairs
Failed/skipped: 13

Dataset({
    features: ['prompt', 'chosen', 'rejected'],
    num_rows: 487
})

487 high-quality preference pairs built automatically! 🎉

🔴 DON'T: Use the same model at the same temperature to generate both the chosen and rejected responses. If the responses are too similar in quality, the preference signal is noise — the model can't learn anything meaningful from the comparison. Use different temperatures, different models, or intentionally degrade one response (e.g., ask the LLM to rewrite it more verbosely or less accurately).

Strategy 3 — Adversarial Response Generation (Most Efficient)

Instead of randomly generating two responses and hoping one is worse, deliberately create bad responses by instructing an LLM to write a flawed version of a good response.

def generate_adversarial_pair(prompt: str, good_response: str) -> dict:
    """
    Take a known-good response and generate a flawed version as the rejected response.
    This guarantees a meaningful quality gap between chosen and rejected.
    """

    # Pick a random flaw type to inject
    import random
    flaw_types = [
        "Make it verbose and padded with unnecessary filler words",
        "Make it vague and unhelpful — hedge every claim",
        "Add one subtle factual error",
        "Make it unnecessarily condescending in tone",
        "Make it incomplete — stop halfway through a key point",
        "Use heavy technical jargon that would confuse a beginner"
    ]
    flaw = random.choice(flaw_types)

    result = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[
            {"role": "system", "content": f"Rewrite the given response but intentionally 
            make it worse by applying this flaw: {flaw}. 
            Keep the core topic but lower the overall quality."},
            {"role": "user", "content": f"Original response:\n{good_response}\n\n
            Rewrite with the flaw applied:"}
        ],
        temperature=0.7
    )

    flawed_response = result.choices[0].message.content

    return {
        "prompt": prompt,
        "chosen": good_response,
        "rejected": flawed_response,
        "flaw_injected": flaw
    }

# Example
pair = generate_adversarial_pair(
    prompt="How do I make a good first impression in a job interview?",
    good_response="Arrive 5–10 minutes early, dress one level above 
    the company's typical dress code, make eye contact 
    and smile when greeting people, prepare two or three thoughtful questions 
    about the role, and follow up with a thank-you email within 24 hours."
)

print(f"Flaw injected: {pair['flaw_injected']}")
print(f"\nChosen:\n{pair['chosen']}")
print(f"\nRejected:\n{pair['rejected']}")

Output:

Flaw injected: Make it verbose and padded with unnecessary filler words

Chosen:
Arrive 5–10 minutes early, dress one level above the company's typical dress code,
make eye contact and smile when greeting people, prepare two or three thoughtful
questions about the role, and follow up with a thank-you email within 24 hours.

Rejected:
Well, there are actually quite a number of various different things that you might
want to potentially consider thinking about when it comes to the topic of making
what could perhaps be described as a good first impression in the context of a
job interview situation...

Cleaning and Validating Your Preference Dataset

Before training, always validate your preference dataset. Bad preference data is worse than no preference data — it can actively misalign your model.

from datasets import load_from_disk

def validate_preference_dataset(dataset_path: str) -> dict:
    """
    Run quality checks on a preference dataset before DPO training.
    Returns a report of issues found.
    """
    dataset = load_from_disk(dataset_path)
    issues = []
    stats = {}

    # Check 1: No empty fields
    for i, sample in enumerate(dataset):
        if not sample["prompt"].strip():
            issues.append(f"Row {i}: empty prompt")
        if not sample["chosen"].strip():
            issues.append(f"Row {i}: empty chosen response")
        if not sample["rejected"].strip():
            issues.append(f"Row {i}: empty rejected response")

    # Check 2: Chosen and rejected are not identical
    identical = sum(
        1 for s in dataset
        if s["chosen"].strip() == s["rejected"].strip()
    )
    if identical > 0:
        issues.append(f"{identical} samples have identical chosen and rejected responses")

    # Check 3: Length ratio check (rejected should not be systematically longer)
    chosen_lens = [len(s["chosen"].split()) for s in dataset]
    rejected_lens = [len(s["rejected"].split()) for s in dataset]
    avg_chosen = sum(chosen_lens) / len(chosen_lens)
    avg_rejected = sum(rejected_lens) / len(rejected_lens)

    stats = {
        "total_samples": len(dataset),
        "avg_prompt_length": sum(len(s["prompt"].split()) for s in dataset) / len(dataset),
        "avg_chosen_length": avg_chosen,
        "avg_rejected_length": avg_rejected,
        "identical_pairs": identical,
        "issues_found": len(issues)
    }

    print("=== Preference Dataset Validation Report ===")
    for key, val in stats.items():
        print(f"  {key}: {val:.1f}" if isinstance(val, float) else f"  {key}: {val}")

    if issues:
        print(f"\n⚠️  Issues found:")
        for issue in issues[:5]:   # show first 5
            print(f"  - {issue}")
    else:
        print("\n✅ No issues found. Dataset is ready for DPO training.")

    return stats

stats = validate_preference_dataset("./my_preference_dataset")

Output:

=== Preference Dataset Validation Report ===
  total_samples: 487
  avg_prompt_length: 14.2
  avg_chosen_length: 67.8
  avg_rejected_length: 94.3
  identical_pairs: 0
  issues_found: 0

✅ No issues found. Dataset is ready for DPO training.

What Is DPO? — Direct Preference Optimization

Before DPO, the dominant method for aligning language models was RLHF (Reinforcement Learning from Human Feedback). RLHF works but is notoriously complex — it requires training a separate reward model, running a PPO (Proximal Policy Optimization) loop, and carefully tuning multiple interacting systems.

DPO (Direct Preference Optimization) was introduced in 2023 and dramatically simplified this process. It bypasses the reward model entirely, training the language model directly on preference data using a clever mathematical reformulation.

💡 Think of it like this: RLHF is like hiring a sports coach (the reward model) to score every practice game, then training your athlete (the LLM) based on those scores in a complex reinforcement loop. DPO cuts out the coach — it says "here are two performances, one better than the other, adjust your technique directly from the comparison." Same outcome, far simpler process. 🏅

How DPO Works — The Intuition

DPO has a beautifully simple objective: increase the probability of generating the chosen response while decreasing the probability of generating the rejected response — but do so in a controlled way that prevents the model from drifting too far from its original behaviour.

📐 What DPO Does During Training

📈

Chosen response

Probability goes UP

📉

Rejected response

Probability goes DOWN

⚖️

Reference model

Anchors training — prevents catastrophic forgetting

The DPO Loss Function — Plain English

The DPO loss compares how your current model scores the chosen response versus the rejected response, relative to how a frozen "reference" version of the same model scores them.

The reference model is the SFT model before any preference training — it acts as an anchor that prevents the model from collapsing to degenerate behaviour. Without it, the model might learn to generate one-word responses as "chosen" and random noise as "rejected" — technically satisfying the preference objective but useless.

DPO Loss = -log( σ( β × (log π(chosen|x)/π_ref(chosen|x) - log π(rejected|x)/π_ref(rejected|x)) ) )

Breaking this down piece by piece:

  • π(response|x) → Your current model's probability of generating this response given the prompt
  • π_ref(response|x) → The reference (frozen SFT) model's probability of the same response
  • log π/π_ref → How much your model has changed relative to the reference model — positive means you're boosting this response, negative means you're reducing it
  • β → A temperature parameter (typically 0.1–0.5) controlling how strongly the reference model constrains training. Higher β → model stays closer to reference.
  • σ → The sigmoid function, which converts the difference to a probability between 0 and 1

The loss is minimised when: the gap between your model and the reference is larger for chosen than for rejected. In plain English: train the model to prefer chosen responses more than the reference model already does.

Implementing DPO in Practice

The good news: you don't need to implement the DPO loss yourself. The trl library (Transformer Reinforcement Learning) provides a DPOTrainer that handles everything.

Step 1 — Install and Import

pip install trl peft transformers accelerate bitsandbytes datasets
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
from peft import LoraConfig, get_peft_model, prepare_model_for_kbit_training, TaskType
from trl import DPOTrainer, DPOConfig
from datasets import load_dataset, load_from_disk

Step 2 — Load the SFT Model as Both Policy and Reference

model_name = "mistralai/Mistral-7B-Instruct-v0.2"

# 4-bit quantization config (QLoRA + DPO)
bnb_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_compute_dtype=torch.bfloat16,
    bnb_4bit_use_double_quant=True
)

# Load the trainable policy model
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    quantization_config=bnb_config,
    device_map="auto"
)
model = prepare_model_for_kbit_training(model)

# Load the tokenizer
tokenizer = AutoTokenizer.from_pretrained(model_name)
tokenizer.pad_token = tokenizer.eos_token
tokenizer.padding_side = "left"    # DPO uses LEFT padding (unlike SFT which uses right)

print(f"Policy model loaded. Memory: {model.get_memory_footprint() / 1e9:.2f} GB")

Output:

Policy model loaded. Memory: 4.12 GB
🟡 TIP — Padding Side Matters for DPO: DPO requires padding_side = "left" (unlike SFT which needs "right"). This is because DPO computes log probabilities across the full sequence, and right-padding would cause the model to include padding tokens in the probability calculation. Getting this wrong causes subtle bugs that are very hard to debug — always double-check this setting before DPO training.

Step 3 — Add LoRA Adapters for Efficient DPO

from peft import LoraConfig, get_peft_model

lora_config = LoraConfig(
    r=16,
    lora_alpha=32,
    lora_dropout=0.05,
    target_modules=[
        "q_proj", "k_proj", "v_proj", "o_proj",
        "gate_proj", "up_proj", "down_proj"
    ],
    bias="none",
    task_type=TaskType.CAUSAL_LM
)

model = get_peft_model(model, lora_config)
model.print_trainable_parameters()

Output:

trainable params: 83,886,080 || all params: 3,827,765,248 || trainable%: 2.1913

Step 4 — Load and Format the Preference Dataset

from datasets import load_dataset

# Option A: Use a public preference dataset
dataset = load_dataset("trl-lib/ultrafeedback_binarized", split="train[:5000]")

# Option B: Load your own dataset from disk
# dataset = load_from_disk("./my_preference_dataset")

# The DPOTrainer expects exactly these three column names:
# 'prompt', 'chosen', 'rejected'
# Let's verify they exist and preview a sample
print("Dataset columns:", dataset.column_names)
print(f"\nDataset size: {len(dataset)} samples")
print("\nSample entry:")
print(f"PROMPT: {dataset[0]['prompt'][:100]}...")
print(f"CHOSEN: {dataset[0]['chosen'][:100]}...")
print(f"REJECTED: {dataset[0]['rejected'][:100]}...")

Output:

Dataset columns: ['prompt', 'chosen', 'rejected']

Dataset size: 5000 samples

Sample entry:
PROMPT: What are the most important skills for a software engineer to develop?...
CHOSEN: The most critical skills for a software engineer combine technical depth and...
REJECTED: Software engineering requires many different skills. You need to know how to...

Step 5 — Configure and Run DPO Training

from trl import DPOConfig, DPOTrainer

dpo_config = DPOConfig(
    # Core DPO parameter — controls reference model constraint
    # Lower beta = more aggressive preference learning (may cause drift)
    # Higher beta = more conservative (stays closer to reference)
    beta=0.1,

    # Training hyperparameters
    output_dir="./dpo-mistral-output",
    num_train_epochs=1,              # DPO usually needs only 1–2 epochs
    per_device_train_batch_size=2,
    gradient_accumulation_steps=8,   # effective batch = 2 × 8 = 16
    learning_rate=5e-5,              # lower than SFT — DPO is a fine adjustment, not a rewrite
    warmup_ratio=0.1,
    lr_scheduler_type="cosine",
    logging_steps=10,
    save_strategy="epoch",

    # Memory optimisation
    bf16=True,
    optim="paged_adamw_32bit",
    max_grad_norm=0.3,
    gradient_checkpointing=True,

    # DPO-specific settings
    max_prompt_length=512,           # maximum tokens in the prompt
    max_length=1024,                 # maximum tokens in prompt + response combined

    report_to="none"
)

trainer = DPOTrainer(
    model=model,
    ref_model=None,        # when ref_model=None, TRL automatically creates a frozen copy
    args=dpo_config,
    train_dataset=dataset,
    tokenizer=tokenizer
)

trainer.train()

Output (sample):

{'loss': 0.6931, 'rewards/chosen': 0.0000, 'rewards/rejected': 0.0000, 'epoch': 0.0}
{'loss': 0.5823, 'rewards/chosen': 0.3241, 'rewards/rejected': -0.2891, 'epoch': 0.2}
{'loss': 0.4901, 'rewards/chosen': 0.5812, 'rewards/rejected': -0.4902, 'epoch': 0.5}
{'loss': 0.4102, 'rewards/chosen': 0.7823, 'rewards/rejected': -0.6721, 'epoch': 0.8}
{'loss': 0.3891, 'rewards/chosen': 0.8901, 'rewards/rejected': -0.7812, 'epoch': 1.0}
Training complete! ✅

Watch the two critical metrics closely:

  • rewards/chosen → Should increase over training. Positive values mean the model assigns higher probability to chosen responses than the reference model does. 📈
  • rewards/rejected → Should decrease (become more negative) over training. Negative values mean the model assigns lower probability to rejected responses than the reference model does. 📉
  • loss → Should steadily decrease, but not too fast. Loss dropping too quickly (< 0.2 in first epoch) usually means overfitting on a small dataset.

Step 6 — Save the DPO-Aligned Adapter

# Save only the LoRA adapter
trainer.save_model("./dpo-adapter")
tokenizer.save_pretrained("./dpo-adapter")

print("DPO adapter saved!")
print(f"Adapter size: ~{sum(p.numel() * 2 for p in model.parameters() if p.requires_grad) / 1e6:.0f} MB")

Output:

DPO adapter saved!
Adapter size: ~168 MB

Evaluating Your DPO-Aligned Model

After training, you need to verify the model actually became better. Two of the strongest evaluation approaches:

Evaluation 1 — Win Rate Against the Reference Model

Generate responses from both the DPO-aligned model and the original SFT model, then use GPT-4 to judge which is better. A DPO-aligned model should win significantly more than 50% of the time.

from peft import PeftModel

# Load both models
base_model = AutoModelForCausalLM.from_pretrained(
    model_name, quantization_config=bnb_config, device_map="auto"
)
dpo_model = PeftModel.from_pretrained(base_model, "./dpo-adapter")

from transformers import pipeline

# Create inference pipelines for both
sft_pipe = pipeline("text-generation", model=base_model, tokenizer=tokenizer,
                     max_new_tokens=256, temperature=0.7, do_sample=True)
dpo_pipe = pipeline("text-generation", model=dpo_model, tokenizer=tokenizer,
                     max_new_tokens=256, temperature=0.7, do_sample=True)

def compare_models(prompts: list, n_samples: int = 50) -> dict:
    dpo_wins = 0
    sft_wins = 0
    ties = 0

    for prompt in prompts[:n_samples]:
        sft_response = sft_pipe(prompt)[0]["generated_text"]
        dpo_response = dpo_pipe(prompt)[0]["generated_text"]

        # Judge with GPT-4
        judgement = judge_response_pair(prompt, dpo_response, sft_response)

        if judgement["winner"] == "A":    # A = DPO model
            dpo_wins += 1
        elif judgement["winner"] == "B":  # B = SFT model
            sft_wins += 1
        else:
            ties += 1

    total = n_samples - ties
    win_rate = dpo_wins / total if total > 0 else 0

    print("=== Model Comparison Results ===")
    print(f"DPO model wins: {dpo_wins} ({dpo_wins/n_samples*100:.1f}%)")
    print(f"SFT model wins: {sft_wins} ({sft_wins/n_samples*100:.1f}%)")
    print(f"Ties:           {ties}     ({ties/n_samples*100:.1f}%)")
    print(f"\nDPO Win Rate (excluding ties): {win_rate:.1%}")

    return {"dpo_wins": dpo_wins, "sft_wins": sft_wins, "ties": ties, "win_rate": win_rate}

test_prompts = [
    "How do I apologise professionally after making a mistake at work?",
    "What is the difference between machine learning and deep learning?",
    "Can you help me write a polite but firm email declining an invitation?",
    # ... add more evaluation prompts
]

results = compare_models(test_prompts)

Output:

=== Model Comparison Results ===
DPO model wins: 34 (68.0%)
SFT model wins: 11 (22.0%)
Ties:            5  (10.0%)

DPO Win Rate (excluding ties): 75.6%

The DPO-aligned model wins 75.6% of head-to-head comparisons against the original SFT model. Alignment is working! 🏆

Evaluation 2 — Reward Margin Tracking

def compute_reward_margin(trainer, eval_dataset) -> dict:
    """
    Compute the average reward margin between chosen and rejected responses.
    Higher margin = model has learned a clearer preference signal.
    """
    eval_results = trainer.evaluate(eval_dataset)

    margin = eval_results.get("eval_rewards/margins", 0)
    chosen_reward = eval_results.get("eval_rewards/chosen", 0)
    rejected_reward = eval_results.get("eval_rewards/rejected", 0)

    print("=== Reward Margin Report ===")
    print(f"Average chosen reward:   {chosen_reward:.4f}")
    print(f"Average rejected reward: {rejected_reward:.4f}")
    print(f"Reward margin:           {margin:.4f}")

    if margin > 1.0:
        print("\n✅ Strong preference signal — model has learned clear preferences.")
    elif margin > 0.5:
        print("\n🟡 Moderate signal — consider more training data or epochs.")
    else:
        print("\n🔴 Weak signal — check your dataset quality or beta parameter.")

    return eval_results

The Complete Pipeline — SFT → DPO in One View

🔄 Complete LLM Alignment Pipeline

🌍

Stage 0 — Pre-trained Base Model

LLaMA 3, Mistral, Gemma etc. — trained on raw internet text. Knows language but doesn't know how to behave.

↓
🎓

Stage 1 — Supervised Fine-Tuning (SFT)

Train on instruction-response pairs. Model learns to follow instructions. Output: SFT model.

↓
📊

Stage 2 — Preference Dataset Creation

Collect (prompt, chosen, rejected) triplets via human annotation, LLM-as-a-Judge, or adversarial generation.

↓
🎯

Stage 3 — DPO Training

DPOTrainer with LoRA. SFT model = policy AND reference. Model learns to prefer chosen over rejected.

↓
🚀

Stage 4 — Merge, Evaluate & Deploy

Merge DPO adapter → evaluate win rate → push to Hugging Face Hub → serve with vLLM.

Pushing Your Aligned Model to Hugging Face

from peft import PeftModel
from transformers import AutoModelForCausalLM
from huggingface_hub import login

login(token="hf_your_token_here")

# Reload base model in fp16 for merging
base = AutoModelForCausalLM.from_pretrained(
    model_name, torch_dtype=torch.float16, device_map="auto"
)

# Load and merge the DPO adapter
dpo_merged = PeftModel.from_pretrained(base, "./dpo-adapter")
dpo_merged = dpo_merged.merge_and_unload()

# Push merged DPO-aligned model to Hugging Face Hub
dpo_merged.push_to_hub("your-username/mistral-7b-dpo-aligned")
tokenizer.push_to_hub("your-username/mistral-7b-dpo-aligned")

print("✅ DPO-aligned model published to Hugging Face Hub!")

Common Mistakes and How to Avoid Them 🪲

  • Using an SFT model that never learned to follow instructions → DPO is a fine-tuning on top of SFT — it adjusts preferences, not foundational instruction-following. If your SFT model is weak, DPO won't rescue it. Always do SFT first.
  • Setting beta too low (< 0.05) → Very low beta means the reference model barely constrains training. The model can overfit to the preference dataset, start repeating chosen responses verbatim, and forget how to handle prompts outside the training distribution. Start with beta=0.1 and adjust from there.
  • Preference data with low quality gaps → If your chosen and rejected responses are almost equally good, the model receives a weak and confusing training signal. Aim for clear quality differences — use the adversarial generation strategy to guarantee a meaningful gap.
  • Training for too many epochs → Unlike SFT, DPO typically needs only 1–3 epochs. More epochs cause the model to become overconfident — it assigns near-zero probability to any response that resembles a rejected sample, even when that response would be perfectly appropriate in a different context.
  • Forgetting left-padding for the tokenizer → DPO requires tokenizer.padding_side = "left". Right-padding (the SFT default) causes incorrect log-probability calculations and produces a model that appears to train but doesn't actually align.
🔴 DON'T: Treat a high DPO win rate on your own evaluation prompts as proof of alignment. Always test on held-out prompts that were not used anywhere in your preference dataset creation pipeline. If the same LLM that generated your preference pairs also evaluates your model, you are measuring how well the model matches one LLM's style — not genuine alignment.
🟢 DO: After DPO training, always run a qualitative sanity check on safety. Ask your DPO-aligned model potentially harmful questions and verify it still refuses appropriately. DPO training on certain datasets can inadvertently weaken safety behaviours if the preference data included rejections of cautious responses.

Quick Summary 📝

What we covered today:

  • The alignment gap → Models can follow instructions but don't know which responses humans prefer. Preference data bridges this gap.
  • Preference dataset structure → Every sample has three fields: prompt, chosen (preferred), and rejected (less preferred).
  • Three dataset creation strategies → Human annotation (highest quality), LLM-as-a-Judge (scalable and practical), adversarial generation (most efficient, guaranteed quality gap).
  • DPO (Direct Preference Optimization) → Trains the model directly on preference pairs without a separate reward model. Simpler, faster, and more stable than RLHF.
  • DPO key concepts → Policy model (being trained), reference model (frozen SFT anchor), beta (constraint strength).
  • Implementation → QLoRA + DPOTrainer + left-padding + paged optimiser. 1 epoch is usually enough. Monitor rewards/chosen and rewards/rejected during training.
  • Evaluation → Win rate against the SFT model using GPT-4 as a judge. A well-aligned DPO model should win 65–80% of head-to-head comparisons.

Next Steps 🚀

  • Read the DPO paper (Rafailov et al., 2023) once you're comfortable with the implementation — the derivation showing how DPO is mathematically equivalent to RLHF under certain assumptions is a beautiful piece of work that will deepen your understanding significantly.

Preference alignment is what separates a raw language model from an AI assistant people actually want to use. A model that can write anything is impressive. A model that consistently chooses to write the right thing is powerful. 🧠✨

Comments