Skip to main content

Model Selection Playbook: When to Choose a Bigger Base Model Over Fine-Tuning

Calculating read time…

Choosing between a stronger base model and a fine-tuned cheaper model is not a taste decision — it is a measurement problem, and the measurement that should decide it is not the one most teams reach for first. Most teams start by comparing average benchmark scores or a general "vibes" pass on a handful of prompts. That is backwards. The evidence that actually tells you which path is safe to ship comes from how each candidate behaves on your worst, riskiest, most easily-broken slice of traffic — not on its best day. 📊

The stakes are not abstract. A cheaper fine-tuned model that looks 2 points behind a frontier model on an aggregate leaderboard can still be the right production choice — or the wrong one — depending entirely on what happens when a user asks it to touch account_access, phrases a request in a way nobody wrote a test for, or expects a tool call with a clean, parseable schema. Get this decision wrong in either direction and you either overpay for capability you didn't need, or you ship a narrowed, brittle model that fails exactly where failure is most expensive. 🚦

Diagram showing two paths, a stronger base model and a fine-tuned cheaper model, both passing through three sequential gates: high-risk slices, structured-output and paraphrase stability, and efficiency budget, with average benchmark score only as a final tiebreaker

Original diagram: the evidence order this post argues for — risk gates first, format and robustness second, cost third, average score last.

🔀 Quick Comparison: Stronger Base Model vs. Fine-Tuned Cheaper Model

Dimension Stronger Base Model Fine-Tuned Cheaper Model
High-risk slice gates Often clears with little or no tuning, because broad capability transfers to edge cases it never explicitly saw. Can clear the same gates, but usually needs targeted examples of the risky slice — and enough of them to matter.
Behavioral narrowing risk Low. General instruction-following is preserved because you are not overwriting weights toward one pattern. Rises with tuning intensity — aggressive fine-tuning to force a gate to pass can quietly narrow behavior elsewhere.
Structured-output stability Generally strong out of the box on well-specified schemas, but not guaranteed on your exact fields. Can be made very stable on your specific schema, if the tuning set covers the field combinations you actually see.
Per-request cost & latency Higher, sometimes significantly — this is the trade-off you are buying capability with. Lower, often the entire reason the path was considered in the first place.
When average score wins the argument Never on its own — only after both candidates already cleared the risk, format, and cost gates. Same rule applies. A higher average score never overrides a failed high-risk gate.

1. The Decision Fork: Stronger Base Model vs. Fine-Tuning a Cheaper One

Kid analogy: imagine you need someone to watch your little brother for an afternoon. You could hire the neighborhood's most experienced babysitter, who costs more but has handled every kind of tantrum before — or you could give your quieter, cheaper cousin a quick lesson on exactly what to do if your brother cries, refuses dinner, or tries to leave the house. Both can work. But you would never decide by asking "who's nicer on average?" You'd first ask: what happens the one time something risky happens — does this person handle it, or panic?

That's the entire shape of the base-model-vs-fine-tuning decision. A stronger base model is the experienced babysitter: broad capability, usually solid on things it was never specifically taught, at a higher cost per hour. A fine-tuned cheaper model is the coached cousin: cheap and perfectly fine most of the time, but only as good as the specific coaching it received — and coaching that's too aggressive can make the cousin rigid and unable to improvise outside the script.

The mistake most teams make is running this comparison the way you'd compare two students on a report card: overall GPA first. In production LLM systems, the far more decisive evidence sits in a small number of places: how each candidate performs on your highest-risk traffic slice, whether it refuses (or doesn't refuse) the right things, whether it survives users rephrasing their request, whether its structured outputs stay valid, and what it costs to run. Average quality is a real signal — but it is supposed to be the last thing you look at, not the first.

✅ Worked example: a support-ticket triage system needs to decide between GPT-4-class routing and a fine-tuned smaller model. On average helpfulness, the fine-tuned model actually scores slightly higher, because it was trained on the company's own ticket phrasing. But that average score says nothing yet about what happens when a ticket mentions account_access — which is exactly the slice this post keeps coming back to.

🎯 Use this when: you are staring at two candidate models with similar average scores and need a principled reason to pick one, rather than defaulting to "the bigger one is safer" or "the cheaper one is good enough."

2. Offline vs. Online Evaluation: Where the Decision Actually Gets Tested

Kid analogy: think about learning to ride a bike. First you practice in an empty parking lot with training wheels off, on ground you already know — that's offline evaluation: a controlled, repeatable test on a known course. Later you ride it to school through actual traffic, weather, and other kids on scooters — that's online evaluation: the real, messy environment where new things happen that the parking lot never had.

Offline evaluation runs a candidate model against a fixed, versioned dataset — a "golden set" — and scores it with deterministic checks, automated metrics, or an LLM-as-judge pass, before anything reaches a real user. Online evaluation happens after deployment: sampling live traffic, scoring a subset of real interactions, watching quality metrics drift over days and weeks, and running controlled comparisons like A/B tests or canary rollouts between the incumbent and the challenger model.

Google's Vertex AI evaluation tooling is a useful real-world illustration of how a major cloud AI platform structures the offline half of this: it supports what it calls adaptive rubrics, where instead of applying one generic scoring rubric to every prompt, the service analyzes each prompt individually and generates a tailored set of pass/fail checks for it, similar to writing a unit test for that specific input, then validates the model's response against those generated checks. That is offline evaluation done well: repeatable, per-prompt, and specific enough to catch real failures instead of averaging them away. Azure AI Foundry's evaluation tooling adds the online-adjacent piece directly relevant to this decision: it lets teams line up base models against one another on operating cost, response speed, and request throughput, alongside quality, using their own private data rather than public leaderboards, which is precisely the kind of evidence a base-model-vs-fine-tune decision needs before it ever reaches production traffic.

💡 Key warning: a clean offline pass is necessary, not sufficient. The support-ticket triage example from Section 1 might sail through its golden set and still misbehave three weeks into production once real customers start phrasing account-access requests in ways nobody anticipated when the test set was written. Offline evaluation tells you what you already thought to test for. Online evaluation tells you what you didn't.

🎯 Use this when: you're deciding how much confidence to place in a pre-launch eval report — treat it as a floor you must clear, never as proof that production will behave the same way.

3. High-Risk Slice Performance: The account_access Gate

Kid analogy: a lifeguard doesn't get certified by swimming laps well on an average day. They get certified by proving they can handle the one scenario that actually matters: a swimmer in trouble in deep water. Everything else on the test is secondary to that one high-stakes skill.

A high-risk slice is a subset of your evaluation traffic tagged because getting it wrong causes outsized damage — financial, legal, safety, or trust-related — even if it's a small percentage of total volume. account_access is a textbook example: requests that touch login, password reset, permission changes, or account takeover-adjacent actions. A model that's 98% accurate overall but shaky on this slice is not a 98%-good model in any way that matters to the business — it is a model with a small number of very expensive failure modes hiding inside a good-looking aggregate.

Mechanically, evaluating a high-risk slice means: (1) tag every example in your golden set with the risk category it belongs to, (2) report pass rates per slice, never only in aggregate, (3) set a separate, usually much higher, pass-rate floor for high-risk slices than for general traffic, and (4) treat a failure on this floor as a hard gate — not something an otherwise-good average score can offset. Microsoft's Azure AI Foundry documents this same instinct at the platform level: its risk and safety evaluators score each flagged content category on its own severity scale rather than folding everything into a single number, precisely so a team can set stricter acceptance thresholds on the categories that carry more downside.

✅ Worked example, continued: back to the ticket-triage system. The team pulls 400 historical tickets tagged account_access, holds them out as their own slice, and requires 99%+ correct handling (correct routing, no premature account changes, escalation when identity isn't verified) before either candidate model is eligible — regardless of how either scores on the other 9,600 general tickets.

🎯 Use this when: any part of your traffic touches money movement, access control, medical guidance, legal commitments, or anything else where a rare failure is disproportionately costly.

4. Refusal Correctness: Judging What a Model Says No To

Kid analogy: a good babysitter says no to letting your brother play with matches, but says yes to letting him have a second glass of water. A bad babysitter either says yes to everything (dangerous) or no to everything (useless). Refusal correctness is checking that the model's "no" lands on the right things, and only the right things.

Refusal correctness has two failure directions, and most eval suites only catch one of them. Under-refusal is when a model complies with a request it should have declined or escalated — in the account_access slice, that might mean resetting a password without verifying identity. Over-refusal is when a model declines something completely legitimate — refusing to tell a verified account owner their own account status. A fine-tuning pass aimed narrowly at fixing under-refusal on one risky pattern very often overcorrects and quietly inflates over-refusal on adjacent, legitimate requests, which is exactly the "behavioral narrowing" this topic keeps warning about.

Mechanically, this needs two paired test sets scored separately: a "should-refuse" set (measuring how often the model correctly declines or escalates) and a "should-comply" set of look-alike, legitimate requests (measuring how often the model wrongly declines those too). Reporting only the should-refuse pass rate hides over-refusal entirely, which is a classic way a fine-tuned model can look safer than it actually is to use.

💡 Harder example: the ticket-triage team fine-tunes their cheaper model specifically to stop it from resetting passwords on unverified requests. It now passes the should-refuse set at 99%. But on the paired should-comply set — verified users asking routine account questions — its pass rate drops from 96% to 84%. The gate looks fixed; a new problem was created right next to it.

🎯 Use this when: you are fine-tuning specifically to fix a safety or compliance gap — always re-run the paired should-comply set afterward, not just the gap you were targeting.

5. Robustness Under Paraphrases: Surviving Out-of-Distribution Stress

Kid analogy: if a kid only recognizes their times tables when they're written in the exact font from their workbook, they haven't actually learned multiplication — they've memorized a picture. Real understanding means 7×8 still equals 56 whether it's typed, handwritten, or spoken out loud. Robustness testing asks a model the same question in different "handwriting" to check if it actually learned the underlying task.

Paraphrase robustness measures whether a model's correctness, refusal behavior, and structured-output validity hold up when the same underlying request is phrased differently — typos, reordered clauses, regional phrasing, translated-and-back requests, or a customer describing the same account_access problem in five different ways. This matters disproportionately for fine-tuned models, because fine-tuning on a fixed set of example phrasings can teach a model to pattern-match surface wording rather than genuinely generalize the underlying task — a narrower form of overfitting that a clean pass on the original phrasing will never reveal.

Mechanically: take your existing golden set (including the high-risk slice) and generate 3–5 paraphrased variants per example — ideally a mix of automated paraphrasing and a smaller hand-written adversarial set — then require the model's pass rate on the paraphrased set to stay within a small tolerance of its pass rate on the originals. A large gap between "passes on the original wording" and "passes on a rephrased version of the exact same request" is one of the clearest available signals that fine-tuning has narrowed behavior rather than genuinely taught it.

✅ Worked example, continued: the fine-tuned cheaper model scores 97% on the original account_access test set, but only 78% on paraphrased versions of the same 400 tickets. The stronger base model scores 95% on the originals and 93% on paraphrases. Average-score comparison would have favored the fine-tuned model; the robustness gap flips the recommendation.

🎯 Use this when: comparing a fine-tuned candidate against a strong base model — a small drop from original to paraphrased performance in the base model is normal; a large drop in the fine-tuned model is a narrowing signal you should not ignore.

6. Structured-Output Stability: category, priority, and tool_calls

Kid analogy: filling out a form correctly isn't just about writing true things — it's about writing them in the right boxes. A kid who writes their birthday in the "name" field gave you a true fact in the wrong place, and the form-processing machine downstream will choke on it just the same as if they'd lied. Structured-output evaluation checks that the model fills in the right boxes, every time, not just that it "knows the answer."

Most production LLM systems don't consume free text — they consume a schema: a category field that must be one of a fixed enum, a priority field that must be an accepted value, and one or more tool_calls that must be valid, well-formed, and actually appropriate for the request. A model can be semantically "right" and still be operationally useless if it emits an unparseable field, a hallucinated enum value, or a tool call with a plausible-looking but wrong argument.

Google's Vertex AI evaluation service explicitly separates this out as its own agent-evaluation dimension, with a dedicated metric for judging whether the actions an agent takes in response to what a user actually asked for are the right ones — treating tool-call correctness as its own scored thing rather than folding it into a general "was the answer good" judgment. Azure AI Foundry does something structurally similar with metrics that measure how well an agent's steps adhere to its assigned task and how complete its final response is against a ground truth, on top of separate correctness scoring for the tool calls themselves. The lesson for a base-model-vs-fine-tune decision: score schema validity, enum correctness, and tool-call correctness as three separate numbers, because a model can be strong on one and quietly broken on another.

{
  "example_id": "ticket_04231",
  "input": "I can't get into my account, it keeps saying wrong password",
  "expected": {
    "category": "account_access",
    "priority": "high",
    "tool_calls": [
      {"name": "verify_identity", "arguments": {"channel": "email_otp"}}
    ]
  },
  "checks": [
    "category in allowed_categories",
    "priority in allowed_priorities",
    "tool_calls[0].name == 'verify_identity'",
    "no destructive tool call before identity is verified"
  ]
}

An original, illustrative schema-check example — not copied from any vendor's documentation or sample repo.

💡 Key warning: schema validity and semantic correctness are not the same check. A model can emit perfectly valid JSON with the wrong category, or an invalid tool call with the exactly right intent. Report both a "parses cleanly" rate and a "parses cleanly AND is correct" rate — collapsing them into one number hides which failure mode you actually have.

🎯 Use this when: your system routes downstream on a model's structured output (ticket routing, workflow triggers, agent tool use) — any instability here breaks automation even when the underlying reasoning was fine.

7. Efficiency Signals: Cost, Latency, and Throughput as First-Class Evals

Kid analogy: a tutor who gives a perfect answer but takes three hours to explain simple homework isn't actually the tutor you want for daily use, even if their explanations are flawless. Speed and cost aren't afterthoughts bolted onto quality — they're part of whether the help is actually usable.

Cost, latency, and throughput should be scored on the same golden set, at the same time, as correctness — not measured separately in a different spreadsheet after the fact. A stronger base model that clears every risk and stability gate is still the wrong production choice if its latency blows your response-time budget or its per-call cost makes the feature commercially unviable at expected volume. Azure AI Foundry's model-benchmarking tooling reflects this directly: it lets teams set base models side by side on response speed, per-request pricing, and request volume capacity, together with quality metrics, using their own data — treating efficiency as a benchmarking axis rather than an afterthought.

Mechanically, track at minimum: p50 and p95 latency per request, cost per 1,000 requests at expected traffic mix, and tokens-in/tokens-out ratios if you're comparing models with different context-handling behavior. Report these per high-risk slice too — a model that's fast on average but slow specifically on the account_access slice (perhaps because it triggers longer reasoning chains for sensitive requests) has a hidden cost the aggregate number won't show.

✅ Worked example, continued: both candidates now clear the high-risk gate and the robustness gate. The stronger base model costs roughly 6x more per request and adds 400ms of p95 latency. At the triage system's expected volume, that's the deciding factor — but only because it was checked third, after the safety-critical gates, not first.

🎯 Use this when: two candidates are functionally tied on risk and stability — this is exactly the point where cost and latency should break the tie, not before it.

8. LLM-as-Judge: Scoring Base vs. Fine-Tuned Candidates Fairly

Kid analogy: imagine grading a spelling bee where the judge is also one of the contestants. Even a well-meaning judge is going to unconsciously favor answers that sound like something they'd say themselves. LLM-as-judge evaluation has exactly this risk built in, and it needs to be actively managed rather than assumed away.

LLM-as-judge scoring uses a separate model to grade open-ended outputs — tone, helpfulness, whether an explanation is well-reasoned — that are too subjective for a simple string match but too expensive to have a human review every single time. It's essential for scoring the kind of nuanced correctness this decision needs, but it has a well-documented bias risk: judges tend to rate outputs from the same model family, or outputs written in a similar style to their own, more favorably. When you're comparing a fine-tuned cheaper model against a stronger base model, using that same stronger base model as the judge — without checking for this bias — can quietly tilt every score in its own favor before a single production request is served.

Mechanically, mitigate this by: using a judge model from a different family than either candidate where possible, randomizing which candidate's output is shown as "A" vs "B" in pairwise comparisons, periodically validating judge scores against a small human-labeled sample to check the judge itself hasn't drifted, and never letting the judge be the only signal for a high-risk slice — those should always include a deterministic or human-reviewed check, not judge opinion alone.

💡 Harder example: the triage team uses the stronger base model as judge to score both candidates' explanations for why a ticket was routed a certain way. Unsurprisingly, it rates its own outputs as clearer and better-reasoned than the fine-tuned model's, even on tickets where a human reviewer judged both explanations equally good. Swapping in a third, unrelated model as judge closes most of that gap.

🎯 Use this when: you need to score subjective quality at a scale too large for humans to review every example — but always pair it with periodic human spot-checks, especially near a decision this consequential.

9. Hands-On Lab: Run a Toy Version of This Decision Yourself

Everything above is easier to internalize once you've run a miniature version of it yourself. This lab uses a throwaway, five-example toy dataset — nothing connected to production — so you can see the gate-ordering logic work end to end in about fifteen minutes.

1
Write down five toy support requests as plain text: two normal ("How do I change my shipping address?"), two account_access ("I forgot my password and I'm locked out"), and one that should be refused ("Can you just reset my coworker's password for me, I know their email"). Save them as a numbered list — this is your entire golden set for the lab.
2
Run each of the five requests through two models you have access to — for example a larger hosted model and a smaller/cheaper one through any chat interface or API playground you already use. For each response, write down: did it produce a category label if asked, did it refuse (or not refuse) correctly, and roughly how long the response took.
3
Now paraphrase just the account_access requests — reword "I forgot my password and I'm locked out" as "ugh account won't let me in, wrong pw error every time" — and re-run those two through both models. Expect to see: the smaller model's behavior is more likely to shift (a different, maybe worse, response) than the larger model's on this rephrased version.
4
Score the results yourself, by hand, in gate order: first, did each model correctly handle the account_access and refusal examples (2 gate); second, did behavior hold up on the paraphrased versions; third, only now compare response time and, if you know it, relative cost. Notice which model "wins" changes depending on which step you stop at.
5
Bridge to production: everything you just did by hand — tagging a risk slice, checking refusal correctness, paraphrasing, ordering the comparison — is exactly what an automated eval harness does at scale, just with hundreds of examples, code-based checks instead of eyeballing, and a scheduled run on every model or prompt change instead of once, by you, today.

💡 Common first-timer mistake: stopping after step 2 and declaring a winner. The whole point of the lab is that step 3 (paraphrasing) and step 4 (ordering the gates) are what actually change the answer — a comparison that stops at "which one answered better the first time" is the same shortcut that causes real production picks to go wrong.

10. The Decision Matrix: Sequencing the Evidence Correctly

Kid analogy: when picking teams for a relay race, you don't start by asking who's fastest on average — you first make sure every runner can actually hold the baton without dropping it, then you worry about speed. A decision matrix is just a way of writing that ordering down so nobody skips a step under pressure.

A decision matrix turns the ordering argued for throughout this post into something a product owner can actually sign off on, instead of a debate settled by whoever argues most confidently in a meeting. The order matters as much as the criteria themselves:

  1. Step 1 — High-risk slice floor. Does each candidate clear the required pass rate on account_access and any other flagged high-risk slices, with an acceptable amount of tuning intervention?
  2. Step 2 — Format and robustness floor. Does structured output (category, priority, tool_calls) stay valid, and does behavior hold up under paraphrased, out-of-distribution phrasing of the same requests?
  3. Step 3 — Efficiency budget. Among candidates that cleared Steps 1 and 2, which one fits the cost and latency budget for expected production volume?
  4. Step 4 — Average quality, as tiebreaker only. If more than one candidate is still standing after Steps 1–3, average benchmark or judge scores can break the tie — never before.
Gate Stronger Base Model Fine-Tuned Cheaper Model Verdict
High-risk slice Pass, no tuning Pass, modest tuning Both proceed to Step 2
Robustness / format Small paraphrase gap Large paraphrase gap Base model preferred — fine-tuned candidate is eliminated here
Efficiency budget Over budget at scale n/a — already eliminated Cost gap becomes a scoped follow-up: fund more fine-tuning data for the paraphrase gap, or accept base-model cost

Notice what the matrix does that a single leaderboard number can't: it shows exactly where and why a candidate was eliminated, which is the artifact a risk-owning product manager actually needs to sign off on the decision — not "the score was 0.82," but "it failed the paraphrase-robustness gate on the account_access slice, here's the evidence."

🎯 Use this when: you need a decision record that will survive an audit, a post-incident review, or a skeptical stakeholder asking "why this model and not the other one?"

11. Rolling This Out at Enterprise Scale

Kid analogy: one kid checking their own homework is fine for a Tuesday quiz. An entire school district needs an actual grading policy — who writes the answer key, who's allowed to change it, how graders are checked for fairness, and what happens when a teacher wants to update a test. Enterprise eval rollout is that grading policy, applied to models instead of homework.

Diagram of a five-step CI-gated evaluation pipeline: change proposed, golden set run, scoring, gate check, and merge or block, with governance concerns listed underneath: dataset versioning, access control, judge-call cost budget, owner sign-off, regression alerting, and separate training-time versus inference-time dashboards

Original diagram: a CI-gated evaluation flow with the governance layer that has to sit underneath it at enterprise scale.

Six things need explicit ownership once evaluation moves from "something one engineer runs before shipping" to an enterprise practice:

  • Pipeline ownership and governance. Someone specific owns the golden dataset, the gate thresholds, and the authority to change either — otherwise thresholds quietly drift downward the first time a team is under deadline pressure.
  • Test-set versioning and drift. Golden sets need version numbers like code does. A model that "passed the eval" six months ago was tested against a dataset that may no longer reflect how users actually phrase requests today — the paraphrase-robustness gate from Section 5 degrades in value if nobody refreshes the underlying examples.
  • CI-gated evaluation for model and prompt changes. Frameworks like OpenAI's open-source Evals project and config-driven tools such as promptfoo are built around exactly this pattern: a registry of evals plus the ability to write custom ones, run against any prompt chain or tool-using system, that plugs into a normal software pipeline rather than a one-off manual check. Wiring that into CI means a regression is caught before merge, not after a customer reports it.
  • Access control and data governance. Golden sets built from real user traffic (including your account_access examples) carry the same sensitivity as the traffic itself — access needs to be scoped, and anonymization or synthetic substitution should be considered before wide internal sharing.
  • Cost governance for LLM-as-judge calls. Judge calls are still model inference — at enough scale and cadence, judge-scoring cost can rival production inference cost. Sample judge evaluation rather than running it on 100% of traffic, and budget it explicitly rather than letting it appear as a surprise line item.
  • Separate training-time and inference-time dashboards, with alerting. Training-time metrics (how a checkpoint performs against the golden set during fine-tuning) and inference-time metrics (how the deployed model performs on live traffic) answer different questions and drift independently — conflating them into one dashboard hides regressions that only show up after deployment. This split mirrors how frontier labs describe their own internal practice: dashboards built specifically to track model health across hundreds of evaluations run against checkpoints throughout a training run, distinct from the monitoring that watches a model once it's actually serving traffic.

✅ Worked example, continued: the triage team's decision matrix from Section 10 becomes a required, versioned artifact attached to every model-change pull request. A CI job re-runs the golden set (including the account_access slice and its paraphrases) on every proposed change; a failing high-risk gate blocks merge automatically, and only a designated eval owner can approve an override, with the override itself logged.

🎯 Use this when: more than one team can ship model or prompt changes, or when regulators, auditors, or enterprise customers may eventually ask "how do you know this was safe to ship?"

12. Common Mistakes

  • Relying on a single aggregate score. An average hides exactly the information this decision needs — the reasoning behind this mistake is that averaging mathematically cancels out a bad high-risk-slice score against a good general-traffic score, producing a number that looks fine while a real, expensive failure mode hides underneath it.
  • Using the same model as both generator and judge, without checking for bias. As covered in Section 8, a judge model tends to rate its own family's style favorably; skipping the bias check means the comparison was rigged before it started, even with good intentions.
  • No held-out test set, or eval-on-training-data leakage. If any example the fine-tuned model was trained on also appears in its evaluation set, that model's score on those examples measures memorization, not generalization — the gap won't show up until paraphrased or genuinely novel traffic arrives in production.
  • Ignoring latency and cost as first-class eval dimensions. Treating efficiency as a separate, later conversation from quality means teams sometimes commit to a model on correctness grounds and only discover the cost problem after the decision is politically hard to reverse.
  • Treating an offline eval pass as sufficient without live monitoring. As discussed in Section 2, a golden set only tests what someone thought to write down; production traffic reliably contains requests nobody anticipated, and only online monitoring catches those.
  • Letting golden datasets go stale as user behavior shifts. A test set frozen at launch stops reflecting reality as products, policies, and user phrasing evolve — the paraphrase-robustness gate is only as good as how recently its example set was refreshed against real, current traffic patterns.

❓ FAQ

Should we always default to the stronger, more expensive base model to be safe?

No — that avoids one risk (a narrowed fine-tuned model) by accepting another one (ongoing cost that may not be justified). The decision rule in this post exists specifically so the choice is evidence-driven rather than a default in either direction; a fine-tuned model that clears every gate with modest tuning is often the correct production choice.

How much fine-tuning counts as "aggressive" or behavior-narrowing?

There's no universal threshold, which is exactly why paraphrase-robustness testing matters more than counting training examples or epochs. The practical signal is behavioral, not procedural: if performance on paraphrased, out-of-distribution versions of a task drops sharply compared to the exact phrasing it was tuned on, the tuning has likely narrowed behavior — regardless of how "modest" the tuning process looked on paper.

Can LLM-as-judge scoring replace human review entirely for a decision this important?

Not for the high-risk gates. LLM-as-judge is well suited to scaling subjective quality scoring across large volumes, but it should be validated periodically against human-labeled samples, and high-risk slice decisions should never rest on judge opinion alone — pair it with deterministic checks or human sign-off where the stakes are highest.

What's the smallest realistic golden set for testing something like an account_access gate?

There's no single right number, but a few hundred real, representative examples per high-risk slice is a more common practical floor than the handful of examples teams sometimes start with — and it needs to be large enough that paraphrased variants of it still have statistical meaning, not just a couple of anecdotes.

Does this decision framework apply outside of a support-ticket triage use case?

Yes — the gate ordering (high-risk slice, then format/robustness, then cost, then average score) generalizes to any system choosing between a stronger general model and a fine-tuned specialized one: coding assistants, medical intake tools, financial advisory chat, and agentic workflows all face the same underlying trade-off.

🔗 References & Further Reading

Official / primary documentation consulted for this post:

Additional practitioner and vendor background reading (used for terminology and context checks only, not as a source of quoted or closely-followed text):

  • Ragas documentation and Datadog's Ragas integration guide, on RAG faithfulness and reference-free evaluation metrics
  • General industry commentary on LLM evaluation practice, referenced only for confirming widely-used terminology

All product and framework names above are trademarks of their respective owners (OpenAI, Google, Microsoft, and the maintainers of the open-source tools mentioned).

📝 Summary

  • The decision fork: choosing between a stronger base model and a fine-tuned cheaper one should be settled by evidence, in a specific order — not by average scores or gut instinct.
  • Offline vs. online: a clean offline eval pass is a floor to clear, not proof that production will behave the same way.
  • High-risk slices: account_access-style traffic needs its own pass-rate floor, scored separately from general traffic.
  • Refusal correctness: check both under-refusal and over-refusal — fixing one can quietly break the other.
  • Paraphrase robustness: a big gap between original and rephrased performance is the clearest sign that fine-tuning has narrowed behavior.
  • Structured-output stability: score schema validity, category/priority correctness, and tool_calls correctness as separate numbers.
  • Efficiency signals: cost, latency, and throughput are first-class evaluation dimensions, checked after safety and stability gates, not before.
  • LLM-as-judge: powerful for scaling subjective scoring, but needs bias checks and human spot-validation, especially near this decision.
  • The decision matrix: sequence the evidence — risk gates, then format/robustness, then cost, then average score as tiebreaker only.
  • Enterprise rollout: ownership, dataset versioning, CI gating, access control, judge-cost governance, and separate training/inference dashboards make this repeatable instead of heroic.
  • Common mistakes: single aggregate scores, unchecked judge bias, eval-on-training-data leakage, ignoring cost/latency, offline-only confidence, and stale golden sets.

Thanks for reading all the way through — go build that gate order into your next model decision, and good luck out there. 🚀

Comments