Prompt anatomy is the practice of breaking a prompt into five distinct parts — Instructions, Context, Input, Constraints, and Output spec — so a large language model can tell what you're asking for apart from what you're giving it and how you want the answer shaped. Most people write prompts as one undifferentiated paragraph and hope the model sorts it out; anatomy means you sort it out first. 🧩
This matters because the failure mode of a badly-structured prompt isn't a crash — it's a plausible-looking wrong answer. A support bot that can't tell your policy document (context) from the customer's question (input) will quote the wrong clause with total confidence. A code-review prompt that buries the "don't touch the config files" rule (a constraint) inside a wall of background text will get it ignored half the time. At production scale, across thousands of calls a day, an unstructured prompt doesn't fail loudly — it just quietly costs accuracy, money, and trust. 📉
📑 In This Post
1. Why Split a Prompt Into Five Parts? 🧠
🧸 Kid analogy: Think of a recipe card. It has a title of what you're making (instructions), a note about who it's for and why ("Grandma's birthday, she's allergic to nuts" — context), the actual groceries on your counter (input), the rule "don't let the oven go over 350°F" (constraints), and a picture of what the finished cake should look like (output spec). If you scrambled all of that into one paragraph of prose, you'd still be able to bake something — just not reliably the same cake twice.
A prompt works the same way. The five parts aren't a rigid syntax the model enforces — they're a mental (and often literal) checklist that keeps you from leaving something implicit that the model then has to guess. This isn't a fringe idea, either: OpenAI's own current API documentation spells out a standard structure for the messages it sends a model, typically in this order — an Identity section describing who the assistant is, an Instructions section defining the rules it must follow, an Examples section showing sample inputs and outputs, and a Context section carrying background information, usually placed near the end since it's the part most likely to change between requests. That's four labeled sections inside a single message, built on the same instinct this post organizes into five. Anthropic's own documentation makes a parallel case in different language, organizing its advice around a progression from plain clarity about the task, through examples and reasoning, up to explicit structural markers and role framing — several of which map directly onto Instructions, Context, and Output spec below.
✅ Worked example: "Summarize the article below in 3 bullet points. Article: """{article text}"""" follows the same delimiter pattern OpenAI's own developer documentation recommends for separating an instruction from the input it acts on — one of the simplest anatomy splits there is.
🎯 Use this when: you're debugging a prompt that "sometimes works" — nine times out of ten, the fix is separating a part the model was forced to infer.
2. Instructions — The Directive 📋
🧸 Kid analogy: Instructions are the one sentence you'd shout across the room if you only had five seconds: "Turn off the oven!" Not the backstory, not the ingredients — just the ask.
Real example: OpenAI's official developer documentation opens its prompting advice with exactly this principle: it contrasts a weaker version of a summarization request, where the ask is left implicit until late in the prompt, against a stronger version that puts the line "Summarize the text below as a bullet point list of the most important points" first, before any supporting material appears.
Mechanically, an instruction should name the task with a verb ("classify," "rewrite," "extract," "compare"), not describe the topic. "This is about customer complaints" is context masquerading as an instruction; "Classify this complaint as billing, shipping, or product-quality" is an instruction. That same OpenAI documentation makes a related point about phrasing: telling the model what it should do, rather than only what it shouldn't do, gives it a positive target to aim for — a negative instruction ("don't ask for the password") leaves the actual desired behavior undefined, so the model fills the gap unpredictably.
💡 Contrast: "Read this review and think about how the customer feels" gives the model a topic, not a task — it doesn't know whether you want a label, a paragraph, or a sarcastic aside. Compare this to the recipe-card analogy above: this is like handing someone groceries with no title on the card at all.
🎯 Use this when: outputs are inconsistent in *kind*, not just quality — one run gives a list, the next gives a paragraph. That's almost always a missing or buried instruction.
3. Context — The Background 🗂️
🧸 Kid analogy: Context is the note on the recipe card that says who the cake is for. The recipe (instruction) doesn't change, but knowing it's for someone with a nut allergy changes which version of the recipe actually gets followed.
Real example: GitHub's own blog documents how GitHub Copilot relies on exactly this layer in production. The post explains that a team can maintain a standing instructions file in a repository, covering things like which framework conventions to follow or how documentation should read, so that background applies automatically to every future request instead of being retyped by hand each time. The piece frames the underlying shift as being about what supporting material actually reaches the model, in a form it can use, rather than about phrasing any single request more cleverly.
Context answers "what does the model need to know that it can't infer?" — a persona ("you are a senior tax accountant"), a domain constraint ("this is for a UK audience, use British spelling"), or standing rules that apply across every call ("always cite the internal policy number"). The trap is treating context as a dumping ground: the more irrelevant background you add, the more the model has to figure out which parts actually matter to this specific request.
✅ Worked example: A copilot-instructions.md file scoped to "this repo uses React function components, Tailwind for styling, and Jest for tests" is context that applies to every future instruction in that repo — exactly the reusable-context pattern GitHub documents.
🎯 Use this when: the same instruction needs to behave differently depending on the audience, domain, or team — that's a sign the missing piece is context, not a better instruction.
4. Input — The Material to Act On 📥
🧸 Kid analogy: Input is the actual groceries sitting on the counter — the flour, eggs, and sugar the recipe will use. It's not the recipe, and it's not the note about who's eating it; it's the raw stuff getting transformed.
Real example: Anthropic's own documentation on structuring prompts recommends wrapping this component in a dedicated marker so the model never confuses it with the instruction sitting next to it. The core reasoning it gives is that keeping different parts of a prompt in visibly separate containers cuts down on the model blending an instruction with the material that instruction is supposed to act on, and it has the side benefit of making that material easy to locate again in code that assembles prompts automatically.
In practice, input is usually the one part of the prompt that changes on every single call — the customer's message, the document to summarize, the code diff to review — while instructions, context, and constraints often stay fixed as a template. Keeping input visually and structurally separate (a tag, a delimiter, a clearly labeled block) is what makes that template reusable at all.
💡 Contrast: Pasting that same review inline, mid-sentence, right after your instruction ("Classify this: The charger stopped working...") works fine for one review. It breaks down the moment the review itself contains the word "classify" or a stray quotation mark — the same ambiguity problem the recipe-card groceries would cause if they got mixed in with the written instructions.
🎯 Use this when: you're building a reusable template where only one section should change per call — isolate that section as input from day one.
5. Constraints — The Guardrails 🚧
🧸 Kid analogy: Constraints are the rule taped to the oven: "Never go over 350°F." The recipe still says what to bake; the constraint just narrows how far you're allowed to go while baking it.
Constraints cover length limits, tone, forbidden content, and behavioral boundaries — anything that shapes the space of acceptable answers rather than naming the task itself. OpenAI's own developer documentation treats vague constraints as a specific, named failure mode: it walks through a description request that only specifies "fairly short," and shows how replacing that with a concrete length range removes the ambiguity entirely. The lesson generalizes past length — a fuzzy constraint gives the model a fuzzy target, and fuzzy targets produce inconsistent output across runs even when the instruction itself never changes.
Role assignment is a form of constraint too. Telling a model up front that it's acting as a tax attorney or a security auditor doesn't change what task it's doing, but it anchors the lens it does that task through — which in turn shapes vocabulary, tone, and the kind of caveats it volunteers without your having to spell each of those out separately.
✅ Worked example: "Respond in under 40 words, in a neutral support-desk tone, and never promise a refund amount" stacks three constraint types — length, tone, and a hard behavioral boundary — onto the same underlying instruction from Section 2.
🎯 Use this when: the model is technically answering the right question but the *style* or *scope* keeps drifting between calls — that's a constraints gap, not an instructions gap.
6. Output Spec — The Shape of the Answer 📐
🧸 Kid analogy: The output spec is the photo on the recipe card showing what the finished cake should look like — round, two layers, white frosting. Without that picture, "bake a cake" could just as easily come back as cupcakes.
Where instructions say what to do, output spec says what shape the answer must arrive in — JSON with named fields, a numbered list, a single word, a specific XML structure. Google's own Gemini API documentation names this directly as a "response format" concern, distinct from the instruction itself, and its current guidance for Gemini 3 models goes further: it recommends placing output-format requirements either in the system instruction or at the very beginning of the prompt, and provides a template that tags a prompt's sections as role, constraints, context, task, and output_format — five labeled parts that line up almost one-to-one with the anatomy this post walks through. Anthropic's own documentation describes a related technique it calls prefilling: you write the opening of the assistant's reply yourself — Anthropic's own example is seeding a response with an unclosed { — and the model continues from there instead of drifting into a different shape or adding a preamble first.
💡 Contrast: Asking for "a JSON-ish summary" without naming the fields is the output-spec equivalent of the recipe photo being blurry — the model will produce *something* structured, but which keys it invents will vary run to run, which is exactly the kind of drift a downstream parser can't tolerate.
🎯 Use this when: a script or another system has to parse the model's reply automatically — an unambiguous output spec isn't a nicety there, it's the difference between a pipeline that runs and one that throws exceptions.
7. Rolling This Out at Enterprise Scale 🏢
A single well-anatomized prompt is a personal skill. Keeping hundreds of prompts anatomized correctly across a team, as models and requirements both keep changing underneath them, is an engineering problem — and it's the one GitHub's own blog points to when it frames standing instruction files and reusable prompt templates as team-level infrastructure rather than one-off tricks. The goal it describes is letting an organization fix its conventions once, in one place, instead of every engineer re-deriving the same background information inside every individual prompt they write.
- Own each of the five parts separately. Store instructions, context, and output spec as versioned template fragments, not one opaque string — so a context update (a new product name) doesn't risk silently touching the output spec.
- Version and diff prompts like code. A change to a constraint (a new length limit) should show up in a diff the same way a code change would, with a reviewer able to see exactly which of the five parts moved.
- Gate changes on a fixed test set. Before a prompt edit ships, run it against a golden set of real (or representative) inputs and compare output-spec compliance, not just "does it look fine" on one example.
- Restrict who can edit context that touches real user data. If context or input fragments are built from live customer records, access control belongs on the template, not just on the database it pulls from.
- Watch for drift when the underlying model changes. A prompt anatomized for one model version can silently degrade after a model upgrade — Anthropic's own newer-model guidance explicitly notes that prompting patterns don't always transfer unchanged across model generations, which is why a fixed regression suite matters more, not less, after an upgrade.
🎯 Use this when: more than one team or more than one engineer touches the same prompt — anatomy without versioning just relocates the ambiguity from the prompt text to your git history.
8. Common Mistakes ⚠️
Writing instructions and context as one paragraph. It reads naturally to a human, but the model has no reliable way to tell "what I want you to do" apart from "background you should know" — so it weights both equally, and sometimes treats background details as part of the task.
Describing constraints as negatives only. "Don't be too long" or "don't sound robotic" defines a boundary without defining the target inside it. OpenAI's own guidance frames this precisely: telling a model what to do instead of only what to avoid gives it a positive target to aim for, rather than an infinite space of "not that" to search through.
Leaving the output spec implicit because "it's obvious." It's obvious to you because you already know what downstream system will consume the answer. The model doesn't know that a script expects exactly three JSON keys unless the output spec says so — and a plausible-but-wrong format is often harder to catch than an outright wrong answer, because it still "looks" like a response.
Never revisiting a template after the model changes. A prompt tuned against one model version can degrade after an upgrade even when nothing else about the request changed, because the new model may follow the same wording differently. Treating a prompt as "done" the day it ships, rather than as something to re-test, is how quiet regressions creep in.
Testing on a handful of easy, hand-picked examples. A prompt that handles three clean sample inputs perfectly can still fall apart on the messy 5% of real traffic — the edge case with unusual formatting, a second language, or an adversarial user trying to override the instructions. Anatomy makes a prompt easier to test, but it doesn't replace actually testing it against realistic variety.
Skipping delimiters between input and the rest of the prompt. Without a clear boundary marker, user-provided input can be read as if it were part of the instructions — which is also the root of most prompt-injection issues, where malicious text embedded in an input tries to impersonate a new instruction.
❓ FAQ
Do all five parts need to appear in every single prompt?
No. A one-off, simple request might only need instructions and input. The five-part anatomy is most valuable for prompts that get reused, templated, or maintained by a team, where implicit assumptions are the most expensive to leave unstated.
Does the order of the five parts matter?
Placement conventions differ slightly by provider, but a consistent theme across official guidance is putting the core instruction early and keeping the input clearly delimited from everything else, rather than burying either one mid-paragraph.
Is this the same thing as "context engineering"?
They're related but not identical. Prompt anatomy is about structuring a single prompt's five parts clearly. Context engineering, as GitHub's own engineering team describes it, is the broader discipline of managing everything that surrounds a prompt at the system level — retrieval, memory, tool definitions, and reusable instruction files across an entire application.
Should I use XML tags, JSON, or plain delimiters to separate the parts?
Any consistent, unambiguous marker works. Anthropic's documentation specifically recommends XML-style tags for its models, while OpenAI's guidance shows simple delimiters like ### or triple quotes achieving the same separation. The mechanism matters less than picking one convention and applying it consistently across your prompts.
What breaks first if I skip the output spec?
Usually the integration layer, not the model itself. The answer may still read as reasonable to a human, but a script expecting a fixed JSON shape, a specific delimiter, or an exact label set will fail silently or throw parsing errors the moment the model's phrasing varies between calls.
🔗 References & Further Reading
- Anthropic — Use XML Tags to Structure Your Prompts
- Anthropic — Prompt Engineering Overview
- Anthropic — Prefill Claude's Response for Greater Output Control
- OpenAI — Prompt Engineering Guide
- OpenAI Help Center — Best Practices for Prompt Engineering with the OpenAI API
- Google — Gemini API: Prompt Design Strategies
- GitHub — Want Better AI Outputs? Try Context Engineering
All product names (Claude, Anthropic, OpenAI, GPT, Gemini, Google, GitHub, GitHub Copilot, Microsoft) are trademarks of their respective owners.
📝 Summary
- A prompt breaks cleanly into five parts: Instructions, Context, Input, Constraints, and Output spec.
- Instructions name the task with a verb and say what to do, not just what to avoid.
- Context supplies background the model can't infer — persona, domain, standing rules — without turning into a dumping ground.
- Input is the one part that usually changes per call, and needs a clear delimiter so it's never mistaken for an instruction.
- Constraints narrow the space of acceptable answers: length, tone, and hard behavioral boundaries.
- Output spec defines the shape of the answer, which matters most the moment another system has to parse it.
- At enterprise scale, each part should be versioned, diffed, and regression-tested independently, especially across model upgrades.
- Most "unreliable prompt" complaints trace back to one of the five parts being left implicit rather than to the model itself.
Next time a prompt gives you an inconsistent answer, don't reach for more words — reach for the missing part. Happy prompting! ✨
Comments
Post a Comment