Why Context Engineering Just Replaced Prompt Engineering — An Enterprise Architect's Take
Context engineering is the discipline of deliberately deciding what a language model sees, in what order, in what shape, and in what volume, before it ever generates a token. It replaced prompt engineering as the primary lever for reliability once production systems moved past single-turn chat into agents that call tools, read documents, and run for dozens of steps — because at that point, the wording of an instruction stopped being the thing that broke, and the composition of everything around that instruction started being the thing that broke. 🧭
This matters because most production LLM failures are context failures wearing a model-quality costume. A coding agent that "forgets" a constraint from ten turns ago, a research agent that contradicts itself over a long session, a support bot that answers confidently from the wrong document — in each case the underlying model is usually fine. It never received the right slice of information, in the right position, at the right moment. Teams that treat this as an engineering discipline — with real production patterns, not folklore — ship agents that survive a quarter of real traffic instead of one good demo. ⚙️
📑 In This Post
- What Is Context Engineering, Really?
- Why It Replaced Prompt Engineering as the #1 Skill
- The Anatomy of a Context Window
- The Four Moves of Context Engineering
- Production Case Study: Context Engineering Inside a Real Agent
- Enterprise Rollout at Scale
- Common Mistakes
- FAQ
- References & Further Reading
- Summary
🔀 Quick Comparison: Prompt vs. Context vs. Fine-Tuning
| Dimension | Prompt Engineering | Context Engineering | Fine-Tuning / Training |
|---|---|---|---|
| What changes | Wording of one instruction | Everything the model sees at inference: history, tools, retrieved data, memory | The model's weights themselves |
| Scope | Single turn, single call | Whole session, whole agent run | Every future call, permanently |
| Cost to change | Seconds — edit text | Hours — adjust retrieval/orchestration logic | Days to weeks — data, compute, evaluation |
| Fixes what kind of failure | "The model misunderstood my ask" | "The model didn't have the right facts, tools, or history" | "The model lacks a skill or behavior entirely" |
| Typical owner | Whoever writes the prompt | Applied-AI / platform engineering team | ML / model infrastructure team |
1. What Is Context Engineering, Really?
📌 Kid analogy first. Picture a kid heading to a group science project at a friend's house. If they show up with an overstuffed backpack — every notebook they own, three years of old homework, a comic book, and the actual project folder buried at the bottom — they'll waste the afternoon just finding the right page. Pack exactly the right folder and nothing else, and the project goes smoothly. Context engineering is deciding, ahead of time, exactly what goes in that backpack.
Technically: a language model has no memory between API calls. Every time it generates a response, it only knows what is physically present in its context window at that instant — the finite sequence of tokens fed in alongside the prompt. Context engineering is the deliberate design of how that window gets filled: what goes in, in what order, in what format, and, just as importantly, what gets deliberately left out.
The term has a fairly specific origin. Developer Simon Willison and former Tesla AI director Andrej Karpathy both began popularizing "context engineering" in mid-2025, as a more honest label than "prompt engineering," which by then had picked up a reputation as a synonym for typing clever phrases into a chatbot. Karpathy framed it as the careful discipline of loading a model's working memory with exactly what the next step requires — no more, no less. Shopify CEO Tobi Lütke amplified the framing shortly after. it has become standard vocabulary across agent frameworks and applied-AI teams.
✅ Worked example: Anthropic's engineering post on effective context engineering for agents frames every system prompt line, every tool, and every message as tokens competing for a model's finite attention — something that must earn its place in the window rather than being included by default. That single reframing (context as a scarce budget, not free storage) is the seed of everything else in this post.
🎯 Use this when: you're deciding whether a new failure in your AI app is a "prompt problem" (wrong instruction) or a "context problem" (missing or misplaced information). In production, it's the latter far more often than teams expect.
2. Why It Replaced Prompt Engineering as the #1 Skill
📌 Kid analogy first. Picture that same kid studying for a test with every note they've ever taken spread open on the desk at once — hundreds of loose pages, out of order, most irrelevant. Even a sharp kid starts making mistakes: skimming the wrong page, missing the one fact the question actually needed. More paper on the desk didn't make them a better student. It made them slower and more confused.
Transformers compute attention between every pair of tokens, a mechanism whose cost scales quadratically with sequence length. Practically, that means every additional token in the window is not free extra memory — it is competition for a limited attention budget. Even as context windows across frontier models grew toward the million-token range, longer did not mean automatically better. Practitioners now use the term "context rot" to describe measurable accuracy degradation as irrelevant, duplicated, or poorly ordered content accumulates in a long context, even when the information the model actually needs is technically still present somewhere in the window.
💡 Contrasting example: A common early mistake in building multi-step research agents is appending every tool result and every past turn straight into the prompt. Teams that do this consistently observe accuracy degrading on later steps of a long task compared with a version that actively prunes and restructures context between steps — the "overstuffed backpack" problem from the analogy above, now showing up as a measurable drop in task success rate.
This is the mechanical reason context engineering displaced prompt engineering as the primary lever for reliability. A perfectly worded instruction sitting on top of a context window full of stale, duplicated, or irrelevant information will still produce an unreliable answer. Fix the context, and the same instruction — often an unremarkable one — starts working consistently.
🎯 Use this when: your agent gets noticeably worse the longer a conversation or task runs. That is a strong signal of context rot, not a weaker model underneath.
3. The Anatomy of a Context Window
📌 Kid analogy first. Think of the context window as a lunch tray with fixed compartments: a slot for "the rules of the cafeteria" (system prompt), a slot for "tools you're allowed to bring" (like a calculator permitted for the test), a slot for "notes you looked up" (retrieved documents), a slot for "what happened earlier today" (conversation history), and finally the actual question being asked right now. Nothing outside the tray exists to the kid in that moment — and nothing outside the context window exists to the model. The infographic above shows a realistic split across these five compartments; in production, that split is exactly what a context-engineering layer is built to control.
The layers, in the order they're typically assembled into a single request:
- System prompt — fixed role, tone, and hard constraints, set once per deployment and rarely changed mid-session.
- Tool definitions — the names, descriptions, and parameter schemas of every function the model may call. Each one occupies real tokens whether or not it gets used on this turn.
- Retrieved context — document chunks pulled from a vector index or search system, relevant to the current query.
- Conversation / task history — prior turns, prior tool calls, and their results.
- The current user turn — the actual question or instruction for this step.
✅ Worked example: Anthropic's public write-up on its multi-agent research system describes a lead agent and several sub-agents, each working from a narrower, purpose-built slice of context rather than one shared, ever-growing transcript — a direct, production-scale version of the lunch-tray idea, where every agent gets its own tray, packed for its own job.
🎯 Use this when: you're debugging why an agent ignored a tool it clearly has access to. Check whether that tool's schema is actually present in this turn's assembled context, and where it sits relative to everything else competing for attention.
4. The Four Moves of Context Engineering
📌 Kid analogy first. A good student doesn't pack a backpack once at the start of the school year and never touch it again. They swap notebooks between classes, throw away a flyer they no longer need, and jot a reminder on a sticky note instead of carrying the whole textbook around. Context engineering techniques are exactly this: ongoing packing, unpacking, and repacking, performed automatically by the system rather than by hand.
LangChain's engineering team, reviewing a range of production agents and research papers, grouped these techniques into four recurring buckets: write, select, compress, and isolate. Each addresses a different failure mode of a growing context window.
Write — persist context outside the live window. Instead of keeping every plan and intermediate decision inside the prompt, an agent saves it to external storage: a scratchpad, a memory file, a database row. Anthropic's multi-agent researcher does this concretely — the lead agent writes its plan to persistent memory before spawning sub-agents, specifically because each sub-agent operates in its own, separate context window and could otherwise lose track of the overall plan if it only lived inside a transcript that might later be truncated or summarized away.
Select — pull in only what's needed, right when it's needed. Rather than preloading an entire knowledge base or an entire codebase into the prompt, the agent retrieves or reads the specific slice relevant to the current step, referenced by an identifier (a file path, a document ID) instead of embedded wholesale. A coding agent that greps a repository for the one relevant function, instead of pasting the whole codebase into context, is doing "select" — it treats the filesystem itself as the source of truth and pulls content in on demand.
Compress — shrink what's already there. As a task runs longer, older tool outputs and verbose intermediate steps get summarized down to their essential conclusions, freeing attention budget for what's happening now. A more aggressive variant, used by production agent Manus, is recitation: the agent maintains a running task list and rewrites it back into the end of its own context on every loop iteration, deliberately exploiting the model's bias toward recent tokens to keep the original goal from drifting out of focus over a long task.
Isolate — split context across separate windows. For large, multi-part tasks, specialized sub-agents each work in their own clean context window on a narrow piece of the problem, then return a condensed result to a coordinating agent, instead of everyone sharing one bloated transcript. The same idea shows up at the tooling level: running code execution or browser automation inside an isolated sandbox means the raw, noisy execution log never has to enter the main context — only the extracted result does.
💡 Contrasting example: Skipping compression on a long-running research agent is precisely how its context balloons with raw, unsummarized tool output in the first place — the fix is pruning the backpack between steps, not after the trip is already over.
🎯 Use this when: an agent needs to run for dozens of steps or tool calls. A single flat, ever-growing prompt will not survive that; pick at least compression plus external memory before you scale task length.
5. Production Case Study: Context Engineering Inside a Real Agent
General-purpose AI agent Manus published a detailed account of the context-engineering lessons behind its production system, after rebuilding its agent architecture four separate times. Two decisions frame everything that follows: Manus builds on top of frontier models' in-context learning rather than fine-tuning its own, specifically so that improvements to the underlying model translate into product improvements within hours instead of weeks; and the team treats context shaping as an experimental discipline, refined empirically rather than derived from a fixed formula. What follows is a stage-by-stage look at the concrete mechanics behind that system.
Stage 1 — KV-cache hit rate as the north-star metric
An autoregressive model generates one token at a time, and at each step it re-uses the internal key/value tensors it already computed for every earlier token in the sequence — a mechanism called KV-caching. If a new request's prefix is byte-for-byte identical to a prefix the engine has already processed, the inference engine can reuse those cached tensors instead of recomputing them from scratch. In an agent loop, where context grows every step but the model's own output stays short, Manus reports an input-to-output token ratio around 100:1 in production — meaning nearly all of the cost lives in re-processing input tokens, over and over, on every single step of the loop.
That is why cache hit rate becomes the single most important cost and latency lever for a production agent, not a minor optimization. The practical rule that falls out of the mechanism: keep the context prefix stable and append-only. Never edit, reorder, or reformat anything earlier in the sequence once it has been sent, because a single byte of difference anywhere in the prefix invalidates the cache for everything after it. Two subtle, easy-to-miss ways teams break this in practice: serializing structured data (like tool arguments) with a JSON library that doesn't guarantee stable key ordering between runs, and embedding a live timestamp inside a system prompt that is supposed to be a cached, unchanging prefix.
# Illustrative pattern for cache-friendly context assembly (original, not from any real codebase)
context = [system_prompt, tool_schemas] # fixed prefix, never edited in place
for step in agent_loop:
context.append(observation) # append-only, preserves the cached prefix
allowed = mask_logits(tool_schemas, state) # restrict choices without editing context
action = model.generate(context, logit_mask=allowed)
context.append(action)
Stage 2 — mask, don't remove
As an agent's toolset grows — especially once it becomes user-configurable — the obvious instinct is to dynamically add or remove tool definitions from context depending on what's relevant to the current step, similar to how RAG selects documents. Manus deliberately avoids this. Editing the tool list mid-session changes the prefix, which invalidates the KV-cache for the rest of the conversation from that point forward, and it can also directly confuse the model when an earlier action in the transcript references a tool that has since disappeared from the visible list. Instead, the full tool list stays fixed inside the cached prefix, and the system restricts which tools are actually legal to choose at a given step by masking the model's output logits during decoding — blocking disallowed tokens at generation time rather than removing information from what the model can see.
Stage 3 — the file system as unlimited, restorable context
Long-running tasks generate large artifacts — downloaded web pages, generated files, big API responses — that are too large to keep in the live context window indefinitely, but summarizing them away permanently risks losing detail the agent needs later. Manus's answer is to write large intermediate results to disk and keep only a lightweight reference, such as a file path, inside the actual context. This gives the agent effectively unbounded external memory while the live window stays small, and critically, the compression is restorable: the agent can re-open the file and read the full content again if a later step needs it, unlike an irreversible summary that has already thrown detail away for good.
Stage 4 — recitation to fight goal drift
On long tasks, the original objective can end up buried near the start of a long transcript — exactly the position most vulnerable to "lost in the middle" effects. Manus counters this by having the agent maintain a running todo-style plan and re-write it back into context near the end of the sequence on every loop iteration. Because models weight recent tokens more heavily, this recitation keeps the global goal inside the model's effective attention even dozens of steps into a task, without needing a bigger context window at all.
Stage 5 — keep failures visible, and avoid pattern lock-in
Two smaller but counter-intuitive lessons round out the case study. First, Manus deliberately leaves failed tool calls and error traces inside the context rather than scrubbing them out to keep the transcript "clean." Seeing a failed action and its error is what lets the model implicitly avoid repeating that same mistake on the next step — a lightweight, in-context form of learning from failure that disappears the moment the evidence is deleted. Second, formatting every step of a repetitive task in exactly the same way, over and over, can cause a model to lock onto the rhythm of the pattern rather than reason about the actual task in front of it, so Manus deliberately introduces small, controlled variation in phrasing and structure to prevent that drift.
🎯 Use this when: you're designing the core agent loop for a system that will run many steps per task, not just answer single questions. Cache-hit rate, masking, and recitation are loop-level decisions that are expensive to retrofit later.
6. Enterprise Rollout at Scale
Moving context engineering from one team's agent to shared infrastructure used by many teams means treating context itself as a governed asset, independent of which cloud, vector database, or orchestration framework sits underneath. The practical checklist:
- Data governance for retrieval corpora: classify and access-control whatever documents feed retrieval the same way you'd classify any other enterprise data source — a pipeline that can retrieve a document into context can also leak it.
- Tenant isolation for agent memory: keep different applications' retrieval indexes, scratchpads, and conversation logs isolated from one another once several teams share the same underlying context-engineering platform.
- Context and prompt versioning: treat system prompts, tool schemas, and retrieval configuration as versioned artifacts with staged rollout and rollback, not files edited directly in production.
- Cost governance via cache-hit rate: as Section 5 shows, an unstable context prefix doesn't just add latency — it can multiply token cost by an order of magnitude. Cache-hit rate belongs on the same dashboard as raw token spend.
- Observability: track context length per call, retrieval precision, and cache-hit rate alongside downstream answer quality. A spike in context length paired with a dip in accuracy is the operational signature of context rot in production, regardless of which vendor's models or infrastructure you're running on.
🎯 Use this when: more than one team starts building agents against the same retrieval sources or agent runtime. That's the point where "our prompt" needs to become "our platform."
🚧 7. Common Mistakes
- Loading an entire knowledge base "just in case." The reasoning is usually "more context can only help." It doesn't — irrelevant documents compete for attention with the facts that actually matter, and can actively pull the model toward a wrong answer.
- Never pruning conversation history. Teams append every turn forever because it's the easiest thing to implement. The agent then spends attention budget re-reading small talk from twenty turns ago instead of the current task.
- Treating a huge context window as a substitute for good curation. A million-token window means a model can hold a lot of text, not that stuffing everything in is efficient, cheap, or accurate — context rot doesn't wait for the window to actually fill up.
- Silently breaking the KV-cache. Non-deterministic serialization of structured data, or a live timestamp embedded in a "static" system prompt, can quietly turn every request into an uncached one — often showing up only as an unexplained jump in cost and latency.
- Dynamically adding and removing tools mid-session. This feels like good context hygiene, but it invalidates the cache for everything after the change and risks the model referencing a tool that's no longer visible. Masking legal choices at decode time is usually the safer pattern.
- Scrubbing errors out of the transcript. Removing failed actions to keep the context "clean" also removes the evidence the model needs to avoid repeating that exact mistake on the next step.
- No context or prompt versioning. Editing a shared system prompt directly in production, with no staged rollout, means a single bad edit silently degrades every agent depending on it, with no easy way back.
- Confusing context engineering with fine-tuning. Teams sometimes reach for expensive retraining to fix a problem that a better-assembled context window would have solved for free, at a fraction of the cost and turnaround time.
❓ FAQ
Is context engineering just a rebrand of prompt engineering?
No. Prompt engineering optimizes the wording of one instruction. Context engineering covers everything else that enters the window — retrieved documents, tool schemas, memory, and history — and treats prompt wording as just one input among several.
Does a 1-million-token context window make context engineering unnecessary?
No, and often the opposite. Larger windows lower the risk of hard truncation, but "context rot" means accuracy can still degrade as irrelevant content accumulates — curation stays necessary even when there's technically room for everything.
Is context engineering the same thing as RAG?
RAG is one technique inside context engineering — specifically the "select" move. Context engineering also covers tool schemas, memory, compression, and multi-agent isolation, which RAG alone doesn't address.
What is KV-cache hit rate, and why should a non-infra person care?
It's the share of a request's input tokens the inference engine can reuse from a previous, identical prefix instead of recomputing. In agent loops with a high input-to-output token ratio, it is often the single largest lever on both cost and latency — roughly a 10x cost difference between cached and uncached tokens in published examples.
Where should I start if my agent's accuracy drops on long conversations?
Start with compression and external memory (Section 4) before reaching for a bigger context window or a fine-tuned model — it's almost always the cheaper and more effective fix.
🔗 References & Further Reading
- Anthropic Engineering — "Effective context engineering for AI agents," anthropic.com/engineering
- Anthropic Engineering — "How we built our multi-agent research system," anthropic.com/engineering
- LangChain — "Context Engineering for Agents," blog.langchain.com
- Manus AI — "Context Engineering for AI Agents: Lessons from Building Manus," manus.im/blog
- Simon Willison — "Context engineering," simonwillison.net (June 2025)
All product names, trademarks, and registered trademarks (Anthropic®, Claude®, LangChain®, Manus®, and others referenced above) are the property of their respective owners. This post synthesizes and explains publicly available information in original wording; it does not reproduce any source's text, structure, or FAQ framing verbatim.
📝 Summary
- Context engineering is deliberately designing everything an LLM sees before it generates a response — not just the prompt.
- It overtook prompt engineering because attention is a finite, quadratic-cost resource; more tokens can make a model worse, not better.
- A context window is assembled from layered sources: system prompt, tools, retrieved documents, history, and the current turn.
- Four recurring techniques — write, select, compress, isolate — cover most production context-engineering patterns.
- Manus AI's production case study shows the mechanics in practice: cache-friendly append-only context, masking tools instead of removing them, the filesystem as restorable memory, recitation against goal drift, and keeping failures visible.
- At enterprise scale, context becomes a governed asset regardless of vendor: data governance for retrieval, tenant isolation, versioning, cache-aware cost governance, and observability tied to accuracy.
That's context engineering, end to end — from the backpack a kid packs for a school project to the cache-aware production loop running a real AI agent. Pack it well, and the model finally gets a fair shot at being as good as it actually is. 🎒
Comments
Post a Comment