Context Engineering: Context Budgets, Compression & AI Agent Context Management
Context is not an unlimited information bucket. It is a limited working space that an AI model uses to understand the task, examine evidence, follow instructions and produce an answer. That creates an engineering problem: when an agent has more information than it can reasonably use, what should stay, what should disappear, what should be compressed, and what should be fetched again? 🧠
This is where context budgets matter. A context budget is the deliberate allocation of the model's available working space across instructions, task state, evidence, conversation history, tool results, memory and the expected response. The goal is not to put as much information as possible into the model. The goal is to put the right information in at the right time.
- What Is a Context Budget?
- Making Room for Instructions, Evidence, and the Answer
- Why More Information Can Reduce Answer Quality
- What Happens When Context Exceeds the Model's Limit?
- How to Decide What to Keep, Remove, or Fetch Again
- Context Compression: Preserving Meaning in Less Space
- How Summaries Lose Facts, and How to Check Them
- When to Return to the Original Source
- Complete Worked Example: A Support Agent
- Enterprise Context-Budget Architecture and Rollout
- Common Mistakes and Why They Fail
- FAQ
- References & Further Reading
- Summary
1. 🧠 What Is a Context Budget?
Definition: a context budget is the amount of model-working capacity that an application deliberately allocates to the different kinds of information needed during an AI task.
The word budget fits because context has competing demands. Instructions need space. Evidence needs space. Recent task state, tool results and memory need space. The model also needs room to produce its response, and for some model families, internal reasoning counts toward the same request capacity.
A context window is the technical capacity of a model for a request. A context budget is the application's planning decision about how to use it. The question is not “How much context can this model accept?” but “How much context should this task receive?”
Original diagram: percentages are for illustration, not a recommended formula.
| Context component | Purpose | Typical question |
|---|---|---|
| Instructions | Define the task and operating rules | What must the agent do and avoid? |
| Task state | Describe what is happening now | What has already happened in this task? |
| Evidence | Provide facts needed for the decision | Which information actually supports the answer? |
| Tool results | Provide fresh observations from systems | What did the external system actually return? |
| Memory | Carry useful information across steps or sessions | What earlier information is still relevant? |
| Answer space | Leave room for the model's response | How much output does this task need? |
In code, a budget is a policy the application enforces, not a comment in a prompt. A minimal version gives every item a priority, protects the ones that must never be dropped, records everything it removes, and returns what it kept in its original order:
def assemble_context(items, window_tokens, reserve_output, count_tokens, log):
limit = window_tokens - reserve_output # answer space is reserved first
chosen, used = set(), 0
for item in sorted(items, key=lambda i: i.priority): # priority 0 = protected
cost = count_tokens(item.text)
if used + cost <= limit:
chosen.add(item.id)
used += cost
elif item.priority == 0:
raise ContextBudgetError(f"protected item {item.id} does not fit")
else:
log.record("dropped", item.id, item.source_id, cost) # never drop silently
kept = [i for i in items if i.id in chosen] # restore the intended order
return kept, used
This is a design pattern, not a universal formula. A real system would try compression before dropping an item, and it would count tokens with the provider's own tokenizer (see section 4).
🎯 Use this when your agent repeatedly carries large histories, retrieval results, memory or tool responses into its next step.
2. 📦 Making Room for Instructions, Evidence, and the Answer
An agent's context is shared by competing needs, and allocation should follow the task. The model should not spend its working capacity separating today's evidence from yesterday's noise because the application failed to curate.
Order matters too. Keep stable, protected material (operating instructions, fixed policy) at a consistent position at the start, put the task-specific evidence next, and place the question and any must-not-miss constraints near the end. Many providers also reuse work for repeated prefixes, so a stable beginning can lower cost and latency.
🎯 Use this when the same context holds instructions, evidence, history and tool data that compete for attention.
3. 🔎 Why More Information Can Reduce Answer Quality
It is tempting to assume more context must produce a better answer. That assumption is unsafe. Additional information can add:
- Distraction: irrelevant material competes with relevant material.
- Duplication: the same fact appears repeatedly in different forms.
- Conflict: old and current information disagree.
- Ambiguity: sources describe similar concepts differently.
- Noise: large tool outputs contain details unrelated to the decision.
- Stale assumptions: earlier task state is carried forward after circumstances changed.
Research supports being careful here. The paper “Lost in the Middle” found that language models often use information less reliably depending on where it sits in a long input, with weaker use of material buried in the middle. Effects vary a lot between models and tasks, and newer models may behave differently, so treat this as a reason to measure your own task rather than a fixed rule. It also does not mean large windows are bad: they are very useful when the extra information is the evidence the task needs.
Context quality depends on both coverage (does it contain what the task needs?) and signal-to-noise ratio (how much else comes with it?).
| Context strategy | Possible result |
|---|---|
| Too little | The model lacks facts required to answer. |
| Too much but well curated | Useful coverage, with higher cost and processing time. |
| Too much and poorly curated | Irrelevance, conflicts and duplication make the task harder. |
| Focused and sufficient | The evidence the task needs, without the baggage. |
🎯 Use this when your team keeps adding retrieved documents or conversation history whenever answer quality drops.
4. 🚦 What Happens When Context Exceeds the Model's Limit?
Definition: a context-window limit is the maximum amount of tokenized information a particular model and interface can process in a request, under that system's rules. Limits differ across models and APIs, and some systems split capacity between input, output and internal reasoning. Never assume the full published capacity is available for retrieved documents.
Count tokens the way your provider does. Different models tokenize text differently, so a character-count guess or a counter built for another model can be badly off. Use the provider's own token-counting facility, and recount after every compression step.
What happens at the limit depends on the application and API. Common behaviors:
- The request is rejected because it is too large.
- The application removes older or lower-priority context before sending.
- The application summarizes or compacts earlier state. Some platforms offer built-in compaction for this.
- The task is split across several model calls.
- The response is capped to preserve capacity.
Overflow handling should be designed deliberately. An agent should not rely on accidental truncation to decide which information disappears.
A production application should track at least:
- Estimated input tokens, and reserved output capacity.
- Size per context category, and number of retrieved items.
- Compression operations performed.
- Items removed or truncated.
- Whether critical evidence was preserved.
🎯 Use this when long-running conversations or retrieval pipelines can grow beyond predictable size.
5. 🔄 How to Decide What to Keep, Remove, or Fetch Again
Original diagram: the questions run left to right; outcomes sit beneath them. Escalate is reachable from any step.
A useful policy has five outcomes:
| Decision | Meaning | Example |
|---|---|---|
| Keep | Required and still valid. | Current contract terms needed for the decision. |
| Remove | Irrelevant, expired or redundant. | An unrelated earlier conversation. |
| Compress | Still useful, but can be smaller. | A long conversation turned into checked task state. |
| Fetch again / fetch original | The original is more reliable than a carried copy: it may have changed, or the decision is high impact. | Re-reading the current policy before a high-impact decision. |
| Escalate | The context cannot safely establish the required fact. | A required document is missing or contradictory. |
Context is a dynamic workspace, not a permanent transcript.
🎯 Use this when the agent's context grows continuously during multi-step work.
6. 🗜️ Context Compression: Preserving Meaning in Less Space
Definition: context compression is representing useful context in a smaller form while preserving what the next task needs. It does not have to mean asking a model to “summarize this.” Forms include:
- Summarization: a long narrative into a shorter one.
- Extraction: keeping only the structured facts the task requires.
- Deduplication: removing repeated information.
- State conversion: turning a long conversation into explicit task state.
- Filtering: removing material outside the current scope.
- Reference preservation: keeping identifiers so the original can be fetched later.
| Original context | Compressed representation | Must preserve? |
|---|---|---|
| Long customer conversation | Issue, affected product, actions tried, current status | Yes, if needed to continue |
| Ten repeated tool responses | Latest verified state plus source identifiers | Usually |
| Old irrelevant conversation | Nothing | No |
| Exact policy clause | Short paraphrase plus source and version reference | Depends on decision risk |
The last row matters most: some information is safe to compress for navigation but unsafe to compress for a final decision. What should usually survive: critical facts, numbers and units, dates and deadlines, exceptions and conditions, decisions already made, unresolved questions, source identifiers, stated uncertainty, and scope.
A security point that is easy to miss. A summarizer reads whatever it is given, including untrusted text. If a retrieved web page or customer message contains an embedded instruction (prompt injection), compressing it into “verified task state” can strip away its origin and make it look like trusted guidance. Mark compressed content with the trust level of its sources, store facts rather than instructions, and never promote a summary into the instruction layer.
A compressed context that drops its source identifier is shorter but harder to verify. Design compression and provenance together.
🎯 Use this when the same task history is becoming too large to carry forward efficiently.
7. 🧪 How Summaries Lose Facts, and How to Check Them
A summary is a transformation, and transformation can lose, weaken or change information. Common losses: numbers (exact values vanish), dates (a deadline becomes “soon”), conditions (“only if” becomes “if”), exceptions, uncertainty (“possibly” becomes fact), attribution, scope (a one-region rule looks universal), version, and conflict between sources.
| Original statement | Risky summary | What was lost? |
|---|---|---|
| “Requests submitted within 30 days are eligible unless the product belongs to category X.” | “Requests within 30 days are eligible.” | The exception for category X. |
| “The reported value is approximately 18.4, based on an incomplete measurement.” | “The value is 18.4.” | Uncertainty and the measurement limitation. |
| “Policy version 7 applies from 1 October.” | “The current policy says...” | Version and effective date. |
These summaries are shorter, but they are not equivalent to the originals. A summary can be fine for navigation and still be inadequate for authorization, compliance, financial calculation or any high-consequence decision.
How to check a summary. The most reliable checks are deterministic. Extract the critical items from the original (numbers, dates, version identifiers, condition words) and confirm each appears in the summary. A crude check like this catches dropped values and dropped conditions:
import re
NUMBERS_DATES = re.compile(
r"\d{1,2} [A-Z][a-z]+ \d{4}" # dates first
r"|\d+(?:[.,]\d+)*(?:\s?(?:%|days?|hours?))?" # numbers, no trailing period
)
CONDITION_WORDS = ("unless", "except", "only if", "provided that")
def missing_critical_facts(original, summary):
needed = {m.group(0).strip() for m in NUMBERS_DATES.finditer(original)}
problems = [f"missing: {n}" for n in sorted(needed) if n not in summary]
if any(w in original.lower() for w in CONDITION_WORDS) \
and not any(w in summary.lower() for w in CONDITION_WORDS):
problems.append("a condition or exception may have been dropped")
return problems
This is a floor, not a guarantee: it cannot judge meaning. A second model reading the summary against the source is a useful extra signal, but it is also probabilistic, so for high-impact workflows show the evidence itself instead of trusting any summary check.
Store facts, not verdicts. A compression record that says “request appears eligible” invites later steps to reuse a conclusion as if it were a fact. Keep the facts, mark each with its source and who or what checked it, and label any conclusion as provisional:
{
"facts": [
{"claim": "Request submitted on day 24 of the 30-day window",
"source_id": "case-4471", "checked_by": "tool:case_lookup"},
{"claim": "Policy 2026-07 allows replacement within 30 days",
"source_id": "policy-2026-07", "source_version": "7", "checked_by": "retrieval"}
],
"open_questions": ["Does the exception for category X apply to this product?"],
"provisional_conclusion": "Possibly eligible, pending the category check",
"trust": "derived_from_mixed_sources",
"original_available": true,
"recheck_required": true
}
The structure is illustrative. What matters is that the record says what was summarized, where it came from, who checked it, and what still needs verification.
🎯 Use this when your agent compresses long conversations, documents, tool histories or memory before carrying them into later steps.
8. 🔍 When to Return to the Original Source
Return to the original source when:
- The exact wording matters.
- A financial, legal, regulatory or operational decision depends on it.
- The summary contains uncertainty, or disagrees with another source.
- An important exception may apply.
- The source has changed since the summary was made, or its version or effective date matters.
- The agent needs evidence that can be cited or audited.
- The summary cannot establish the required fact.
- The proposed action is hard to reverse or has significant consequences.
Use the summary for continuity. Use the original source for verification.
This is not absolute. Low-risk tasks may not need the source every time. Weigh the task's consequences, the reliability of the summary, the freshness of the source and the cost of retrieval.
| Situation | Reasonable context strategy |
|---|---|
| Low-risk conversational continuity | A checked summary may be sufficient. |
| Exact factual lookup | Retrieve the relevant source when precision matters. |
| Conflicting information | Return to authoritative sources and resolve the conflict. |
| High-impact action | Verify critical evidence against current authoritative sources before executing. |
| Missing source | Enter an information-gap state rather than inventing certainty. |
🎯 Use this when a compressed representation is being used to justify an action rather than merely continue a conversation.
9. 🏗️ Complete Worked Example: A Support Agent
Scenario: a customer asks, “Can I get a replacement for this product, and can you process it now?” The agent can reach the current service policy, the customer's contract, the product record, the earlier conversation, a shipment-status tool, an old policy document and a previous agent summary.
- Define the task. There are two jobs: decide eligibility, and possibly execute a replacement.
- Build the initial context. Current policy, applicable contract terms, relevant product facts, and only the conversation needed to understand the issue.
- Exclude the rest. The old policy must not compete with the current one; unrelated history takes no space.
- Use the previous summary carefully. It gives continuity, but eligibility facts must trace to original sources.
- Fetch fresh state. Call the shipment or product tool instead of trusting an old statement.
- Check the budget. Drop duplicated history and irrelevant tool output; compress long conversation into checked task state.
- Verify critical evidence. If eligibility depends on an exception clause, re-read the current policy section.
- Separate answer from action. “Appears eligible” does not grant permission to execute.
- Apply authorization. Execution needs its own permission or approval, enforced outside the model.
Here is what that does to the budget, using a hypothetical 32,000-token window (numbers are illustrative):
| Component | Naive assembly | Curated assembly |
|---|---|---|
| Operating instructions | 1,500 | 1,500 |
| Policies | 9,000 (current plus ten old versions) | 1,200 (current section only) |
| Conversation history | 12,000 (twenty conversations) | 1,500 (recent turns) plus 800 (checked task state) |
| Contract and product data | 3,500 (full records) | 900 (relevant excerpts) |
| Tool output | 6,000 (raw) | 600 (needed fields) |
| Memory | 2,500 | 300 |
| Input total | 34,500, over the limit | 6,800 |
| Reserved for answer | none left | 2,000 |
The curated version uses well under a third of the window and is easier to audit. The unused headroom is intentional. The original sources stay retrievable for verification.
🎯 Use this when an agent must maintain a long-running task while still deciding from current authoritative information.
10. 🏢 Enterprise Context-Budget Architecture and Rollout
In production, context budgeting is a pipeline, not one large prompt template.
| Layer | Responsibility |
|---|---|
| Task analysis | Determine what the current task actually requires. |
| Retrieval | Find candidate evidence from approved sources. |
| Selection | Filter by relevance, authority, scope, freshness and access rules. |
| Compression | Reduce size while preserving required meaning, provenance and trust labels. |
| Budgeting | Ensure the assembled context leaves room for the response and other processing. |
| Verification | Return to authoritative sources when precision or risk requires it. |
| Action control | Keep authorization and consequential action enforcement outside the model. |
Useful observability metrics:
- Context size before selection, after selection, and after compression.
- Percentage of retrieved items discarded.
- Number of source re-fetches.
- Summary verification failures.
- Context-overflow events.
- Tool calls triggered after compressed context.
- Cases where a human reviewer asked for the original source.
- Cost and latency of context preparation.
These measurements turn context engineering from informal prompt-writing into something that can be tested and improved.
- Ownership: name an owner for the budget policy, for the compression rules and for the source-of-truth systems they depend on.
- Versioning and change control: treat budgets, priority classes, compression logic and verification rules as versioned production artifacts. A change can alter what the agent sees without any code change elsewhere.
- Release gate: before shipping a change, check what new information can now enter or be dropped, whether any protected item could be crowded out, what the summary checks cover, and whether tracing still explains every drop.
- Cost governance: set token budgets and alerts per workflow. Re-fetching, long contexts and extra compression calls all cost money and latency.
- Kill switch: keep a tested way to switch compression off and fall back to fetch-from-source or human escalation if summaries start causing wrong decisions.
Enterprise note: the person who tunes the budget to save cost should not be the only person who can approve dropping evidence that decisions depend on.
11. ⚠️ Common Mistakes and Why They Fail
1. Treating the context window as a target to fill. A larger capacity is a limit, not a recommended payload size.
2. Keeping the entire conversation forever. Long conversations hold repeated questions and obsolete assumptions. Preserve task state deliberately instead of using the transcript as the only memory.
3. Compressing everything automatically. Some information holds exact conditions, numbers or exceptions that should be verified against the original.
4. Using summaries as permanent authority. A summary represents a source. It does not inherit the source's authority.
5. Losing provenance during compression. If the system keeps a conclusion but forgets where it came from, later verification is hard.
6. Compressing untrusted content into trusted state. A summary can strip the origin from injected text. Carry trust labels with compressed content.
7. Ignoring freshness. A compact summary can outlive the policy, customer state or record it describes.
8. Silently truncating context. When something must go, the application should know what disappeared and log it.
9. Miscounting tokens or forgetting the answer. Counting with the wrong tokenizer, or not reserving output space, causes overflows or clipped answers.
10. Fetching everything again. The opposite mistake: re-retrieving large sources when a checked compact state would do. Balance freshness, accuracy, cost and relevance.
11. Treating context management as a model-only problem. A model cannot enforce database permissions or approval rules because they were described in text. The surrounding application must.
12. ❓ FAQ
13. 🔗 References & Further Reading
- OpenAI — Conversation state
- OpenAI — Compaction
- OpenAI — Counting tokens
- OpenAI Cookbook — Context Engineering: Short-Term Memory Management with Sessions
- Anthropic — Effective context engineering for AI agents
- Liu et al. — Lost in the Middle: How Language Models Use Long Contexts (Transactions of the Association for Computational Linguistics, 2024)
- OWASP — Top 10 for LLM Applications (prompt injection is LLM01)
- NIST — AI Risk Management Framework
- NIST — AI RMF: Generative AI Profile (NIST AI 600-1)
Originality & attribution: Product and standards names belong to their owners. Platform features change quickly, so check current documentation before relying on a specific capability.
14. 📝 Summary
- Context budget: the deliberate allocation of working space among instructions, evidence, task state, memory, tool results and output.
- More is not automatically better: extra information can add distraction, duplication, stale facts and conflicts.
- Limits matter: count tokens with the provider's tokenizer and reserve answer space; never assume the whole window is for input.
- Keep, remove, compress, fetch or escalate: manage context actively by relevance, freshness, authority and scope.
- Compression is transformation: a shorter form helps only if the facts the next decision needs survive, with sources and trust labels attached.
- Check summaries: numbers, dates, conditions, exceptions, uncertainty, scope and version are what get lost; check them deterministically where you can.
- Verify high-impact decisions: use the summary for continuity and the original source for verification.
- Govern it: own, version, observe and be able to switch off your budget and compression logic.
Think of context as a working desk, not a warehouse. Put today's instructions on the desk. Bring the evidence today's task needs. Remove what no longer matters. Compress what must stay but not in full. And when an important fact becomes uncertain, go back to the original source.
A powerful agent is not created by giving the model more information. It is created by engineering a context that gives the model the right information, at the right time, in the right form, with enough room to reason and respond. That is the difference between having a large context and managing context intelligently.
Comments
Post a Comment