Skip to main content

Context Engineering: Context Budgets, Compression & AI Agent Context Management

Calculating read time…

Context is not an unlimited information bucket. It is a limited working space that an AI model uses to understand the task, examine evidence, follow instructions and produce an answer. That creates an engineering problem: when an agent has more information than it can reasonably use, what should stay, what should disappear, what should be compressed, and what should be fetched again? 🧠

This is where context budgets matter. A context budget is the deliberate allocation of the model's available working space across instructions, task state, evidence, conversation history, tool results, memory and the expected response. The goal is not to put as much information as possible into the model. The goal is to put the right information in at the right time.

1. 🧠 What Is a Context Budget?

🧒 Kid Analogy
Imagine you have one school bag. You need to carry your homework, lunch, books, pencils and today's project. If you fill the entire bag with old notebooks, there may be no room for the things you actually need today.
✅ Practical idea
A simple lookup (“What is this order's status?”) needs only a few hundred tokens of context. A contract-dispute investigation may legitimately need tens of thousands. Both are well budgeted, because each got what its task required and nothing more.

Definition: a context budget is the amount of model-working capacity that an application deliberately allocates to the different kinds of information needed during an AI task.

The word budget fits because context has competing demands. Instructions need space. Evidence needs space. Recent task state, tool results and memory need space. The model also needs room to produce its response, and for some model families, internal reasoning counts toward the same request capacity.

A context window is the technical capacity of a model for a request. A context budget is the application's planning decision about how to use it. The question is not “How much context can this model accept?” but “How much context should this task receive?”

A stacked bar showing one illustrative split of a context window: instructions 10 percent, task state 8, evidence 25, recent history 12, tool results 15, memory 5, reserved answer space 15 and headroom 10

Original diagram: percentages are for illustration, not a recommended formula.

Context componentPurposeTypical question
InstructionsDefine the task and operating rulesWhat must the agent do and avoid?
Task stateDescribe what is happening nowWhat has already happened in this task?
EvidenceProvide facts needed for the decisionWhich information actually supports the answer?
Tool resultsProvide fresh observations from systemsWhat did the external system actually return?
MemoryCarry useful information across steps or sessionsWhat earlier information is still relevant?
Answer spaceLeave room for the model's responseHow much output does this task need?

In code, a budget is a policy the application enforces, not a comment in a prompt. A minimal version gives every item a priority, protects the ones that must never be dropped, records everything it removes, and returns what it kept in its original order:

def assemble_context(items, window_tokens, reserve_output, count_tokens, log):
    limit = window_tokens - reserve_output          # answer space is reserved first
    chosen, used = set(), 0
    for item in sorted(items, key=lambda i: i.priority):   # priority 0 = protected
        cost = count_tokens(item.text)
        if used + cost <= limit:
            chosen.add(item.id)
            used += cost
        elif item.priority == 0:
            raise ContextBudgetError(f"protected item {item.id} does not fit")
        else:
            log.record("dropped", item.id, item.source_id, cost)   # never drop silently
    kept = [i for i in items if i.id in chosen]     # restore the intended order
    return kept, used

This is a design pattern, not a universal formula. A real system would try compression before dropping an item, and it would count tokens with the provider's own tokenizer (see section 4).

🛡️ Safety Check
Threat: critical instructions or authorization context get buried among large amounts of variable information. Control: keep application-controlled instructions structurally separate and protect them in the budget. Residual risk: separation alone does not guarantee correct model behavior, so authorization must be enforced outside the model.

🎯 Use this when your agent repeatedly carries large histories, retrieval results, memory or tool responses into its next step.

2. 📦 Making Room for Instructions, Evidence, and the Answer

🧒 Kid Analogy
If your teacher gives you a two-page assignment, the pages in front of you need to contain the instructions, the information needed to solve the problem and enough space to write your answer. You cannot use every page for background reading and then complain that there is no room left for the assignment.
✅ Green — Practical Example
Question: “Is this customer's replacement request eligible under the current policy?” The application needs the operating instructions, the customer's contract terms, the current replacement policy, the service event, a little recent conversation, perhaps a current product-status tool result, and enough capacity left for the answer. Now imagine adding twenty unrelated customer conversations, ten old policy versions, hundreds of lines of tool output and every earlier interaction. The context grew; the task did not get clearer.

An agent's context is shared by competing needs, and allocation should follow the task. The model should not spend its working capacity separating today's evidence from yesterday's noise because the application failed to curate.

Order matters too. Keep stable, protected material (operating instructions, fixed policy) at a consistent position at the start, put the task-specific evidence next, and place the question and any must-not-miss constraints near the end. Many providers also reuse work for repeated prefixes, so a stable beginning can lower cost and latency.

🛡️ Safety Check
Threat: critical information is crowded out by bulk material that happens to be easy to retrieve. Control: reserve deliberate space for required task information before adding optional material. Residual risk: a well-organized context still depends on the model reading it correctly, so critical decisions need independent checks.

🎯 Use this when the same context holds instructions, evidence, history and tool data that compete for attention.

3. 🔎 Why More Information Can Reduce Answer Quality

🧒 Kid Analogy
Imagine asking your teacher a simple question and handing over every notebook you have owned since primary school. The extra notebooks may contain useful information, but they can also make it harder to find the one page that answers today's question.
✅ Green — Practical Example
An agent must determine the current delivery date. The application retrieves the order, the current shipment status and a warehouse event. It also includes an old delivery estimate, a previous order and an unrelated support conversation. If the old estimate conflicts with the current shipment status, the extra material makes the task harder, not easier.

It is tempting to assume more context must produce a better answer. That assumption is unsafe. Additional information can add:

  1. Distraction: irrelevant material competes with relevant material.
  2. Duplication: the same fact appears repeatedly in different forms.
  3. Conflict: old and current information disagree.
  4. Ambiguity: sources describe similar concepts differently.
  5. Noise: large tool outputs contain details unrelated to the decision.
  6. Stale assumptions: earlier task state is carried forward after circumstances changed.

Research supports being careful here. The paper “Lost in the Middle” found that language models often use information less reliably depending on where it sits in a long input, with weaker use of material buried in the middle. Effects vary a lot between models and tasks, and newer models may behave differently, so treat this as a reason to measure your own task rather than a fixed rule. It also does not mean large windows are bad: they are very useful when the extra information is the evidence the task needs.

Context quality depends on both coverage (does it contain what the task needs?) and signal-to-noise ratio (how much else comes with it?).

Context strategyPossible result
Too littleThe model lacks facts required to answer.
Too much but well curatedUseful coverage, with higher cost and processing time.
Too much and poorly curatedIrrelevance, conflicts and duplication make the task harder.
Focused and sufficientThe evidence the task needs, without the baggage.
💡 Remember
“More context” and “more useful context” are different engineering goals.

🎯 Use this when your team keeps adding retrieved documents or conversation history whenever answer quality drops.

4. 🚦 What Happens When Context Exceeds the Model's Limit?

🧒 Kid Analogy
Your school bag has a maximum size. If you keep adding books after the bag is full, something has to happen: some books must come out, the load must be reorganized, or you need a bigger bag. The bag cannot hold unlimited material.

Definition: a context-window limit is the maximum amount of tokenized information a particular model and interface can process in a request, under that system's rules. Limits differ across models and APIs, and some systems split capacity between input, output and internal reasoning. Never assume the full published capacity is available for retrieved documents.

Count tokens the way your provider does. Different models tokenize text differently, so a character-count guess or a counter built for another model can be badly off. Use the provider's own token-counting facility, and recount after every compression step.

What happens at the limit depends on the application and API. Common behaviors:

  1. The request is rejected because it is too large.
  2. The application removes older or lower-priority context before sending.
  3. The application summarizes or compacts earlier state. Some platforms offer built-in compaction for this.
  4. The task is split across several model calls.
  5. The response is capped to preserve capacity.

Overflow handling should be designed deliberately. An agent should not rely on accidental truncation to decide which information disappears.

🛡️ Safety Check
Threat: important evidence is silently dropped when context becomes too large. Control: define explicit priority classes and deterministic context-management rules before the request reaches the model, and log every drop. Residual risk: the wrong item can still be prioritized, so critical decisions should expose the evidence and sources they used.

A production application should track at least:

  • Estimated input tokens, and reserved output capacity.
  • Size per context category, and number of retrieved items.
  • Compression operations performed.
  • Items removed or truncated.
  • Whether critical evidence was preserved.
💡 Engineering principle
Never make “whatever fits” your context-management policy.

🎯 Use this when long-running conversations or retrieval pipelines can grow beyond predictable size.

5. 🔄 How to Decide What to Keep, Remove, or Fetch Again

🧒 Kid Analogy
Before leaving for school, you check your bag. Today's homework stays. Yesterday's finished worksheet comes out. A textbook you need tomorrow may stay if there is room. A book too heavy to carry all day can stay at home until the teacher asks for it.
Flow chart asking whether an item is needed, still valid, exact or high impact, and whether it fits the budget, ending in remove, fetch again, fetch original, keep, compress or escalate

Original diagram: the questions run left to right; outcomes sit beneath them. Escalate is reachable from any step.

A useful policy has five outcomes:

DecisionMeaningExample
KeepRequired and still valid.Current contract terms needed for the decision.
RemoveIrrelevant, expired or redundant.An unrelated earlier conversation.
CompressStill useful, but can be smaller.A long conversation turned into checked task state.
Fetch again / fetch originalThe original is more reliable than a carried copy: it may have changed, or the decision is high impact.Re-reading the current policy before a high-impact decision.
EscalateThe context cannot safely establish the required fact.A required document is missing or contradictory.
✅ Green — Practical Example
A troubleshooting agent has a 30-message conversation. Messages 1 to 10 describe the original problem, 11 to 25 hold investigation details, and 26 to 30 contain the confirmed current state. Instead of carrying all 30 forever, the system keeps a compact checked state and fetches the original messages if exact wording or history becomes necessary.

Context is a dynamic workspace, not a permanent transcript.

🎯 Use this when the agent's context grows continuously during multi-step work.

6. 🗜️ Context Compression: Preserving Meaning in Less Space

🧒 Kid Analogy
Imagine ten pages of class notes about a science experiment. Instead of carrying all ten pages everywhere, you make one smaller page with the goal, the important measurements, the final result and anything you still need to remember. The small page helps only if it keeps the facts that matter.

Definition: context compression is representing useful context in a smaller form while preserving what the next task needs. It does not have to mean asking a model to “summarize this.” Forms include:

  • Summarization: a long narrative into a shorter one.
  • Extraction: keeping only the structured facts the task requires.
  • Deduplication: removing repeated information.
  • State conversion: turning a long conversation into explicit task state.
  • Filtering: removing material outside the current scope.
  • Reference preservation: keeping identifiers so the original can be fetched later.
Original contextCompressed representationMust preserve?
Long customer conversationIssue, affected product, actions tried, current statusYes, if needed to continue
Ten repeated tool responsesLatest verified state plus source identifiersUsually
Old irrelevant conversationNothingNo
Exact policy clauseShort paraphrase plus source and version referenceDepends on decision risk

The last row matters most: some information is safe to compress for navigation but unsafe to compress for a final decision. What should usually survive: critical facts, numbers and units, dates and deadlines, exceptions and conditions, decisions already made, unresolved questions, source identifiers, stated uncertainty, and scope.

A security point that is easy to miss. A summarizer reads whatever it is given, including untrusted text. If a retrieved web page or customer message contains an embedded instruction (prompt injection), compressing it into “verified task state” can strip away its origin and make it look like trusted guidance. Mark compressed content with the trust level of its sources, store facts rather than instructions, and never promote a summary into the instruction layer.

A compressed context that drops its source identifier is shorter but harder to verify. Design compression and provenance together.

🛡️ Safety Check
Threat: compression removes a condition or exception that changes a business rule, or launders untrusted text into trusted state. Control: preserve critical facts, source references and trust labels, and re-check the original before high-impact decisions. Residual risk: a summary can still contain subtle errors, so compression must not automatically become final authority.

🎯 Use this when the same task history is becoming too large to carry forward efficiently.

7. 🧪 How Summaries Lose Facts, and How to Check Them

🧒 Kid Analogy
You tell a friend about a long school trip: “We visited the museum and had lunch.” You forgot to say the museum was closed for two hours and lunch moved because of rain. Your summary is shorter, but the missing details change how someone understands the day.

A summary is a transformation, and transformation can lose, weaken or change information. Common losses: numbers (exact values vanish), dates (a deadline becomes “soon”), conditions (“only if” becomes “if”), exceptions, uncertainty (“possibly” becomes fact), attribution, scope (a one-region rule looks universal), version, and conflict between sources.

Original statementRisky summaryWhat was lost?
“Requests submitted within 30 days are eligible unless the product belongs to category X.”“Requests within 30 days are eligible.”The exception for category X.
“The reported value is approximately 18.4, based on an incomplete measurement.”“The value is 18.4.”Uncertainty and the measurement limitation.
“Policy version 7 applies from 1 October.”“The current policy says...”Version and effective date.

These summaries are shorter, but they are not equivalent to the originals. A summary can be fine for navigation and still be inadequate for authorization, compliance, financial calculation or any high-consequence decision.

How to check a summary. The most reliable checks are deterministic. Extract the critical items from the original (numbers, dates, version identifiers, condition words) and confirm each appears in the summary. A crude check like this catches dropped values and dropped conditions:

import re
NUMBERS_DATES = re.compile(
    r"\d{1,2} [A-Z][a-z]+ \d{4}"                       # dates first
    r"|\d+(?:[.,]\d+)*(?:\s?(?:%|days?|hours?))?"      # numbers, no trailing period
)
CONDITION_WORDS = ("unless", "except", "only if", "provided that")

def missing_critical_facts(original, summary):
    needed = {m.group(0).strip() for m in NUMBERS_DATES.finditer(original)}
    problems = [f"missing: {n}" for n in sorted(needed) if n not in summary]
    if any(w in original.lower() for w in CONDITION_WORDS) \
       and not any(w in summary.lower() for w in CONDITION_WORDS):
        problems.append("a condition or exception may have been dropped")
    return problems

This is a floor, not a guarantee: it cannot judge meaning. A second model reading the summary against the source is a useful extra signal, but it is also probabilistic, so for high-impact workflows show the evidence itself instead of trusting any summary check.

Store facts, not verdicts. A compression record that says “request appears eligible” invites later steps to reuse a conclusion as if it were a fact. Keep the facts, mark each with its source and who or what checked it, and label any conclusion as provisional:

{
  "facts": [
    {"claim": "Request submitted on day 24 of the 30-day window",
     "source_id": "case-4471", "checked_by": "tool:case_lookup"},
    {"claim": "Policy 2026-07 allows replacement within 30 days",
     "source_id": "policy-2026-07", "source_version": "7", "checked_by": "retrieval"}
  ],
  "open_questions": ["Does the exception for category X apply to this product?"],
  "provisional_conclusion": "Possibly eligible, pending the category check",
  "trust": "derived_from_mixed_sources",
  "original_available": true,
  "recheck_required": true
}

The structure is illustrative. What matters is that the record says what was summarized, where it came from, who checked it, and what still needs verification.

🛡️ Safety Check
Threat: a summary removes a limitation and the agent treats the remaining statement as universally true. Control: preserve uncertainty, scope, exceptions and sources, then validate critical claims against the original. Residual risk: the verification can itself fail, so high-impact workflows should expose evidence rather than rely on an opaque summary.

🎯 Use this when your agent compresses long conversations, documents, tool histories or memory before carrying them into later steps.

8. 🔍 When to Return to the Original Source

🧒 Kid Analogy
If your friend says the school notice lets you skip an exam, you should not make an important decision from your friend's summary. You would check the actual notice, especially if the decision matters.

Return to the original source when:

  1. The exact wording matters.
  2. A financial, legal, regulatory or operational decision depends on it.
  3. The summary contains uncertainty, or disagrees with another source.
  4. An important exception may apply.
  5. The source has changed since the summary was made, or its version or effective date matters.
  6. The agent needs evidence that can be cited or audited.
  7. The summary cannot establish the required fact.
  8. The proposed action is hard to reverse or has significant consequences.
✅ Green — Practical Example
A summary says, “The customer appears eligible for a replacement.” Before executing it, the agent retrieves the current policy and the customer contract. The originals confirm the exact conditions. The summary helped the agent navigate; the authoritative sources established the decision.

Use the summary for continuity. Use the original source for verification.

This is not absolute. Low-risk tasks may not need the source every time. Weigh the task's consequences, the reliability of the summary, the freshness of the source and the cost of retrieval.

SituationReasonable context strategy
Low-risk conversational continuityA checked summary may be sufficient.
Exact factual lookupRetrieve the relevant source when precision matters.
Conflicting informationReturn to authoritative sources and resolve the conflict.
High-impact actionVerify critical evidence against current authoritative sources before executing.
Missing sourceEnter an information-gap state rather than inventing certainty.

🎯 Use this when a compressed representation is being used to justify an action rather than merely continue a conversation.

9. 🏗️ Complete Worked Example: A Support Agent

Scenario: a customer asks, “Can I get a replacement for this product, and can you process it now?” The agent can reach the current service policy, the customer's contract, the product record, the earlier conversation, a shipment-status tool, an old policy document and a previous agent summary.

  1. Define the task. There are two jobs: decide eligibility, and possibly execute a replacement.
  2. Build the initial context. Current policy, applicable contract terms, relevant product facts, and only the conversation needed to understand the issue.
  3. Exclude the rest. The old policy must not compete with the current one; unrelated history takes no space.
  4. Use the previous summary carefully. It gives continuity, but eligibility facts must trace to original sources.
  5. Fetch fresh state. Call the shipment or product tool instead of trusting an old statement.
  6. Check the budget. Drop duplicated history and irrelevant tool output; compress long conversation into checked task state.
  7. Verify critical evidence. If eligibility depends on an exception clause, re-read the current policy section.
  8. Separate answer from action. “Appears eligible” does not grant permission to execute.
  9. Apply authorization. Execution needs its own permission or approval, enforced outside the model.

Here is what that does to the budget, using a hypothetical 32,000-token window (numbers are illustrative):

ComponentNaive assemblyCurated assembly
Operating instructions1,5001,500
Policies9,000 (current plus ten old versions)1,200 (current section only)
Conversation history12,000 (twenty conversations)1,500 (recent turns) plus 800 (checked task state)
Contract and product data3,500 (full records)900 (relevant excerpts)
Tool output6,000 (raw)600 (needed fields)
Memory2,500300
Input total34,500, over the limit6,800
Reserved for answernone left2,000

The curated version uses well under a third of the window and is easier to audit. The unused headroom is intentional. The original sources stay retrievable for verification.

💡 What changed?
The system did not make the model smarter by giving it more information. It improved the model's working environment by choosing information deliberately.
🛡️ Safety Check
Threat: an outdated summary or policy influences a consequential replacement action. Control: prioritize current authoritative sources, preserve provenance, verify critical conditions and enforce authorization independently. Residual risk: the model may still misread the evidence, so consequential actions need validation, auditability and appropriate human oversight.

🎯 Use this when an agent must maintain a long-running task while still deciding from current authoritative information.

10. 🏢 Enterprise Context-Budget Architecture and Rollout

In production, context budgeting is a pipeline, not one large prompt template.

LayerResponsibility
Task analysisDetermine what the current task actually requires.
RetrievalFind candidate evidence from approved sources.
SelectionFilter by relevance, authority, scope, freshness and access rules.
CompressionReduce size while preserving required meaning, provenance and trust labels.
BudgetingEnsure the assembled context leaves room for the response and other processing.
VerificationReturn to authoritative sources when precision or risk requires it.
Action controlKeep authorization and consequential action enforcement outside the model.

Useful observability metrics:

  • Context size before selection, after selection, and after compression.
  • Percentage of retrieved items discarded.
  • Number of source re-fetches.
  • Summary verification failures.
  • Context-overflow events.
  • Tool calls triggered after compressed context.
  • Cases where a human reviewer asked for the original source.
  • Cost and latency of context preparation.

These measurements turn context engineering from informal prompt-writing into something that can be tested and improved.

🚦 Rollout and Governance
  • Ownership: name an owner for the budget policy, for the compression rules and for the source-of-truth systems they depend on.
  • Versioning and change control: treat budgets, priority classes, compression logic and verification rules as versioned production artifacts. A change can alter what the agent sees without any code change elsewhere.
  • Release gate: before shipping a change, check what new information can now enter or be dropped, whether any protected item could be crowded out, what the summary checks cover, and whether tracing still explains every drop.
  • Cost governance: set token budgets and alerts per workflow. Re-fetching, long contexts and extra compression calls all cost money and latency.
  • Kill switch: keep a tested way to switch compression off and fall back to fetch-from-source or human escalation if summaries start causing wrong decisions.
💡 Engineering principle
If you cannot observe what entered the context, what was removed, what was compressed and what was fetched again, you cannot explain why an agent behaved differently after a change.

Enterprise note: the person who tunes the budget to save cost should not be the only person who can approve dropping evidence that decisions depend on.

11. ⚠️ Common Mistakes and Why They Fail

1. Treating the context window as a target to fill. A larger capacity is a limit, not a recommended payload size.

2. Keeping the entire conversation forever. Long conversations hold repeated questions and obsolete assumptions. Preserve task state deliberately instead of using the transcript as the only memory.

3. Compressing everything automatically. Some information holds exact conditions, numbers or exceptions that should be verified against the original.

4. Using summaries as permanent authority. A summary represents a source. It does not inherit the source's authority.

5. Losing provenance during compression. If the system keeps a conclusion but forgets where it came from, later verification is hard.

6. Compressing untrusted content into trusted state. A summary can strip the origin from injected text. Carry trust labels with compressed content.

7. Ignoring freshness. A compact summary can outlive the policy, customer state or record it describes.

8. Silently truncating context. When something must go, the application should know what disappeared and log it.

9. Miscounting tokens or forgetting the answer. Counting with the wrong tokenizer, or not reserving output space, causes overflows or clipped answers.

10. Fetching everything again. The opposite mistake: re-retrieving large sources when a checked compact state would do. Balance freshness, accuracy, cost and relevance.

11. Treating context management as a model-only problem. A model cannot enforce database permissions or approval rules because they were described in text. The surrounding application must.

🛡️ The Honest Limit
No current technique guarantees that compression preserves every fact or that a model uses every item it is given. The goal is risk and blast-radius reduction: keep evidence traceable, verify before consequential actions, enforce authorization outside the model, and make context decisions visible enough to investigate.

12. ❓ FAQ

Is a context budget the same as a context window?
No. The context window is the model or API's technical capacity for a request. A context budget is the application's deliberate allocation of that capacity among instructions, evidence, history, memory, tool results and output requirements.
Is more context always better?
No. More information can improve an answer when it supplies missing evidence, but irrelevant, duplicated, stale or conflicting information can make context harder to use effectively.
What should happen when context becomes too large?
The application should apply an explicit strategy such as removing irrelevant information, compressing useful state, splitting the task or retrieving information only when needed, and it should log what was dropped. It should not depend on accidental truncation.
Is summarization the same as context compression?
Not exactly. Summarization is one compression technique. Compression can also involve extraction, deduplication, structured task state, filtering and preserving references to original sources.
When should an agent retrieve the original source again?
When exact wording, current status, exceptions, source version, unresolved uncertainty or high-impact decision evidence matters. A summary is useful for continuity, but the original may be required for verification.
Can a summary become stale?
Yes. If the original source changes after the summary is created, the summary may no longer represent the current state, which is why freshness and revalidation matter.
Can summarizing untrusted content make it safe?
No. A summary of untrusted text is still derived from untrusted text, and it can hide where the text came from. Keep trust labels on compressed content and do not treat it as instructions.
Should every compressed context retain the original source?
For important information, yes: keep a source identifier or retrieval path so the system can verify the compressed form instead of treating it as an untraceable fact.

13. 🔗 References & Further Reading

Originality & attribution: Product and standards names belong to their owners. Platform features change quickly, so check current documentation before relying on a specific capability.

14. 📝 Summary

  • Context budget: the deliberate allocation of working space among instructions, evidence, task state, memory, tool results and output.
  • More is not automatically better: extra information can add distraction, duplication, stale facts and conflicts.
  • Limits matter: count tokens with the provider's tokenizer and reserve answer space; never assume the whole window is for input.
  • Keep, remove, compress, fetch or escalate: manage context actively by relevance, freshness, authority and scope.
  • Compression is transformation: a shorter form helps only if the facts the next decision needs survive, with sources and trust labels attached.
  • Check summaries: numbers, dates, conditions, exceptions, uncertainty, scope and version are what get lost; check them deterministically where you can.
  • Verify high-impact decisions: use the summary for continuity and the original source for verification.
  • Govern it: own, version, observe and be able to switch off your budget and compression logic.
🧠 The mental model to remember

Think of context as a working desk, not a warehouse. Put today's instructions on the desk. Bring the evidence today's task needs. Remove what no longer matters. Compress what must stay but not in full. And when an important fact becomes uncertain, go back to the original source.

A powerful agent is not created by giving the model more information. It is created by engineering a context that gives the model the right information, at the right time, in the right form, with enough room to reason and respond. That is the difference between having a large context and managing context intelligently.

Comments