How AI Selects Context: Relevance, Freshness, Conflicting Sources, and Evidence
Context engineering for an AI agent is the disciplined process of deciding what information the agent should receive for a particular task, where that information came from, how much authority it has, and what the agent is allowed to do with it. The difficult part is not finding information. It is selecting the right evidence without accidentally turning untrusted data into instructions. 🧭
In a production agent, a poor context decision can do more than produce an unhelpful answer. It can expose sensitive information, trigger an inappropriate tool, carry contaminated memory into a later task, or lead an operator to believe an action rested on authoritative evidence when it did not. Current agent-security guidance therefore treats context providers, tool outputs, permissions and trust boundaries as part of the security architecture. 🛡️
Original diagram: selection sits between retrieval and the agent, and the action gate sits outside the model.
- Quick Comparison: Where Context Comes From
- The Context Selection Problem
- Start With the Information the Task Actually Needs
- Retrieval Finds Candidates; Selection Chooses Evidence
- Source Provenance: Keep Facts Connected to Their Origins
- When Available Information Cannot Answer the Question
- The Working Context Is a Curated Workspace
- Worked Example: Selecting Context From Several Sources
- Enterprise Rollout: Governing the Context Pipeline
- Common Mistakes and Why They Fail
- Honest Limits
- FAQ
- References & Further Reading
- Summary
1. 🔀 Quick Comparison: Where Context Comes From
| Approach | What enters the agent workspace | Main engineering question | Security concern |
|---|---|---|---|
| Pre-loaded context | Known instructions, policies and stable references | What is genuinely required for this agent role? | Stale or over-broad standing information |
| Just-in-time retrieval | Candidate documents or records selected during a task | Which evidence is relevant now? | Untrusted or incorrectly scoped retrieved material |
| Short-term memory | Current task state and recent observations | What must survive to complete this task? | Old assumptions becoming invisible dependencies |
| Long-term memory | Persistent preferences, facts or prior observations | Should this fact still be trusted and retained? | Memory poisoning, stale information, missing provenance |
| Tool output | Results returned by databases, APIs or services | What did the tool actually establish? | Treating returned data as authority or as an instruction |
2. 🧩 The Context Selection Problem
Definition: context selection is the process of deciding which available information goes into the agent's working context for a particular task, while preserving the source, scope, trust status and handling rules attached to that information.
Available information is not the same thing as required information. An enterprise agent may have access to thousands of records, several knowledge repositories, tool results, earlier task state and persistent memory. The pipeline should not treat all of them as equally relevant, current, authoritative or safe. Mechanically:
- Identify the task and the requested outcome.
- Identify the facts needed to complete it.
- Identify which systems can legitimately provide those facts.
- Retrieve candidate evidence.
- Filter candidates by authorization, scope, freshness and provenance.
- Assemble a compact working set.
- Keep policy enforcement and authorization outside the agent's discretion.
- Record which sources and tools influenced the action.
Microsoft's agent safety guidance in Microsoft Agent Framework names context services (memories, user profiles, retrieval results) and tool-accessed services as trust boundaries. It treats tool results as untrusted, advises allow-listing and validating the arguments the model supplies for tools, and recommends gating high-risk tools behind human approval. It also notes that the system role has the highest trust and must never contain untrusted input, and that you should attach only context providers you trust.
🎯 Use this when your agent has more potential information than any single task requires.
3. 🎯 Start With the Information the Task Actually Needs
Definition: task-oriented context design starts by describing the minimum information the requested operation needs to be performed correctly and safely. Four fields are enough to begin with: the objective, the required evidence, the allowed actions, and the decision boundary (when the agent must stop and ask). Writing them as data, not prose, lets code check them:
{
"objective": "Decide replacement eligibility",
"required_evidence": ["current_policy", "customer_contract", "service_event"],
"allowed_actions": ["read", "propose"],
"decision_boundary": "Escalate if required evidence is missing or two authoritative sources conflict"
}
Google Cloud's published security-operations architecture is a useful real-world pattern. It describes an agentic workflow that looks up alerts, consults incident-response plans and runbooks held in a RAG knowledge database, keeps persistent state in a separate memory system, and uses human-in-the-loop approval. The lesson is architectural, not vendor-specific: the context for “investigate this endpoint” should not automatically become the context for “change this endpoint,” because investigation and action need different information and different authorization.
Best-practice design:
- Define information categories, not just a list of repositories.
- Map each category to an approved source.
- Assign sensitivity and access rules before retrieval happens.
- Give changeable information a freshness rule.
- Define what evidence is too weak to proceed.
🎯 Use this when the same agent performs several tasks with different information and authorization needs.
4. 🔎 Retrieval Finds Candidates; Selection Chooses Evidence
Definition: retrieval is a search that identifies candidates. Context selection is the later decision about which candidates enter the working set. They are often drawn as one box, but they solve different problems: retrieval can return something on-topic yet obsolete, out of scope, contradictory, or outside the user's authorization. A selection policy can examine:
- Relevance: does it directly support this task?
- Authority: is this source allowed to establish the fact?
- Freshness: is it still valid for the decision?
- Scope: does it apply to this customer, region, product or transaction?
- Provenance: can its origin and transformation history be identified?
- Sensitivity: may the agent receive it?
- Conflict: does another authoritative source disagree?
Selection is code, not a hope. A minimal version checks authorization first, then source type, scope and freshness, and records why each rejected item was rejected:
def select_context(task, candidates, user, policy):
selected, rejected = [], []
for c in candidates:
# In production, enforce this at the retrieval layer so unauthorized
# content never leaves the store, and use the END USER's permissions.
if not policy.can_read(user, c):
rejected.append((c.source_id, "not_authorized")); continue
if c.source_type not in task.allowed_source_types:
rejected.append((c.source_id, "wrong_source_type")); continue
if not c.applies_to(task.scope):
rejected.append((c.source_id, "out_of_scope")); continue
if c.is_expired(task.as_of):
rejected.append((c.source_id, "stale")); continue
selected.append(attach_labels(c)) # trust label + provenance
gaps = [r for r in task.required_evidence
if not any(s.satisfies(r) for s in selected)]
conflicts = find_conflicts(selected) # e.g. policy vs contract
return selected, rejected, gaps, conflicts
The shape matters more than the details: authorization before ranking, reasons recorded for every rejection, and gaps and conflicts returned as first-class results instead of being quietly resolved by the model. Treat the rejection log itself as access-controlled, because the existence of an item the user cannot read can be sensitive too.
A selected item should carry metadata that says how it may be used:
{
"source_id": "policy-2026-041",
"source_type": "approved_policy",
"source_version": "7",
"retrieved_at": "2026-10-03T09:12:00Z",
"scope": "service_entitlement",
"trust_label": "reference_data",
"allowed_use": ["answer", "cite"],
"authority": "policy_owner"
}
🎯 Use this when your retrieval layer can return more candidates than the agent should see.
5. 🧾 Source Provenance: Keep Facts Connected to Their Origins
Definition: provenance is the information that lets a piece of context be traced to its origin, scope, version and transformation history. It should travel with the information instead of vanishing during retrieval, summarization or memory creation.
| Provenance field | Why it matters |
|---|---|
| Source identity | Shows where the information originated. |
| Owner | Identifies who controls the source. |
| Version / date | Separates current information from historical information. |
| Scope | Shows which customer, product, geography or process it applies to. |
| Transformation | Records whether it was extracted, summarized, merged or otherwise changed. |
| Trust classification | Separates evidence from authority and from untrusted content. |
Original diagram: information changes form as it moves, but its origin and trust level must not be lost.
Provenance matters most when information crosses boundaries. A retrieved document becomes a memory item; the memory influences a later tool call; the tool call changes a business system. Without lineage, an operator sees the final action and cannot reconstruct why the agent believed the underlying fact. Two rules keep lineage honest: every derived item keeps pointers to its sources, and a derived item inherits the lowest trust of anything it was built from. A summary of untrusted text must not become “reference data” just because a summarizer touched it.
For governance context, the NIST AI Risk Management Framework describes continuous, lifecycle risk management. That supports the idea that information and decisions in an AI system need accountable ownership and not just an application log.
🎯 Use this when information can be reused across sessions, agents, workflows or business decisions.
6. 🚦 When Available Information Cannot Answer the Question
Definition: an information gap exists when the available, authorized and sufficiently relevant context does not establish the fact the task requires. A mature pipeline supports a controlled stop condition:
- Identify the missing fact.
- Check whether an approved source can provide it, and retrieve it if so.
- If not, decide whether the task can safely continue without it.
- If it cannot, explain the gap.
- Escalate or request human input when the workflow needs a decision.
Do not leave “is the evidence sufficient?” entirely to the model's judgment. Compare the evidence against the task's required_evidence list in code, and have the pipeline return a structured result the rest of the system can act on:
{
"status": "insufficient_evidence",
"missing": ["contract_schedule_b"],
"attempted_sources": ["contract-repository"],
"next_step": "request_from_authorized_reviewer",
"actions_blocked": ["approve_exception"]
}
This matters most for tool-using agents, because uncertainty in context should not turn into autonomy in action. OWASP's guidance on excessive agency (LLM06 in the 2025 edition of its LLM Top 10, LLM03 in the 2026 edition) names excessive functionality, permissions and autonomy as the root contributors to damaging agent actions.
🎯 Use this when an agent works on business records where “unknown” is a legitimate and important state.
7. 🧰 The Working Context Is a Curated Workspace
Definition: the working context is the controlled set of task information available to the agent for an execution step. Each item should have an explicit role (instructions, evidence, state, tool result, memory) and know what it is, where it came from, what scope it has, how long it stays valid and how it may be used:
context_item = {
"content": "...",
"role": "evidence",
"source_id": "...",
"trust": "untrusted_data",
"scope": "...",
"expires_at": "...",
"allowed_actions": ["read", "cite"]
}
The field names are illustrative. The principle is that the system, not the model, holds this metadata, so it can still be enforced when the model misreads a document.
Safe context assembly loop:
- Load application-controlled instructions.
- Load only the task state the current step needs.
- Retrieve authorized candidate evidence.
- Attach provenance and trust labels.
- Remove expired, unrelated or unauthorized items.
- Place evidence in a clearly separated data area.
- Expose only the tools the current operation needs.
- Apply external authorization and approval before consequential actions.
🎯 Use this when your agent combines instructions, memory, retrieved documents and tool responses in one execution.
8. 🏗️ Worked Example: Selecting Context From Several Sources
Scenario: a service agent receives “Can this customer's replacement request be approved under the current service policy?” Retrieval returns six candidates:
| Source | Status | Selection reasoning |
|---|---|---|
| Current policy | Authoritative | Primary evidence for eligibility. |
| Customer contract | Authoritative | Establishes customer-specific terms. |
| Old policy PDF | Historical | Traceable, but not current approval authority. |
| Support ticket | Operational evidence | May establish what happened, not the governing policy. |
| Technician note | Unstructured | Potentially useful evidence; not automatically authoritative. |
| Inventory record | Operational data | Relevant only if availability is part of the decision. |
- Define the decision. Determine eligibility under the current policy, not merely summarize the customer's history.
- Select governing evidence. The current policy and the customer contract become the central context.
- Add supporting evidence. The ticket joins only if it establishes facts such as the failure date.
- Preserve provenance. Each selected item keeps its source, owner, version and scope.
- Check for conflict. If the policy and contract disagree, the agent does not silently choose. The organization's precedence rule decides, or the case escalates.
- Check sufficiency. If the policy requires a document that is missing, enter an information-gap state.
- Separate recommendation from action. A proposed approval is not the approval transaction.
- Gate the tool. If policy requires human authorization, the action waits for it.
What the selection log would show for this task:
selected: current-policy (v7), customer-contract-8831, ticket-4471 (failure date only)
evidence only (not authority): technician-note-221
rejected: old-policy-2024 -> stale
inventory-record -> not needed for this decision
gaps: none
conflicts: none
next: propose approval; wait for human authorization
🎯 Use this when an agent must combine several repositories before making a business decision or proposing an action.
9. 🏢 Enterprise Rollout: Governing the Context Pipeline
Treat enterprise context engineering as an operational control plane, not a collection of prompts and retrieval queries.
| Area | What good looks like |
|---|---|
| Ownership | Named owners for instructions, retrieval sources, tool definitions, memory policy, data classification and action authorization. |
| Versioning and change control | Instructions, tool schemas, retrieval policies, source mappings and memory rules are versioned. A context change can alter behavior with no application-code change. |
| Readiness review | Identify what changed, affected tools, sources and users, access and classification impact, failure and escalation paths, logging coverage, approval rules; get owner sign-off; deploy with rollback. |
| Data classification | Every source has a classification and access rule. “The agent can technically read it” is never the policy that decides what enters context. |
| Secrets and credentials | Secrets stay out of ordinary context. Tools obtain credentials through controlled identity and secret management. |
| Memory retention and deletion | Persistent memory has an owner, provenance, purpose, retention and deletion rules, so it does not become an undocumented business dependency. |
| Cost governance | Monitor repeated retrieval, unnecessary tool calls, oversized histories and long multi-step workflows against business value. |
| Observability | Trace which sources were selected and rejected, which tools ran, which authorization checks occurred, which approvals were requested and what was executed. Redact sensitive content in traces. |
| Alerting | Unusual tool usage, access outside expected scope, repeated approval attempts, abnormal outbound activity, unexpected context-configuration changes, tools used outside their purpose. |
| Incident response | A tested kill switch: revoke credentials, disable tools, quarantine a context source, stop workflows, and preserve traces. |
🎯 Use this when an agent is moving from an experiment into a business process with real customer, financial, operational or regulatory consequences.
10. ⚠️ Common Mistakes and Why They Fail
1. Treating retrieved or tool content as trusted instructions. Retrieved material can contain useful facts without having authority to control the agent. Microsoft's guidance classifies tool results and external context as untrusted input.
2. Selecting by similarity score alone. The closest match is not necessarily the authoritative one. Check authority, version and scope, not just relevance.
3. Authorizing only after retrieval, or with the agent's own access. If unauthorized content reaches the context and is hidden afterward, it can still influence the answer. Authorize at retrieval, using the end user's permissions, and keep selection-time checks as a second layer.
4. Assuming newest means most authoritative. A recent note can be less authoritative than an older approved policy. Use explicit source precedence and flag conflicts instead of picking silently.
5. Granting broad credentials “for convenience”. Broad access turns a context mistake into a larger incident. OWASP names excessive functionality, permissions and autonomy as root causes of excessive agency.
6. Relying on standing instructions alone as a security boundary. Instructions are one layer. Authorization and enforcement belong in the application and tool layers.
7. Stuffing the workspace instead of curating it. More information creates more selection work. Assemble context around the task, not around the maximum data available.
8. Unbounded memory with no provenance or expiry. A remembered statement can outlive the conditions that made it true. OWASP's 2026 article on memory and context poisoning explains why persistent state is itself an attack surface.
9. Letting derived items inherit more trust than their inputs. A summary or memory built from untrusted text is still untrusted. Propagate the lowest trust level and keep source pointers.
10. Shipping context changes with no review gate or tracing. Changing retrieval rules or tool descriptions changes behavior without code changes. Version, trace and keep rollback.
11. Approval fatigue. If people approve every harmless action they start approving mechanically. Concentrate review where reversibility, sensitivity or impact makes human judgment meaningful.
12. No kill switch. If you cannot quickly disable an agent's credentials or tools, the incident-response plan is incomplete.
11. 🛡️ Honest Limits
No current context-engineering technique fully eliminates prompt injection or every other way an agent can be manipulated. The practical objective is to reduce the probability of unsafe behavior and, more importantly, contain the blast radius when a component misbehaves.
Design for imperfect behavior. The surrounding system decides which data can be accessed, which tools can be called, which arguments are acceptable, which actions need approval, what gets logged and how the agent can be stopped. That posture matches current platform and security guidance on trust boundaries, validation, least privilege, approval and lifecycle governance.
12. ❓ FAQ
13. 🔗 References & Further Reading
- Microsoft Learn — Agent Framework: Safety and trust boundaries
- Google Cloud Architecture Center — Agentic AI for security operations workflows
- OWASP Top 10 for LLM Applications 2025 — LLM06 Excessive Agency (renumbered LLM03 in the 2026 edition)
- OWASP — Top 10 for LLM Applications (current edition)
- OWASP GenAI Security Project — Memory Is a Feature. It Is Also an Attack Surface
- NIST — AI Risk Management Framework
- NIST — AI RMF: Generative AI Profile (NIST AI 600-1)
Originality & attribution: Product and standards names belong to their owners. Platform guidance changes quickly, so check current documentation before relying on a specific capability.
14. 📝 Summary
- Selection: decide what information an agent receives for a specific task and how it may be used. Available is not the same as required.
- Start from the task: write the objective, required evidence, allowed actions and decision boundary as data.
- Retrieval vs selection: retrieval finds candidates; selection checks authorization, authority, scope, freshness and conflicts, and logs why items were rejected.
- Provenance: keep origin, version, scope and trust with every item, and let derived items inherit the lowest trust of their inputs.
- Information gaps: check evidence against the required list in code, and stop or escalate rather than assume.
- Working context: label every item and keep controlled policy separate from variable or untrusted data.
- Enterprise rollout: own, version, classify, observe and be able to stop the context pipeline.
- Honest limit: no technique eliminates every attack, so aim for defense in depth and a contained blast radius.
An agent should not receive information merely because the system can retrieve it. It should receive information because the application has deliberately decided it is relevant, authorized, appropriately scoped and useful for the task, with enough provenance preserved to understand where it came from. That is the difference between having context and engineering context.
Comments
Post a Comment