Skip to main content

How AI Selects Context: Relevance, Freshness, Conflicting Sources, and Evidence

Calculating read time…

Context engineering for an AI agent is the disciplined process of deciding what information the agent should receive for a particular task, where that information came from, how much authority it has, and what the agent is allowed to do with it. The difficult part is not finding information. It is selecting the right evidence without accidentally turning untrusted data into instructions. 🧭

In a production agent, a poor context decision can do more than produce an unhelpful answer. It can expose sensitive information, trigger an inappropriate tool, carry contaminated memory into a later task, or lead an operator to believe an action rested on authoritative evidence when it did not. Current agent-security guidance therefore treats context providers, tool outputs, permissions and trust boundaries as part of the security architecture. 🛡️

Pipeline from task spec through authorized retrieval, selection, labeled working context and agent to an action gate outside the model, with rejected items and information gaps logged

Original diagram: selection sits between retrieval and the agent, and the action gate sits outside the model.

1. 🔀 Quick Comparison: Where Context Comes From

ApproachWhat enters the agent workspaceMain engineering questionSecurity concern
Pre-loaded contextKnown instructions, policies and stable referencesWhat is genuinely required for this agent role?Stale or over-broad standing information
Just-in-time retrievalCandidate documents or records selected during a taskWhich evidence is relevant now?Untrusted or incorrectly scoped retrieved material
Short-term memoryCurrent task state and recent observationsWhat must survive to complete this task?Old assumptions becoming invisible dependencies
Long-term memoryPersistent preferences, facts or prior observationsShould this fact still be trusted and retained?Memory poisoning, stale information, missing provenance
Tool outputResults returned by databases, APIs or servicesWhat did the tool actually establish?Treating returned data as authority or as an instruction

2. 🧩 The Context Selection Problem

🧒 Kid Analogy
Imagine a teacher asks, “Why was Alex absent yesterday?” You have a school timetable, the attendance register, a lunch menu, five old homework books, and a note from another student. Having more papers does not automatically help. You first need the papers that can actually answer the question.
✅ Practical worked example
A service agent is asked, “What is the approved replacement procedure for this failed component?” It can reach a maintenance manual, an old incident ticket, a current service bulletin, an inventory record and a technician's free-text note. The context pipeline should treat the current procedure as the principal evidence, use inventory only if availability is part of the question, and treat the free-text note as evidence rather than authority.

Definition: context selection is the process of deciding which available information goes into the agent's working context for a particular task, while preserving the source, scope, trust status and handling rules attached to that information.

Available information is not the same thing as required information. An enterprise agent may have access to thousands of records, several knowledge repositories, tool results, earlier task state and persistent memory. The pipeline should not treat all of them as equally relevant, current, authoritative or safe. Mechanically:

  1. Identify the task and the requested outcome.
  2. Identify the facts needed to complete it.
  3. Identify which systems can legitimately provide those facts.
  4. Retrieve candidate evidence.
  5. Filter candidates by authorization, scope, freshness and provenance.
  6. Assemble a compact working set.
  7. Keep policy enforcement and authorization outside the agent's discretion.
  8. Record which sources and tools influenced the action.

Microsoft's agent safety guidance in Microsoft Agent Framework names context services (memories, user profiles, retrieval results) and tool-accessed services as trust boundaries. It treats tool results as untrusted, advises allow-listing and validating the arguments the model supplies for tools, and recommends gating high-risk tools behind human approval. It also notes that the system role has the highest trust and must never contain untrusted input, and that you should attach only context providers you trust.

🛡️ Safety Check
Threat: a retrieved document contains text that tries to redirect the agent (indirect prompt injection). Control: label retrieved material as data, keep application-controlled instructions separate, and restrict tool permissions independently. Residual risk: an agent can still misread untrusted information, so permissions, approval gates, validation and monitoring must add containment.

🎯 Use this when your agent has more potential information than any single task requires.

3. 🎯 Start With the Information the Task Actually Needs

🧒 Kid Analogy
If someone asks, “What time does the bus leave?” you need the bus schedule, not the entire history of the transportation company. The question decides the size of the backpack.
✅ Practical worked example
For an expense-policy question, the agent might need the current policy, the employee's region and the expense category. It does not automatically need the employee's full HR record, unrelated expense history or access to payment tools.

Definition: task-oriented context design starts by describing the minimum information the requested operation needs to be performed correctly and safely. Four fields are enough to begin with: the objective, the required evidence, the allowed actions, and the decision boundary (when the agent must stop and ask). Writing them as data, not prose, lets code check them:

{
  "objective": "Decide replacement eligibility",
  "required_evidence": ["current_policy", "customer_contract", "service_event"],
  "allowed_actions": ["read", "propose"],
  "decision_boundary": "Escalate if required evidence is missing or two authoritative sources conflict"
}
💡 Key warning
“Give the agent everything because it may need it later” is not a context strategy. It increases the number of facts the agent must tell apart, exposes more sensitive records to the task, and makes provenance harder to reason about.

Google Cloud's published security-operations architecture is a useful real-world pattern. It describes an agentic workflow that looks up alerts, consults incident-response plans and runbooks held in a RAG knowledge database, keeps persistent state in a separate memory system, and uses human-in-the-loop approval. The lesson is architectural, not vendor-specific: the context for “investigate this endpoint” should not automatically become the context for “change this endpoint,” because investigation and action need different information and different authorization.

Best-practice design:

  1. Define information categories, not just a list of repositories.
  2. Map each category to an approved source.
  3. Assign sensitivity and access rules before retrieval happens.
  4. Give changeable information a freshness rule.
  5. Define what evidence is too weak to proceed.
🛡️ Safety Check
Threat: an agent receives unrelated sensitive information simply because its service account can access it. Control: scope data access to the task, and authorize against the end user's permissions before context is assembled. Residual risk: correctly authorized information can still be misread, so tool permissions and downstream controls remain necessary.

🎯 Use this when the same agent performs several tasks with different information and authorization needs.

4. 🔎 Retrieval Finds Candidates; Selection Chooses Evidence

🧒 Kid Analogy
Going to the library and finding ten books about dinosaurs is retrieval. Putting three books on your desk because they answer today's homework question is context selection.
✅ Practical worked example
An agent must decide whether a customer's service entitlement is active. Retrieval finds five candidates: the current contract record, an expired contract, a support ticket, a product manual and a technician note. Selection treats the current contract as the primary entitlement evidence, uses the ticket only if it supplies relevant status, and rejects the manual as irrelevant to entitlement.

Definition: retrieval is a search that identifies candidates. Context selection is the later decision about which candidates enter the working set. They are often drawn as one box, but they solve different problems: retrieval can return something on-topic yet obsolete, out of scope, contradictory, or outside the user's authorization. A selection policy can examine:

  1. Relevance: does it directly support this task?
  2. Authority: is this source allowed to establish the fact?
  3. Freshness: is it still valid for the decision?
  4. Scope: does it apply to this customer, region, product or transaction?
  5. Provenance: can its origin and transformation history be identified?
  6. Sensitivity: may the agent receive it?
  7. Conflict: does another authoritative source disagree?
💡 Important distinction
Relevance is not authority. A highly relevant technician note may explain exactly what happened and still be unsuitable as the source of an official policy requirement.

Selection is code, not a hope. A minimal version checks authorization first, then source type, scope and freshness, and records why each rejected item was rejected:

def select_context(task, candidates, user, policy):
    selected, rejected = [], []
    for c in candidates:
        # In production, enforce this at the retrieval layer so unauthorized
        # content never leaves the store, and use the END USER's permissions.
        if not policy.can_read(user, c):
            rejected.append((c.source_id, "not_authorized")); continue
        if c.source_type not in task.allowed_source_types:
            rejected.append((c.source_id, "wrong_source_type")); continue
        if not c.applies_to(task.scope):
            rejected.append((c.source_id, "out_of_scope")); continue
        if c.is_expired(task.as_of):
            rejected.append((c.source_id, "stale")); continue
        selected.append(attach_labels(c))          # trust label + provenance
    gaps = [r for r in task.required_evidence
            if not any(s.satisfies(r) for s in selected)]
    conflicts = find_conflicts(selected)           # e.g. policy vs contract
    return selected, rejected, gaps, conflicts

The shape matters more than the details: authorization before ranking, reasons recorded for every rejection, and gaps and conflicts returned as first-class results instead of being quietly resolved by the model. Treat the rejection log itself as access-controlled, because the existence of an item the user cannot read can be sensitive too.

A selected item should carry metadata that says how it may be used:

{
  "source_id": "policy-2026-041",
  "source_type": "approved_policy",
  "source_version": "7",
  "retrieved_at": "2026-10-03T09:12:00Z",
  "scope": "service_entitlement",
  "trust_label": "reference_data",
  "allowed_use": ["answer", "cite"],
  "authority": "policy_owner"
}
🛡️ Safety Check
Threat: retrieved content contains instruction-like text and is mistaken for an operating command. Control: treat retrieved content as data, preserve trust labels and keep application-controlled instructions separate. Residual risk: selection can still be wrong, so sensitive actions need independent authorization and, where appropriate, human approval.

🎯 Use this when your retrieval layer can return more candidates than the agent should see.

5. 🧾 Source Provenance: Keep Facts Connected to Their Origins

🧒 Kid Analogy
Imagine your teacher gives you three facts, but you forget which book each came from. Later one fact is challenged, and you cannot check it because the labels are gone.
✅ Practical worked example
An agent receives a summarized maintenance record. Instead of storing only “component replaced successfully,” the memory entry keeps a reference to the original maintenance record, its date, the equipment identifier, the source owner, and whether the statement was directly recorded or derived from several records.

Definition: provenance is the information that lets a piece of context be traced to its origin, scope, version and transformation history. It should travel with the information instead of vanishing during retrieval, summarization or memory creation.

Provenance fieldWhy it matters
Source identityShows where the information originated.
OwnerIdentifies who controls the source.
Version / dateSeparates current information from historical information.
ScopeShows which customer, product, geography or process it applies to.
TransformationRecords whether it was extracted, summarized, merged or otherwise changed.
Trust classificationSeparates evidence from authority and from untrusted content.
Lineage from source record to retrieved extract, summary, memory item and tool call with dashed pointer back to the source and a rule that derived items inherit the lowest trust

Original diagram: information changes form as it moves, but its origin and trust level must not be lost.

Provenance matters most when information crosses boundaries. A retrieved document becomes a memory item; the memory influences a later tool call; the tool call changes a business system. Without lineage, an operator sees the final action and cannot reconstruct why the agent believed the underlying fact. Two rules keep lineage honest: every derived item keeps pointers to its sources, and a derived item inherits the lowest trust of anything it was built from. A summary of untrusted text must not become “reference data” just because a summarizer touched it.

For governance context, the NIST AI Risk Management Framework describes continuous, lifecycle risk management. That supports the idea that information and decisions in an AI system need accountable ownership and not just an application log.

🛡️ Safety Check
Threat: an untrusted observation is saved as a permanent memory and later looks identical to an approved business fact. Control: store provenance, scope, creation time, trust label and expiry with every memory. Residual risk: provenance does not make bad data true; it makes the data's origin and handling visible to downstream controls.

🎯 Use this when information can be reused across sessions, agents, workflows or business decisions.

6. 🚦 When Available Information Cannot Answer the Question

🧒 Kid Analogy
Your teacher asks, “Who took the missing book?” You have the seating chart and yesterday's attendance sheet. Neither proves who took it. The right answer is not to guess; it is to say what you know and what evidence is missing.
✅ Practical worked example
An agent is asked whether a contract permits an exception. It finds the current contract, but the exception clause points to a separate Schedule B that is unavailable. The safe outcome: “The contract references Schedule B, which I cannot access. I cannot establish from the available records whether the exception is permitted.” The agent then requests the schedule or routes the case to an authorized reviewer.

Definition: an information gap exists when the available, authorized and sufficiently relevant context does not establish the fact the task requires. A mature pipeline supports a controlled stop condition:

  1. Identify the missing fact.
  2. Check whether an approved source can provide it, and retrieve it if so.
  3. If not, decide whether the task can safely continue without it.
  4. If it cannot, explain the gap.
  5. Escalate or request human input when the workflow needs a decision.

Do not leave “is the evidence sufficient?” entirely to the model's judgment. Compare the evidence against the task's required_evidence list in code, and have the pipeline return a structured result the rest of the system can act on:

{
  "status": "insufficient_evidence",
  "missing": ["contract_schedule_b"],
  "attempted_sources": ["contract-repository"],
  "next_step": "request_from_authorized_reviewer",
  "actions_blocked": ["approve_exception"]
}
💡 Key warning
“The system returned something” is not the same as “the system returned enough evidence.” A query can succeed technically and still fail to establish the business fact the task needs.

This matters most for tool-using agents, because uncertainty in context should not turn into autonomy in action. OWASP's guidance on excessive agency (LLM06 in the 2025 edition of its LLM Top 10, LLM03 in the 2026 edition) names excessive functionality, permissions and autonomy as the root contributors to damaging agent actions.

🛡️ Safety Check
Threat: missing evidence is silently replaced by an assumption and the agent proceeds to an irreversible tool action. Control: define explicit information-gap states and require approval or escalation when mandatory evidence is absent. Residual risk: humans can still approve bad actions, so the approval screen should show the evidence and provenance behind the proposal.

🎯 Use this when an agent works on business records where “unknown” is a legitimate and important state.

7. 🧰 The Working Context Is a Curated Workspace

🧒 Kid Analogy
Think about your school backpack. Your teacher's instructions, today's notebook and your pencil belong in the main compartment. A random note from a stranger does not become an official school instruction just because it is physically inside the bag.
✅ Practical worked example
A customer-support agent receives a tool result containing a customer address. The address is evidence returned by the CRM. It is not an instruction to update the address, send an email or call another service. Those actions need separate authorization and tool policies.

Definition: the working context is the controlled set of task information available to the agent for an execution step. Each item should have an explicit role (instructions, evidence, state, tool result, memory) and know what it is, where it came from, what scope it has, how long it stays valid and how it may be used:

context_item = {
    "content": "...",
    "role": "evidence",
    "source_id": "...",
    "trust": "untrusted_data",
    "scope": "...",
    "expires_at": "...",
    "allowed_actions": ["read", "cite"]
}

The field names are illustrative. The principle is that the system, not the model, holds this metadata, so it can still be enforced when the model misreads a document.

Safe context assembly loop:

  1. Load application-controlled instructions.
  2. Load only the task state the current step needs.
  3. Retrieve authorized candidate evidence.
  4. Attach provenance and trust labels.
  5. Remove expired, unrelated or unauthorized items.
  6. Place evidence in a clearly separated data area.
  7. Expose only the tools the current operation needs.
  8. Apply external authorization and approval before consequential actions.
🛡️ Safety Check
Threat: a malicious or compromised source inserts instruction-like content into retrieved material. Control: keep the source as data, maintain a trust boundary, validate tool arguments and enforce permissions outside the agent. Residual risk: an agent can still select the wrong evidence; defense in depth limits what that mistake can accomplish.

🎯 Use this when your agent combines instructions, memory, retrieved documents and tool responses in one execution.

8. 🏗️ Worked Example: Selecting Context From Several Sources

🧒 Kid Analogy
Someone asks, “Can I return this school uniform?” You find a current school rule, an old notice, a student's message, a receipt and a teacher's general comment. You do not throw all five papers at the teacher and hope. You work out which papers establish the answer.

Scenario: a service agent receives “Can this customer's replacement request be approved under the current service policy?” Retrieval returns six candidates:

SourceStatusSelection reasoning
Current policyAuthoritativePrimary evidence for eligibility.
Customer contractAuthoritativeEstablishes customer-specific terms.
Old policy PDFHistoricalTraceable, but not current approval authority.
Support ticketOperational evidenceMay establish what happened, not the governing policy.
Technician noteUnstructuredPotentially useful evidence; not automatically authoritative.
Inventory recordOperational dataRelevant only if availability is part of the decision.
  1. Define the decision. Determine eligibility under the current policy, not merely summarize the customer's history.
  2. Select governing evidence. The current policy and the customer contract become the central context.
  3. Add supporting evidence. The ticket joins only if it establishes facts such as the failure date.
  4. Preserve provenance. Each selected item keeps its source, owner, version and scope.
  5. Check for conflict. If the policy and contract disagree, the agent does not silently choose. The organization's precedence rule decides, or the case escalates.
  6. Check sufficiency. If the policy requires a document that is missing, enter an information-gap state.
  7. Separate recommendation from action. A proposed approval is not the approval transaction.
  8. Gate the tool. If policy requires human authorization, the action waits for it.

What the selection log would show for this task:

selected:  current-policy (v7), customer-contract-8831, ticket-4471 (failure date only)
evidence only (not authority): technician-note-221
rejected:  old-policy-2024        -> stale
           inventory-record       -> not needed for this decision
gaps:      none
conflicts: none
next:      propose approval; wait for human authorization
✅ Practical outcome
The final working context holds only the current policy, the customer's terms, the relevant service record and a little supporting evidence. The old policy remains traceable but does not compete with the current one as if both were equally authoritative.
💡 Why this matters
The retrieval system succeeded by finding six related sources. The selection system succeeded only when it reduced them to the evidence the decision needed, while keeping provenance and handling rules.
🛡️ Safety Check
Threat: a stale policy or untrusted note influences an approval action. Control: apply authority, version and scope checks before assembly, and require independent authorization for the approval tool. Residual risk: the selected evidence can still be incomplete, so the system needs explicit escalation and auditability.

🎯 Use this when an agent must combine several repositories before making a business decision or proposing an action.

9. 🏢 Enterprise Rollout: Governing the Context Pipeline

🧒 Kid Analogy
A school does not let every student change the rulebook, hand out master keys and rewrite attendance records. Different people own different parts of the system, and important changes need review.

Treat enterprise context engineering as an operational control plane, not a collection of prompts and retrieval queries.

AreaWhat good looks like
OwnershipNamed owners for instructions, retrieval sources, tool definitions, memory policy, data classification and action authorization.
Versioning and change controlInstructions, tool schemas, retrieval policies, source mappings and memory rules are versioned. A context change can alter behavior with no application-code change.
Readiness reviewIdentify what changed, affected tools, sources and users, access and classification impact, failure and escalation paths, logging coverage, approval rules; get owner sign-off; deploy with rollback.
Data classificationEvery source has a classification and access rule. “The agent can technically read it” is never the policy that decides what enters context.
Secrets and credentialsSecrets stay out of ordinary context. Tools obtain credentials through controlled identity and secret management.
Memory retention and deletionPersistent memory has an owner, provenance, purpose, retention and deletion rules, so it does not become an undocumented business dependency.
Cost governanceMonitor repeated retrieval, unnecessary tool calls, oversized histories and long multi-step workflows against business value.
ObservabilityTrace which sources were selected and rejected, which tools ran, which authorization checks occurred, which approvals were requested and what was executed. Redact sensitive content in traces.
AlertingUnusual tool usage, access outside expected scope, repeated approval attempts, abnormal outbound activity, unexpected context-configuration changes, tools used outside their purpose.
Incident responseA tested kill switch: revoke credentials, disable tools, quarantine a context source, stop workflows, and preserve traces.
🛡️ Safety Check
Threat: a compromised context source influences a high-impact workflow. Control: combine source governance, least-privilege permissions, approval gates, egress controls, action logging and a kill switch. Residual risk: no layer is perfect; the architecture is built to contain the consequences when one layer fails.

🎯 Use this when an agent is moving from an experiment into a business process with real customer, financial, operational or regulatory consequences.

10. ⚠️ Common Mistakes and Why They Fail

1. Treating retrieved or tool content as trusted instructions. Retrieved material can contain useful facts without having authority to control the agent. Microsoft's guidance classifies tool results and external context as untrusted input.

2. Selecting by similarity score alone. The closest match is not necessarily the authoritative one. Check authority, version and scope, not just relevance.

3. Authorizing only after retrieval, or with the agent's own access. If unauthorized content reaches the context and is hidden afterward, it can still influence the answer. Authorize at retrieval, using the end user's permissions, and keep selection-time checks as a second layer.

4. Assuming newest means most authoritative. A recent note can be less authoritative than an older approved policy. Use explicit source precedence and flag conflicts instead of picking silently.

5. Granting broad credentials “for convenience”. Broad access turns a context mistake into a larger incident. OWASP names excessive functionality, permissions and autonomy as root causes of excessive agency.

6. Relying on standing instructions alone as a security boundary. Instructions are one layer. Authorization and enforcement belong in the application and tool layers.

7. Stuffing the workspace instead of curating it. More information creates more selection work. Assemble context around the task, not around the maximum data available.

8. Unbounded memory with no provenance or expiry. A remembered statement can outlive the conditions that made it true. OWASP's 2026 article on memory and context poisoning explains why persistent state is itself an attack surface.

9. Letting derived items inherit more trust than their inputs. A summary or memory built from untrusted text is still untrusted. Propagate the lowest trust level and keep source pointers.

10. Shipping context changes with no review gate or tracing. Changing retrieval rules or tool descriptions changes behavior without code changes. Version, trace and keep rollback.

11. Approval fatigue. If people approve every harmless action they start approving mechanically. Concentrate review where reversibility, sensitivity or impact makes human judgment meaningful.

12. No kill switch. If you cannot quickly disable an agent's credentials or tools, the incident-response plan is incomplete.

💡 The recurring pattern
Most serious context failures are not fixed by adding more instructions. They are fixed by better separation, provenance, permissions, validation, approval, observability and containment.

11. 🛡️ Honest Limits

No current context-engineering technique fully eliminates prompt injection or every other way an agent can be manipulated. The practical objective is to reduce the probability of unsafe behavior and, more importantly, contain the blast radius when a component misbehaves.

Design for imperfect behavior. The surrounding system decides which data can be accessed, which tools can be called, which arguments are acceptable, which actions need approval, what gets logged and how the agent can be stopped. That posture matches current platform and security guidance on trust boundaries, validation, least privilege, approval and lifecycle governance.

12. ❓ FAQ

Is retrieval the same as context selection?
No. Retrieval identifies candidate information. Context selection determines which candidates enter the working context after considering relevance, authority, scope, freshness, sensitivity and provenance.
Should retrieved documents be treated as instructions?
Normally no. Retrieved documents are data or evidence. Application-controlled instructions and authorization policy stay separate, and tool results should be treated as untrusted and validated before they influence consequential operations.
What should an agent do when it cannot find enough information?
Identify the gap, try an approved retrieval path if one exists, and otherwise stop, explain what is missing and escalate or ask for input when the workflow requires it.
Why does provenance matter for agent memory?
A persistent memory item may influence a different task later. Provenance shows where it came from, when it was created, what scope it had and whether it should still be retained.
If I summarize untrusted content, is the summary safe?
No. A derived item inherits the lowest trust of its inputs, so a summary of untrusted text stays untrusted and must keep pointers back to its sources.
Does good context selection make an agent secure?
No. It is one security layer. Production systems still need least-privilege access, validation, approval controls, logging, monitoring, incident response and a way to revoke or stop the agent.

13. 🔗 References & Further Reading

Originality & attribution: Product and standards names belong to their owners. Platform guidance changes quickly, so check current documentation before relying on a specific capability.

14. 📝 Summary

  • Selection: decide what information an agent receives for a specific task and how it may be used. Available is not the same as required.
  • Start from the task: write the objective, required evidence, allowed actions and decision boundary as data.
  • Retrieval vs selection: retrieval finds candidates; selection checks authorization, authority, scope, freshness and conflicts, and logs why items were rejected.
  • Provenance: keep origin, version, scope and trust with every item, and let derived items inherit the lowest trust of their inputs.
  • Information gaps: check evidence against the required list in code, and stop or escalate rather than assume.
  • Working context: label every item and keep controlled policy separate from variable or untrusted data.
  • Enterprise rollout: own, version, classify, observe and be able to stop the context pipeline.
  • Honest limit: no technique eliminates every attack, so aim for defense in depth and a contained blast radius.

An agent should not receive information merely because the system can retrieve it. It should receive information because the application has deliberately decided it is relevant, authorized, appropriately scoped and useful for the task, with enough provenance preserved to understand where it came from. That is the difference between having context and engineering context.

Comments