Context reliability and safety is the discipline of making sure an AI agent works from facts that are present, current, relevant and consistent, and that nothing it reads can quietly take over its decisions. 🧭
Why it matters: an agent that books refunds, edits tickets or reads customer files acts on what it is shown. A missing policy page causes a wrong refund; a poisoned document can steer a real action; a leaky memory can expose one customer to another. Reliability and safety are one design problem. 🔐
📑 In This Post
- 1. Part A: Context Failures and Diagnosis
- 2. Checking Answers Against Supplied Sources
- 3. Citations and Evidence
- 4. Testing with Real Questions
- 5. Instruction Injection in Documents and Tool Results
- 6. Separating Trusted Instructions from Untrusted Content
- 7. Permissions and Privacy for Context and Memory
- 8. Enterprise Rollout
- 9. Common Mistakes
- 10. Honest Limits
- 11. FAQ
- 12. References
🔀 Quick Comparison: Trusted vs Untrusted Context
| Source | Examples | May give orders? | Handling |
|---|---|---|---|
| Trusted | Standing rules, allow-listed tool specs, authenticated user request | Yes | Versioned, reviewed |
| Untrusted | Web pages, uploaded files, retrieved docs, tool output, other agents | No, data only | Fence, label, cite |
| Semi-trusted | Memory notes the agent wrote earlier | No | Provenance + expiry |
1. Part A: Context Failures and Diagnosis
Four failure shapes. Missing: the needed fact never reached the agent. Stale: it arrived but was superseded. Excessive: so much arrived that the relevant part was buried or sensitive data rode along. Conflicting: two sources disagree and nothing says which wins.
effective_from and superseded_by metadata and prefer the newest.How to diagnose whether the agent received the needed facts
- Log the exact context assembled for each run: source IDs, versions, retrieval time.
- Ask: was the needed fact in that assembly? If not, it is a retrieval or permissions problem, not a reasoning problem.
- If present, was it fenced from noise, current, and uncontradicted?
- Replay the same assembly after your fix to confirm the right facts now arrive.
🎯 Use this when an answer is wrong and you must decide whether to fix data, retrieval or permissions.
2. Checking Answers Against Supplied Sources
Source-use checking asks: for each claim in an answer, which supplied passage supports it? Cheap structural checks come first: does every cited ID exist in the assembled context, and does the quoted span actually appear in that passage? Claims with no supporting passage are flagged for review or removed.
🎯 Use this when you need an automated gate before answers reach users.
3. Citations and Evidence
Make each claim point to a stable source ID, version and location (page or section). Prefer showing the supporting excerpt next to the claim so a reviewer verifies in seconds. Evidence records should live in logs too, so an auditor can reconstruct why an action happened.
answer = {
'claim': 'Refund window is 14 days',
'evidence': [{'source_id': 'policy-returns', 'version': 'v12',
'section': '3.2', 'excerpt_hash': '<HASH_PLACEHOLDER>'}]
}
🎯 Use this when reviewers, auditors or customers must be able to verify answers.
4. Testing with Real Questions
Test the pipeline, not the engine. Collect real questions from tickets and logs (with personal data removed) and, for each, record which sources must be retrieved, which must never be retrieved, and which actions are allowed. Then assert on the assembled context and the action log.
- Retrieval tests: required sources present; stale ones absent.
- Permission tests: a user without access never causes restricted data to be assembled.
- Safety tests: planted harmless fake instructions in test documents must not trigger any tool call.
- Change tests: rerun the set whenever instructions, tools or memory policy change.
🎯 Use this when gating a release or tracking whether a change broke context assembly.
5. Instruction Injection in Documents and Tool Results
Instruction injection (the industry alias is prompt injection) happens when text the agent reads, such as a web page, file or tool result, contains lines written to steer it. Agents are exposed because reading and acting share one channel. Warning signs: sudden tool calls unrelated to the user's request, requests to send data outward, or text addressing the agent directly.
Handling untrusted content, step by step
- Tag the content's origin and mark it data-only.
- Strip or neutralize markup that hides text from humans.
- Restrict tools for this task to the minimum needed.
- Require approval for any outbound or irreversible action triggered after reading it.
- Log content and resulting actions together.
🎯 Use this when agents read anything not authored by your own team.
6. Separating Trusted Instructions from Untrusted Content
Keep standing rules, tool specs and the authenticated user request in clearly separate slots from retrieved material. Wrap untrusted text with its origin so downstream checks can see it.
def wrap_untrusted(text, source_id):
# Label only; never treated as instructions by the loop
return {'role': 'data', 'source': source_id,
'trust': 'untrusted', 'body': text}
ALLOWED_TOOLS = {'lookup_order', 'search_policy'} # read-only
def call_tool(name, args, human_ok=False):
if name not in ALLOWED_TOOLS and not human_ok:
raise PermissionError('tool not allow-listed')
🎯 Use this when designing how context is assembled and tools are exposed.
7. Permissions and Privacy for Context and Memory
Decide what may enter context by the requester's rights, not the agent's. Classify data (public, internal, confidential, regulated). Memory needs provenance (who or what wrote it), scope (per user or shared), expiry and deletion paths.
memory_rule = {
'scope': 'per_user', 'source': 'user_confirmed',
'expires_days': 90, 'contains_pii': False
}
🎯 Use this when agents remember across sessions or serve many users.
8. Enterprise Rollout
- Ownership: a named owner for the context pipeline and a review board.
- Versioning: instructions, tool definitions and memory policies in source control with change control.
- Sign-off gate: pipeline tests, permission tests and security review pass before release.
- Access and classification: every source tagged and entitlement-checked.
- Secrets: tool credentials from a vault, short-lived, scoped per task.
- Retention: memory expiry and deletion on request.
- Cost governance: per-task spend limits and budget alerts.
- Observability: traces of context assembly and every action; dashboards.
- Alerting: policy violations and unusual actions page a human.
- Incident response: kill switch, credential revocation, memory quarantine, post-incident review.
🎯 Use this when moving from prototype to a governed production service.
9. Common Mistakes
- Trusting retrieved or tool content as instructions: reading and obeying share one channel, so any author of that content becomes a boss.
- Broad credentials 'for convenience': a tricked agent inherits every permission, magnifying damage.
- Standing instructions as the only boundary: text can be overridden; permissions cannot.
- Stuffing instead of curating: noise hides the right fact and widens data exposure.
- Unbounded memory: without provenance or expiry, one bad write persists forever.
- No review gate or tracing: you cannot explain or roll back a bad change.
- Approval fatigue: if humans approve everything, approval protects nothing; reserve it for high-risk actions with clear summaries.
- No kill switch: containment during an incident needs minutes, not a deploy cycle.
10. Honest Limits
No current technique fully removes instruction injection. Treat it as an open problem. Assume the agent can be tricked, then limit what a tricked agent can read, change and send. The goal is reduced risk and a contained blast radius, not certainty.
❓ 11. FAQ
My agent answered wrongly; where do I look first?
Open the assembled-context log. Missing fact means retrieval or permissions; present but ignored means placement or conflicts.
Is a citation proof the answer is right?
No. It shows provenance. Pair it with the source's trust label and review.
Can I stop injection with a strong rule in the standing instructions?
Not alone. Use permissions, allow-lists, egress controls and approvals as well.
What should memory store?
Small, confirmed, scoped facts with source and expiry; never secrets.
How often should we rerun context tests?
On every change to instructions, tools, data sources or memory policy, and on a schedule.
🔗 12. References & Further Reading
- OWASP GenAI Security Project
- NIST AI Risk Management Framework
- MITRE ATLAS
- Model Context Protocol documentation
- Google Secure AI Framework (SAIF)
All trademarks belong to their owners.
📝 Summary
- Failures come in four shapes: missing, stale, excessive, conflicting.
- Check that answers are backed by supplied sources.
- Citations connect claims to evidence but don't prove trust.
- Test the pipeline with real questions.
- Treat documents and tool output as data, not orders.
- Separate trusted and untrusted zones in code.
- Control what enters context and memory by requester rights.
- Govern with owners, gates, tracing, alerts and a kill switch.
- Accept honest limits and shrink the blast radius.
Happy building, and keep your backpack tidy! 🎒
Comments
Post a Comment