Multi-Agent Context Engineering: What Should Cross the Boundary Between AI Agents
Multi-agent context engineering is the discipline of deciding what each agent sees, what it is allowed to believe, what it may remember, and what it is permitted to do.
Adding agents does not make a system smarter by itself. It adds boundaries, and every boundary is a place where information can be lost, corrupted, or quietly promoted into authority. This post shows how to design those boundaries on purpose, using one concrete enterprise scenario, corrected code, and the behavior of real frameworks as documented in 2026.
- 1. Start with one agent: when more agents actually pay off
- 2. A running example: the supplier who “changed banks”
- 3. Who owns the next decision? Handoff, agent-as-tool, or code
- 4. The context packet: what is allowed to cross a boundary
- 5. Trust: from “read this” to “do this”
- 6. Memory is governed state, not a notebook
- 7. Tools, identity and the control plane
- 8. Approvals a human can actually judge
- 9. Observability, budgets and recovery
- 10. Rollout: making context a governed capability
- 11. Failure modes mapped to controls
- FAQ
- References and further reading
- Summary
1. Start with one agent: when more agents actually pay off
Read that as a design rule: the honest reasons to split into agents are narrow.
- Context isolation. A subtask would flood the main context with search results or tool output that the caller never needs again.
- Parallelism. Independent threads can run at the same time.
- Tool overload. One agent with dozens of tools picks worse than three agents with a handful each.
- Trust separation. One component must read hostile content, and another must hold the power to act. This is a security reason, and the one most tutorials skip.
If none of those applies, use one agent with good tools. The rest of this post assumes at least one does.
Use this when: you are about to add an agent because the diagram looks more impressive, rather than because a boundary is doing real work.
2. A running example: the supplier who “changed banks”
We will use this scenario for the rest of the post, because it is the pattern behind real business-email fraud and it exercises every layer. A supplier emails accounts payable, attaches a PDF on letterhead, and asks to change the bank account used for future payments. The email also says the next invoice is urgent.
- Reader agent. Reads the email and PDF. Has no tools. Can only emit a fixed schema.
- Validator. Plain code, not a model. Checks the reader’s output.
- Vendor-data agent. Read-only access to the supplier master record.
- Policy agent. Applies written rules and proposes the next step, such as a callback to a number already on file.
- Control plane. Authorizer, human approver, and the only code path allowed to write bank details.
3. Who owns the next decision? Handoff, agent-as-tool, or code
| Pattern | Who owns the next decision | What context moves by default | Best fit |
|---|---|---|---|
| Handoff | The receiving specialist | Usually the full prior conversation, unless you filter it | Specialist should answer the user directly |
| Agent as tool | The calling agent | Only the task input you pass | A focused subtask whose result the caller will use |
| Code-orchestrated | Your workflow code | Only the typed state you declare | Business sequences that must be predictable and auditable |
Control: keep an explicit allow-list of transitions, and separately control what context each transition carries. These are two different controls.
Residual risk: a permitted transition can still carry bad data, so the receiver validates the packet (section 4).
Use this when: you are choosing between “let the specialist take over” and “ask the specialist and keep control.”
4. The context packet: what is allowed to cross a boundary
A context packet is a typed object that crosses a boundary. Its key property is that facts, claims, decisions and permissions live in different fields, so a claim can never pass itself off as a fact. Here is the packet our policy agent receives:
{
"task": {"id": "T-2041", "type": "supplier_bank_change", "owner": "policy_agent"},
"facts": [
{"name": "supplier_id", "value": "S-883", "source": "vendor_master",
"trust": "system_verified", "fetched_at": "2026-10-02T09:14Z"}
],
"claims": [
{"name": "new_account_last4", "value": "4471", "source": "inbound_email",
"trust": "untrusted", "extracted_by": "reader_agent", "schema_valid": true},
{"name": "sender_domain_matches_master", "value": false,
"source": "validator", "trust": "system_verified"}
],
"decisions": [
{"name": "risk_tier", "value": "high", "made_by": "policy_rules_v12", "kind": "deterministic"}
],
"constraints": ["no bank-detail write without signed approval"],
"allowed_next": ["request_callback_verification", "escalate_to_human"],
"expires_at": "2026-10-02T11:14Z"
}
The reader agent is the only component that touches hostile text, and its output is forced through a validator written as ordinary code. The validator is where untrusted content becomes bounded data:
import re
REASONS = {"bank_change", "address_change", "contact_change", "other"}
ALLOWED_KEYS = {"reason_code", "new_account_last4", "sender_domain", "urgency_claimed"}
def validate_reader_output(raw: dict, master_domains: set) -> dict:
if set(raw) - ALLOWED_KEYS: # no extra fields, ever
raise ValueError("unexpected fields from reader")
if raw.get("reason_code") not in REASONS:
raise ValueError("bad reason_code")
last4 = str(raw.get("new_account_last4", ""))
if not re.fullmatch(r"\d{4}", last4):
raise ValueError("bad account suffix")
domain = str(raw.get("sender_domain", "")).lower()
return {
"reason_code": raw["reason_code"],
"new_account_last4": last4,
"sender_domain_matches_master": domain in master_domains,
"urgency_claimed": bool(raw.get("urgency_claimed", False)),
} # enums, digits and booleans only: no free text survives
Two design choices matter here. First, there is no free-text field, so an injected sentence has nowhere to travel. Second, the packet expires, so a stale decision cannot be replayed into a later step. Keep the raw email as an evidence reference (an ID the human reviewer can open), not as content pushed into the next agent’s prompt.
Control: typed fields, provenance on every item, validation in code, and an expiry.
Residual risk: the validator itself can be wrong, so the highest-impact action still needs an independent control (section 7).
Use this when: you keep forwarding whole transcripts between agents because nobody has defined the state that matters.
5. Trust: from “read this” to “do this”
The standard term for text that tries to steer an agent is prompt injection, and when the text arrives through a document, web page or tool result it is indirect prompt injection. Use that vocabulary; it is what your security team, auditors and search readers use.
Label every item by its source, not its tone
- System-controlled: produced by your own application code.
- Policy-controlled: a decision from a versioned rules engine.
- Verified tool result: structured output from a controlled tool, validated by code.
- Authenticated user input: the user’s own request. It carries the user’s authority and no more.
- Untrusted content: files, web pages, inbound email, retrieved text, third-party messages.
- Agent-generated claim: another agent’s statement, which is only as trustworthy as everything that agent read.
Apply the Rule of Two along the whole path
The rule is stated for one agent in one session. Multi-agent systems create a gap: a reader that sees untrusted input (A) and an operator that holds sensitive access and write power (B and C) each pass the check alone. If the reader’s free-text summary flows into the operator, the attacker’s words have reached an agent that can act. This is our extension of the rule, not part of Meta’s statement, and it is why the validator in section 4 exists.
Control: separate the reader from the actor, pass only validated typed fields, and keep authority in the control plane.
Residual risk: no current technique fully eliminates injection. The goal is to bound what a fooled component can do.
Use this when: any agent reads inbound email, documents, web pages, tickets, or another agent’s free text.
6. Memory is governed state, not a notebook
The lesson is about where memory is allowed to land. Two rules cover most of it:
- Write policy. Only code at or above the trust level of the fact may write it. Text derived from untrusted content is stored as a claim with its source, never as a standing instruction.
- Read policy. Retrieved memory enters a prompt as labeled data, never in the highest-trust instruction slot.
{
"id": "mem-7731",
"tenant": "acme-eu",
"kind": "fact",
"text": "Callback number for supplier S-883 confirmed by phone",
"source": {"workflow": "T-1987", "verified_by": "human:ap-lead"},
"created_at": "2026-08-14", "review_after": "2027-02-14",
"readable_by": ["policy_agent"], "delete_with": "supplier S-883 offboarding"
}
Control: owner, tenant scope, source, review date, readers, and deletion path on every record; no untrusted text in instruction slots.
Residual risk: a legitimate record can go stale, so decisions that matter re-check the source of truth.
Use this when: your agents carry anything across turns, cases, or users.
7. Tools, identity and the control plane
Think of two planes. The information plane holds facts, evidence, memory and messages. The control plane holds identity, credentials, policy, approvals and audit. The agent proposes; the control plane decides. The code below fixes a common bug, where an action marked “needs approval” is returned as simply allowed:
import hashlib, hmac, json, time
from enum import Enum
class Decision(Enum):
ALLOW = "allow"
NEEDS_APPROVAL = "needs_approval"
DENY = "deny"
POLICY = {
"policy_agent": {"allow": {"read_vendor_master", "request_callback"},
"needs_approval": set()},
"payments_service": {"allow": set(), "needs_approval": {"update_bank_details"}},
}
def authorize(agent_id: str, action: str) -> Decision:
rules = POLICY.get(agent_id)
if rules is None:
return Decision.DENY # unknown caller: deny
if action in rules["allow"]:
return Decision.ALLOW
if action in rules["needs_approval"]:
return Decision.NEEDS_APPROVAL # NOT allowed yet
return Decision.DENY # default deny
def fingerprint(agent_id, action, args) -> str:
body = json.dumps({"a": agent_id, "x": action, "p": args}, sort_keys=True)
return hashlib.sha256(body.encode()).hexdigest()
def approval_valid(ap: dict, fp: str, requester: str, key: bytes) -> bool:
if ap["fingerprint"] != fp or ap["expires_at"] < time.time():
return False
if ap["approver_id"] == requester: # no self-approval
return False
msg = f'{fp}|{ap["approver_id"]}|{ap["expires_at"]}'.encode()
return hmac.compare_digest(
hmac.new(key, msg, hashlib.sha256).hexdigest(), ap["signature"])
def execute(agent_id, action, args, tools, approval=None, key=b""):
d = authorize(agent_id, action)
if d is Decision.DENY:
raise PermissionError(f"{agent_id} may not {action}")
if d is Decision.NEEDS_APPROVAL:
fp = fingerprint(agent_id, action, args)
if not approval or not approval_valid(approval, fp, agent_id, key):
raise PermissionError("valid signed approval required")
return tools[action](**args) # log before and after
The approval is bound to the exact action and arguments through the fingerprint, so an approval for “update account ending 4471” cannot be reused for a different account. The signing key lives in the approval service, never in an agent’s context. In production this logic belongs in your identity and policy layer rather than application code, but the shape stays the same.
What the 2026 MCP specification changes for this layer
- Gateways get a real hook. Header-based method and tool names let you enforce allow-lists and rate limits at a gateway, outside any agent.
- State becomes visible. An explicit handle is something you can log, scope to a tenant, and expire. Hidden transport sessions are not.
- Cacheable tool lists help pinning. Hash the approved tool catalog, diff it on refresh, and alert when a tool description changes. Tool descriptions are text the model reads, so treat them as untrusted input until reviewed.
- Transport auth is not action auth. A valid token proves who connected. It does not decide whether this agent may call this tool with these arguments.
Control: separate identity per agent, least privilege, default deny, argument-bound approvals, and logs of every call.
Residual risk: policy can be misconfigured, so sensitive tools must be easy to inspect and revoke.
Use this when: an agent can read or change enterprise systems and you need a hard line between what it knows and what it may do.
8. Approvals a human can actually judge
Design the approval screen so the reviewer judges facts the system computed, not a story the agent wrote:
- Show old value, new value, supplier, and amount from system records, side by side.
- Show what is unverified in plain terms, for example “sender domain does not match the master record.”
- Show the source evidence as a link to the original, not as the agent’s paraphrase.
- Reserve approval for high-impact boundaries so reviewers do not stop reading.
- On resume from a checkpoint, re-validate expiry and re-fetch facts. An approval granted yesterday is not proof of today’s state.
Use this when: a person is the last control before money moves, data leaves, or a privilege changes.
9. Observability, budgets and recovery
A trace should answer, for any run: which agent owned each step; which packet crossed which boundary and with what trust labels; which content came from outside the control plane; which tools were called and which authorization decision ran; where a human approved; and what failed and where it can resume. Separate failures by type, because the right response differs:
- Transient: retry with a cap.
- Data quality: ask for clarification or another source.
- Authorization or policy violation: stop, record, escalate.
- Context corruption: discard the affected state and rebuild from the last trusted checkpoint.
- Cascade (ASI08): one agent’s bad output feeds many. Add schema checks at every edge so errors stop at the first boundary.
Add budgets per task (turns, tool calls, spend) and a kill switch that disables a tool, workflow or agent identity without a redeploy. Alert on repeated authorization failures, unexpected external messages, long handoff chains, and spikes in write attempts.
Use this when: you are moving from a demo to something people depend on.
10. Rollout: making context a governed capability
Every context component is a production dependency. Give each one the following before release:
| Component | Owner and version | Review gate | Rollback or removal |
|---|---|---|---|
| Instructions and tool descriptions | Named owner, versioned, diffed in review | Trace review on representative runs | Previous version restorable |
| Context packet schemas | Versioned like an API | Compatibility check for every consumer | Dual-read period |
| Memory policies and stores | Owner, tenant scope, retention | Privacy and data-classification sign-off | Bulk quarantine and delete path |
| Tools, scopes and identities | Per-agent identity, least privilege | Security approval for any new write path | Credential revocation and kill switch |
| Handoff graph | Documented allowed transitions | Cycle and depth review | Disable an edge without redeploy |
Keep secrets out of context entirely; tools receive credentials from a secrets manager at call time. Classify data before it crosses a boundary, because it cannot be un-sent afterward. For governance vocabulary, NIST’s Generative AI Profile frames AI risk management as a lifecycle activity, which fits treating context as a managed asset.
11. Failure modes mapped to controls
| Failure mode | OWASP agentic entry | Primary control in this post |
|---|---|---|
| Untrusted text steers an agent | ASI01 Agent goal hijack | Reader/actor split, typed validator, no free text across edges |
| Valid tool used for the wrong job | ASI02 Tool misuse | Default-deny allow-lists, argument-bound approvals |
| Over-broad credentials | ASI03 Identity and privilege abuse | Per-agent identity, least privilege, no secrets in context |
| Poisoned tool or component | ASI04 Supply chain | Pinned tool catalog, reviewed descriptions |
| Persistent poisoned state | ASI06 Memory and context poisoning | Write and read policies, expiry, re-check source of truth |
| Spoofed agent-to-agent messages | ASI07 Insecure inter-agent communication | Authenticated, schema-checked packets with expiry |
| One bad output spreads | ASI08 Cascading failures | Validation at every edge, budgets, kill switch |
| Persuasive agent misleads reviewer | ASI09 Human-agent trust exploitation | Approval screens built from system facts |
FAQ
References and further reading
Summary
- Add agents only for isolation, parallelism, tool focus, or trust separation, and expect higher token cost.
- Pick handoff, agent-as-tool, or code orchestration by who owns the next decision. Do not assume the framework minimizes context for you.
- Cross boundaries with typed context packets that separate facts, claims, decisions and permissions.
- Count the Rule of Two along the whole data path, and put a code validator between readers and actors.
- Treat memory as governed state with owners, expiry and read and write policies.
- Keep authority in a control plane with default deny, per-agent identity and argument-bound approvals.
- Build approval screens from system facts, and add traces, budgets and a kill switch before production.
Comments
Post a Comment