Graph Engineering for AI Agents: A Practical Guide to Task Assignment, Multi-Agent Coordination & Handoffs
Graph engineering is the discipline of deciding which task happens where, which agent owns it, what must happen before the next step, how work is routed, what information crosses a boundary, and what happens when something goes wrong.
Imagine a software-change request arriving in an enterprise engineering system. One AI agent understands the request, another checks security, another examines tests, another summarizes the evidence, and a human finally approves or rejects the release. The difficult part is not simply creating four agents. The difficult part is making the workflow between them predictable.
That is where workflow graphs become useful. A graph makes relationships visible: what depends on what, what can run in parallel, where a decision changes the route, where one agent hands control to another, where state is stored, and where execution must stop.
More agents do not automatically create a better system. Better task boundaries and clearer control flow usually matter more than the number of agents.
📑 In This Post
- What the workflow graph actually controls
- How to assign tasks to agents
- How multiple agents coordinate
- How to design safe handoffs
- How state, context, and authority move through the graph
- Branches, retries, failure recovery, and completion
- How to test a workflow graph
- Security and the danger of excessive agency
- Enterprise rollout
- Common mistakes
- FAQ
01
What the Workflow Graph Actually Controls
🟧 Child-friendly analogy: Think about a school field trip. The students are the work, the teachers are responsible for different groups, and the trip plan says which group goes where and when. If Group A must finish registration before Group B can enter the bus, that dependency is part of the plan. The people may change; the route and rules still matter.
An AI workflow graph represents the relationships among work units. A node can represent an AI agent, a deterministic function, a validation step, a database operation, a human approval, or another executable component. An edge describes how execution moves from one node to another.
The important insight for beginners is this: the graph is not the model. The model produces reasoning or content. An agent is a model-plus-capabilities arrangement. The graph describes how work is organized around that agent or among several agents.
The exact boundary differs across architectures. In one implementation, graph orchestration may be part of the application. In another, a framework runtime owns state transitions. In another, an agent runtime may itself expose handoffs and tool calls. OpenAI's Agents SDK, for example, describes agents, handoffs, agents-as-tools, guardrails, sessions, and tracing as runtime primitives, while LangGraph describes a graph in terms of nodes, edges, state, conditional routing, and execution. Those are concrete architecture choices, not universal definitions for every agent system.
| Part | Main responsibility | Simple question |
|---|---|---|
| Model | Generates or evaluates content within a model invocation. | “What should I produce or decide from this input?” |
| Agent | Combines a model with instructions, tools, state or runtime behavior. | “What job am I responsible for?” |
| Workflow graph | Defines dependencies, routes, transitions, joins, and completion conditions. | “Where does execution go next?” |
| Harness | Runs and controls an agent, including runtime policies and tool interaction. | “How is this agent actually operated?” |
| Sandbox | Provides an isolated execution environment when the architecture needs one. | “Where can this work safely execute?” |
This distinction becomes important during failures. Suppose an agent correctly determines that a security check is required, but the workflow never routes to the security node. That is a graph problem. Suppose the graph routes correctly but the agent cannot perform its job because it lacks the needed information. That is more likely an agent or context problem. Suppose the agent receives a sensitive tool it should not have. That is an authority and harness problem.
You need to explain an AI workflow to another engineer and want everyone to agree on which part owns routing, state, tools, and permissions.
02
How to Assign Tasks to Agents
🟧 Child-friendly analogy: A teacher does not tell five children, “Everyone do the whole project.” Instead, one child gathers facts, another draws, another checks spelling, and another presents the result. Clear ownership reduces duplicated work and arguments about who is responsible.
The first graph-engineering decision is therefore not “How many agents should I create?” It is “What is the smallest meaningful task that deserves its own responsibility boundary?”
A task deserves a separate agent when the work has a distinct objective, different tools or knowledge, a different risk profile, or a clear reason to change independently from the surrounding steps. A task does not deserve its own agent merely because another label sounds impressive.
A useful task contract can contain six things:
- Objective: what this node is trying to accomplish.
- Input: what information it may rely on.
- Output: what downstream nodes can consume.
- Authority: which tools and actions it may use.
- Failure behavior: what happens when the task is incomplete or uncertain.
- Completion evidence: what proves that the task is actually finished.
Notice something subtle: the task contract separates what the agent may think about from what the agent may do. That distinction becomes extremely valuable as workflows become more capable.
For our fictional running example, imagine this production-change workflow:
🟩 Worked example: a fictional production-change review
- Intake Agent turns the change request into structured work items.
- Test Agent examines the available test evidence.
- Security Agent reviews approved code and security evidence.
- Risk Coordinator combines the independent findings.
- Human Approver makes the release decision when policy requires human authority.
- Release Executor may perform the approved action only after the workflow reaches the correct authorization boundary.
Safety Check
Risk → A task may be assigned to an agent that has more authority than the task requires.
Control → Define the minimum tools, data scope, and downstream permissions for each task, and enforce those restrictions outside the model's own instructions.
Remaining risk → A well-scoped agent can still make an incorrect decision, so consequential actions need independent validation, policy checks, or human review where appropriate.
A workflow is becoming difficult to reason about because every node seems able to do everything. Split responsibility before splitting agents.
03
How Multiple Agents Coordinate
🟧 Child-friendly analogy: Think about a restaurant kitchen. The pastry station does not wait for the grill station to finish if the two dishes are independent. But both may need to finish before the head chef can plate the final order. That is parallel work followed by a join.
Coordination is where a graph begins to look like a real workflow rather than a collection of agents.
There are three common coordination patterns:
| Pattern | How it behaves | Useful when |
|---|---|---|
| Sequential | A finishes before B begins. | B genuinely depends on A. |
| Parallel | Independent tasks execute without waiting for one another. | Independent evidence can be gathered at the same time. |
| Fan-out / join | One node distributes work, then a downstream node waits for the required results. | Several specialists contribute to one combined decision. |
The crucial word is dependency. Do not make a step sequential just because it appears in a human-written checklist. Ask whether the later task actually requires the earlier result.
For our running example, the Test Agent and Security Agent can often inspect the same approved change snapshot independently. That creates a natural fan-out. The Risk Coordinator then waits for both results.
🟩 Worked example: coordination logic
The graph should also answer an uncomfortable question: what if one branch finishes and another never does? A join without a timeout, cancellation rule, or partial-result policy can leave the workflow waiting indefinitely.
This is one reason graph engineering is different from simply calling several agents. A graph should explicitly describe the lifecycle of the work.
Frameworks implement this idea differently. LangGraph, for example, documents normal edges and conditional edges, and notes that multiple outgoing edges can cause destination nodes to execute in parallel. Its documentation also distinguishes graph state from longer-lived stores and provides persistence and fault-tolerance mechanisms. Those details are implementation-specific, but they illustrate the broader engineering concept: task relationships need an executable representation.
You are trying to reduce latency or agent chatter. First prove that tasks are independent enough to run in parallel; then add concurrency.
04
How to Design Safe Handoffs
🟧 Child-friendly analogy: Imagine a relay race. A handoff is not “Everyone now owns the baton.” One runner releases it and the next runner accepts responsibility. A good handoff makes it obvious who is responsible next.
A handoff is a transfer of control or responsibility from one agent to another. It is not merely “sending a message” to another agent.
This distinction matters because there are at least two very different multi-agent designs:
| Design | Who owns the user-facing flow? | Typical purpose |
|---|---|---|
| Handoff | The receiving specialist can become the active agent. | Routing the conversation or task to the specialist who should own the next stage. |
| Agent as a tool / subtask | The original orchestrator remains in control. | Using a specialist for a bounded piece of work while keeping one coordinator responsible for the result. |
OpenAI's current Agents SDK documentation explicitly distinguishes these two patterns: a handoff transfers the active conversation to a specialist, while an agent used as a tool lets the original agent retain control and consume the specialist's result. The useful engineering lesson is broader than the particular SDK: decide whether the specialist is taking ownership or merely supplying evidence.
A robust handoff should answer four questions:
- Why is the handoff happening?
- What information should cross the boundary?
- What authority does the receiving agent gain?
- What happens if the receiving agent cannot complete the task?
The do-not-transfer part is easy to overlook. A handoff can accidentally become a data-sharing mechanism that passes the entire transcript, every tool result, or secrets that the receiving agent does not need.
OpenAI's handoff documentation also exposes implementation controls for customizing what is passed onward, including input filtering, handoff metadata, and callbacks. It specifically notes that authorization checks tied to parsed handoff data should occur before side effects in the handoff callback. The exact API is SDK-specific, but the design principle is general: authorization should be checked at the boundary where authority changes.
Safety Check
Risk → An agent passes control to a specialist whose tools are more powerful than the original task requires.
Control → Treat every handoff as an authority boundary: validate the destination, filter transferred information, and re-check downstream authorization.
Remaining risk → Even a correctly authorized specialist can misuse legitimate access or misunderstand the transferred task, so consequential operations still need policy enforcement and monitoring.
The next specialist should clearly become responsible for a stage of the workflow. Use a bounded subtask instead when you only need another agent's analysis.
05
How State, Context, and Authority Move Through the Graph
🟧 Child-friendly analogy: Think of a school office folder. The folder holds official progress, each teacher gets the pages needed for their subject, and the school administrator controls which records can be changed. Giving someone the whole filing cabinet just because they need one page is poor organization and poor security.
Graph systems become confusing when three different things are treated as the same:
- State: what the workflow knows about the current run.
- Context: the information selected for a particular agent or task.
- Authority: what the runtime actually permits the component to do.
A state object might contain a change identifier, current route, review results, timestamps, approval status, and failure count. Context is the smaller, task-specific slice presented to an agent. Authority is enforced by the runtime and downstream systems.
That means this statement is dangerous: “The agent was told not to deploy.” Instructions describe expected behavior; they do not necessarily create technical authorization. The stronger statement is: “The agent has no deployment capability or downstream permission at this stage.”
🟩 Worked example: state versus context
Workflow state: change ID, security result, test result, approval state, retry count.
Security-agent context: approved code snapshot, requested change, relevant security policy, test summary.
Security-agent authority: read-only access to the approved snapshot and approved analysis tools.
Release-agent authority: available only after the required approval transition and downstream authorization checks.
Persistence is another graph-engineering concern. A long-running workflow may need to resume after an interruption. LangGraph documents checkpoint-based persistence for thread-scoped graph state and separate stores for longer-lived application data. This is one architecture's implementation, but the conceptual split is useful: the data required to resume this run is not necessarily the same as knowledge that should survive across runs.
Safety Check
Risk → Sensitive data from one user, workflow, or tenant gets carried into another agent or run because the system passes broad state by default.
Control → Define explicit state boundaries, task-specific context projection, tenant-aware authorization, and retention rules.
Remaining risk → Data can still leak through generated summaries, tool results, logs, or persistence layers, so inspect the entire data path rather than only the prompt.
A multi-agent system keeps getting “larger” context instead of more relevant context. Model what should cross each edge rather than forwarding everything.
06
Branches, Retries, Failure Recovery, and Completion
🟧 Child-friendly analogy: A delivery route may say, “If the store is open, deliver the package. If it is closed, leave it at the approved pickup point.” The route is not broken because there are two possible outcomes. The important thing is that both outcomes were designed in advance.
A production graph should not describe only the happy path.
For every meaningful node, ask:
- What if the input is invalid?
- What if the dependency times out?
- What if the agent returns an incomplete result?
- What if the same task is attempted twice?
- What if the user or operator cancels the workflow?
- What proves the workflow may safely stop?
Retries are particularly important. A retry is not automatically safe. If a node sends an email, submits an order, changes a database row, or triggers an external operation, repeating the same call can create a duplicate side effect.
A graph also needs a real completion rule. “The last agent produced some text” is often too weak. Completion might mean:
- required reviews are present,
- required checks passed,
- the authorization state permits the next operation,
- the output satisfies a defined schema, and
- the workflow recorded a final status.
This is where graph engineering starts to overlap with reliability engineering. The workflow should be able to answer not only “What did the model say?” but also “What state is the system in now?”
For example, LangGraph currently documents configurable retries, per-node timeouts, and error handlers as distinct fault-tolerance mechanisms. That is a useful reminder that failure policy belongs to the execution design, not merely to the agent's instructions.
A workflow has external side effects or can run for a long time. Design the retry, timeout, cancellation, and resume behavior before production traffic arrives.
07
How to Test a Workflow Graph
🟧 Child-friendly analogy: Testing one actor in a play is not enough. You also need to check whether the actors enter in the correct order, whether the right person gets the message, and what happens when someone misses a cue.
A common mistake in agent systems is to test only the quality of individual agent responses. That can miss failures in the workflow itself.
Graph testing should cover at least five dimensions:
| Test area | Question | Example assertion |
|---|---|---|
| Routing | Did the run take the expected route? | Security review is mandatory for authentication changes. |
| Dependencies | Did downstream work wait for required evidence? | Release planning does not start before required review results exist. |
| Handoffs | Did only the intended information and authority cross the boundary? | The security agent receives no deployment capability. |
| Recovery | Does a failed node recover or stop safely? | A timeout routes to a defined recovery state. |
| Completion | Can the workflow prove it is done? | Final status is recorded only after mandatory gates pass. |
A strong test suite includes both positive paths and negative paths. For example:
🟩 Worked example: graph test cases
- Normal change → all required reviews → approval → execution.
- Security review fails → no execution node is reachable.
- Test service times out → bounded retry → recovery route.
- A specialist returns malformed structured output → validation catches it.
- A handoff is attempted without required authorization → transition is rejected.
- The operator cancels after review but before execution → execution never starts.
Tracing is especially useful here because the unit of debugging is not always the individual model call. You may need the full path: input → route decision → node execution → tool call → handoff → state update → next node. OpenAI's Agents SDK documents built-in tracing for agent runs; LangGraph exposes execution and event-streaming mechanisms. Again, the implementation varies, but the operational requirement is consistent: a production graph should leave enough evidence to reconstruct what happened.
An agent appears to “fail randomly.” First reconstruct the route and state transitions. The root cause may be the graph rather than the model output.
08
Security: Bound the Damage of Bad Routing
🟧 Child-friendly analogy: Imagine giving a child a school keycard. The safe design is not “Tell the child to behave.” The safer design is a card that only opens the rooms needed for that activity.
Multi-agent workflows create another security dimension: one component can influence what another component does. That can be useful for specialization, but it also creates propagation paths for mistakes or manipulated instructions.
OWASP's 2025 guidance describes excessive agency in terms of excessive functionality, excessive permissions, and excessive autonomy. It also recommends minimizing tools, minimizing permissions, avoiding unnecessarily open-ended capabilities, enforcing authorization downstream, and using human approval for higher-impact actions where appropriate.
For workflow graphs, this leads to a simple rule:
A routing decision should never be allowed to silently expand an agent's authority beyond what the next task requires.
Indirect prompt injection is especially relevant to agent graphs because untrusted content can influence later actions. NIST has described agent hijacking as a form of indirect prompt injection in which malicious instructions are placed into data an agent consumes. In a graph, that means the malicious influence may begin in one node and affect the behavior of later nodes.
Safety Check
Risk → External content influences an agent that can trigger a privileged downstream action.
Control → Separate untrusted content from control decisions, minimize tool scope, enforce authorization at the downstream system, constrain transitions, and add review gates for consequential actions.
Remaining risk → No single prompt rule, filter, approval dialog, or sandbox eliminates instruction-injection risk. A compromised component may still generate misleading recommendations or waste resources, so defense in depth remains necessary.
This is also why “human in the loop” should be designed carefully. A human approval step is strongest when the person sees a trustworthy, independently derived summary of the action, not merely a model-generated statement saying “everything looks safe.”
In high-impact workflows, the downstream service should still enforce its own authorization. The graph can decide when an action is eligible; it should not be the sole place that decides whether the caller is authorized.
Your graph can reach production systems, private data, financial actions, communications systems, identity controls, or infrastructure. Treat every edge that changes authority as a security boundary.
09
Enterprise Rollout
A graph that works in a notebook is not automatically ready for enterprise operation. The production challenge is less about drawing a larger graph and more about creating ownership, observability, control, and change discipline.
A practical rollout can happen in stages.
- Stage 1 — deterministic skeleton: build the route with simple placeholder nodes before adding complex agent behavior.
- Stage 2 — bounded specialists: introduce agents with small, explicit responsibilities.
- Stage 3 — state and persistence: define what survives interruption and what must be discarded.
- Stage 4 — security gates: apply authorization, approval, logging, and data boundaries.
- Stage 5 — failure testing: inject timeouts, malformed outputs, missing dependencies, cancellation, and rejected approvals in a controlled test environment.
- Stage 6 — controlled production rollout: use bounded traffic, monitoring, rollback procedures, and explicit ownership.
Each production run should ideally answer: who started it, which version of the graph ran, which nodes executed, what decisions changed the route, what external actions occurred, and where the workflow ended.
Production ownership checklist
- Graph owner and service owner are identified.
- Tool permissions are reviewed independently of agent instructions.
- Important transitions and external side effects are logged.
- State retention and sensitive-data handling are defined.
- Budgets and retry limits are bounded.
- A failed run can be investigated without relying on a model's own explanation.
- A dangerous transition can be disabled without redesigning the entire application.
One additional enterprise concern is graph versioning. A change to routing logic can affect in-flight work. Some systems apply a new graph version to future and existing runs differently. That behavior is architecture-specific and should be verified before deployment rather than assumed.
A graph becomes business-critical. At that point, graph definitions should be treated like production software, with ownership, testing, observability, controlled change, and incident procedures.
10
Common Mistakes
The most expensive mistakes are often structural rather than model-specific.
1. Creating an agent for every sentence in the workflow.
Cause: confusing modularity with specialization.
Consequence: more handoffs, more context transfer, more latency, and harder debugging.
Correction: split by meaningful responsibility, not by arbitrary step count.
2. Letting every agent see the entire transcript.
Cause: “it is easier than selecting the useful information.”
Consequence: unnecessary context, accidental disclosure, and unclear ownership of information.
Correction: define explicit context projections for each boundary.
3. Using instructions as the only security boundary.
Cause: assuming an agent will always follow its role description.
Consequence: a model mistake or manipulated input can still reach powerful tools.
Correction: enforce least privilege in the runtime and downstream systems.
4. Treating a handoff as a data dump.
Cause: forwarding everything to avoid missing information.
Consequence: sensitive data crosses boundaries unnecessarily and debugging becomes difficult.
Correction: transfer a deliberate task package.
5. Retrying side effects blindly.
Cause: applying the same retry pattern to read and write operations.
Consequence: duplicates or repeated external actions.
Correction: make side-effecting operations idempotent where possible, or route failures to reconciliation.
6. Testing agents but not the graph.
Cause: measuring answer quality while ignoring route correctness.
Consequence: the “right” agents produce the wrong overall workflow.
Correction: test paths, joins, handoffs, permissions, retries, cancellation, and completion.
7. Having no cancellation strategy.
Cause: designing only success and failure outcomes.
Consequence: work may continue after the user or operator no longer wants it to.
Correction: make cancellation a first-class workflow state where the architecture requires it.
8. Assuming the graph itself is the security boundary.
Cause: concentrating authorization only in orchestration code.
Consequence: a downstream integration may still accept an unauthorized request through another path.
Correction: enforce authorization at the system that owns the protected resource.
A production workflow is not “finished” when the happy path works. It is finished when the important alternative paths have explicit behavior too.
11
❓ FAQ
Q1. Does every AI agent application need a workflow graph?
No. A simple agent that answers a question with a small number of deterministic tool calls may not need a separately modeled graph. Graph engineering becomes more valuable when the system has multiple meaningful stages, branches, parallel work, specialists, approvals, retries, or recovery behavior that should be made explicit and testable.
Q2. Should the model decide the next node, or should code decide it?
Either can be appropriate. Model-driven routing is useful when the task is open-ended and the model needs to choose among legitimate specialist paths. Code-driven routing is useful when policy or business rules must be deterministic and auditable. Many systems combine both: code defines the allowed routes while the model helps decide which allowed route is relevant.
Q3. What is the difference between a handoff and asking another agent for help?
A handoff usually changes who owns the next stage of work. Asking another agent as a bounded subtask can keep the original coordinator in control. The right choice depends on whether the specialist should own the conversation or simply contribute an intermediate result.
Q4. What should be transferred during a handoff?
Transfer the smallest task package that lets the receiving component do its job correctly: the objective, relevant inputs, necessary evidence, useful constraints, and any structured metadata needed for the next transition. Avoid treating the entire transcript or all available state as the default handoff payload.
Q5. How do I know whether my graph is production-ready?
You should be able to explain the allowed routes, task ownership, state boundaries, authority boundaries, failure behavior, retry rules, cancellation behavior, completion criteria, and operational evidence. Then test those behaviors explicitly. A workflow that produces good answers but cannot explain or control its important transitions is not yet operationally mature.
12
🔗 References & Further Reading
Product and organization names remain the property of their respective owners.
- OpenAI Agents SDK — Agent orchestration — documents model-driven versus code-driven orchestration and the distinction between agents-as-tools and handoffs.
- OpenAI Agents SDK — Handoffs — documents handoff configuration, input filtering, handoff metadata, and boundary-related behavior.
- LangGraph — Graph API overview — documents nodes, normal edges, conditional edges, entry points, and parallel execution behavior.
- LangGraph — Persistence — documents checkpoint-based graph state and separate stores for longer-lived application data.
- LangGraph — Fault tolerance — documents retries, timeouts, and error handling at the node level.
- OWASP GenAI Security Project — LLM06:2025 Excessive Agency — provides guidance on minimizing functionality, permissions, and autonomy and enforcing downstream authorization.
- NIST — Strengthening AI Agent Hijacking Evaluations — discusses indirect prompt injection and agent hijacking risks.
13
📝 Summary
- Assign tasks by responsibility, not by agent count.
- Use dependencies to decide what must be sequential and what can run in parallel.
- Treat handoffs as explicit ownership and authority boundaries.
- Keep state, task-specific context, and authorization conceptually separate.
- Design retries, cancellation, recovery, and completion before production.
- Test the graph itself—not only the quality of individual agent responses.
- Use least privilege and downstream authorization so a bad model decision has a limited blast radius.
An AI workflow becomes easier to trust when every important transition has a clear owner, a defined input, a bounded authority, an observable outcome, and an explicit next step.
Comments
Post a Comment