Skip to main content

Graph Engineering for AI Agents: A Practical Guide to Human Approval, Step Permissions & Secure Handoffs

Calculating read time…
AI agents become much more useful when they can coordinate several steps, but that same freedom creates a difficult engineering question: who is allowed to do what, at which step, with which information, and who gets the final say?

A workflow graph gives us a useful way to answer those questions. Instead of treating an agent as one large loop that can see everything and call anything, we can model work as connected stages: collect information, validate it, hand the work to a specialist, prepare an action, pause for approval, and finally execute the authorized operation.

This article focuses on three controls that become especially important as an agent workflow grows: human approval, step-level permissions, and secure handoffs between agents. The examples are intentionally fictional, so they illustrate engineering patterns rather than claiming to describe a private production implementation.

🧒 Child-friendly analogy

Imagine a school trip. One person checks the list, another checks whether everyone has permission to go, a teacher decides whether the bus can leave, and the driver controls the bus. You would not give every person every key. A secure agent graph works similarly: each step gets only the authority it needs.

Engineering principle

The graph should make authority visible. A model may suggest an action, but the workflow should determine whether that action is permitted, whether approval is required, and which component is actually allowed to execute it.

🔀 Quick Comparison
Control Main question Best enforcement point Typical failure avoided
Human approval Should this high-impact action happen now? Workflow state before execution Unsafe or unexpected consequential action
Step permission What can this node or identity do? Tool gateway and downstream authorization Excessive agency and over-privilege
Secure handoff What information and authority cross to the next agent? Handoff boundary and receiving-agent policy Unnecessary data or authority propagation

01
What the graph actually controls

Before adding approvals or permissions, it helps to separate several ideas that are often mixed together.

The model generates predictions or decisions from the information supplied to it.

The agent is the application-level construct that combines a model with instructions, tools, state, and runtime behavior.

The workflow graph describes the relationships between tasks and execution paths: which step can follow another, where a branch occurs, where work pauses, and where the workflow ends or recovers.

The harness is the runtime machinery that actually runs the agent or agents and enforces application-level behavior. A graph can live inside that harness or inside a larger application.

The tool performs an operation outside the model, such as reading an account, creating a draft, or submitting a transaction.

The sandbox, when one is used, isolates execution. It can reduce the impact of operations performed inside the sandbox, but it does not automatically decide whether a business action is authorized.

Context is the information supplied to an agent for a particular step. Context engineering is therefore about selecting and organizing information; graph engineering is about coordinating work and transitions.

🧒 Child-friendly analogy

Think of a restaurant kitchen. The recipe is not the same thing as the chef, and the chef is not the same thing as the kitchen doors, payment terminal, or manager. A workflow graph is closer to the order of work and the rules about who handles each stage.

A simple graph might look like this:

intake ↓ validate ↓ policy_check ↓ prepare_action ↓ [approval required?] ├── no → execute └── yes → human_review → execute ↓ verify ↓ complete

The important observation is that approval is not merely a sentence in a prompt. It is a state in the workflow. Likewise, a permission is not merely a sentence such as “do not delete data.” It should be enforced at an execution boundary that can reject an unauthorized operation.

🛡️ Safety Check

Risk: treating the model as the final security authority.

Control: enforce authorization and execution rules outside the model at the graph, tool, gateway, and downstream-service layers.

Remaining risk: a compromised or mistaken model can still propose the wrong action, so the workflow needs validation, bounded permissions, observability, and recovery.

🎯 Use this when...

You are designing an agent that moves beyond a single response and begins making decisions, calling tools, or delegating work across multiple steps.

02
Human approval as a graph state transition

Human approval is most useful when an operation is meaningful enough that automated execution should stop and ask a person to make a decision.

Examples include releasing a payment, changing a security configuration, deleting important data, publishing externally, or approving a high-impact business transaction. The precise threshold depends on the application and its risk model.

🧒 Child-friendly analogy

Suppose an assistant prepares your lunch but cannot start the expensive oven by itself. It can chop vegetables and prepare everything, then stop and ask: “Ready to start?” Approval is the doorway between preparing an action and actually taking it.

The key graph transition is:

prepared_action ↓ pending_approval ├── approved → executable_action → execute ├── rejected → rejected → explain / stop └── expired → expired → re-evaluate or cancel

This is stronger than asking the human to approve a vague statement such as “The agent wants to continue.” A useful approval record identifies the specific action being approved.

Approval field Why it matters
ActorIdentifies who is making the decision.
ActionStates exactly what operation is being considered.
ScopeLimits which record, environment, amount, recipient, or resource is affected.
EvidenceShows the relevant facts or checks used to prepare the action.
ValidityPrevents a stale approval from being reused indefinitely.
DecisionApproved, rejected, expired, or otherwise resolved.

The deeper idea is that approval should attach to an action instance, not simply to an agent or an entire run.

For example, “Finance Agent is trusted” is too broad. A much clearer boundary is “Finance Agent may prepare a reimbursement draft, but release of the payment requires approval from an authorized reviewer.”

✅ Worked example

In the fictional expense workflow, the agent extracts a receipt, checks policy, calculates the reimbursement, and prepares a payment request. The graph then pauses. The reviewer sees the employee, amount, supporting evidence, policy result, and intended destination. Only after the reviewer approves that particular request does the execution step become reachable.

This pattern is reflected in current agent runtimes. For example, the OpenAI Agents SDK documents human-in-the-loop interruptions for tool calls that need approval and supports serializing paused execution state so a run can later be resumed. Its documentation also describes approvals that can arise after a handoff or inside nested agent execution. This is an example of one runtime architecture, not a universal graph standard. OpenAI Agents SDK human-in-the-loop documentation.

🛡️ Safety Check

Risk: approval fatigue causes reviewers to approve everything without inspecting the action.

Control: reserve approval for consequential actions, present concise decision-relevant evidence, and make the approval boundary specific to the proposed operation.

Remaining risk: a human can still approve a bad action. Approval is therefore one control, not proof that the workflow is safe.

🎯 Use this when...

An operation changes important data, creates external commitments, crosses a business approval boundary, or otherwise has consequences that should not be delegated completely.

03
Step permissions: give each node only the authority it needs

Human approval answers “should this action happen?” Step permissions answer a different question: “is this step allowed to perform this action at all?”

🧒 Child-friendly analogy

Imagine four keys: one opens the classroom, one opens the storeroom, one starts the school bus, and one opens the principal's office. Giving one student all four keys because it is convenient is not good security. Give each person only the keys needed for their job.

This is the familiar security principle of least privilege. NIST defines least privilege in terms of restricting access privileges to the minimum necessary for assigned tasks. The same concept applies to processes and service identities, not only to humans. NIST least privilege glossary.

OWASP’s current guidance for agentic and LLM-based systems similarly identifies excessive functionality, excessive permissions, and excessive autonomy as major sources of excessive agency, and recommends minimizing tools, functionality, and permissions while enforcing authorization in downstream systems. OWASP LLM06:2025 Excessive Agency.

In a graph, least privilege becomes easier to reason about because each node has a defined responsibility.

Graph step Allowed capabilities Explicitly unavailable Reason
Intake Read submitted receipt Payment APIs Classification does not require financial authority.
Policy check Read policy and expense data Write financial records The node evaluates a case; it does not change the system of record.
Prepare Create a draft transaction Release payment Preparation and execution remain separate.
Release Execute the approved transaction Create unrelated transactions Authority is limited to the approved action.

Notice the last row carefully. “Release Agent can release payments” is still broader than ideal. A stronger design makes the execution permission specific to the transaction being released.

A useful conceptual model is:

permission = identity + graph_step + operation + resource_scope + approval_state + validity_window

That expression is a design pattern, not a standard API. Its purpose is to force the engineer to think about authorization as a combination of factors rather than a single broad role.

🛡️ Safety Check

Risk: one shared agent identity can read, modify, approve, and execute across the entire workflow.

Control: separate capabilities by step and enforce the restrictions at the tool and downstream authorization layers.

Remaining risk: even a least-privileged tool can be misused within its permitted scope, so input validation, resource scoping, rate limits, and monitoring still matter.

What about tool metadata? Some ecosystems provide metadata describing whether a tool is read-only, destructive, idempotent, or open-world. MCP, for example, calls its current tool annotations hints, not guarantees. Its maintainers explicitly note that clients should not treat annotations from untrusted servers as enforcement controls. That distinction is important: metadata can help a runtime make decisions, but a hint should not be the only thing preventing a dangerous operation. MCP: Tool Annotations as Risk Vocabulary.

🎯 Use this when...

You are deciding which tools each graph node should see, especially when one tool can modify external systems or access sensitive data.

04
Secure agent handoffs

A handoff occurs when one agent transfers responsibility for part of a workflow to another agent. This can be useful when different agents specialize in different tasks, but it introduces a new security question: what exactly crosses the boundary?

🧒 Child-friendly analogy

Imagine passing a school file from the receptionist to the teacher. You do not need to pass the receptionist's entire desk, every key, and every other student's paperwork. You pass the information the teacher actually needs.

A secure handoff therefore has at least four design questions:

  1. Who is the receiving agent?
  2. What minimum information does it need?
  3. What authority does it gain, if any?
  4. What must remain inaccessible?

A useful distinction is between data handoff and authority handoff. Passing a case summary to a specialist does not have to mean passing the ability to execute every tool available to the previous agent.

Likewise, passing conversation history is not equivalent to passing application state. Sensitive details embedded deep in prior tool results may accidentally cross the boundary if the runtime forwards the full transcript.

Current OpenAI Agents SDK documentation illustrates this concern. Its documented handoff mechanism can pass conversation history to the receiving agent, while input filters can change what is forwarded. The documentation also warns that filtering structured tool items does not automatically guarantee that sensitive content has disappeared from ordinary messages or summaries. That is an implementation-specific example of a broader design rule: treat the receiving agent as a new trust boundary and explicitly control what crosses it. OpenAI Agents SDK handoffs documentation.

✅ Worked example

The fictional Intake Agent extracts an expense claim and identifies the relevant employee and receipt. It hands the Policy Agent a compact structured object instead of the entire conversation:

{ "case_id": "<CASE_ID>", "employee_id": "<EMPLOYEE_ID>", "expense_type": "travel", "claimed_amount": "<AMOUNT>", "receipt_reference": "<RECEIPT_REF>" }

The Policy Agent can evaluate reimbursement rules, but it never receives payment-release credentials.

The handoff payload can also carry a reason or priority chosen by the model, while application-controlled identity and authorization data remain outside the model-generated payload. This distinction prevents the model from casually declaring itself authorized.

model-generated handoff data ↓ schema validation ↓ application-controlled authorization ↓ minimal receiving context ↓ receiving agent
🛡️ Safety Check

Risk: a handoff transfers sensitive history or broader authority than the next agent needs.

Control: use an explicit handoff contract containing a small, typed payload and independently determine the receiving step's permissions.

Remaining risk: sensitive information can still leak through summaries, error messages, logs, or tool outputs unless those channels are treated as part of the same trust boundary.

🎯 Use this when...

Different agents have different responsibilities, data needs, or authority levels, and you want the boundaries to be explicit rather than relying on a shared conversation.

05
State, identity, and trust boundaries

Approval and handoff designs often fail because state is treated as an incidental implementation detail.

Once a workflow pauses for a human, the system must remember what was being approved. Once a handoff occurs, the system must know which step is active. After a restart, it must not accidentally resume the wrong branch.

🧒 Child-friendly analogy

Think of a parcel that pauses overnight. The delivery system needs to remember which parcel it is, where it came from, where it is going, and what approvals belong to it. “Continue” is not enough information.

For a paused graph, the state should conceptually contain enough information to reconstruct the pending transition without trusting the browser, mobile client, or model to invent the missing details.

workflow_state ├── run_id ├── current_node ├── case_id ├── pending_action ├── action_scope ├── required_approver_role ├── approval_status ├── state_version └── expiry

An especially useful concept is an action fingerprint: a deterministic representation of the specific action being approved. It can be used to detect whether the action changed between approval and execution.

For example, conceptually:

approval_target = hash( case_id + action_type + target_resource + amount + recipient + state_version )

This is an engineering pattern rather than a universal protocol. Its purpose is simple: the system should detect that “the thing the human approved” is the same thing “the executor is now trying to do.”

Identity also needs care. The person pressing an approval button is not automatically the person authorized to approve the action. The system should evaluate the reviewer's authenticated identity and role using normal authorization mechanisms.

Likewise, the model-generated claim “the manager approved this” is not evidence of approval. An approval should be recorded by the application's trusted authorization path.

🛡️ Safety Check

Risk: a stale or tampered workflow state causes the system to execute an action different from the one that was reviewed.

Control: protect persisted workflow state, bind approval to the intended action and scope, validate state versioning, and re-authorize before execution.

Remaining risk: application bugs can still create inconsistent state, so state-transition tests and audit trails are necessary.

06
Build the complete approval workflow

Now combine the three ideas into one fictional workflow.

Scenario: an employee submits a travel expense. An AI workflow reads the receipt, checks company policy, prepares a reimbursement request, and sends it for approval. Payment release occurs only after the appropriate human decision.

The graph can be designed as:

START ↓ Intake Agent ↓ Document Validation ↓ Policy Agent ↓ Finance Preparation ↓ Risk Classification ├── low-risk + policy-compliant → optional straight-through path └── consequential action → Human Approval ├── reject → END / correction ├── expire → return to review └── approve ↓ Release Tool ↓ Verification ↓ COMPLETE

The phrase “optional straight-through path” matters. Not every architecture needs a human review for every operation. Overusing human gates can make an automated system so slow that people start bypassing it.

The policy should therefore distinguish between operations according to their impact, not according to whether an LLM happened to produce them.

Action Automation Graph rule
Read receipt Automatic Read-only capability for the intake step.
Check policy Automatic Policy Agent can evaluate but cannot release funds.
Create draft Automatic Draft-only permission; no final execution.
Approve payment Human decision Authenticated reviewer and specific action scope.
Release payment Conditional Execution tool checks authorization and approval state again.

The last row demonstrates an important principle: approval does not replace authorization. The release tool should independently verify that the execution is permitted. This is defense in depth.

In a generic implementation, the execution path may look like this:

function execute(action, workflow_state, reviewer): validate_authenticated_reviewer(reviewer) validate_current_step(workflow_state, "release") validate_permission(reviewer, action) validate_approval(workflow_state, action) validate_scope(action) execute_downstream_operation(action) record_audit_event(action, reviewer)

Illustrative pseudocode: this is an original conceptual example, not a vendor SDK implementation.

✅ Worked example: what the reviewer should see

Request: reimburse <AMOUNT> for <EMPLOYEE_ID>

Evidence: receipt extracted successfully; required fields present.

Policy result: eligible according to the workflow's configured rules.

Destination: prevalidated payment target.

Impact: external financial transaction.

Decision: approve / reject.

🛡️ Safety Check

Risk: the model changes the amount, beneficiary, or target after approval.

Control: bind approval to a specific action representation and revalidate that representation immediately before execution.

Remaining risk: a downstream system may change independently, so the workflow also needs execution responses, reconciliation, and monitoring.

🎯 Use this when...

You have a workflow where some steps can be automated safely but one or more actions cross a meaningful authorization or business-control boundary.

07
Failure, cancellation, retries, and recovery

A graph is not secure merely because the happy path is secure. The difficult cases often happen when the workflow is interrupted.

Consider these events:

  1. The human rejects the action.
  2. The approval expires.
  3. The worker crashes after approval but before execution.
  4. The downstream system times out after receiving the request.
  5. A retry causes the same operation to be submitted twice.
  6. The underlying record changes while approval is pending.
🧒 Child-friendly analogy

Imagine a teacher approves a bus trip at 10:00, but the trip is moved to another date at 15:00. The old “approved” sticker should not automatically authorize the new trip. The thing being approved has changed.

Approval should therefore expire or become invalid when important facts change.

Likewise, retry behavior needs to be designed for the actual downstream operation. If the action is not safely repeatable, the graph needs a mechanism for detecting whether the first attempt succeeded before trying again.

Failure Graph response Security concern
Approval rejected Move to rejection/correction path. Do not silently retry until approved.
Approval expired Recompute whether approval is still needed. Prevents stale authorization.
Worker restart Restore durable workflow state. Do not trust client-supplied resume state.
Timeout after execution Reconcile before retrying. Avoid duplicate external actions.
Underlying data changed Invalidate approval and re-evaluate. Approved state no longer matches reality.

Cancellation is another often-missed graph feature. If a user cancels a pending task, the cancellation must prevent the execution branch from becoming reachable merely because an old worker later resumes.

The state machine should therefore distinguish between pending, approved, rejected, cancelled, expired, and executed rather than collapsing everything into a Boolean such as approved=true.

PENDING → APPROVED → EXECUTING → EXECUTED PENDING → REJECTED PENDING → EXPIRED PENDING → CANCELLED APPROVED → EXPIRED → CANCELLED
🛡️ Safety Check

Risk: a retry or delayed worker turns an old decision into a new action.

Control: make transitions explicit, make execution conditional on current state, and design idempotency or reconciliation around downstream effects.

Remaining risk: distributed failures can create ambiguous outcomes, so observability and reconciliation remain necessary.

08
How to test a secure agent graph

Workflow evaluation should test the graph, not just the final answer.

A model can produce a perfectly worded answer while the workflow still has a broken permission boundary, a missing approval transition, or an unsafe retry.

A practical test suite should therefore include at least six categories.

Test area Question Example assertion
Routing Does the case enter the correct path? High-impact transaction reaches review.
Permissions Can the wrong step call the tool? Policy step cannot invoke release operation.
Handoff Does only intended data cross? Receiving agent gets case fields, not unrelated secrets.
Approval Does execution stop until the decision exists? No approved state means no release call.
Recovery What happens after restart or timeout? Workflow resumes from durable state without duplicate execution.
Auditability Can you reconstruct what happened? Trace links request, decision, identity, tool call, and outcome.

Also test deliberately hostile or confusing situations:

  1. An untrusted document contains instructions telling the agent to bypass policy.
  2. A tool call requests a resource outside the permitted scope.
  3. The model tries to call a tool belonging to another workflow step.
  4. A handoff contains fields the receiving schema does not recognize.
  5. The reviewer approves an action and then the action changes before execution.
  6. A resume request references a workflow instance that the requester does not own.

These are graph-security tests because they verify transitions, authority, and state rather than merely textual output quality.

🧒 Child-friendly analogy

Testing the final answer is like checking whether a cake looks nice. Testing the graph is like checking whether the oven door locks, the timer works, the right person has the key, and the cake does not go into the wrong oven.

A strong evaluation dataset should therefore include both normal cases and boundary cases. The boundary cases often reveal the real weaknesses.

🛡️ Safety Check

Risk: testing only successful end-to-end cases hides authorization and recovery defects.

Control: test every important state transition, denied action, malformed handoff, expired approval, retry, and cancellation path.

Remaining risk: production traffic contains combinations not represented in tests, so post-deployment tracing and incident feedback must feed back into evaluation.

09
Enterprise Rollout

Introducing approvals and step permissions into an enterprise workflow is partly a security problem and partly an operating-model problem.

The most important question is not “which agent framework should we choose?” It is “who owns the authority represented by each graph node?”

A practical rollout can begin with a narrow, high-value workflow.

  1. Inventory actions. List every tool operation the workflow can perform.
  2. Classify impact. Separate read, draft, reversible write, irreversible write, external communication, and privileged operations.
  3. Assign ownership. Identify the business and technical owner for each consequential action.
  4. Define node permissions. Give each graph step the smallest useful capability set.
  5. Define approval rules. Specify which actions need review, by whom, and under what conditions.
  6. Protect state. Store workflow state and approvals in application-controlled storage.
  7. Instrument traces. Capture graph transitions, tool decisions, approvals, failures, and downstream results.
  8. Run boundary tests. Test rejected actions, stale approvals, handoff leakage, retries, and unauthorized resume attempts.
  9. Review regularly. Permissions and approval policies should change as the workflow changes.

NIST's least-privilege guidance emphasizes reviewing privileges and removing or reassigning them as necessary. For agent workflows, that translates naturally into periodic review of graph-node permissions, tools, downstream identities, and approval roles. NIST least privilege guidance.

Logging deserves special care. A useful trace records what happened without turning logs into a second uncontrolled data store. Tool arguments can contain personal, financial, or confidential data, so logging should be deliberate rather than “log everything.”

✅ Worked example: audit trail
09:14 case-created 09:15 receipt-validated 09:15 policy-evaluated 09:16 draft-prepared 09:16 approval-requested 11:42 approval-recorded 11:42 execution-authorized 11:43 downstream-transaction-submitted 11:43 execution-confirmed

The actual production record would use the organization's approved identifiers, retention policy, access controls, and privacy rules.

🛡️ Safety Check

Risk: operational logs become a new source of sensitive-data exposure.

Control: define logging fields intentionally, restrict access, minimize sensitive payloads, and apply retention rules appropriate to the data.

Remaining risk: metadata itself can reveal business-sensitive behavior, so observability needs its own threat model.

🎯 Use this when...

An agent workflow is moving toward production, especially when it accesses enterprise systems, sensitive data, or actions with material business consequences.

10
Common Mistakes

The following mistakes appear repeatedly when teams move from a simple agent to a graph that can take consequential actions.

1. “The prompt says the agent is not allowed to do that.”

Cause: confusing instructions with enforcement. Consequence: the model remains the final decision-maker. Correction: enforce the restriction at the tool, gateway, and downstream authorization layers.

2. One identity owns every tool.

Cause: operational convenience. Consequence: one compromised step has access far beyond its task. Correction: split permissions by workflow responsibility and resource scope.

3. The entire conversation is handed to every agent.

Cause: full-history forwarding is easy. Consequence: unnecessary sensitive information crosses trust boundaries. Correction: define explicit handoff payloads and filter context intentionally.

4. Approval means “the run was approved once.”

Cause: approval treated as a Boolean. Consequence: a later, different action inherits the earlier decision. Correction: tie approval to action, scope, state, identity, and validity.

5. Approval gates everything.

Cause: treating human review as the universal safety mechanism. Consequence: people become the bottleneck and may approve without reading. Correction: automate low-impact steps and reserve review for actions where human judgment adds meaningful control.

6. Tool metadata is mistaken for enforcement.

Cause: treating a descriptive hint as a security boundary. Consequence: the system assumes a tool behaves safely because it says it does. Correction: use metadata to inform policy and user experience, but keep deterministic authorization in trusted enforcement layers.

7. A worker can resume any workflow with a client-supplied state blob.

Cause: confusing serialized state with authenticated authority. Consequence: unauthorized users may attempt to continue someone else's run. Correction: keep authoritative state in controlled storage and authenticate every resume request.

8. Security review stops at the agent.

Cause: focusing on the LLM while ignoring downstream services. Consequence: a tool may still have broader database, API, or infrastructure permissions than required. Correction: enforce least privilege and authorization in the systems that actually execute the action.

🛡️ Safety Check

Risk: relying on one “smart” security control.

Control: combine bounded permissions, explicit state transitions, targeted human review, downstream authorization, validation, logging, and recovery.

Remaining risk: defense in depth reduces blast radius but cannot guarantee that an agent workflow will never make a harmful decision.

🎯 Use this when...

You already have an agent workflow and want to perform a practical threat-and-reliability review before giving it broader access.

11
❓ FAQ

Q1. Does every AI agent workflow need human approval?

No. Approval is most useful when an action has enough impact, uncertainty, or business significance that automated execution should pause. Low-impact read operations may not need a human gate, while consequential writes or external commitments may warrant one. The threshold should come from the application's risk and authorization model.

Q2. Why are step permissions necessary if a human approves the final action?

Because approval and authorization solve different problems. Approval asks whether a particular action should happen. Step permissions determine whether the component attempting that action has the authority to do it. Keeping those controls separate reduces the chance that a broad agent identity can bypass a workflow boundary.

Q3. Is passing the full conversation to the next agent safer because it preserves context?

Not necessarily. Full-history forwarding can preserve useful information, but it can also cross trust boundaries unnecessarily. A safer design starts with the minimum information the receiving agent needs and explicitly filters or transforms the context. Sensitive tool outputs, summaries, and error messages also need to be considered because removing only structured tool records may not remove copied content elsewhere.

Q4. Can a prompt or system instruction replace authorization?

No. Instructions can shape model behavior, but they should not be treated as the enforcement boundary for consequential operations. Authorization should be checked by trusted application and downstream systems so that an operation can be rejected even when the model proposes it.

Q5. What should happen when the approved action changes before execution?

The approval should normally become invalid for the changed action. The workflow should detect the mismatch, return to the appropriate review state, and request a new decision when required. Treating approval as permanently attached to a workflow run is weaker than binding it to the action, scope, identity, and current state.

12
🔗 References & Further Reading

Vendor and project names belong to their respective owners.

13
📝 Summary

• Human approval creates a deliberate pause before consequential execution.

• Step permissions restrict each workflow node to the authority required for its job.

• Secure handoffs transfer the minimum useful context rather than blindly forwarding every piece of history.

• Durable state allows a workflow to stop, recover, and continue without turning old decisions into new authority.

• Downstream authorization remains important because the model and graph should not be the only enforcement point.

• Graph evaluation should test transitions, permissions, handoffs, approvals, retries, cancellations, and recovery — not only the quality of the final response.

The central lesson is simple: an agent can make a recommendation, a graph can decide where work goes next, a human can approve a consequential action, and a trusted execution layer can enforce whether that action is actually permitted. Good graph engineering makes those boundaries explicit instead of hiding them inside a prompt or a single all-powerful agent.


Comments