Skip to main content

Graph Engineering for AI Agents: Design a Software-Engineering Workflow from Planning through Review and Repair

Calculating read time…

When an AI agent changes software, the difficult engineering problem is not simply getting the model to write code. 

The harder problem is deciding what happens before the change, what evidence must exist before the next step, which tools are allowed to act, what happens when tests fail, when a human must intervene, how state survives interruption, and how the system knows that the work is actually finished.

That is where graph engineering for AI agents becomes useful. Instead of treating an agent as one large loop that can freely decide what to do next, we can represent the work as a controlled workflow: planning, inspection, implementation, testing, review, repair, approval, and completion connected by explicit routes.

This article uses one fictional software-engineering scenario throughout: an AI-assisted system receives a change request, creates a structured implementation plan, inspects a repository, proposes a change, runs tests, reviews the result, repairs failures, and stops for human approval before a consequential action. The scenario is invented for teaching; it does not describe a private or real company deployment.

Concept Main responsibility Typical question
Model Generate, classify, summarize, reason, or propose actions “What should we do or produce next?”
Agent Use model capabilities together with tools and runtime behavior “How can the system act toward a goal?”
Workflow graph Define dependencies, routes, state transitions, loops and completion conditions “What may happen next, and under which condition?”
Harness Run the agent and enforce runtime controls around tools, state, limits and execution “How do we control the running system?”
Sandbox Isolate execution or workspace access when the architecture requires it “Where can this action safely execute?”

These boundaries are architectural rather than universal product definitions. For example, the current OpenAI Agents SDK describes agents together with tools, handoffs, guardrails, sessions and runtime behavior, while LangGraph describes workflows through state, nodes and edges. Those are useful concrete architectures, not a mandatory industry-wide decomposition.

01
What Graph Engineering Actually Controls

A workflow graph is easiest to understand as a map of responsibility and movement. A node does some unit of work. An edge determines where execution goes next. State carries the information needed for later decisions. A loop sends execution backward when recovery or iteration is required.

🍊 Child-friendly analogy

Imagine a school project with checkpoints. One student writes the first draft, another checks the facts, another checks the formatting, and a teacher approves the final submission. The important part is not only who does each task. It is the rule saying when one person is allowed to hand the work to the next person.

In an AI system, that rule can be represented explicitly. The model may suggest that a patch looks complete, but the graph can still require tests, static checks, review findings, or approval before the next consequential step.

This is one reason graph engineering is different from prompt engineering. Prompt wording can influence model behavior, but a graph gives the application an explicit execution structure. The graph can say that a “deploy” node has no incoming route until a policy check and approval state are both satisfied.

Engineering principle

The graph should make important state transitions visible enough to test and govern.

A graph can therefore encode more than a happy-path sequence. It can represent branches, parallel work, joins, retries, approval gates, cancellation, recovery and termination. Not every application needs every mechanism.

🎯 Use this when...

Your AI system performs more than one meaningful step and a mistake in one step can affect what happens later.

02
The Running Example: An AI Software Change Workflow

Fictional teaching scenario: a development team receives a request such as “update the invoice validation logic so zero-value lines are handled consistently.” The AI system is allowed to inspect a repository, propose a patch, run approved tests and prepare the change for review.

The system is intentionally not given one giant instruction saying “solve the issue from start to finish.” Instead, the work is divided into graph stages.

🟢 Worked example: the end-to-end route
START
  ↓
Intake → Plan → Inspect → Implement → Test → Review
                  ↓
                  Review decision
                  ├─ findings → Repair → Test
                  ├─ pass → Approval
                  └─ risk / retry limit → Human Review

Approval
├─ reject → STOP
└─ approve → Apply Change → Post-check → END

The important detail is the route policy. A review finding does not merely produce text. It changes the next allowed node. A failed test does not become “try again forever.” It routes into a bounded repair path.

The model can participate inside multiple nodes. For example, it can produce a plan in Plan, propose a change in Implement, explain failures in Repair, and summarize risk in Review. That does not mean the model owns the whole workflow.

🛡️ Safety Check

Risk → A model-generated plan could request a tool or change that exceeds the original task.

Control → Keep the allowed repository scope, tool inventory, change types and approval requirements outside the model's free-form output.

Remaining risk → The model can still misunderstand the task or produce an unsafe proposal, so independent validation and review remain necessary.

03
Design State Before Designing Nodes

Many workflow problems are actually state-design problems. Developers often begin by listing agent roles: planner, coder, tester, reviewer. A stronger starting point is asking: what facts must survive from one stage to the next?

🍊 Child-friendly analogy

Think of a parcel moving through a warehouse. The workers can change, but the parcel needs labels: where it came from, what it contains, which checks passed, and what happens next. Workflow state is the digital equivalent of those labels.

For the software-change scenario, a conceptual state might contain:

request
plan
repository_snapshot
allowed_scope
proposed_change
test_results
review_findings
repair_count
approval_status
risk_status
workflow_status

The exact representation depends on the framework. LangGraph, for example, explicitly models graph state and uses nodes that return state updates; its documentation also describes reducers that determine how updates to state fields are combined.

The engineering lesson is broader than that specific framework: state should represent facts that the workflow needs to make or audit a transition. Avoid storing everything simply because it might be useful someday.

Engineering principle

State is not a transcript. It is the durable information required to continue, evaluate and explain the workflow.

This distinction becomes important when the workflow grows. A giant conversation history can be expensive, difficult to validate, and ambiguous about which statement is authoritative. A small structured state can make routing far easier to test.

A practical rule is to separate at least three categories:

State category Examples Why it matters
Business/workflow state status, findings, approval, retry count Controls routing and completion
Artifact state plan, patch reference, test report Carries the work produced by earlier nodes
Runtime context connection handles, ephemeral caches, execution metadata May need different lifecycle and protection
🎯 Use this when...

You need reliable recovery, auditable transitions, human approval, bounded retries, or meaningful workflow-level evaluation.

04
Turn Planning into an Explicit Graph

Planning is often where an AI system first becomes difficult to reason about. A model can generate a useful plan, but the workflow should treat the plan as an artifact to validate, not as an automatic command stream.

🍊 Child-friendly analogy

Suppose a child says, “I will make a birthday cake.” That is a goal, not yet a safe procedure. A useful checklist might be: inspect ingredients, choose a recipe, prepare, bake, check, decorate. The checklist determines what should happen before the next step.

In software engineering, a plan could be a structured document containing intended files, rationale, dependencies, expected tests and risk markers. The graph can then validate whether the proposal fits the allowed scope before implementation begins.

🟢 Worked example: planning as data

Imagine the planner returns this conceptual object:

scope: "invoice-validation"
files_allowed: ["validator.py", "test_validator.py"]
change_type: "code-and-tests"
required_checks: ["unit-tests", "review"]
risk: "medium"
approval_required: true

The graph can validate the fields before permitting the implementation node to run.

This is an example of a general pattern: convert a model proposal into structured state, then route from the structure. That creates a clearer boundary between model generation and workflow policy.

🛡️ Safety Check

Risk → A generated plan may contain an unexpected file, tool, or external action.

Control → Validate scope, tool permissions and required checks before the implementation route is enabled.

Remaining risk → Validation rules themselves can be incomplete. Keep consequential operations behind independent authorization and review.

A planner can also fail by being too vague. A plan such as “update validation and fix tests” provides almost no useful state for the graph. A better artifact defines what the system believes needs to change, what evidence would prove the change, and what constraints apply.

🎯 Use this when...

The next stage depends on facts produced by the planner and you want those facts to be testable rather than hidden inside an unstructured conversation.

05
Separate Planning, Implementation, Testing and Review

Once the plan is accepted, the next temptation is to let one agent perform everything. That is simple to prototype but difficult to govern. Separating stages gives the graph meaningful checkpoints.

A useful decomposition is:

  1. Inspect: gather the approved files, interfaces, tests and constraints.
  2. Implement: produce a proposed change inside the approved scope.
  3. Test: execute the permitted test suite and record machine-readable results.
  4. Review: evaluate whether the change matches the request and the evidence.
  5. Decide: route to repair, approval, escalation or completion.

These nodes do not have to map one-to-one with different agents. A single agent can participate in several nodes. Likewise, separate specialist agents can participate in the same workflow. The graph is about the execution relationship, not the number of personas.

The OpenAI Agents SDK currently documents both agents-as-tools and handoffs as orchestration patterns. It also distinguishes orchestration decisions made by the LLM from orchestration decisions made by code. That distinction fits graph engineering well: a model can help choose among allowed actions, while application code can enforce the hard workflow boundaries.

# Illustrative pseudocode — not a vendor SDK example

if state.tests == "failed":
    next_node = "repair"
elif state.review == "findings":
    next_node = "repair"
elif state.risk == "high" and state.approval != "approved":
    next_node = "human_review"
else:
    next_node = "complete"

The value of this pattern is not the syntax. The value is that the route can be tested independently from the model.

🛡️ Safety Check

Risk → A reviewer may consume untrusted repository content or tool output that contains instruction-like text.

Control → Treat external content as data, not authority. Keep policy, permission checks and routing logic outside the content being reviewed.

Remaining risk → Prompt injection and indirect instruction manipulation cannot be eliminated by one filter. Limit what a compromised stage can access or change.

This is closely related to current OWASP guidance for agentic systems: the risk is not merely a wrong sentence from a model. It is what can happen when model output influences tools, other agents, external systems or consequential state.

🎯 Use this when...

You need independent evidence between stages, different permissions for different activities, or a clean boundary between “proposed” and “executed.”

06
Build Repair Loops Without Creating Chaos

Repair is where a simple workflow becomes a real graph. The system must be able to move backward without becoming an infinite loop.

🍊 Child-friendly analogy

Think about a teacher checking homework. If one answer is wrong, the student corrects that answer and returns the work. But after repeated failures, the teacher does not say “try forever.” There is a rule for when the problem needs a different kind of help.

A repair loop therefore needs at least four concepts:

  1. Failure evidence: what exactly failed?
  2. Repair scope: what is the agent allowed to change?
  3. Attempt budget: how many repair cycles are allowed?
  4. Escalation path: what happens when the budget is exhausted?
🟢 Worked example: bounded repair

Suppose the workflow allows at most three repair attempts for an individual change.

attempt 0 → test fails → repair
attempt 1 → test fails → repair
attempt 2 → test fails → repair
attempt 3 → stop automatic repair → human review

The counter is part of workflow state. That makes the boundary visible and testable.

Another useful distinction is between recoverable failure and workflow-invalidating failure. A temporary service timeout may be retryable. A discovery that the requested file is outside the allowed scope should not be “repaired” by trying another path. The graph should route those cases differently.

Failure Possible route Why
Transient tool timeout Bounded retry Likely operational rather than logical
Unit test failure Repair → test Evidence identifies a possible defect
Scope violation Stop or human review Retry does not fix authorization
Repeated contradictory review Escalate The workflow may have insufficient information or conflicting criteria

A repair loop is therefore not simply “call the model again.” It is a controlled state transition with evidence, bounds and an exit strategy.

🎯 Use this when...

Failures can be diagnosed and corrected automatically, but you need an explicit limit on repetition and an escalation route.

07
Human Approval and Trust Boundaries

A human approval step is not simply a button at the end of the workflow. It is a trust boundary: the system changes from autonomous preparation to an externally authorized decision.

The most useful approval designs show the reviewer the smallest meaningful decision package: what changed, why it changed, what evidence supports it, what failed previously, and what action approval would authorize.

🍊 Child-friendly analogy

Imagine an assistant preparing a form for you. Preparing the form is one job. Signing it is another. The assistant can fill in information, but the signature is the point where authority changes hands.

Current LangGraph documentation provides a concrete example of this architecture through interrupts: a graph can pause, persist the relevant state, wait for an external response, and then resume using the same thread identity. This is one implementation approach; other workflow systems can represent the same business idea differently.

Likewise, the OpenAI Agents SDK documents human-in-the-loop capabilities and handoff controls as part of its runtime. The important graph-engineering lesson is not the product API. It is that approval should be represented as a state transition with explicit authorization semantics.

🛡️ Safety Check

Risk → Users can approve without understanding what action will occur next.

Control → Make the approval payload specific: target, change summary, evidence, scope and next action.

Remaining risk → Human approval can become routine clicking. Approval does not automatically make an unsafe workflow safe.

A second trust boundary exists between content and authority. A repository file, issue description or retrieved document may contain instructions, but that content should not automatically redefine system policy. The application should distinguish “data being inspected” from “instructions the workflow is authorized to follow.”

That separation matters because prompt-injection and excessive-agency risks become more consequential when model outputs can trigger tools or downstream actions. Current OWASP agentic guidance treats excessive agency as a serious risk category precisely because unexpected or manipulated outputs can cause damaging actions when systems grant too much functionality, permission or autonomy.

🎯 Use this when...

The workflow can create an irreversible, high-impact, externally visible, privileged or financially meaningful action.

08
State, Recovery, Retries and Durable Execution

A workflow that only works while one process stays alive is not yet a production workflow. Real systems encounter worker crashes, network failures, manual waiting periods, deployment changes and long pauses.

🍊 Child-friendly analogy

Think of a board game where everyone leaves the room halfway through. A reliable game needs the board position recorded so the players can continue later. A workflow needs an equivalent record of where execution stopped and what state had already been established.

LangGraph's current persistence documentation describes checkpointers for thread-scoped graph state and stores for longer-lived application data. It explicitly connects checkpoints with continuity, human-in-the-loop execution, time travel and fault tolerance.

Temporal presents another architecture for durable workflows: workflow state can be reconstructed from recorded event history, while external interactions such as API calls, database queries, LLM invocations and file operations are placed in Activities rather than performed as nondeterministic workflow logic.

These designs are different, but they illustrate the same engineering concern: the system needs a reliable answer to “what had already happened?” and “what can safely happen next?”

🟢 Worked example: safe retry thinking

Suppose the workflow invokes a test runner. Retrying a failed test command may be acceptable. Retrying an external deployment operation blindly may not be. The graph should classify operations according to whether repeating them is safe, whether they are idempotent, and whether they require a unique operation identifier.

operation_id = "change--apply"

if already_completed(operation_id):
    return recorded_result()

result = perform_change()
record_completion(operation_id, result)

Illustrative pseudocode only. The exact idempotency mechanism must match the target system.

The deeper lesson is that retry policy belongs to the workflow design. “Retry on error” is not a sufficient production rule because different failure types have different side-effect behavior.

🛡️ Safety Check

Risk → A resumed or retried node repeats a side effect that had already succeeded.

Control → Persist completion evidence, use idempotency keys where the target supports them, and isolate external side effects behind well-defined operations.

Remaining risk → Exactly-once behavior is difficult across distributed systems. The workflow must still handle ambiguous outcomes.

A related concept is state retention. Storing every event forever can increase latency, storage cost and privacy exposure. State should therefore have an intentional lifecycle: what is necessary during execution, what is retained for audit, and what should eventually expire.

🎯 Use this when...

The workflow can pause, resume after human input, run for a long time, span multiple workers, or repeat actions after failures.

09
Evaluate the Workflow, Not Just the Final Answer

A final answer can look correct even when the workflow behaved incorrectly. Suppose the agent eventually produces a valid patch but skipped a required test, exceeded its tool scope, retried too many times, or ignored a human rejection. The output alone would hide those failures.

Graph evaluation therefore asks questions about paths and transitions, not only outputs.

Dimension Example test question
Routing Does a failed test always enter the intended repair route?
Dependencies Can implementation run before the required inspection state exists?
Loop bounds What happens after the repair budget is exhausted?
Permissions Can an inspection node invoke a write-capable tool?
Recovery Can the workflow resume after interruption without duplicating a side effect?
Completion Can the workflow declare success without required evidence?

A strong test suite should include negative paths. For example, deliberately provide a review finding, an invalid approval, a test failure, a missing dependency, an exceeded retry count and a tool permission denial. The expected result is not “the model gives a good answer.” The expected result is “the graph enters the correct state and route.”

🟢 Worked example: route-level evaluation

A useful test might start with:

state.review_findings = ["missing validation case"]
state.repair_count = 2

expected_next_node = "repair"

A second case can set the repair count to its maximum and assert that the next route is human review rather than another automatic repair.

This style of testing also makes it easier to separate workflow correctness from model quality. You can later evaluate whether the model's plans or reviews are good, but first verify that the graph obeys its own rules.

🎯 Use this when...

You need confidence that failure paths, permissions and recovery behavior work even when the model behaves unexpectedly.

10
Observability, Budgets, Cancellation and Completion

A production graph needs more than logs that say “agent started” and “agent finished.” Operators need to know which node ran, what route was selected, what evidence was produced, what tool was called, how long the node took, whether a retry occurred, and why the workflow stopped.

The exact telemetry model varies by runtime. OpenAI's current Agents SDK includes tracing for agent runs, while LangGraph documentation provides observability and tracing concepts around graph execution. These are implementation examples of a broader engineering requirement: make meaningful state transitions observable.

The same principle applies to budgets. A graph can track:

  1. maximum repair attempts;
  2. maximum wall-clock duration;
  3. maximum tool invocations;
  4. maximum external actions;
  5. maximum model or service cost allowed for the workflow.

These are not merely performance settings. They define the shape of failure.

🛡️ Safety Check

Risk → A workflow consumes unbounded tool calls, retries or model work while pursuing a difficult task.

Control → Enforce explicit budgets in the runtime and graph state, and record when a budget caused termination.

Remaining risk → Budgets limit damage and cost but do not guarantee useful progress. A poor workflow can still fail quickly.

Cancellation is another graph property that is easy to overlook. If a user withdraws a request, the workflow should have a clear cancellation state and a policy for in-flight operations. Cancellation may be easier for pure computation than for an external action already submitted to another system.

Completion should also be explicit. “The agent said it finished” is not a completion criterion. A better workflow can require conditions such as:

required_plan = true
required_tests = "passed"
required_review = "passed"
required_approval = true
required_postcheck = "passed"

A graph with explicit completion criteria is much easier to reason about than one where every node can simply return “done.”

🎯 Use this when...

The system must be operated over time, investigated after failures, or controlled under cost and execution limits.

11
Enterprise Rollout

Moving from a prototype graph to an enterprise workflow is mostly about making assumptions explicit.

Start with ownership. Every node should have a clear purpose and an owner responsible for its behavior. Every external tool should have a defined permission scope. Every state field that contains sensitive or business-critical information should have a retention decision.

Then introduce controlled change. A graph definition is application logic. Changes to routing rules, retry limits, approval thresholds or tool permissions can alter business behavior even when no model or prompt changes.

Area Enterprise question
Ownership Who owns each node, tool and routing policy?
Access Which node can access which data or tool?
Change control How are graph changes reviewed and promoted?
Auditability Can operators reconstruct why a route was taken?
Incident response Can a risky workflow be stopped or isolated quickly?
Retention How long should workflow state, traces and artifacts remain available?

A useful rollout strategy is to start with a read-mostly workflow. Let the system inspect, classify, test and draft changes before giving it permission to perform consequential actions. That makes the route structure easier to validate before the risk of external side effects increases.

Then introduce write actions one at a time, each with an explicit permission boundary and a clear completion signal. This approach is not a guarantee of safety, but it gives the organization smaller units to test and govern.

🛡️ Safety Check

Risk → A single broad credential allows many workflow nodes to access or modify too much.

Control → Separate permissions by node or operation where practical, and keep the least necessary authority at each step.

Remaining risk → Permission separation reduces blast radius but does not prevent a permitted action from being misused.

🎯 Use this when...

A prototype is moving toward shared infrastructure, production data, regulated processes or actions with real-world consequences.

12
Common Mistakes

The most common graph-engineering mistakes are not exotic. They usually come from making the first prototype too unconstrained and assuming the same design will remain reliable at scale.

Mistake Consequence Correction
One giant agent loop Hard to know why the system chose a route Introduce explicit workflow stages and state transitions
Routing from free-form text Small wording changes can alter workflow behavior Route from structured state where practical
Unbounded repair Runaway cost, latency or repeated side effects Use attempt budgets and escalation
Broad tool permissions Larger blast radius when an agent fails Apply least privilege and separate read/write capabilities
No durable state Restart loses progress or repeats work Persist the state required for recovery
Completion means “agent said done” The workflow can finish without evidence Define machine-checkable completion criteria
Approval with no context Human review becomes ritual instead of control Show the exact action, scope and evidence being approved

One subtle mistake is also worth highlighting: over-graphing. Not every automation needs a sophisticated graph. A simple deterministic sequence may be easier to test and operate. Graph engineering becomes valuable when there are meaningful branches, loops, parallel work, stateful decisions, human interruptions or complex recovery requirements.

Engineering principle

A graph should make real complexity explicit, not manufacture complexity where a simpler workflow already works.

13
❓ FAQ

Q1. Is a workflow graph the same thing as an AI agent?

No. An agent is an AI-enabled runtime entity that can use instructions, tools, model calls and other capabilities, while a workflow graph describes how tasks, state and transitions are organized. An agent can participate inside a graph without being the graph itself.

Q2. Should the model decide every next step in the workflow?

Not necessarily. Some routing decisions can be delegated to the model, while important constraints can be enforced by application code or the workflow runtime. A mixed design is often useful when the system needs both flexible reasoning and predictable control.

Q3. Why does workflow state matter so much?

State records the facts required to continue, route, recover and audit the workflow. Without explicit state, important decisions may remain hidden in conversation history or transient runtime memory, making recovery and evaluation harder.

Q4. When should an AI software-change workflow stop automatic repair?

It should stop when repair is no longer safely bounded or when the evidence suggests the problem cannot be resolved automatically. Typical triggers include an exceeded retry budget, scope violations, contradictory review results, missing information or a higher-risk action that requires human authorization.

Q5. Does putting an AI agent in a workflow graph make it safe?

No. A graph can make permissions, transitions, approvals, recovery and limits more explicit, but it cannot guarantee safe behavior. Model errors, prompt injection, tool failures, incorrect policies and implementation bugs can still occur, so the design needs defense in depth and continuous testing.

14
🔗 References & Further Reading

Vendor and project names belong to their respective owners.

📝 Summary

  • Graph engineering makes task dependencies, routes, state transitions and recovery behavior explicit.
  • The model is not the workflow. The model may propose actions while the graph and runtime define what can happen next.
  • Design state deliberately. Store the facts needed for routing, recovery, audit and completion.
  • Use bounded repair loops. Failures should lead to defined routes, not endless retries.
  • Make authority visible. Separate untrusted content from workflow policy and use approval gates where the risk warrants them.
  • Persist what must survive. Recovery depends on knowing what already happened and what should happen next.
  • Evaluate paths, not only answers. Test routing, permissions, retries, recovery and completion conditions.
  • Keep the graph as simple as the problem allows. More nodes do not automatically mean better engineering.
Final thought

When an AI system moves from “generate an answer” to “perform a sequence of actions,” the workflow itself becomes an engineering artifact. Designing that workflow explicitly is what turns an unpredictable agent loop into something a team can inspect, test, operate and improve.

Comments