Skip to main content

Graph Engineering for AI Agents: Practical Guide to Missing Information, Replanning, State & Resume

Calculating read time…

An agent workflow rarely fails because the model simply "doesn't know something."

In production, work can stop because a required input is missing, a tool returns an unexpected result, an earlier assumption becomes invalid, a person has to approve an action, or the process itself is interrupted by a restart or infrastructure failure.

Graph engineering is the discipline of designing those situations explicitly. Instead of assuming that a workflow moves from beginning to end in one straight line, we model what happens when information is missing, a route changes, state must be preserved, or execution has to continue later.

🧒 Child-friendly analogy

Imagine a delivery driver following a map. The driver reaches a road that is closed. A good system does not throw away the entire trip. It records where it is, checks what changed, chooses another road, and continues. A workflow graph does something similar for digital work.

In this article, we will use one original teaching scenario throughout: an AI-assisted supplier-invoice exception workflow. The example is fictional and exists only to explain graph behavior. The goal is not to make one particular framework the answer; it is to understand the engineering patterns that can be implemented in different graph runtimes and applications.

We will look at four closely related problems:

  • What should the graph do when required information is missing?
  • How can it change its plan when reality differs from the original plan?
  • How should it represent and persist execution state?
  • How can it safely resume after a pause, approval, crash, or restart?
Engineering principle

A resilient graph does not assume the path will remain valid. It makes change, interruption, and recovery part of the design.

01
What Changes When an Agent Workflow Meets Reality?

A simple workflow is easy to imagine:

Receive invoice
        ↓
Extract information
        ↓
Match purchase order
        ↓
Check policy
        ↓
Approve
        ↓
Post result

The difficulty begins when one of those steps does not behave as expected.

Suppose the invoice contains a vendor name and amount but no purchase-order number. Or the purchase order exists but is already closed. Or the policy service is temporarily unavailable. Or a manager must approve a high-value exception. Or the worker process crashes immediately after a downstream action succeeds.

These are not all the same failure. A graph needs to distinguish at least four different kinds of situations:

Situation What happened? Typical graph response Why it matters
Missing information A required input is absent or ambiguous. Ask, retrieve, validate, or route to an exception path. Guessing can turn uncertainty into a business error.
Changed conditions A planned step is no longer valid. Re-evaluate the state and choose another route. The old plan may now be unsafe or impossible.
Execution failure A service, worker, model call, or tool fails. Retry, compensate, pause, or fail safely. A retry is not always safe for side effects.
Interruption A person must decide, or execution stops and restarts later. Persist a resumable state and continue from a defined boundary. The workflow must remember what has already happened.

The important idea is that the graph is not merely a list of tasks. It is a map of possible execution states and transitions.

🎯 Use this when...

A workflow can encounter uncertainty, waiting, changing external conditions, human approvals, retries, or process restarts.

02
Handling Missing Information Without Guessing

The first graph-engineering question is deceptively simple:

“Does the workflow have enough information to perform the next action safely?”

🧒 Child-friendly analogy

Imagine someone asks you to deliver a parcel but gives you the street name and forgets the house number. Making up a house number is not “being helpful.” The sensible action is to ask for the missing detail or use another trusted source to find it.

For an AI workflow, missing information should become an explicit graph condition rather than an invitation for the model to invent an answer.

Consider our fictional invoice workflow. The next step is to compare an invoice with its purchase order. The graph requires a reliable purchase-order identifier or another approved matching strategy.

✅ Worked example — fictional scenario

The invoice contains:

  • Supplier: North Ridge Components
  • Invoice total: ₹184,500
  • Invoice number: INV-2026-1842
  • Purchase-order number: missing

The graph should not silently invent a purchase-order number. Instead, it can try an approved lookup using supplier and invoice metadata. If that lookup produces multiple candidates or no candidate, the graph can enter a needs_information or manual_resolution state.

Notice the separation of responsibilities. The model may help identify that the purchase-order field appears absent. The graph decides that the workflow cannot safely proceed. A tool may search approved systems for the missing information. A human may provide it. The state records what happened.

A useful design is to make missing information a typed state rather than a vague sentence in the conversation.

Illustrative state shape — not a vendor SDK

{
"status": "needs_information",
"missing": [
{
"field": "purchase_order_id",
"reason": "No reliable match found"
}
],
"attempted_sources": [
"invoice_metadata",
"approved_po_lookup"
],
"next_allowed_actions": [
"ask_requester",
"manual_resolution"
]
}

This is more useful than storing only “the agent needs more information.” The graph can now make deterministic routing decisions from structured state.

There are three useful ways to resolve missing information:

  1. Retrieve it: call an approved source of truth.
  2. Ask for it: request information from the user or another authorized participant.
  3. Escalate it: route the work to a human or exception queue when automated resolution is inappropriate.

The graph should also distinguish unknown from not applicable. An empty value does not necessarily mean the same thing as a field that is genuinely irrelevant to this transaction.

🛡 Safety Check

Risk → The model fills a missing business field with a plausible value and the graph treats it as verified.

Control → Separate inferred values from verified values, validate required fields before consequential actions, and route unresolved ambiguity explicitly.

Remaining risk → A trusted source can itself contain incorrect or stale information. Verification rules still matter.

Engineering principle

“Missing” is a state. It should be represented, observed, and routed—not hidden inside model output.

🎯 Use this when...

The next step depends on facts that may be unavailable, ambiguous, incomplete, expired, or accessible only from another system.

03
Replanning: When the Original Route No Longer Works

Missing information and replanning are related, but they are not identical.

🧒 Child-friendly analogy

Missing information is like not knowing which road to take. Replanning is like knowing the road you chose is closed after you start driving. You need to look at the current situation again and select a different route.

A graph should replan when the current state makes the existing route invalid, inefficient, unauthorized, or impossible.

For example, our fictional workflow initially plans:

Invoice received
      ↓
Find purchase order
      ↓
Three-way match
      ↓
Policy check
      ↓
Approve exception
      ↓
Post

During matching, the purchase-order system reports that the order was closed two months ago and the invoice amount does not correspond to the remaining balance.

The graph now has new information. Continuing with the original approval route may be incorrect. A safer route might be:

✅ Worked example — fictional scenario

The graph changes from:

automated match → policy check → posting

to:

exception classification → request supporting document → human review → revised decision

The important point is that the graph does not merely “try harder.” It changes the workflow because the evidence changed.

This leads to an important distinction between retrying and replanning.

Mechanism Question Example
Retry “Is the same action worth trying again?” A temporary timeout occurs while reading a supplier profile.
Replan “Does the workflow need a different next action?” The purchase order is closed, so the normal matching route is no longer valid.
Resume “How do we continue from known durable progress?” A reviewer approves the exception after the system has been offline.

A practical replanning loop looks like this:

Illustrative graph logic

1. Read current state.
2. Validate the assumptions behind the current plan.
3. Record what changed.
4. Select a valid next route.
5. Enforce limits on retries and replanning.
6. Persist the new route and state.
7. Continue from the new graph position.

A critical engineering question is who is allowed to replan?

In some architectures, the model can propose a next step and the application or graph runtime validates and routes it. In another architecture, the graph can contain deterministic routing rules and call the model only for classification or bounded decision support. Neither boundary should be treated as universal. The security property comes from making authority explicit.

🛡 Safety Check

Risk → A model-generated plan causes the workflow to take an action outside the intended business or security boundary.

Control → Constrain the set of legal routes, validate authorization in application or downstream systems, and place human approval around high-impact actions where appropriate.

Remaining risk → A graph can still contain incorrect business rules or incomplete authorization logic. Dynamic planning does not eliminate those risks.

OWASP's current guidance on excessive agency identifies excessive functionality, excessive permissions, and excessive autonomy as important causes of damaging agent behavior. Its recommended mitigations include narrowing tools and permissions and enforcing authorization in downstream systems rather than trusting an agent to decide what it is allowed to do.

🎯 Use this when...

The workflow can discover new evidence that invalidates the original plan, especially when external systems, policies, approvals, or dependencies can change.

04
Tracking State: Knowing Where the Work Really Is

A graph cannot reliably recover if it does not know what happened before the interruption.

🧒 Child-friendly analogy

Think about building a LEGO model. If you stop halfway and come back tomorrow, you need more than the final picture on the box. You need to know which pieces were already attached and which step you were doing. Workflow state serves a similar purpose.

For graph engineering, “state” should be treated as a deliberate data model. It can include different categories of information, and those categories should not automatically be mixed together.

State category What it describes Invoice example
Execution state Where the graph is and which route is active. “Waiting for reviewer”
Business state Facts about the work item. “PO closed; invoice total ₹184,500”
Decision state Decisions and their provenance. “Manager approved exception at 14:05”
Control state Retry counts, deadlines, locks, approvals, or cancellation status. “Approval pending; one reminder sent”
Reference state Identifiers used to reconnect with external systems or stored artifacts. “invoice_id=..., review_request_id=...”

A strong state model answers questions such as:

  • What has already completed?
  • What is currently waiting?
  • What information is still missing?
  • Which route is valid now?
  • Which actions are still permitted?
  • What external side effects may already have happened?

That last question is especially important. A graph can know that it was executing a “post invoice” node, but that does not prove whether the downstream posting system committed the operation before the worker failed.

Illustrative state

{
"workflow_id": "WF-2026-00981",
"status": "awaiting_review",
"current_route": "exception_review",
"completed_steps": [
"extract_invoice",
"locate_po",
"classify_exception"
],
"pending_step": "manager_review",
"retry_counts": {
"supplier_lookup": 1
},
"external_refs": {
"invoice_id": "",
"review_request_id": ""
},
"side_effects": {
"invoice_posting": "not_started"
}
}

The exact schema will differ by system. The important design choice is that state is explicit enough to support routing and recovery.

05
Resuming Work After Pause, Failure, or Restart

Resume is where graph engineering becomes operational engineering.

🧒 Child-friendly analogy

Suppose you are writing a long homework assignment and the computer shuts down. A good autosave lets you continue near the point where you stopped. A bad system makes you start the entire assignment again—and you may accidentally repeat work you already completed.

A resumable graph needs a durable boundary. The implementation differs across runtimes, but the concept is consistent:

Execute step
    ↓
Record durable progress
    ↓
Commit or record relevant side-effect boundary
    ↓
Move to next graph state

The word boundary matters. Saving state somewhere in memory is not the same as establishing a durable recovery point.

Current LangGraph documentation describes persistence through checkpointers for graph-state snapshots and stores for application-defined data. Its documentation also describes interrupts that save graph state and allow execution to be resumed later using a stable thread identifier.

OpenAI's current Agents SDK documentation similarly exposes a serializable RunState for paused runs and approval flows, allowing an interrupted run to be persisted and resumed later.

Microsoft's current durable-agent documentation describes another architectural approach: a workflow runtime can checkpoint progress, persist sessions, wait for external events, and restore work after failures or scale events. This illustrates that the durability boundary can live in a workflow engine rather than inside the agent framework itself.

✅ Worked example — fictional scenario

Our invoice workflow pauses because a manager approval is required.

The graph stores a durable state such as:

  • Current route: exception_review
  • Pending action: manager_decision
  • Required reviewer identity or role: stored server-side
  • Evidence references: stored separately from transient model output

The application can stop running. Hours later, the reviewer responds. The runtime restores the stored state, verifies that the reviewer is authorized, applies the decision to the pending request, and resumes the graph.

That last sentence hides an important security rule: do not let the client manufacture the state that is resumed.

The current OpenAI Agents SDK human-in-the-loop guidance explicitly describes server-side authorization and validation around pending approvals and advises against accepting replacement tool calls, arguments, approval records, or serialized state directly from the client.

🛡 Safety Check

Risk → A user modifies a workflow identifier, approval payload, or serialized state and tricks the system into resuming a more privileged action.

Control → Keep authoritative workflow state server-side, authenticate the reviewer, authorize access to the specific run, validate pending request identifiers, and atomically consume approval requests before resuming.

Remaining risk → Authorization logic can still be implemented incorrectly, and stale approvals can become unsafe when business conditions change. Revalidate critical assumptions before consequential actions.

06
The Boundaries: Model, Agent, Graph, Harness, Tools, and State

One reason graph discussions become confusing is that several different components are often described as though they were one thing.

🧒 Child-friendly analogy

Think of a restaurant. The model is like a person suggesting what to order. The agent is the worker coordinating the task. The graph is the kitchen's sequence and routing board. The harness is the equipment and operating controls around the worker. Tools are things such as the payment terminal or inventory system. State is the order ticket and its status. A sandbox, when used, is a controlled room where certain work can happen without touching the rest of the restaurant.

Component Primary responsibility Example in this article
Model Produces language or other model outputs used by the application. Suggests which information appears missing.
Agent Combines a model with tools, instructions, context, and runtime behavior to perform a task. Investigates an invoice exception using approved capabilities.
Workflow graph Defines task relationships, routes, transitions, waits, joins, loops, and completion conditions. Routes missing information to retrieval, user input, or review.
Harness Runs and controls the agent, including runtime constraints, tool invocation, context handling, and other controls depending on the architecture. Enforces runtime rules and manages tool interactions.
Tool Performs an external operation or accesses an approved capability. Reads an invoice or checks purchase-order status.
State Records durable facts and execution information needed to route or recover work. Stores approval status, route, completed steps, and external references.

These boundaries can overlap in real implementations. A workflow graph can live inside an agent harness or inside a broader application. A workflow runtime can own persistence while an agent framework provides model and tool behavior. The architecture determines the exact boundary.

This distinction also helps prevent a common mistake: assuming that adding a graph automatically makes an agent reliable. The graph can improve routing and state control, but it does not magically fix incorrect model output, bad tools, weak authorization, or poor external data.

07
Production Patterns for Reliable Graphs

Once the basic concepts are clear, several engineering patterns become useful.

Pattern 1 — Treat uncertainty as data.

Do not hide uncertainty inside free-form model text. Record whether a value is verified, inferred, unavailable, conflicting, or awaiting confirmation when that distinction affects a decision.

Pattern 2 — Make legal routes explicit.

Even when a model proposes what to do next, the application should define what transitions are actually permitted. This changes the graph from “anything the model says” to “a bounded set of possible next states.”

Pattern 3 — Separate retry from recovery.

A transient network error may justify retrying a read operation. A payment, database update, email, or external mutation may require a stronger side-effect strategy. The graph should know which category a node belongs to.

Pattern 4 — Make consequential operations idempotent where possible.

If recovery can cause the same operation to be attempted again, the operation should ideally tolerate safe repetition or have a durable idempotency mechanism. Otherwise, the graph can resume while accidentally duplicating the business effect.

Pattern 5 — Record completion separately from intention.

“The agent intended to post the invoice” is not the same as “the invoice was posted successfully.” Completion should be based on an observable result or durable acknowledgement.

Pattern 6 — Bound replanning.

A graph that can replan forever has a loop without a business limit. Put boundaries around repeated planning, escalation, time, budget, and tool usage.

Pattern 7 — Give every wait a meaning.

“Waiting” should not be a vague condition. Distinguish waiting for a person, waiting for a service, waiting for a timer, waiting for a file, and waiting for a business event. Each can have different timeout and escalation behavior.

Pattern 8 — Make cancellation a real route.

A long-running workflow may become unnecessary. Users can withdraw a request, a deadline can expire, or an external object can be deleted. Cancellation should produce an explicit terminal or compensating state rather than leaving abandoned work indefinitely.

Illustrative lifecycle

RECEIVED
↓
VALIDATING
├── missing data ───────→ NEEDS_INFORMATION
│                              ↓
│                         WAITING_FOR_INPUT
│                              ↓
└──────────────────────── VALIDATING
↓
PLANNING
↓
EXECUTING
├── transient failure → RETRY
├── changed condition → REPLAN
├── approval needed → WAITING_FOR_APPROVAL
└── completed → VERIFY
↓
COMPLETED

Any active state may also move to:
CANCELLED
FAILED_REVIEW
MANUAL_ESCALATION

This is intentionally conceptual. Different graph runtimes express these transitions in different ways. The engineering idea is more important than the syntax.

🛡 Safety Check

Risk → Replanning becomes a route for progressively increasing privileges or repeatedly calling powerful tools.

Control → Define route-level permissions, tool-level permissions, retry ceilings, time limits, approval requirements, and downstream authorization.

Remaining risk → Defense in depth reduces blast radius but cannot guarantee that every model-generated decision or external input will be benign.

08
Enterprise Rollout

A graph that works in a notebook can behave very differently after it becomes a business process.

Start with a small state machine, not a giant autonomous loop. Identify the business states, allowed transitions, waiting points, high-impact operations, and terminal conditions before adding more agent behavior.

Define ownership. Someone should own workflow definitions, someone should own downstream integrations, and someone should be responsible for security controls and incident response. These responsibilities may belong to the same team in a small system, but they should still be explicit.

Log state transitions, not secrets. Useful operational records include workflow ID, state transition, route selected, tool identity, result category, retry count, approval status, timestamps, and error classification. Sensitive payloads should follow the application's data-retention policy and should not be copied into logs merely because they are convenient.

Design for observability. A production operator should be able to answer: “Where is this workflow?”, “Why is it waiting?”, “What failed?”, “What already happened?”, and “What is allowed to happen next?”

Test state transitions independently of model quality. Route logic, approval behavior, cancellation, retries, joins, timeouts, and recovery can often be tested with deterministic fixtures even when model behavior itself is variable.

Retain enough information to investigate incidents. NIST's AI Risk Management Framework materials emphasize testing, evaluation, verification, validation, and lifecycle-oriented risk management. A workflow graph should therefore be designed so that consequential behavior can be reconstructed from appropriate records rather than from memory or an operator's guess.

Choose persistence according to the workflow. A short request-response interaction may not need a sophisticated durable workflow engine. A process that can wait for people, timers, external events, or infrastructure recovery needs stronger persistence semantics. Current documentation from multiple agent and workflow platforms demonstrates this architectural distinction.

Engineering principle

Durability is a system property, not just a database feature. Recovery depends on how state, side effects, retries, and external identity fit together.

09
Common Mistakes

Mistake 1 — Treating every failure as a retry.

Cause: Retry is the easiest recovery mechanism to implement.
Consequence: Duplicate side effects or repeated expensive work.
Correction: Classify failures and decide which operations are safely repeatable.

Mistake 2 — Letting the model “fill in” business-critical gaps.

Cause: A conversational interface rewards fluent answers.
Consequence: A plausible value becomes an unverified business fact.
Correction: Represent missing information explicitly and require approved evidence before consequential actions.

Mistake 3 — Replanning without boundaries.

Cause: The system keeps searching for another route whenever a step fails.
Consequence: Endless loops, rising cost, or unexpected tool usage.
Correction: Limit replanning attempts, time, tools, budget, and authority.

Mistake 4 — Saving conversation history but not execution state.

Cause: Developers equate memory with recovery.
Consequence: The system remembers what was discussed but not which side effects actually occurred.
Correction: Store explicit execution state and external references alongside appropriate conversational context.

Mistake 5 — Resuming from a client-provided snapshot.

Cause: The browser or API client already has a convenient representation of the workflow.
Consequence: A client-controlled representation can become a privilege-escalation or integrity problem.
Correction: Keep authoritative state on the server and re-authorize the action when resuming.

Mistake 6 — No cancellation route.

Cause: The workflow was designed only for success and failure.
Consequence: Abandoned work continues consuming resources or remains stuck indefinitely.
Correction: Model cancellation, expiration, escalation, and cleanup explicitly.

Mistake 7 — Assuming an approval means the whole situation is still safe.

Cause: The approval was valid when requested.
Consequence: External state may change while the workflow waits.
Correction: Revalidate critical business and authorization conditions immediately before high-impact execution.

10
❓ FAQ

Q1. Is missing information the same as a failed tool call?

No. A failed tool call means the system attempted an operation but did not obtain the expected result. Missing information means the workflow does not currently have enough reliable input to continue safely. A graph may respond to the first with retry or alternate execution and to the second with retrieval, user input, or escalation.

Q2. Does replanning mean letting the model freely decide the entire workflow?

No. Replanning is a graph behavior, not a requirement for unrestricted model autonomy. A model can propose a route, a deterministic graph can choose among predefined branches, or the application can combine both approaches. The important property is that the resulting transition stays inside the system's allowed boundaries.

Q3. Is conversation memory enough to resume an interrupted agent?

Not necessarily. Conversation history tells you what was communicated. Resume logic may also need execution position, pending approvals, retry metadata, external identifiers, business state, and side-effect status. A durable graph should preserve the information required to continue safely, not merely the visible conversation.

Q4. How should a graph handle a human approval that arrives hours later?

Persist a durable workflow state that identifies the pending decision, authenticate and authorize the reviewer, validate that the decision belongs to the stored request, and resume from the server-owned state. Before a consequential action, recheck conditions that could have changed while the workflow was waiting.

Q5. When do I need a durable workflow engine instead of a simple agent loop?

A simple loop can be sufficient for short-lived work with limited state and low recovery requirements. A durable workflow approach becomes more valuable when work can span long waits, process restarts, human approvals, external events, retries, parallel branches, or side effects that must not be repeated accidentally. The right boundary depends on the application's reliability requirements.

11
🔗 References & Further Reading

The following primary sources were consulted to verify framework capabilities, durability patterns, human-in-the-loop behavior, and security guidance. The explanations, analogies, examples, diagrams-as-text, tables, and FAQ wording in this article are original.

Vendor and project names belong to their respective owners. The article does not imply endorsement, private access, or a claim that one architecture is universally preferable.

12
📝 Summary

  • Missing information should become an explicit state, not an invitation to guess.
  • Replanning means reevaluating the current situation when the original route is no longer valid.
  • Retrying and replanning solve different problems.
  • State should describe execution progress, business facts, decisions, controls, and important external references.
  • Resume requires a trustworthy durable boundary, not merely a copy of conversation history.
  • Side effects must be considered when designing retries and recovery.
  • Authorization should be enforced by the application and downstream systems, not assumed from model intent.
  • Bounded routes, cancellation, observability, and explicit completion rules make graph behavior easier to operate.
Final thought

A production agent does not become dependable merely because it can make a plan. Dependability comes from knowing when the plan is invalid, recording what has actually happened, controlling what may happen next, and having a safe way to continue when the world changes.

Thanks for reading. Keep experimenting, keep testing the failure paths, and treat the workflow graph as an engineering system—not just a diagram of successful steps.

Comments