Skip to main content

Graph Engineering for AI Agents: How Workflow Concepts Map to LangGraph and Other Frameworks

Calculating read time…

Graph engineering is the practice of making an AI workflow explicit: what work happens, what depends on what, where a decision changes the route, where parallel work can happen, where state is preserved, where a human can intervene, and what conditions allow the workflow to finish.


Imagine a fictional enterprise invoice-review workflow. An invoice arrives, important fields are extracted, deterministic validation checks run, two independent reviews happen, the results are joined, and the workflow either prepares the invoice for payment or asks a human reviewer to approve an exception. That is not merely “an agent with a good prompt.” It is a workflow with dependencies, routes, state, controls, and completion rules.

Frameworks such as LangGraph, OpenAI Agents SDK, Google ADK, and CrewAI can express parts of this problem, but they do not describe the same abstraction in the same way. LangGraph presents a graph-oriented runtime with explicit state, nodes, and edges. OpenAI Agents SDK centers on agents, tools, handoffs, guardrails, sessions, and a managed agent loop. Google ADK now exposes graph-based and dynamic workflow styles alongside higher-level workflow agents. CrewAI Flows use event-driven methods, listeners, routers, and shared flow state.

Engineering principle

Do not ask, “Which framework has a node called X?” Ask, “Where does this architecture represent the dependency, route, state, authority, and recovery rule?”

Graph idea LangGraph OpenAI Agents SDK Google ADK CrewAI
Work unit Node Agent, tool call, handoff, or application-controlled step Workflow node, agent, function, tool, or human-input node Flow method, listener, router, Crew, or task invoked by a Flow
State Explicit graph state with schemas and reducers Sessions and run/context mechanisms; conversation state is a major use Session state plus workflow/event data Flow state shared by Flow methods
Branch Conditional edges or Command Handoffs or agents-as-tools; routing can be model-driven Router nodes and explicit route edges @router() and listeners
Parallelism Fan-out through graph edges and state updates Usually composed through application orchestration when explicit parallel dependency graphs are required Parallel branches and explicit joins Multiple starts/listeners and conjunctions such as and_()
Human checkpoint interrupt() with resume Tool approval interruptions and resumable run state Human-input workflow nodes and application-level controls Human-feedback Flow mechanisms

This table is a conceptual mapping, not a claim that the frameworks implement identical runtimes or semantics. A framework's abstraction boundary matters.

01
The Mental Model: What a Workflow Graph Actually Controls

A useful beginner mental model is a railway network. The model may decide what it wants to say or which tool it would like to call, but the graph defines which stations exist, which tracks connect them, which junctions are permitted, where multiple tracks can operate independently, and where the train must stop for a human decision.

🟧 Child-friendly analogy

Think of a school trip. One teacher decides which activity to do, but the timetable says when the bus leaves, which class goes where, when the groups meet again, and when an adult must approve the next step. The teacher's judgment is not the same thing as the timetable.

Technically, a workflow graph is a representation of execution relationships. A node performs some work. An edge represents a permitted transition. State carries information needed by later steps. A route is a decision that selects one or more next steps. A join waits for multiple dependencies. A cycle deliberately sends execution back to an earlier point.

This is why graph engineering is different from simply writing an agent prompt. A prompt can influence model behavior, but it does not by itself create a durable dependency structure, a reliable join rule, an auditable state transition, or a deterministic authorization boundary.

Do not conflate these layers

Model: generates or selects outputs using the supplied input.

Agent: combines a model with instructions, tools, state/context, and a runtime loop.

Workflow graph: defines relationships among tasks and execution routes.

Harness: surrounds agent execution with runtime controls such as limits, approvals, tool policies, retries, and observability.

Sandbox: when present, isolates code, filesystem, network, or workspace execution according to its configuration.

Safety Check — Risk → Control → Remaining risk

Risk: The model is treated as if it were the workflow controller and therefore receives permissions it does not actually need.

Control: Keep routing, authorization, and irreversible actions explicit in application or workflow logic where practical.

Remaining risk: A model can still influence an allowed action through malformed or adversarial input, so tool validation and least privilege remain necessary.

🎯 Use this when...

A task has more than one meaningful execution step, especially when order, branching, parallel work, human decisions, recovery, or auditability matters.

02
Core Graph Concepts: Nodes, Edges, State, Routes, Joins and Loops

Before learning a framework, learn the language of the problem. Once these concepts are clear, framework syntax becomes much easier to recognize.

Node — “Do this piece of work.”

A node is a unit of execution. It might contain deterministic code, a model call, a tool call, a validation step, a retrieval operation, a human checkpoint, or another composed workflow. The important question is not whether a node contains an LLM. The question is what responsibility that execution unit owns.

Edge — “After this, what can run?”

An edge connects execution units. A fixed edge means the next step is known. A conditional edge means the next step depends on state, an event, a classification, a policy result, or another decision.

State — “What does the workflow know right now?”

State is more than chat history. A workflow may need an approval status, extracted invoice number, risk score, retry count, branch result, external correlation ID, or completion flag. The exact state model depends on the framework and architecture.

Branch — “Which route should continue?”

Suppose a validation step produces PASS, REVIEW, or REJECT. A branch turns that result into an execution route.

Fan-out — “These independent jobs can happen separately.”

After invoice extraction, a compliance check and duplicate-invoice check may not depend on one another. The workflow can start both. That is fan-out.

Join — “Wait until these dependencies are ready.”

If the next decision requires both checks, the workflow needs a join condition. A join is not the same as simply putting two tasks next to each other. It represents a dependency relationship.

Cycle — “Repeat this path under controlled conditions.”

A cycle is useful for bounded retry, iterative review, or correction. A production workflow should define why it loops, how many times it may loop, and what happens when the exit condition is not reached.

Start ↓ Extract invoice ↓ Validate fields ├── invalid ─────→ Reject / repair path │ └── valid ↓ ┌───┴───────────┐ ↓ ↓ Duplicate check Policy check ↓ ↓ └───────┬───────┘ ↓ Join ↓ Risk decision ├── low → prepare payment └── high → human approval ├── approve → execute └── reject → stop

Illustrative workflow notation — this is a teaching example, not vendor-specific syntax.

The key insight is that these concepts are architecture primitives, not brand-specific vocabulary. Different frameworks expose them through different APIs, decorators, classes, or managed runtime behavior.

Engineering principle

Draw the workflow in plain language before choosing framework APIs. The framework should encode a design; it should not become the design.

03
Quick Comparison: How the Concepts Map Across Frameworks

The most useful comparison is not “which framework is most powerful?” It is “where does each framework put the responsibility?”

Concept What to look for Typical mapping Important distinction
Node A unit that performs work Graph node, Flow method, workflow node, agent/tool step An “agent” is not automatically equivalent to a graph node.
Route A decision that changes execution Conditional edge, router, handoff, listener selection Some routes are model-selected; others are deterministic.
State Data needed across steps Graph state, session state, Flow state, run context Conversation history is only one kind of state.
Join A dependency on multiple completed paths Graph fan-in, JoinNode, conjunction, application orchestration A sequence is not automatically a join.
Approval Pause, decision, resume Interruptions, approval flow, human-feedback step Approval is a control point, not merely another model message.
Worked example: translating one concept

Suppose your design says: “After the risk assessment, if risk is high, stop and request approval.”

In LangGraph, this can be modeled with an explicit state-driven route and an interrupt. In OpenAI Agents SDK, a sensitive tool can be configured to require approval and the run can pause for review. In Google ADK, a workflow graph can route execution to a human-input step. In CrewAI, a Flow can use its human-feedback mechanisms and route subsequent listeners based on the result.

The design intent is the same; the runtime semantics are not.

🎯 Use this when...

You are evaluating frameworks, migrating an existing workflow, or explaining an architecture to a team that uses different agent stacks.

04
One Running Example: Invoice Exception Review

Let us make the concepts concrete with one explicitly fictional scenario. The business wants to process supplier invoices, but exceptions must be reviewed before payment.

🟧 Child-friendly analogy

Imagine a school office receiving a permission slip. One person reads it, another checks whether the information is complete, two people independently check different rules, and a teacher decides when something unusual happens. Each person has a job; the process tells them when that job starts.

Step 1 — Intake

The application accepts an invoice reference and the document payload. This step should establish identifiers and validate basic request shape before expensive processing begins.

Step 2 — Extraction

A document-processing component extracts invoice number, supplier, amount, currency, and purchase-order reference. The workflow stores the structured result as state.

Step 3 — Deterministic validation

Required fields, numeric formats, and known identifiers are validated using ordinary code. This is a useful boundary: the workflow does not need a model to answer a question that ordinary validation can answer exactly.

Step 4 — Fan-out

A duplicate check and a policy check start independently. The graph now represents concurrency rather than pretending the checks are a linear dependency.

Step 5 — Join

The decision stage should not run until the required checks have produced their results.

Step 6 — Route

Low-risk invoices can proceed to a preparation step. High-risk or ambiguous cases move to a human review route.

Step 7 — Approval

The workflow pauses and records the information necessary for a reviewer to understand the pending decision. The approval result becomes part of the workflow state.

Step 8 — Final action and completion

Only after the required conditions are satisfied does the workflow invoke the consequential business action. An audit record should capture what happened and why.

Safety Check — Risk → Control → Remaining risk

Risk: A model-generated interpretation is allowed to trigger payment directly.

Control: Separate recommendation from execution. Put authorization, amount limits, and final business validation at the tool or application boundary.

Remaining risk: Incorrect extracted data or manipulated external content may still influence the recommendation, so downstream validation remains important.

🎯 Use this when...

You need one scenario to explain graph concepts to developers, architects, security teams, and business stakeholders without starting with framework syntax.

05
LangGraph: The Graph-First Mental Model

LangGraph is particularly useful for learning graph engineering because its current documentation describes the workflow directly in terms of state, nodes, and edges. Its documentation also emphasizes durable execution, human-in-the-loop behavior, persistence, streaming, and the ability to combine deterministic and agentic steps.

The basic mental model is straightforward:

state = information shared by the workflow node = work performed against that state edge = rule describing what executes next graph = the compiled workflow

The current Graph API documentation describes StateGraph as the main graph class. Nodes can update state, while edges determine the next node. A graph must be compiled before execution, and compilation is also where runtime concerns such as checkpointers can be configured.

Worked example: a tiny LangGraph skeleton
from typing import TypedDict from langgraph.graph import StateGraph, START, END class ReviewState(TypedDict): invoice_id: str status: str def validate(state: ReviewState): # Illustrative logic only. return {"status": "valid"} builder = StateGraph(ReviewState) builder.add_node("validate", validate) builder.add_edge(START, "validate") builder.add_edge("validate", END) graph = builder.compile()

Illustrative teaching code. It demonstrates the shape of the graph, not a complete invoice-processing application.

The important lesson is not the number of lines. It is the separation of responsibilities. The node owns the validation work. The edge says validation leads to completion. In a larger graph, the route and state model become explicit parts of the design.

Conditional routing

LangGraph supports conditional edges and also provides Command patterns for routing and state updates. That makes it possible to keep a route decision near the step that produces the decision while still making the path explicit.

Illustrative pseudocode — not a copy of vendor documentation
def decide(state): if state["risk"] == "high": return "human_review" return "prepare_payment"

The graph engineer's responsibility is now to decide whether the route should be model-driven, rule-driven, or a hybrid. For consequential enterprise actions, deterministic policy gates often deserve a stronger boundary than a free-form model choice.

Human-in-the-loop

LangGraph's interrupt mechanism can pause execution and expose information for a human decision. A resume value can then continue the graph. The framework documentation also highlights an important engineering detail: nodes may be re-executed when an interrupt resumes, so side effects before an interrupt should be safe to repeat or moved to an appropriate boundary.

Safety Check — Risk → Control → Remaining risk

Risk: A node performs a non-idempotent external action before an approval interrupt and is re-run when the workflow resumes.

Control: Place consequential side effects after the approval point, or make pre-interrupt operations idempotent.

Remaining risk: The business operation may still fail after approval, so the workflow also needs error handling and reconciliation.

Why this matters for graph engineering

LangGraph makes a useful conceptual boundary visible: the graph controls orchestration while individual nodes can contain either deterministic logic or agentic behavior. This is exactly the distinction a graph engineer needs when deciding which parts of a workflow should be flexible and which parts should remain explicit.

Primary sources: LangGraph overview, Graph API, Interrupts, and Persistence.

06
OpenAI Agents SDK, Google ADK and CrewAI: Different Ways to Express the Same Ideas

This is where many beginners become confused. They learn a graph vocabulary and then expect every framework to expose a Node, Edge, and Join class. That expectation is too narrow.

OpenAI Agents SDK — agent-oriented orchestration

The OpenAI Agents SDK centers on a small set of agent-oriented primitives: agents, tools, handoffs or agents-as-tools, guardrails, sessions, human-in-the-loop mechanisms, and tracing. Its Runner manages an agent run, including turns, tool execution and handoffs.

Worked example: graph concept → Agents SDK concept

A graph designer might say: “Triage → choose specialist → specialist completes the user-facing task.”

In the Agents SDK, that design maps naturally to a triage agent with handoffs. When the specialist should help but the manager should retain control, agents-as-tools is the corresponding pattern.

That is a conceptual translation, not a claim that handoffs are literally graph edges. A handoff changes the active agent within the SDK's run model; the framework's own runtime defines the semantics.

Human approval in the Agents SDK

The current SDK documents tool-level approval flows in which a sensitive tool can require approval, execution can pause, the interruption can be resolved, and the run can continue. This is a good example of a workflow concept being exposed as an approval mechanism rather than a generic graph edge.

Conceptual mapping
Sensitive action ↓ Tool call requires approval ↓ Run pauses ↓ Human decision ├── approve → continue └── reject → stop / recover

For explicit fan-out and fan-in dependency graphs, do not assume that an agent-oriented SDK automatically becomes a workflow engine. The architectural choice may be to compose the SDK inside application orchestration or a dedicated workflow runtime.

Google ADK — graph-based workflows plus dynamic workflows

Current Google ADK documentation explicitly describes graph-based agent workflows as execution graphs composed from nodes and edges. It also documents dynamic workflows for control flow that is easier to express with ordinary program logic, and it positions older template workflow agents such as sequential, parallel, and loop agents as higher-level patterns that have been superseded by newer workflow structures in ADK 2.x.

The graph route model is especially useful for explaining sequence, branch, fan-out, and join. A route-producing node can select the next path, while a join node can wait for parallel work before continuing.

Worked example: explicit branch
classify ↓ route ┌─┴──────────────┐ ↓ ↓ standard_case exception_case ↓ human review

This is close to the graph vocabulary used in LangGraph, but the actual APIs, node model, event model, and runtime constraints are ADK-specific.

CrewAI — event-driven Flow orchestration

CrewAI Flows describe workflows through entry points, listeners, routers, and shared Flow state. A method marked with @start() can begin a path. A method using @listen() responds to another method's output. A router can select different downstream paths.

@start() def inspect_invoice(self): ... @listen(inspect_invoice) def calculate_risk(self, result): ... @router(calculate_risk) def choose_path(self): ...

Illustrative syntax shape. Refer to current CrewAI documentation before turning examples into application code.

CrewAI also provides shared Flow state and human-feedback mechanisms. An important conceptual difference is that you can think of the workflow as an event-driven sequence of methods and listeners rather than starting from a graph builder API.

Safety Check — Risk → Control → Remaining risk

Risk: A routing method is given broad external side effects because it is “only part of the flow.”

Control: Keep routing functions focused on selecting the route; isolate consequential operations behind narrowly scoped tools or service methods.

Remaining risk: An incorrect route can still lead to a valid but unintended downstream action, so route tests and authorization checks should exist independently.

🎯 Use this when...

You need to understand why “agent orchestration” is broader than “graph syntax,” or you are comparing framework architectures without reducing them to marketing labels.

07
How to Translate a Workflow Design Between Frameworks

Framework migration becomes much easier when you translate semantics first and syntax second.

Step 1 — Write the workflow as verbs

For example: receive → extract → validate → check duplicate and policy → join → assess → approve or reject → execute → record.

Step 2 — Mark deterministic versus agentic work

Circle the steps that genuinely require model reasoning. Everything else should be considered for ordinary code first.

Step 3 — Define the state contract

List the values that downstream steps need. Avoid turning state into a dumping ground for every message, tool result, and temporary object.

Step 4 — Define route predicates

Write the actual condition for each route. “The model feels this is risky” is not equivalent to “risk score exceeds an approved threshold.”

Step 5 — Mark irreversible actions

Sending money, deleting data, publishing content, changing production configuration, or modifying customer records should be treated differently from generating a draft.

Step 6 — Define recovery before implementation

For every external action, answer: what happens if the action times out? What happens if the process restarts? Can the step run twice? How do we identify the original request?

Step 7 — Then choose the framework expression

Only now should the design be encoded using graph nodes and edges, agent handoffs, workflow routes, Flow listeners, or another framework abstraction.

Architecture intent ↓ Responsibilities ↓ State contract ↓ Dependencies ↓ Routes ↓ Approvals ↓ Recovery rules ↓ Framework implementation
Engineering principle

A workflow is portable at the semantic level before it is portable at the API level.

08
State, Recovery, Human Decisions and Tool Boundaries

This is where graph engineering becomes production engineering. A graph that works on a developer laptop can still fail in production if its state, retries, approvals, and tools are poorly bounded.

State is an engineering contract

A useful state field should have a reason to exist. Ask who writes it, who reads it, what values are legal, how long it should live, and whether it contains sensitive information.

State should not silently become authority

A state field such as approved=true is not automatically proof that the correct person approved the action. Production systems should associate consequential decisions with authenticated identities, authorization checks, timestamps, and audit information as appropriate.

Recovery must respect side effects

Retries are easy to discuss and surprisingly hard to get right. A retry of a pure computation is different from a retry of an external mutation. A graph that retries a payment request without an idempotency strategy can create a second payment attempt.

Worked example: making an external action safer

Instead of treating “execute payment” as a generic callable action, define an operation identity such as payment_request_id. The downstream service can use that identity to distinguish a repeat of the same logical request from a new request.

This is a general distributed-systems design principle, not a special feature of any particular agent framework.

Human approval is part of the graph

A human checkpoint is not simply “send a message to a person.” The workflow should know what decision is pending, what data the reviewer sees, what action the approval authorizes, when the decision expires, and what happens after rejection.

pending_review ↓ show evidence + requested action ↓ human decision ┌──┴─────────────┐ ↓ ↓ approved rejected ↓ ↓ execute close / repair
Safety Check — Risk → Control → Remaining risk

Risk: Approval becomes a checkbox exercise because reviewers see incomplete evidence.

Control: Present the exact proposed action, relevant evidence, identity of the requesting workflow, and meaningful consequences.

Remaining risk: Human review can still be mistaken or rushed. Approval reduces a class of risk; it does not make the process infallible.

Tool boundaries matter as much as graph edges

A graph can route correctly and still be unsafe if every node shares broad credentials. Give each tool only the permissions required for its responsibility. Validate tool arguments before execution. Do not allow free-form model output to become an authorization token merely because it happens to contain words such as “approved.”

🎯 Use this when...

A workflow can restart, wait for a human, touch external systems, or perform an action that cannot safely be repeated.

09
Testing the Graph Instead of Only Testing the Model

A common mistake is to evaluate only the final answer. A workflow can produce a convincing final message while taking an unauthorized route internally.

Test the route

For every branch, provide inputs that should activate each route. Include boundary values, missing values, malformed values, and contradictory inputs.

Test the join

Verify that downstream work does not begin when a required dependency is missing. Also test partial failures: one parallel branch succeeds while another times out.

Test the loop

A loop should have a termination rule. Test both successful termination and forced termination after the maximum permitted number of attempts.

Test state recovery

Pause a run, restart the process, resume it, and verify that the workflow has not lost critical state or repeated an unsafe side effect.

Test approval boundaries

Verify that an unapproved action cannot reach the execution tool through another route, another agent, a retry path, or a fallback handler.

Test completion

Define what “done” means. A workflow should not be considered complete merely because the final model response sounds finished.

Workflow evaluation checklist [ ] Each route has positive and negative tests [ ] Join waits for all required dependencies [ ] Loops have explicit termination [ ] Retries are bounded [ ] Non-idempotent actions are protected [ ] Approval can be tested independently [ ] Unauthorized routes are rejected [ ] State survives the intended recovery scenario [ ] Completion has a machine-checkable condition

This is also where the framework choice affects testing. The current OpenAI Agents SDK documents deterministic testing utilities for orchestration-owned behavior. Other frameworks provide their own execution, tracing, evaluation, or visualization surfaces. The principle remains the same: test the workflow contract, not only the model's prose.

Safety Check — Risk → Control → Remaining risk

Risk: Tests cover successful paths but never attempt invalid or adversarial routes.

Control: Add negative-path tests, permission tests, restart tests, and malformed-input tests.

Remaining risk: Real-world distributions change. Periodic production monitoring and incident-driven regression tests remain necessary.

🎯 Use this when...

Your workflow has business consequences. The more expensive or irreversible the outcome, the less acceptable it is to rely only on end-to-end “looks good” tests.

10
Enterprise Rollout: Ownership, Security and Operations

Moving from a local graph to a production workflow changes the engineering questions.

Ownership

Every production graph needs an owner who understands business impact, dependencies, credentials, external services, and recovery procedures.

Change control

A route change can be as significant as a code change even when the model prompt barely changes. Review additions to tools, permissions, branches, approval rules, and completion criteria.

Secrets

Keep API keys, service credentials, certificates, and database secrets outside prompts and ordinary workflow state whenever practical. Use the platform's approved secret-management mechanisms.

Observability

A production trace should allow an operator to answer: which route executed, what state existed, which tools were called, where time was spent, what failed, and whether a human decision was required.

Data retention

Do not automatically retain every intermediate model output forever. Decide what evidence is needed for operations, compliance, debugging, and incident response, and set appropriate retention.

Budget and latency

Graph design changes cost. Fan-out can increase concurrency and total model/tool work. Loops can multiply it. Human pauses reduce real-time responsiveness. A route that looks elegant on paper can be expensive in production.

Incident response

Know how to stop or disable a dangerous workflow path. High-impact systems should have operational controls that do not depend on asking the model to behave differently.

Safety Check — Risk → Control → Remaining risk

Risk: A compromised agent can reach more tools than necessary.

Control: Apply least privilege per tool, validate inputs, separate read and write capabilities, and use approvals for high-impact actions.

Remaining risk: Instruction injection and unexpected model behavior cannot be eliminated by one control. The design should limit the damage a compromised component can cause.

Enterprise rollout sequence
  1. Document the workflow and ownership.
  2. Separate deterministic policy from model-driven reasoning.
  3. Define state and recovery contracts.
  4. Add route, permission, approval and failure-path tests.
  5. Instrument traces and operational metrics.
  6. Pilot with bounded permissions and reversible actions.
  7. Expand scope only after production evidence supports the next boundary.
🎯 Use this when...

The workflow will touch real customer data, enterprise systems, external services, production environments, or financially meaningful actions.

11
Common Mistakes

Mistake 1: Treating the framework API as the architecture

Cause: Starting with classes and decorators. Consequence: The resulting graph reflects framework syntax rather than business dependencies. Correction: Draw the workflow first, then encode it.

Mistake 2: Making the model responsible for every route

Cause: Assuming the model is the most flexible component, so it should decide everything. Consequence: Critical policy becomes difficult to test and audit. Correction: Use deterministic gates where the rule is deterministic.

Mistake 3: Calling conversation history “workflow state”

Cause: Assuming that because an agent remembers a conversation, it therefore remembers every operational fact needed for recovery. Consequence: Missing or ambiguous state during restarts. Correction: Define explicit operational state.

Mistake 4: Retrying irreversible actions blindly

Cause: Treating all retries as harmless. Consequence: Duplicate external operations. Correction: Use idempotency, operation identifiers, reconciliation, or a safer workflow boundary.

Mistake 5: Building parallel work without a join contract

Cause: “Run these two tasks at the same time” sounds complete. Consequence: The downstream step may use partial results or race conditions. Correction: Specify exactly which outputs are required before the next stage.

Mistake 6: Making approval decorative

Cause: Adding a human screen without changing authorization. Consequence: Approval exists visually but is not enforced in the execution path. Correction: Make the approval result a real prerequisite for the consequential action.

Mistake 7: Forgetting cancellation and expiry

Cause: Designing only the “approved” and “completed” path. Consequence: Workflows become stuck or remain actionable long after their business context has changed. Correction: Model timeout, cancellation, expiration, rejection, and operator intervention explicitly.

Mistake 8: Assuming one framework abstraction should be copied literally into another

Cause: Translating API names rather than responsibilities. Consequence: A visually similar workflow behaves differently. Correction: Translate node responsibility, state semantics, route ownership, recovery behavior, and authority boundaries—not class names.

12
❓ FAQ

What is the difference between an agent and a workflow graph?

An agent is a runtime construct that uses a model, instructions, tools, state or context, and an execution loop to work toward a task. A workflow graph describes how multiple execution steps and decisions relate to each other. An agent can live inside a graph node, but the two concepts are not interchangeable.

Is LangGraph the only way to build agent workflow graphs?

No. LangGraph is explicitly graph-oriented, but other systems express workflow concepts through different abstractions. Google ADK provides graph-based and dynamic workflows, CrewAI provides event-driven Flows, and the OpenAI Agents SDK provides agent-oriented orchestration through agents, tools, handoffs, guardrails, sessions, and run control.

When should a workflow use a deterministic node instead of an agent?

Use deterministic logic when the rule is naturally deterministic, especially for validation, authorization, formatting, threshold checks, and state transitions that should be predictable. Use an agent where the task genuinely benefits from model-based interpretation or adaptive reasoning. Many useful production workflows combine both.

How should I think about human approval in an agent graph?

Treat approval as an execution checkpoint. Define what action is pending, what evidence the human sees, who is allowed to decide, how the decision is recorded, what happens on approval or rejection, and how the workflow resumes or expires.

What should I test in a workflow graph that I would not necessarily test in a normal chatbot?

Test routes, joins, loops, state recovery, retries, permissions, approval enforcement, cancellation, external side effects, and completion conditions. A chatbot can fail mainly at the answer level; a workflow can fail while still producing a convincing answer because an internal route or external action was wrong.

Publishing note: The FAQ JSON-LD above is included because this article specification requests matching machine-readable FAQ metadata. Do not interpret it as a promise of a Google FAQ rich result; Google deprecated the FAQ rich-result feature for ordinary Search starting May 7, 2026.

13
🔗 References & Further Reading

Framework names, product names, APIs, and trademarks belong to their respective owners. 

14
📝 Summary

  • A workflow graph makes execution relationships explicit. Nodes perform work; edges define transitions; state carries operational information.
  • LangGraph is graph-first. Its state/node/edge model maps directly to graph engineering concepts.
  • OpenAI Agents SDK is agent-oriented. Handoffs, agents-as-tools, sessions, guardrails, approvals, and runner behavior provide orchestration without making the whole architecture a graph API.
  • Google ADK explicitly supports graph-based workflows. It also provides dynamic workflows and higher-level workflow abstractions, giving engineers multiple ways to represent control flow.
  • CrewAI Flows are event-driven. Starts, listeners, routers, and shared state express workflow relationships without requiring a graph-builder mental model.
  • Production graph engineering is about more than routing. State recovery, idempotency, permissions, approvals, observability, cancellation, and completion criteria matter just as much.
  • Design semantics before syntax. Once the dependencies, routes, state, authority boundaries, and recovery rules are correct, framework implementation becomes a translation exercise.
Final takeaway

Good agent engineering asks, “What should the model decide?” Good graph engineering adds, “What must the system decide, what may happen next, what must wait, and what must never happen without the right authority?”


Comments