Skip to main content

AI Agent Harness Engineering: State Management in an AI Agent Harness — Sessions, Context, and Resuming Work

Calculating read time…

State management in an AI agent harness is the engineering discipline of deciding what must survive a run, what the model should see on the next turn, what can be safely forgotten, and how an interrupted task can continue without repeating or corrupting work. 

A useful harness does not simply “remember everything.” It keeps different kinds of state at different lifetimes, scopes them to the right user and task, reconstructs a bounded context for the model, and verifies recovery before allowing more tool actions.

Imagine an agent reviewing an invoice. It has already read the purchase order, compared three line items, called an internal validation tool, and is waiting for a human approval before sending a correction request. The process stops because the worker restarts. When the agent comes back, the important question is not merely, “Does it remember the chat?” The real question is: Can the harness reconstruct the right state, prove what already happened, identify what still needs to happen, and safely continue?

Engineering principle

Continuity is not the same thing as memory. A transcript can preserve conversation history while still being a poor recovery record. A durable state record should tell the harness what stage the task reached, which external effects were confirmed, and what may safely happen next.

📑 In This Post

AT A GLANCE
Session, Context, State, and Checkpoint

Thing Main question Typical contents Primary job
Run state What is happening right now? Current step, pending tool call, transient data, interruption Drive the current execution safely
Session Which ongoing interaction is this? Session ID, conversation history, session metadata Connect multiple runs that belong together
Model context What should the model see now? Selected history, task facts, tool results, instructions, retrieved material Give the model the information needed for this decision
Durable business state What must remain true after the process ends? Workflow status, approvals, external IDs, artifact references Support reliable recovery and business correctness
Checkpoint Where can execution restart? Versioned snapshot, completed steps, pending work, verification markers Resume without guessing or duplicating effects

01
What State Actually Means in an Agent Harness

🧒 Child-friendly analogy

Think about a board game. The conversation is what the players said. The game state is whose turn it is, which pieces have moved, which rules have been unlocked, and whether someone is waiting for a dice roll. You cannot reliably resume the game just by remembering the conversation.

In an agent harness, state is the set of information that the runtime needs to understand the current execution and make a correct next decision. That can include conversation history, workflow status, tool results, approval status, external identifiers, task checkpoints, and references to artifacts. The exact fields depend on the architecture.

The model is not the durable owner of that state. The model produces a proposed next action from the context it receives. The harness decides how that context is assembled, whether a tool may run, whether a prior action is already complete, and whether the run is allowed to continue.

🟢 Worked example — fictional scenario

An internal invoice agent receives INV-1042. Its durable workflow state might say:

task_id: "TASK-1842" invoice_id: "<INVOICE_ID>" stage: "awaiting_approval" purchase_order_checked: true validation_check: "passed" correction_request: "not_sent" approval_required: true checkpoint_version: 7

Notice what is missing: a giant copy of every earlier message. The workflow state captures durable facts that matter for recovery, while the model context can be reconstructed from those facts plus selected history.

This leads to a useful rule: store facts and transitions explicitly; do not depend on the model to infer the entire workflow from a transcript.

🎯 Use this when...

A task can pause, span multiple worker runs, call tools, require approval, or create external side effects.

02
Session vs Context vs Durable State

🧒 Child-friendly analogy

A session is the folder for one conversation. Context is the handful of pages you put on the desk for today’s question. Durable state is the official record that says what actually happened. They are related, but they are not interchangeable.

Session. A session groups related work. Some frameworks persist conversation history behind a session identifier. For example, the current OpenAI Agents SDK supports session backends that retrieve history before a run and store new items after it; it can also resume an interrupted run using the same session storage. See the OpenAI Agents SDK session documentation.

Context. Context is the actual material presented to the model for a specific inference step. A session may contain far more material than should be sent every time. A harness can retrieve recent history, durable task facts, selected tool results, and other relevant information, then construct a bounded context.

Durable state. Durable state is the authoritative record of facts that must survive worker restarts, context changes, and model changes. It should not depend on the wording of a previous model response. A business workflow database, task store, event log, object store, or other persistence layer can serve this purpose depending on the architecture.

Checkpoint. A checkpoint is a recoverable execution position. It tells the harness enough to decide what to do next. It may be represented inside a workflow system, a serialized run state, or a combination of database records and external workflow metadata.

Architecture note

There is no universal requirement that one product must provide all four layers. OpenAI documents both client-managed sessions and server-managed conversation continuation. These are examples of different implementation boundaries, not a single industry standard.

🛡 Safety Check

Risk → The application treats a session ID as proof that the caller is authorized to view or modify that session.

Control → Authenticate the caller separately and authorize access to the session and its underlying state. In its session documentation, OpenAI explicitly notes that a session identifier by itself is not authentication or authorization. See the session trust-boundary notes.

Remaining risk → A correctly authorized caller can still receive stale, poisoned, over-retained, or cross-workflow state if the state model itself is poorly designed.

03
Choose State by Lifetime, Scope, and Authority

🧒 Child-friendly analogy

Keep your lunch order, your school timetable, and your passport in different places. You would not carry the passport into every lunch conversation, and you would not trust a sticky note to be the official passport record.

A practical state model starts with three questions:

  1. How long must it survive? One tool call, one run, one session, or an entire business process?
  2. Who may see it? One invocation, one user, one tenant, a team, or an operational service?
  3. Who is authoritative? The model, the application database, an external system, or a human approval record?
State class Example Lifetime Design concern
Ephemeral run state Current tool result, pending interruption Current execution Serialization and safe replay
Session state Conversation history and session metadata Conversation or work thread Isolation, retrieval, retention
Workflow state approval_status, stage, external reference Business process Authority and consistency
Artifact state Report, file, test output Until retention policy Integrity, ownership, expiry

A strong harness also records state transitions, not just final values. “Approval received” is stronger than a mutable boolean because the harness can attach a timestamp, actor identity, approval scope, and the exact action that the approval authorizes.

state_transition: from: "validated" to: "awaiting_approval" actor: "harness" reason: "correction_request_requires_human_approval" action_scope: "send_correction_request" version: 8

The exact schema will differ by system. The design idea is the important part: make authority-bearing facts explicit and versionable.

🎯 Use this when...

You need to decide what belongs in a database, what belongs in a transcript, and what can disappear after the run.

04
Rebuilding a Useful Context for the Model

🧒 Child-friendly analogy

Suppose you have a huge school library, but your teacher asks one question. You do not carry every book to the desk. You bring the few pages that help answer today’s question.

Context reconstruction is the harness deciding which stored information becomes the model’s next input. This is where context engineering touches harness engineering: the harness owns the retrieval and assembly mechanism, while the selected context determines what information is available to the model for the current decision.

A practical reconstruction pipeline can look like this:

  1. Load the authenticated session and workflow records.
  2. Check the latest checkpoint and its schema version.
  3. Fetch only the history and artifacts relevant to the next step.
  4. Remove or summarize material that no longer earns its space.
  5. Attach current tool observations rather than replaying stale guesses.
  6. Apply the current permissions and policy context.
  7. Build the final bounded model input.
🟢 Worked example — rebuilding TASK-1842

The harness discovers that TASK-1842 is waiting for approval. It does not need the entire raw conversation. It can construct a compact context containing:

  • Current task objective: correct a validated invoice mismatch.
  • Confirmed facts: purchase order matched; validation passed.
  • Current state: correction request not yet sent.
  • Approval record: approval required and still pending.
  • Recent tool evidence: last validation result and timestamp.
  • Relevant user instruction: continue the review, but do not send the correction before approval.
context = { "goal": "resolve_invoice_mismatch", "stage": "awaiting_approval", "verified_facts": ["po_match", "validation_passed"], "pending_effect": "send_correction_request", "approval": "pending", "latest_tool_observation": "<VALIDATION_RESULT>" }

This is intentionally illustrative pseudocode, not a vendor SDK. The important pattern is to separate authoritative facts from model-generated narration.

For long-running work, context reduction becomes important. OpenAI documents compaction mechanisms that reduce context size while preserving state needed for later turns; its guidance distinguishes server-side compaction from explicit compaction, and notes that compacted output should be treated as machine state rather than edited like a human-written summary. See OpenAI's compaction guidance. Anthropic’s current guidance similarly discusses multi-context-window workflows and recommends maintaining structured task state and incremental progress rather than relying only on a growing conversation. See Anthropic's long-horizon state guidance.

Engineering principle

A summary is a derived view, not automatically the source of truth. When summarization changes what the model sees, keep critical workflow facts outside the summary and reconstruct them from authoritative state.

🛡 Safety Check

Risk → A stale or attacker-influenced summary becomes the next turn’s “truth.”

Control → Keep critical facts separately typed, versioned, and revalidated at recovery boundaries. Treat model-generated summaries and retrieved external text as untrusted inputs unless verified.

Remaining risk → A malicious or incorrect fact can still enter the durable state through a legitimate workflow path, so write operations and state transitions require their own validation and authorization.

05
How Pause, Resume, Retry, and Recovery Really Work

🧒 Child-friendly analogy

A bookmark only tells you the page number. A good save file also records which quests are complete, which items you already collected, and whether you were standing before or after opening a locked door.

Safe resumption is a reconciliation problem. The harness must compare its checkpoint with the outside world before continuing. This matters because a crash can happen after a tool succeeded but before the harness recorded success.

Consider this sequence:

1. Harness asks tool: create correction request. 2. Tool creates request CR-771. 3. Network connection drops. 4. Harness never receives the confirmation. 5. Worker restarts. 6. Checkpoint says: correction request = unknown. 7. Harness must reconcile before calling "create" again.

The safest recovery is not “ask the model what probably happened.” It is to query an authoritative status source, use an idempotency key when the external system supports it, or route the ambiguous case to a controlled recovery path.

🟢 Worked example — recovery decision
  1. Load checkpoint version 8.
  2. Verify the session belongs to the authenticated user and correct tenant.
  3. Inspect the external correction system using the request's idempotency or correlation identifier.
  4. If CR-771 exists, record the confirmed external ID and continue from the next safe step.
  5. If it does not exist, retry only within the allowed policy and retry budget.
  6. If the external system cannot prove either outcome, stop and surface the ambiguity for review.

OpenAI's current Agents SDK documentation provides a concrete example of durable run state for human-in-the-loop interruptions: an interrupted run can be serialized into run state and resumed with the same session backend. The documentation also describes safeguards around exclusive session access and reconciliation after interrupted work, illustrating why “resume” is more than replaying the last message. See Run State.

A recovery routine should therefore answer four questions before allowing another consequential tool call:

  1. Identity: Is this the correct user, session, tenant, and workflow?
  2. Position: What is the last verified checkpoint?
  3. Effects: Which external actions are confirmed, unknown, or incomplete?
  4. Authority: Does the current run still have permission to perform the next action?
🎯 Use this when...

Your agent can be interrupted during tool calls, approvals, deployments, payments, record updates, or any other externally visible operation.

06
Security: State Is Part of the Attack Surface

🧒 Child-friendly analogy

Imagine a notebook that your robot reads every morning. If anyone can quietly write “the teacher already approved this,” the robot may act on a lie the next day. A notebook that influences future actions is part of the control system, not ordinary scratch paper.

Persistent memory and state can increase usefulness while creating another path for bad or incorrect information to influence later runs. OWASP's 2026 guidance explicitly treats memory and context poisoning as a distinct agentic security concern and recommends treating retained context as security-relevant state. See OWASP's memory and context poisoning discussion.

For harness engineers, that means state should be protected with the same seriousness used for other control-plane data. The main questions are:

  • Scope: Can one user or tenant read another session?
  • Integrity: Can stored state be modified without authorization?
  • Provenance: Do you know whether a fact came from a user, a trusted system, a tool result, or model inference?
  • Retention: Are you keeping sensitive state longer than the workflow needs it?
  • Secrets: Are credentials, tokens, or raw personal data being persisted into conversation history or summaries?
🛡 Safety Check

Risk → Persistent state carries an instruction-like value from an untrusted document or earlier tool result into a later run.

Control → Separate data from authority: store provenance, validate state-changing fields, constrain which state can affect permissions, and require policy checks at the point of consequential action. Never treat persisted text as an authority grant.

Remaining risk → Context poisoning can still shape model reasoning, and legitimate state may become stale. No single filter, prompt, or storage mechanism eliminates this class of risk.

A particularly important boundary is the difference between state that helps the model reason and state that grants authority. A note saying “customer approved the change” should not itself be sufficient to authorize an API call. The approval record should be verified by the harness from an authoritative source.

The same principle applies to secrets. Keep credentials in a dedicated secret mechanism rather than conversation history or general-purpose agent memory. When the agent needs a credential, the harness should mediate access rather than teaching the model the secret.

Retention is another engineering decision. Some state can expire quickly; other records may be required for audit or business continuity. The right retention period is determined by the application, legal requirements, and risk model—not by the convenience of the agent.

07
Testing State Management and Recovery

🧒 Child-friendly analogy

Do not test only whether the robot can finish when nothing goes wrong. Turn the lights off halfway through the game. Then ask whether it can reopen the same game without moving the pieces twice.

Harness evaluation should include both happy-path continuity and failure-path recovery. A session feature can appear to work while failing exactly when a production system needs it most.

Minimum state-management test matrix:

Test Injected condition Expected property
Resume after restart Terminate worker after checkpoint Resume from durable state without losing verified progress
Unknown tool outcome Drop response after external write Reconcile before retrying; no unintended duplicate effect
Session isolation Attempt to load another user's session Authorization failure before context assembly
State schema migration Load an older checkpoint Deterministic upgrade or controlled rejection
Context compaction Replace large history with compact state Critical task facts and recovery behavior remain correct

For every recovery test, assert more than the final text. Inspect the state transition, tool-call count, external side effects, authorization decision, and trace. A fluent answer can hide a duplicate API call.

assert checkpoint.version == expected_version assert session.tenant_id == authenticated_tenant assert correction_request.external_status in {"confirmed", "not_created"} assert tool_effect_count <= allowed_effects assert resume_trace.contains("recovery_reconciliation")

This is illustrative pseudocode. In production, the actual assertions should inspect your workflow store, tool adapter telemetry, and authoritative external systems.

🛡 Safety Check

Risk → Evaluation checks only that the agent eventually produced a correct-looking response.

Control → Test side effects, state transitions, permissions, recovery, and trace evidence—not only natural-language output.

Remaining risk → Test suites cover known failure classes. Production can still encounter novel races, provider outages, malformed state, or unexpected tool behavior.

🎯 Use this when...

Your team is preparing an agent for long-running tasks, multi-worker execution, human approval, or business-critical tool calls.

08
Enterprise Rollout

For an enterprise harness, state management becomes an ownership problem as much as a coding problem. Decide who owns the session lifecycle, workflow records, retention policy, recovery procedures, and approval boundaries.

A practical rollout sequence:

  1. Define the state contract. Document fields, provenance, allowed transitions, versions, and retention.
  2. Pick the persistence boundary. Decide which state belongs in session storage, workflow storage, artifact storage, or external systems of record.
  3. Instrument recovery. Log checkpoint creation, restoration, state-version changes, tool reconciliation, approvals, and recovery failures.
  4. Gate consequential actions. Enforce authorization and policy at tool execution time, even when earlier context suggests approval.
  5. Exercise incidents. Practice corrupted state, unavailable stores, stale checkpoints, cross-tenant access attempts, and partial external effects.

Use versioning for state schemas. When the application changes from checkpoint_version: 7 to a new format, the harness should either migrate old state deterministically or stop safely with an actionable recovery path. Quietly guessing at missing fields is a poor production strategy.

Observability should make one question easy to answer: why did this run take this action using this state? A useful trace can connect the session, checkpoint, model call, tool call, approval event, external identifier, and final outcome without exposing secrets.

Engineering principle

Resumeability is a production property, not a UI feature. A “Continue” button is only trustworthy when the system can prove where execution stopped and what external effects already happened.

09
Common Mistakes

1. Treating the whole transcript as the state.
The transcript is useful evidence, but it is not necessarily a reliable workflow record. Correction: store explicit workflow facts, approvals, external references, and checkpoint versions.

2. Using a session ID as authorization.
A session identifier groups data; it does not automatically establish that the caller is allowed to access that data. Correction: authenticate and authorize separately. OpenAI's current session documentation makes this distinction explicitly. See the session trust-boundary notes.

3. Letting summaries become business truth.
A summary can omit a critical exception or preserve a malicious instruction. Correction: keep critical facts in typed, authoritative state and use summaries mainly as derived context.

4. Retrying an external write without reconciliation.
A timeout does not prove the action failed. Correction: reconcile first, then retry under an idempotency or deduplication strategy where supported.

5. Persisting secrets inside memory.
Long-lived context tends to spread further than intended. Correction: keep credentials in dedicated secret storage and expose only the minimum authority needed at tool execution time.

6. Assuming compaction means “just summarize the chat.”
Long-running systems need to preserve the state required to continue correctly, not merely produce a shorter transcript. Current platform documentation reflects this distinction. See OpenAI's compaction guidance.

7. Resuming two workers against the same mutable history.
Concurrent restoration can create duplicate or conflicting state transitions. Correction: use locking, optimistic concurrency, workflow ownership, or another explicit single-writer/reconciliation design appropriate to your architecture.

8. Checking only final output.
An agent can produce a convincing final sentence after making the wrong tool calls. Correction: evaluate state transitions, permissions, side effects, and recovery traces.

🎯 Use this when...

You are reviewing an existing agent that “usually remembers,” but nobody can clearly explain how it recovers after failure.

10
❓ FAQ

Q1. Is session history the same thing as agent memory?

No. Session history is one form of stored interaction state. Memory can include durable facts, summaries, preferences, artifacts, or other records. A harness should define which memories are authoritative and which are only contextual hints.

Q2. Does resuming a session mean the model will see everything from the previous run?

Not necessarily. A session can contain stored history while the harness selects, filters, summarizes, or compacts what is sent to the model. The current OpenAI Agents SDK, for example, supports session history retrieval plus per-run controls over how much history is included. See the session documentation.

Q3. What should happen when the last tool call may or may not have succeeded?

Treat the outcome as unknown until reconciled. Query an authoritative system when possible, use an idempotency or deduplication mechanism when supported, and stop for controlled review when the system cannot establish the outcome safely.

Q4. Should important state be written into the prompt so the model can remember it?

The model can use important state through its context, but the prompt should not become the authoritative database. Keep critical workflow facts in durable state and project the minimum necessary subset into the next model context.

Q5. What is the simplest state design for a small agent?

Start with a clearly scoped session record, a small workflow state object, explicit checkpoint versions, and a deterministic recovery path. Add compaction, distributed stores, or more advanced memory only when the workload actually requires them.

REFERENCES
🔗 References & Further Reading

The architecture notes and examples in this article are original teaching material. Product names, APIs, and specifications belong to their respective owners.

SUMMARY
📝 Summary

  • Session connects related runs; it is not automatically authorization.
  • Context is what the model sees for the current decision; it should be reconstructed deliberately.
  • Durable state records facts and workflow transitions that must survive restarts.
  • Checkpoints make recovery explicit instead of asking the model to guess what happened.
  • Safe resumption reconciles uncertain external effects before retrying consequential actions.
  • State is a security surface: scope it, protect its integrity, control retention, and track provenance.
  • Evaluate the harness, not just the answer: test state transitions, permissions, side effects, and recovery traces.


Comments