Running Agent Harnesses in Production: Reliability, Cost, and Long-Running Tasks
Running an AI agent in production is not simply a matter of keeping the model call alive.
The real engineering problem is keeping the entire run understandable, bounded, recoverable, observable, and safe when the work takes longer than one request, one process, one context window, or one human interaction.
Imagine a fictional enterprise agent asked at 9:00 AM to review a large batch of supplier invoices, compare them with purchase records, identify exceptions, prepare a report, and wait for a finance reviewer before sending any consequential update. At 9:18 AM the browser disconnects. At 9:27 AM the worker process restarts. At 10:05 AM an external service times out. At 10:40 AM a reviewer approves one action but rejects another.
A demo agent might simply start over. A production harness should know what work was completed, what remains uncertain, which actions were already attempted, where durable state lives, which operations may be retried, which actions require approval, and when the run must stop.
A long-running agent is reliable only when the harness can preserve progress, bound execution, verify outcomes, and recover without blindly repeating side effects.
01 — What changes when an agent enters production?
02 — Separate request lifetime from work lifetime
03 — Durable state, checkpoints, and recovery
04 — Long-running work across context windows
05 — Cost control is a runtime responsibility
06 — Retries, duplicate actions, and idempotency
07 — Permissions, tools, external input, and safety
08 — Observability, evaluation, and completion criteria
09 — Cancellation, approvals, and human recovery points
01
What changes when an agent enters production?
A useful starting point is to stop treating “the agent” as one piece of software. A production system usually has several responsibilities distributed across multiple components.
🟧 Child-friendly analogy
Think of an agent like a very capable worker in a workshop. The model is the worker's reasoning engine. The tools are the equipment. The harness is the workshop manager: it decides what the worker is allowed to do, records what happened, pauses the job when necessary, and makes sure the job does not continue forever. A sandbox, when one exists, is a fenced work area where execution can happen with tighter isolation.
Model. The model produces predictions and tool-selection decisions from the information supplied to it. It does not automatically become a durable workflow engine merely because it can perform many reasoning or tool-use turns.
Agent. The agent is the behavior built around the model: instructions, available tools, policies, state, and the loop that allows additional work after an intermediate result.
Harness. The harness is the runtime and control layer around the run. Depending on the architecture, it may own tool dispatch, context assembly, state persistence, retries, permissions, human approvals, tracing, cancellation, budgets, verification, and recovery.
Application. The surrounding application may own authentication, user interfaces, business workflows, tenant isolation, API endpoints, queues, and external system integration. Some of these responsibilities may also be implemented inside the harness.
Sandbox. A sandbox is an optional isolated execution environment. It can reduce the blast radius of code execution or filesystem access, but a sandbox does not automatically solve authorization, business-policy errors, data leakage, or incorrect external side effects.
↓
Application / API boundary
↓
Harness: policy + state + loop + budgets + verification
↓
Model ↔ tool adapters ↔ external services
↓
Optional sandbox / isolated execution environment
↓
Durable state + traces + business artifacts
This boundary is architectural rather than universal. One platform may package several of these responsibilities together, while another may expose them separately. Current official documentation illustrates that difference: the OpenAI Agents SDK describes a higher-level runtime with tools, sessions, human-in-the-loop support, and tracing, while Microsoft documents durable agent hosting as an additional durability layer around agent execution. OpenAI Agents SDK documentation and Microsoft Durable Extension documentation.
🟩 Worked example — fictional enterprise agent
Our fictional Supplier Reconciliation Agent receives a batch of invoices, retrieves approved purchasing data, compares records, creates an exception report, and prepares proposed actions. It can read enterprise systems automatically, but sending a consequential update requires a human approval step.
The important engineering question is not “Can the model complete the task?” It is “Can the runtime continue, stop, recover, and prove what happened when the task is interrupted?”
Your agent performs multiple tool calls, takes longer than a normal request, changes external state, or needs to survive worker restarts and human waiting periods.
02
Separate request lifetime from work lifetime
One of the most important production design decisions is deciding whether the user's request and the agent's work have the same lifetime. Often they should not.
🟧 Child-friendly analogy
Ordering a pizza and baking the pizza are different things. The phone call can end quickly, but the oven still needs time. A production API request is often the phone call. The durable work is the baking.
A user-facing HTTP request might need to return quickly because of client timeouts, load-balancer limits, mobile connectivity, or the simple fact that the user does not need to keep a browser tab open while the job runs.
That creates a useful separation:
| Architecture shape | State ownership | Typical fit | Main production concern |
|---|---|---|---|
| Interactive run | Request or session context | Short conversations and bounded tool use | Latency and request failure |
| Background run | Run record plus continuation state | Work that outlives the initiating request | Reconnect, status, and failure handling |
| Durable workflow | Persistent checkpoints and business state | Hours, days, waiting periods, external events | Correct resume and duplicate side effects |
These patterns are not mutually exclusive. A durable workflow can contain short interactive model calls. A background request can still be non-durable if its continuation state disappears with the worker. The architectural question is where the durable boundary actually lives.
There are also vendor-specific implementations. For example, OpenAI documents a background mode for Responses API processing; the platform documentation notes that this mode stores response data for roughly ten minutes to support polling. That makes it a useful example of request decoupling, but it should not be casually interpreted as a complete durable business-workflow system. OpenAI data-controls documentation.
🛡 Safety Check
Risk → A client reconnects and accidentally starts a second copy of the same job.
Control → Give the business operation a stable run or job identity and make the start operation explicitly idempotent.
Remaining risk → A client can still present stale or conflicting commands, so the server must validate authorization and current run state on every consequential transition.
The work may continue after the initiating request ends, or when users need to leave and later return to the same run.
03
Durable state, checkpoints, and recovery
Long-running execution becomes much easier to reason about when the harness explicitly separates conversation history from workflow state.
🟧 Tricky concept: memory is not the same as progress
A diary can tell you what happened, but it does not automatically tell you what the next safe step is. A checkpoint is closer to a bookmark with a rule: “Step 4 finished successfully; Step 5 is pending; approval is required before Step 6.”
A durable agent can maintain several kinds of state:
- Run identity: a stable identifier for the business job or agent run.
- Workflow position: the current phase, completed steps, outstanding work, and pending approvals.
- Business artifacts: reports, extracted records, intermediate files, and other durable outputs.
- Tool outcome records: enough information to reconcile what an external system actually accepted.
- Operational metadata: attempt number, timestamps, policy decisions, usage, error categories, and cancellation state.
The harness does not have to store all of this in one database. The important property is that the source of truth is explicit and recoverable.
Microsoft's current durable-agent documentation describes persistent agent sessions, checkpointing, recovery after failures, distributed scaling, and pause/resume behavior for human interaction. Those are architecture-specific implementations of the broader durable-execution pattern, not requirements that every agent must use the same technology. Microsoft Durable Extension and Durable Task extension for Agent Framework.
For the fictional Supplier Reconciliation Agent, a useful durable state might look like this:
phase: "exception_review"
completed_batches: 7
pending_batch: 8
report_uri: <ARTIFACT_LOCATION>
approval_status: "waiting"
last_confirmed_external_effect: "none"
retry_count: 2
cancel_requested: false
The exact schema is application-specific. The principle is not: recovery should begin from an authoritative state record rather than from the model's best guess about what happened.
Checkpoint design matters. A checkpoint placed before a consequential side effect can cause the same action to be attempted again after recovery. A checkpoint placed only after everything is finished provides too little progress information. A useful boundary is often after a step whose result has been validated and whose side effects can be safely identified.
🟩 Worked example — one safe transition
The harness retrieves purchase records, validates the response, writes a durable batch result, and only then advances the workflow checkpoint.
If the worker crashes after the validation but before the checkpoint is committed, the recovery path must determine whether the durable result already exists. It should not assume that “the tool call probably succeeded” or “the step probably failed.”
🛡 Safety Check
Risk → Recovery treats conversational history as authoritative business state.
Control → Keep business-critical state and confirmed side effects in explicit durable records.
Remaining risk → A corrupt or incomplete state record can still produce incorrect recovery, so critical transitions should have consistency checks and reconciliation procedures.
Your agent can be interrupted after any tool call, approval, external event, or worker restart.
04
Long-running work across context windows
There is another boundary that becomes important before many developers expect it: a long-running agent may outlive its useful context. Even if the workflow state is durable, the model cannot necessarily keep every previous message, tool result, intermediate explanation, and document in every subsequent model call.
Context engineering is therefore one responsibility that can appear inside a harness: choosing what information is supplied to the model for the current step. It is not the definition of harness engineering itself.
🟧 Child-friendly analogy
Imagine a large project folder. A new engineer should not carry every email, meeting transcript, and scratch note into every meeting. They need the current task, the important decisions, the latest artifacts, and the rules that still apply. The harness has to prepare that smaller working set without losing authoritative facts.
Common techniques include:
- Compaction or summarization: reduce old conversational detail while retaining useful decisions and progress.
- Artifact handoff: store structured progress outside the conversation and reload the relevant artifact later.
- Selective retrieval: provide only the records needed for the current step rather than re-sending everything.
- Task decomposition: turn one huge objective into smaller independently verifiable phases.
The distinction between context and state is critical. A compact summary can help the model remember what happened, but it should not become the only source of truth for something such as “invoice 347 was already posted successfully.” That fact should live where the harness can verify it.
Anthropic has publicly described long-running agent work in which an initializer creates structure for a task, later sessions make incremental progress, and structured artifacts carry information between sessions. Its 2026 work also discusses harness design for multi-hour autonomous development and the use of context resets. These are documented vendor-specific designs that illustrate the general problem: a long-running run may need explicit mechanisms to bridge context boundaries. Anthropic — Effective harnesses for long-running agents and Anthropic — Harness design for long-running application development.
The same idea appears in different forms across agent runtimes. OpenAI's Agents SDK, for example, documents sessions as a persistent memory layer for maintaining working context across agent runs. Again, that is one implementation of session persistence rather than a universal harness architecture. OpenAI Agents SDK Sessions.
🟩 Worked example — crossing a context boundary
Suppose the fictional reconciliation job finishes seven batches and is about to process batch eight. The harness does not need to replay every previous tool response. It can provide the model with:
Current task: reconcile batch eight.
Durable facts: batches one through seven are confirmed complete.
Relevant policy: proposed external updates require approval.
Current artifact: exception report version three.
Pending decision: reviewer has not yet approved any write action.
Use context to help the model reason; use durable state to tell the system what is actually true.
The same task may span multiple model sessions, context resets, worker processes, or human handoffs.
05
Cost control is a runtime responsibility
An agent can become expensive without doing anything malicious. Repeated reasoning, repeated retrieval, oversized tool results, unnecessary retries, and accidental loops can all turn a small task into a large one.
That makes cost a harness concern rather than merely a procurement concern.
🟧 Child-friendly analogy
Imagine giving someone a taxi card and saying, “Keep driving until you are completely satisfied.” A careful system gives them a destination, a route budget, and rules for stopping. An agent needs equivalent limits for model calls, tool use, retries, elapsed time, and concurrency.
A useful production pattern is to assign every run a budget vector rather than one vague timeout.
| Budget | What it limits | Example stop condition | Why it matters |
|---|---|---|---|
| Time | Wall-clock execution | Maximum run window reached | Prevents endless or abandoned work |
| Model turns | Number of model-loop iterations | Turn limit reached | Bounds reasoning loops |
| Tool calls | External actions or retrievals | Tool-call budget exhausted | Bounds downstream workload |
| Retry budget | Repeated attempts | Retry ceiling reached | Stops retry storms and cost amplification |
The actual values are workload-specific. A simple read-only research workflow may tolerate more model turns but no external writes. A high-impact financial workflow may need a small write-action budget and explicit approvals.
Cost control also includes controlling the size and repetition of context. If a tool returns a ten-megabyte payload and the harness forwards most of it into every subsequent model call, the runtime has created a cost and latency multiplier. Tool adapters should therefore return information sized for the decision being made, with pagination, filtering, summarization, or structured fields where appropriate.
A production harness should also distinguish expected retry from unexpected repetition. Repeating a transiently failed read may be correct. Repeating a completed write because the response was lost is a different problem.
OWASP's 2025 LLM risk guidance explicitly discusses Unbounded Consumption as a risk involving uncontrolled inference, resource depletion, unexpected cost, and service degradation. Its mitigations include rate limiting, resource controls, timeouts, monitoring, and limits on queued work. These are directly relevant to an agent harness because the harness is where repeated execution can often be bounded. OWASP — LLM10:2025 Unbounded Consumption.
max_wall_time: <WORKLOAD_LIMIT>
max_model_turns: <TURN_LIMIT>
max_tool_calls: <TOOL_LIMIT>
max_retry_attempts: <RETRY_LIMIT>
max_parallel_jobs: <CONCURRENCY_LIMIT>
stop_on_budget_exhaustion: true
When a budget is exhausted, the harness should produce an explicit state such as budget_exhausted, not quietly continue or pretend that the task completed normally.
🛡 Safety Check
Risk → A malicious or merely difficult request causes excessive model and tool usage.
Control → Apply per-run quotas, tenant limits, rate limits, queue limits, and hard execution budgets.
Remaining risk → A bounded run can still consume the full allowed budget, so monitoring and abuse detection are still required.
The agent can repeat work, fan out into many tool calls, process user-controlled input, or run concurrently at meaningful scale.
06
Retries, duplicate actions, and idempotency
Failures are normal in distributed systems. Network timeouts happen. Workers restart. Queue messages are delayed. External APIs return errors. A production harness should therefore distinguish retryable uncertainty from confirmed failure.
🟧 Child-friendly analogy
You ring a shop and ask whether your order was placed. The phone cuts out after the shopkeeper says, “Yes, I entered it.” Calling again and saying “place the order again” may create two orders. A safer approach is to ask whether order number <ORDER_ID> already exists.
This is why idempotency matters. An operation is idempotent when repeating it produces the same intended final outcome rather than creating an unwanted duplicate effect.
For agent tools, idempotency can be implemented in several ways depending on the service:
- Stable idempotency key: the same logical action carries the same unique request identity.
- Check-before-write: the tool checks whether the intended effect already exists.
- Transactional boundary: the state record and the business change are committed in a way that allows safe reconciliation.
- Reconciliation endpoint: the harness can query the downstream system after a timeout to discover the actual result.
Microsoft's reliability guidance explicitly emphasizes idempotent task design for retries, and its durable-agent material warns that applications remain responsible for preventing duplicate side effects. These are useful examples of a general distributed-systems rule rather than a requirement to use Microsoft's infrastructure. Azure Batch reliability guidance and Microsoft long-running agent resilience guidance.
Do not retry every error. Authentication failures, policy denials, malformed requests, and explicit business rejections usually need a different path from transient infrastructure errors. A generic “try again three times” wrapper around every tool can turn a small failure into repeated side effects.
🟩 Worked example — handling a lost response
The harness asks a downstream service to create an exception ticket with an idempotency key derived from <RUN_ID> and the business exception identifier.
The network connection fails before the response arrives. Recovery does not blindly issue a second create request. It first asks the downstream system whether that logical ticket already exists, then records the confirmed state.
↓
timeout / uncertain outcome
↓
reconcile downstream state
↓
already completed? ── yes → record confirmed result
└─ no → retry with same logical identity
🛡 Safety Check
Risk → A retry duplicates a real-world action such as sending, deleting, booking, or updating.
Control → Make consequential tools idempotent where possible and reconcile uncertain outcomes before another attempt.
Remaining risk → Some external systems cannot provide reliable idempotency or reconciliation, so the harness may need a human review state rather than an automatic retry.
A model decision can cause an external side effect and the network response cannot be assumed to be the same as the business outcome.
07
Permissions, tools, external input, and safety
Long-running execution increases the importance of security because the agent has more opportunities to encounter untrusted data, make tool calls, and make mistakes.
A harness should treat external content as data, not as automatically trusted instructions. An invoice note, web page, email body, document, repository file, or tool response may contain text that attempts to influence the model. The harness should not assume that content is safe merely because it arrived through an approved connector.
The important protection is not a magic “prompt injection filter.” It is damage containment: limit the functions, permissions, identities, data visibility, network access, and approval authority available to the run.
OWASP's current Excessive Agency guidance describes three root causes that are especially relevant to agent harnesses: excessive functionality, excessive permissions, and excessive autonomy. Its mitigations include minimizing available extensions, narrowing their functionality, using least privilege, and independently verifying or approving high-impact actions. OWASP — LLM06:2025 Excessive Agency.
🟧 Child-friendly analogy
A child may be allowed to borrow a library book without being handed the keys to the library. Giving an agent a tool should work the same way: provide exactly the authority needed for the task, not every capability that happens to be available behind the same connector.
Least privilege should exist at several layers:
- Tool surface: expose only the functions the run needs.
- Data scope: limit records, tenants, folders, projects, or accounts that the tool can access.
- Identity: use an identity appropriate to the run instead of an unnecessarily powerful shared account.
- Action authority: separate read, propose, and commit operations when practical.
- Network boundary: restrict unnecessary outbound connectivity and internal service reachability.
A sandbox can help with code execution, but it should be considered one layer. A sandbox does not decide whether an API call is authorized, whether a record belongs to the user, whether the model misunderstood a business rule, or whether the final result is correct.
🛡 Safety Check
Risk → Untrusted document content influences a model and causes an unauthorized tool action.
Control → Treat external content as untrusted data, restrict tools and downstream permissions, validate tool arguments, and require approval for high-impact actions.
Remaining risk → No single prompt rule, filter, sandbox, or approval step eliminates instruction injection or model error. Defense in depth reduces the possible impact when one layer fails.
For long-running jobs, permissions should also be reconsidered over time. A run that starts with read-only access may later reach a phase where it needs a write operation. That does not mean the harness should simply start the whole job with write access.
🟩 Worked example — propose before commit
The fictional reconciliation agent may automatically collect records and generate an exception report. Instead of giving the model direct authority to send supplier notices, the harness creates a proposed action. A human reviewer sees the intended target, reason, and affected records. Only an explicit approval causes the commit-capable tool to become available.
Your agent consumes external content, uses connectors, accesses enterprise data, or can cause a business side effect.
08
Observability, evaluation, and completion criteria
A production agent cannot be operated safely if the team sees only the final answer. A useful trace should let an engineer reconstruct the important lifecycle of the run without exposing secrets unnecessarily.
At minimum, think about recording:
- Run identity and tenant/user identity, subject to privacy and retention policy.
- State transitions, including checkpoint and approval changes.
- Tool calls, arguments after appropriate redaction, result status, latency, and downstream request identifiers.
- Budget usage, including model-call count, tool-call count, elapsed time, retries, and relevant provider usage metadata.
- Failure classification, distinguishing policy denial, validation error, transient infrastructure failure, timeout, cancellation, and business rejection.
Do not treat “log everything” as a security strategy. Logs may themselves contain confidential records, credentials, or sensitive user content. Redaction, retention limits, access control, and trace sampling are part of the operational design.
Modern agent SDKs show why tracing and testing are increasingly treated as runtime concerns. OpenAI's Agents SDK documents built-in tracing for model generations, tool calls, handoffs, guardrails, and custom events. Its testing guidance also distinguishes deterministic harness-level testing from integration testing for external model and infrastructure behavior. OpenAI Agents SDK and OpenAI Agents SDK Testing.
Evaluation should include the harness, not just the model. A production test can ask:
- Does the harness stop when the turn budget is exhausted?
- Does a worker restart resume from the correct checkpoint?
- Does a duplicate event avoid duplicating a business side effect?
- Does an unauthorized tool call remain blocked even when the model requests it?
- Does cancellation prevent new side effects after the cancellation boundary?
- Does the system report “failed” rather than claiming success when a required verification step did not pass?
🟧 Tricky concept: final text is not completion
A model can confidently say “The reconciliation is complete” while a required downstream write failed, an approval is still pending, or one batch was never processed. Completion is a harness decision based on observable criteria, not merely a sentence generated by the model.
A useful completion contract can include:
all_required_steps_confirmed = true
required_artifacts_persisted = true
required_external_effects_reconciled = true
pending_approval = false
budget_exhausted = false
run_status = "completed"
This is illustrative pseudocode rather than a vendor SDK contract. The point is to move completion from “the model sounded finished” to “the system's explicit success conditions were satisfied.”
🛡 Safety Check
Risk → A misleading final response hides an incomplete or failed external action.
Control → Define completion as verified state transitions, required artifacts, and reconciled side effects.
Remaining risk → Verification itself can be incomplete or wrong, so high-impact workflows may need independent controls outside the agent.
You need to debug failures, prove what happened, evaluate regressions, or decide whether an agent actually completed its business objective.
09
Cancellation, approvals, and human recovery points
A long-running agent needs a meaningful answer to one simple question: What happens when a person says “stop”?
Cancellation is often more complicated than setting a boolean because there may be an in-flight model request, an external tool call, a queue message, or a process that cannot be interrupted immediately.
🟧 Child-friendly analogy
Pressing the stop button on a washing machine may stop it from starting another cycle, but it does not magically reverse water already pumped into the drum. A cancellation request must therefore distinguish “do not start anything else” from “undo what already happened.”
A practical cancellation model is:
- Requested: persist the cancellation request.
- Quiescing: stop scheduling new model/tool work where possible.
- Reconcile: determine whether any in-flight operation actually completed.
- Final state: record whether the run was cancelled, partially completed, or requires operator recovery.
The same state-machine mindset helps with human approvals. A pending approval should be a durable system state, not merely a message left in a chat window.
OpenAI's Agents SDK documents human-in-the-loop interruptions that can pause execution, serialize run state, and resume after an approval or rejection. That is one concrete implementation of a broader pattern: a production harness needs a persistent boundary at which human decisions can safely interrupt and resume a run. OpenAI Agents SDK Human-in-the-loop and OpenAI Run State.
Approval fatigue is another production concern. If the user receives a confirmation request for every harmless read operation, users may approve automatically. A more useful design is to align approval thresholds with impact: no approval for low-risk reads, explicit approval for high-impact writes, and possibly grouped review for several related low-risk actions.
🟩 Worked example — approval as a durable boundary
The Supplier Reconciliation Agent reaches a point where it has prepared five proposed supplier notices. The harness stores the proposals, the target records, the policy decision, and the current run state. The model does not need to “remember” the approval request from an old conversation. The approval record itself is durable and bound to the run and authorized reviewer.
🛡 Safety Check
Risk → An old, duplicated, or unauthorized approval is replayed against a new run.
Control → Bind approvals to the run, user, action, and relevant state version; verify authorization before resuming.
Remaining risk → A correctly authorized human can still approve an incorrect action, so the approval screen should expose meaningful scope and expected impact rather than only a generic “Approve” button.
The task can wait for a human, can be cancelled, or can cross a boundary where a person or external event must decide what happens next.
10
Enterprise Rollout
A production harness should normally be introduced in stages. The goal is to prove the runtime behavior before granting the system broad authority.
Stage 1 — Observe. Run against realistic workloads without consequential writes. Measure latency, model usage, tool-call frequency, failure categories, context growth, and recovery behavior.
Stage 2 — Read-only production work. Allow retrieval and analysis against real data while keeping external mutations disabled or simulated. Test cancellation, worker restarts, duplicate events, and budget exhaustion.
Stage 3 — Controlled writes. Introduce narrowly scoped commit tools. Add approval gates, idempotency, downstream reconciliation, and explicit ownership for failures.
Stage 4 — Long-running workloads. Allow work to survive process restarts and user disconnects. Validate durable state, context handoffs, queue behavior, and stale-run cleanup.
Stage 5 — Operational maturity. Establish change control, incident response, budget alerts, trace retention, access review, runbook ownership, and regression evaluation.
| Ownership area | Production question | Evidence to retain |
|---|---|---|
| Application team | What business outcome defines success? | Completion contract and business artifacts |
| Platform team | Can runs survive infrastructure failures? | Checkpoint history and recovery records |
| Security team | Can the run exceed intended authority? | Policy decisions, access records, tool permissions |
| Operations team | How is an unhealthy run detected and recovered? | Alerts, traces, runbooks, incident records |
Secrets and credentials deserve separate attention. The model should generally not receive raw credentials simply because a tool needs them. The harness or tool adapter should obtain and present only the authority needed to perform the operation.
Retention should also be designed deliberately. A long-running job can produce much more trace data than a short chat. Retain what is needed for debugging, audit, security investigation, and business requirements, then remove what is no longer needed.
Incident response should include operational controls for agent runs: disable a problematic tool, stop new runs, quarantine a workflow version, revoke an identity, or require approval for a previously automatic action. The goal is not merely to restart the service. It is to contain the authority of the running system.
🛡 Safety Check
Risk → A production incident is handled by restarting workers, while the underlying bad policy or tool permission remains unchanged.
Control → Maintain operational kill switches, tool-level disablement, access revocation, version rollback, and documented incident procedures.
Remaining risk → An already-completed external side effect cannot always be undone, so containment and reconciliation must be designed before incidents occur.
The agent is moving from prototype to a shared enterprise environment with real users, real data, and real downstream effects.
11
Common Mistakes
1. Restarting from the beginning after every failure.
Cause: the application stores only the original request.
Consequence: repeated model work, repeated tool calls, and possible duplicate effects.
Correction: persist workflow state and resume from confirmed progress.
2. Treating conversation history as the business database.
Cause: the transcript feels like a convenient memory store.
Consequence: important state becomes ambiguous, oversized, difficult to query, or inconsistent with external systems.
Correction: keep authoritative business state in structured durable records.
3. Retrying every failure automatically.
Cause: retry logic is attached to the whole agent loop instead of individual failure classes.
Consequence: policy errors repeat, costs increase, and writes may duplicate.
Correction: classify failures and retry only operations that are safe to retry or can be reconciled.
4. Giving the agent a broad connector because it is easier.
Cause: one connector exposes reads, writes, and administrative operations.
Consequence: a mistaken or manipulated decision has a larger blast radius.
Correction: narrow the tool surface, downstream identity, data scope, and action authority.
5. Using a single giant model turn for the entire job.
Cause: decomposition and checkpoint design were skipped.
Consequence: failure near the end wastes earlier work and makes recovery difficult.
Correction: divide the task into independently verifiable phases.
6. Forgetting cancellation.
Cause: the prototype assumes the run finishes once started.
Consequence: users cannot reliably stop an expensive or incorrect workflow.
Correction: make cancellation a first-class state transition and reconcile in-flight operations.
7. Measuring only model quality.
Cause: evaluation focuses on the final response.
Consequence: the team misses failures in orchestration, permissions, recovery, state, and side effects.
Correction: test the entire harness lifecycle separately from model behavior.
8. Declaring success because the final message sounds confident.
Cause: completion is inferred from language rather than verified state.
Consequence: incomplete work is reported as successful.
Correction: require explicit completion criteria and independent verification where the action matters.
9. Letting approval become a meaningless button.
Cause: every action requires confirmation, so users learn to click through.
Consequence: human oversight becomes nominal rather than useful.
Correction: reserve approvals for meaningful impact and show reviewers what will actually happen.
10. Assuming a sandbox solves the whole security problem.
Cause: isolation is confused with authorization and correctness.
Consequence: data, API, identity, or business-rule risks remain outside the sandbox boundary.
Correction: treat sandboxing as one control inside a broader defense-in-depth design.
12
❓ FAQ
1. Does a long-running agent always need a durable workflow engine?
No. A short, bounded agent run may only need request-scoped state and ordinary retries. Durable execution becomes valuable when work must survive process restarts, long waits, human approvals, external events, or extended execution. The correct durability mechanism depends on the architecture and the failure modes that matter.
2. Is conversation history enough for recovering an agent run?
Usually not for important business workflows. Conversation history can help reconstruct context, but durable workflow state should explicitly record confirmed progress, pending work, approvals, artifacts, and important external effects. Recovery should rely on authoritative state rather than a model's interpretation of an old transcript.
3. How should an agent harness control cost?
Bound the main sources of work: elapsed time, model turns, tool calls, retries, concurrency, and context growth. Measure actual usage, detect repeated work, keep tool results appropriately sized, and stop explicitly when a budget is exhausted. Cost controls should be part of runtime policy rather than an after-the-fact billing exercise.
4. What is the safest way to retry a tool call?
First classify the failure. Then determine whether the operation is idempotent or whether the downstream result can be reconciled. For uncertain consequential writes, confirm the actual external state before repeating the action. Some failures are better routed to human recovery than automatically retried.
5. Can prompt injection be eliminated by a good harness?
No single technique eliminates instruction injection or model error. A strong harness limits the impact through least-privilege tools, narrow permissions, validation, isolation where appropriate, approval for high-impact actions, monitoring, and clear recovery procedures. The goal is to reduce blast radius and make failure detectable and recoverable.
13
🔗 References & Further Reading
The following primary sources were used to verify architecture-specific claims and current guidance relevant to this article:
- OpenAI Agents SDK documentation — runtime capabilities including tools, sessions, human-in-the-loop support, and tracing.
- OpenAI Agents SDK Sessions — persistent session history and resumed runs.
- OpenAI Agents SDK Human-in-the-loop — interruption, approval, serialization, and resume patterns.
- OpenAI Agents SDK Testing — deterministic harness testing and integration boundaries.
- OpenAI API data controls — current endpoint-specific application-state and background-mode behavior.
- Anthropic — Effective harnesses for long-running agents — published long-running harness design and cross-session artifacts.
- Anthropic — Harness design for long-running application development — current discussion of multi-hour autonomous work and harness structure.
- Anthropic — Scaling Managed Agents — architecture-specific discussion of long-horizon managed agents and harness boundaries.
- Microsoft Agent Framework — Durable Extension — persistent state, checkpoints, recovery, and distributed hosting.
- Microsoft Durable Task extension for Agent Framework — durable agent sessions, long-running workflows, and known operational constraints.
- OWASP — LLM06:2025 Excessive Agency — excessive functionality, permissions, autonomy, and mitigation strategies.
- OWASP — LLM10:2025 Unbounded Consumption — resource limits, timeouts, rate limiting, monitoring, and cost-related abuse.
Product and organization names belong to their respective owners. Architecture-specific claims are attributed to the relevant primary sources.
14
📝 Summary
- A production agent run needs a clear boundary between the model, agent behavior, harness controls, tools, application services, and any sandbox.
- Request lifetime and work lifetime do not have to be the same.
- Durable state should record confirmed progress instead of relying on conversation history alone.
- Context management helps the model continue work, but durable business state remains the authoritative source for important facts.
- Cost is a runtime property: bound time, turns, tools, retries, concurrency, and context growth.
- Retries need idempotency or reconciliation when external side effects are involved.
- Least privilege, narrow tools, validation, approvals, and defense in depth reduce the impact of model mistakes and instruction injection.
- A final model message is not proof of completion; the harness should verify completion criteria.
- Cancellation, human approval, tracing, evaluation, and incident response should be designed as first-class runtime behavior.
The production harness is what turns a sequence of model decisions into an operationally controlled task: bounded when it runs, explicit about authority, durable when it must wait, observable when it fails, and recoverable when reality does not follow the happy path.
Comments
Post a Comment