Designing a Reliable Agent Loop: Stopping Rules, Retries, and Progress Tracking
A production harness must decide when the work is actually complete, when another step is justified, when a failed operation can safely be retried, when progress has stalled, and when the run should stop rather than continue consuming time, money, permissions, or external side effects.
Imagine a support-ticket agent that must investigate a customer's issue, read the relevant account information, consult internal documentation, prepare a response, and update the ticket. A weak implementation says, "Keep asking the model to continue until it sounds finished." A stronger harness asks much more concrete questions: What evidence proves completion? What state has already been changed? Is this retry safe? Has the agent actually made progress? Is human approval required? What happens when a tool times out after possibly completing the operation?
That is the engineering problem this article addresses. The focus is not prompt wording. The focus is the runtime control system around the model: stopping rules, retry policy, progress state, recovery, cancellation, and verification.
📝 Child-friendly analogy
Think of the agent loop as a child cleaning a room with a checklist. The child can decide what to pick up next, but an adult still decides when the room is actually clean enough, when the child has repeated the same action too many times, and when it is time to stop because something needs adult help.
The exact boundary between model, agent, harness, application, and sandbox varies by architecture. Some platforms own the loop for you; others let the application own the entire loop. The engineering principles below are therefore expressed as responsibilities rather than as one universal vendor architecture.
Running example used throughout
Fictional scenario: a support-ticket agent must investigate a ticket, retrieve relevant facts, determine whether the issue can be resolved, prepare a reply, obtain approval before sending the reply when required, and close the ticket only after the required evidence exists.
- What the Agent Loop Actually Does
- Designing Stopping Rules
- Why “The Model Says Done” Is Not a Completion Check
- Designing Safe Retries
- Progress Tracking: State, Evidence, and Milestones
- Recovery, Cancellation, and Partial Progress
- Observability and Evaluation of the Loop
- Worked Design: A Reliable Support-Ticket Loop
- Enterprise Rollout
- Common Mistakes
- ❓ FAQ
- 🔗 References & Further Reading
- 📝 Summary
A reliable agent loop does not ask only, “Should the model continue?” It asks, “Does the system have sufficient evidence to continue, stop, retry, wait, fail, or recover?”
| Concern | Harness responsibility | Typical failure when omitted |
|---|---|---|
| Stopping | Define explicit completion, cancellation, blocking, and budget conditions. | Runaway loops, premature completion, or expensive repeated work. |
| Retrying | Classify failures and decide which operations are safe to replay. | Duplicate side effects or repeated failures. |
| Progress | Persist milestones and verified state independently from model prose. | Lost work, loops over the same step, misleading “percentage complete.” |
| Recovery | Know what happened before failure and what can safely resume. | Restarting from the beginning or repeating irreversible actions. |
01
What the Agent Loop Actually Does
An agent loop is the repeated runtime cycle that lets a model observe information, choose a next action, receive the result, and continue until a terminal condition is reached. The loop may be implemented by an SDK, a workflow engine, an application, or a combination of these.
One current example is the OpenAI Agents SDK. Its documented runner repeatedly invokes the current agent, checks whether the output is final, handles handoffs when applicable, executes tool calls, and continues. It also exposes a configurable maximum turn count and raises a specific exception when that limit is exceeded. That is an architecture-specific implementation, not a universal definition of every agent framework. OpenAI Agents SDK: Running Agents
📝 Child-friendly analogy
A loop is like a person following a treasure map. The model can choose which clue to inspect next. The harness is the rulekeeper that says whether there is still a map to follow, whether the person has reached the destination, whether they have walked in circles, and whether they are allowed to open the next door.
A useful conceptual sequence looks like this:
receive task
↓
load trusted state
↓
invoke model
↓
inspect model decision
↓
┌──────────────┬────────────────┬───────────────────┐
│ final result │ tool operation │ needs intervention│
└──────┬───────┴───────┬────────┴─────────┬─────────┘
│ │ │
verify execute pause
│ │ │
└───────→ update state ←──────────┘
│
apply stop rules
│
continue or finish
This simple diagram reveals an important boundary. The model may recommend an action, but the harness should still own deterministic runtime decisions such as cancellation, time budgets, maximum turns, retry policy, authorization checks, and whether a tool result is sufficient evidence to mark a milestone complete.
This does not mean the harness must be a huge framework. For a short workflow, a small loop around a model call and a few tools can be enough. OpenAI's current documentation explicitly distinguishes cases where developers may want to own the loop directly from cases where the Agents SDK can manage turns, tools, guardrails, handoffs, or sessions for them. OpenAI Agents SDK overview
You are deciding where logic belongs. Put deterministic runtime controls in the harness; let the model make judgments where uncertainty is unavoidable; let tools perform bounded external operations; keep durable business state outside ephemeral model prose.
02
Designing Stopping Rules
The first question in a reliable loop is not “How many turns should I allow?” It is “What conditions mean this run has reached a legitimate terminal state?” A maximum-turn limit is a guardrail. It is not a definition of success.
For a production harness, useful terminal states commonly include:
- Completed: required success criteria have been verified.
- Blocked: the task needs a human, missing information, or an external condition that the agent cannot safely resolve.
- Cancelled: the user, operator, timeout policy, or control plane stopped the run.
- Budget exhausted: the run reached a defined operational boundary such as turns, tool calls, wall-clock duration, or another application-specific budget.
- Failed: the system reached a condition from which continuing would be unsafe or unlikely to help.
- Stalled: the run keeps producing activity without producing meaningful state change.
🟢 Worked example — defining success
For the fictional support-ticket agent, “done” is not simply “the model generated a helpful paragraph.” A stronger completion predicate might require:
- The required ticket facts were retrieved.
- The issue classification is present.
- A response draft exists.
- Any required human approval has been recorded.
- The final external action, if any, has a verified result.
- The ticket status reflects the actual outcome.
Notice the difference between activity and progress. Calling another tool is activity. Changing a required state from “missing” to “verified” is progress.
A useful stopping hierarchy is:
if completion_is_verified:
stop(COMPLETED)
elif cancellation_requested:
stop(CANCELLED)
elif human_input_required:
stop(WAITING_FOR_HUMAN)
elif hard_budget_reached:
stop(BUDGET_EXCEEDED)
elif fatal_failure:
stop(FAILED)
elif progress_is_stalled:
stop(STALLED)
else:
continue()
Illustrative pseudocode: this is a design sketch, not a vendor SDK implementation. The exact ordering should be adapted to the application's safety and business rules.
Why put verified completion before a generic turn limit? Because a good agent may finish earlier than the maximum. Conversely, a badly behaving agent may continue indefinitely unless a hard limit exists. You need both positive stopping conditions and defensive limits.
🛡 Safety Check
Risk → The model continues because it believes more investigation is useful, even though an external action has already succeeded.
Control → The harness records deterministic completion evidence and checks it before allowing another action cycle.
Remaining risk → A faulty verifier can still declare success incorrectly, so critical operations may require independent read-back checks or human review.
You are tempted to use a single “max iterations” value as the whole control strategy. A turn limit is necessary in many systems, but it should be a safety ceiling rather than the only definition of completion.
03
Why “The Model Says Done” Is Not a Completion Check
Language is not the same thing as state. A model can produce a confident sentence such as “The ticket has been resolved,” while the underlying ticket remains open, the message was never sent, or a tool result was incomplete.
This distinction becomes even more important when the agent is allowed to modify external systems. Anthropic's description of effective agents emphasizes that agents can use intermediate environment results as ground truth and can stop at completion conditions or checkpoints rather than assuming that the model's own narration is sufficient. Anthropic: Building Effective AI Agents
📝 Tricky concept: observation versus assertion
If a child says, “I put the book away,” the statement is an assertion. Looking at the shelf and seeing the book there is evidence. An agent harness should prefer the second kind of signal whenever the task depends on an external state.
This suggests a useful separation:
| Signal | What it means | Reliability role |
|---|---|---|
| Model message | The model believes a task condition has been met. | Useful reasoning signal, not automatically proof. |
| Tool result | An external system returned an observation. | Often stronger evidence about external state. |
| Application state | A deterministic record maintained by the application. | Best place for durable workflow milestones. |
| Independent verification | A second mechanism checks a consequential result. | Useful for high-impact or ambiguous operations. |
For example, after the support agent invokes a “send response” operation, the harness should not blindly assume success because the tool call returned without an exception. Depending on the tool contract, the response may contain an operation identifier, status, or confirmation that can be persisted and later reconciled. The exact mechanism is architecture-specific.
Completion should be a predicate over trusted state, not merely a sentence generated by the model.
This also makes testing easier. You can ask, “Which state fields must be true for this workflow to be completed?” That question is much easier to test than “Did the final answer sound confident enough?”
A task changes external state, requires approval, or has a business definition of “done” that can be represented independently of model prose.
04
Designing Safe Retries
Retries are necessary because agent runs depend on unreliable components: networks fail, services throttle requests, tools time out, and models occasionally return malformed or unusable output. But a retry is not automatically safe.
📝 Child-friendly analogy
Retrying a question is like asking someone to repeat what they said. Retrying “transfer the money” is different. The first request might already have worked even if you did not hear the confirmation. Repeating it could create a second transfer.
That difference is the boundary between harmless recovery and dangerous duplication.
Current OpenAI Agents SDK documentation is unusually explicit about this distinction. Its runner-managed retry facility considers error type, retry-after information, timeout and network conditions, whether a response has already started, and replay-safety information. It also blocks several classes of replay by default. This is an example of a framework implementing a broader principle: retry policy must consider whether replaying the operation is safe. OpenAI Agents SDK: Models and runner-managed retries
| Failure type | Typical treatment | Why |
|---|---|---|
| Temporary network problem | Often retry with bounded backoff. | The underlying operation may not have failed permanently. |
| Rate limiting | Honor provider guidance such as a retry-after signal when available. | Immediate repetition can make throttling worse. |
| Validation error | Usually stop or ask for correction rather than blindly retry. | The same invalid request is likely to fail again. |
| Permission failure | Stop, escalate, or follow a controlled credential-recovery path. | Repeating an unauthorized action does not grant authority. |
| Tool timeout after a write | Treat outcome as potentially unknown until reconciled. | A timeout does not necessarily prove that the remote action did not happen. |
| Malformed model output | Retry or reject according to the output contract and retry budget. | The failure may be local to the generation attempt. |
A second important distinction is model-call retry versus agent-loop continuation. Suppose a model request fails once and the runtime retries the same request. That is not necessarily another agent turn. If the model successfully returns a tool call and the tool then fails, the harness may either retry the tool, surface the error to the model, or terminate. These are different control layers.
Without explicit separation, retry multiplication can become surprisingly large:
outer agent loop
├── model retry policy
│ └── request attempts
│
└── tool retry policy
└── tool attempts
If every layer has an independent retry allowance, the total number of attempts can grow far beyond what a developer expected. A production design should therefore have a visible retry budget and should know which layer consumed it.
🛡 Safety Check
Risk → A non-idempotent external action is replayed because a timeout was mistaken for proof of failure.
Control → Classify the tool operation before retrying; use idempotency or reconciliation mechanisms for operations that can change state.
Remaining risk → A remote system may expose incomplete or misleading status information, so critical actions may still require human review or compensating procedures.
The Model Context Protocol also provides a useful vocabulary here. Its tool annotations include hints such as read-only behavior, destructive behavior, and idempotency. Importantly, the MCP project explicitly describes these annotations as hints, not guarantees. Clients should not turn a declared annotation into their only safety boundary. Model Context Protocol: Tool Annotations as Risk Vocabulary
A robust retry decision can therefore be expressed conceptually as:
retry_allowed =
failure_is_transient
AND retry_budget_remaining
AND operation_replay_is_safe
AND run_not_cancelled
AND overall_budget_remaining
Illustrative design expression: this is a conceptual policy, not a universal formula or vendor API.
Backoff is another important control. If many agents immediately retry the same unavailable service, they can amplify an outage. Bounded backoff and jitter spread attempts over time. The exact values should come from the service's reliability requirements rather than being copied mechanically from a sample.
Your agent touches external APIs, databases, messaging systems, file operations, or anything where repeating an action can change state twice.
05
Progress Tracking: State, Evidence, and Milestones
Longer agent runs create a second reliability problem: the system needs to know what has already happened. Conversation history alone is not a sufficient progress database.
A useful distinction is:
- Transcript: what the model and tools said during the interaction.
- Durable state: structured information the application trusts and can use after a restart.
- Trace: an observability record showing how the run evolved.
- External state: the actual state of downstream systems such as a ticket, order, database record, or message.
📝 Child-friendly analogy
The transcript is the child's diary. The progress state is the checklist on the refrigerator. The trace is the security camera record. The actual room is the external system. A diary can say “I cleaned the room,” but the room itself remains the ultimate thing that needs checking for work that affects reality.
For the fictional support-ticket agent, a durable progress record might conceptually contain:
run_id task_id current_phase completed_milestones required_milestones blocked_reason retry_counts last_verified_state last_external_operation approval_status created_at updated_at
The field names above are an illustrative schema, not a vendor standard. The important property is that the harness can answer, without asking the model to reconstruct history:
- What has definitely completed?
- What remains?
- What is currently blocked?
- What external action was last attempted?
- Can the next action be safely replayed?
Avoid relying too heavily on “progress percentage.” An agent saying “80% complete” may sound useful, but percentage requires a stable denominator. Many agentic tasks discover new work while executing, making a numeric percentage misleading.
🟢 Worked example — milestone tracking
Instead of tracking:
“Ticket resolution: 67%”
track explicit facts:
- Ticket loaded: yes
- Customer account verified: yes
- Root-cause evidence collected: yes
- Response drafted: yes
- Approval required: yes
- Approval received: no
- External response sent: no
This makes the next step obvious: the system is not “67% done.” It is waiting for approval before a consequential action.
Progress tracking also helps detect stalls. A run can make many model calls without changing any important milestone. A simple stall detector can look for repeated state snapshots, repeated tool calls with the same arguments, or repeated failure categories over a bounded window.
meaningful_progress =
newly_verified_fact
OR newly_completed_milestone
OR newly_resolved_blocker
OR safely_completed_external_action
This definition is intentionally conservative. Repeating a thought, reformulating a query, or producing another natural-language explanation does not automatically count as progress.
🛡 Safety Check
Risk → The model writes its own progress state and marks an unverified action as complete.
Control → Let deterministic application logic update critical milestone fields from validated tool results and approval events.
Remaining risk → Incorrect tool data can still contaminate state, so trust boundaries and validation must extend to tool outputs.
Runs can be long, interrupted, resumed, reviewed by humans, or expensive enough that losing partial progress would be painful.
06
Recovery, Cancellation, and Partial Progress
A production agent should assume that runs can stop unexpectedly. The process may crash. A browser tab may close. A human may cancel the request. A tool may become unavailable. A deployment may restart the worker.
Reliable recovery begins with a simple question:
Recovery should resume from verified state, not from the assumption that the previous attempt either completely succeeded or completely failed.
This is especially important for tools with external side effects. Consider a support response that is being sent through an external messaging system. The agent invokes the tool. The network connection times out. What happened?
There are at least three possibilities:
- The operation definitely failed before execution.
- The operation definitely succeeded.
- The operation's outcome is unknown.
The third case is the dangerous one. A naïve retry converts uncertainty into a potential duplicate side effect.
Possible recovery strategies include:
- Read-after-write verification: query the downstream system to determine whether the expected state already exists.
- Idempotency key: use a stable operation identifier when the downstream interface supports deduplication.
- Reconciliation: place ambiguous operations into a state that requires explicit resolution before continuing.
- Human review: require an operator when the impact of an incorrect replay is high.
These are architectural strategies rather than universal API features. Your downstream system determines which ones are actually available.
Cancellation deserves similar attention. Cancellation should not merely mean “stop asking the model for another response.” A well-designed harness should determine what cancellation means for in-flight work. For example, stopping a waiting model call is different from undoing an already accepted payment, job submission, deployment, or message.
Current MCP ecosystem documentation also illustrates why cancellation and long-running execution are explicit protocol concerns rather than accidental application behavior. The MCP ecosystem supports cancellation and, in newer task-oriented extensions, durable task states such as working and input-required. These capabilities are architecture-specific and should not be treated as universal requirements for every agent. MCP Tasks specification
📝 Child-friendly analogy
If someone stops you while you are building a Lego model, you do not start from zero next time. You look at the model that is already on the table. Good recovery works the same way: continue from verified pieces instead of blindly repeating earlier work.
A practical state machine might therefore distinguish:
RUNNING WAITING_FOR_TOOL WAITING_FOR_HUMAN WAITING_FOR_RECONCILIATION COMPLETED FAILED CANCELLED
These labels are illustrative. The key idea is to make recovery state explicit enough that the next process invocation knows what happened.
🛡 Safety Check
Risk → A cancelled or crashed run is restarted from the original task without inspecting prior external actions.
Control → Persist checkpoints and external-operation identifiers; resume from a known state or enter explicit reconciliation.
Remaining risk → Recovery logic can itself contain bugs, so exercise cancellation and crash-recovery scenarios in evaluation before production rollout.
Your agent may pause for human approval, run longer than a typical request, use external tools, or operate in environments where workers can restart.
07
Observability and Evaluation of the Loop
A reliable loop cannot be judged solely by its final response. Two runs can produce the same final answer while one uses three tool calls and the other uses thirty, one retries a failed action safely while the other repeats a write, and one finishes with verified state while the other merely sounds confident.
This is why traces matter. The OpenAI Agents SDK currently records events such as agent runs, model generations, tool calls, guardrails, and handoffs in its tracing system. Again, that is a product-specific implementation, but it illustrates an important architectural idea: the harness should produce enough execution evidence to explain how a result was reached. OpenAI Agents SDK: Tracing
Useful loop-level observability fields can include:
- Run identifier
- Turn number
- Current state or phase
- Tool name and outcome category
- Retry count and reason
- Milestone transition
- Stop reason
- Approval or escalation event
- Cost, latency, and resource measurements where available
Be careful about sensitive data in traces. Logging a tool name and outcome can be useful without storing the entire content of every customer record. Current agent SDK documentation explicitly provides configuration around whether sensitive data is included in traces, which reinforces the broader principle that observability has its own data-governance boundary. Tracing configuration reference
Evaluation is the other half of the problem. Anthropic's current guidance on agent evaluations emphasizes that agents operate over multiple turns, modify state, and use tools, making evaluation more complex than checking a single generated answer. Anthropic: Demystifying Evals for AI agents
🟢 Worked example — reliability test matrix
| Test scenario | Expected harness behavior |
|---|---|
| Task completes after one tool call | Stop immediately after verified completion. |
| Transient tool failure | Retry according to bounded policy. |
| Validation failure | Do not spin on the same invalid operation. |
| Repeated identical state | Detect stall and stop or escalate. |
| Human approval required | Pause instead of repeatedly asking for the same action. |
| Tool timeout after possible write | Enter reconciliation rather than blindly replaying. |
| Run cancellation | Stop new work and leave recoverable state. |
Useful metrics can include unnecessary turns, tool-call count, retry frequency, duplicate-action rate, stalled-run rate, successful recovery rate, time-to-completion, approval wait time, and resource consumption. These are engineering measurements, not model benchmark scores.
🛡 Safety Check
Risk → Tracing becomes a second source of sensitive-data exposure.
Control → Define trace fields intentionally, minimize sensitive payloads, apply retention rules, and restrict access.
Remaining risk → Even metadata can reveal sensitive workflow information, so trace access should remain subject to the application's security model.
08
Worked Design: A Reliable Support-Ticket Loop
Let us now combine the ideas into one coherent design. Everything in this scenario is fictional and is intended to demonstrate engineering mechanics rather than describe a real company's deployment.
🟢 Fictional scenario
A support agent receives a ticket saying, “My subscription was charged twice.” The agent can read ticket details, inspect billing information, search approved documentation, prepare a response, request human approval, send the approved message, and update the ticket status.
Step 1 — Define the required outcome.
The harness defines that a run is complete only after the ticket has an evidence-backed resolution path and all required downstream actions have succeeded or the workflow has explicitly ended in a human-review state.
Step 2 — Give the model a bounded set of tools.
The agent does not receive arbitrary database credentials or an unrestricted shell. It receives narrowly scoped operations such as:
get_ticket(ticket_id) get_billing_facts(customer_id) search_approved_knowledge(query) create_reply_draft(ticket_id, content) request_human_approval(ticket_id, action_summary) send_approved_reply(ticket_id, approval_id) update_ticket_status(ticket_id, status)
Those names are illustrative. The important design property is bounded functionality and explicit permissions.
Step 3 — Track milestone state outside the model.
ticket_loaded = true billing_facts_verified = true resolution_path_verified = true reply_drafted = true approval_required = true approval_received = false reply_sent = false ticket_closed = false
The model can use these facts to decide what to do next, but the application remains responsible for updating the authoritative fields based on validated events.
Step 4 — Separate “continue” from “retry.”
Suppose search_approved_knowledge() temporarily fails. The harness may retry the tool according to its policy. If the tool succeeds, the agent continues. That is different from the model deciding that it needs another tool call because the result revealed missing information.
Step 5 — Stop at the approval boundary.
When a message is ready but human approval is required, the run should enter a waiting state instead of repeatedly generating alternative messages. The state should clearly record why execution stopped and what event is required to resume.
🛡 Safety Check
Risk → A model sees “approval required” as another instruction to solve and repeatedly attempts the consequential operation.
Control → Make approval a harness-controlled state transition. The tool is unavailable or rejected until the required approval record exists.
Remaining risk → Authorization logic can still be incorrectly configured, so access to consequential tools must be reviewed independently of model instructions.
Step 6 — Reconcile ambiguous outcomes.
Assume send_approved_reply() times out. The harness does not immediately call it again. Instead it records:
last_operation = "send_approved_reply" last_operation_status = "UNKNOWN" reconciliation_required = true
The system can then use a safe verification path, such as checking the message system for a corresponding operation identifier, before deciding what action is possible next.
Step 7 — Close only after verification.
The final “close ticket” action is permitted only after the required external response is confirmed. This prevents the harness from confusing “response generation completed” with “customer interaction completed.”
completion_is_verified =
billing_facts_verified
AND resolution_path_verified
AND reply_drafted
AND approval_received
AND reply_sent
AND ticket_closed
Illustrative pseudocode: the expression represents the design idea, not a drop-in application implementation.
This workflow is intentionally more constrained than “let the agent work until it feels finished.” That constraint is a feature, not a weakness. The harness has made the important state transitions explicit.
You are designing agents for business processes where completion has a concrete definition and some actions have greater consequences than others.
09
Enterprise Rollout
A reliable loop becomes an enterprise concern when it is connected to production systems, multiple users, customer data, financial or operational workflows, or long-running tasks.
The organization does not need a giant governance program for every agent. It does need clear ownership of the controls that matter.
| Control area | Useful ownership question | Evidence to retain |
|---|---|---|
| Stopping | Who defines the business completion predicate? | Versioned workflow rules and test cases. |
| Retries | Who decides which operations are replay-safe? | Tool contracts, retry policies, idempotency or reconciliation procedures. |
| Permissions | Does the agent have more authority than the task requires? | Tool inventory, scopes, identities, approval gates. |
| Observability | Can an operator reconstruct why a run stopped or retried? | Trace identifiers, events, stop reasons, operational metrics. |
| Recovery | What state survives process failure? | Checkpoint schema, recovery tests, reconciliation records. |
| Change control | What happens when tools, policies, or completion rules change? | Versioned configuration and regression evaluations. |
Security should be treated as a property of the whole loop. OWASP's current guidance on excessive agency highlights three recurring causes of risk: excessive functionality, excessive permissions, and excessive autonomy. The same guidance recommends limiting tools and privileges and adding appropriate approval for high-impact actions. OWASP LLM06:2025 — Excessive Agency
For the loop specifically, this means the following controls should be considered together:
- Least privilege: expose only the tools and permissions required by the task.
- Bounded autonomy: require deterministic checks or human approval before high-impact actions.
- Resource limits: cap turns, retries, tool calls, time, or other relevant consumption.
- Traceability: record consequential decisions and state transitions.
- Recovery testing: deliberately test crashes, timeouts, cancellation, duplicate attempts, and partial completion.
🛡 Safety Check
Risk → A harness is technically reliable but has enough authority to turn a single bad decision into a high-impact action.
Control → Treat loop reliability and authorization as separate controls: reliable execution should operate inside a least-privilege boundary.
Remaining risk → Even strong runtime controls cannot eliminate all model or integration failures. They reduce the blast radius and make failures observable and recoverable.
10
Common Mistakes
The following mistakes are common because they make prototypes appear simple. They become expensive once real tools and real state are involved.
1. Using maximum turns as the definition of success
Cause: a hard limit is easy to implement. Consequence: the agent can reach the limit without knowing whether the task is complete. Correction: define a completion predicate and keep maximum turns as a separate defensive boundary.
2. Retrying every error
Cause: “retry” feels like a universal recovery mechanism. Consequence: validation errors repeat forever and write operations may happen twice. Correction: classify errors before retrying.
3. Treating a timeout as proof that nothing happened
Cause: the caller saw no successful response. Consequence: a replay can duplicate an external action. Correction: distinguish known failure from unknown outcome and reconcile the downstream state.
4. Keeping progress only in the conversation
Cause: the model already sees the transcript. Consequence: process restarts and long contexts become harder to recover reliably. Correction: maintain compact, structured milestones outside the transcript.
5. Counting activity as progress
Cause: more tool calls look like more work. Consequence: an agent can loop while appearing busy. Correction: define meaningful progress as a verified state or milestone change.
6. Giving the model authority to decide all runtime limits
Cause: the model is already making decisions. Consequence: the same component being controlled also controls the boundary. Correction: enforce hard limits in deterministic runtime code.
7. Stacking independent retry systems
Cause: SDK, HTTP client, tool wrapper, and agent loop each get their own retries. Consequence: latency and attempts multiply unexpectedly. Correction: define retry ownership explicitly and measure total attempts.
8. Assuming tool metadata is a complete safety boundary
Cause: a tool is labeled read-only, idempotent, or non-destructive. Consequence: runtime trusts a hint that may be inaccurate. Correction: treat tool metadata as useful context while enforcing critical guarantees independently.
🛡 Safety Check
Risk → A looping agent repeatedly invokes powerful tools because each layer assumes another layer will stop it.
Control → Place deterministic ceilings at the harness boundary and enforce authorization at the tool boundary.
Remaining risk → A compromised or misleading tool can still influence the model's next decision, so untrusted tool output should be treated as data rather than authority.
11
❓ FAQ
Q1. Is a maximum-turn limit enough to prevent an agent from looping forever?
A maximum-turn limit provides a hard upper bound on one type of loop activity, but it does not define successful completion, prevent duplicate external actions, or recover partial progress. A reliable harness combines a positive completion condition with defensive limits, retry budgets, cancellation, and state-aware recovery. Some agent frameworks expose maximum-turn controls directly; the exact semantics depend on the framework.
Q2. How should I decide whether an agent action can be retried?
Start with the failure type and the replay consequence. Temporary transport problems can be retryable, while validation or authorization failures normally need correction rather than repetition. For external writes, determine whether the operation is idempotent, has a deduplication mechanism, or can be reconciled safely before replay. A timeout after a possible write should be treated as an unknown outcome until verified.
Q3. Should the model generate and maintain the progress state?
The model can propose what should happen next, but important progress fields should normally be maintained by application logic from validated events. This reduces the chance that the model declares a milestone complete simply because it believes it completed it. The transcript can remain useful as context while durable state serves as the machine-readable record of progress.
Q4. What should the harness do when a tool times out after a potentially successful write?
Do not automatically repeat the write. First represent the outcome as ambiguous, then use a read-back, operation identifier, idempotency mechanism, reconciliation workflow, or human review appropriate to the system. The correct method depends on what the downstream service guarantees.
Q5. Can a well-designed harness guarantee that an agent will always behave correctly?
No. A harness can bound execution, constrain permissions, classify failures, persist state, detect stalls, require approvals, and make actions observable. Those controls reduce risk and improve recoverability, but they do not guarantee that the model, tool, verifier, or surrounding application will always behave correctly. Reliability is therefore a layered engineering property rather than a single feature.
12
🔗 References & Further Reading
The following primary sources were used to verify architecture-specific and current claims. The explanations, teaching sequence, examples, pseudocode, table wording, and FAQ content in this article are original.
- OpenAI Agents SDK — Running Agents — used to verify the documented runner loop, final-output termination behavior, handoffs, tool execution, and maximum-turn handling.
- OpenAI Agents SDK — Models and Runner-Managed Retries — used to verify current retry-policy concepts, backoff configuration, timeout handling, and replay-safety considerations.
- OpenAI Agents SDK — Tracing — used to verify current tracing coverage and trace configuration concepts.
- Anthropic — Building Effective AI Agents — used to verify the multi-step agent-loop model, environment feedback, checkpoints, and stopping-condition concepts.
- Anthropic — Demystifying Evals for AI Agents — used to verify the current emphasis on evaluating multi-turn, stateful agent behavior.
- Model Context Protocol — Tool Annotations as Risk Vocabulary — used to verify the meaning and limitations of read-only, destructive, idempotent, and open-world tool hints.
- OWASP — LLM06:2025 Excessive Agency — used to verify the security discussion around excessive functionality, permissions, and autonomy.
- Google Search Central — FAQ structured data guidance — used to verify the current FAQ rich-result guidance.
Vendor and standards references are included for factual verification and architecture-specific examples. Product names and trademarks belong to their respective owners.
13
📝 Summary
- Stopping rules define when the agent should finish, wait, fail, cancel, or stop because it is no longer making meaningful progress.
- A maximum-turn limit is a defensive runtime boundary, not a business definition of completion.
- Completion should be verified against trusted state and external evidence rather than inferred only from model prose.
- Retries must be classified by failure type and replay safety; not every failure should be retried.
- Progress should be explicit through milestones and verified state rather than vague percentages or repeated model activity.
- Unknown outcomes matter: after a timeout on a potentially successful write, reconciliation may be safer than replay.
- Observability and evaluation should measure the trajectory of the run, not only the final text.
- Security remains layered: bounded loops, least privilege, validation, approvals, and tracing reduce risk but do not make an agent completely safe.
A reliable agent is not an agent that never stops. It is an agent that knows why it is continuing, why it is stopping, what it has verified, and what it is allowed to do next.
Comments
Post a Comment