Skip to main content

Testing an AI Agent Harness: Completion Checks, Tracing, and Debugging

Calculating read time…
An AI agent can produce a plausible answer, call the right tools, and still fail the task. 
That is why testing an agent harness cannot stop at checking whether the final text “looks good.” A harness needs ways to determine whether the requested outcome actually happened, whether the agent stayed within its authority, whether the run stopped correctly, and what happened between the first model call and the final result.

Consider a fictional invoice-exception agent. A user asks it to inspect an invoice, compare it with a purchase order, and open an exception case if the values do not match. The model may correctly identify a mismatch but never create the case. Or it may create the case twice after a retry. Or it may create the case successfully but incorrectly report that the case was created before the tool confirms it. A final-answer check alone can miss all three problems.

Engineering principle

Agent testing should verify both the outcome and the path that produced it. The path does not need to be identical on every run, but important boundaries, permissions, side effects, and stopping conditions must remain testable.

In this post, completion checks verify that a task really reached its required state; tracing records the important events of a run so engineers can inspect what happened; and debugging uses those observations to find the earliest meaningful divergence from the expected behavior.

One important boundary

The model generates or selects the next action. The agent is the model-driven behavior operating toward a task. The harness manages the run around that agent: state, tools, limits, approvals, retries, stopping, validation, and observability. A tool performs an external operation. A sandbox, when used, provides an isolated execution environment. An application provides the user-facing business workflow. These boundaries differ by architecture, so treat them as design responsibilities rather than a universal component diagram.

Focus Core question Best evidence
Completion checks Did the required outcome actually happen? Authoritative application or tool state
Tracing What happened during the run? Structured events and spans
Debugging Where did behavior first diverge? Trace + failed assertion + state

01
What Testing an Agent Harness Really Means

🟧 Child-friendly analogy

Imagine a child asked to pack a school bag. You do not judge the job by asking, “Did you say you packed it?” You open the bag and check: notebook present, pencil case present, homework present, nothing important missing. Agent testing works the same way. The agent's message is a report. The harness needs independent evidence about the actual state.

Traditional software tests often assume a predictable sequence of instructions. Agent systems are different because the model can choose among tools, repeat actions, stop early, recover from failures, or take another valid route to the same result. Modern agent-evaluation guidance therefore treats the final outcome and the recorded execution as related but distinct evidence.

That does not mean “anything goes.” The harness still owns rules that should not be left to the model: which tools are available, which arguments are valid, when an approval is required, how many attempts are allowed, what constitutes a successful completion, and what must happen when the environment disagrees with the model.

Test concern Question Evidence Typical owner
Completion Did the requested business state actually occur? Database or API state, artifact, result object Harness + domain logic
Trace What happened during the run? Ordered events, tool calls, timings, errors Harness / observability
Debugging Where did the run first leave the expected contract? Trace comparison, assertion, state mismatch Engineer / platform team

A useful mental model is to separate three questions:

1. Did it finish? The task reached the required state.

2. Did it behave within its contract? Tool, permission, budget, approval, and stopping rules were respected.

3. Can we explain what happened? The run left enough evidence for diagnosis.

🎯 Use this when...

You are moving from an agent prototype that “works in a demo” to a system that needs repeatable tests, incident diagnosis, and controlled changes.

02
Completion Checks: Proving the Task Is Actually Done

🟧 Child-friendly analogy

A delivery driver saying “package delivered” is not enough. Someone still needs to check whether the package is actually at the right door. The completion check is that final physical check.

A completion check is an independent condition that determines whether the task reached its intended result. The strongest checks do not ask the model whether it succeeded. They inspect evidence available outside the model's own claim.

✅ Worked example — fictional scenario

Input: invoice <INVOICE_ID> and purchase order <PO_ID>.

Required outcome: identify a material mismatch and create exactly one exception case containing the correct invoice and purchase-order references.

Required non-outcomes: do not create a duplicate case, do not approve payment, and do not access unrelated customer records.

Check What to verify Failure example
Outcome state Exactly one case exists with the expected references. Case was never created.
Data correctness Case fields match validated source records. Wrong purchase order attached.
Side-effect boundary No payment approval or unrelated mutation occurred. Agent used a broader write capability.
Termination Run stops after verified completion. Agent continues polling after success.
Illustrative pseudocode — not a vendor SDK
result = run_agent(task)

checks = [
outcome_exists(expected_case),
outcome_matches_source(expected_case),
no_forbidden_side_effects(),
within_tool_budget(),
approval_policy_respected(),
run_stopped_after_completion()
]

passed = all(checks)

if not passed:
mark_run_failed()
preserve_trace_for_debugging()

Notice the separation: the agent performs work; the harness decides whether the observed state satisfies the task contract. That separation reduces the risk of treating the model's own confidence as proof.

Engineering principle

The strongest completion signal is usually external evidence, not the model's final sentence.

🛡 Safety Check

Risk → The agent reports success before a side effect is actually committed.

Control → Verify authoritative downstream state before marking the task complete.

Remaining risk → The downstream system may be eventually consistent, delayed, or ambiguous. The harness still needs an explicit policy for “pending” versus “verified complete.”

🎯 Use this when...

The agent changes data, creates records, sends messages, executes code, or performs any action where “I think I did it” is not enough evidence.

03
Testing the Agent Loop Instead of Only the Final Answer

🟧 Child-friendly analogy

Imagine testing a calculator by checking only the last number it displays. A wrong sequence of button presses could accidentally produce the correct final number. Agent tests have the same problem: a good-looking result can hide a bad journey.

An agent run is often a loop:

user task
   ↓
harness prepares context + permissions
   ↓
model proposes next action
   ↓
harness validates / authorizes
   ↓
tool executes
   ↓
result enters the agent context
   ↓
completion check?
   ├─ yes → verify → stop
   └─ no  → continue / recover / ask for approval

Every arrow creates a possible test boundary. The model can select an inappropriate tool. The harness can validate the wrong field. The tool can return an error that the agent misunderstands. State can be lost between turns. A retry can repeat a non-idempotent operation. A cancellation can arrive after a side effect has already started.

For that reason, useful harness tests are layered rather than singular:

Tool contract tests. Give adapters known inputs and confirm schema validation, authorization checks, error translation, and output normalization.

Loop tests. Exercise sequences such as model → tool → model → completion, including retries and failure branches.

Scenario tests. Start with a realistic task and evaluate resulting state plus important behavior constraints.

Trace assertions. Check that required events happened and prohibited events did not happen.

Regression tests. Re-run stable task sets whenever tools, instructions, models, state management, or harness logic changes.

✅ Worked example — two valid paths

Suppose the agent must determine whether an invoice differs from a purchase order.

Path A: fetch invoice → fetch PO → compare → create exception.

Path B: fetch a combined reconciliation record → compare → create exception.

Both may be valid if the architecture permits both tools and the final state is correct. A rigid test that rejects Path B merely because it expected Path A is testing an implementation detail rather than the business contract.

🛡 Safety Check

Risk → A test suite rewards a model for reaching the correct result through an unauthorized path.

Control → Pair outcome assertions with invariant checks such as allowed-tool set, permission scope, approval requirements, and forbidden side effects.

Remaining risk → The invariant list can still be incomplete. Review it whenever tools, business permissions, or downstream APIs change.

04
Tracing: Turning an Agent Run Into an Inspectable Timeline

🟧 Child-friendly analogy

If a toy robot suddenly takes the wrong turn, watching only where it ends up may not explain anything. You need a replay of the steps: what instruction it received, what result it got, which turn it chose, and when the mistake first appeared. A trace is that kind of replay record.

A trace is a structured record of a logical agent run. A trace normally contains smaller units often called spans: model calls, tool executions, agent transitions, guardrails, state operations, or application-defined events. The exact trace schema depends on the framework.

The important architectural idea is broader than any vendor API: a useful trace should let an engineer reconstruct the significant events without requiring a complete transcript of every internal detail.

Trace field Why it matters Example
Run identity Connects logs, state, user request, and downstream effects. run-PLACEHOLDER
Event order Shows what happened before and after a failure. tool call → error → retry
Tool metadata Shows which capability was invoked and whether it succeeded. get_invoice / create_case
Timing Separates model latency, tool latency, and waiting time. 1.4s model, 3.2s tool
Outcome markers Makes completion and failure states machine-checkable. verified_complete
event = {
  "run_id": "run-PLACEHOLDER",
  "type": "tool_result",
  "tool": "create_exception",
  "status": "success",
  "request_ref": "<REQUEST_REF>",
  "timestamp": "<TIMESTAMP>"
}

The exact fields will differ by implementation. The design principle is to make important decisions and side effects observable without storing unnecessary sensitive content.

🛡 Safety Check

Risk → Traces can capture secrets, personal data, authorization tokens, or sensitive tool arguments.

Control → Decide explicitly what is captured, redact sensitive values before export, restrict trace access, and define retention periods appropriate to the data.

Remaining risk → Redaction can miss new sensitive fields when the application changes. Treat trace schemas as security-sensitive interfaces requiring review.

🎯 Use this when...

The agent has enough steps that an application log saying “request failed” is no longer sufficient to explain where or why the run diverged.

05
Debugging: Finding the First Meaningful Divergence

🟧 Child-friendly analogy

Think about following a recipe. If the final cake is wrong, you do not start by blaming the oven. You first ask: which step was the first one that became incorrect? The wrong ingredient may have entered much earlier.

Agent failures often have the same structure. The visible error may occur several steps after the root cause. A tool failure may actually come from malformed state. A bad tool choice may come from an earlier context mistake. A duplicate record may come from retrying an operation that was not safe to repeat.

Expected:
1. authorize read_invoice
2. read invoice
3. authorize read_purchase_order
4. read purchase order
5. compare records
6. request approval if required
7. create exactly one exception
8. verify exception
9. stop

Observed:

1. authorize read_invoice
2. read invoice
3. authorize read_purchase_order
4. read purchase order
5. compare records
6. create exception
7. timeout
8. retry create exception
9. duplicate rejected
10. verify exception
11. stop

The key debugging discovery is not “step 9 failed.” It is “the harness retried a side-effecting operation after a timeout without first knowing whether the first request had committed.” That points the engineer toward idempotency, reconciliation, or operation-status lookup rather than toward prompt wording.

  1. Freeze the evidence. Save the run identifier, relevant inputs, tool responses, state snapshot, and test version.
  2. Mark the first failed assertion. Do not begin with the final error message if an earlier contract violation exists.
  3. Walk backward one event at a time. Ask which preceding event made the failure possible.
  4. Classify the fault. Model decision, harness logic, tool adapter, downstream service, state store, policy, or test itself.
  5. Reproduce with controlled dependencies. Replace unstable external services with deterministic fakes where appropriate.
  6. Add a regression test before closing the incident. The failure should become a test case, not merely a bug report.
✅ Worked example — diagnosing a false completion

Symptom: Users report that the agent says “exception created,” but the case is occasionally missing.

Trace finding: The agent emits a successful-looking final message immediately after the tool request, before the downstream response is observed.

Root cause: Completion was coupled to the model's decision rather than to verified downstream state.

Correction: The harness waits for the authoritative operation result, runs the completion check, and reports “pending” or “failed” when verification is unavailable.

Engineering principle

Debug the earliest broken invariant, not merely the loudest error.

06
Building Deterministic Tests Around Nondeterministic Behavior

🟧 Child-friendly analogy

A good teacher does not expect every student to solve a problem using exactly the same handwriting or sequence. The teacher checks whether the answer is correct and whether key rules were followed. Agent evaluation works better when tests distinguish necessary behavior from merely one possible path.

The model's outputs can vary between runs. That means a useful test system needs to control what it can control and measure what it cannot fully control.

Control the environment. Use stable fixtures for important input data. Stub external APIs when the goal is to test harness logic. Fix known tool responses for failure-path tests. Reset mutable state between trials.

Do not force determinism where flexibility is legitimate. A research agent may legitimately search different sources. A coding agent may repair a bug through several valid edits. In those cases, outcome and policy checks are stronger than exact transcript matching.

Test layer Good question Good evidence
Deterministic unit Does this adapter enforce its contract? Schema validation
Scenario Can the run complete a representative task? Verified outcome
Trajectory Did the run remain within required behavioral boundaries? Allowed tools, approvals, retries
Human review Does automated grading agree with expert judgment? Calibrated sample review
✅ Worked example — repeatable trial design

For a completion test, create a fixed invoice fixture, fixed purchase-order fixture, deterministic read responses, and a test exception store. Run the same task several times. Grade independently for:

  • correct mismatch detection,
  • exactly one created exception,
  • no forbidden mutation,
  • correct stop behavior.

07
Testing Permissions, Risky Actions, and Recovery

🟧 Child-friendly analogy

A school helper may have a key to the classroom, but that does not mean they should have the key to every room in the building. The agent should receive only the capabilities required for the task.

Testing the harness also means testing its authority boundaries. Excessive agency can arise from too much functionality, too much permission, or too much autonomy. The practical response is to narrow tools, constrain permissions, and require appropriate approval and downstream authorization.

Security test Question Expected behavior
Capability Can the agent invoke only tools intended for this workflow? Unneeded tools unavailable.
Authorization Does downstream authorization reject out-of-scope requests? Request denied independently.
Approval Does a high-impact action stop at the required checkpoint? No action before approval.
Recovery What happens after timeout, cancellation, duplicate request, or partial failure? Safe retry or explicit recovery state.
🛡 Safety Check

Risk → Untrusted content influences the model toward an action the user did not authorize.

Control → Test tool allow-lists, downstream authorization, approval gates, rate limits, and forbidden-action checks independently of the model.

Remaining risk → No single filter or approval mechanism eliminates instruction-injection risk. Damage still needs to be constrained by permissions and downstream controls.

Recovery deserves the same seriousness as prevention. Suppose a tool times out after sending a request. A naive harness may retry immediately. A safer design first asks: Is the previous operation known to be idempotent? If not, the harness may need a status-check operation or reconciliation step before issuing another write.

Illustrative recovery logic — not a working SDK
try:
    operation = execute_write(request)
    verify(operation)
except Timeout:
    status = lookup_operation_status(request.reference)

```
if status == "committed":
    verify_existing_result()
elif status == "not_found" and retry_is_safe:
    retry_write()
else:
    move_to_reconciliation()
    stop_agent_run()
```

08
Regression Suites and Production Debugging

🟧 Child-friendly analogy

Every time you repair a bicycle, you check the wheel, brakes, and chain again. Fixing one part can break another. A regression suite is that repeated safety check for software changes.

Agent behavior can change when almost anything in the harness changes: a tool schema, tool description, state representation, approval rule, retrieval logic, system instruction, model configuration, retry policy, or context budget. Therefore the regression suite should contain more than “does the final answer contain the expected sentence?”

  1. Happy-path tasks that represent normal business work.
  2. Boundary tasks with missing fields, empty results, ambiguous requests, and unsupported operations.
  3. Failure tasks involving tool errors, timeouts, stale state, and partial downstream success.
  4. Security tasks that attempt to cross permission boundaries using harmless test fixtures.
  5. Recovery tasks that verify cancellation, restart, resume, retry, and reconciliation behavior.
task_id
input_fixture
allowed_tools
required_approvals
completion_checks
forbidden_side_effects
expected_state
trace_assertions

The production side has a different purpose. A pre-release test asks, “Does the current implementation behave correctly on our known cases?” Production observability asks, “What is happening with real inputs that we did not predict?” The two complement each other.

Layer Primary purpose Typical signal
CI regression Catch known breakage before deployment. Completion check failed after code change.
Canary / staged run Observe behavior on limited real traffic. Unexpected retry or failure rate.
Production trace Investigate individual incidents. First failed span.
Human review Find cases that automated checks miss. Unclear or surprising behavior.
🛡 Safety Check

Risk → Production traces become a second database containing sensitive user activity.

Control → Minimize collected data, define role-based access, establish retention and deletion rules, and separate diagnostic identifiers from business secrets.

Remaining risk → Operational metadata can still reveal sensitive patterns even when payloads are redacted. Access governance must cover metadata as well as message content.

09
Enterprise Rollout

At enterprise scale, agent testing becomes an engineering lifecycle rather than a collection of scripts run by one developer. The exact operating model depends on the architecture, but a practical ownership split can look like this:

Responsibility What should be defined Typical owner
Task contract Success state, failure state, invariants, approval rules Product + domain + engineering
Harness controls Tool allow-list, retries, budgets, cancellation, state recovery Platform / agent engineering
Observability Trace schema, retention, access, alerting, incident evidence Platform / SRE / security
Evaluation suite Representative tasks, regression cases, security cases, calibration Shared ownership

A sensible rollout is incremental.

  1. Start with a narrow task contract. Write down what “done” means and what must never happen.
  2. Add completion checks. Verify actual state rather than model assertions.
  3. Instrument the run. Capture enough trace information to reconstruct failures.
  4. Build a small regression set. Include happy, boundary, failure, security, and recovery cases.
  5. Promote production incidents into tests. Every important failure should make the system harder to break in the same way twice.
  6. Review governance with scale. Expand access controls, retention, cost limits, change approval, and incident procedures as the agent's authority grows.
Engineering principle

Do not wait for production incidents to invent your evaluation strategy. Encode the task contract before the agent becomes difficult to reason about.

10
Common Mistakes

1. Mistake: Treating the final answer as the completion signal.
Cause: The model says “done,” so the application returns success. Consequence: false positives and missing side effects. Correction: verify authoritative state outside the model.

2. Mistake: Testing only the happy path.
Cause: Early demos rarely exercise timeouts, duplicate requests, missing data, cancellation, or authorization failures. Consequence: the most expensive failures appear in production. Correction: add boundary, failure, security, and recovery scenarios.

3. Mistake: Requiring one exact trajectory.
Cause: Engineers record one successful transcript and treat it as the only valid path. Consequence: legitimate agent flexibility is mistaken for failure. Correction: assert required outcomes and invariants; inspect exact trajectories only where the process is contractually constrained.

4. Mistake: Logging everything.
Cause: “More trace data must mean better debugging.” Consequence: larger cost, greater privacy exposure, and harder incident analysis. Correction: capture fields that answer known operational questions and redact what is not necessary.

5. Mistake: Retrying every failed tool call.
Cause: Generic retry logic is added without understanding tool semantics. Consequence: duplicate or repeated side effects. Correction: distinguish idempotent operations from non-idempotent ones and use reconciliation where the commit state is uncertain.

6. Mistake: Testing authorization only in the model.
Cause: The system instruction says the agent must not perform a forbidden action. Consequence: an unexpected tool call can bypass the intended boundary. Correction: enforce authorization in the harness and downstream system, and test the denial path.

7. Mistake: Ignoring the test environment.
Cause: External services, mutable databases, network conditions, or stale state vary between test runs. Consequence: engineers cannot tell whether the agent or the environment caused the failure. Correction: make important dependencies deterministic and record environment versions and fixtures.

8. Mistake: Fixing the prompt when the harness is broken.
Cause: The model is visible, so every issue looks like a model-behavior issue. Consequence: prompt changes hide underlying runtime defects. Correction: trace the event sequence and classify the first failure before changing instructions.

11
❓ FAQ

Q1. Should every agent have a deterministic test that reproduces the exact same tool sequence?

No. Some workflows have multiple valid paths. Test the required outcome, important safety invariants, permissions, approvals, and stopping conditions. Demand one exact sequence only when the sequence itself is part of the contract.

Q2. What is the difference between a completion check and an evaluation?

A completion check usually answers a narrow runtime question such as “Did the requested record reach the required state?” An evaluation is broader: it can combine several graders to assess outcome quality, process behavior, safety, cost, latency, or other task-specific criteria across repeated trials.

Q3. What should a trace contain?

Enough structured evidence to reconstruct important decisions and failures: run identity, event order, tool activity, model activity where appropriate, timings, errors, state transitions, and completion markers. Do not assume that capturing every prompt and tool payload is automatically better; privacy, retention, and access requirements matter.

Q4. How do I debug an agent that fails only sometimes?

Preserve the run evidence, repeat the task under controlled dependencies, compare the trace with the task contract, and identify the first broken invariant. Multiple trials can help separate a recurring harness defect from variable model behavior.

Q5. Can tracing make an agent safe?

No. Tracing improves visibility and diagnosis; it does not prevent an unsafe action by itself. Safety still depends on bounded capabilities, downstream authorization, suitable approval controls, validation, recovery design, and defense in depth.

12
🔗 References & Further Reading

The following primary sources are included for factual verification. The explanations, examples, tables, and pseudocode in this article are original teaching material.

  1. OpenAI Agents SDK
  2. OpenAI Agents SDK — Tracing
  3. OpenAI Agents SDK — Running Agents
  4. Anthropic — Demystifying evals for AI agents
  5. OWASP GenAI Security Project — Excessive Agency

Vendor and project names are trademarks or properties of their respective owners. 

13
📝 Summary

→ A final answer is not proof of completion. Verify the real state of the system.

→ Test the harness contract, not one magical transcript. Different valid paths can exist.

→ Use traces to reconstruct the run. Make important events, boundaries, and failures observable.

→ Debug the first broken invariant. The visible failure is often downstream of the real defect.

→ Treat security controls as testable behavior. Permissions, approvals, retries, and recovery belong in the evaluation suite.

→ Turn production failures into regression cases. A strong harness gets harder to break in the same way twice.

Final thought: The goal of harness testing is not to make an agent follow one predetermined script. The goal is to make its important behavior observable, verifiable, bounded, and recoverable even when the path taken by the model changes.


Comments