Skip to main content

Posts

Showing posts with the label Harness Engineering

Testing an AI Agent Harness: Completion Checks, Tracing, and Debugging

An AI agent can produce a plausible answer, call the right tools, and still fail the task.   That is why testing an agent harness cannot stop at checking whether the final text “looks good.” A harness needs ways to determine whether the requested outcome actually happened, whether the agent stayed within its authority, whether the run stopped correctly, and what happened between the first model call and the final result. Consider a fictional invoice-exception agent . A user asks it to inspect an invoice, compare it with a purchase order, and open an exception case if the values do not match. The model may correctly identify a mismatch but never create the case. Or it may create the case twice after a retry. Or it may create the case successfully but incorrectly report that the case was created before the tool confirms it. A final-answer check alone can miss all three problems. Engineering principle Agent testing should verify both the outcome and the path that produced it. ...

Running Agent Harnesses in Production: Reliability, Cost, and Long-Running Tasks

Running an AI agent in production is not simply a matter of keeping the model call alive. The real engineering problem is keeping the entire run understandable, bounded, recoverable, observable, and safe when the work takes longer than one request, one process, one context window, or one human interaction. Imagine a fictional enterprise agent asked at 9:00 AM to review a large batch of supplier invoices, compare them with purchase records, identify exceptions, prepare a report, and wait for a finance reviewer before sending any consequential update. At 9:18 AM the browser disconnects. At 9:27 AM the worker process restarts. At 10:05 AM an external service times out. At 10:40 AM a reviewer approves one action but rejects another. A demo agent might simply start over. A production harness should know what work was completed, what remains uncertain, which actions were already attempted, where durable state lives, which operations may be retried, which actions require approval, and ...