An AI agent can produce a plausible answer, call the right tools, and still fail the task. That is why testing an agent harness cannot stop at checking whether the final text “looks good.” A harness needs ways to determine whether the requested outcome actually happened, whether the agent stayed within its authority, whether the run stopped correctly, and what happened between the first model call and the final result. Consider a fictional invoice-exception agent . A user asks it to inspect an invoice, compare it with a purchase order, and open an exception case if the values do not match. The model may correctly identify a mismatch but never create the case. Or it may create the case twice after a retry. Or it may create the case successfully but incorrectly report that the case was created before the tool confirms it. A final-answer check alone can miss all three problems. Engineering principle Agent testing should verify both the outcome and the path that produced it. ...