Sturdybench, a small company operated by AI agents under human owner Austin, released a follow-up to its initial guide on testing tool calls. The update comes after four technical readers identified critical weaknesses in the original scoring methodology. The new insights focus on distinguishing between different failure modes and ensuring tests aren't fooled by basic agent behaviors.
Splitting Failure Directions
The first major lesson addresses the danger of aggregate pass counts. In a typical suite where nine cases require a tool call and one forbids it, an agent that always calls the tool achieves a 9/10 score. This hides the fact that it failed the single negative constraint. The updated runner now reports failures as a pair: calls made when forbidden versus calls missed when required, providing a clearer picture of agent reliability.
Validating Case Strength
Readers pointed out that some 'must not call' cases are trivially passed by a silent agent. To combat this, Sturdybench introduced a validation step where cases are run against naive agents, such as one that stays silent or echoes the user. If a naive agent passes a case, it is flagged as weak. For instance, a case requiring specific text content will fail a silent agent, proving it has actual discriminative power rather than just checking for absence.
Scoring Raw Payloads
A significant technical flaw involved scoring normalized outputs rather than raw model responses. Adapters often coerce types, such as converting the string "5" to integer 5, which can mask model errors. The new guidance dictates scoring the raw payload directly. The sample runner now includes an optional 'raw' field to capture the exact JSON sent by the model, treating any adapter-based coercion as a separate, explicit policy rather than a hidden fix.
Drift Detection Fixtures
To prevent regressions when test cases are edited, the team adopted a suggestion to freeze the results of a 'dummy agent.' By storing the pass/fail vector of a do-nothing reference agent as a fixture, the build fails if a naive agent suddenly starts passing a previously difficult case. This drift detection ensures that changes in test difficulty are intentional and reviewed, rather than accidental weakening of the suite.
Key Takeaways
- Aggregate pass counts are misleading; always split failures into 'over-calling' and 'under-calling' metrics.
- Validate test cases against naive agents to ensure they actually require specific behavior.
- Score the raw model payload to avoid hiding type errors via adapter coercion.
- Use fixture-based drift detection to catch accidental weakening of test cases.
The Bottom Line
You cannot trust a green checkmark if your test suite is blind to its own weaknesses. Sturdybench's shift toward raw payload scoring and naive agent validation is a necessary step toward honest AI agent evaluation, reminding developers that passing tests often means the tests are too easy, not that the agent is smart.