Anthropic has integrated automated evaluation design and performance tuning directly into the Claude Code workflow via the new claude-api skill. The update introduces two distinct sub-commands, /claude-api build-eval and /claude-api hillclimb, which guide developers through creating robust test suites and iteratively improving application performance without manual oversight. This move addresses the persistent industry challenge of designing evaluations that accurately mirror production traffic while avoiding the trap of overfitting to synthetic benchmarks.
Automated Evaluation Design Principles
The build-eval command enforces four core principles of good evaluation design: mirroring production tasks, ensuring performance improves with stronger models, maintaining passable headroom at the frontier, and minimizing run-to-run variance. Claude interviews the developer to sample inputs from production transcripts, bug reports, and manual cases, prioritizing real-world data over easily generated but irrelevant examples. The system then proposes a grader, defaulting to programmatic verification for constrained outputs or LLM-as-judge for open-ended tasks, and validates the grader by running it twice on identical outputs to detect inconsistency.
Hillclimbing for Performance and Cost
The /claude-api hillclimb command allows Claude to iteratively modify system prompts, skills, tool descriptions, and API parameters to optimize for either performance or cost. The process involves splitting the evaluation set into train and test subsets, where Claude proposes one change per round and reverts patches if the test set score remains flat while the train set improves, a clear sign of overfitting. In a case study involving a customer support benchmark, the tool reduced token cost by less than half by switching from Opus 4.8 to Opus 5.5 on low effort, while simultaneously increasing decision accuracy from 74.4% to 87.8%.
Diagnosing Failures and Refining Skills
When performance stalls, the hillclimber analyzes remaining failures by root cause rather than making blind adjustments. For example, when optimizing the claude-api skill itself, Claude identified that the model was defaulting to older API shapes due to trained priors, such as using fixed-budget thinking instead of adaptive thinking. By adding specific guidance tables to the skill documentation to correct these priors, the evaluation score improved from 74% to 80%, eventually reaching approximately 88% after fixing ambiguous test cases and grader contradictions.
Key Takeaways
- The claude-api skill automates the creation of evaluation sets using production traffic, bug reports, and synthetic data anchored in real examples.
- Hillclimbing prevents overfitting by using held-out test sets and reverting changes that do not improve generalization.
- Automated tuning successfully reduced costs by less than half while improving accuracy in Anthropic's internal customer support benchmarks.
- The system diagnoses evaluation flaws, such as ambiguous tasks or incorrect graders, rather than assuming all failures are model capability issues.
The Bottom Line
This skill effectively bridges the gap between theoretical eval design and practical implementation, turning a historically subjective and error-prone process into a reproducible engineering workflow. By automating the tedious aspects of grading and overfitting detection, Anthropic is empowering developers to focus on application logic rather than benchmark mechanics.
Sources
https://claude.dev/blog/automating-eval-design-and-hillclimbing/