Hex engineers have demonstrated that rigorous evaluation frameworks can dramatically reduce error rates in generative AI features, cutting wrong-edit rates from 21% to just 3%. The team achieved this 7x improvement by implementing a hill-climbing development process for their new "Quick Edits" feature, which allows users to restyle charts using smaller, cheaper models like GPT-6 Luna or Claude Haiku 4.5. This approach moves beyond using evals as a final QA gate, integrating them into the core development loop to drive iterative improvements.

The Validator Beat the Prompt

A surprising finding from Hex’s process is that the majority of performance gains came from improving the harness around the model, specifically the validator code, rather than tweaking the system prompt. Over 50% of the hill-climb improvements were driven by expanding the validator to accept "almost-right" outputs, such as fixing missing quotation marks or rounding numbers to allowed values. This shift reduced hard errors by 99% for Haiku and increased pass rates by 10 percentage points, proving that robust post-processing logic is often more effective than complex prompt engineering.

Building a Massive Eval Suite

To support this rigorous testing, Hex constructed a suite of 1,800 eval cases derived from real user requests and extended by LLMs like GPT-6 Astra and Claude Fable 5.1. The team intentionally included ambiguous and difficult cases, such as "make it pop" or requests involving data changes disguised as style changes, to ensure the model could correctly identify when to hand off tasks to the main Hex agent. Crucially, they maintained a 40% holdout set of cases that were hidden from the coding agents during development to prevent overfitting and ensure generalizability.

Metrics That Matter for User Trust

The team prioritized minimizing "wrong edits"β€”where the model changes a chart in an unrequested wayβ€”over minimizing "incorrect handoffs," where a simple task is escalated unnecessarily. They accepted a slightly lower overall pass rate with GPT-6 Luna because it made 20% fewer wrong edits at half the cost compared to GPT-5.6 Luna, recognizing that user trust erodes quickly when AI makes unexpected changes. This strategic decision highlights the importance of aligning eval metrics with actual product experience and user sentiment rather than chasing raw accuracy scores.

Key Takeaways

  • Validators and harness logic drove more reliability gains than prompt tweaks.
  • A 40% holdout set was essential to prevent overfitting during hill-climbing.
  • Minimizing wrong edits was prioritized over minimizing incorrect handoffs.
  • Running 10 attempts per case was the sweet spot for reducing evaluation noise.

The Bottom Line

Stop obsessing over prompt engineering. The real leverage in building reliable AI features lies in robust validators and harnesses that gracefully handle near-misses, not in trying to force the model to be perfect every time.