Grading LLM outputs manually is a bottleneck that kills iteration speed, but most automated judges are either too expensive or too lenient. A developer on DEV.to recently detailed a method for calibrating an LLM judge to match their personal grading standards for technical product datasheets, achieving a 25x cost reduction. The project involved testing four AI systems against 47 customer questions derived from 27 PDFs, including three scanned documents without text layers.
The Benchmark Results
The developer ran a custom pipeline using Claude Sonnet via API with page citations against the Claude app (Opus), the ChatGPT app, and Chatbase. The custom pipeline and the Claude app tied for the top spot, each answering 46 out of 47 questions correctly with zero hallucinations. ChatGPT and Chatbase lagged significantly, answering only 35 and 33 questions correctly respectively. Notably, no system invented a spec value, but the weaker models failed by incorrectly stating information was missing when it was present in the documents.
Where Models Fail on Technical Data
The analysis revealed specific failure modes in handling technical documentation. Scanned pages were often treated as blank by systems relying solely on text layers. Complex queries requiring cross-referencing between a product's advertised range and specific mode limitations caused errors. Unit conversion discrepancies also tripped up models, where converting a listed value led to confident but wrong answers. Additionally, conflicts between website marketing copy and official datasheets required a strict rule: the datasheet value always wins, even if it contradicts the web page.
Building a Calibrated Judge
To grade the 188 generated answers without spending hours on manual review, the developer built a hybrid judge. Initial attempts using a single LLM judge showed poor agreement with human grading, largely due to errors in the developer's own answer key. By introducing Cohen's kappa to measure agreement beyond chance, the developer refined the process. The final solution combined code-based checks for numerical accuracy with Jev, a System One model from TypeSafe that handles natural language interpretation and returns typed probabilities.
The Hybrid Architecture
The hybrid judge uses code to extract and match numbers, ensuring precision on specs like tolerance ranges. Jev handles the semantic grading, answering four specific questions per request: verdict, brush-off detection, unasked extra info, and rule adherence. This architecture costs approximately $0.03 to grade all 188 answers, compared to $0.80 for a Sonnet-only judge. The system achieved a high level of agreement with the human grader, with a kappa score of 0.85 after refining the policy to flag rather than penalize verbose but correct answers.
Key Takeaways
- Hybrid judges combining code checks and specialized LLMs offer superior cost-efficiency.
- Human answer keys often contain errors that must be corrected before calibrating an AI judge.
- Technical datasheets require strict rules for resolving conflicts between marketing pages and official specs.
- Cohen's kappa is essential for validating judge agreement beyond simple percentage accuracy.
The Bottom Line
Stop trusting raw LLM judges. If you want reliable grading, you need to calibrate against a human standard using statistical rigor, not just vibes. The developer's approach demonstrates that you don't need a frontier model to grade another frontier model. By offloading numerical checks to deterministic code and using a smaller, faster model for semantic nuance, you can build a judge that is both cheap and accurate. The real work wasn't in prompting the judge, but in auditing the gold standard and defining clear business rules for edge cases.