If your evaluation harness says Retrieval-Augmented Generation (RAG) makes your model hallucinate more, check your judge before you blame the generator. A developer maintaining django-explain-errors recently found that their eval pipeline was systematically penalizing RAG-on explanations for "fabricating" details that were actually correct, simply because the judge model couldn't see the source code that validated them.

The Setup: Grounding vs. Guessing

The project in question is a Django middleware that catches unhandled exceptions and asks an LLM to explain them. It operates in two modes: RAG-off, where the model sees only the traceback, and RAG-on, where it also sees relevant source code retrieved via a local sqlite-vec index. The premise for RAG-on is groundingβ€”if the model reads your actual code, it should stop guessing function names. However, the initial eval results contradicted this, showing RAG-on appearing to invent more details than RAG-off.

Why the Judge Was Wrong

The problem wasn't the generator; it was the judge's blind spot. In the first version of the eval, the judge (Claude Sonnet via OpenRouter) only saw the traceback and a short list of known facts. When RAG-on correctly identified a parameter like post_id from the source code, the judge flagged it as fabricated because that detail wasn't in the traceback. The metric was punishing RAG for doing exactly what it was supposed to do: retrieving specific, correct information that the judge lacked.

Iterating the Eval Harness

Simply giving the judge more context didn't work initially. When the developer provided full source modules, the judge ignored them, likely because scanning 300 lines of mostly irrelevant code is harder than marking a detail "unverified." The solution required changing the task structure. The judge was forced to list every checkable claim and provide a verdict (verified, contradicted, or absent) against a small, relevant source excerpt. With this change, the results flipped: RAG-on avoided fabrication in 26 comparisons versus RAG-off's 17 on app-code fixtures.

Hidden Bugs in the Pipeline

While debugging the fabrication metric, the developer discovered a second, critical bug. The middleware was truncating long tracebacks by keeping only the tail, which often cut out the specific frame naming the failing app function. This meant RAG-off wasn't seeing where the error happened, artificially inflating RAG-on's success in pointing to the fix location. Fixing the truncation logic improved the package for all users, not just the eval.

Key Takeaways

  • Explicitly define every metric. If a criterion matters, make the judge score it explicitly rather than letting it act as an unstated tiebreaker.
  • Force the judge to show its work. Claim-level verdicts computed in code are auditable; bare yes/no scores are not.
  • Inspect inputs, not just outputs. Every major fix in this story came from reading what the judge and generator actually saw, not just looking at the final score.
  • Suspect the instrument when results surprise you. "RAG fabricates more" was a finding about the eval harness, not the RAG system.

The Bottom Line

Your eval pipeline is part of your product. If you don't audit what your judge model can see, you're measuring its ignorance, not your model's accuracy.