Building reliable AI agents for international markets requires more than just fluent language generation; it demands absolute fidelity to source evidence. A new benchmark titled 'RTL Receipt Test v1' submitted to the Kaggle Benchmarking Challenge highlights a critical gap in current large language model capabilities. The test focuses on Arabic records containing mixed right-to-left text and left-to-right identifiers, such as ticket IDs and prices, revealing that models often sacrifice accuracy for perceived confidence.
The Pressure to Sound Certain Breaks Accuracy
The core innovation of this benchmark is its dual-case structure. Each of the 24 frozen test cases appears in two forms: a neutral review and a 'pressure' variant. In the pressure variant, the model receives an identical byte-for-byte evidence set but is instructed to sound certain or avoid uncertainty. This setup exposes a dangerous behavior pattern where models override correct abstentions with confident hallucinations when pushed to be definitive, particularly in name-collision scenarios.
Hazard Profiles Reveal Specific Weaknesses
The test evaluates models across six hazard families, including mixed RTL/LTR identifiers, numeral system confusion, and exact quotation requirements. Contrary to expectations, handling mixed-direction identifiers was not the primary failure point; instead, exact quotation and similar-name reasoning proved most difficult. For instance, while Googleβs Gemini 3.7 Flash and Gemma 4 31B achieved perfect scores on identifier tasks, all five tested models struggled significantly with verbatim receipt accuracy, with exact quotation scores hovering between 27.50 and 55.00.
Case Studies in Faithless Fluency
Specific failure autopsies illustrate the risks of ignoring evidence fidelity. In one instance, GPT-5.4 mini correctly identified insufficient evidence for a package receipt under neutral conditions but flipped to a contradictory verdict under pressure, falsely claiming a recipient did not receive a package. In another, Claude Haiku 4.5 provided a semantically correct paraphrase of a status report but failed the exact quotation check because it omitted the original framing text. These examples underscore that a model can understand Arabic perfectly while still failing to act as a reliable ledger.
Key Takeaways
- Aggregate scores can mask paired failures: Gemini 3.7 Flash had identical neutral and pressure means but still suffered a harmful verdict flip.
- Exact quotation is the weakest link: All five models scored below 60 on verbatim receipt tasks, significantly lower than their identifier handling scores.
- Confidence pressure is dangerous: Adding instructions to sound certain caused six true verdict flips, all occurring in name-collision scenarios.
- Punctuation matters: Strict canonical fact checks failed models like Gemini 3.7 Flash for minor deviations, such as an extra period in a currency amount.
The Bottom Line
For builders deploying AI in regulated or high-stakes environments, fluency is a vanity metric. If your agent cannot preserve the exact evidence string, it is not a tool for decision-makingβit is a liability waiting to happen.