In a diagnostic benchmark titled 'Reward Evidence,' GPT-5.4 nano demonstrated a peculiar failure mode: perfect financial tracking coupled with flawed opportunity assessment. Across 40 synthetic paid-work cases, the model correctly extracted every cash receipt but missed nine availability labels. This discrepancy highlights a critical gap in automated work assistants, where distinguishing between an offer, eligibility, and actual payment delivery requires more than just parsing dollar signs. The study, developed with OpenAI Codex assistance and published on DEV.to, serves as a warning that a clean money ledger does not equate to a correct decision on whether to pursue a gig.
Benchmark Design and Methodology
The Reward Evidence dataset consists of 40 original synthetic cases organized into 20 matched pairs, designed to test seven specific extraction fields including reward form, currency, and confirmed cash receipts. The evaluation framework distinguishes between cash and service credits, prize pools and individual awards, and pending versus delivered payments. To ensure rigor, the author implemented a strict grader that requires exact correctness on all seven fields and complete evidence sets. The benchmark was run using the Kaggle Benchmarks SDK, with temperature set to 0 for deterministic outputs, though the author noted that effective provider settings were not independently verified.
Model Performance Breakdown
The results revealed a stark hierarchy among the tested models. Gemini 3.7 Flash achieved a perfect score of 40/40 on the grounded task, maintaining 100% accuracy on both availability and cash receipt fields. In contrast, GPT-5.4 nano scored 25/40 on exact correctness, with its primary failure stemming from nine availability mismatches despite having a perfect cash receipt score. gpt-oss-20b performed better than nano with 30/40 exact correctness, though it suffered from two invalid schema responses. Notably, all 120 requests across the three models completed with zero request errors, indicating that infrastructure stability was not the bottleneck.
Critical Failure Modes
The specific errors made by GPT-5.4 nano illustrate the nuance required for paid-work automation. In one instance, the model correctly recorded a delivered USD 60 award as 6000 cents but mislabeled the still-open program as 'CLAIMED,' failing to recognize that the amount received did not establish an exclusive assignment. Another case involved an explicitly eligible, funded task with a future deadline that was incorrectly labeled as 'ELIGIBILITY_REQUIRED.' These errors demonstrate that while the model could parse the financial magnitude, it failed to interpret the temporal and contractual status of the opportunity.
Rubric Integrity and Corrections
The author disclosed a necessary rubric correction during the pilot phase, which initially yielded a 37/40 score for both prompts. Review revealed two unsupported gold labels where offers without selection terms were incorrectly marked as accepted, and a superseded announcement’s terms were carried over. The dataset was frozen as v2 with these corrections, preserving the original pilot artifacts for transparency. This self-correction process underscores the importance of evidence discipline for benchmark authors, not just the models being tested. The final v2 dataset SHA256 hash is 08e5fe18c05520bf1d6e18316c01641fe209822e34d1ccb1a2bb4ced0425f2bd.
Key Takeaways
- GPT-5.4 nano achieved 100% accuracy on cash receipt amounts but failed 9 availability labels in 40 cases.
- Gemini 3.7 Flash outperformed other models with a perfect 40/40 score on the grounded task.
- The benchmark distinguishes between advertised terms, eligibility, and delivered payments as separate fields.
- Invalid schemas for gpt-oss-20b resulted in zero field credit, highlighting the strictness of the grader.
- The study is a small diagnostic using synthetic data, not a general ranking of model capabilities.
The Bottom Line
If you’re building a paid-work agent, don’t trust the ledger alone; you need separate logic to verify availability and eligibility, or your bot will happily claim money for jobs it can’t actually do.