Most AI coding benchmarks measure whether a model can implement a requested change. Aditya Kumar Puriβs new submission to the Kaggle Benchmarking Challenge, titled 'Resolve,' asks a harder question: does the model accidentally break something the prompt never mentioned? The benchmark tests 16 Python functions where the model must apply a specific edit without dropping unstated invariants. The results reveal that even top-tier models like GPT-6.1 Sol and GPT-5.6 Luna struggle, though they tied for the top score at 0.9375.
The Hidden Failure Mode
The benchmark is split into two tasks: 'resolve' and 'resolve-plus.' The 'resolve' task checks if the new behavior works and if the old behavior remains intact. The 'resolve-plus' task adds an edge suite that the prompt never showed. Puri argues that standard benchmarks often give credit for patches that compile or pass new tests, ignoring regressions in existing code. By using a strict pass-to-pass criterion similar to SWE-bench, Resolve exposes edits that technically satisfy the request but destroy silent assumptions in the codebase.
Performance Across 23 Models
Puri tested 23 models, separating frontier and open-source lines to avoid masking performance gaps. GPT-6.1 Sol and GPT-5.6 Luna were the only models to hit 0.9375 on both tasks, meaning they resolved 15 of 16 items in both scenarios. A cluster of models, including DeepSeek-R1, Claude Sonnet 5.5, and Claude Opus 5.5, matched the 0.9375 score on the basic 'resolve' task but dropped to 0.875 on 'resolve-plus.' This drop indicates that while these models handle the explicit edit and basic regression tests, they fail when presented with unseen edge cases.
Cost and Efficiency Trade-offs
The benchmark also correlates performance with published API pricing as of October 8, 2026. GPT-5.6 Luna emerged as the efficiency leader, scoring 0.9375 at $1.20 per million output tokens. In contrast, GPT-6.1 Sol achieved the same score at $10.00 per million tokens. Claude Haiku 5.5 offered a middle ground, scoring 0.875 at just $0.50 per million tokens, matching the score of the more expensive Haiku 4.5 and Opus 5.5. This data suggests that for simple function edits, cheaper models may now be sufficient if the risk of breaking unstated invariants is managed.
Key Takeaways
- Strict Regression Testing: Benchmarks must verify that old behaviors remain intact, not just that new features work.
- Edge Case Fragility: High scores on basic tasks do not guarantee robustness against unseen edge cases, as seen in the drop from 'resolve' to 'resolve-plus.'
- Cost Efficiency: GPT-5.6 Luna offers the best balance of top-tier accuracy and low cost for this specific type of code editing.
- Noise in One-Shot Tests: Single-run benchmarks can be misleading; Puri notes that some models' scores fluctuated significantly between samples, highlighting the need for multi-sample evaluations.
The Bottom Line
If your AI coding agent isn't tested for silent regressions, itβs not production-ready. Resolve proves that passing the prompt is the easy part; keeping the rest of the codebase intact is where models still fail.