A new benchmark submission to the Kaggle Benchmarking Challenge has exposed significant instability in how Large Language Models handle API deprecation states. The study, authored by developer ajipelumi, tested six distinct models across five vendors, including Anthropic's Claude Sonnet 5, Google's Gemini 3.8 Flash, and OpenAI's GPT-5.4 Mini. The core finding is stark: on questions requiring nuanced knowledge of API lifecycles, six out of seven model runs contradicted themselves, providing different answers to the same query across three samples.
The Trap of Binary Deprecation
The benchmark deliberately avoided simple yes/no questions about whether a feature is deprecated. Instead, it focused on the critical middle ground where a feature is 'Deprecated but live'βannounced as ending but still functional. For example, Cloudflare's foundation_dns is deprecated but works until November 23, 2026. The study argues that collapsing this state with 'Removed' features hides two opposite errors: hallucinating code that no longer runs, or forcing unnecessary migrations for features that still work. By tracking announcement dates and end-of-life dates separately, the benchmark highlighted that models struggle most when the honest answer requires a qualifier like 'mostly' or 'except'.
Noise Floors and False Determinism
A crucial component of the study was a control run where Google's Gemini 2.5 Flash was tested twice under identical conditions. This revealed a significant 'noise floor' in the metrics, showing that even the same model can produce consistency scores varying by 0.074 and accuracy varying by 0.115 purely due to non-determinism. The author discovered that the Kaggle Benchmarks SDK, despite appearing to support temperature settings, often defaults to support_temperature: False, causing providers to ignore the requested temperature and apply their own defaults. This means many existing benchmarks may be measuring random variance rather than true model capability.
Instability Is Not Ignorance
Contrary to the assumption that models fail because they lack training data, the study found that instability is not simply a function of ignorance. GPT-5.4 Mini recognized every item in the test set (a cutoff rate of 0.0) but was the least consistent model, with an instability rate of 0.846. Meanwhile, the most consistent models still answered 31% of questions differently across samples. The benchmark suggests that partial truthsβwhere a feature is current for some operations but removed for othersβare the primary destabilizers. For instance, Cloudflare's DNS API allows updating record names and TTLs but prohibits changing record types since June 2026, a nuance that caused six of seven models to flip-flop.
Key Takeaways
- Non-Determinism is the Norm: Even with
temperature=0requested, LLMs are not deterministic. Benchmarks must use multiple samples and modal answers to be valid. - Nuance Breaks Models: Models fail most on 'deprecated but live' features where the answer is conditional, not binary.
- Recognition β Reliability: A model can perfectly recall a feature's existence (like GPT-5.4 Mini) and still be wildly inconsistent in its usage advice.
- Benchmark Design Matters: Without a control run (duplicating a model), it is impossible to distinguish true performance differences from statistical noise.
The Bottom Line
Current LLM benchmarks are dangerously flawed if they assume determinism; without control runs to establish a noise floor, we are measuring random variance rather than true capability. Developers must treat model consistency on nuanced API states as a bug, not a feature, and demand multiple samples before trusting any single answer.