A new submission to the Kaggle Benchmarking Challenge has revealed a startling uniformity in large language model reasoning. The benchmark, titled Liquidity Event Identification, tested whether LLMs could correctly distinguish between subtle market structure events like liquidity sweeps, rejections, and breakouts. The result? Every single model that successfully completed the evaluation achieved a perfect 100% accuracy score. This outcome challenges the assumption that complex financial reasoning remains a significant bottleneck for current state-of-the-art models.

The Benchmark Design

Created by developer iamclassified, the benchmark focused on seven controlled price-action scenarios designed to isolate reasoning capabilities from trading profitability. The scenarios covered specific technical setups, including equal-high and equal-low liquidity sweeps, breakouts followed by acceptance, and brief breaks followed by rejections. Crucially, the test included near-identical scenarios where a minor detail changed the correct interpretation, as well as cases with conflicting multi-timeframe evidence and deliberately misleading explanations. The goal was not to see if an AI could make money, but to measure its ability to parse market logic.

Model Performance and Failures

The evaluation included 20 different models available through Kaggle, spanning a wide range of providers and capability levels. The list included heavyweights like GPT-6 Astra, GPT-5.6 Sol, Claude Opus 5, and Gemini 3.7 Flash, alongside smaller or open-weight models like Qwen3 Next 80B and DeepSeek R1. Out of the 20 runs, 16 models completed the benchmark successfully, while four encountered execution errors. The failing models were Qwen3 Next 80B, GPT-OSS 120B, DeepSeek R1, and Claude Sonnet 4.6. These failures were technical execution issues rather than reasoning errors, so they were excluded from the accuracy calculation.

Implications for LLM Evaluation

With 112 out of 112 completed scenario evaluations marked correct, the benchmark failed to differentiate between model tiers. The author argues that this perfect score is not evidence of superior trading intelligence, but rather proof that the current scenario set is insufficiently challenging. The subtle distinctions in liquidity-based price action, which are difficult for human traders to master, appear trivial for modern LLMs in a controlled text environment. This suggests that current financial reasoning benchmarks may be saturating quickly, requiring more ambiguous and complex test cases to reveal actual performance gaps.

Key Takeaways

  • All 16 successfully executed model runs achieved 100% accuracy on the Liquidity Event Identification benchmark.
  • Four models failed due to execution errors, not reasoning failures, highlighting infrastructure stability issues on Kaggle.
  • The benchmark included advanced models like GPT-6 Astra and Claude Opus 5, showing top-tier consistency.
  • The author plans to expand the scenario set to include stronger multi-timeframe conflicts and misleading contexts.
  • Perfect scores indicate that current controlled financial reasoning tests are too easy to differentiate model capabilities.

The Bottom Line

The 100% accuracy rate isn't a triumph for AI financial literacy; it's a failure of the test design. If every model from GPT-6 to Grok scores perfectly on liquidity reasoning, the benchmark is measuring reading comprehension, not market intuition. We need benchmarks that break models, not ones that let them coast.