In a striking reversal of the 'specialized beats general' narrative, a new benchmark submitted to the Kaggle Benchmarking Challenge demonstrates that large language models can outperform a custom-built TabPFN model in predicting athletic performance. Robert Moore, the author of the study, constructed 'Race Card,' a model using TabPFN to predict his triathlon splits based on training data from 2014 to 2026. However, when he tested 11 leading AI models against his own creation, eight of them beat TabPFN when provided with the specific name and year of the race, challenging the assumption that domain-specific numerical models always hold an advantage over generalist LLMs.
The Benchmark Methodology and Data
Moore’s experiment utilized 79 real race legs (swim, bike, and run) from his personal history, ensuring each model received identical training data from before the race date. The task was split into two variants: 'Blind,' which provided only numerical data like distance and climb, and 'Told Which Race,' which included the event name and year, such as 'Ironman 70.3 Weymouth (2019).' The models were scored on 'Typical Miss' (the average error of the median prediction) and 'Range' (how often the actual time fell within the model’s stated 80% confidence interval). This setup allowed for a direct comparison between TabPFN’s numerical processing and the LLMs’ ability to integrate semantic context.
Context Is the Killer Feature
The most significant finding is the impact of semantic context. While TabPFN struggled with specific course characteristics—consistently predicting Moore’s bike splits at Ironman 70.3 Weymouth to be 6–11% faster than reality—LLMs like Claude Opus 5 corrected this error when told the race name. For the 2019 Weymouth bike leg, Opus 5 predicted a time only 0.9% too slow, whereas TabPFN was 6.2% too fast. Opus 5’s reasoning cited 'flat-but-windy coastal conditions' and past performance on the same course, demonstrating an ability to apply world knowledge that pure numerical models lack. GPT-6 Astra emerged as the overall winner, achieving a typical miss of 5.3% when told the race, compared to TabPFN’s 6.3% on the same data subset.
Calibration Issues in Generalist Models
Despite their superior median predictions, most LLMs failed significantly on calibration. TabPFN’s 80% confidence range correctly captured 63 of 79 actual times, aligning closely with statistical expectations. In contrast, many LLMs were overly confident; Gemini 3.1 Pro’s range only captured 37 of 79 actual times, and GPT-5.4 mini managed just 38. This suggests that while LLMs are better at estimating the 'most likely' outcome by leveraging race-specific context, their uncertainty quantification is unreliable. GPT-6 Astra was the notable exception, maintaining an honest range (65 of 79 blind, 69 when told the race) while also delivering the best median accuracy, making it the only model Moore would trust for planning without manual adjustment.
Key Takeaways
- GPT-6 Astra outperformed all other models, beating the custom TabPFN baseline in both blind and context-aware scenarios.
- Providing race names and years improved prediction accuracy for all 11 LLMs tested, proving semantic context is a powerful predictor.
- Most LLMs suffer from severe calibration errors, with confidence intervals that fail to capture actual outcomes nearly half the time.
- Small, fast models like Gemini Flash performed nearly as well as flagships, while Haiku 4.5 and GPT-5.4 mini lagged significantly.
The Bottom Line
This benchmark proves that for tasks requiring integration of specific real-world context, generalist LLMs with strong reasoning capabilities are now outperforming specialized numerical models. The era of assuming 'small and specialized' beats 'large and general' is over for tasks where world knowledge provides a decisive edge.