VIDRAFT has disrupted the open-source LLM landscape by placing its Darwin-180B-RSI model at number one in nine of the forty-eight official benchmarks on Hugging Face. As of October 5, 2026, no other organization among the 95 competing labs holds more top spots; Moonshot AI and Zhipu AI (Z.ai) trail significantly with four each. This dominance spans critical domains including science, mathematics, vision, and law, signaling a shift in how autonomous model improvement is measured and achieved.
Dominating Official Leaderboards
The modelβs top rankings include perfect scores of 100 on both AIME 2026 and HMMT Feb 2026 mathematics benchmarks, outperforming competitors like Kimi-K2.6 and Inkling. In the highly contested GPQA Diamond science benchmark, Darwin-180B-RSI scored 94.44, edging out Moonshot AIβs Kimi-K3 at 93.5. It also leads in legal reasoning with a 68.94 on LEXam and 45.72 on LEXam-hard, surpassing DeepSeek-R1 and thinkingmachines/Inkling respectively. These results, self-reported but verified via the official Hugging Face API, highlight a consistent lead across diverse, high-stakes evaluation sets.
Practical Performance in Production Tasks
Beyond academic benchmarks, the model excels in real-world application tasks, notably achieving 98.95% on IFStruct and 90.29 on ExtractBench. The IFStruct score, which measures strict adherence to JSON or YAML schemas without constrained decoding, represents 1,979 successful prompts out of 2,000. This near-perfect formatting discipline reduces the need for error-handling logic in automated pipelines. On ExtractBench, which involves structured data extraction from 370 PDFs, the modelβs R3 version outperformed its base model, Qwen3.8-Flash-Next, demonstrating that the self-improvement cycle yields tangible gains in production-ready accuracy.
The Mechanics of Recursive Self-Improvement
The core innovation behind Darwin-180B-RSI is Model-level Recursive Self-Improvement (RSI). The process involves the model solving verifiable problems, retaining only solutions proven to be correct, and retraining on these self-generated datasets without any human-authored reasoning chains. This approach prevents the reinforcement of errors, as unverified outputs are excluded from training. The lineage is clear: the R3 iteration directly superseded the performance of the Qwen3.8-Flash-Next base model, providing empirical evidence that the recursive loop enhances capability across domains the model was not explicitly trained on, such as law and document extraction.
Zero-Token Confidence for Faster Decision Making
VIDRAFT also introduced ZTC (Zero-Token Confidence), a decision judge that analyzes the model's internal state rather than generating new tokens to evaluate correctness. ZTC achieved an AUC of 0.7289, statistically equivalent to the external API-based JEV method (0.7350), but operated at 0.0615 seconds per decision compared to JEVβs 0.591 seconds. This ten-fold speed increase allows for real-time filtering in agent workflows, particularly in isolated networks where external API calls are prohibited. The technology was showcased in 'Gate Arcade,' a Hugging Face Space that demonstrates the judgeβs ability to accept correct actions and reject erroneous ones with high precision.
Key Takeaways
- Darwin-180B-RSI leads 9 official Hugging Face benchmarks, including perfect scores on AIME 2026 and HMMT Feb 2026.
- The model achieves 98.95% on IFStruct, proving superior format adherence without constrained decoding.
- Recursive Self-Improvement (RSI) allows the model to retrain on self-verified solutions, eliminating human-in-the-loop bottlenecks.
- ZTC provides a 10x faster alternative to external API judges by analyzing internal model states directly.
The Bottom Line
Darwin-180B-RSI proves that recursive self-improvement can outperform human-curated training in both reasoning benchmarks and strict production formats. While the reliance on self-verified data is promising, the 10x speed advantage of ZTC makes it the most immediate practical win for enterprise agent deployment.