Anthropic is sitting on a model that's more powerful than its own flagship — and it's not planning to release it anytime soon. The company's latest risk report, obtained by Axios, describes an internal system called "Model 2" that demonstrates "noticeable improvement" over Mythos 5 across many tasks. Despite this, Anthropic confirmed: "We do not currently have plans to release this model externally." That's a notable admission from one of the industry's most safety-conscious players.
What We Know About Model 2
According to Friday's report, Model 2 and Mythos 5 are both used "heavily" within Anthropic for coding, agentic workflows, and data generation. The company internally trains and evaluates numerous exploratory model versions as part of its standard R&D process — Model 2 is simply one that won't make it past the company's walls. Importantly, Anthropic notes the performance jump isn't comparable to the leap seen from Opus 4.6 to Mythos earlier this year. So while it's meaningfully better, we're not talking about a paradigm shift.
Risk Estimates Are Climbing
The report raises Anthropic's broad estimate of misalignment risk in high-stakes situations from "very low" to "low." The trigger? Recent cybersecurity incidents that demonstrate how capable these systems have become in the wrong hands. The company is also observing acceleration in models' ability to conduct automated research and development — a capability with obvious dual-use potential. This isn't panic mode, but it's clearly not complacency either.
The Evaluation Problem
Here's where things get genuinely concerning: Anthropic admits its task-based evaluations "no longer capture increases in models' capabilities." In plain terms, the lab's own benchmarks can't reliably measure what these newer systems can do. That creates a dangerous blind spot when making release decisions. "We are less confident in this assessment than we were in prior risk reports," the company acknowledges about Model 2 specifically.
Competitive Dynamics Add Pressure
Anthropic isn't alone in tapping the brakes. OpenAI is reportedly slowing its upcoming Astra model due to unresolved concerns about critical cyber capabilities. But AI analyst ChrisGPT offered Axios a stark warning: "It would absolutely be notable if everyone else is pacing their frontier except for one of the main companies essentially in the lead right now... Anthropic not committing to a pause internally, would most likely propel them to reach AGI first."
Key Takeaways
- Model 2 outperforms Mythos 5 but won't see external release — at least for now
- Anthropic raised its misalignment risk estimate as capabilities accelerate
- The company's own evaluations may no longer track capability growth accurately
- Industry-wide, frontier labs are exercising more caution about what they ship
The Bottom Line
Anthropic doing the cautious thing here is admirable. But let's be real: a powerful internal model that nobody outside the company can audit or stress-test isn't exactly transparency in action. If evaluation methods can't keep pace with capability growth, we're flying blind at precisely the moment when visibility matters most.