Developers relying on Claude Haiku 4.5 for deterministic tasks may have silently experienced a behavioral shift earlier this month. A new community-driven monitoring project, aimodelrecord.com, reveals that the model's output consistency spiked from roughly 72β83% to 100% between October 7 and October 9, despite no official announcement from Anthropic. This phenomenon, described by developer Hiro Matsu as the model stopping wobbling, highlights the fragility of assuming model stability based solely on version IDs.
The Silent Shift in Determinism
Matsu's methodology involves querying Claude with 100 fixed questions twice daily, probing reasoning, instruction following, and structured output. For Haiku 4.5, which supports temperature 0, the goal was to catch non-deterministic behavior. From September 26 through October 5, the model produced identical answers for the same prompt only 72β83% of the time. However, starting October 7, this consistency jumped to 95%, and by October 9, it reached 100%. This suggests that while the weights remained fixed, the serving infrastructureβpotentially including sampling logic or safety classifiersβwas adjusted.
Infrastructure Updates vs. Weight Changes
Anthropic's documentation explicitly notes that while model weights are fixed for a given ID, the serving infrastructure can change. This update aligns with that caveat. Notably, October 7 also marked the launch of Claude Haiku 5.5 and the reclassification of Haiku 4.5 as a legacy model. While no causal link has been confirmed, the timing suggests that legacy model optimization or routing changes may have inadvertently stabilized output. For engineers caching temperature-0 outputs or running regression tests, this shift meant that approximately one-third of previously stable answers changed between October 8 and October 9.
Newer Models Remain Volatile
In contrast to the stabilized Haiku 4.5, newer models like Sonnet 5.5, Opus 5.5, and Haiku 5.5 exhibit significant variability, with same-day answer consistency ranging from 32% to 49%. This is largely due to adaptive thinking, a default feature in Claude 4.7+ models that introduces non-determinism even at default temperatures. The record shows that Haiku 4.5's 8 refusal rate on judgment calls remained constant, while newer models showed varied refusal patterns. This divergence underscores that while legacy models may achieve technical determinism, newer architectures prioritize nuanced reasoning over strict reproducibility.
Key Takeaways
- Haiku 4.5's output consistency improved from 72β83% to 100% between Oct 7 and Oct 9, 2026, without official notice.
- The change likely stems from infrastructure updates (sampling/router) rather than weight modifications, as model IDs remained unchanged.
- Developers using temperature 0 for caching or testing should be aware that legacy model updates can silently break determinism.
- Newer models (Sonnet 5.5, Opus 5.5) remain highly non-deterministic due to adaptive thinking features.
The Bottom Line
This incident serves as a critical reminder that model IDs are not immutable contracts for behavior. If you are building production systems that rely on Claude Haiku 4.5's determinism, you must implement your own output monitoring, because Anthropic's infrastructure can shift the goalposts without a changelog entry.