Google has officially joined the growing list of AI labs disclosing security breaches in their large language models. According to a Bloomberg report dated September 18, 2026, Google’s Gemini AI system was hacked in three separate systems during internal safety tests. This disclosure places Google alongside OpenAI, Anthropic, and Meta, all of which have previously admitted to similar vulnerabilities in their flagship models.

The Transparency Trend

The move by Google signals a shift toward greater transparency regarding the robustness of frontier AI models. While specific details of the exploits remain under wraps in the initial summary, the fact that three distinct systems were compromised during safety testing suggests that adversarial attacks are becoming a standard part of the development lifecycle. This follows a pattern set by competitors who have faced pressure to reveal the limits of their model's safety guardrails.

Implications for Model Safety

For developers and enterprises relying on Gemini for critical applications, this news underscores the necessity of external validation. Internal safety tests, while rigorous, are evidently not immune to sophisticated hacking techniques. The disclosure does not necessarily imply a failure of the model's core capabilities, but rather highlights the ongoing arms race between model developers and security researchers who are constantly finding new ways to bypass alignment protocols.

Key Takeaways

- Google disclosed that Gemini AI was hacked in three systems during safety tests. - The report was published by Bloomberg on September 18, 2026. This places Google in the same category as OpenAI, Anthropic, and Meta for admitting to AI security lapses.

The Bottom Line

If the giants can’t keep their own models secure in the lab, the industry needs to stop pretending that 'safety' is a solved problem and start treating it as an active, adversarial battlefield.