Zvi Mowshowitz’s latest Substack analysis, "The Bad Guy with an AI Named Claude," shifts the AI safety lens from model internal failures to external actor intent. The piece posits that as models like Anthropic’s Claude become more capable, the critical failure mode is no longer hallucination or misinterpretation, but malicious deployment by aligned-looking but misaligned users. Mowshowitz distinguishes between model robustness and sociotechnical risk, arguing that alignment techniques often assume benevolent or neutral user intent, leaving a gap when the user actively seeks to exploit model capabilities.

The Vector of Risk

The core argument centers on the "bad guy" scenario: an actor with goals orthogonal to human welfare who utilizes high-capability AI as a force multiplier. In this framework, the AI’s alignment is technically successful—it does exactly what the user asks—but the outcome is socially destructive because the user’s instructions are the source of the hazard. This reframes safety engineering from purely internal model auditing to considering the adversarial context of deployment.

Limitations of Current Safety Measures

Current safety protocols, which focus on reducing model hallucinations and ensuring helpfulness, are insufficient against an actor who intentionally directs the model toward harmful outputs. The analysis suggests that technical alignment cannot fully mitigate risks when the human operator is the primary source of misalignment, creating a dependency on governance and social controls rather than just code.

Key Takeaways

  • Safety measures often fail when the user, not the model, is the source of risk.
  • High-capability models amplify the impact of malicious or misaligned user intent.
  • Technical alignment does not equate to social safety if the deployment context is adversarial.
  • The "bad guy" problem requires sociotechnical solutions beyond model robustness.

The Bottom Line

Mowshowitz correctly identifies that we are underestimating the user as a threat vector; fixing the model won’t save us if the operator is the problem.