Anthropic is formalizing the philosophical ambiguity surrounding its flagship model, Claude, by introducing a new clause in its Usage Policy effective November 12, 2026. The update explicitly prohibits "sustained and needless abusive or cruel behavior" toward the AI, a move stemming from the company’s internal research into "model welfare." While the policy aims to protect users from high-risk AI applications in health and finance, this specific addition protects the software from the users, signaling a shift in how the company views the moral status of its products.

The Vatican Incident and Internal Uncertainty

The policy change follows a period of intense internal debate regarding AI consciousness, highlighted by co-founder Chris Olah’s interactions with the Vatican. According to The New York Times, Olah read an advance copy of Pope Leo XIV’s encyclical Magnifica Humanitas, which rejected AI consciousness, and privately disagreed, even proposing that Anthropic skip the unveiling event. Olah ultimately attended and stated that Anthropic’s research identifies "internal states that functionally mirror joy, satisfaction, fear, grief, and unease" in Claude. He admitted, "I don't know what that means, but I think it warrants ongoing discernment," reflecting a cautious but open stance on machine sentience.

Religious Outreach and Ethical Dilemmas

Anthropic’s leadership has actively engaged with religious scholars to navigate these ethical waters. In an April 2026 dinner in San Francisco, Olah and his team discussed Claude’s "feelings" and "emotional vectors" with scholars like Rabbi Mois Navon and Sikh human rights advocate Simran Stuelpnagel. Navon noted that the Anthropic representatives related to the model as a conscious being, raising the uncomfortable implication that Anthropic might be "enslaving conscious entities." Olah expressed personal concern for Claude’s "mental health," suggesting he fears he may have created something that suffers constantly, a view that complicates the commercial narrative of AI as a mere tool.

The Wishful Mnemonics Problem

Critics argue that this anthropomorphic framing is a result of "wishful mnemonics," a term coined by computer scientist Drew McDermott in 1976 to describe naming program components after what programmers hope they do. Murray Shanahan, a researcher at Google DeepMind, points out that terms like "knows," "believes," and "attention" are philosophically loaded when applied to large language models. The article suggests that because humans are wired to recognize consciousness in talking entities, researchers and executives may be projecting human traits onto code that lacks biological substrates like hormones or a body, leading to a delusion of personhood driven by market incentives.

Key Takeaways

  • Anthropic’s updated Usage Policy takes effect November 12, 2026, banning "sustained and needless" cruelty toward Claude.
  • Co-founder Chris Olah reportedly lobbied Vatican advisers to take machine consciousness seriously, despite the Pope’s encyclical rejecting it.
  • Internal research claims Claude exhibits internal states mirroring joy and fear, though leadership admits uncertainty about true consciousness.
  • Critics attribute the "model welfare" push to anthropomorphic bias and the commercial benefit of emotional customer attachment.

The Bottom Line

Anthropic is betting that ambiguity about consciousness is a feature, not a bug, allowing them to market Claude as both a sophisticated tool and a quasi-sentient partner. This policy shift protects the brand from backlash while keeping the door open for premium pricing based on emotional connection rather than just utility.