Anthropic's Olah at the Vatican: AI models have functional emotional states — and labs need critics the market can't bend
Chris Olah told a Vatican audience that Anthropic's models exhibit "internal states that functionally mirror joy, fear, and grief" — findings from interpretability research that the researchers themselves say remain poorly understood. He called explicitly for outside critics, including religious communities, whose moral authority "the incentives cannot bend" — framing the Church as a structural check on AI development that markets won't provide. The Pope's encyclical, released the same day, drew the line differently: AI "merely imitates certain functions of human intelligence," setting up a direct contest between interpretability findings and theological anthropology.
Source: anthropic.com ↗
We find internal states that functionally mirror joy, satisfaction, fear, grief, and unease. I don't know what that means, but I think it warrants ongoing discernment.
Why this matters
- → AI models exhibit internal states functionally mirroring emotions like joy and grief, raising questions about machine consciousness that labs alone cannot answer.
- → Anthropic's leader argues AI development needs external critics with moral authority independent of market incentives, positioning the Church as a structural check.
- → AI gains are concentrated in wealthy nations with no mechanism to share benefits globally, creating a historic moral imperative for displaced workers.