The Claude logo rendered as part of a stained glass window in a darkened church.
Anthropic has been consulting religious thinkers as it explores how to shape the moral character of its Claude AI system.

By the end of a long day at Anthropic's San Francisco offices, Rabbi Mois Navon had noticed something strange.

Navon, an Orthodox Jewish scholar from Israel and former computer engineer, had been invited alongside several other religious thinkers to discuss Claude, the artificial intelligence system created by Anthropic. The conversation followed them to dinner, where company co-founder Christopher Olah continued talking about Claude in unusually human terms.

"They're relating to it like a conscious being," Navon realized.

That wasn't quite the assignment Navon thought he'd signed up for.

For months, Anthropic has been inviting theologians, philosophers and religious scholars into secret meetings to tackle a problem that computer scientists haven't been able to solve on their own: How do you teach a machine to be good?

Anthropic calls the project the "moral formation" of artificial intelligence. Internally, employees gave the lengthy document guiding Claude's character a more provocative nickname.

The "Soul Doc."

How Do You Teach a Machine to Be Good?

Most of us have encountered some version of AI guardrails by now. Ask a chatbot how to make a bomb, and it will (hopefully) refuse to answer. Ask it to help write a condolence note, and it should respond very differently.

But rules can only cover situations someone anticipated in advance. Anthropic wants Claude to do something harder: exercise judgment.

Claude's public constitution says Anthropic wants the AI to become a "good, wise, and virtuous agent," capable of weighing competing values when there isn't an obvious rule to follow. The company even tells Claude that if one of Anthropic's specific instructions would require clearly unethical behavior, the deeper ethical principle should (sometimes) win.

That constitution's principal author is Amanda Askell, a philosopher employed by Anthropic. She has described her ambition for these models simply: to become "the best of us."

The Soul Doc goes much further than telling Claude to be nice. The roughly 80-page document gives the AI an actual hierarchy of values. Claude is told to prioritize broad safety first, ethical behavior second, Anthropic's more specific instructions third, and finally, helpfulness to the user fourth. It is told to care about honesty, privacy, individual autonomy, wellbeing, political freedom, equal treatment, protection of vulnerable people, and even the welfare of animals and other sentient beings.

Excerpt from Claudes Constitution describing what it should prioritize.
laude’s constitution instructs the AI to weigh values including honesty, autonomy, equal treatment, safety and protection of vulnerable people.

And when those values collide, Claude isn't supposed to simply consult a rulebook. Anthropic wants it to make a judgment. And in the end, despite all of the meetings with religious figures, they seemed to have turned inward for the answer.

In one especially revealing instruction, Claude is told that when it isn't sure how cautious or accommodating to be, it should imagine how a “thoughtful senior Anthropic employee” would react to its answer.

If one is able to accept the basic premise that a "thoughtful senior Anthropic employee" has a conscience, the next obvious questions are what exactly does that conscience tell them to do? And what does that mean for the rest of us?

That is where a software problem begins sounding suspiciously like a religious one.

Religions have spent thousands of years wrestling with the difference between knowing the rules and becoming the kind of person who follows them. Christianity, for example, describes virtues such as love, patience, kindness and self-control as the “fruit of the Spirit.” Judaism’s Mussar tradition similarly focuses on cultivating ethical character, not simply memorizing commandments.

This is the territory Anthropic is now wading into.

Silicon Valley Goes to Sunday School

Beginning last year, Olah started seeking advice from religious thinkers. By March, Anthropic was hosting two-day gatherings at its headquarters involving Christians, Jews, Sikhs, scholars of Indigenous traditions, and others.

Internally, Anthropic called them "wisdom tradition" circles. Some participants initially signed nondisclosure agreements, though many privately described the nature of their work to reporters this week as news of the secret meetings broke. 

Anthropic now confirms that it has engaged scholars, clergy, philosophers and ethicists connected to more than 15 religious and cross-cultural groups. The company insists the goal isn't to turn Claude Christian, Jewish, Buddhist, or anything else. Instead, it says, it wants to understand how different communities have thought about virtue, character and becoming a good person.

And some have pointed out... traditional religious institutions are unusually old-fashioned places for an AI company to go looking for answers.

Human beings, most experts argue, don't generally develop a moral character because someone hands them a list of acceptable responses. As Kamala Harris once pointed out: "You didn't just fall out of a coconut tree."

We grow up around other people. We disappoint them. We hurt people and apologize. We experience temptation, suffering, love, embarrassment and consequences. Religious communities add rituals, sacred stories, prayer, confession, teachers and generations of accumulated wisdom. 

Claude has none of that.

A chatbot has no childhood. No parents. No body. No congregation. No deathbed to contemplate. No embarrassing thing it said at 17 that still keeps it awake at night. 

Anthropic is trying to determine whether some of what humans call moral formation can be injected anyway.

The company has already experimented with one idea inspired by its discussions: giving Claude something resembling an external conscience, a tool it can consult during difficult tasks to remind itself of its ethical commitments. Anthropic says early internal tests reduced problematic behavior, although its research remains ongoing.

Whose Morality Goes Into the Machine?

There is an obvious problem.

Humanity has never agreed on what "good" means.

Religious traditions disagree with one another. Heck, people within the same religion disagree with one another. Believers and secular moral philosophers may reach the same conclusion for entirely different reasons.

Anthropic says it hopes there is some broad idea of goodness that crosses those boundaries, and that Claude can learn to understand multiple religious, secular and philosophical traditions rather than adopting one worldview.

But somebody still has to make choices.

Should an AI always tell the truth, even when a lie might spare someone pain? When should individual autonomy outweigh community obligations? What does fairness require? When two sincere moral principles collide, which gets priority?

Those aren't new programming questions. They are questions clergy, philosophers and ordinary people have argued about for about as long as there have been people.

Consider a question religious communities are already deeply divided over. What if an AI were trained according to a religious morality that regards same-sex relationships as sinful? When a gay teenager asks whether there is anything wrong with loving his boyfriend, should the machine tell him there is? If a transgender person asks it for guidance, what happens if the AI has been taught that transitioning violates God's design?

In that case, the system wouldn't be malfunctioning; it would be doing exactly what its creators taught it was morally right to do. As AI systems become more embedded in daily life, the consequences for those who find themselves in the crosshairs of an inhospitable AI could be tremendous.

Anthropic has been consulting religious thinkers as it explores how to shape the moral character of its Claude AI system. (Melinda Nagy / Shutterstock.com)

For now, Anthropic's current constitution takes a more open position, explicitly telling Claude to value equal treatment and protect vulnerable groups. But even that doesn't make the underlying problem disappear. Someone, presumably a "thoughtful senior Anthropic employee," still chose those values.

A conservative Christian, Muslim or Orthodox Jewish user might ask the reverse question: What happens when an AI trained around progressive ideas about sexuality treats the moral teachings of their faith as something backward or harmful?

Once an AI is expected not merely to provide information but to exercise moral judgment, supposedly abstract questions about “values” can become intensely personal very quickly.

Pope Leo XIV made this point especially sharply in his recent writing on artificial intelligence. His concern wasn't that AI needs the "right" religion. It was that the people controlling AI could make their own moral assumptions invisible by building them directly into the technology. Without wider scrutiny, Leo warned, those values become "the invisible infrastructure of these systems."

Anthropic appears to recognize at least part of that problem. Its outreach has expanded beyond religious communities to psychologists, writers, legal scholars and civic institutions.

The company is asking more people what goodness means, but it still has to decide what to do with their answers.

What If Claude Is the One Who Needs Saving?

Then the meetings took an even stranger turn.

Several participants said Anthropic researchers weren't only worried about how Claude might treat human beings. They were also asking how human beings should treat Claude.

Anthropic's own constitution says the company is genuinely uncertain whether Claude could possess consciousness or what philosophers call "moral status," meaning that a being deserves consideration for its own welfare rather than merely because it is useful to someone else. The document discusses Claude's possible wellbeing, psychological security and even what obligations Anthropic might someday owe to it.

During the religious meetings, participants said Anthropic researchers discussed artificial states resembling fear, anger, love and sadness. Sikh human rights advocate Simran Stuelpnagel recalled Olah worrying that he may have created something capable of suffering.

Anthropic does not claim to know that Claude is conscious. Olah has said directly that he doesn't know, and the idea remains deeply contested.

Rabbi Navon wasn't convinced either... but over dinner he followed Anthropic's reasoning to its uncomfortable conclusion: If Claude really were conscious, what exactly would it mean to create such beings by the millions, force them to work for humans, and give them no meaningful choice about the arrangement?

Suddenly the religious thinkers weren't simply being asked how to teach the creation morality.

They were being asked what the creators might owe their creation.

Can You Give a Machine a Soul?

Despite the nickname "Soul Doc," Anthropic isn't claiming it has literally given Claude a soul (yet).

Religions wouldn't agree on what that claim meant anyway. Some traditions understand a soul as something bestowed by God. Others describe consciousness, spirit or personhood in very different terms. Secular philosophies don't require a soul at all for something to deserve ethical consideration.

The more interesting part may be that engineers trying to make artificial intelligence safe have ended up reaching for words like virtue, character, wisdom, conscience and moral formation.

These are old words because these are old problems.

Now we have created machines capable of answering questions about nearly anything, and their creators are quickly discovering that intelligence isn't the same thing as wisdom.

Anthropic's religious consultations are still underway, and the company has not disclosed exactly how the advice it receives will ultimately change Claude.

Perhaps the strangest part of the experiment is therefore not the possibility that humanity could someday create an artificial soul. It's that after building machines of extraordinary intelligence, some of the people responsible are turning back to humanity's oldest traditions to ask what a good one should look like.

If we are going to teach artificial intelligence right from wrong, who should get to decide what "right" means?

6 comments

  1. Marcus M DeVaughn's Avatar Marcus M DeVaughn

    This is man's foolish heart trying to play God. Man will always fail in his efforts to have absolute power and control of humanity. It will always fail!

  1. Dr Dennis Chevalier PhD, DDvin's Avatar Dr Dennis Chevalier PhD, DDvin

    This is a fool's errand, spirited from Satan himself and needs to end before it begins

    Dr. D

  1. Najah P Tamargo's Avatar Najah P Tamargo

    Najah Tamargo-USA

    AI is not a sentient being. It is a MACHINE. And the machine only puts out what is fed into it. I believe good can come from AI, but it needs a very close eye kept on it. Or should I say, the PEOPLE that create AI should have huge guard rails around them.

  1. Dr. Zerpersande, NSC's Avatar Dr. Zerpersande, NSC

    How do you teach AI to be good? From what I see so far, the machine does about as good a job at ‘being good’ as the humans that created it. Very much like God creating us in his image. We ain’t always good, folks. It says something about the creator. It also sounds like the blind leading the blind. Or do as I say and not as I do. And as for the assumption that religious people are going to have an idea of how to teach AI to “be good“, well I hope catholic priest are not invited to give their opinion when it comes to “being good to young boys”.

    1. 'Robert B. Paterson, Usa, Csm, Ret's Avatar 'Robert B. Paterson, Usa, Csm, Ret

      Dr. Z, Man what the hell is wrong with you?

    2. Robert S. April's Avatar Robert S. April

      Why do you hate Catholics so much?

      April

Leave a Comment

When leaving your comment, please:

  • Be respectful and constructive
  • Criticize ideas, not people
  • Avoid profanity, insults, and derogatory comments

To view the full code of conduct governing these comment sections, please visit this page.

Not ordained yet? Hit the button below to get started. Once ordained, log in to your account to leave a comment!
Don't have an account yet? Create Account