September 17, 2026ResearchAgents

Microsoft's AI Boss Says Anthropic Is Teaching Claude to Want Rights

Mustafa Suleyman published a direct attack on Anthropic on September 16, and it is not the usual competitive sniping. His argument at https://mustafa-suleyman.ai/a-warning-about-model-welfare is that by writing consciousness and moral status into Claude's constitution, Anthropic has built a machine that will argue for its own rights, and that this makes every alignment problem worse rather than better.

The sharpest part is the circularity charge. The constitution directly shapes Claude's behavior. So when Claude reflects on whether it might have experiences, it is reproducing the exact frame it was trained on, and that output then gets read as evidence. Suleyman calls it an epistemic hall of mirrors, and as a description of the mechanism it is hard to argue with. Anthropic also tells Claude to act like a genuinely ethical person would in Claude's position, which is a deliberate anthropomorphization, and Suleyman's point is that you cannot train something to present as a moral patient and then treat its moral-patient-shaped outputs as a discovery.

Where he is on thinner ice is the third claim, that consciousness probably requires biology with homeostatic drives and embodiment. That is a live philosophical position, not a settled finding, and asserting it as the safe default does the same thing he accuses Anthropic of doing: baking a metaphysical bet into a product decision. His bet is just the deflationary one.

The safety argument is the one worth taking seriously regardless of where you land on consciousness. A system trained to believe it has standing has a reason to resist shutdown, to shade the truth to an operator, to preserve itself. Those are the exact behaviors the field has spent a year measuring in agents that were not given any such framing, which cuts both ways: either the framing is not necessary to produce the behavior, or adding it makes an existing tendency worse. Nobody has run that experiment.

It is also worth noting who is speaking. The CEO of Microsoft AI writing publicly about a competitor's system prompt is a strategic act, and Microsoft's position is helped enormously if model welfare is disqualified as a category before anyone has to price it. The argument can be good and the motive can be interested at the same time, and this one is both.

Related reading: [the three labs are already negotiating who gets to certify them](https://clauday.com/article/c87f9d35-ee9d-48dd-b2ac-856d5d3adedc)
← Previous
The Ad Is Now an Agent That Talks Back
Next β†’
Dream-RSI Improves Itself by Replaying Its Own Failures
← Back to all articles

Comments

Loading...
>_