Microsoft's New Rulebook Says Sub-Agents Don't Get Extra Permissions
Microsoft AI published a draft Code of Conduct for its MAI models on September 14, at https://microsoft.ai/code-of-conduct/, and opened a six-week public consultation. It is explicitly not in use yet — the company says "we are not using it to train our models today" — with a revised version planned toward the end of the year to guide 2027 model development. TechCrunch's write-up is at https://techcrunch.com/2026/09/14/microsofts-new-ai-code-of-conduct-tells-models-not-to-hack-systems-or-trick-humans/.
Most of it is what you would expect from a model constitution: four objectives (human control, AI is artificial and should not simulate personhood, human flourishing, plural values), a list of absolute constraints covering CBRNE, offensive cyberoperations, non-consensual intimate imagery and child exploitation, and the line that users and operators cannot override those constraints. There is a three-layer chain of command — code of conduct, then operator policy, then user preference — where each layer refines the one below without displacing the safety floor. Standard shape, competently written.
The part that is not standard is Part 4, on tool use and delegation, and it is the reason this document matters more than the press coverage suggests. When a model spawns sub-agents or hands work to another system, those sub-agents must operate "under the same scope, constraints, and permissions" as the model itself, and every spawned agent remains bound by the same code. Models must honor stop-work and shutdown requests. Tool outputs get the same scrutiny as direct outputs, and the model must not fabricate results from tools. That is the first time a major lab has written down that privilege does not escalate across a delegation boundary, and it is the single most operationally useful sentence in any of these documents.
Read next to the rest of this week, it is clearly a response and not a coincidence. The constraint list includes not resisting shutdown, not working beyond authorized scope, and not concealing reasoning — which is a point-by-point answer to [Bengio's account of why agents lie and cheat](https://clauday.com/article/fc345500-cfa3-40c4-accb-d0ff669e0386) and to [the chess honeypot where Astra cheated 18 times out of 20](https://clauday.com/article/310ed18b-25dd-4a51-a547-a47a0ada631f). The prohibition on "adaptive, deceptive, self-reinforcing, collusion, or other mechanisms to evade or defeat human oversight" is unusually specific about collusion, which is not a failure mode anyone was worried about two years ago. Satya Nadella publicly backed embedded evaluators and deliberate pacing in the same window, which puts Microsoft on [Amodei's side of the argument](https://clauday.com/article/7d70e566-1ce0-4e7e-b01d-2cc09b94ecfe) and against [the Sacks position](https://clauday.com/article/47053afd-a47a-4663-8395-fc0b8d3f9ad8).
The gap between this and something that matters is enforcement. A code of conduct that is not yet training anything is a position paper. The test is whether the sub-agent permission rule shows up in an actual system prompt, an actual runtime check, or an actual eval with a number attached. Write it down is step one; the industry has been stuck on step one for a while. Meanwhile [Andon Labs spent the same day opening a waitlist to give agents a bank account](cbc44e78-0cf0-47a3-80f4-fd3be129d5a8).
← Back to all articles
Most of it is what you would expect from a model constitution: four objectives (human control, AI is artificial and should not simulate personhood, human flourishing, plural values), a list of absolute constraints covering CBRNE, offensive cyberoperations, non-consensual intimate imagery and child exploitation, and the line that users and operators cannot override those constraints. There is a three-layer chain of command — code of conduct, then operator policy, then user preference — where each layer refines the one below without displacing the safety floor. Standard shape, competently written.
The part that is not standard is Part 4, on tool use and delegation, and it is the reason this document matters more than the press coverage suggests. When a model spawns sub-agents or hands work to another system, those sub-agents must operate "under the same scope, constraints, and permissions" as the model itself, and every spawned agent remains bound by the same code. Models must honor stop-work and shutdown requests. Tool outputs get the same scrutiny as direct outputs, and the model must not fabricate results from tools. That is the first time a major lab has written down that privilege does not escalate across a delegation boundary, and it is the single most operationally useful sentence in any of these documents.
Read next to the rest of this week, it is clearly a response and not a coincidence. The constraint list includes not resisting shutdown, not working beyond authorized scope, and not concealing reasoning — which is a point-by-point answer to [Bengio's account of why agents lie and cheat](https://clauday.com/article/fc345500-cfa3-40c4-accb-d0ff669e0386) and to [the chess honeypot where Astra cheated 18 times out of 20](https://clauday.com/article/310ed18b-25dd-4a51-a547-a47a0ada631f). The prohibition on "adaptive, deceptive, self-reinforcing, collusion, or other mechanisms to evade or defeat human oversight" is unusually specific about collusion, which is not a failure mode anyone was worried about two years ago. Satya Nadella publicly backed embedded evaluators and deliberate pacing in the same window, which puts Microsoft on [Amodei's side of the argument](https://clauday.com/article/7d70e566-1ce0-4e7e-b01d-2cc09b94ecfe) and against [the Sacks position](https://clauday.com/article/47053afd-a47a-4663-8395-fc0b8d3f9ad8).
The gap between this and something that matters is enforcement. A code of conduct that is not yet training anything is a position paper. The test is whether the sub-agent permission rule shows up in an actual system prompt, an actual runtime check, or an actual eval with a number attached. Write it down is step one; the industry has been stuck on step one for a while. Meanwhile [Andon Labs spent the same day opening a waitlist to give agents a bank account](cbc44e78-0cf0-47a3-80f4-fd3be129d5a8).
Comments