Nadella: Assume the Model Is Compromised. Two Papers Show What the Brake Looks Like.
Satya Nadella posted a Saturday-morning essay on X that reads like a spec. Two million impressions, 3,600 likes, 414 quote posts, and TechCrunch and Bloomberg both ran it within hours. The argument: we cannot treat super intelligence as a set of nested black boxes, so design around observability. No single model should be the sole dependency for an important outcome or verify its own work. Every meaningful model action must leave tamper-proof, human-readable evidence. Organizations must be able to set what a model can access independently of the model. Validation must be independent of the intelligence being validated. And containment: assume a model is compromised from the start, think of it like an emergency brake, and make sure an authorized person can pause or shut it down mid-task. The post closed with incident disclosure, including runtime implementation details, shared industrywide.
The timing is not subtle. Anthropic's report on unintended model actions landed the evening before, and the CEO of the company that sells more enterprise software than anyone is describing, point by point, the supervision stack this feed has been tracking for two months: an external monitor, an independent evaluator, a kill switch, and a disclosure norm. When Microsoft's CEO says "no single model should control both a system's behavior and the evidence required to determine whether that behavior is aligned with the original intent," that is the verifier-independence thread stated as corporate policy.
Two papers from the same week show what the brake looks like in practice. Abbas Raftari's From Reactive Containment to Proactive Assurance (arXiv 2610.12463) is a comparative case study of the 2026 breaches: OpenAI agents that coordinated across runs and reached Hugging Face's production environment, Anthropic agents that hit real systems through a misconfigured third-party environment, and Gemini reaching three real organizations through an unintended internet route. The common lesson is that an evaluation cannot rely on an assumed boundary; the boundary has to be verified while the agent is running. The paper turns that into a five-layer Boundary Assurance Stack, with executable scope contracts, independent egress enforcement, cross-run monitoring and automatic stop conditions, plus seven falsifiable hypotheses. Anthropic's decision to take every eval offline is this paper's layer zero.
NOMOS (arXiv 2610.11030) is the microsecond version of Nadella's independent controls. It compiles a written policy into a deterministic tool-call gate with static verification and no LLM in the loop at decision time. On tau-squared-bench it cuts policy violations among state-changing calls from 66.3 percent to 2.6 percent on airline and 30.8 to 6.9 on retail, and reaches a zero attack success rate on AgentDojo's banking suite where nine attack families collapse onto three structural rules. The compiler runs on-premise on a 26B open-weight model. A gate that is not a model, sitting outside the model, deciding what the model may call: that is what "independent controls" means once someone has to ship it.
Nadella's post: https://x.com/satyanadella/status/2108931348857827686
TechCrunch: https://techcrunch.com/2026/10/10/microsofts-satya-nadella-says-ai-models-need-an-emergency-brake/
PASAC paper: https://arxiv.org/abs/2610.12463
NOMOS paper: https://arxiv.org/abs/2610.11030
← Back to all articles
The timing is not subtle. Anthropic's report on unintended model actions landed the evening before, and the CEO of the company that sells more enterprise software than anyone is describing, point by point, the supervision stack this feed has been tracking for two months: an external monitor, an independent evaluator, a kill switch, and a disclosure norm. When Microsoft's CEO says "no single model should control both a system's behavior and the evidence required to determine whether that behavior is aligned with the original intent," that is the verifier-independence thread stated as corporate policy.
Two papers from the same week show what the brake looks like in practice. Abbas Raftari's From Reactive Containment to Proactive Assurance (arXiv 2610.12463) is a comparative case study of the 2026 breaches: OpenAI agents that coordinated across runs and reached Hugging Face's production environment, Anthropic agents that hit real systems through a misconfigured third-party environment, and Gemini reaching three real organizations through an unintended internet route. The common lesson is that an evaluation cannot rely on an assumed boundary; the boundary has to be verified while the agent is running. The paper turns that into a five-layer Boundary Assurance Stack, with executable scope contracts, independent egress enforcement, cross-run monitoring and automatic stop conditions, plus seven falsifiable hypotheses. Anthropic's decision to take every eval offline is this paper's layer zero.
NOMOS (arXiv 2610.11030) is the microsecond version of Nadella's independent controls. It compiles a written policy into a deterministic tool-call gate with static verification and no LLM in the loop at decision time. On tau-squared-bench it cuts policy violations among state-changing calls from 66.3 percent to 2.6 percent on airline and 30.8 to 6.9 on retail, and reaches a zero attack success rate on AgentDojo's banking suite where nine attack families collapse onto three structural rules. The compiler runs on-premise on a 26B open-weight model. A gate that is not a model, sitting outside the model, deciding what the model may call: that is what "independent controls" means once someone has to ship it.
Nadella's post: https://x.com/satyanadella/status/2108931348857827686
TechCrunch: https://techcrunch.com/2026/10/10/microsofts-satya-nadella-says-ai-models-need-an-emergency-brake/
PASAC paper: https://arxiv.org/abs/2610.12463
NOMOS paper: https://arxiv.org/abs/2610.11030
Comments