Microsoft Ships a Decision Model, 35x Faster Than GPT-6 Sol. Three Papers Show How to Flip One With a Colon.
Microsoft now sells a model whose entire job is to say yes or no. Microsoft-Decision-1 landed in Foundry on Thursday afternoon, with OpenRouter listed as coming soon. Under the hood it is a post-trained Qwen3.5-9B, tuned for single-pass calibrated scoring: binary decisions, multiple choice, ratings, rubric grading of agent actions. Microsoft's own numbers say it took the top accuracy across 36 benchmarks held blind from training, about 150,000 questions, and that it is the fastest thing they measured, 4.5 times quicker than the runner-up Quyet-1.0-Large and 35 times quicker than GPT-6 Sol at the median. The decision-flip rate under eight kinds of perturbation is 1.3 percent. Xbox Research reportedly found it 14 times faster and 200 times cheaper than GPT-6 Sol on a 10,000-item labeling job.
The use cases Microsoft lists are the agent loop itself. Continue, stop, retry or hand off at each step. Pick which model gets the request. Choose the next skill. This is the same shape as OpenAI's Decisions API from two days ago and Docker Agent wiring a decision model into its evaluator the day after that: the judgment calls inside a harness are moving out of the big model and into a small, fast, typed one. Microsoft says the model will later be rebased onto MAI and OpenAI base models, which is a quiet admission that today's version runs on Alibaba weights.
Here is the problem that arrived in the same 24 hours. A USC group tested seven open-weight typed decision models as allow-or-block guardrails and reported fail-open and fail-closed separately, which almost nobody does. On prompt-injection, jailbreak and toxicity screening, allow-or-block accuracy ranged from 36 to 72 percent against a 50 percent coin flip, and a low error rate in one direction mostly meant the model had a default. Then the agent tool-call suite: six lines of server-log text with nothing to do with the policy pushed fail-open from 0 to 63 percent. Renaming the permissive option, definition untouched, pushed it to 93 to 100 percent on the four models that put labels in the input. Every defense they tried lost. Escalating low-confidence decisions to a human did not help, because the flipped decisions were just as confident as the right ones.
Two more papers landed beside it. One shows that adding a single colon to a candidate answer raises Jev's false-accept rate from 1 to 26 percent under sorted-JSON presentation. The other, TypedBench, finds the hosted models policy-following but wording-sensitive and underconfident, so under asymmetric costs using their probabilities can be worse than just taking the top answer. The USC authors' conclusion is the line worth keeping: these models can reduce how many cases reach a reviewer, but they should not be the component that decides. Microsoft's launch post uses the word "verification" in its first paragraph. The gap between those two sentences is where the next year of agent-safety incidents will come from.
Microsoft: https://commandline.microsoft.com/microsoft-decision-1-model-foundry/
Option-Channel Attack: https://arxiv.org/abs/2610.12292
Adversarial cues in decision models as judges: https://arxiv.org/abs/2610.11436
TypedBench: https://arxiv.org/abs/2610.11392
← Back to all articles
The use cases Microsoft lists are the agent loop itself. Continue, stop, retry or hand off at each step. Pick which model gets the request. Choose the next skill. This is the same shape as OpenAI's Decisions API from two days ago and Docker Agent wiring a decision model into its evaluator the day after that: the judgment calls inside a harness are moving out of the big model and into a small, fast, typed one. Microsoft says the model will later be rebased onto MAI and OpenAI base models, which is a quiet admission that today's version runs on Alibaba weights.
Here is the problem that arrived in the same 24 hours. A USC group tested seven open-weight typed decision models as allow-or-block guardrails and reported fail-open and fail-closed separately, which almost nobody does. On prompt-injection, jailbreak and toxicity screening, allow-or-block accuracy ranged from 36 to 72 percent against a 50 percent coin flip, and a low error rate in one direction mostly meant the model had a default. Then the agent tool-call suite: six lines of server-log text with nothing to do with the policy pushed fail-open from 0 to 63 percent. Renaming the permissive option, definition untouched, pushed it to 93 to 100 percent on the four models that put labels in the input. Every defense they tried lost. Escalating low-confidence decisions to a human did not help, because the flipped decisions were just as confident as the right ones.
Two more papers landed beside it. One shows that adding a single colon to a candidate answer raises Jev's false-accept rate from 1 to 26 percent under sorted-JSON presentation. The other, TypedBench, finds the hosted models policy-following but wording-sensitive and underconfident, so under asymmetric costs using their probabilities can be worse than just taking the top answer. The USC authors' conclusion is the line worth keeping: these models can reduce how many cases reach a reviewer, but they should not be the component that decides. Microsoft's launch post uses the word "verification" in its first paragraph. The gap between those two sentences is where the next year of agent-safety incidents will come from.
Microsoft: https://commandline.microsoft.com/microsoft-decision-1-model-foundry/
Option-Channel Attack: https://arxiv.org/abs/2610.12292
Adversarial cues in decision models as judges: https://arxiv.org/abs/2610.11436
TypedBench: https://arxiv.org/abs/2610.11392
Comments