Agent workflows are full of small, thankless decisions: which model should handle this request, should the previous step be retried, and whether an anomaly deserves a human. Microsoft's answer, published October 9, is Microsoft-Decision-1 — a small model that generates no text and exists to judge.
It doesn't write, it decides
Unlike LLMs, decision models are purpose-built for structured output: given a fixed set of answer options, the model returns a calibrated probability for each, and software can act on that number directly. Microsoft's official use-case list covers model routing, classification, prioritization, verification, workflow control, agent guardrails and AI judging — the high-frequency, low-glamour plumbing of agent infrastructure. The base is Qwen3.5-9B, post-trained for single-pass scoring; Microsoft says it will rebase the model on others, including MAI and OpenAI models. It is available now on Microsoft Foundry and OpenRouter.
One detail worth savoring: Microsoft owns the Phi family and partners closely with OpenAI, yet picked Alibaba's Qwen3.5-9B as the first base — the 9B class happens to sit right at the latency and cost sweet spot for decisions.
The numbers, as officially measured
Microsoft evaluated the model across 36 benchmarks kept blind from training, nearly 150,000 questions in total, and reports the highest accuracy among the models it tested. On speed it measured 2.5x quicker than H2O-Lightning-4B v1.1 (the runner-up) and 35x quicker than GPT-6 Sol at P50 latency. OpenRouter's independent model page lists a 0.27s best-provider P50 and a 32,768-token context window.
On robustness, the team perturbed identical requests in eight ways; decisions flipped on 1.3% of perturbations on average, with zero flips when option descriptions were paraphrased or options reordered. Safety testing covered 5,250 requests across 11 benchmarks. Pricing is the category's signature weapon: $0.042 per million input tokens, output free.
Microsoft is its own first customer
The most persuasive part of the announcement is the internal case work. Xbox Research used the model to process more than 10,000 pieces of open-ended feedback and reviews — competitive in quality with GPT-6 Sol while running over 14x faster and 200x cheaper. The Copilot team uses it to grade chat and agentic responses, competitive with GPT5.6 Luna at 100x the speed. Microsoft Discovery's adaptive replanning scored it 46x more consistent than the LLM-based score. In other words, Microsoft did its own accounting first, then took the model to market.
Four vendors in two weeks: a category forms
Zoom out: in the past two weeks Cloudflare open-sourced Clef, Amazon released Strands Decider, and Perplexity shipped pplx-decider — add Firelex's Jeff from late September and now Microsoft, and "decision models" are graduating from paper concept to a distinct layer of the agent stack. The logic is straightforward: the cost and latency pain of LLM-as-judge is real, and once an agent makes a million small calls a day, "using GPT-6 to decide whether to call GPT-6" becomes a joke. Generation and judgment are splitting up: big models handle the hard, small models handle the fast.
The real thing to watch is what comes next: when judgment becomes cheap enough to embed everywhere, guardrails, routing and quality control shift from afterthoughts to factory defaults. (Official announcement: https://commandline.microsoft.com/microsoft-decision-1-model-foundry/ ; OpenRouter model page: https://openrouter.ai/microsoft/microsoft-decision-1 )