What Microsoft announced
CEO Satya Nadella unveiled the model on Friday, October 9, 2026. His post, as quoted by Seeking Alpha, says it outperforms both LLMs and other decision models in latency and quality. Microsoft also says it is already testing the model internally. The areas named are incident response, quality control and scientific discovery. Nadella's post says the model is available in Foundry now, with OpenRouter access coming soon. No date has been given for OpenRouter.
The Foundry model catalog also lists a microsoft-decision-1 entry for text classification, with text input, JSON output and a 32.768k context window, marked Generally available (GA). That listing gives developers something concrete to check. The catalog snippet does not show regions, quotas or terms.
What it does
Microsoft's launch post, as summarized by FourWeekMBA, says the model takes a fixed set of answer options and returns a calibrated probability score for each one. The same summary says it supports yes/no, multiple-choice and rating options, plus rubric-based grading of AI responses and agent actions, through a structured API call.
Windows Report describes the practical uses as classification, prioritization, evaluating AI responses and controlling agent workflows. A calling application could use the probabilities to decide whether to proceed, retry, escalate or ask a human to review.
That points to a pattern many teams already use: a cheap, fast "judge" or router sits next to a larger generative model. The large model writes. The small model decides what happens next. This is my reading of the design, not a Microsoft-prescribed architecture.
Under the hood
Per the FourWeekMBA summary of Microsoft's post, Microsoft post-trained Qwen3.5-9B for fast, single-pass decision scoring. It plans to rebase the model on other models, including Microsoft AI (MAI) and OpenAI models. In other words, Decision-1 is a post-trained small open-weight base model, not a ground-up Microsoft design. That is a sensible way to get speed and low cost. It also means the "rebase" roadmap matters: quality and behavior could change when the underlying model does.
The benchmark claims
All of these figures are Microsoft's, as reported by Windows Report and FourWeekMBA:
- Accuracy: The highest accuracy in a 36-benchmark comparison, spanning nearly 150,000 questions, with benchmarks kept blind from training.
- Speed: The fastest model measured: 4.5 times quicker than Quyet-1.0-Large, the runner-up, and 35 times quicker than GPT-6 Sol.
- Robustness: Microsoft perturbed the same request in eight ways, and the model changed its decision on 1.3% of perturbations on average.
- Safety testing: 5,250 safety requests across 11 benchmarks.
- Internal workloads: Windows Report says Xbox Research used the model to categorize more than 10,000 pieces of gaming feedback. It reports quality comparable to GPT-6 Sol at over 14 times the speed and 200 times lower cost. Microsoft's Copilot team reportedly found it competitive with GPT-5.6 Luna for evaluating AI responses.
The "35x" figure is a comparison with one specific model, GPT-6 Sol. It is not a claim that Decision-1 beats every LLM by that margin. The sources I could access do not describe the hardware, prompt lengths, batch sizes or concurrency used in the tests. I also did not find the benchmark names or per-benchmark scores. A general-purpose model that writes out reasoning will almost always be slower than a model that returns one scored answer. So a large latency gap is plausible, but it is not proof of a better result for your workload.
Pricing and access
Windows Report and FourWeekMBA both report pricing of $0.042 per million input tokens, with output tokens free. I could not verify that rate on an official Foundry pricing page. Check current terms in Foundry before building a cost model. Also confirm regional availability, rate limits and data-handling terms, because none of these appear in the sources I could access.
How to evaluate it
Microsoft has not published a deployment procedure. The steps below are my suggestions, based on how the model is described.
- Pick one narrow task. Good candidates are ticket triage, feedback categorization, or pass/fail grading of AI outputs against a rubric.
- Build a labeled test set from your own data. Vendor benchmarks will not reflect your edge cases.
- Run it against your current baseline. Compare accuracy, latency and cost.
- Look at the errors, not just the average. A wrong "proceed" and a wrong "escalate" have very different consequences.
- Test for stability. Rephrase the same inputs and see how often the decision flips. Microsoft's own 1.3% figure is a reasonable number to try to reproduce.
- Keep a human in the loop for consequential actions until you have your own data.
Bottom line
Decision-1 is a credible idea: a small, cheap model that does one job and returns probabilities you can set thresholds on. The reported pricing, if it holds, would make high-volume routing and evaluation much cheaper. But the headline numbers come from Microsoft's own tests. Independent testing on real workloads will say more than the launch post does.
References
- Microsoft-Decision-1 AI Model Is Here With 35x Faster Performance Than GPT-6 Sol Windows Report · 2026-10-09T19:12:48+00:00
- Microsoft Launches Decision-1 at $0.042 per Million Tokens - FourWeekMBA fourweekmba.com
- Microsoft's Nadella unveils tech giant's new fast decision-making AI model (MSFT:NASDAQ) seekingalpha.com