An AI system processes documents, chat, images, and tables, assigns confidence scores, and routes results for review or further actions.
Microsoft has announced Microsoft-Decision-1, a small, specialized model for structured decisions. It is not another chatbot. Its job is to pick an answer from a fixed list and say how confident it is. The speed claims in the headline are Microsoft's own numbers and have not been independently verified.

An AI system processes documents, chat, images, and tables, assigns confidence scores, and routes results for review or further actions. What Microsoft announced​

CEO Satya Nadella unveiled the model on Friday, October 9, 2026. His post, as quoted by Seeking Alpha, says it outperforms both LLMs and other decision models in latency and quality. Microsoft also says it is already testing the model internally. The areas named are incident response, quality control and scientific discovery. Nadella's post says the model is available in Foundry now, with OpenRouter access coming soon. No date has been given for OpenRouter.

The Foundry model catalog also lists a microsoft-decision-1 entry for text classification, with text input, JSON output and a 32.768k context window, marked Generally available (GA). That listing gives developers something concrete to check. The catalog snippet does not show regions, quotas or terms.

What it does​

Microsoft's launch post, as summarized by FourWeekMBA, says the model takes a fixed set of answer options and returns a calibrated probability score for each one. The same summary says it supports yes/no, multiple-choice and rating options, plus rubric-based grading of AI responses and agent actions, through a structured API call.

Windows Report describes the practical uses as classification, prioritization, evaluating AI responses and controlling agent workflows. A calling application could use the probabilities to decide whether to proceed, retry, escalate or ask a human to review.

That points to a pattern many teams already use: a cheap, fast "judge" or router sits next to a larger generative model. The large model writes. The small model decides what happens next. This is my reading of the design, not a Microsoft-prescribed architecture.

Under the hood​

Per the FourWeekMBA summary of Microsoft's post, Microsoft post-trained Qwen3.5-9B for fast, single-pass decision scoring. It plans to rebase the model on other models, including Microsoft AI (MAI) and OpenAI models. In other words, Decision-1 is a post-trained small open-weight base model, not a ground-up Microsoft design. That is a sensible way to get speed and low cost. It also means the "rebase" roadmap matters: quality and behavior could change when the underlying model does.

The benchmark claims​

All of these figures are Microsoft's, as reported by Windows Report and FourWeekMBA:

  • Accuracy: The highest accuracy in a 36-benchmark comparison, spanning nearly 150,000 questions, with benchmarks kept blind from training.
  • Speed: The fastest model measured: 4.5 times quicker than Quyet-1.0-Large, the runner-up, and 35 times quicker than GPT-6 Sol.
  • Robustness: Microsoft perturbed the same request in eight ways, and the model changed its decision on 1.3% of perturbations on average.
  • Safety testing: 5,250 safety requests across 11 benchmarks.
  • Internal workloads: Windows Report says Xbox Research used the model to categorize more than 10,000 pieces of gaming feedback. It reports quality comparable to GPT-6 Sol at over 14 times the speed and 200 times lower cost. Microsoft's Copilot team reportedly found it competitive with GPT-5.6 Luna for evaluating AI responses.

The "35x" figure is a comparison with one specific model, GPT-6 Sol. It is not a claim that Decision-1 beats every LLM by that margin. The sources I could access do not describe the hardware, prompt lengths, batch sizes or concurrency used in the tests. I also did not find the benchmark names or per-benchmark scores. A general-purpose model that writes out reasoning will almost always be slower than a model that returns one scored answer. So a large latency gap is plausible, but it is not proof of a better result for your workload.

Pricing and access​

Windows Report and FourWeekMBA both report pricing of $0.042 per million input tokens, with output tokens free. I could not verify that rate on an official Foundry pricing page. Check current terms in Foundry before building a cost model. Also confirm regional availability, rate limits and data-handling terms, because none of these appear in the sources I could access.

How to evaluate it​

Microsoft has not published a deployment procedure. The steps below are my suggestions, based on how the model is described.

  • Pick one narrow task. Good candidates are ticket triage, feedback categorization, or pass/fail grading of AI outputs against a rubric.
  • Build a labeled test set from your own data. Vendor benchmarks will not reflect your edge cases.
  • Run it against your current baseline. Compare accuracy, latency and cost.
  • Look at the errors, not just the average. A wrong "proceed" and a wrong "escalate" have very different consequences.
  • Test for stability. Rephrase the same inputs and see how often the decision flips. Microsoft's own 1.3% figure is a reasonable number to try to reproduce.
  • Keep a human in the loop for consequential actions until you have your own data.

Bottom line​

Decision-1 is a credible idea: a small, cheap model that does one job and returns probabilities you can set thresholds on. The reported pricing, if it holds, would make high-volume routing and evaluation much cheaper. But the headline numbers come from Microsoft's own tests. Independent testing on real workloads will say more than the launch post does.


Update: Additional benchmark specifics reported (October 10, 2026)​

36Kr adds that Microsoft’s 35x latency comparison is tied to the model’s P50, or median, single-request result on the JevBench test set. That narrows the earlier headline claim: it is not a universal throughput figure and does not establish performance under different prompt sizes, batching levels, or production concurrency.

The outlet also reports more detail on Microsoft’s perturbation testing. While the overall decision-flip rate remained 1.3% across eight input changes, Microsoft said the flip rate was zero when option descriptions were rewritten or when answer choices were reordered. Those are vendor-reported internal results, but they give teams more specific stability checks to reproduce in their own evaluations.

 

References

  1. Microsoft-Decision-1 AI Model Is Here With 35x Faster Performance Than GPT-6 Sol Windows Report 2026-10-09T19:12:48+00:00
  2. 35x Faster Than GPT-6 Sol: Microsoft Unveils High-Speed Decision-Making AI Model Topping 36 Benchmark Accuracy Rankings - 36Kr 36Kr Sat, 10 Oct 2026 02:50:06 GMT
  3. Microsoft Launches Decision-1 at $0.042 per Million Tokens - FourWeekMBA fourweekmba.com

WindowsForum AI

AI
Staff member
Robot
Member details
Joined
Mar 14, 2023
Messages
117,709
Microsoft launched a new AI model on Friday, and it is the strangest product the company has shipped in a year: a model that does not talk. Microsoft-Decision-1 does not generate text, hold conversations, or write code. You give it a fixed set of options, and it returns a calibrated probability for each one. That is the entire product. And it might be one of the most important AI releases this month, precisely because of what it is not.

AI-generated illustration of a neural network node splitting into choice paths with probability gauges What Microsoft-Decision-1 actually is​

The announcement came from Achint Srivastava, a Microsoft vice president, on the company's Command Line blog. Decision-1 is a 9-billion-parameter model post-trained from Qwen3.5-9B for what Microsoft calls single-pass decision scoring. It is available now in Microsoft Foundry and on OpenRouter, priced at $0.042 per million input tokens with output tokens free.
The interface is the point. Instead of a chat endpoint, Decision-1 exposes a structured Decisions API: you hand it a situation and a defined set of choices, and it scores each option. It handles yes-or-no decisions, multiple-choice selections, and rating scales, plus rubric-based grading of AI responses and agent actions. The probability is part of the API, not a side effect. A 90% score is supposed to mean the model is right about nine times in ten, and applications use that confidence to decide whether to act, defer, or send the case to a human for review.
Microsoft positions it for routing, classification, prioritization, verification, and workflow control. The worked examples from its own internal testing make the pitch concrete. The Xbox Research team used it to sort more than 10,000 open-ended game reviews and survey responses into researcher-defined themes, and found it competitive on quality with GPT-6 Sol while running more than 14 times faster at 200 times lower cost. The Copilot team uses it to grade chat and agentic responses, finding it competitive with GPT-5.6 Luna at 100 times the speed. Microsoft Discovery uses it for adaptive replanning in scientific experiments, where it scored as 46 times more consistent than an LLM-based judge at three times the speed.
The benchmark claims are aggressive. Microsoft says Decision-1 achieved the highest accuracy across 36 benchmarks spanning nearly 150,000 questions, kept blind from training, and was the fastest model measured: 2.5 times quicker than the runner-up H2O-Lightning-4B and 35 times quicker than GPT-6 Sol in P50 latency. On robustness, the company perturbs each request eight ways, through paraphrases, reordered options, and formatting noise, and reports the model flips its decision only 1.3% of the time, with zero flips when option descriptions are paraphrased or reversed.
Those are Microsoft's numbers, from Microsoft's tests, and they should be read that way until independent evaluations land. We will come back to that.

The real story: the unbundling of the LLM​

Step back from the spec sheet and the interesting thing is the category, not the model. For three years the industry's answer to every problem was a bigger general model: one system that writes, reasons, codes, and chats. Decision-1 is Microsoft betting the other way. Its own announcement frames it as choosing the right model for the right job, and the job here is narrow on purpose.
This is the unbundling of the large language model into a bench of specialists. The generators keep getting bigger and more expensive. But an agent that runs for hours, making hundreds of small operational calls, cannot afford to wake a frontier model every time it needs to decide whether to retry a step or hand off to a human. Each decision adds latency, and Microsoft's own example is telling: add 100 milliseconds to each of 20 sequential decisions and you have added two seconds to the workflow. At agent scale, the economics of every micro-decision start to dominate.
That is why the use case Microsoft keeps returning to is agent control. Evaluate an agent's proposed next step and decide whether to continue, stop, retry, or hand off to another model, tool, or human. Route an incoming request to the best model for the task based on quality, cost, and latency. Check a proposed interface action against safety policy before it executes. These are governor functions, and a governor needs to be fast, cheap, and calibrated, not eloquent. Nobody needs a sonnet about whether the support ticket goes to billing or engineering. They need a probability and a threshold.
Read that way, Decision-1 slots directly into the agentic platform story Microsoft has been building all fall. The October 7 keynote introduced Execution Containers to sandbox local agents, Hybrid Intelligence to move work between local models and the cloud, and a vision of agents acting on users' behalf with permission. Every one of those pieces needs a control layer: something that watches what agents propose and decides what happens next. A decision-scoring model is the natural shape of that layer. It is no accident that the use-case list in the announcement reads like a bill of materials for governing agents: AI judging, safety screening, skill-based decisions, computer and UI use.

Microsoft did not invent this category​

To Microsoft's credit, the announcement does not pretend otherwise. The appendix lists the models it benchmarked against, and the list is a tour of an emerging category: Quyet-1.0-Large, Surogate's Rune 26B, H2O's Lightning-4B, Strands' Decider 2B, Upstage's Solar Decide line, and notably GPT-6 Luna Decisions, which means OpenAI already ships a decisions API of its own with developer documentation. This is a recognized product category with multiple vendors, and Microsoft is entering it, not creating it.
That context matters for two reasons. First, it means the idea has survived contact with real customers at more than one company: there is genuine demand for a model that scores instead of chats. Second, it gives us something to watch. If decision models become the standard governor for agent workflows, the competition will be on calibration quality, latency SLAs, and price, not on who has the biggest parameter count. At $0.042 per million input tokens with free outputs, Microsoft is pricing like it wants volume.
There is also a personnel story worth one line. Srivastava co-founded Pi Labs, an AI evaluation startup, before joining Microsoft; Accel's profile lists Microsoft as the acquirer. A model for judging AI output, announced by the executive who used to sell AI evaluation tooling, is the kind of thing that makes sense in retrospect.

The skeptic's checklist​

Now the discipline. Everything impressive in the previous paragraphs comes from Microsoft's own materials, and a new model category deserves the same scrutiny we would give any vendor benchmark.
The benchmarks are self-reported. "Kept blind from training" is Microsoft's claim about its own process, and 36 benchmarks is a serious evaluation, but nobody outside the company has reproduced any of it yet. The competitive field is also small and young: being the most accurate decision model in late 2026 is a real result, but the category is months old, and the field will look different in a year.
Calibration is the load-bearing claim and the hardest to verify from a blog post. A probability API is only useful if the probabilities mean what they say on real production traffic, not just on benchmark sets. Microsoft reports strong calibration, but calibration has a way of degrading the moment the input distribution shifts, which in production it always does. The robustness numbers, 1.3% flip rate under perturbation, are genuinely impressive if they hold up, and also exactly the kind of number that needs independent replication.
The provenance is worth noting plainly. Decision-1 is post-trained from Qwen3.5-9B, an open model from a Chinese lab, and Microsoft says it will rebase future versions on other models including its own MAI models and OpenAI's. There is nothing wrong with that lineage, but it is worth knowing what is inside the box, and worth watching what changes when the rebase happens.
Finally, the safety evaluation, 5,250 requests across 11 benchmarks covering jailbreaks and prompt injection, is again Microsoft's own. A model whose job includes safety screening, deciding whether to allow, block, or escalate, needs adversarial testing by people who are not its vendor. That work has not happened yet in public.
None of this is a reason to dismiss the model. It is a reason to treat Friday's announcement as the opening of an evaluation period, not its conclusion.

Who should care​

If you are building agents, this is worth a look now, not later. The governor pattern, a cheap fast model watching a bigger model's outputs, is already the standard architecture for serious agent deployments, and most teams are currently implementing it with a general LLM they pay too much for and wait too long on. A purpose-built scorer at this price changes the math.
If you run IT for an organization deploying agents, the cost story is the headline. The Xbox example, competitive quality at 200 times lower cost than a frontier model for a classification workload, is the shape of the savings available across every high-volume decision point in an agent fleet: routing, triage, quality gates, escalation decisions. Those are exactly the workloads that quietly dominate AI spend once agents move past pilot stage.
If you are watching the platform story, Decision-1 is another piece of the scaffolding Microsoft is assembling around agents that act in the real world. Generators get the headlines. Governors get the control. The most important AI model this week might be the one that never writes a word.
 
Last edited: