Microsoft-Decision-1 was announced by Microsoft on its official blog on October 9, 2026, is available now in Microsoft Foundry, and is coming soon to OpenRouter. It does not write long text or reason openly. Its only job is to assign a calibrated probability score to each option in a fixed set, so software can act on the score immediately.
A different job from a general-purpose LLM
Microsoft-Decision-1 was post-trained from Qwen3.5-9B for fast, single-pass decision scoring, and Microsoft plans to rebase it onto its own MAI models and OpenAI models. It handles yes/no questions, multiple choice, and rating scales, and can grade AI responses and agent actions against a rubric, all through a simple structured API. Why build a separate model just for judging? Microsoft's reasoning is practical: every decision adds latency, and 100 extra milliseconds across 20 sequential decisions adds two seconds to a workflow. Asking a large general model to think through each judgment is too slow and too expensive at scale.
Two numbers to check first: latency and stability
Across 36 benchmarks and nearly 150,000 questions kept blind from training — spanning routing, ranking, long context, multilingual, and safety tasks — Microsoft-Decision-1 posted the highest accuracy in Microsoft's own testing. Its P50 latency was about 35 times faster than GPT-6 Sol and 4.5 times faster than runner-up Quyet-1.0-Large. Stability got its own test: the same request was perturbed eight ways, through paraphrasing, reordering, key changes, and formatting noise, and decisions flipped only 1.3% of the time on average, with zero flips when option descriptions were paraphrased or options were reversed. The probability is part of the API output, not an afterthought: a 90% prediction should be right about nine times out of ten, so applications can decide whether to act, defer, or ask for human review. On safety, Microsoft tested 5,250 requests across 11 benchmarks covering harmful content, jailbreaks, and prompt injection, and reports the model refused harmful behavior while keeping utility high.
Where Microsoft already uses it internally
The Xbox research team used it to sort more than 10,000 pieces of player feedback into fixed themes, matching GPT-6 Sol on quality while running over 14 times faster at roughly one two-hundredth of the cost. The Copilot team uses it to grade chat and agent responses, reporting quality competitive with GPT-5.6 Luna at 100 times the speed. On-call engineers use it for knowledge retrieval during incident response, and Microsoft Discovery uses its rubric grades to drive adaptive replanning, scoring 46 times more consistently than an LLM-based grader and speeding the loop nearly fourfold. Microsoft is not alone on this path: OpenAI has already turned the same kind of judgment into an interface, covered in our earlier piece on the Decisions API public beta. The difference is that Microsoft has made judging a standalone model rather than a mode of a large one.
Pricing is $0.042 per million input tokens, with output tokens free. The model fits decisions where the options are already fixed: model routing, classification, prioritization, validation, workflow control, and checking an agent's proposed next step. It does not fit open-ended questions, where there is nothing to score. Teams running large volumes of structured judgments should test two things on their own data before adopting it: whether decisions survive rephrasing, and whether its stated probabilities match real accuracy. Leaderboard rank comes after that.