Elias Virtanen
October 11, 2026
13 min read
Microsoft introduced a new artificial intelligence model built for one narrow job: making fast, structured calls instead of writing essays. Microsoft-Decision-1, unveiled on October 9, 2026, skips free-form text generation entirely. Feed it a situation, a question, and a fixed set of answer options, and it returns a typed, numerical decision or a calibrated probability for each choice in a single pass. No paragraphs, no chain-of-thought, no rationale. Just a score.
Chairman and CEO Satya Nadella announced the model directly, framing it as infrastructure rather than a chatbot. The company is already running it internally for incident response, quality control, and scientific discovery, according to Nadella’s own statement. Within hours, outlets including TestingCatalog, The Decoder, and Yahoo Tech were dissecting benchmark claims, pricing, and what the launch says about where enterprise AI spending is heading next.
Don’t miss new tech stories on Google
Add Tech Insider once in the Google app and our stories appear in your news suggestions.
What Is Microsoft-Decision-1, Exactly
Microsoft-Decision-1 is a decision-scoring model, a category that sits apart from general-purpose large language models like GPT-6 Sol or Gemini 4 Argon. Rather than generating a written answer, it takes a structured input, a scenario, a question, and a closed list of options, and returns a probability or score attached to each option. That output format is the whole point: it is built for routing, classification, prioritization, verification, workflow control, agent guardrails, and AI judging, the tasks that sit underneath most production agent systems rather than in front of them.
The model is post-trained on Alibaba’s Qwen3.5-9B, an open-weight base Microsoft picked rather than building the decision engine on its own MAI family or on OpenAI’s models, at least for this first release. Decision-1 is available now in Microsoft Foundry, where it carries a public preview tag, and through OpenRouter. Pricing is set at $0.042 per million input tokens, with output tokens offered free, a structure that only makes sense once you understand the output is a handful of numbers rather than pages of generated text.
Satya Nadella’s Announcement and Early Testing
Nadella introduced the model in his own words: “Introducing Microsoft-Decision-1, our new model for fast decision-making. It delivers top performance on structured decision tasks, outperforming both LLMs and other decision models in latency and quality,” he wrote in a post announcing the launch. He added that Microsoft is “already testing it across Microsoft for everything from incident response to quality control to scientific discovery.”
That framing matters because it puts Decision-1 in the same bucket as Microsoft’s other recent moves to control AI behavior at scale, including the six AI safeguards and “emergency brake” controls Nadella pushed out earlier this year. A model that can cheaply and quickly verify whether an agent’s next action is safe, correct, or compliant is a natural companion to a trust framework built around stopping agents before they go wrong, rather than after.
How Decision-1 Scores Options Instead of Writing Answers
The mechanical difference between Decision-1 and a conventional chat model is single-pass scoring. A standard LLM asked “should I approve this refund request?” will generate tokens one at a time, building toward an answer and often a justification. Decision-1 instead takes the whole input, the case details and the fixed answer set, yes, no, escalate, in one shot and returns calibrated numbers against each option immediately. There is no token-by-token generation of the decision itself.
Supported Question Types
According to reporting from TestingCatalog, Decision-1 handles yes/no questions, multiple-choice selections, numeric ratings, and rubric-based grading, the last of which lets it score an AI-generated response against a predefined checklist. That rubric-grading capacity is what makes the model useful as an automated judge for other AI systems’ output, not just for human-style business decisions.
Context Window and Consistency
TestingCatalog also reported a context window of up to 32,000 tokens, enough to fit a lengthy support ticket, a multi-step agent transcript, or a batch of candidate outputs ahead of the scoring call. The same reporting cited an internal consistency figure: across eight input perturbations designed to test whether small wording changes flip the model’s answer, Decision-1’s decision changed only 1.3% of the time, a stability number that matters a great deal if a company is wiring the model into automated approval or escalation pipelines where a flip-flopping judge would be worse than a slow one.
Inside the Benchmarks: Accuracy Across 36 Tests
Microsoft’s own benchmark claim, as reported by MarkTechPost and corroborated by TestingCatalog, puts Decision-1’s average accuracy at 83.5% across 36 blind benchmarks covering close to 150,000 held-out questions, numbers the company says make it the most accurate decision model tested in that batch. The closest rival in the same evaluation, a model called Quyet-1.0-Large, scored 81.9% on accuracy, according to TestingCatalog’s account of the results.
Speed is where the reported gaps widen, and where some caution is warranted. MarkTechPost’s write-up cites a 85-millisecond median (p50) latency figure and describes Decision-1 as roughly 2.5 times faster than a model called H2O-Lightning-4B in head-to-head testing. TestingCatalog separately reported internal Microsoft testing showing even larger gaps in specific deployments: Xbox Research reportedly measured Decision-1 running 14 times faster and 200 times cheaper than GPT-6 Sol on its workload, while Microsoft’s own Copilot team reported results around 100 times faster than GPT-5.6 Luna with comparable quality. Multiple outlets, including Neowin and BigGo Finance, have circulated a broader figure putting Decision-1 at roughly 35 times faster than GPT-6 Sol overall. That multiplier has not been confirmed in an official Microsoft benchmark document reviewed for this story, so it should be read as a figure reported by those outlets rather than an independently verified Microsoft claim.
Pricing and Availability in Foundry and OpenRouter
Decision-1 is live in Microsoft Foundry under public preview status and accessible through OpenRouter, giving developers a route into the model outside Microsoft’s own cloud stack. The pricing model, $0.042 per million input tokens with free output tokens, is unusual for a hosted model, but it reflects what the model actually produces: a handful of probability scores rather than generated prose. For a company running millions of routing or verification calls a day, that pricing structure turns what would be a meaningful LLM API bill into a comparatively small line item, assuming the accuracy holds up in production.
The public preview label on Microsoft Foundry is worth noting on its own. It signals Microsoft expects the API, scoring format, or pricing to shift before general availability, which is typical for a brand-new model category rather than a sign of instability. Enterprises piloting Decision-1 for guardrail or compliance use cases should plan for interface changes before locking in production workflows.
A Sample Decision Request, Illustrated
Decision-1’s documented behavior, a scenario plus a fixed option list in, a probability per option out, maps to a request shape similar to the simplified illustration below. This is a conceptual example of the input/output pattern reported for the model, not a verbatim copy of Microsoft’s API schema.
POST /v1/decisions
{
"model": "microsoft-decision-1",
"context": "Support ticket: customer reports a double charge on a $42 order placed 6 days ago.",
"question": "Should this refund be auto-approved?",
"options": ["approve", "escalate", "deny"]
}
Response:
{
"scores": {
"approve": 0.81,
"escalate": 0.17,
"deny": 0.02
},
"latency_ms": 85
}
That single-call, fixed-option pattern is why decision models are being pitched as cheaper, faster substitutes for using a full LLM, like the ones compared in our GPT-6.1 Sol vs. Gemini 4 Argon breakdown, purely as a gatekeeper inside an agent pipeline.
The New “Decision Model” Category Microsoft Just Joined
Decision-1 did not create a new product category so much as join one already moving fast. According to The Decoder, a startup called Jev kicked off the trend in mid-September 2026 with a decision-scoring model of its own, and Decision-1 is reported to outperform Jev’s 1.13.0 release on the same benchmark set. Cloudflare has its own open-source entry, Clef, also built on a Qwen base, though The Decoder’s reporting notes Clef was notably absent from Microsoft’s own published comparison set, an omission worth flagging rather than ignoring.
OpenAI is in the race too, with a Decisions API that reduces complex evaluations to structured outputs rather than open text, a direct analog to what Decision-1 does. And Amazon has already shipped a comparable idea under a different name: Strands Decider, a 2-billion-parameter open agent model Amazon released with roughly 106-millisecond response times for the same class of routing and guardrail tasks. Microsoft’s own Strands Agents SDK coverage on this site already flagged a 115-millisecond “decider” layer as a meaningful building block for agent frameworks, months before Decision-1 arrived with a bigger marketing push behind it.
Why Qwen3.5-9B and Not an In-House Microsoft Model
The choice to post-train Alibaba’s Qwen3.5-9B rather than Microsoft’s own MAI line, or OpenAI’s technology, is one of the more pointed details in this launch. It suggests Microsoft wanted a fast, cheap, flexible base it could specialize quickly for one job, rather than waiting on a larger, slower-moving model family to be adapted for narrow scoring tasks. Reporting from The Decoder indicates Microsoft plans to rebase Decision-1 on its own MAI models and on OpenAI’s technology in future versions, which reads as an admission that the Qwen base was a fast first step rather than the intended long-term foundation.
That sequencing, open-weight base now, in-house base later, lines up with a broader pattern across the AI capex story this year. Microsoft, Meta, and Google have collectively pushed AI infrastructure spending past $725 billion, and a chunk of that spending is going toward building and training proprietary foundation models the company can eventually swap in under products like Decision-1 without depending on a third party’s weights.
Decision Models at a Glance
Several figures in that table come from different parties measuring different things in different ways, which is exactly why decision models are hard to compare cleanly right now. There is no shared, independently audited leaderboard yet the way there is for general-purpose chat models.
Speed Claims by Source
Market Impact: Cutting the Cost of Agent Guardrails
The real audience for Decision-1 is not someone chatting with a bot. It is every engineering team that currently burns a full-size LLM call just to answer a yes/no question inside a larger agent workflow, approve or deny, escalate or close, flag or ignore. That pattern is expensive at scale because general chat models are priced and built for generating long, flexible text, not for returning a single calibrated number. A dedicated decision model priced at a fraction of a cent per million tokens, with free output, changes the economics of running those checks on every single agent action rather than sampling a subset.
That shift matters most for companies already running agentic systems at volume. Nadella’s own examples, incident response, quality control, scientific discovery, all describe high-frequency, structured-judgment tasks rather than open-ended conversation. If Decision-1’s accuracy and latency claims hold up outside Microsoft’s own benchmarks, the practical effect is that agent guardrails stop being a cost center teams try to minimize and start being something they can run on every step without a second thought.
Historical Context: From LLM-as-Judge to Dedicated Scorers
The “LLM-as-judge” pattern, using a general chat model to grade or rank another model’s output, has been a standard technique in AI evaluation for a couple of years. It works, but it is slow and expensive at scale: every judgment call requires spinning up a full generative model capable of writing essays, then throwing away everything except a single verdict buried in the output. Decision models like Decision-1, Jev, Clef, and Amazon’s Strands Decider represent the next logical step, purpose-built scorers that skip the generation step entirely and go straight to the number.
That evolution mirrors what happened earlier with search relevance and content moderation, both of which moved from general models to small, specialized classifiers once volume made the general-purpose approach too costly. Decision models are effectively bringing that same specialization to agent orchestration, a sign that the “just call GPT for everything” phase of agent design is giving way to more deliberately layered systems, a general reasoning model for the hard parts, a decision model for the fast, repetitive parts.
Competitive Pressure on OpenAI and Amazon
Microsoft’s timing is notable. The Decision-1 launch landed within days of reports that OpenAI had opened its own Decisions API to a wider developer base, placing two of the largest AI vendors in direct competition over a product category that barely existed two months ago. That overlap puts pressure on pricing in particular: with Decision-1’s output tokens free and input priced at $0.042 per million, any rival charging meaningfully more for comparable accuracy will need to justify the gap with either quality or integration advantages, such as deeper hooks into an existing agent framework like Google’s Gemini Agent or Microsoft’s own Copilot stack.
Amazon’s position is different because Strands Decider is already open-willing to self-host. That split, hosted-and-priced versus open-and-free, is likely to define how this category shakes out over the next year: cloud vendors compete on latency, accuracy, and integration, while open releases compete on cost and control
What Comes Next: Five Predictions
- General availability will follow a short preview window. Microsoft Foundry’s public preview tag on Decision-1 suggests the company is gathering real usage data before locking the API and pricing for general availability, likely within the next few months rather than a full year.
- A rebase onto MAI or OpenAI technology is coming. Reporting already points to Microsoft’s plan to move Decision-1 off the Qwen3.5-9B base and onto its own or OpenAI’s models, which would reduce Microsoft’s reliance on an externally developed open-weight model for a core piece of its agent infrastructure.
- Independent, audited benchmarks will emerge within weeks. The gap between MarkTechPost’s 2.5x figure and the widely circulated 35x figure against GPT-6 Sol is wide enough that third-party evaluators, the kind that already test models like GPT-6.1 Sol and Gemini 4 Argon, will likely run their own head-to-head comparisons.
- Cloudflare’s Clef will get folded into more comparison tests. Its absence from Microsoft’s initial benchmark set is already drawing attention, and expect Clef to show up explicitly the next time any vendor in this category publishes numbers.
- Pricing in this category keeps falling. With Decision-1 undercutting full LLM calls by a wide margin and Amazon offering a comparable open-source option for free, expect OpenAI’s Decisions API and any other entrant to adjust pricing downward rather than compete purely on accuracy.
How This Fits Microsoft’s Broader AI Strategy
Decision-1 does not exist in isolation. It arrives on the heels of Microsoft’s push for AI trust safeguards, its enormous capital spending alongside Meta and Google, and a steady cadence of product launches meant to keep pace with OpenAI and Google in agentic AI rather than just chat. A cheap, fast, accurate decision layer is the kind of unglamorous infrastructure piece that makes the flashier agent products, copilots that can act rather than just answer, actually safe and affordable to run at scale. It is less a headline feature than a foundation for every other AI product Microsoft wants to ship next.
Whether Decision-1 becomes the default choice for that role, or gets overtaken by Jev, Clef, Amazon’s Strands Decider, or whatever OpenAI ships next for its Decisions API, will likely come down to accuracy numbers that hold up outside vendor-run benchmarks, and to how quickly Microsoft folds the model into its own Copilot and Azure AI Foundry tooling so developers do not have to go looking for it.
Frequently Asked Questions
What is Microsoft-Decision-1?
It is a specialized AI model from Microsoft, announced October 9, 2026, built to make fast, structured decisions, scoring fixed answer options like yes/no or multiple-choice, rather than generating open-ended text.
What model is Decision-1 based on?
Microsoft post-trained Alibaba’s Qwen3.5-9B to create Decision-1, rather than building it on Microsoft’s own MAI models or OpenAI’s technology, at least for this initial release.
How much does Microsoft-Decision-1 cost?
Microsoft prices Decision-1 at $0.042 per million input tokens, with output tokens offered free, reflecting the fact that its output is a short set of scores rather than generated text.
Where can I access Decision-1?
Decision-1 is available now in Microsoft Foundry under public preview status, and also through OpenRouter.
Is Decision-1 really 35 times faster than GPT-6 Sol?
That figure has circulated widely in outlets such as Neowin and BigGo Finance, but it has not been confirmed in an official Microsoft benchmark document. Other reported multipliers vary by test: MarkTechPost cites roughly 2.5x versus H2O-Lightning-4B, while TestingCatalog reports internal Microsoft teams measuring gaps as high as 100x to 200x in specific workloads.
What is Decision-1 used for?
Documented use cases include routing, classification, prioritization, verification, workflow control, agent guardrails, and AI judging, essentially any task where a system needs a fast, structured call rather than a written response.
How does Decision-1 compare to Amazon’s Strands Decider?
Both are decision-scoring models aimed at agent guardrail and routing tasks. Strands Decider is a 2-billion-parameter open-ision-1 is a hosted, Qwen3.5-9B-based model priced per token rather than released as open weights
Who else makes decision models like this?
The category also includes Jev, the startup reported to have started the trend in mid-September 2026, Cloudflare’s open-lopers around the same time Microsoft launched Decision-1
