Microsoft Decision-1 Explained: A Qwen Fine-Tune Priced to the Cent Like Jev ($0.042 per Million Tokens), the systemone API, Foundry Setup, and How It Compares with Jev, Clef, OpenAI's Decisions API and Amazon's Strands Decider

Logeshwaran
—

Microsoft Decision-1 is Microsoft's first decision model, released on October 9, 2026 in public preview on Microsoft Foundry. It does not write text. You hand it a situation and a fixed list of answers, and it returns a calibrated probability for each one, in a single pass, for $0.042 per million input tokens with output free. Two facts the launch coverage skipped. First, Microsoft did not build it from scratch: Decision-1 is a post-trained version of Alibaba's open-weight Qwen3.5-9B, and Microsoft plans to rebase it soon on other models, including Microsoft AI (MAI) and OpenAI models. Second, that price is not a coincidence. TypeSafe's Jev, the model that created this category in September, costs exactly $0.042 per million input tokens with output free. Microsoft matched it to the tenth of a cent, and then borrowed the endpoint name too: you call Decision-1 at /providers/microsoft/v1/systemone.

Jake runs a phone repair shop. Messages arrive all day through the shop's website and WhatsApp: "how much for a cracked S24 screen", "my repair from March has stopped working again", "you charged me twice", and a steady trickle of spam. Last month his nephew wired them into a chat model that wrote a polite paragraph back to each one and filed it under the wrong heading about one time in five. Ethan, who looks after AWS and a few APIs for local businesses, told him the job was never writing: it was sorting. "You want a yes or a no, or one label out of four, with a number that says how sure it is. That is a decision model." This page is what Ethan then worked through: what Microsoft Decision-1 is, what Microsoft actually claims and what it leaves out, the price against Jev, Cloudflare's Clef, OpenAI's Decisions API and Amazon's Strands Decider, how to deploy it on Microsoft Foundry and call the systemone API, how to write questions that get reliable answers, what to run if you want the model on your own machine or on AWS, and the errors you will meet first.

⚡ Quick Answer

• What it is → a 9B decision model, post-trained from Qwen3.5-9B, that answers yes/no, pick-one and rate-on-a-scale questions with probabilities instead of prose. 32K-token input, text and JSON only. What a decision model does.

• Price → $0.042 per million input tokens, output free; the same as Jev, 2.4 times cheaper than OpenAI's Decisions API, slightly dearer than Cloudflare's Clef-flash at $0.038. The comparison table and a worked bill.

• Microsoft's claims → highest accuracy across 36 benchmarks and nearly 150,000 questions, P50 latency about 35 times faster than GPT-6 Sol; no per-benchmark scores published. What is proven and what is not.

• How to call it → deploy from the Foundry model catalog, then POST to /providers/microsoft/v1/systemone with a state and typed questions. The request, the response and the curl.

• Not on AWS, not downloadable → no weights, no Bedrock listing. Amazon's own answer is the open Strands Decider 2B, which runs on a CPU. The AWS and local routes.

• Who it is for → routing, classification, grading model output, agent guardrails, incident triage: anything that is a choice, not a paragraph. Six things to build.

If "decision model" is new to you, our plain-English guide to decision models explains the idea from zero in ten minutes.

If the words tokens, API or model are still fuzzy, start with our guide to what an LLM is; everything below assumes only that.

What Microsoft Decision-1 is, and what a decision model does

In one line: Decision-1 is a decision-scoring model that returns calibrated probability scores for fixed answer options instead of generated text. You give it a state, which is the situation, as text or JSON, and one or more questions, each with a fixed set of allowed answers. It returns, for each question, the probability of each answer. It never produces a sentence, never explains itself, and never invents a fifth option when you offered four.

The industry name for this kind of model is System One, a term TypeSafe borrowed from Daniel Kahneman when it launched Jev in September 2026: System 1 is the fast, intuitive judgment, System 2 the slow, deliberate reasoning. Decision-1 adopts the vocabulary: its request pattern works like a system-one API, and the endpoint is literally named systemone. Decision-1 answers three kinds of question:

You askTypeYou get backJake's shop example
Is this true?noulOne probability from 0 to 1"Is this message spam?"
Which option is it?choiceOne selected option plus a probability for every option"Quote, warranty claim, billing or other?"
How much?scoreA value on an ordered scale plus a probability for every level"How angry is this customer, from calm to furious?"

Those are Microsoft's names; "noul" is TypeSafe's word for a yes/no question, and OpenAI calls the same thing a "predicate." In the other framings people search for, it handles yes/no, multiple-choice, rating, classification and rubric-based questions, takes inputs of up to 32K tokens, and is text-only. What it does not do is just as important. Generation, open-ended question answering, conversation, translation and summarization are out of scope, and it never writes an explanation or a rationale.

Jake: "So it is a model that cannot talk. Why is that good?"

Ethan: "Because a model that can talk will talk its way into a wrong answer and make it sound right. This one can only point at one of the doors you showed it, and tell you how sure it is. For sorting your inbox, pointing is the whole job."

How Microsoft built it: a Qwen under the hood, and no weights to download

The recipe is simple: Microsoft took Alibaba's open-weight Qwen3.5-9B and post-trained it for fast, single-pass decision scoring, and it plans to rebase the model soon on others, including Microsoft AI (MAI) and OpenAI models. So the first Microsoft decision model is a Microsoft training run on an Alibaba base, which puts it in the same family as most of the open decision models that appeared in September: Cloudflare's Clef-flash is a Qwen3.5-9B too, Amazon's Strands Decider sits on Qwen3.5-2B, and H2O's Lightning on Qwen3.5-4B.

"Single-pass" is the technical heart of every model in this category. A chat model answers a multiple-choice question by generating tokens one at a time, which is slow and lets it wander. A decision model reads the state and the options once and reads the answer straight off the probability distribution, so the cost is one forward pass and the output is a set of numbers. That is where the speed and the price come from, and why output tokens are free: there are none.

Unlike Clef, Strands Decider or H2O-Lightning, Decision-1 is hosted only. Access is through Foundry only: there are no weights on Hugging Face and no download. If you want a decision model on your own hardware, the section on local alternatives below covers what exists today.

Microsoft's numbers: what is claimed, what is measured, and what is missing

Microsoft makes four kinds of claim, and it helps to separate them.

Accuracy. "Microsoft-Decision-1 achieved the highest accuracy in our 36-benchmark comparison, spanning nearly 150,000 questions across benchmarks kept blind from training." Separately: "We also took several of the top public models on the popular open leaderboard JevBench and tested them across 36 additional public and private benchmarks." An editor's note says the post "was updated from the original to add benchmarks for Jev on accuracy and calibration." What the post does not contain is a single per-benchmark score, so there is no number you can compare with the figures other vendors publish, and no independent JevBench result for Decision-1 is public yet.

Speed. "Microsoft-Decision-1 P50 latency is ~35x faster than GPT-6 Sol," and among the decision models it tested it was "2.5 times quicker than H2O-Lightning-4B v1.1, the runner-up." Microsoft measured both on its own deployment. Its argument for why speed matters is sound: "adding just 100 milliseconds to each of 20 sequential decisions adds two seconds to the overall workflow."

Robustness. Across eight kinds of perturbation, "Microsoft-Decision-1 changes its decision on 1.3% of perturbations on average with zero flips when option descriptions are paraphrased or when options are reversed or shuffled." That last clause matters more than it looks; option order steering the answer is the best-known weakness of this whole category, and our Clef guide shows it happening.

Safety. "We tested Microsoft-Decision-1 on 5,250 requests across 11 benchmarks, covering harmful content, jailbreak attempts, and prompt injection, and found that the model successfully refused harmful behavior while retaining a high degree of utility." Again, no figures.

The appendix names the six models Microsoft benchmarked against, and they are a useful map of the field. Here is what each one is, from its own model card or repository:

Model Microsoft tested againstWhoBuilt onWhat its own card claims
GPT-6 Luna DecisionsOpenAIGPT-6 Luna, via the Decisions APIPublic beta; "about 10x faster than the Responses API"; $0.10 per million input tokens
H2O-Lightning-4BH2O.aiQwen3.5-4B, Apache 2.0#1 open-weight composite on JevBench v1.6.1 at 72.5 against Jev's 71.5; P50 about 28 to 33 ms on an H100 or RTX PRO card
Strands-Decider 2BAmazon's Strands LabsQwen3.5-2B torso plus a pointer head, Apache 2.03rd of 33 in JevBench's 2B class; about 115 ms on an RTX 3090 and 153 ms on an M3 Mac
Surogate Rune 26B-A4B (v3)InvergentGemma 4 26B-A4B, Apache 2.057.44 on the Decision Index, 0.45 behind Jev; 0.09 to 0.18 s per decision; reads images
Quyet-1.0-LargeChinh NguyenGemma 4 31B, Apache 2.031.3B parameters, one 80 GB GPU, up to 10 options per question; no public scores
deck-31BAn independent developerGemma 4 31B, frozen, no training215 of 231 public JevBench items in a local run; P50 1.05 s; needs an 80 GB FP8 GPU

Two absences stand out. Jev itself was added only in the editor's note, and Cloudflare's Clef, the other hosted model with a global network behind it, is not in the list at all. When you read "highest accuracy," read it as highest among these six plus Jev, on Microsoft's benchmarks, measured by Microsoft.

Jake: "Thirty-six benchmarks and no scores? If I told a customer my repairs passed thirty-six tests and refused to show one, they would walk."

Ethan: "They would. The others publish their numbers on a public board, and Microsoft has not yet. It may be the best of the lot; today you cannot check, so you test it on your own fifty messages before you trust it. That is true of every model on this page."

How Microsoft uses it in-house: the four stories behind the launch

The most convincing part of the launch is not the benchmark paragraph but four internal deployments, each with a number attached:

  • Xbox Research sorted "more than 10,000 open-ended feedback items" into fixed themes, with quality competitive with GPT-6 Sol, "over 14 times faster and 200 times less expensive."
  • The Copilot team uses it to measure the quality of chat and agentic responses, "competitive with GPT5.6 Luna and 100 times faster."
  • Incident response: on-call engineers' AI retrieval over logs, tickets and messages "performed better and faster than an LLM for knowledge retrieval."
  • Microsoft Discovery, the scientific agent, grades each experiment against a rubric before replanning; the decision model was "46 times more consistent than the LLM-based score at three times the speed."

Read these as a job description. Sorting feedback into themes, grading responses against a rubric, routing an incident, deciding whether to replan: every one is a choice among options you already know. None is "write me something." If your task fits that description, a decision model is the right tool; if it does not, no decision model will help.

Microsoft Decision-1 pricing against Jev, Clef, OpenAI and Strands

Input tokens cost $0.042 per million. Output tokens are free. Here is how that sits against every hosted decision model you can call today, and the two free routes:

ModelPer million input tokensOutputContextInputsWhere
Microsoft Decision-1$0.042Free32KText, JSONMicrosoft Foundry
TypeSafe Jev 1.13$0.042Free64KTextTypeSafe API
OpenAI Decisions API (gpt-6-luna)$0.10FreeNot stated in the guideText, inline base64 imagesOpenAI, public beta
Cloudflare Clef-flash (9B)$0.038 (was $0.09 until October 9)Free24K hostedText, imagesWorkers AI; open weights
Cloudflare Clef (27B)$0.24Free64KText, imagesWorkers AI; open weights
Cloudflare Clef-omni$0.15Free64KText, images, audio, videoWorkers AI; open weights
Amazon Strands Decider 2BFree; your hardware—4K serving windowTextLocal CPU or GPU; Apache 2.0
H2O-Lightning-4BFree; your hardware—Not stated on the cardText, imagesSelf-hosted with vLLM; Apache 2.0

Three things to read off that table. Microsoft priced Decision-1 at exactly Jev's number, so the choice between them is not about money; it is about ecosystem, context length and whose benchmarks you believe. Cloudflare undercut both by four tenths of a cent the same day Microsoft launched, and its free tier on Workers AI covers thousands of short decisions a day at no charge. And OpenAI's hosted option costs 2.4 times Microsoft's, in return for a far larger base model and image input.

A worked bill, with every number an assumption to swap for your own. Jake's shop gets about 300 messages a week. Each one, with its question block, is about 300 tokens, so a month is roughly 390,000 input tokens.

WorkloadTokens a monthDecision-1 or JevOpenAI DecisionsClef-flash
Jake's shop: 300 messages a week at 300 tokensabout 390K$0.02$0.04$0.01, or $0 inside the free tier
A support desk: 50,000 tickets a month at 800 tokens40M$1.68$4.00$1.52
An agent guardrail: 10 million checks a month at 400 tokens4 billion$168$400$152

At Jake's scale the price is a rounding error whichever you choose, and the deciding factors are accuracy on his messages and how little code it takes. At ten million decisions a month the gaps are real money, and they are still small next to what a generative model would charge to read the same tokens and write an answer; Microsoft's Xbox team put its own saving at 200 times. On Foundry, the standard deployment type is GlobalStandard; a DataZoneStandard deployment keeps processing inside a data zone and is available "in selected regions." Check the Foundry pricing page for your deployment type before you forecast a bill.

Microsoft Foundry: what it is, and how to deploy Decision-1 in five steps

If you have not touched Azure's AI tools for a year, the names have changed under you. The names went in this order: Azure AI Studio became Azure AI Foundry, which is now Microsoft Foundry, with the old portal living on as "Foundry (classic)." Foundry is the catalog and deployment layer for models on Azure: OpenAI's models, Microsoft's own MAI models and partner models sit in one catalog, you deploy the one you want into your own Azure resource, and you call it with an Azure identity or an API key. Decision-1 is listed there as a "Direct from Azure" model, version 1, in public preview.

Before you deploy, you need: an Azure subscription with a valid payment method, a Foundry project, permission to create deployments (the Cognitive Services Contributor role covers it), and a way to sign in, with Microsoft Entra ID recommended over an API key. Then:

  1. Open the Foundry portal at ai.azure.com and go to your project.
  2. Select Model catalog, search for Microsoft-Decision-1 and select it.
  3. Select Deploy, choose the deployment type (GlobalStandard everywhere, or DataZoneStandard where offered) and a deployment name, then confirm.
  4. On the deployment details page, copy the resource endpoint, the deployment name and, if you use one, the API key.
  5. Open the Foundry Playground for the model and send one question before you write code; Microsoft's own demo there checks urgency, picks a support team and scores priority from a single customer message.

The same deployment from the command line, with your own names in the placeholders:

az cognitiveservices account deployment create \
  --name <ACCOUNT_NAME> \
  --resource-group <RESOURCE_GROUP> \
  --deployment-name <DEPLOYMENT_NAME> \
  --model-name "Microsoft-Decision-1" \
  --model-format Microsoft \
  --model-version "1" \
  --sku-name GlobalStandard \
  --sku-capacity 1

Change --sku-name to DataZoneStandard to keep inference inside the data zone nearest your application, where that type is available. The latency advice is worth repeating because it is the whole point of the model: create the Foundry resource in a region close to your users and benchmark it, since "a deployment type doesn't guarantee lower latency."

The systemone API: the request, the response and your first call

Requests go to your Foundry resource's decision endpoint:

<your-foundry-resource-endpoint>/providers/microsoft/v1/systemone

Two rules trip up first attempts. The model field in the body is your deployment name, not the catalog name, for example a deployment called pi-decision-1. And the response's own model field reports the underlying model, so it can read microsoft-decision-1 while your request said something else. A routing request for Jake's inbox looks like this:

{
  "model": "jake-decision-1",
  "state": "Hi, you fixed my phone screen in March and now the touch has stopped working on the left side. Is this covered?",
  "questions": {
    "intent": {
      "type": "choice",
      "instructions": "What is this customer asking for? Select exactly one.",
      "criteria": {
        "quote": "Asking the price or time for a new repair",
        "warranty": "A problem with a repair already done by the shop",
        "billing": "Charges, invoices, refunds or payment problems",
        "other": "Anything else, including spam and greetings"
      }
    },
    "needs_human_today": {
      "type": "noul",
      "instructions": "Does this message need a reply from a person today?",
      "criteria": {
        "true": "The customer has a live problem or is upset",
        "false": "Routine, informational or spam"
      }
    },
    "frustration": {
      "type": "score",
      "instructions": "How frustrated is the customer?",
      "criteria": ["Calm", "Mildly annoyed", "Frustrated", "Angry"]
    }
  }
}

Each question has a type, free-text instructions, and criteria: a map of option name to description for choice, an optional true/false pair for noul, and an ordered list from lowest to highest for score. Independent questions about the same state belong in one request: combining them means fewer requests and keeps the decision context consistent. The response carries an answers object keyed by your question names. A choice answer contains the selected option and a probability for every option; a noul answer is a probability from 0 to 1; a score answer is "the probability-weighted average of the level indexes," so a four-level scale runs from 0 to 3 and an answer of 2.4 means the model leans past "Frustrated" toward "Angry." The response also includes the model name and token usage.

Sending it with Microsoft Entra ID, the recommended sign-in:

az login
export AZURE_ENDPOINT="<your-foundry-resource-endpoint>"
export DEPLOYMENT_NAME="jake-decision-1"
export AZURE_ENTRA_TOKEN=$(az account get-access-token \
  --resource https://cognitiveservices.azure.com \
  --query accessToken --output tsv)

curl "$AZURE_ENDPOINT/providers/microsoft/v1/systemone" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $AZURE_ENTRA_TOKEN" \
  -d @request.json

With an API key instead, replace the Authorization header with -H "api-key: $AZURE_API_KEY". Microsoft's Python quickstart uses nothing but the standard library plus azure-identity: it builds the same JSON, posts it with urllib, checks that each answer's type is choice and that the returned option is one you offered, and times each call so you get accuracy and median latency on your own labeled examples. That last habit is the one to copy. Microsoft's own instruction is to "replace EXAMPLES with representative, labeled examples from your workload before you use the measurements to make deployment decisions."

Jake: "So I write the four doors, and it tells me which one and how sure?"

Ethan: "And when it is not sure, the number says so. Your warranty message might come back 0.71 warranty, 0.22 quote, 0.07 billing. That 0.71 is the useful part. Below some line you pick, the message goes to you instead of being filed."

Writing questions that work: the rules, in plain words

A decision model is only as good as the options you hand it, and the failure modes are worth knowing up front. The model can reflect biases from its base model and training data; scores can change with how you phrase or order the questions and options, and a poorly framed question still returns a score; calibration is strongest on familiar task types, and the model can rely on outdated knowledge; and as a safety filter it can miss subtle harmful content or flag benign content. The best practices follow from those, and they apply to every model in this category:

  • Validate on your own data first. Fifty labeled examples from your real traffic tell you more than any vendor table.
  • Describe each option clearly and use "clear, neutral wording." The description is what the model matches against; a vague option attracts vague matches.
  • Add an escape hatch. "Include an abstention option, such as cannot tell, when the model shouldn't make a forced choice." Jake's "other" bucket is that hatch.
  • Set thresholds by the cost of being wrong. A missed warranty claim costs a customer; a mis-filed spam message costs nothing. Different thresholds.
  • Randomize option order in testing and check whether the answers move. Microsoft reports zero flips on reversed or shuffled options; verify it on your questions.
  • Use scores for ordering, not as absolute ratings. Microsoft's words: "prefer scores for relative ordering and thresholds rather than as absolute, calibrated ratings."
  • Keep people in the loop where it matters. For decisions about credit, employment, housing, healthcare or legal matters, use the model for decision support with meaningful human review, never as the sole decision-maker, and tell affected users when AI contributed.

Turned into a routine you can finish in an afternoon:

  1. Pull fifty real items from your own traffic and label them by hand, including the awkward ones you would normally skip.
  2. Write your questions with one option description per outcome and an abstention option, then run all fifty and record the top probability for each.
  3. Count the mistakes, then look at the probabilities on the wrong answers; most should sit low. Pick a threshold below which items go to a person, and check how many of the fifty that catches.
  4. Shuffle the option order, run the fifty again, and compare; any answer that moves points to a weak description.
  5. Run the same fifty through a second model, Jev, Clef-flash or a local Strands Decider, before you commit, since the request shape barely changes.

One more from the launch post that is easy to miss: a "90% prediction should be right about nine times out of 10 on representative cases." That is what calibrated means, and it is the property that lets you set a threshold and trust it. It is also the property most worth checking yourself, because no calibration figure has been published.

Decision-1 vs Jev vs Clef vs OpenAI's Decisions API vs Strands Decider

Five decision models now compete for the same jobs, and they differ in ways that matter more than price:

Microsoft Decision-1TypeSafe Jev 1.13Cloudflare Clef / Clef-flashOpenAI Decisions APIAmazon Strands Decider 2B
ReleasedOctober 9, 2026, public previewSeptember 2026October 1, 2026; Clef-omni October 9Public beta, GA "in the coming weeks"October 1, 2026
BaseQwen3.5-9BUndisclosedQwen3.8-27B / Qwen3.5-9BGPT-6 LunaQwen3.5-2B
Question typesnoul, choice, scorenoul, choice, scorenoul, choice, scorepredicate, choice, scorenoul, choice, score
InputsText, JSONTextText, up to 4 images; Clef-omni adds audio and videoText, inline base64 imagesText
Context32K64K64K; flash 24K hostedNot stated in the guide4K
Price per M input$0.042$0.042$0.24 / $0.038$0.10Free, local
WeightsNot publishedClosedOpen, Apache 2.0ClosedOpen, Apache 2.0, with training data
Published speed claimP50 about 35x faster than GPT-6 Sol70 to 500 ms end to endUp to 2x faster since October 9About 10x faster than the Responses APIAbout 115 ms on an RTX 3090
Endpoint/providers/microsoft/v1/systemone/v1/systemoneWorkers AI run/v1/decisions/v1/systemone (local server)

Every speed claim in that table is the vendor's own, measured on the vendor's own setup, and they are not comparable with one another. What the table does settle is fit. If your inputs include photos, Decision-1 is out and Clef or OpenAI is in. If you need more than 32K tokens of state, Jev or Clef. If you are already on Azure with Entra ID and Foundry governance, Decision-1 slots in with no new vendor. If the data cannot leave your building, Strands Decider or the open Clef weights. If you want the largest model behind the decision and will pay 2.4 times for it, OpenAI's beta. And the request shape is now nearly a standard: four of the five accept a state and a list of typed questions, so moving between them is mostly a URL change. Our guides cover each one in depth: Clef, Strands Decider, and the open-source Jev family.

Is Microsoft Decision-1 on Amazon Bedrock? The AWS reader's route

No. Decision-1 is an Azure model: it lives in Microsoft Foundry, and Amazon Bedrock has no decision-model category at all as of October 2026. If your stack is on AWS you have three honest options.

First, call Decision-1 anyway. It is an HTTPS endpoint, so a Lambda function or an agent running on Bedrock AgentCore can post to Foundry like any other API, with the cross-cloud latency and a second cloud bill that implies. Second, use Amazon's own decision model. Strands Decider 2B, released by Amazon's Strands Labs on October 1, 2026 by Marc Brooker, Mike Chambers and Fabio Nonato de Paula, is a 2-billion-parameter Qwen3.5-2B with its language head replaced by a pointer head of about a million parameters, open under Apache 2.0 with its training data and scripts, answering in about 115 ms on an RTX 3090 and 153 ms on an M3 Mac. It is not a Bedrock model either; it runs where you run it, on an EC2 instance, in a container, or on the laptop in front of you, and Microsoft thought enough of it to include it in its benchmark set. Our Strands Decider guide covers the install on Windows, Mac, Linux and Kali and a test on a CPU-only laptop. Third, if you want a hosted model with a free tier and no cloud account at all, Cloudflare's Clef-flash on Workers AI is reachable from any AWS Region.

Jake: "Amazon made one of these and gave it away, and Microsoft made one and charges for it?"

Ethan: "Amazon's is small enough to run on your shop PC; that is the gift. Microsoft's is more than four times bigger and runs on their servers; that is the service. For your three hundred messages a week, either costs less than one screen protector."

Can you run Microsoft Decision-1 locally? No, and here is what you can run

You cannot. Microsoft has published no weights for Decision-1, there is nothing on Hugging Face under Microsoft's name, and every description of the model is of a hosted service. Anything labeled "Decision-1 GGUF" is another model wearing the name. The underlying Qwen3.5-9B is open, but the post-training that turns it into a decision model is Microsoft's and is not released.

The open field is wide, though, and most of it runs on ordinary hardware:

Open decision modelSizeRuns onOur guide
Strands Decider 2B2BCPU, NVIDIA or Apple Silicon; pip install strands-deciderStrands Decider
Ollama's Nimble and Tev10.8B to 9BAny machine Ollama runs onOllama decision models
OpenJev, Kev-4B and friendsSmall, mostly under 10Bllama.cpp on Windows, Mac, LinuxOpenJev with llama.cpp
Liquid AI d1-3B3Bllama.cpp or Transformers; license capped at $10M revenueLiquid d1-3B
H2O-Lightning-4B4BvLLM on a data-center or workstation GPUIts Hugging Face card ships a serve.sh
Cloudflare Clef-flash and Clef9B / 27BAbout 41 GB and 85 GB of VRAM at full precisionClef
Quyet-1.0-Large, Rune 26B-A4B, deck-31B26B to 31BOne 80 GB GPUResearch-grade; see their cards

For Jake's kind of workload, the 2B and 4B models are the sensible local choice: they answer in a fraction of a second on a laptop and the data never leaves the shop. The 27B-and-up tier exists for people who want to beat Jev on a leaderboard, and it needs hardware most offices do not own.

What to build with it: six jobs that are choices, not paragraphs

The intended uses are agent controls, model routing, intent analysis, data labeling, AI judging, incident response routing, data validation and content classification. Here is what each looks like as a question, in the shape Decision-1 takes:

  • Inbox and ticket routing (choice): "Which team should handle this?" with billing, engineering and support as options, exactly Microsoft's quickstart. Jake's four doors are the small-business version.
  • Agent guardrails (noul): "Does this tool call touch production data?" or "Are the arguments of this tool call grounded in what the user asked?" asked before the tool runs, with a threshold that blocks or escalates.
  • Grading model output (score): "How well does this answer follow the rubric?" on a four-level scale, which is how Microsoft's Copilot team measures quality and how Microsoft Discovery decides whether to replan.
  • Prompt-injection and content gates (noul): "Does this document contain instructions aimed at the AI rather than the reader?" The model can miss subtle cases, so pair it with your other filters.
  • Model routing (choice): "Is this request simple, standard or hard?" to send cheap requests to a small model and hard ones to a large one. A few hundred tokens at $0.042 per million to save dollars on the other side.
  • Incident triage (choice plus score): classify the incident type, score its severity from cosmetic to critical, and page accordingly, in one request.

The pattern across all six: the application already knows the possible outcomes, and the only question is which one. The moment you need the model to produce something new, a summary, a reply, a fix, you are back to a generative model: anything that creates, summarizes, rewrites or explains content is a job for a generative LLM. The two combine well. Decision-1 decides whether a request is worth answering and who should answer it; the generative model writes the answer.

Microsoft Decision-1 errors, and what each one means

The errors you are most likely to hit, and what causes each:

ErrorCauseWhere to look first
400 Bad RequestInvalid question type or shapeA score with a map instead of a list, a choice with one option, a typo in type
401 UnauthorizedCredential missing, invalid or expiredEntra tokens expire; re-run az account get-access-token. With a key, the header is api-key, not Authorization
403 ForbiddenIdentity lacks access to the deploymentRole assignment on the Foundry resource
404 Not FoundWrong endpoint or pathThe path is /providers/microsoft/v1/systemone, not the OpenAI-style /openai/v1 base
422 Unprocessable EntityDeployment name or question definition not supportedYou sent Microsoft-Decision-1 as model; it must be your deployment name
429 Too Many RequestsDeployment rate limit exceededExponential backoff, or request more quota; batch independent questions into one request

Three more that are not errors but feel like them. A score that comes back as 2.4 on a four-level scale is correct; it is a weighted average, not an index. A choice whose top probability is 0.4 is the model telling you it does not know; route it to a person rather than lowering your standards. And if every answer looks suspiciously confident, check whether your option descriptions are doing the work the state should: a description that quotes typical customer phrasing can match on wording alone.

The limits to know before you rely on it

Plan around these: text and JSON input only, no images or audio; inputs up to 32K tokens; no explanations or rationales in the response; public preview, with future versions planned on MAI and OpenAI models, which means the model behind the same deployment name may change; not designed or evaluated as the sole automated decision-maker in consequential decisions about people; calibration strongest on familiar task types; and sensitivity to wording and option order, which Microsoft reports as a 1.3% average flip rate under perturbation. None of these is unusual for the category. All of them argue for the same habit: label fifty real examples, run them, and read the numbers before you connect anything to production.

Jake's router has been live for a week. The four options and the two extra questions cost him under a cent a month, the "needs a human today" flag catches the angry ones within a minute, and the paragraph-writing model now only writes when a message has been sorted into a door where a reply is wanted. The one he still reads himself is the pile the model marks under 0.6, which turned out to be about one message in twelve.

Frequently asked questions about Microsoft Decision-1

What is Microsoft Decision-1?

Microsoft Decision-1 is Microsoft's first decision model, released October 9, 2026 in public preview on Microsoft Foundry. Given a situation and a fixed set of answers, it returns a calibrated probability for each answer instead of generating text. It handles yes/no, pick-one and rating questions, takes up to 32K tokens of text or JSON, and costs $0.042 per million input tokens.

Is Decision-1 a Microsoft product?

Yes. Microsoft trained, hosts and sells it through Microsoft Foundry. The base model underneath is Alibaba's open-weight Qwen3.5-9B, which Microsoft post-trained for decision scoring; Microsoft plans to rebase later versions on its own MAI models and on OpenAI models.

What is a decision model, or decision AI?

A decision model is an AI model that chooses among options you define instead of writing text. You give it a state and typed questions, and it returns probabilities: one number for a yes/no question, a probability per option for a choice, a position on a scale for a score. The category is also called System One, after Daniel Kahneman's term for fast, intuitive judgment.

How does a decision model make a decision?

In a single forward pass. The model reads the state and the list of options once, and the answer is read directly from its probability distribution over those options rather than generated token by token. That is why it is fast, why output tokens are free, and why it cannot invent an answer that was not on the list.

How much does Microsoft Decision-1 cost?

$0.042 per million input tokens, with output tokens free. A 300-token routing request costs about a thousandth of a cent, and ten million 400-token decisions a month cost about $168. Deployment type and region can change the figure, so check the Foundry pricing page for your deployment.

How much does Microsoft Foundry cost?

Foundry itself bills by what you deploy and use; there is no platform fee for calling a model. For Decision-1 the cost is $0.042 per million input tokens. You need an Azure subscription with a valid payment method, and other models in the catalog carry their own per-token prices.

Is Microsoft Decision-1 free?

No. There is no free tier for Decision-1 on Foundry, and it needs a paid Azure subscription. The price is low enough that a small business's monthly bill is cents. If you want a decision model at zero cost, Cloudflare's Workers AI free tier covers thousands of Clef-flash decisions a day, and Amazon's Strands Decider 2B runs free on your own hardware.

How do I use Microsoft Decision-1?

Deploy it from the Foundry model catalog, copy the resource endpoint and deployment name, then POST a JSON body with model set to your deployment name, a state, and a questions object to /providers/microsoft/v1/systemone, authenticating with a Microsoft Entra ID token or an API key. The response's answers object carries the probabilities.

What question types does Decision-1 support?

Three: noul for yes/no, which returns a probability from 0 to 1; choice for picking one option, which returns the selected option and a probability for every option; and score for an ordered scale, which returns the probability-weighted average of the level indexes plus a probability per level. Several questions can share one request.

What is the difference between Microsoft Decision-1 and Jev?

Both cost $0.042 per million input tokens with output free and use the same request shape. Jev, from TypeSafe, launched first, has a 64K context and an undisclosed base model; Decision-1 has a 32K context, is built on Qwen3.5-9B, and runs inside Microsoft Foundry with Azure identity and governance. Microsoft claims higher accuracy on its own benchmarks but has published no scores.

Microsoft Decision-1 vs OpenAI Decisions API: which is cheaper?

Decision-1, at $0.042 per million input tokens against $0.10 for OpenAI's Decisions API on gpt-6-luna, about 2.4 times less. OpenAI's version is in public beta, accepts inline base64 images, and runs on a far larger model. Both charge only for input.

Is Microsoft Decision-1 open source?

No. Microsoft has not published the weights, and the model is available only as a hosted service through Microsoft Foundry. The base model, Qwen3.5-9B, is open, but Microsoft's decision post-training is not. Open alternatives include Cloudflare's Clef, Amazon's Strands Decider and H2O's Lightning-4B.

Can I run Microsoft Decision-1 locally?

No, because there are no weights to download. For a local decision model, Amazon's Strands Decider 2B runs on a CPU or any GPU, Ollama's Nimble and Tev1 run anywhere Ollama does, and Cloudflare's Clef-flash weights run on a GPU with about 41 GB of VRAM. Any file claiming to be Decision-1 is a different model.

Is Microsoft Decision-1 available on Amazon Bedrock?

No. Decision-1 is on Microsoft Foundry, and Bedrock has no decision models. From AWS you can call Foundry over HTTPS, run Amazon's open Strands Decider 2B on your own instance, or use Cloudflare's Clef on Workers AI from any Region.

What is Microsoft Foundry, and is it the same as Azure AI Foundry?

Microsoft Foundry is the current name of what was Azure AI Studio and then Azure AI Foundry: the Azure catalog where you find, deploy and call models, including OpenAI's, Microsoft's own and partners'. The old portal is now called Foundry (classic). Decision-1 is listed in the Foundry catalog as a "Direct from Azure" model.

How fast is Microsoft Decision-1?

Microsoft reports that its P50 latency is about 35 times faster than GPT-6 Sol and 2.5 times faster than H2O-Lightning-4B, the fastest other decision model it measured, but publishes no millisecond figures. Latency depends on your region and deployment type, so deploy close to your users and measure.

Does Microsoft Decision-1 explain its answers?

No. It returns typed numerical decisions without a written rationale; it does not generate explanations or rationales. Use the probabilities to decide whether to act automatically, ask another model, or send the item to a person, and use a generative model if you need the reasoning spelled out.

What does systemone mean in the Decision-1 API?

It is the endpoint name, /providers/microsoft/v1/systemone, borrowed from the System One label TypeSafe gave this model category when it launched Jev with a /v1/systemone endpoint. Ollama's decision models, Amazon's Strands Decider and llama.cpp use the same name, so client code moves between them with little more than a URL change.

πŸ“š ALSO READ

Where to go next, from the open decision models you can run tonight to the generative models Decision-1 is built to sit in front of:

📌 Bookmark this if you are choosing a decision model this month.

A new category of model arriving every two weeks can feel like one more thing to keep up with. It is simpler than it looks. Decision models do one job, choosing, and they now come in five flavors that cost about the same and speak nearly the same language, so the right question is not "which is best" but "which fits where my code already lives." Jake did not need to understand Qwen or Foundry to get his inbox sorted; he needed four good option descriptions and a threshold. Ethan's advice to him is the advice for anyone reading this: label fifty real examples, run them through two of these models, and let your own numbers decide.

📌 If you keep one line from this page

Microsoft Decision-1 is a hosted Qwen3.5-9B that returns probabilities instead of prose, at exactly Jev's $0.042 per million input tokens; Microsoft claims the top accuracy on 36 benchmarks but has published no scores, so test it on your own fifty examples first.

Call it at /providers/microsoft/v1/systemone with your deployment name as the model, and route anything under your confidence threshold to a person.

Revision note. Written October 10, 2026, the day after Microsoft released Decision-1. "Don't use a sledgehammer to crack a nut," says the English proverb; a yes-or-no question never needed a model that writes essays.

#AI

Related