Most chatbot and intake workflows hide a lot of small decisions: is this a sales lead or a support request, is it urgent, which team should get it, is the answer grounded in the source? Asking a large language model to write a paragraph to answer each one is slow and costly. A new category of model does only the deciding. It reads the situation, you supply the options, and it returns a probability for each. Two big vendors shipped one this week.
What Microsoft released
Microsoft’s post calls decision models “an important new category in AI,” built to deliver structured outputs that software can act on instead of text. Microsoft-Decision-1 is available in Microsoft Foundry and through OpenRouter. Its Foundry model page says it is built on Alibaba’s open-weight Qwen3.5-9B and post-trained by Microsoft, and supports yes/no, multiple-choice, rating, classification and rubric-based questions. It lists the model as generally available, with text input, JSON output and a 32,768-token context window. The page is explicit about limits: it does not generate text or explanations, accepts no images, audio or video, and should not be the sole basis for decisions about credit, employment, housing, insurance, education, healthcare or legal rights. Microsoft’s post adds that it will soon rebase the model on others, including Microsoft AI (MAI) and OpenAI models.
“Input tokens cost $0.042 USD per million tokens. Output tokens are free.”Microsoft, Introducing Microsoft-Decision-1, October 9, 2026
The speed and accuracy claims are Microsoft’s own. Its post says Decision-1 had the highest accuracy in its 36-benchmark comparison of nearly 150,000 questions, that it was the fastest measured, 2.5 times quicker than the runner-up, H2O-Lightning-4B v1.1, and that its P50 latency is about 35 times faster than GPT-6 Sol. Those are Microsoft’s benchmarks, not an independent test, and we haven’t reproduced them.
What OpenAI released
OpenAI’s community announcement from October 6, 2026 says the Decisions API makes decisions up to 10 times faster than GPT-6 Luna through the Responses API. The docs describe a dedicated POST /v1/decisions endpoint and three question types: a predicate (the probability that a statement is true), a choice (one option from a fixed set, with confidence scores) and a score (a rating against ordered levels). It accepts text and images. The docs say the API is in public beta, that OpenAI expects to reach general availability “in the coming weeks,” and that gpt-6-luna is the only model available now. They also say it supports Zero Data Retention and HIPAA use for eligible customers, with data residency and regional processing in the United States and Europe (EEA + Switzerland). Regional processing premiums apply. Canada isn’t mentioned.
Side by side
| Item | Microsoft-Decision-1 | OpenAI Decisions API |
|---|---|---|
| Status | Generally available on Foundry | Public beta; GA expected in the coming weeks |
| Input price | $0.042 | $0.10 (plus regional processing premiums where they apply) |
| Inputs | Text only, up to 32K tokens | Text and images |
| Underlying model | Post-trained Qwen3.5-9B | gpt-6-luna |
| Regions documented | None on the pages we read | US and Europe (EEA + Switzerland) |
As an illustration from the listed rates, 10,000 routing requests of 1,000 input tokens each is 10 million input tokens: $0.42 at Microsoft’s price or $1.00 at OpenAI’s, before any regional premium. That is our arithmetic, not a measured bill, and real requests include your instructions and options. Both are small next to a chatbot reply. Some developers on OpenAI’s forum noted that the Decisions API has no cached-input discount yet, which matters if you resend long instructions.
How Alberta teams could use one (our advice)
What follows is our advice, not Microsoft’s or OpenAI’s. A decision model fits the front door of a chatbot or intake workflow. Ask it whether a message is a new lead, a booking, a complaint or spam; whether it is urgent; which person should get it; and whether a drafted reply is supported by the source. Then act automatically when confidence is high and send ambiguous cases to a person. Microsoft’s own page describes this confidence-based escalation.
Before you rely on one, test it. Collect a few hundred of your own past messages, label them by hand, and compare the model’s choices with yours, including how often a confident answer is wrong. Keep humans in the loop for anything consequential, such as credit, hiring, housing or insurance, which Microsoft says its model isn’t designed to decide alone. And check data residency for both vendors. Neither documents a Canadian region, so keep personal information out of trials until you have confirmed where it is processed. Our AI receptionist Alberta page covers call and chat intake, AI automation for Alberta trades covers the lead-to-quote flow, and an AI consulting engagement can build the labelled test set.
