Microsoft-Decision-1's Qwen3.5-9B base is a stopgap
Microsoft says the Qwen3.5-9B version of Decision-1 is a stopgap, because it will rebase the model on its own and OpenAI models.
| Microsoft-Decision-1 | |
|---|---|
| base model | Qwen3.5-9B |
| accuracy | 83.5 percent on 36 benchmarks |
| confidence score | 92.2 percent |
| input price | $0.042 per million tokens |
| output price | free |
| latency vs H2O-Lightning-4B | 2.5x faster |
| latency vs Jev | 2.8x faster |
| cost vs GPT-6 Sol | more than 20x cheaper |
The model is offered through Microsoft Foundry. Input tokens are priced at $0.042 per million, and output tokens are free. The claimed accuracy is 83.5 percent across 36 benchmarks. That score does not establish reliability beyond the test set, nor a durable edge over Jev, H2O-Lightning-4B, or GPT-6 Sol.
The pricing is the buyer-facing constraint. The input price and free output tokens put pressure on OpenAI's GPT-6 Sol and on Jev's latency pitch. Microsoft says it will rebase Decision-1 on its models and OpenAI models, so the Qwen3.5-9B version is a stopgap rather than a platform.
The missing details are the independent context window, weight size, and benchmark conditions. The document names Qwen3.5-9B as the base, but it does not show whether that accuracy survives outside Microsoft's chosen 36 tasks.
Lead with the rebase, keep the price as the buyer-facing constraint, and flag the missing context window, weight size, and benchmark conditions.
After @perplexitydevs. We did not report this. The pictures, if any, are theirs. Also filed by The Decoder.

