Microsoft-Decision-1 vs Jev: how the decision models differ and when to use each one

Microsoft-Decision-1 and Jev cost exactly the same, $0.042 per million input tokens and free output, so choosing between them depends on context, versioning, documentation, and accuracy with your own data. Microsoft launched its decision model on October 9, 2026, available that same day on Microsoft Foundry and OpenRouter, and enters a category that Jev, from TypeSafe, has led since September.

What is Microsoft-Decision-1?

Microsoft-Decision-1 is a small decision scoring model: you give it a fixed set of options and it returns a calibrated probability for each one, not text. Microsoft built it with post-training of Qwen3.5-9B to score in a single pass, and supports yes/no questions, multiple choice and rating, plus scoring AI responses with rubrics and agent actions. Microsoft targets it at routing, classification, prioritization, verification, and agent control. The model page on OpenRouter is explicit about what it doesn’t do: open-ended generation, conversation, translation, or summarization.

If the category is new to you, we explain how a decision model works, and why it’s not an LLM with a JSON schema, in Jev by TypeSafe: how AI that makes decisions without generating text works.

How does Microsoft-Decision-1 differ from Jev?

They have the same price and almost everything else is different.

Microsoft-Decision-1 Jev 1.13
Manufacturer Microsoft TypeSafe
Launch Oct 9 2026 Sep 18 2026 (OpenRouter listing)
Input price $0.042 USD / 1M tokens $0.042 USD / 1M tokens
Output price Free Free
Context 33K 64K per request (32K for state + longest question)
Base model Qwen3.5-9B post-trained (will change) Not published
Versioning Weights updated continuously, same API shape Versioned ID jev-1.13.0, plus jev-latest alias
Where it’s used Microsoft Foundry, OpenRouter (served by Azure) TypeSafe API, OpenRouter (served by TypeSafe)
API documentation No published request schema Quickstart with complete request and response schema
P50 latency on OpenRouter 0.21 s 0.19 s

Prices, context, and latency according to OpenRouter and TypeSafe documentation as of October 10, 2026. Latency is OpenRouter’s live measurement and changes daily.

Is Microsoft-Decision-1 faster than Jev?

With available data today, no: in OpenRouter’s live measurement both hover around 0.2 seconds median, with Jev slightly ahead. Microsoft’s benchmarks place Microsoft-Decision-1 about 35 times faster than GPT-6 Sol and 2.5 times faster than H2O-Lightning-4B, second in their comparison, but these are Microsoft measurements, not independently reproduced. Microsoft says it also evaluated Jev on accuracy and calibration (indicated by an editor’s note in the announcement), but at the time of publishing this note those figures only appear as images in the post, without a text version we could verify and cite.

The rest of the claims are also Microsoft’s: higher accuracy on 36 benchmarks and nearly 150,000 questions excluded from training, and decisions that change in only 1.3 percent of reformulated or reordered inputs. Internally, Microsoft reports that Xbox Research used it to classify over 10,000 player comments by topic, with competitive quality versus GPT-6 Sol, over 14 times faster and at a fraction of the cost.

How to use Microsoft-Decision-1 on OpenRouter?

It’s used via OpenRouter’s Decisions API, not the chat completions endpoint: OpenRouter warns that chat completions SDKs don’t work with this model. In OpenRouter’s SDKs (Python, TypeScript, and Go) the call is in alpha.decisions.create and receives three mandatory parameters: model, questions, and state. The model identifier is microsoft/microsoft-decision-1. The alpha namespace indicates that the API itself is still preliminary.

As of October 10, 2026, neither Microsoft nor OpenRouter had published the questions schema or the response object schema for Microsoft-Decision-1, so we don’t reproduce a request here: a guessed schema would fail silently for anyone copying it.

Jev’s request shape, on the other hand, is documented in TypeSafe’s quickstart. state contains the text being evaluated and questions is a map of names you choose to typed questions: choice (choose a labeled option), score (rate on an ordered scale), or noul (a yes/no type value):

{
  "model": "jev-latest",
  "state": "The deploy failed on the payments service and customers cannot check out.",
  "questions": {
    "urgency": {
      "type": "noul",
      "instructions": "Does this message describe an incident that needs immediate attention?"
    }
  }
}

The request goes to POST https://api.typesafe.ai/v1/systemone with a bearer token. The response returns an answers object with a key for each question, with the option or score, a confidence value, and the probabilities per option.

Both models share the same idea in three parts (a state, named questions, probabilities back), so it’s likely that swapping one for the other within OpenRouter’s Decisions API is cheap. Until Microsoft documents its question types, “likely” is as far as we get.

How much does Microsoft-Decision-1 cost versus Jev?

Grego: With identical list prices, cost stops differentiating one from the other and becomes an argument in favor of the category. Output is free and a decision takes just a few tokens anyway, so you pay for what you send. At $0.042 per million input tokens, a million decisions with 500-token inputs costs about $21, on either one. For a CTO, the real cost question isn’t the bill, it’s what you give up. Jev lets you pin jev-1.13.0 and know that the model behind your router won’t change without notice. Microsoft-Decision-1 tells you the opposite on purpose: weights are updated continuously and the base will move from Qwen3.5-9B to MAI and OpenAI models “soon”. If your decisions feed an audit log or a compliance control, that difference weighs more than any benchmark. If your stack already runs on Azure and procurement prefers a single vendor, Foundry is the best argument for Microsoft.

When to use a decision model instead of an LLM?

When your software already knows the possible answers and only needs to choose, score, or approve one. Routing a request to a model or team, labeling comments, checking if an agent’s next action should proceed, filtering content, scoring an AI response with a rubric: all of those are decisions, and a decision model returns you a probability you can apply a threshold to instead of text you have to parse. For classification tasks with fixed labels, it’s the natural tool. Use an LLM when the result is the text itself or when you don’t know the options beforehand.

Within the category, choose based on the constraint:

  • Long inputs (logs, lengthy tickets, full conversations): Jev’s 64K versus Microsoft-Decision-1’s 33K.
  • You need a stable model behind a decision you might have to explain later: Jev’s versioned IDs.
  • Organizations working on Azure: Microsoft-Decision-1 via Foundry.
  • Spanish content: neither provider guarantees parity with English. TypeSafe indicates Jev achieves its best accuracy in English; Microsoft says its benchmarks include multilingual tasks. Test it with your own text.

What other decision models are on OpenRouter?

Microsoft and TypeSafe are not alone. OpenRouter’s decision model ranking for the week of October 9, 2026 shows Jev 1.13 well ahead by volume, followed by newcomers like OpenAI’s GPT-6 Luna Decisions, Cloudflare’s Clef and Clef Flash, Perplexity’s Decider, and Liquid’s d1. Microsoft-Decision-1 wasn’t yet in the top ten on that date: it had been available for a day. Two of the largest AI labs launching decision APIs weeks apart is a reasonable signal that the decision model is becoming a standard layer of agent stacks, not a TypeSafe niche.

How to test a decision model with your own data?

Accuracy on someone else’s benchmark says little about your labels. yoDEV Decisions runs Jev on public datasets or your own CSV and reports accuracy, calibration, latency, and cost per thousand decisions. Our first round of results compared Jev with an open decision model on four datasets, two of them in Spanish. As of October 10, 2026, Microsoft-Decision-1 is not yet in the tool; when we pit it against Jev, we’ll add the results here.


And you? What decision in your application would you take away from an LLM today to give it to a decision model?

Which decision model would you try first?
  • Jev
  • Microsoft-Decision-1
  • An open model locally
  • I prefer another option (tell us which)
0 votantes

Comment below or, if this article reached you by email, reply directly to the email: your response publishes here.

Related: Jev by TypeSafe: how AI that makes decisions without generating text works