Designing with AI

Designing with AI

How Jev works: calibrated decision models

Issue #69 | Jev from TypeSafe AI returns typed decisions with probabilities and generates no text. This article explains how it works and what it trades away, and tests the idea on open Qwen models.

Victor Dibia, PhD's avatar
Victor Dibia, PhD
Sep 28, 2026
∙ Paid

TypeSafe AI recently released Jev, a model that promises low latency (TypeSafe reports 70 to 500 ms per call) and inference so cheap that output tokens are not metered[1][2]. At the same time, TypeSafe says Jev “achieves similar levels of intelligence on System One tasks compared to existing LLMs”[1]. This is fascinating.

A claim like this raises some natural questions. How does it work: what is the architecture, and how does it differ from the typical autoregressive model? What are the key tradeoffs, and when does it make sense to use a model like Jev? Can I build, train or use one for my own use cases, and what does performance look like?

This article works through those questions in order: how Jev and models like it work, experiments I ran on open Qwen models to test the idea, and what we know about Jev itself.

An interactive version of this post is on my website here.

How decision models like Jev work

Jev is a “System One” model (the name echoes Kahneman’s System 1, fast intuitive judgement) from TypeSafe AI, described as “a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out”[1]

Figure 1. The running example goes through the four steps on an open 7B model (the setup is in the experiments section). In step 3, the raw probability of edit personal details is only 0.36, because the model split the first word between "Edit" and "edit"; after normalising in step 4 it is 0.98.

Models like Jev provide output in four steps:

  1. Start with a decision. Given a state (the text the decision is about) and a list of options, pick one. For example: a customer writes to a bank, “I want to change my address.”, and we need to pick their intent from a list of 10 potential intents.

  2. Turn it into a sentence to complete. Write a prompt that lists the options and ends exactly where the answer goes:

Classify the customer message into exactly one of these intents.
Intents: Refund not showing up; age limit; edit personal details; ... (10 in all)

Message: I want to change my address.
Intent: ___
  1. Score every option. Append each option to the prompt and ask the model how likely that exact text is as the completion (the product of the probabilities of its tokens, the word pieces a model reads and writes). The options are known in advance, so nothing is generated, and all of them are scored in parallel. We call this step scoring: reading the model’s probabilities for text we supply, instead of generating text.

  2. Normalise. Divide these probabilities by their total so they add up to 1. The results are the option probabilities, and the highest is the answer.

Jev packages this as an API with a few distinctive traits:

  • No text generation. “System One models do not write replies, produce code, or generate explanations of their reasoning”[7].

  • Many questions, one read of the input. The model “processes the state once and evaluates all questions against it in parallel”[2].

  • Three question types, all lists. A choice picks one of up to 255 options[8]. A score rates the state on 2 to 10 ordered levels[9]; I call it the rating-scale question, because “scoring” already means step 3. A noul (”short for bernoulli”[10]) answers yes or no with one probability.

  • Fast and cheap. TypeSafe reports 70 to 500 ms per call, against 3 to 329 s for frontier large language models (LLMs) on the same tasks[1], and charges $0.042 per million input tokens with output free[2].

Note: TypeSafe has not disclosed how Jev is implemented, but are likely to follow as similar approach above.

Calibrated decision models

While Jev is one implementation, we can think of this class of models as calibrated decision models. Decision means the goal is to select from a list of options, not to generate text. Calibrated refers to the model’s confidence, the probability it gives to the option it chose. A model is calibrated when its confidence matches its accuracy: if it gives a thousand answers at 90% confidence, about 900 of them are right.

In my opinion, confidence is important here because it lets us build systems and apps that do something ML models are notoriously bad at: say “I don’t know” or “I’m not sure” when confidence is low.

The avid reader is likely getting suspicious here: calibration is a hard problem. Early evidence suggests Jev is well calibrated. One independent benchmark measured its expected calibration error (ECE) at 0.0588 on 662 prompt-injection messages[12], meaning its confidence and its accuracy differed by about 6 percentage points on average. A stock model came close on the same task, as What we know about Jev shows.

Use cases for decision models

Decision models have a few properties that make them useful in production. First, the model can only select from a provided list of options, which removes a whole class of hallucinations and parsing failures: there is no answer that is not one of your options. Second, when the model is well calibrated, its confidence lets you build more thoughtful experiences that reflect certainty/confidence. The system can act when it is sure, ask a clarifying question when it is not, or hand off to a human. Third, it is fast and cheap enough to call on every step of a workflow, not just occasionally.

All of this suits tasks with structured output, and many production processes have that shape. A decision model fits when:

  • The answer can be listed: a category, an action, a yes or no, a level on a scale.

  • The decision is frequent or on the hot path, so latency and cost matter.

  • One judgement is enough. Everything happens in one pass, so tasks that need step-by-step reasoning are a poor fit; TypeSafe lists arithmetic and counting among Jev’s weak spots[13].

  • You can check the confidence against labelled examples from your own task.

Computer use. In the computer-use chapter of my book, Designing Multi-Agent Systems[14], latency and cost are two challenges I highlight: complex tasks need many model calls, one per action, so delays and costs compound. One of the mitigations the chapter points to is smaller, faster models tuned for interface understanding and action prediction. A decision model fits that description: it reads the state of the page and selects the next action from the list of possible actions. Jev takes text only[2], so here the state would be the page’s text rather than a screenshot.

Agent evaluation. Teams running agents need both fleet-wide metrics and task-specific ones, and observability platforms need to score agents quickly. Evaluation matters even more inside optimisation loops, where an evaluator is called thousands of times, and a fast yes/no or rating-scale question is a good fit for both.

The one caveat is evaluation of the decision model itself. We need evals that tell us not just whether it is accurate, but whether it is calibrated, measured on labelled examples. The experiments below show why: accuracy can look fine while the confidence is badly off.

Why scoring is fast

To see why scoring works, and why it is fast, we first need to look at how a language model produces text.

An autoregressive model takes text in and produces text out. The input is split into tokens, pieces of text such as a word or part of a word, from a fixed vocabulary (152,064 of them for Qwen2.5-7B, used later[15]). Each token becomes a vector, and the vectors pass through a stack of identical layers. After the last layer, every position is turned into one number per vocabulary entry, and a softmax turns those numbers into probabilities: the model’s prediction of the next token. One run of the input through all the layers is a forward pass.

One forward pass produces a next-token probability distribution at every position of the input, all at once.

There are two ways to turn those probabilities into a decision: generate the answer, or score the options.

Generating a decision

To generate text, we run a loop around the model: run a forward pass, pick a token from the distribution at the last position, append it to the input, and repeat until the model produces a stop token. There are two ways to get a decision this way:

  1. Generate a label. Ask for the answer. The model generates a few tokens, for example edit personal details, and you match that text to one of your options.

  2. Generate probabilities. Ask for a JSON object with a probability for each option. The model generates it token by token: every option name and every number.

This loop is serial and hence slow: each forward pass needs the token the previous one picked, so the passes cannot run at the same time. Modern models make each pass cheaper, mainly with a KV cache, which stores the work already done on earlier tokens so that each pass only processes the one new token. But the passes still run one after another. A short label takes only a few generation steps, while a probability for every option takes one step per token of the whole answer. Both also depend on the model behaving: a generated label can match no option, and generated JSON may not parse.

Scoring a decision

The model first reads the prompt: one forward pass over all its tokens, which also fills the KV cache. Unlike generation, this pass runs in parallel, because every token of the prompt is known before we start. Scoring relies on the same fact: the options are known in advance too. A second forward pass reads every option side by side, each one reading the cached prompt, and normalising their probabilities gives the answer. That is two passes, however long the options are, and scoring cannot return anything off the list, because you supplied the text of every option.

Figure 2. The intent question from Figure 1 is answered both ways with the same prompt, on the 7B instruct model. Generating takes one pass per token of the answer, each waiting for the one before; scoring takes two, whatever the answer. Shown: mode: generate the answer.

What scoring gives up

  • Options are judged separately. Each option is scored on its own, and one option cannot see another. The model never compares edit personal details with transaction charged twice directly; the comparison happens only when we normalise.

  • Different wordings compete. A score is the probability of a piece of text, not of an answer being right. In Figure 1, the model split its first word between “Edit” and “edit”, and only one of them is counted[16]. Word options the way the model would say them.

  • Cost grows with the number of options. Scoring reads every token of every option, so its cost grows with questions × options. A generated label costs about the same however many options there are: the model compares the options inside one forward pass and generates only the winner. Giving each option a one-token code, such as a letter, and reading only that avoids the second pass[3][4]; Can a cheaper readout do as well? measures what it costs.

Experiments with Qwen2.5 1.5B and 7B

To test all of this, I ran quick experiments on Modal using open Qwen2.5 models. The goal was not the fastest possible numbers, but to compare the approaches under identical conditions and report what we see and the conditions under which we see it[5].

  • Models: Qwen2.5 1.5B and 7B[15], each as a base model (trained only to predict the next token) and an instruct model (further trained to follow instructions and chat, the kind most APIs serve).

  • Hardware and software: one NVIDIA A10G GPU on Modal, Hugging Face transformers 4.46.3, 16-bit weights, one request at a time. No serving engine such as vLLM, which would make generation faster but add its own variables.

  • Scoring code: my own script, implementing the two passes described in Scoring a decision. An option’s option score is the sum of its tokens’ log-probabilities (the log of their product), and normalising the scores gives the option probabilities. I checked it against a slow version that recomputes the prompt for every option, on a 0.5B model; the two agreed to within 0.00003 in 32-bit arithmetic, and in the 16-bit arithmetic used for the runs they picked the same answer.

  • Data: for the intent question, 1,200 Banking77 test messages, sampled with a fixed seed, with 2, 10, 25 or 77 options (for fewer than 77, the correct intent plus randomly chosen wrong ones)[17]. For a yes/no question, “is this message a prompt injection?”, all 662 messages of the deepset prompt-injections dataset[18].

  • Cost: about $4 of GPU time across every run, as billed by Modal. The code, the raw results and a cost ledger are in the site’s repository.

Is scoring faster than generation?

On the 7B instruct model, reading a 339-token prompt took 0.115 s, and each generated token added 36.9 ms. A generated label (11 tokens) took 0.47 s, and generated probabilities for 10 options (125 tokens) took 4.57 s. Scoring the same question took 0.23 s.

In every setting where generated probabilities were run, scoring was 7× to 54× faster than generating probabilities on the 7B, and 13× to 132× on the 1.5B. With 77 options, generated probabilities would be about 890 tokens, roughly 33 s on the 7B by the fitted line (not run).

Against generating a single label, the answer depends on the number of options, the trade-off from What scoring gives up. With 5 questions about the same state on the 7B, scoring was 9.8× faster at 2 options, 3.5× at 10 and 1.7× at 25, and at 77 options it was slower (0.5×). A longer state narrows the gap too: 3.5× at 256 state tokens, 2.6× at 1,024 and 1.4× at 4,096, because processing the state starts to dominate.

Figure 3. Each point shows how many times faster scoring was than generating a label, at 2, 10, 25 and 77 options. A request asks several questions about the same 256-token state, and the state is read once for all of them, which is why more questions favour scoring. Shown: questions: 5 questions; model: Qwen2.5 7B.

Part of the slowdown at many options comes from our setup, not the technique. Our code makes a copy of the prompt’s KV cache for every option before reading it; serving engines such as vLLM share a single copy, so they would move the break-even point to more options. A serving engine would also make generation several times faster, so all the speedups here are upper bounds. Either way, every token of every option still has to be read, so scoring’s cost keeps growing with the number of options. More questions about the same state, by contrast, are cheap: the state is read once, and each question only adds its own options.

Is scoring as accurate as generation?

Scoring was at least as accurate as generating a label.

At 10 options the two approaches picked the same answer on 98.5% of messages. Generated labels were measured on a subset of 400 messages, scoring on all 1,200. Model size matters more than the approach: at 77 options, scoring was right 39% of the time with the 1.5B instruct model and 64% with the 7B.

Are the models calibrated?

Base models were close to calibrated; instruct models were not. At 10 options, the base models’ ECE was 0.05 (1.5B) and 0.02 (7B), while the instruct models’ was 0.18 and 0.10. Scale helped but did not fix it: at 77 options, the 7B instruct model’s average confidence was 95%, and it was right 64% of the time. The base models go against earlier findings that neural networks and language models are poorly calibrated[19][20]; the instruct models agree with them, and match what OpenAI reported for GPT-4: well calibrated after pretraining, and significantly less so after the post-training that turns it into a chat model[21].

Figure 4 shows where the instruct model goes wrong. We sort its answers into groups by confidence and compare each group’s average confidence with the fraction that were right, which gives a reliability diagram; a calibrated model sits on the diagonal. The instruct model puts 1,152 of its 1,200 answers above 90% confidence, at an average confidence of 99.8%, and only 90.6% of those are right.

Figure 4. On the 7B instruct model with 10 options, 1,152 of 1,200 answers have confidence above 90% as the model gives them, and 90.6% of those are right. After temperature scaling (below), 872 do, and 97.4% of those are right. Accuracy is 88.8% in both views. Shown: model: 7B instruct; probabilities: as the model gives them.

Calibrating with temperature scaling

Temperature scaling divides every option score by one fitted number, the temperature, before normalising. It never changes which option wins, so accuracy stays the same, but it spreads the confidence out. Fitted on 600 labelled messages and tested on the other 600, it brought the 7B instruct model’s ECE from 0.10 to 0.03, about as good as the 7B base model (0.02). Because it keeps the order, it only moves where the 90% line falls. Before scaling, about 108 of the answers above 90% were wrong; after, about 23. The catch is that it needs labelled examples from your own task.

One more choice affects calibration: how an option’s tokens are combined into a score. Averaging their log-probabilities instead of summing them made the 7B base model underconfident (44% average confidence, 85% right), while summing made the instruct model overconfident. Pick one way of computing the score, and calibrate that one.

User's avatar

Continue reading this post for free, courtesy of Victor Dibia, PhD.

Or purchase a paid subscription.
© 2026 Substack Inc · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture