fullauto.online

Models

JEV: what TypeSafe's decision model actually does

A decision model, not a chat model: one typed, calibrated call in 70–500 milliseconds, output tokens free. What it is for, where it breaks, and why agents have already run 8.1 billion tokens through it.

Published
26 Sep 2026
Reading
9 min
Class
models

half-life 60dfrom 26 Sep 2026

An agent makes hundreds of decisions a day that are not really language problems. Is this error worth retrying? Which of these three files did the user mean? Does this diff clear the bar? Every one of those can be a three-second, three-cent call to a frontier model — which is one of the ways agent bills reach four figures while the loop crawls.

JEV is TypeSafe AI's answer to the boring half of that work: a model built to make one narrow, typed decision per call, in under half a second, for a rounding error. It has been available since 15 Sep 2026, and agents have already pushed 8.1 billion tokens through it on OpenCode Zen alone. Here is what it actually is.

What it is

JEV 1.13 — named for William Stanley Jevons, and for Kahneman's System One: fast, intuitive judgement — is a decision model. It does not chat, does not write code, does not plan. You give it a question with typed constraints and it returns one decision. A Choice call picks among up to 255 named options; a Score call grades one thing against stated criteria. High-cardinality choices run through a second scoring stage so 255 options stay tractable.

The output that matters is not the pick but the calibration. JEV's probabilities are trained against outcomes rather than predicted from word frequency, so a 0.82 output is meant to be an 82% claim about the world. When an agent is deciding whether to escalate, retry or drop, that number is the useful part.

How it decides

You encode the state once; each question then runs as its own isolated branch over that shared state, so one request can carry many independent decisions without them contaminating each other. The budget is deliberately small — about 32K tokens of state plus the longest question, 64K per request. This is not a model you point at a repository. It is a model you point at a decision.

The numbers

MeasureFigureSource, date
Latency70–500 ms end to endTypeSafe, Sept 2026
Price$0.042 / MTok in, output freeTypeSafe, Sept 2026
Official evals61.7–76.0% accuracy at 0.3–0.5 s/caseTypeSafe, Sept 2026
Vendor cost claim193.6× faster, 444.6× cheaper than frontier LLMs on their workflowsTypeSafe, Sept 2026
OpenCode Zen usage8.1B tokens, 7.6K users, rank #38 that weekOpenCode data, 24 Sep 2026
Replications1,600+ cataloguedhanxiao.io, 26 Sep 2026

Output tokens being free is the quiet part. A decision is a few tokens long, so the effective cost of a call is a rounding error — and the comparison in the fourth row is the one to hold on to. The frontier is not being beaten here. It is being routed around, most of the time.

Where it fits

Inside loops. Routing, triage, gating, classification: the narrow decisions an agent makes hundreds of times a day. TypeSafe ships Python and JavaScript SDKs and an agent-skills repository, plus an MIT-licensed System One adapter that runs JEV and a frontier model behind the same interface — which is the right way to find out whether it earns its place in your stack rather than in a blog post.

Where it breaks

TypeSafe documents the failure modes more honestly than most vendors (their model-jaggedness notes, Sept 2026): JEV reads questions literally, is weak at maths and exact magnitudes, mishandles date and time comparisons, and stumbles on indirection. The mitigations are documented too, but the shape of the advice is “ask a bounded question and do not reconstruct magnitudes from the scores”.

An independent re-scoring (warmersun.com, Sept 2026) reran four of TypeSafe's own eval workflows against an agent judge and found roughly 67.8% mean agreement for JEV, against roughly 73–74% for the best frontier LLMs. So on judgement-heavy decisions the frontier is still ahead. JEV's case is not that it is smarter. It is that it is 193.6× faster and 444.6× cheaper at “good enough, and here is how sure I am”.

JEV on our leaderboard

We have put JEV on the model leaderboard as a proper entry: rank 34 of 34, composite 47, scored 26 Sep 2026. The composite measures general agent work — coding, terminal sessions, tool calls — and on those axes JEV scores what it is: a model that never writes code or drives a shell. Its speed score, 97, is the highest on the board. If you read the rank and stop, you will conclude JEV is the worst model we track. Read the stat sheet and you will see it is the only model we track that does this job at all.

If you want to try it

Reach for it where a decision is narrow, frequent and priced per call. Skip it where the answer needs to be prose, the question needs arithmetic, or a wrong answer costs more than a rounding error. It is served on OpenCode Zen and TypeSafe's own API. The cheap experiment is the System One adapter: run your existing router's calls through it for a day and count the agreement rate.

The frontier is for the hard 20%. Most of the other 80% is a decision problem wearing a chat interface's clothes.

Sources