fullauto.online

← Model leaderboard

Model · stat sheet

JEV 1.13

TypeSafe AI · closed weights·decision
Overall
47
Rank
#34 / 34
Coding35
Terminal30
Reasoning60
Tool use50
Context40
Speed97
CodingTerminalReasoningTool useContextSpeed

Not a chat model — a System One decision model: one typed, calibrated decision per call in 70–500ms, output tokens free. The composite understates it; read the breakdown.

What it is

JEV 1.13 is not a chat model and not an agent model — it is TypeSafe AI's "System One" decision model, released to early access on 15 Sept 2026 and named for William Stanley Jevons (and Kahneman's System One: fast, intuitive judgement). It takes a question with typed constraints — a Choice call with up to 255 named options, a Score call, or a Noul call — and returns one calibrated decision with probability readouts, in 70 to 500 milliseconds end to end. Its probabilities are trained against outcomes rather than next-token likelihoods, so a 0.82 output is meant to be an 82% claim about the world, not a style.

How the composite is built — and what it hides

This board's composite measures general agent work (Coding 24% · Terminal 20% · Reasoning 20% · Tool use 15% · Context 13% · Speed 8%). By those axes JEV scores 47 and ranks last of 34 — and that ranking is honest about what the axes measure while misleading about what JEV is for. Axis by axis:

Coding — 35

measured on real edit-and-run workloads: multi-file changes, tests passing, diffs that hold up under review.

Terminal — 30

measured on long agentic sessions in a shell: planning, running tools, recovering from failed commands.

Reasoning — 60

measured on planning and hard decompositions — the questions a model has to answer before it acts.

Tool use — 50

measured on calling tools correctly: right arguments, right ordering, no hallucinated functions.

Context — 40

measured on usable context under load: what survives at the top of a long session.

Speed — 97

throughput and latency; higher is faster.

Where it fits

Inside other people's loops. Routing, triage, gating, classification, scoring — the narrow decisions an agent makes hundreds of times a day, where a three-second frontier call is slow and expensive and a regex is too dumb. The adoption numbers say teams are already doing this: 8.1B tokens and 7.6K unique users through OpenCode Zen by 24 Sept 2026 (OpenCode data), sitting at rank #38 that week, and 1,600+ community replications catalogued on hanxiao.io as of 26 Sept 2026. TypeSafe ships Python/JS SDKs, an agent-skills repository and an MIT-licensed System One adapter for A/B-ing JEV against an LLM router.

Limits

Read TypeSafe's own jaggedness documentation before trusting it on anything it hasn't seen: it reads questions literally, struggles with maths and exact magnitudes, makes date/time comparison errors, and stumbles on indirection. An independent re-scoring (warmersun.com, Sept 2026) found ~67.8% mean agreement with an agent judge on TypeSafe's own four eval workflows, against ~73–74% for the best frontier LLMs — so for judgement-heavy decisions the frontier is still ahead; JEV's case is that it's 193.6× faster and 444.6× cheaper (vendor claim, Sept 2026) at "good enough and calibrated".

Price and access

$0.042 per 1M input tokens; output tokens free. That makes the blended rate $0.0315 per 1M — the cheapest paid call on this board by roughly an order of magnitude. Closed weights; available via OpenCode Zen and TypeSafe's API.

Alternatives on this board

  • Claude Haiku 4.5 — when the narrow call still needs language generation.
  • Grok 5 Mini — a small general model when the decision needs a paragraph of reasoning.

Sources

Scores are fullauto.online's composite index (0–100): Coding 24% · Terminal 20% · Reasoning 20% · Tool use 15% · Context 13% · Speed 8%. Editorial, not a vendor benchmark; 2026 tiers are early reads. Last scored 26 Sep 2026 · back to the leaderboard.