Jev Is Fast. But Can You Trust Its Answers?

summarized

TLDR

Jev, Typesafe's classification model, is fast and cheap but has a serious calibration problem: on a benchmark, it failed to admit ignorance on no-fit items 50% of the time, while most LLMs did so 97-100%. This means routing code using its confidence threshold could confidently send mismatched queries down the wrong intent flow. A conflicting pilot study shows better calibration, leaving the reliability question unresolved.

Key points

Jev's confidence calibration error was 0.246, the worst among models tested.

On no-fit items, Jev admitted ignorance only 49.7% of the time, versus 97-100% for most LLMs.

Jev costs $0.07 per thousand decisions, about 3x cheaper than the cheapest LLM tested.

Jev handles up to 255 options; beyond that it returns an HTTP 400 error.

A conflicting pilot study by Abdel Stark found Jev with 87% accuracy and calibration error of 0.054.

Tools mentioned

Techniques

  • confidence thresholding
  • calibration error measurement
  • option shuffling robustness test
  • no-fit detection
Transcript (captions)

0:00 Picture a banking bot getting this message. What time does your leads branch close on Sunday? Your intent list has 77 options. Lost card, failed transfer, exchange rates. None of them

0:10 is branch hours. What you want is for the bot to say it's not sure and hand the message to a person. An independent test threw questions like that at nine models. Seven of them backed off almost

0:21 every time. One backed off only about half the time. The rest of the time it picked an option with a confidence above 0.5, which is where its own docs say to ask instead of guess. A lot of routing

0:32 code works exactly like that. If confidence is low, escalate. If it's high, act. A confident wrong answer never reaches the escalate branch. Instead, it heads down the lost card

0:42 flow and someone who only wanted opening hours is now answering security questions. Behind that result is a model that works nothing like a chatbot. You hand it a message and a list of options

0:52 and it makes one pass. It doesn't write a single word. You get a probability back for each option plus a confidence score. Roughly confidence tracks how peaked that spread is. If one option

1:03 soaks up nearly all the probability, the spread is peaked and confidence is high. If it's smeared across a dozen options, confidence is low. So for our branch hours message, a well- behaved model

1:14 should give you something flat and smeared because nothing deserves the weight. Meet Jev from Typesafe. Their docs tell you what to do with that number. In the example code, an if

1:24 statement checks whether confidence is below 0.5 and the comment inside reads, "The customer hasn't said what they want. Ask, don't guess. Hold on to that 0.5."

1:34 The independent benchmark coming up uses the same line. So when Jev gets graded on admitting it doesn't know, it's held to Typesafe's own rule. Typesafe's launch post goes much further. According

1:44 to the post, Jev is 193.6 six times faster and 444.6 times cheaper than GPT6 Astra and Fable 5.1. Both are Typesafe's own numbers and nobody's reproduced them independently. They even say some bias

1:58 could exist and the numbers may be on the higher end which is more than most launch posts admit. For a test Typesafe didn't write, there's Nibsard's decision model benchmark on GitHub which puts Jev

2:09 up against other models. No vendor sponsors it. The costs are paid and listed and the raw logs are public. The measured API bill came to $28.34, less than lunch for the meeting where

2:20 you'd pick a model. Everything here comes from version two of the report, dated the 18th of September, 2026 in the repo's history. Version two replaces the earlier ones, and those are still in the

2:31 repo unchanged, so you can see what got corrected. Jev ran against eight LLMs at temperature zero with thinking turned down as low as each one allows. On the 77wave banking test, Jev got 76.3%

2:45 right. Midpack, in other words, GPTOS 120B led with 81.3 and GLM 5.3 was right behind at 80.4. That's 200 items with no confidence intervals. So, nobody's checked whether a 5point gap would hold

3:00 up. Jev answers in a median of about 264 to 276 milliseconds and that stays flat whether you give it two options or 255. On banking it was 274. GPTOS 120B running on Cerebras, a provider

3:17 built for fast inference came in at 331. Call it 1.2 times faster. So where does a giant multiplier come from? The report says it only shows up against LLM left in their slowest default mode. Even with

3:30 thinking at minimum, a couple of models here were slow. Claude Haiku 4.5 took 4360 milliseconds, about 16 times. Jev GLM 5.3 took 3,425,

3:43 and it can't switch thinking off at all, which pushes its time up. And the 193.6 came from GPT6 Astra and Fable 5.1, which this benchmark didn't run. Cost is where Jev earns its keep. On banking, it

3:57 runs seven cents per thousand decisions. Even the cheapest LLM GPT5 4 nano cost 19 cents. GPTO OS 120B cost 32 and GLM 5.3 cost $242. So, Jev is roughly three times cheaper than the cheapest LLM and

4:16 about 35 times cheaper than GLM. There's no single tidy multiplier and a few LLM costs are flagged for incomplete usage data. shuffle a model's options and its answer shouldn't move. To test that, the

4:28 benchmark shuffled each item's options three times. Jev flipped on 13% of items and GPT OS 120B on 15. GPT5 4 mini flipped on 37%. So, reordering a list changed its mind more than a third of

4:42 the time. Jev wins this one, although 13 against 15 is a small gap on a small sample. Keep adding options and Jev hits a wall. Step by step, the benchmark raised the count to 512, and every LLM

4:55 handled it. Jev handles up to 255, and at 256, it sends back an HTTP 400 error. Too many choices. Typesafe documents the limit, so it's no surprise. For a list like these 77 banking intents, 255 is

5:10 plenty. For picking one product out of a big catalog, you'd better hope the catalog read the docs. Otherwise, you're writing a short list step before Jev ever sees it. Back to the branch hours

5:20 question. Synthetic items of that kind are part of the benchmark. Questions where no option is good or where the message doesn't settle the answer. On those items, a model counts as admitting

5:30 ignorance if its confidence is 0.5 or lower. Saying the words, "I don't know," isn't required. Letting the number drop is enough. Jev did that 49.7% of the time. Seven of the eight LLMs did it

5:43 between 97.3 and 100% of the time. GPT5 4 mini the eighth got 64.7 a result the RIDMY's oneline summary leaves out even with that one counted Jev comes last a broader measure points the same way

5:58 calibration error asks a simple question when a model says it's 80% sure is it right about 80% of the time Jev's expected calibration error was 0.246 246, the worst of any model measured.

6:11 Roughly, its stated confidence missed its actual hit rate by about 25 percentage points on average. Put its reliability diagram next to GPTOS 120BS and you can see the gap. Take that back

6:22 to the routing code. Your escalation rule is the vendor's rule below 0.5. Ask on the benchmarks no fit items. That rule would fire about half the time at best. Whatever didn't escalate would go

6:35 down the path Jev picked, carrying a confident score that tells your code everything's fine. Jev's wins are real, though. On cost, it's the cheapest model on the board by a wide margin. Shuffle

6:45 the options, and it changes its answer less than anything else tested. Latency stays flat from two options to 255, and it beats even the fast LLM on Cerrus, just not by much. All of this comes from

6:57 a small test. 200 banking items, 25 items per option count step, three shuffles per item, and no significance testing. Every LLM was a mid2026 model with thinking at its minimum. None of it

7:11 tells you how Jeff compares with GPT6, Opus 5.5, or the GPT6 Astra, and Fable 5.1 models behind Typesafe's own numbers. A second study disagrees. Abdel Stark's pilot dated the 17th of

7:25 September ran Jev on a 72 label banking 77 setup and got 87% accuracy with a calibration error of 0.054. By that measure, Jev looks well calibrated. Its sample was 300 examples.

7:39 GLI was the only comparison and its author says plainly it's a pilot and shouldn't be read as a leaderboard. Why the two results differ is still unresolved and the setups aren't the

7:49 same. My verdict then? If you're routing lots of tickets across a fixed list of 255 options or fewer and cost or shuffle stability is what hurts, Jev is worth a trial. I just wouldn't make its

8:01 confidence your only escalation trigger. If accuracy matters most or you need more than 255 options, GPTOS 120B on Cerebras scored highest here. It's about 60 milliseconds slower and cost about

8:15 four and a half times as much. Only your data can settle what's left. One test found Jev bluffing when nothing fits, and another found its confidence well-c calibrated. Pull a batch of real

8:25 messages that match none of your options. Run them through and count how often confidence drops below 0.5. Collect more and you can trust the count more. So, which result holds for yours?

Frontier News · by Hyperjump Technology