This AI Claims a 200× Speedup. Here’s the Catch.

summarized

TLDR

Jev, a system one model from Typesafe, claims up to 200x speedup on decision tasks by selecting from predefined answers instead of generating open-ended text. The real value is in narrow, checkable judgments where the model returns probabilities your code can act on, but the speedup is conditional and the confidence values don't guarantee correctness. It's worth piloting for a single repeated decision, not for broad automation.

Key points

Jev is a system one model that selects from user-defined answers rather than generating open-ended text.

Typesafe claims Jev can be up to 200 times faster on decision tasks, but this is based on selected workflow evaluations.

Jev returns probabilities for each answer, but a confidence of 0.9 does not mean a 90% chance the decision is correct.

The published speed comparisons use an adapter for structured decisions and assume a compact input, which favors Jev's approach.

Typesafe prices Jev at $42 per million input tokens with output uncharged at launch, but sustainability is unproven.

Tools mentioned

Techniques

  • System one model
  • Structured decision output
  • Probability-based confidence
  • Batch question processing
  • Confidence thresholding
Transcript (captions)

0:00 Types safe says Jev can be up to 200 times faster on decision tasks. The catch is that it can't write an answer. It chooses from answers you define. For software that needs to decide what

0:11 happens next, that trade could be useful. A support request arrives. Someone's card was charged twice. Your application needs to send it to the right team. It doesn't need an essay

0:21 about the emotional significance of being charged twice. Billing can handle that. Jeb is built for this kind of narrow judgment with probabilities your code can use to decide whether to act.

0:31 But a valid answer can still be the wrong answer. Choosing technical support would fit the format and miss the problem. So the practical question is where this model saves useful time and

0:40 where a fast mistake would just create more work. The answer starts with the output. We'll follow that support request through a small example, then check what the published speed

0:49 comparisons actually measure. Typesafe introduced Jev in September 2026. Its founder, Diego Almeida, co-authored the instruct GPT paper back in 2022, part of the research behind instruction

1:01 following chat models. His team is now applying a different training objective to decisions that software can use. The company calls Jev a system one model. That's its name for quick bounded

1:12 judgments inside software. You supply the possible kinds of answer before sending the question, which gives the rest of your application a defined set of results to work with. According to

1:21 the documentation, a request has state, a model name, and questions. State means the information the model can examine, our customer's message, perhaps the account record, and the policy that

1:32 applies. The questions describe what you want it to decide from that information. So, the quality of the input is part of the system. If the duplicate charge exists in your payments database, but

1:42 you never include that record, the model can't inspect it through this request. It can recognize what the customer is asking for without establishing whether the payment really happened. The first

1:52 question uses a type called choice. We define billing technical support and other as possible destinations. Jev's documented response includes one selected option, a probability for each

2:03 option and a separate confidence value. Those option names come from our application. For example, a response might put 92% on billing, five on technical support, and three on other.

2:14 These are made up values for the walkthrough, not a live Jev result. They let us trace exactly what the surrounding code would do without pretending we measured the model. The

2:23 code selects billing because it has the highest probability. For this example, it routes automatically only when that top probability reaches 90%. Otherwise, it sends the request for review. That

2:34 threshold is our teaching choice, and it's applied to the top probability, not to Jeb's separate confidence field. Notice what hasn't happened. No money has moved. A routing decision has put a

2:44 ticket in a queue. Refunding a payment would require its own checks against the actual transaction, the account permissions, and the refund policy. Understanding a request and having

2:54 authority to execute it are different parts of the application. A regular language model can also return structured data. Typesafe's own comparison adapter supports providers

3:04 structured output modes. So, this isn't a comparison between usable software and an assistant that insists on writing a poem. The question is how much work the model does to produce the requested

3:14 decision and its probabilities. In ordinary auto reggressive generation, each output token depends on the tokens already produced. Even a short structured response takes a sequence of

3:24 generation steps that's distinct from reading the input where substantial work can happen in parallel. Jev's launch description says its sampler produces the requested outputs together. Picture

3:34 the response as a form with several fields. A text generator fills the response through a sequence of tokens. Jev's interface gives your code the typed results from one query. That

3:44 picture explains the dependency being removed. It isn't a diagram of Jev's unpublished internal layers or its training hardware. The company calls its training method reinforcement learning

3:54 for calibrated decisions or RLCD. The stated aim is to produce useful decisions with probabilities that reflect how often those decisions are right. The public material explains that

4:05 objective. It doesn't give us enough detail to independently reproduce the trained model. There's a practical consequence that matters more than the acronym. Suppose our support workflow

4:15 needs a destination, a refund request check, and an urgency rating. If we send three requests one after another, the application waits for three separate round trips. Those questions may all be

4:26 answerable from the same original message. The documentation recommends asking independent questions together. Each is evaluated against the same state and then your code combines the answers.

4:36 The same ticket goes into three independent judgments and all three results come back together. The code can then use the destination to decide which of the other answers matter. For simple

4:46 arithmetic, assume each call takes 2/10 of a second. Three serial calls would take 6/10 before other overhead. A single batched call taking the same 2/10 would remove two weights. That's an

4:58 illustration of the scheduling benefit, not a measured Jev latency or a promise that every batch costs the same. But this has a boundary. If the first result tells you which customer record to

5:08 fetch, the next question can't use that record until it arrives. Those steps are dependent. You can combine questions whose evidence is already available. You can't parallelize away a missing piece

5:18 of information. Extra questions also aren't the same thing as free input. Typesafe's choice documentation says additional questions and options still cost tokens even when response time

5:28 barely changes. So batch the questions your workflow might need and keep track of what the service actually builds. Faster scheduling doesn't make irrelevant questions useful. Typesafe

5:37 also gives up open-ended text output. Jev can choose a tool from a defined list, but it won't invent a new function body or compose the customer's reply. If the next step needs an explanation in

5:48 ordinary language, a template, a person, or a text generating model still has a job to do. The other question types help define what that job split looks like. One is called new. It answers a yes or

6:00 no question with the probability of yes. Ask whether the customer explicitly requested a refund and you get a value between zero and one. Your code decides what range is enough to take the next

6:11 step. A value near one means the model leans strongly toward yes. Near zero means it leans strongly toward no. A value near the middle means neither answer dominates. It doesn't mean the

6:22 customer wants half a refund. And this type doesn't return a separate confidence value. The documentation is explicit about that. Score is for an ordered scale. You might define an issue

6:32 as cosmetic, disruptive with a workaround, or completely blocking. Those descriptions matter because rate the issue leaves the standard unclear. A useful scale tells the model what each

6:43 level means before it tries to place the example on it. The documented score is calculated from a distribution across those levels. Suppose the middle level gets 70% and the blocking level gets 30%

6:54 with nothing on the cosmetic level. Number the levels 0 1 and 2. The weighted result is 1.3 between the middle and highest levels. Now return to the confidence field on choice and

7:05 score. According to Typesafe, it summarizes how concentrated the probability distribution is. A strong peak means the model favors one answer. A spread out distribution means more

7:15 ambiguity. It's derived from those probabilities rather than a second observer checking whether the answer is true. That's why a confidence value of 0.9 should not automatically be read as

7:25 a 90% chance that this particular decision is correct. The probability attached to a specific option and the confidence statistic are different quantities. The documentation doesn't

7:36 make them interchangeable and our policy shouldn't either. Calibration is the further question. Do predicted probabilities line up with observed outcomes? Say you collect 100 similar

7:46 decisions assigned an 80% chance. A well-calibrated group would be right about 80 times. You learn that by checking the outcomes across the group, not by admiring the precision of one

7:56 decimal. Our second mock ticket says the export crashed and the customer wants their money back. Billing gets 48%. Technical support gets 47 and other gets five. The same routing rule now sends it

8:08 to review. No format error occurred. The problem is that one destination may not capture the whole request. There are two useful responses to that ambiguity. You could keep a human review path or ask

8:19 separate questions about a technical failure and a refund request, allowing both to be true. That changes the design to fit the task. Raising the confidence threshold alone won't repair a question

8:29 that forces overlapping needs into one category. Our third mock ticket asks to restore a deleted account. Another option lets the system admit that its normal destinations don't fit. Without

8:39 that option, the available labels still force a selection. A beautifully formatted answer can be your first clue that you designed the wrong menu. This is the limit of the zero hallucinations

8:49 claim. Typesafe says it's zero in that chart comes from guaranteed schema matching, not an empirical count of factual mistakes. A value can belong to the allowed set and still describe the

9:00 wrong situation. The closed menu prevents an invented label. It doesn't prove the selected label is correct. With that distinction established, the speed evidence is easier to read. Types

9:10 safe reports end to end calls in roughly 70 to 500 milliseconds. It's published near 200 times comparison comes from selected workflow evaluations. That number is a result under a particular

9:21 comparison, not the acceleration factor for every job a developer could give it. The evaluation site shows four workflows. security incidents, agent traces, invoice processing, and customer

9:32 service. Each separates judgments from rules implemented in code that matches the product's intended use. It also means a coding benchmark or a long form writing task would be answering a

9:42 different question about capability. Look closely at how the reference answers are created. The site says it uses the average responses of two larger models at high reasoning effort. Those

9:52 are model generated reference labels. Agreement with that reference can help compare systems, but it doesn't establish that every final action agrees with a human reviewed business policy.

10:02 The published comparisons also use an adapter that asks the language models for compatible structured decisions. Types safe notes that asking for probabilities tends to be slower and

10:11 more expensive than asking for discrete answers alone. If your application only needs one label, that simpler baseline deserves to be measured, too. There's a second distinction inside the launch

10:20 material. The short side-by-side demo uses a compact input, which the company says favors its approach. The workflow tests involve more complex calls. Don't combine the best latency from one demo

10:32 with the strongest multiplier from another and present them as one experiment. On the evaluation page captured for this episode, Jev's aggregate agreement is 67.8%.

10:41 In customer service, it's 76% while invoice processing is 61.8. Those are vendor results against the reference labels, not our accuracy measurements. The variation is the reason to inspect

10:53 the task you actually need. There is some evidence outside the vendor's own evaluation. Mike Taylor at every published an early test of Jev as a writing checker. He reported hundreds of

11:03 judgments in under a second while also saying he wanted a more thorough accuracy check before production that supports a promising use case with a clear limit on the conclusion. Every

11:13 also reports a small comparison using 12 synthetic passages. Jev took a median of.35 seconds per passage against 8.83 for fable 5.1 at high effort. That's roughly 25 times faster in that test. It

11:27 is separate evidence, not an independent reproduction of type safes 200 times claim. The quality result is worth keeping beside the stopwatch. Jev caught six of seven intended defects. The

11:38 comparison model caught all seven. one missed defect persisted across repeated runs. A cheap checker can still be useful, but the example shows exactly why it returns a probability can't

11:48 substitute for measuring what it misses. A guardrail works in much the same way. Suppose another model drafts a support reply. A checker can compare that draft with the relevant policy and ask whether

11:59 it promises a refund the policy doesn't allow. The proposed reply and the source policy need to be present, otherwise the checker is missing part of the evidence. If the checker flags a problem, your

12:10 application can stop the reply for review or ask for a revision. If it passes the reply, that's a model judgment that still needs evaluation. Test examples where the policy is

12:20 genuinely violated as well as harmless messages that use similar words so you can see both missed problems and unnecessary blocks. But for actions with clear mechanical rules, use those rules

12:31 directly. A payment amount above the permitted limit doesn't require a language model's opinion. Neither does a user lacking access to an account. Let code enforce facts the system already

12:41 knows and reserve the model call for the part that needs language interpretation. This division also makes failures easier to locate. If the ticket went to billing because of a model answer, save that

12:51 answer with the question and the relevant input. If code rejected a refund because the transaction had already been reversed, save that rule's result. One final success or failure

13:00 flag would hide two very different causes. Extraction has a similar boundary. If you need an email address from a document, a close set model can choose among candidate addresses your

13:10 program already found. Type safe documents that pattern. Extract candidate spans ask which one fits the question and return the original value. The model selects from evidence rather

13:21 than generating a new address that can preserve the exact spelling, but it can't rescue a candidate that your first pass never found. If the right address was missed, choosing the best remaining

13:31 one is still wrong. include a no match route and test the extraction stage as well as the model selection. The whole chain has to retain the answer. The same idea helps with tools. Define the

13:42 operations and the argument values that your application can accept. Then let the model choose within those boundaries. Jev's current documentation allows up to 255 options in a choice.

13:53 Larger cataloges need some selection structure around that limit. For a large tool catalog, you could retrieve a short list first or choose a category and then choose within it. That saves a giant

14:03 flat menu, but introduces another place to make a mistake. If the first selection excludes the right tool, the second selection can't bring it back. Measure the complete route to the

14:12 correct action, and a fast service call still sits inside a larger application. Reading a database, finding the policy, waiting for network traffic, and performing the chosen action all take

14:23 time. Reducing the model slice can make the experience noticeably better while leaving the slowest remaining step exactly where it was. Here's a simple example. Suppose the old model takes 4

14:34 seconds and everything else takes 1 second. Replace the model with a call that takes a tenth of a second. The model step is 40 times faster, but the whole request drops from 5 seconds to

14:44 1.1, about 4 and a half times faster. The price deserves the same treatment. Typesafe publishes an input price of $42 per billion tokens, equivalent to 4.2 cents per million with output uncharged

14:57 at launch. Any savings multiplier depends on the competing model and the work each one performs. The input rate is what you can use to calculate your own bill. Suppose your total build input

15:07 averages 2,000 tokens per request and you make a million requests. That's 2 billion input tokens. At the published rate, the input bill is $84. This is arithmetic using an assumed request

15:19 size, not a measurement of what a particular support application would consume. Types safe itself says the long-term sustainability of its price will take time to demonstrate. So the

15:29 current rate is a starting point for a pilot. If the application only works economically at one launch day price, that dependency should be visible in your calculation before you commit the

15:38 design around it. There's another bill that can dwarf the model bill. Review. Imagine a 100,000 tickets and a 1% review rate. That still leaves a,000 tickets for people to handle. Those

15:49 figures are hypothetical, but the multiplication is real. Small percentages matter when they sit on top of a large queue. Thresholds decide how much work lands in that queue. Increase

15:59 the required certainty, and you may reduce automatic mistakes while also sending more cases to review. The useful graph shows error rate alongside the share of requests you automate. A system

16:10 that gets everything right by reviewing everything hasn't solved the same problem. The consequences also differ by action. Routing a ticket incorrectly and authorizing the wrong refund don't have

16:20 the same cost. In a simplified example, suppose reviewing a case cost $1 and a wrong action cost 100. If the action is right 95% of the time, its expected error cost is $5. So under those

16:34 assumptions, review is cheaper. The break even probability would be 99% before adding other costs or benefits. That isn't a recommended production threshold. It shows how the consequences

16:44 set the requirement and why a plausible looking confidence value can't choose your business policy for you. Typesafe has funding to pursue the idea. Its investor DCVC announced a $40 million

16:56 seed round. Jev is offered through early access and public documentation is available. Those facts support that this is a product you can investigate. They don't independently establish the speed

17:06 or reliability of your eventual integration. For a first pilot, choose one repeated judgment with a result you can check. Routing historical support tickets is a better defined starting

17:16 point than asking a new model to run the entire support department. Preserve the original messages and the expected destinations, including ambiguous cases that a person would reasonably review.

17:26 Keep separate examples for designing the questions and evaluating them. If every troublesome ticket becomes an example in the prompt, testing on those same tickets no longer tells you how the

17:35 system handles new work. Include both ordinary requests and the awkward boundaries, mixed intents, missing context, and inputs that belong in none of your categories. Compare the same

17:45 task against what you already use. That might be a few rules, an existing classifier, or a language model returning one structured label. Match the requested output and measure the

17:55 quality your application requires. Asking only one competitor for a detailed probability distribution would change the work being compared. Run the proposed router in shadow mode first. It

18:06 records its decision while the existing process still handles the request. Measure latency from your deployment region, including slow responses and failures. Group the mistakes by cause.

18:16 Then choose the automation threshold from the errors you can tolerate and the review capacity you actually have. I'd start Jev with a narrow decision that can be checked and keep the surrounding

18:25 code in charge of execution. The evidence makes that pilot worth considering. It doesn't support replacing a general coding assistant or trusting every return label. The catch

18:35 in the title is the boundary that makes the product interesting. Fewer possible outputs can make a useful decision much cheaper. If you're deciding how to arrange those decisions into a larger

18:45 workflow, our loop versus graph engineering video follows that next question. For Jev itself, the test is concrete. Does this one decision become faster and cheaper at the error rate

18:55 your application can accept? Measure that and you'll know whether the headline matters to your software.

Frontier News · by Hyperjump Technology