Jev is the FIRST of a Whole New Class of AI Models (Here's How to Actually Use It)

summarized

TLDR

Jev introduces a new class of AI model called system one models, which are master decision makers that output structured multiple-choice answers with confidence scores, rather than free-form text. It is dozens to hundreds of times cheaper than large language models, enabling fast, parallel decision-making in workflows like PR triage and game testing. The presenter argues Jev is best used alongside LLMs, with Jev handling classification and routing decisions while LLMs handle deeper reasoning.

Key points

Jev is a system one model, not a large language model, and outputs structured multiple-choice decisions with confidence scores.

Jev is dozens to hundreds of times more cost-effective than top LLMs like GPT-6 Astra and Claude Fable 5.1.

Jev makes routing decisions in about 0.2 seconds per decision, enabling real-time classification in workflows.

The presenter used Jev in a PR triage workflow with Archon, combining Jev for classification and routing with LLMs for deeper review.

Jev can be accessed via OpenRouter or typesafe.ai, and an MCP server is available for easy integration.

Critics argue Jev is not truly innovative, just repackaging existing classification model ideas.

Tools mentioned

Techniques

  • System one models
  • Classification and routing
  • Structured output
  • Confidence scoring
  • Parallel decision-making
Transcript (captions)

0:00 So we have a new AI model that was just released within the last week called Jev. But Jev is special. Jev is not just another large language model. It introduces an entirely new class of AI

0:11 models called system one models which are master decision makers. Now that might not sound super exciting at face value, but having a master decision maker is actually incredibly useful,

0:23 especially because of how fast and cheap is. We'll of course talk about that as well. And I'm genuinely excited for this. Like I haven't been excited for an AI model for a while now. I got to be

0:33 honest, I've been kind of burnt out with all the new LLMs coming out because every single time we have a new model, there's always the rush to incorporate it in our workflows. There's all the

0:42 hype on the internet for look what this does for your second brain. Look what kind of beautiful websites you can make. Look at these beautiful scenes that I'm generating in Blender. And you see that

0:50 time and time again. It's just like, man, I want something new. And we have something new and genuinely innovative here with Jev. Now, Jev has been out for a little bit, so you might have already

1:00 seen a video or two on it, but I've specifically waited to make content on Jev cuz I wanted to actually build with it as I try it out initially before I go and just talk about it. And so, of

1:10 course, we'll start with a highle overview of how Jev works, its limitations, and the really incredible benefits. But then, I want to get into how I've been building with Jev and show

1:19 you some really cool things I've been using with it already. And then there's one other thing I want to hit on, which is there's actually a good amount of criticism for Jev as well. The idea that

1:29 it's not really truly innovative and it's more just taking ideas that already exist in the industry for classification models. It's kind of true. It's something interesting to talk about and

1:39 so we'll hit on that as well. So the most important thing to understand with Jev is it is not a large language model. You can't have a conversation with it like you can with chat, GBT, Claude,

1:50 Gemini because it doesn't generate text. Its sole responsibility is to take a bunch of input around a situation and generate a decision. And so with this, there's a brand new training algorithm

2:03 that they called RLCD, reinforcement learning for calibrated decisions. Is this different from RLHF or reinforcement learning with human feedback which is the algorithm used to

2:14 train generative AI models? And so the promise with Jev and system one models is that we are solving a lot of reliability and hallucination issues we see with generative AI specifically for

2:25 making decisions. And I love what they say here. Our claim may sound too good to be true, but the bitterest lesson in AI is that optimizing for the right task gets you an unfair advantage. And I

2:36 agree with this 100%. And I love how they sign off here. May your intelligence be ever reliable. And if you follow my channel, you know that that's a really important thing for me

2:45 that I'm always chasing. Like when I'm building archon and my AI coding workflows, I'm all about adding as much determinism in these workflows as we possibly can. And Jev is another tool in

2:55 our tool belt to do that. Now, before I go and explain more with Jev, I want to show you what it looks like in practice. So you can see what the inputs and outputs look like. It's quite different

3:04 from a large language model because it's never going to be free form text. It's always going to be a decision. In fact, it looks a lot like structured output that we have with LLMs, though there are

3:13 a lot of differences and we'll talk about that as well. So, the best way to describe the input to Jev is it's really a two-part input. You have the situation that it is analyzing and then you have a

3:24 set of questions that each have a multiplechoice answer. It is never generating free form text. Remember, it is always making a decision based on the situation. And so for this basic example

3:34 right here, you can imagine Jev is integrated into some kind of customer support agent. And so this person, they're having trouble connecting their Stripe account. They're losing sales and

3:43 they need help ASAP. And so we're asking Jev a series of questions that it's going to answer in parallel. And it's always going to pick from the criteria that we give it, multiple choice. And so

3:53 the output quite simply here is going to be its answer, its choice for each one of the questions as well as its confidence score. And so here it made a routing decision. It's 64% confident

4:03 that it should go to the billing department this complaint. And for the frustration, a score of one, that means that the customer is deemed frustrated but civil. So doing sentiment analysis

4:13 as well. Okay, so that's cool, Cole. But large language models can already do this. They can already output structured JSON to make decisions in a workflow. And yes, that is true. But there are

4:22 three massive benefits to Jev. It's the perfect trifecta. Jev is faster, more cost-effective, and more reliable when it comes to making decisions compared to large language models. At least what

4:34 they claim and what I've seen, it has a 0% failure rate for structured output. So, it never has any kind of malformed JSON that would make a future step in your workflow fail. And Jeev is

4:44 incredibly cost effective. This graph right here is logarithmic. And so, Jev is dozens, even hundreds of times more cost effective than all of the best large language models. There are a

4:55 couple of models that do actually make better decisions if you give them the time than Jev, but it is incredibly cheap. Like look at the cost per million tokens of Jev compared to something like

5:06 GBT6 Astra or Claude Fable 5.1. And so if you combine the incredible cost effectiveness with the speed that we have with Jev, it's able to make decisions 20 to 200 times faster. It's

5:19 40 to a,000 times cheaper. That together means that you can build these systems having a lot of intelligence, making snap decisions extremely fast, and even making hundreds or thousands of

5:30 decisions in parallel. So, there are a million different use cases for Jev that are super powerful, very practical. But the coolest one that I've been working with right now I want to show you just

5:40 really quickly is using Jev to play video games exactly as a user would, making decisions as quickly as us or even faster. And so I've been experimenting a lot with using Jeb

5:52 within my AI software factory to test out games as I'm building them as the large language model is actually writing the code. And so you can see in real time it's making all these decisions

6:02 with confidence scores for all the different actions that I'm proposing for it. And it seems like I'm playing. It's attacking, moving to enemies, dodging, but I have hands off the keyboard. This

6:10 is so cool to watch. And I've tried to get large language models to do this kind of thing, but it just doesn't work because they can't process things fast enough and it would be way too

6:20 expensive. So, this, by the way, is actually my game. This is running on local host right now. This isn't just some like Twitter demo that I have up, though. I'm going to show some of those

6:27 as well. The important thing here is that Jev can't create this game, but it can definitely interact with it in a way that a large language model never could. So, it's not like you're going to

6:37 totally swap all your large language models for Jev right now. It doesn't work that way. You're just going to put Jev in your automations where you have those decision points or any kind of

6:46 classification step. And so the best workflows for AI coding or any kind of business use case going forward is going to be a combination of Jev for the decision-making and LLMs for the other

6:57 reasoning. The sponsor of today's video is Firecrawl. Every AI agent that I build eventually needs access to the web. And there are a lot of agents out there that have these capabilities out

7:08 of the box like Cloud Code. But if I'm building my own AI agent with Pyantic AI or langraph or PI, I have none of that. And even if you are using clawed code, the search capabilities built right in

7:19 are very inefficient and tokenheavy if you haven't noticed before. And Firecrawl has the solution for this. They call it the context API for AI agents. That's exactly how I use it. And

7:31 they have an MCP server that makes it extremely easy to bring their context API into any AI coding assistant or other AI agent. So, for example, with cloud code here, I just copy this

7:41 command, go into a new terminal, paste it in just a single line to get the MCP server added. So, now when I go into cloud for the first time, I simply have to do /mcp to then set up the

7:51 authentication and then I'm good to go. So, I'm showing you the full flow here in just like 20 seconds. Authorize and then back over to the terminal. Authentication is successful and I can

8:00 now start sending in my requests. And the search capabilities of Firecol are powerful. It's not just a Google search. It's able to directly generate queries that answer my question instead of just

8:10 performing a really broad web search like you'd usually see in something like cloud code. So right here I asked what are people running into when upgrading to Pantic AI version 2 a specific but

8:20 powerful example because it has to look through a lot of context to answer this but it's able to do so in only three calls to the firecrawl mcp server. So firecrawl gives me exactly the context I

8:31 need. It can also give me the full page as clean markdown. The MCP server is easy to use anywhere and they also have an SDK if we want to build firecrawl directly into our custom agent tools.

8:40 So, it's super easy to use whether you're building your own agent or using something out of the box like a coding agent. And firecrawl is free to get started with a,000 credits a month and

8:49 you don't even need an API key to use their MCP server. I'll have a link to them in the description. Now, of course, I've been doing a lot of testing with this myself, building larger AI coding

8:58 workflows with my open source tool, Archon, combining Jev with LLMs, using the right model for the right step. And so, with this Jev PR triage workflow, essentially what it does is we have

9:11 classification and routing at the start that figures out what kind of review we need to perform on a pull request and then go and do that review. And of course, for the first two steps here,

9:20 classification and routing, I'm going to be using Jeb because we're just making decisions here. And so, it's very coste effective this workflow because, of course, Jev itself is cost effective,

9:31 but then also we get to decide what kind of review we're doing because we don't always need a super deep AI review on every single poll request. And so, this is just one simple example with a lot of

9:43 testing that I've been doing with Archon. So, also let me know in the comments if you want me to make more content on this. I'm definitely going to continue to explore using Jev within AI

9:51 coding workflows for testing things like I showed you with my game for classification steps with things like issue triaging and pull request review. The possibilities are endless. And if

10:02 you've been following my channel in Archon and you're curious, this is the exact workflow that you just saw in the Archon UI. So I just simply call this Python script that makes the

10:11 classification with Jev. And so for my Jev usage, I'm going directly through Open Router. So, Open Router was super fast to make the Jev model available along with all the other LLMs that you

10:22 can use there. And then also, if you want, you can go directly through typesafe.ai. Typesafe is the company that created Jev. And so, of course, I'll link to

10:30 this in the description. By the way, they're not sponsoring this video at all. I am genuinely excited for this model. I hope that you are too. Just going through some of the use cases with

10:38 me here. And if you're not sold yet, let me show you some more use cases. So, this is some experimentation I've been doing myself. I've seen a lot of other people on the internet do this as well,

10:47 using Jev as an LLM router. It's a really common use case where you have a bunch of different LLMs that you want to pass the right requests to, right? Like sometimes for the sake of cost, the

10:58 simpler requests go to the faster model, ones that require deeper reasoning, we want to send to the strong model. Traditionally, you've used yet another LLM to make the routing decision. But

11:08 again, with Jev, even with tiny LLMs, it is going to be faster and cheaper and of course more reliable. So the situation we give as input to Jev is the query that we want to route. And then the

11:20 multiple choice that it has to answer is which one of these models should we route it to the strong, the coding, the open or the fast. And this is just a quick visualization I put together to

11:29 show you all the testing that I've been doing. But for a deeper question, it routes to the strong model with a confidence of 100%. Uh this, you know, convert this bash to PowerShell. A

11:38 little bit of a coding task. It has a 98% confidence going to the coding model. And some faster ones here like convert 72 Fahrenheit to Celsius. Yeah, this can definitely be handled by a

11:48 cheaper model like GPT 5.6 Luna is the faster one here on open router compared to I don't know like what's the strong one here? Yeah, Claude Sonnet 5 for example. Now these numbers aren't the

11:58 best. This is just a really small subset of all the testing that I've been doing. But the really cool thing to show you here is that out of the dozens of the tests that I have visualized here, it

12:07 costed me 4/10en of a penny to do all of this routing with Jev. And the average time it took was 2/10 of a second to make each one of these routing decisions. Super cool. Okay, so that's

12:18 enough of my testing. I hope you liked it, but let me show you really quickly what other people have been sharing on the internet as well. So this person on X posted using Jev to play the classic

12:27 game Doom. And it actually looks a lot like my own testing with my own game where we have the decisions that are being displayed on one side right here in real time as Jev is playing the game.

12:37 And then of course the game itself. It's so cool how it's able to play something like this. And yeah, it's a basic game because you can't just give it like millions of decisions, but this is still

12:47 incredibly impressive. And of course, I'll link to all of these resources in the description. Another really powerful use case for Jev that you've probably thought of at this point is using Jev

12:57 for browser automation. So more traditional tools like Playrights and Versell's Asian browser CLI, it's always driven by an LLM. And I use these every single day as I'm building web apps,

13:08 full stack apps, but it's always slow, right? Like the slowest part of my AI coding workflow for any kind of full stack app is always when it has to validate things visually and navigate

13:18 the browser. But now we don't need an LM to do it. We can use Jev because every single situation is the current layout of the site and it just has to decide with multiple choice the next action to

13:30 take like click this button or type in this input. And so this is just one really cool open source repo that I've seen. There's probably going to be a lot of Jev browser use tools released in the

13:39 next couple of weeks, but this is one of them. It works incredibly well. And then I also found this really cool visualization of Jev playing Pong. And so that's the top row right here. the

13:49 game is slowed down basically to the rate that the model can handle and Jev can pretty much handle it at a human rate and then other LLMs down here like 3.8 Flashcloud haiku 4.5. You can see

14:00 the game has to be slowed down a lot for it to actually process where the paddle needs to be as the ball is coming. And then one last resource I want to show really quick is this open source repo

14:11 that curates a list of projects and just general use cases for Jev. So I'll scroll down in the read me too current coverage. We got classification and routing. Of course that's going to be

14:20 the most common one. I mean that's literally what Jeb is made for. but then using that for agent decisions, verifications and guard rails, calibration and research, games and

14:28 simulation, finance and trading. There are so many cool use cases to poke around here. So yeah, I'll link to it in the description. Just check this out. You just get your imagination going here

14:37 as you go through these different use cases like I'm trying to do for you in this video. Because when you you think of Jev as this glorified classification model, it's not very exciting at first.

14:48 But once you realize what you can do with it, man, the world becomes your oyster here. And speaking of Jev being a glorified classification model, that's the last thing I want to hit on really

14:58 quickly cuz it's the biggest criticism that I've seen for Jev. And I've actually seen it quite a bit. A lot of people say that we've had the idea of a classification model in the AI industry

15:07 for decades and so we're just reinventing the wheel here. And that's it's true to an extent, especially when we use it for very basic things like this example I have in the type safe

15:18 playground. But what really makes Jev powerful is how general it is. Like I think the best way to describe it is it feels like there's still the intelligence of a large language model

15:28 operating behind the scenes producing the structured output answering our different questions. Like I just had this really silly one right here. Which country has the coolest buildings? I

15:37 gave it a few options and it's actually says Germany with 63% confidence. A very opinionated thing. I mean, who knows what it's going off of, but if I run this over and over and over again, the

15:46 numbers change a little bit, but it actually always says Germany, which is interesting. Very cool. So, anyway, anyway, I've built a lot of classification models in the past, but

15:56 it's always for a very specific task and you have a very specific data set. So, I've used, you know, like TensorFlow and PyTorch to build classification where you give it a chest position and it says

16:06 who's winning or it looks at an animal and identifies which animal it is. But then if you have it do a different task, it can't do it at all because it's so specific. But with Jev, it's

16:16 classification in the general sense. You can give it any kind of situation for customer support or opinions on countries if you really want and it's able to give you a response here. Now,

16:27 this is kind of a silly example, but for anything more objective, like what kind of pull request review level does this need or what model should we route to here? Or what's the next best action in

16:38 this video game? Like, Jev can just handle any of that. And it's so accurate. So, I hope that you found this interesting. All the super cool use cases for Jev. I would encourage you to

16:47 try this right now either through Open Router. They also have a wait list that I was able to get into within a day. I'll link to that in the description as well. And I will certainly be doing a

16:56 lot more content especially with Archon and how I'm using it in my AI coding workflows. So stay tuned for that. So with that, if you appreciated this video, you're looking forward to more

17:05 things on AI coding and Jev and system one models, I would really appreciate a like and a subscribe. And with that, I will see you in the next

Frontier News · by Hyperjump Technology