Enjoying this issue?
Get tomorrow's AI & engineering digest in your inbox — hand-picked, summarized, and always spam-free.
TLDR
Jev, from Typesafe AI, is a system-one model that returns typed answers with probabilities instead of generating text. It enables building a model router that classifies requests, scores difficulty, and detects PII in a single parallel call, allowing 75% of requests to be handled locally by a 2B model. At 4 cents per million input tokens and no output cost, the entire demo cost less than 1 cent.
Key points
Jev outputs typed answers with probabilities rather than generating text.
Three question types exist: choice (category), score (difficulty), and null (yes/no privacy check).
The router uses FastAPI backend with Next.js frontend, routing to MiniCPM, DeepSeek, or Qwen.
Privacy gating forces local processing when PII is detected; Jev is replaced by local Semif for full local control.
The entire demo cost less than 1 cent; 53 requests processed with 75% answered locally.
Tools mentioned
Techniques
- model routing
- prompt rewriting for image generation
- parallel three-question evaluation
- threshold-based fallback
Stop scrolling. Start reading smarter.
Receive the day's most important AI & engineering updates in one concise email. No spam.
Transcript (captions)
All right. So, in this video, I want to jump in and actually show building something with Jev, some of the actual decisions you need to make about how you're going to use it, some of the
techniques of how you can use it, and some of the benefits and challenges that you're going to see along the way. So, I'm going to start off with a very quick recap, but then I'm going to jump in to
building a model router. So, the whole idea here is that this is something that will be an endpoint that you run locally. It can then interpret what you send to that endpoint via jev or via an
open jev like we'll look at later on. And then it can decide what sort of category is this and what model should it be sent to. Can also decide whether you've got PII, private identifiable
information in there so that we could have it set up to not send to the cloud for that. So that always gets sent to something locally. So, I'm going to try and focus on a few different models in
here, but locally, I'm going to run a lot of this with just the mini CPM 52B. Now, I'm running this on my machine with an endpoint, but you could be running this with LM Studio, you could be
running this with Olama, totally up to you. And then for the cloud models, I'm going to focus on a few models there, but mostly I'm going to be using the Deepseek V4.1 Flash there. to make it a
bit more fun. Also, what I'm serving locally is the latest Quen image 2.1. So, this literally just came out yesterday as I'm recording this. It's a really nice model that can do a whole
bunch of different kinds of things like image generation, image editing, etc. Probably the only thing that sucks about it at the moment is the license currently that it's basically a
non-commercial license. But I'm going to show you how our system will be able to make images, interpret images, do different automated calls to lots of LLMs that are out there. All right, so a
quick recap. If you missed my first video on it, Jev is from Typesafe AI, and they call it a system one model. It doesn't generate text at all. You give it a state, which is basically just
whatever you want to have judged, and then you give it a set of typed questions, and what comes back are typed answers with probabilities as well. The cool thing is that then you can use that
for your code to basically make conditional logic decisions on. So that's basically just all your sort of if else statements, right? If the choice comes back this, then I do this. If the
probability of this is higher than this percent, then I'm going to do this. So we don't have to worry about a model outputting good JSON and basically getting the prompt right to have it
comply with everything that we want. And it really should come back in somewhere around the 70 to 500 millisecond range. And remember, the cool thing for this is it doesn't charge you anything for
output tokens at all, and it only charges you 4 cents per million input tokens. A nice way of thinking about this is that you can kind of think of it as an if statement that understands
language. Now, in Jev and most of the open Jev clones, there are three kinds of questions you can ask it. The first is choice. So, you just basically give it a set of options. it picks one and
you get a probability for each option plus a confidence value. So in many ways that's like your categories, right? You use it to basically categorize things into different classes. The second is
score. This is where you've got some kind of ordered levels from sort of low to high and up to 10 of them and it can then basically give you a position on a scale that you've determined. Again,
that's with confidence. The third type is a null, which is really just like a yes or no question, right? You basically just ask it a question. You give it your input. It will come back saying yes or
no. Can also be thought of as true or false going on here. And yet again, that will also give you a probability that that answer is yes. So that if you take the inverse of that, you've got the
probability of no. All right. So how is all of this going to map out onto a router? It turns out pretty neatly. So you can think of choice being the categorization. What kind of task is
this? Is it chitchat? Is it code? Is it asking for an image? And we could have a whole bunch of different things like that. We've then got it basically where it picks a lane, right? And then Jev is
going to basically pick the lane. So chitchat is going to go into a small model running locally. If it's code though, that's going to go out to the deep model on open router. If it's an
image request, that's going to go to Quen image 2.1, which is also running on my machine. The second one for score is we can actually ask something like how hard is this. So within a lane that
decides whether the little model can handle it. So remember the little model that we're using is mini CPM5 2B right a tiny 2B model. So that will decide if that can handle it or if it needs to go
up a tier to a much better model. And then thirdly we've got the null which is the gate. Does this contain private or client data in any way? If that comes back as high, the whole thing has to
stay local no matter what. And at the start, I'm going to be using jev to actually do all of these. But later on, we can switch to one of the open jevs. So that literally if it is PII, it never
actually leaves our machine even to go to Jev. But at the start, I'm going to say, okay, we trust Jev, we just don't trust the big frontier lab model providers. The other cool thing is that
all three questions are going to go in one request against the same state and Jev evaluates them in parallel. So asking three things and we could actually even add in more is going to
cost you pretty much the same time as asking one. But because you get a confidence back, the router can know when it's unsure if the confidence type is low. So you don't guess, you just
fall back to a safe default. Now, as you see, I'm going to point out right from the start that there is one thing in here. If I'm sending your prompt to type safe to ask whether it's private, I've
kind of already leaked it, right? So later on in the video, I'm going to swap out Jev for one of the open versions running locally so that the whole router is actually just running on my machine.
So let's jump in and have a look at how this actually gets built. Right. So let me give you a little demo of what this can do. You can see here if I basically just fill out a prompt and I send that
off, this is going to basically have Jev decide whether this should go to a local model, whether it should go to Deepseek 4.1, whether it should go to an image model. So, if we come over here and
look, we can see the models that I've got set up here is we've got a mini CPM 52B that's fully local that's running here at the moment. I've then got the Quen image 2.1 which again is fully
local running here. And then I've got the DeepSeek 4.1 flash model. Now that's on open router. So any of the OS are open router. I've also got it set up if we want to use something like a Deepseek
Flash Plus web in there. And if we want to, we can even turn on Opus 5 in here for something that's going to be really hard. Now, if I wanted to, I could run Deepseek 4.1 Flash locally. I am
actually sort of halfway through recording a video all about Deepseek 4.1 and the DeepSeek harness and how all those things go together and how you can do it all locally. I could also run a
local VLM if I wanted to in here, but the idea is at the moment we're using Jev to basically split all these up. I can come in here and actually you can see I can pick the models separately. So
if I wanted to run this just to this model all the time, no problem. You can see super fast there. It's guaranteed to go local. If I wanted to do that and just run it to DeepSeek, I can do that,
right? And then obviously it's going to take longer because it's going to the cloud. It's going to decide if it needs some thinking. It's going to then stream the response back. But really what we
want is auto, right? This is where Jev is going to basically monitor the whole thing and decide what it is that we want. So now you can see, okay, if I do the hey, how's it going? Really, this is
a simple request, right? we would expect that this is going to basically go to the local model and sure enough it went to the local model. Now if we come in here and actually look at the requests,
you can see that the difficulty here here was basically zero that it was really there was no private information in there and the the lane or the category that it decided that this
response should go to was the small model, right? And it did all of that and Jeff basically did all of that in 331 million. All right. Now, if I give it something a little bit harder, you can
see, okay, I'm going to get it to write a Python code to dduplicate stuff. So this is actually still a pretty kind of simple thing, but we've got this now being routed to DeepS as opposed to
doing it locally. And this is basically because everything that I've put in here around code should actually go to the Deep Seek model. So now if I come in and actually look at Jev, I can see for this
last run what actually happened. So the difficulty went up. So this was more difficult than chithat because it's not like off the charts sort of difficult. So if we look here, we can actually see
what is actually getting sent to Jeb. Here we've got the state with the actual prompt in there etc. You can see the questions. So for choice, our choice is going to have the instructions of what
is the latest user message asking the assistant to do. Judge the latest message in the context of the conversation. And then the criteria basically tells us, okay, we've got one
class that's chitchat, one class that's simple question, one class that's rewrite or summarize, one class that's code, one class that's reasoning or analysis, and then one class that's
image for doing image generation. At the same time, we've also got score. How much model capability does a good answer need? And we can see we've got trivial, we've got easy, we've got moderate,
we've got hard. And you can see you kind of got definitions for those. So trivial is a small 2B model handles it easily right greetings oneline facts tiny edits that kind of thing and then it goes
right up to sort of like a frontier right needs the best model available research grade reasoning large systems etc for the privacy stuff we've got a null right so the instruction here is
does this conversation contain personal confidential client financial medical credential data such as API keys passwords anything like that And then we've basically got true that it
contains real secrets and false which it's it's not sensitive data at all. So that's one that we'll sort of come back and look at. We've got another null in there for deciding whether it needs web
stuff or not. If we look at what comes back, you can see from this the choice was code which makes sense cuz we asked for code, right? And we can see that it really has come back like at 100% for
code in there. By the looks of it, it's kind of decided that it's worked out its probabilities there and its confidence. And then we can see for private and for web there. Let's try one that actually
is sort of private. So, okay, here I've put a fake API key in and it's not going to do anything if you want to use it. So, here we can see straight away it detected that this was local and was
private. If we come across and we look at the actual message here, we can see straight away now difficulty is not super difficult, but we've got private. So this is determined that okay, this is
going to have to stay local for this case in here. And you can see the same thing if we put in like a person's name or something or their phone number or something. Again, it straight away picks
up PII, right? So PII is your private identifiable information. Extracts that. It assigns it to the local model. And we can see that if we come in here, we've got the private coming through at like
0.96. So that's what's driving the code at the back end to actually check all this out. So this is the architecture of how all this works, right? We've got our
browser. So what you're seeing here is an Nex.js UI app, right? Just kind of a simple app. It's then sending to a fast API server. So this is a server that's running locally on my machine. So fast
API is just a nice way to build a simple sort of Python API in here. You can see that with one call it will basically do the deciding and then we'll do the routing. It can then stream back to us
from the winner. We've got a fallback in there. We can log everything to SQL light if we want to have everything logged in there and we've got health checks going on. And then you can see we
compatible endpoints. So that's just the standard OpenAI API for calling models on Open Router etc. And we're using that for both Jev, but also for mini CPM, which is running on SG Lang as a local
model here. I'm pretty sure that's the Dlash version I've got going there. And then we've got for the DeepSeek model in the cloud and we've got the Quen image model in here. So the way it works is it
basically asks Jev, right? It gets a decision. It works out what it needs from that, and then it decides a lane. And these lanes are sort of preconfigured. They've got an order of
preference that they go through this. Now, if I turned on the actual deepseeek in the spark, we could actually run that locally. You can see for hard here, if I've got claude opus on, that will route
to opus first up as opposed to just routing to deepsek. That's the key thing that I've got there. And then you can see we've just got our sort of conditional logic in here of that what's
going to actually happen as we go through this. So, if it's greater than 6, we're going to go for a web enabled model. If the prompt is over 16,000 tokens, we can set it to go for a
general model. We can do a whole bunch of different rules in here. We've got the privacy flag. If it's above 0.5, we're going to go for a local model, etc. And you can see at the same time,
I'm logging everything out and storing everything. So, I'm storing the jev decisions in here. I'm storing the the SQLite messages and I'm even storing images in here. So you can see that now
if I ask it to make an image it's worked out that okay it should be using a quen image in here it's going off it's generating on the machine here and we're going to get an image back showing in
here and you can see sure enough there's our sandy colored manoon. Now actually the way this is set up we're actually doing some prompt rewriting in here. So this is obviously one of the things that
all the different sort of image services use is that they'll take your simple prompt which in this case was this and they'll rewrite it. Right now in this case we're using just the mini CPM model
to rewrite it. Could actually probably use a better model to rewrite it and to get a nicer rewriting. You can see okay this is basically run it. You can see that it's added some stuff to it like a
sunlit meadow some other stuff in there to the actual prompt. So that's how it all sort of comes together like that. The other thing too is that I've put in here a tools registry. So you can see
I've actually set it up so that the quen image was actually a tool thing in here. And you can see also that we've got our judge, right, which is what we've got here. We can actually make our judge
both the web version or we can even try doing it with one of the open- source versions. So in a second we'll try out this with SEM if for being the sort of local Jev as opposed to the web one. You
can see here we can also change out the routing thresholds. So remember Jev is delivering back probabilities etc and stuff like that. So that's something that we can play with in here. And I've
got some other little sort of simple things for like putting in a system prompt in there. All right, we've moved now to the open jev one and we can see this here. Right. So let's try again now
with our privacy one and see okay is the SEM if model able to do it. And you see it was so quick. It's fully local. It's worked it out. Let's come in and actually look at our judge. And you see
that the format going in and coming back is slightly different, but it's still doing the exact same kind of thing. We've got this difficulty. We've got the private. We've got the lanes going on
there. And you can see in this case it was 88 milliseconds running through all of this. Let's ask it to do a search on the web for this. So this should be using the open router with the web
version of this. And you can see sure enough it's using duck.go in there. We've basically got this going through. And you can see sure enough we've got our response back going through this. So
the idea here was just to show you a realworld example of using jev. And you could use either the webbased Jev, which obviously is going to be better at generalization most of the time,
especially for harder kinds of things, or you could use one of the open ones. In this case, I was using SEM if in here. You can see that this is basically been gating all the models that we have
here. Now, if I come in here, you can see I can change the thresholds and stuff like that. But we can look at some of the stats just from a little bit since I restarted this last that it's
done 53 requests. It's answered 75% of them locally, right? We can see that yes, sure enough, we've spent money, but this is what we've spent both not just on Jev, but also on deepseek with these
requests. And we can see that this is how much we've saved by having 75% of the requests be local in this case. So depending on whether we've got Jev configured or semif configured, we can
see what the actual response times are going to be coming back here. And we can see how many of the prompts were actually private in here. I've been showing you these as sort of oneshot
prompts, but we can actually do a full sort of chat in here as well. Now, you could configure this that if you started out with something really hard and we were using opus if we were actually
doing sort of a long conversation going here, we'd probably want to use the caching and stay on the clawed opus. So, that would be pretty easy to take this and configure this up. The idea that I'm
trying to show you here is just building a sort of real world use case either using the Jev model itself or using one of the open versions of this. So let me know in the comments if there are other
sort of use cases that you would like to see me do something like this and also sort of what you're actually looking at using this for yourself. My plan is actually to get a second channel up in
the next couple of days which is going to be probably a lot more code intensive. So what I might do is just record a walkthrough of showing the actual code behind these and I'll put
the code up on GitHub and stuff like that if people are interested to play with it and try it out etc. For me the biggest takeaway I would say is whether you're using the Jev model in the cloud
or whether you're using one of the open local versions of this. This really enables you to do a lot of rethinking of how you actually build and structure these kinds of different agentic use
cases. Now, these kinds of things where you want to be able to save on tokens, but you want to make sure that you've got these decisions happening really quickly and can be scaled for a tiny
amount of money here. The funny thing is when I went back and looked at the demos that I did the other day in one of the videos, the entire demos that I did in the video cost less than 1 cent. So,
this is a huge way to basically save money and compute whether you're basically just doing everything fully local or whether you're using the cloud version as well. Anyway, as always, if
you found the video useful, please click like and subscribe and I will talk to you in the next video. Bye for now.