Open Jev Models Are Here!!

summarized

TLDR

Open implementations of Jev, a universal classifier that reads logits from the prefill stage, are proliferating rapidly with at least 20-30 projects. The closest open version, SEMif, uses a frozen Qwen 3.5 4B model and scores 74.3 on Jevbench versus Jev's 75.3, but the gap widens on hard tasks where no open model matches Jev's accuracy. The real takeaway is that the logit-readout trick is simple enough to replicate, but the generalization from a well-funded lab's training still matters for complex classification.

Key points

Jev is a universal classifier that reads logits from the prefill stage rather than generating tokens.

At least 20-30 open projects have emerged within 24 hours attempting to replicate Jev.

SEMif, a frozen Qwen 3.5 4B model, scores 74.3 on Jevbench compared to Jev's 75.3 overall.

Bespoke Nimble uses a LoRA on Qwen 3.5 9B trained with only 3,000 contrastive examples.

Jev's terms of service prohibit running benchmarks against other models.

Open versions perform well on easy tasks but struggle with multi-hop reasoning and date arithmetic.

Tools mentioned

Techniques

  • logit readout from prefill stage
  • contrastive data curation
  • entailment-based classification
  • diffusion-based classification
  • hidden state projection onto options
Transcript (captions)

0:00 Okay, in my last video I talked about Jev and that project really has taken off massively over the past few days and like many things I guess in AI you've got the extremes of people who

0:12 absolutely love it and think it's the best thing ever and on the other extent people saying that this is nothing special nothing new but really in this video I want to talk about what if you

0:22 could do all of this locally if you could run the weights locally if the weights were open models etc. So, in this video, I'm going to look at five, possibly six versions of this idea that

0:32 are open projects. And I got to say, the amount of open projects around this kind of thing has just exploded over the past day or so. Yesterday, when I was looking at this, there were like a couple of

0:43 these things. This morning, going through it, there's well over 20, maybe even 30 different projects out there of people trying to replicate this model. And some of them are actually getting

0:53 pretty close results, too, Jev. And one of the things that I noticed researching this, and shout out to Susan Jang for pointing this out, that part of the terms of service of actually using the

1:03 Jev model are that you won't run any benchmarks on it to compare it to other models, which is kind of insane in itself. But anyway, let's talk about what actually makes Jev special. So

1:14 really, Jev is a universal classifier with perhaps a little bit of regression thrown in there. And of course, classification is not new. It's been around a long time. We've had things

1:24 like the BERT models, more recently the modern BERT models, and even things like text CNN's can be really good and really fast for text classification. But generally, the challenge with most of

1:35 those is that each one needs to be trained or fine-tuned on your specific use case. So, you change your labels, you're basically having to train again. And that's something that I've done

1:45 plenty of times myself over the years when I was training simple NLP models to be able to do different kinds of text classification for my previous startup where we were constantly looking at

1:56 reviews and trying to predict sentiment trying to predict what people were interested in etc. Now there actually have been a couple of different sort of attempts at creating a universal

2:06 classifier with no training. Even going back to sort of 2019 there were papers that reframed classification as entailment. And it turned out that entailment tasks were something that the

2:16 BERT models could actually do quite well. So for example, there are different kinds of entailment, but I'll give you a really simple example is where if you've got two sentences, is

2:25 the second sentence random and not related to the first sentence or is it related to the first sentence? If I have one sentence about pizza and then the next sentence is the boy loves pizza, we

2:37 would say that that entails. And you can actually sort of think about that those kinds of things could be trained for a more generalized sort of yes no kind of thing which is one of the tasks that Jeb

2:46 is actually very good at. So the question was well why didn't that eventually become Jev 7 years ago or something like that. Well really the models back then just didn't really have

2:55 enough world knowledge to be able to generalize to all the different things. Clearly, whatever the model that Jev is using, it has enough sort of world knowledge in its weights that it is able

3:06 to generalize to lots of different ideas. And I think where Jev is really kind of interesting is that it's caused a lot of people to start thinking about what can we do just with really fast

3:17 classification and what kind of tasks can be offloaded away from a reasoning LM to things that can be just done with simple choices or simple sort of yes no answers. Now, in my last video, I said

3:30 that we'd probably see open versions of this. And sure enough, like within 24 hours, we saw a whole bunch of these come out. And as I'm recording this, like I said earlier, there's probably at

3:40 least 20 plus, maybe even 30 plus different open implementations of this idea. So, how close do these open implementations actually get? Well, someone's already built a benchmark for

3:52 this called Jevbench, and it scores these systems on intelligence, calibration, speed, and cost. And you can see looking at this on the overall score, Jev is at 75.3. And right behind

4:05 it is a project called SEM. If so, I think this was originally called OpenJV. Now, this project is basically a frozen Quen 3.5 4B model with no training at all. Then we've got a diffusion model

4:18 called DJ at 74.3 followed by Lelaya and then followed by a whole bunch of different models scoring in the 60s. And this leaderboard that we're looking here is actually a

4:27 waiting that they're using for Jevbench. But you could actually emphasize things like accuracy where DJ actually goes up. And the big thing here is that all of these you've got some form of trade-off

4:37 between speed, cost, and accuracy. If we look at things like GPT 5.6 six Luna. Here we can see that its intelligence level is much higher than the other models. So in that sense, the system 2

4:50 model still wins for different hard tasks, but it's costing roughly six times as much as Jev and is much much slower here. Let's look at a few of these projects, see what differentiates

5:00 them, and see how you could actually run these yourself. Okay, so what I've done here is just build a demo with six of these. So, all six of these are running on an Nvidia RTX Pro 6000 Blackwell. And

5:14 a big thanks to Dell for sponsoring the compute here. This is running on a Dell T2 Pro Max workstation here. So, basically what I've done is for all of these is just pull the repos, set them

5:25 up, put them on the GPU, and we can actually test them out. Okay, so the first one here is SEM if which was originally called OpenJV. So, the idea here is pretty simple. You take a normal

5:36 open model, you put the question and the options in the prompt and then rather than letting it generate an answer, you look at the logits for the option tokens that come out and you basically softmax

5:46 over those. And if you remember in the last video, my guess was that Jev was basically using the prefill stage and then doing prediction off that. And it seems that many people converged on that

5:57 same idea. So basically, you process the prompt once, read off the scores, and answer the tokens. So you can see here we're going to put this in. If we run both, we can see okay the generated JSON

6:09 that we get back and then the actual just scores that we get back here. So if we want to we can just generate the scores out and then see okay the probability distribution that's coming

6:19 off the softmax there. And this allows you to basically do a lot of different sort of things where you've got the yes no questions where you've got some kind of policy etc. And you can see that okay

6:31 if I change that to basically only modify some things now you see okay it's definitely getting confused right now if we totally change it so we're asking it to actually make a change you can see

6:43 now it really understands that okay it's required there so this is the typical kind of use that we would use Jev for and that actually brings us to bespoke nimble it's the same sort of readout

6:54 trick but they've basically put a lura on a quen 3.59B model and built the whole thing with under 3,000 training examples. And just to make it clear, they've pointed out

7:04 that they didn't distill from Jev. The really interesting thing here is actually how they made the data. They call it contrastive data curation. You write two examples that are almost

7:14 identical and you change one fact so that the right answer flips. So the example is a refund rule. Only mirror can authorize refunds on this account. In one version, the refund was signed by

7:25 Mirror, so it's authorized. In the other one, it was signed by Noah. Everything else is exactly the same. And what that teaches the model is which piece of evidence should actually change the

7:35 decision. Now, on their own hold out set, this is getting around 90% where their base model gets around 66 and Jev is getting around 93, but actually on the JeffBench hard tier, it was only

7:46 getting around 44. So, again, this is strong on what it's been trained on, which totally makes sense. But the other thing I would say is the fact that they've only sort of done some training

7:55 on 3,000 examples. Probably we shouldn't expect it to generalize that much. And just to make it clear, they've pointed out that they didn't distill from Jev. The really interesting thing here is

8:05 actually how they made the data. They call it contrastive data curation. You write two examples that are almost identical and you change one fact so that the right answer flips. So their

8:17 example is a refund rule. Only Mirror can authorize refunds on this account. In one version, the refund was signed by Mirror, so it's authorized. In the other one, it was signed by Noah. Everything

8:27 else is exactly the same. And what that teaches the model is which piece of evidence should actually change the decision. Now, on their own hold out set, this is getting around 90% where

8:38 their base model gets around 66 and Jev is getting around 93. But actually on the Jeffbench hard tier, it was only getting around 44. So again, this is strong on what it's been trained on,

8:48 which totally makes sense. But the other thing I would say is the fact that they've only sort of done some training on 3,000 examples. Probably we shouldn't expect it to generalize that much. All

8:58 right. So the demo here really shows how they sort of flipped the facts for making the training data in here. So you can see here that okay, we've got this signed by mirror in here. If we run the

9:10 score, is it authorized? Yes, because remember they only should be authorized if they're signed by mirror. Okay, we've got mirror is the one to authorize account 42. This was signed by mirror.

9:21 So, it should be true, right? It should process the refund. And we can see that okay, we've got our nice JSON response back there. If we flip the fact here now, you can see we've actually changed

9:32 it to was signed by Noah, right? And you can see it comes back false there. If I change it to Sam, it should still be false, right? Cuz I've basically done that. And remember this is done with a

9:43 Laura adapter on the 9B model. We can see that actually Quen 3.5 seems to be doing quite well with this question at least both with the Laura on and with the Laura off. Okay, if we look at this

9:55 second one here, we can see Priya bought a conference ticket. She needs basically pre-approval from the manager. Priya's manager sent a written approval for the ticket on March 1. Actually on this one

10:07 it should be true but we can see in this case that the adapter is actually getting it wrong. The normal Quen model gets it right for this part but it doesn't get it right for the second

10:17 part. Big thing here it seems is that you can see that the model has jumped on preapproval rather than just approval. So if we change it to preapproval both of them get approved now. It's basically

10:29 both scored as true. The untrained model gets it right, but only just. And the other one is still getting it wrong. And you see to actually get it to be both right. I actually had to change this to

10:41 pre-approval for the conference ticket before she bought and then straight away both of them can get it right in here. So these things can be a little bit hit and miss here. Okay. Next up is the

10:51 Decider from Mapa. So this is built on the Quen 3.5 2 base model. So this is actually a much smaller model. And this one's a little bit different in that rather than

11:02 reading the sort of normal output logits, it reads the hidden state at each end of slot and projects that onto the options. So it does the same three types as Jev. So you've got choice,

11:14 score, null. It can take up to a 32k context window. And the cool thing with this one is it actually serves back the same wire format as the type safe API. So you can see here we're basically

11:26 putting in I was charged twice for an order last night and the duplicate is still pending. I need this fixed before my card closes this Friday etc. You can see that what it's doing is we got

11:37 choice for basically multiple criteria being billing technical sales other urgency is in there escalate whether this should sort of escalate up and then actually measuring the sort of

11:49 frustration level. So if we run it, you can see one, it's actually very quick and I find actually if I run it a few times, yeah, it gets actually quicker like that. Here you can see sure enough

11:59 it's got that was billing that that should be handled this week. The urgency it's probably should be handled today, but cuz the person's saying by the end of the week and you can see it's really

12:08 not sure too much whether it's this week, whether it's today. Should it be escalated? Yes. And then a frustration score, it's sort of halfway between very frustrated and frustrated. So, it's

12:20 cited on very frustrated here. The cool thing with this one is just the speed. So, if we run like 10 of them, you can see we're averaging 33 milliseconds per answer here. Now, one of the things I

12:31 noticed about this model is that it's often on the fence. So, with Jev in my testing, it was pretty decisive in its sort of confidence about something being one thing or the other thing. With this

12:44 model, it's very indecisive and it tends to be sort of in the middle for a lot of the different tasks. And while for certain things it does get it totally right, I got to say one of the things

12:54 that is nice about Jeb is just how the majority of classifications that I made in it seem to have a very high confidence, it was confident of its choices, not just a model hedging in

13:04 there. Okay, so the other three that I've put in here, the first is Alex Ortega's open jev. And this one really builds on the whole entailment idea from before. One of the cool things that this

13:15 is doing here though is that we can actually put images into it quite easily. So you can see this is an image going in and we can run these through and get scores for all of these. So you

13:27 can see sure enough the payment failed. So we've basically got each of these statements about this here and it can go through and give us the sort of true or false for each of those. And you can see

13:38 I can come in and just add another statement there. And we can see that okay, the try button is green. No, it's not. It's blue, right? And sure enough, it changes it there. If I change it to

13:48 actually asking that the try button is blue, now it's saying yes. And you can use this idea here on images as well. And the cool thing for this one is they've actually got like it playing

13:57 Doom, right? So you can see that if you pass in these things, you give it the choice of turning left, turning right, shooting, whatever it actually is, the model can actually start to play some of

14:08 this stuff. Another demo that's really cool in here is the whole sort of Minecraft. So if you basically have it move around, they can get it to actually make different kinds of decisions via

14:21 choices, right, based on what's there. So I got to say that this one I think is really kind of cool of how it's able to do it with images as well. Next up is diffusion gemma. So this is a totally

14:32 different class of model rather than just doing the auto reggressive thing of calculating the prefill etc on the attention heads. This is basically doing it all through diffusion. So it runs a

14:42 den noising step in here and then it's reading the log probes of those slots and it just benefits from the fact that it's just so quick in here. So you can see we've got a whole bunch of different

14:53 tickets here and we're trying to basically work out which team should handle this ticket. So this is the choice kind of thing in here. We've got the null of does the customer need a

15:02 reply within the hour. We've got how is the customer feeling? And then we've got some choices there. So with this we can run just one ticket at a time. You can see the speed on this is kind of insane

15:13 which is definitely true with diffusion Gemma. I never got around to making a video about this, but in my testing, you could get super fast tokens from this. The only thing I felt with this model

15:24 was that it had kind of been undertrained compared to say the Gemma 4 31B or the Quen models of things around that size. But speedwise, this thing's insane. So, you can see this is how many

15:38 decisions it can make per second. If I just refresh it and give it all the tickets here, we've got 16 tickets. I run all of them. You can see, boom, it's just done, right? So, in that case, it

15:49 was basically able to fly through these. And you can see it's done a pretty good job of some of these. I don't know, like this one, I probably think I would maybe say it's urgent. I don't know. I guess

15:58 it really depends on how you want to evaluate these. And just passing more of a prompt in is probably going to influence these things as well. But see, sure enough, things like this, it knows

16:08 that it's technical. It knows that the user is annoyed. We can see the latency times coming back here. And you can see that we can actually get really high concurrency here as we go through this.

16:18 So this one is definitely the one to try if you can run this model and you want to do sort of insane amounts of questions on this. It would be very interesting to see the Gemma diffusion

16:28 doing some of the other tasks in here. Now you should remember of course that this is not a small model. This is a 26B model here and in their test they were basically running it on a DJX Spark.

16:38 Here I'm running it on the RTX Pro 6000. So lastly, one that I would say is kind of like a bonus in here is NanoJV. So this is a tiny 600 million parameter model. And you can see that here we're

16:51 basically giving it a game of snake where it's given choices. So you can see that if I just play this, you can see that it's got these choices whether it goes east, west, north, south, and then

17:03 it's given this sort of grid format of where things are so that it can actually do it. So if I speed it up to fast, you can see that actually it does a really good job at being able to play this

17:15 game. So it just shows you I think this one that if you can repurpose ideas to just be simple classification tasks like this one where it's just basically given a bunch of different options in here

17:27 that can actually let this model do a whole bunch of different things going through. You can see at the end it basically was running out of options and was going to basically hit its tail no

17:36 matter what. Okay, so one more that I wanted to add in here is Leila. So if you remember back earlier, I talked about the whole idea of sort of BERT models as being the old way to do this.

17:47 Well, Leila actually is a BERT model. It's built on modern BERT. So this is actually a sort of much newer version, not the original 2018 version of BERT. And it's actually built on modern BERT

17:58 large. So it's about 420 million parameters, which remembering back to 2018, that was huge then, but now is just absolutely tiny. And the cool thing is this one can do multilingual as well.

18:09 So I think it covers around about 100 different languages in here. It takes the three same types of questions. So you've got choice, score, null, and then answers them all in a single forward

18:19 pass. So the cool thing here is they've actually put all of this into a Google Collab. So you can come in here and actually try it out yourself. And because the model is actually so small,

18:29 it's actually just running on its T4 GPU in here. Along with the actual Lelaya model, there's also a really nice write up of a person who actually created this over a year ago. And in many ways, it's

18:40 quite understandably bitter about no one paying attention to his research, yet a very well-funded lab gets all the credit for these kinds of things. And while I do think Leila is quite impressive, when

18:51 we change the Jevbench weightings to just taking account to accuracy here, we can see that Jev is really competing with other large language models that a lot of these models below it and

19:02 especially something like Leila are actually doing a lot worse most likely because they just weren't trained on the sort of generalization. And I will say that a wellunded lab should be expected

19:14 to be able to train a model that's going to generalize a lot better than someone who's just making a fine-tuned version of an older BERT model, etc. All right, so which one of these should you

19:23 actually use? If you want to do this today with no training, you could do the SIM if style readout on a model that you're already using already. If you want a small self-contained server that

19:34 speaks the Jev format, then probably decider might be worth trying out. And then especially if you've got some kind of narrow policy decision, take Nimble's recipe and run it on your own data. The

19:46 whole sort of idea of swapping something in, swapping something out to actually do the contrastive thing. And lastly, if you need images, check out one or the actual decider vision one there. The

19:58 challenge is going to be that for really hard stuff, the sort of multihop stuff, the date arithmetic, really none of these are working that great. And so what might be your best option there is

20:09 to actually sort of cascade those things. Let a fast model take everything first and then when it's not very confident of something pass that one up to a reasoning model on a low or sort of

20:19 medium effort. And I think that can be a really powerful trick of sort of being able to get the benefits of prices and speed of system one on most of what you're doing and then only go to system

20:31 two where you really need it. Overall, I would say the open versions are closer than I expected on sort of the easy and standard stuff, but there's definitely a real gap on some of the hard stuff. So,

20:41 as always, I'll put the links to all of these in the description. You can go and check them out. If you haven't seen the original Jeb video I made, check that out over there as well. And then, I'm

20:50 really conscious that as I've been recording this, there are probably more of these things coming online all the time. So, let me know in the comments if you found some that worked really well

20:58 for you or if any of these are breaking on certain things that you can't get to work at all. And I've been looking at this myself and starting to think about the mini CPM 52B model could actually be

21:09 a really nice model to try for some of these tasks and for doing your own sort of luras and stuff for this kind of thing. So, I'm going to probably have a play with that. If people are really

21:19 interested in that, let me know as well. and maybe it's something that I look at doing a video about over the next few days. Anyway, as always, if you found the video useful, please click like and

21:28 subscribe and I will talk to you in the next video.

Frontier News · by Hyperjump Technology