Jev - The Ultimate Classification Model?

summarized

TLDR

Jev, a model from Typesafe AI, is built exclusively for fast, cheap classification tasks and outputs only structured values (choice, score, or null) instead of generating text. It costs 4.2 cents per million input tokens with free output tokens and responds in 70–500 milliseconds, making it a practical alternative to fine-tuned BERT models for system-one decisions like routing support tickets or detecting PII. The model is trained with a new method called RLCD (reinforcement learning for calibrated decisions) and cannot break its output schema, though it can still select the wrong option.

Key points

Jev supports only three output types: choice, score, and null (yes/no with probability).

The model costs 4.2 cents per million input tokens and zero for output tokens.

It responds in 70 to 500 milliseconds per query.

Jev is trained with a new method called RLCD (reinforcement learning for calibrated decisions).

The model cannot hallucinate in the sense of breaking the output schema but can still pick wrong options.

Tools mentioned

Techniques

  • RLCD (reinforcement learning for calibrated decisions)
  • system-one classification
  • parallel sampling
Transcript (captions)

0:00 Okay, so if you pretty much look at what every frontier lab has been doing over the past 2 years, it's all been in one direction. They've all been focused on reasoning. And to get that reasoning,

0:13 mostly they've been focused on longer chains of thought and thinking budgets. And while those models are great, they'll happily sit there for multiple minutes before they even give you an

0:22 answer back. And of course, if you wanted to see that chain of thought, the Frontier Labs are not going to let you see it, even though you're paying for it. So if you read Daniel Canman's book,

0:30 thinking fast and slow, you know that all this kind of reasoning stuff is system two thinking. It's slow, deliberate, effortful. Whereas system one, on the other hand, is fast,

0:41 intuitive, sort of like a gut call that you make really quickly without any deliberation. And here's where the subject of today's video, I think, is really interesting. Most decisions that

0:51 actually sit inside of software aren't system 2 problems. They're often just really simple classification problems. What kind of support ticket is this? Is this message urgent? Did an agent's

1:02 output actually break a rule? You shouldn't even need 30 seconds for that kind of thinking, let alone the sort of minutes that some of the models end up taking. And while it's great that you

1:13 can ask one of these reasoning models for a one-word answer and then wrap it in JSON, you're paying a huge latency cost for generating all that text before you get your one-word label back. So

1:25 this week, a new lab that's just come out of stealth has gone the complete opposite way of the reasoning models. They've built a model that you can't chat with at all, let alone have it

1:35 reason and deliberate over things for a long time. It just doesn't generate text in the standard auto reggressive way. And the funny thing actually enforcing that is that the output tokens of this

1:45 model are actually free. But actually working out why those tokens are free might be the thing that actually tells us a lot about how this thing is built. All right. So the company is called

1:55 Typesafe AI and the model is called Jev. So the founder is Dio Almida and actually he was at OpenAI for quite a while before and funnily enough was one of the top authors on the instruct GPT

2:07 paper. That was the paper that actually led to the whole sort of instruction tuning of models which has just progressed onwards from that to the reasoning models that we see today. So

2:17 if you don't know that paper basically took a raw language model basically like what we would call a base model now and actually made it follow instructions and out of all that research came chat GPT.

2:29 So this is someone who clearly helped make models good at talking to people. And the question that he says that's been bugging him for the last four years is that as models have become superhuman

2:39 at chat where's all the automation here? And interestingly, he's saying that chat is the wrong interface for software. Software doesn't want a paragraph. It wants a value that it can basically use

2:50 straight away. Now, Dio and the rest of the team at Typescafe AI have spent the last two years in stealth working on this model called Jev. All right. So, what is Jev? So, they specifically refer

3:02 to this as a system one model. And probably the easiest way of looking at this is it's kind of like a function call. You pass in two things. The first is what they call the state which is

3:13 just your unstructured text or data and that could be anything that you want classified. It could be a support ticket, could be an agent trace, it could be a log file, etc. And then the

3:23 second thing that you pass in is a set of typed questions. And there are only three kinds of questions that you can ask here. There's choice where you give it a list of options and it picks one.

3:34 There's score where it rates something on a scale that you defined. And then there's one called null, which is a yes or no question where it basically comes back with a probability that the answer

3:45 is yes. So if you look at what comes back here, there's no text to pass for the choice question. You get the option it picked and also a probability of every other option. And I think this is

3:56 really kind of interesting because this is one of the things that a lot of us have kind of been trying to fake with JSON where we ask models to do stuff and return JSON output. But in many ways,

4:06 the number that you're getting back there is just generated text. And if you look at things as LLM as a judge, you'll often see that that text skews in a very certain way. I.e. it's just a model

4:18 writing the character 0.9 rather than actually a real representation of probability. The way they want you to use this model is to think about it as being a smart if statement. You don't

4:30 ask one big fuzzy question like rate this startup pitch. You ask small gut questions in there. What's the feasibility of this? What type of market are they going after? And of course, if

4:41 you're asking a choice question, you have to actually give it the options to choose from. Then the cool thing is that you can actually combine all of those in normal code. So when something changes,

4:51 you can just change a number in your code. You don't need to go and tweak any prompts or anything. All right. So what I've done is basically just code up a little demo app to try this out. I'm

5:03 using the model on open router. So, this is a crazy sort of price model, right? You're looking at 4.2 cents per million tokens in and 0 cents per million tokens out. And you can see if we come in here,

5:15 they've got some code to actually how to call it on using the open AI uh API endpoints. And so, we've been playing around for a couple of hours with different kinds of demos and trying out

5:26 different things. So, the first key thing here is that there are three kinds of things that you can do, right? One of them is a choice. So you can see here I've got some text that I'm going to

5:35 pass in and then it's got a choice of five different classes out here. And it's very good at being able to detect which of these is the right one. And so what it will do if we look at the raw

5:48 output here is that it basically has a type choice. In this case the choice was French. And then it gives us the probabilities for each of these and a confidence score. And you can see that

5:59 it did that pretty quickly. So, I'm probably on the opposite side of the world to where the model is actually being hosted, but you'll see that it's very quick at being able to respond even

6:10 for things like this where I wanted to try it on Thai, but using with Thai characters, it's going to give it its way straight away that it's the Thai language. So, I asked it to basically

6:19 make like a romanized version of that. And you could see that even that it has no problems being able to get that right. The second thing that you can do with this is you can get it to give a

6:31 score. So here you can see that we're giving a score between zero and two. We're sort of doing sentiment here. We're doing a classification task again. We've got sentiment coming out. And if

6:42 we go for positive sentiment, you'll see it comes out two of two. If we go for negative, zero. If we go for mixed, it's pretty good at being able to do that. I can see each time this is costing me

6:56 0.0014 of a cent right in here. If you can break down what you're trying to do to lots of different classifications, you really can extremely cheaply do a

7:11 lot of stuff with this. And you can see in this case where I tried to make it a little bit more positive than in the middle, I it does change coming back each time. So it still is a stochcastic

7:22 process going on here. And we can see the confidence is less than before. Each time I ping it, I'm getting something slightly different but pretty close to the same sort of ballpark scores there.

7:34 When we look at the API for that, you can see that okay, the probabilities that it's coming back. So you can see in this case, we've passed in a score, we've got that back coming out of there.

7:43 So scoring is a really nice function in there. And then lastly, we've got the true or false or the yes no kind of probabilities here. So this is a null and it's basically going to give back a

7:53 probability for what we put in. Now if I ask it something simple like that, it comes back 87% yes. Is this text asking a question? Okay, the answer would be no here. So you can see that it's 2% yes.

8:07 So basically it's no, right? You could kind of think of this as being flattened out by a sigmoid function. So you've got zero to one coming out of this. And even if we do things like where okay, we've

8:17 got a question but we've got no question mark, it's very confident that that is still a question without a question mark. So it's not like it's just looking for a question mark there. And you can

8:27 see different kinds of questions will get different responses. So interestingly, if I just ask it, can you help? It's a lot less sure if it can because I haven't been specific. But if

8:38 I ask it, can you help find my cat? Well, then it's 95% sure that it can do it. And if I just change one word in there to can you help me feed my cat, it goes back down a bit. So, it is

8:51 interesting how it actually does this. You will see that like again, this is still a stochastic process. It comes back slightly different each time, but it is pretty consistent of you either

9:02 being able to say one thing or the other thing. And you can see if I just give it two words, then it starts losing its confidence about this. So, it does seem to me the longer that the input that you

9:13 put in, the more confident it gets with its responses out. All right. Now, looking at some sort of more practical kinds of things that we could do here, you can see that we could put in like a

9:24 message of, I was charged twice for an order. Please refund the duper. Get it to work out which team it's going to route it to. Here it's billing. Kind of obvious. Sales one, it gets 100% sales.

9:37 What if we do something a little bit more ambiguous? You can see now it's not very confident. It's actually coming back that it's unclear as as the class, but it's still thinking that this is

9:46 more of a technical kind of thing. And you can see that each time we run that, we do get a slightly different response, you know, going through this. Now, on top of these, we're actually doing

9:55 multiple things at the same time. So, you can see that we've also got a null in here for was a refund requested. Obviously, if I select this, the answer is yes. Is it sort of timesensitive?

10:07 Right? In this case, it's saying timesensitive 7%. If I change that to please refund the duplicate charge right now instead of just right now, you can see the time sensitive is jumping up to

10:19 sort of like 60 70% in this case. We can even try it with things like injecting stuff in there. So, doing any sort of thing where someone's trying to do a prompt injection or something like that.

10:31 This seems to do pretty well at being able to get that kind of thing. Sarcasm is also something that it seems to be doing an okay job at it. It would be interesting also to test it on humor and

10:43 other kind of things, but it is good that it's not being tricked by a lot of different things that normally would trick this kind of thing. Other tasks that you can get it to do is things like

10:53 code reviews, like where you can basically ask it, is something safe? Is something not safe? It seems to do a good job at that. doing things like classification on different types of

11:03 content. That seems to work really well. You can see here it's able to discover PII information, personal identifiable information here. Does a good job at that kind of thing. It does a good job

11:14 at sort of spam detection as well. So the fourth example here is looking at agent tool selection. So this is a little bit like the cactus model that we looked at a while ago in that it's doing

11:26 some kind of sort of function calling. But it's important to understand here that it's not actually extracting anything out and passing it to the tool, right? It's just telling us which tool

11:37 to use as opposed to getting the right arguments out to pass to the tool. So in that sense, function calling models are still able to basically generate out what should be the input going to the

11:49 function. Last up, one of the things that I thought was really interesting to test it is stringing actions together, right? So where we basically give it something and we then want to basically

12:01 run a number of different tasks over this to see how long does it take, how does it actually process these. So, here's 20 different tasks going through and you can see that it's flying along.

12:14 Now, I could have done them in parallel for some of them at least, but you can see that just going through it like that, we've gone through all different 20 tasks in here, gotten the answers

12:25 out, gotten the results back for these, and it's cost just over 12th of 1 cent. I really feel like this is where you're going to find really interesting things, right? when you can string lots of

12:37 different classifications together to be able to get some kind of response out that's really valuable for you in a really quick time. Okay, so how does this all actually work? Well, they don't

12:50 really tell us a lot here. There's no paper, there's no architecture diagram. They don't give us a lot of information. But what they do give us is three pieces that they mention that there's a new

13:00 model architecture, a parallel sampler, and then perhaps the most interesting bit is that they're training with a new method they're calling RLCD. This is reinforcement learning for calibrated

13:12 decisions. Now, how RLCD actually works, they don't really tell us, but they compare that to the existing LLM using both RLHF, so that's reinforcement learning from human feedback, and also

13:24 reinforcement learning from verifiable rewards, which is a whole sort of GRPO and a lot of the ways that the modern models are being trained at the moment. So, it does seem probable to me that

13:34 this is some kind of transformer, but perhaps what it's actually doing is just using the prefill stage to calculate heads, etc., and then doing a prediction out which is just a classification or a

13:45 regression depending on the different tasks that you've got there. One of the cool things here is definitely the speed because there are no tokens being generated one after another. Everything

13:55 comes back in a single pass and that means you're done in about 70 to 500 milliseconds which you can add to the basically round trip time and it's still extremely fast for doing these kinds of

14:07 tasks. Now, one of the things I find fascinating is that they claim it can't hallucinate. And it seems what they're really saying there is that this is not going to break the schema that you give

14:17 it to basically come back with things. So like no broken JSON, no invented tool names, it can only deal with what you've actually given it. It's not going to actually generate sort of anything new

14:27 there. Now, it does seem that it can still pick the wrong option and like a wrong answer meaning that it could have the right type of question but actually give you a wrong answer for that. And it

14:38 seems that this is where they're relying on that RLCD to be able to deliver highquality answers back. Now, we don't know if this is just like a smaller LLM that's been trained in a different way

14:49 or if this is actually a different architecture, etc. It is kind of amazing though that it can do things like play doom where it can do it so quickly that you can basically have it responding to

15:03 these inputs. And the thing that I find fascinating here is that doing that and running sort of 10 queries per second even in that rate ends up only costing about $7 per hour. So overall, I would

15:14 say if you're doing any kind of classification with any kind of BERT model or things like that, this is definitely worth trying out to see how it goes for your particular use cases.

15:26 And I do wonder if this really delivers like they say it does. Is this going to be the end of sort of fine-tuning small BERT models for doing purely classification tasks? Anyway, they make

15:38 the point here that this is early days, that they're really focused on any sort of decision that you need to be able to automate. And my guess is that they've probably kicked off a whole bunch of

15:47 people trying to make open-source versions of this. So, over the next month or so, we might see some really interesting open- source models that come along with similar ideas to this.

15:58 Anyway, let me know in the comments if you've actually tried Jev, where you could see yourself actually using it. And I'd love to hear from people who've actually tested it and found that it

16:07 didn't work for certain things. What were those use cases that it didn't work for? And as always, if you found the video useful, please click like and subscribe. And I will talk to you in the

16:16 next video. Bye for now.

Frontier News · by Hyperjump Technology