Run a Jev-Style Model on Your Own GPU

summarized

TLDR

The core insight is that many AI decisions don't need text generation—just scoring logits against a fixed menu of options. Open models like Kev (a drop-in replacement for Jev) and Semif (logit scoring on any capable model) achieve ~81–87% accuracy on routing tasks while running locally, for near-zero per-decision cost. The practical tradeoff is a few percentage points under Jev's hosted version in exchange for keeping data on your own hardware.

Key points

Jev is a hosted decision model that returns only a probability for a fixed set of answers.

Local Kev models (0.8B, 4B, 9B) are trained on Qwen 3.5 and aim for Jev's interface.

Kev 4B and 9B score about 87% on seen tasks, close to Jev's 88%.

Semif applies logit scoring to any open model, reaching ~81% without fine-tuning.

Vonn uses a 395M-parameter BERT encoder for CPU-friendly decisions under 15ms.

Tools mentioned

Techniques

  • logit scoring
  • key-value cache sharing across multiple questions
  • calibration via Brier score
  • entailment-based training for image+text decisions
Transcript (captions)

0:00 Watch this. A message lands in a support inbox. My card just got charged twice. What do I do? A model reads it and picks one action out of four. Refund, escalate, reply, or ignore. It answers

0:12 escalate 87%. That took half a second on the graphics card in this machine. Nothing about that customer left the room. That's the whole pitch here. Take the fast decisions a cloud AI makes and

0:24 run them on the GPU already own. The hosted version everyone's talking about is called Jev. It's very good. It's also closed. Using it means shipping your data, your users' messages to someone

0:35 else's servers. So, how close can you get on hardware you control? By the end, you'll know which model to run and what it costs. Start with what Jev even is because it's strange. It doesn't write

0:45 text at all. You hand it your program state and a question with a fixed list of allowed answers and it hands back one answer with a probability. Type safe, the lab

0:55 behind it calls this a system one model after the fast automatic side of how we think. A normal language model writes a word at a time. So, even a yes or no spins up the whole sentence machine. Jev

1:06 skips that and just decides. Software makes thousands of tiny judgment calls inside a single request and that's the reason anyone wants this. Which bucket is the form done? Which button next?

1:17 Most of them never needed a paragraph because it never writes a paragraph, it's almost free. 4 cents for a million words of input and answers back in well under a second. Put real numbers on

1:28 that. A million decisions from Jev is a few cents milliseconds each. Ask a chat model to make the same million one-wordy answer at a time and you're into real dollars and long minutes of waiting. Jev

1:39 takes three shapes of question, a multiple choice, a yes or no with a probability, and a rating on a scale. Almost every routing or moderation call you make is one of those three. Teams

1:50 like one more thing, the answer is guaranteed to be a valid type. Because the choice is is to fit your menu, it can't return something your code didn't expect. Type-safe even markets zero

2:00 hallucination, though its own page admits that covers the format, not whether the pick is right. The person behind it, Diogo Almeida, was on the original team that built ChatGPT. His

2:11 pitch is almost a confession. Chat didn't deliver, so he built a model that decides instead of talks. So, why not just use it? Because Jev is hosted. To ask it anything, you send your program

2:22 state to Type-safe servers, and that state is usually the sensitive part. Your users' messages, your internal documents, the customer record you're trying to sort. For a hospital, a bank,

2:32 or anyone under a data agreement, that's a hard no. If the decision can't run in the building, the automation just doesn't happen. So, here's the plan for the rest

2:41 of this video. I'll show you how the trick works, the three open models worth running, what to point them at, and the exact price you pay in accuracy and in GPU. Start with the trick, because once

2:51 you see it, the local versions make sense. The keyword is logit. Before a language model says anything, it gives every possible next word a raw score called a logit. Normally, it turns

3:02 those scores into probabilities, picks a word, says it, and repeats. Writing the sentence is the slow, expensive part. A type decision throws a sentence away. You give the model your

3:12 allowed answers: refund, escalate, reply, ignore. And instead of letting it write, you read the logit for each answer straight off the model. The highest score wins. That's one forward

3:23 pass through the network, so it can't invent a fifth option, and no sentence ever gets produced. You can even ask several questions about the same state at once. The model reads your document

3:33 one time, keeps its notes on it, the key-value cache, and every question reads from those same notes. One read, many decisions. That shared read is why throughput jumps. On one open build,

3:44 scoring decisions from scratch runs a couple per second. Reuse the notes across questions with the same context, and it climbs to about 20 a second on the same card.

3:54 An independent write-up of Jev's architecture puts it in one line. Give it the state, the questions, and the allowed answers, and it returns the answers in parallel without generating

4:03 text. Put numbers on it. Say the model reads a 200-word support ticket. Reading it is the cost. Scoring your four options on top is a few milliseconds, basically a rounding error, which means

4:14 a small model that could never write a good reply can still be excellent at picking the right bucket. The hard part was the writing, not the decision, and we just deleted it. One more piece,

4:24 calibration. When the model says 87%, you want it right about 87 times in a hundred, like a good weather forecast. Jev's trained for that, so the number becomes a dial.

4:35 Auto handle the confident calls, send the shaky ones to a person. There's a wrong way to do this, and one project shows it clearly. GitHub's own local Jev asks a normal model to write

4:45 the probabilities out as JSON. It works on the wire, but the model is guessing its confidence in words, not measuring it. Their readme says so. The numbers are self-reported, not read from the

4:56 logits. So, that's the mechanism. Read the scores, don't write the answer, and trust the number only as far as it's calibrated. Every local model here is chasing that. Now, which ones earn your

5:06 disk space? Start with the one I'd hand most people. It's called Kev. Yes, Kev instead of Jev. It's from Jared Palmer, who built serious developer tools like Formic and Turbo Repo, so it isn't a

5:17 toy. Kev is a small family of decision models built on Qwen 3.5, the open weights from Alibaba. Three sizes, 0.8 billion parameters, 4 billion, and 9 billion. You download the weights and

5:30 run a little server. Palmer built it from a public write-up of Jev's architecture and put the whole thing together with an AI coding agent, which tells you how reproducible this idea has

5:40 become. A couple of weekends, not a research lab. Kev copied Jev's interface exactly, and that's what makes it the easy pick. You take Type-Secure's official software kit, change the

5:50 address from their cloud to your own machine, and the same code just works. Your app can't tell the difference. Palmer also published his homework, which is rarer than it should be. On the

5:59 tasks Kev was trained on, the four and nine billion models score about 87% right on Jeff's heels at 88. But watch what happens on tasks it hasn't seen. The nine billion holds up around 85. The

6:12 four billion slips to the high 70s. The smallest one, point eight billion, falls to 65, barely better than a coin flip on a two-way call. The calibration tells the same story. There's a score

6:24 for how honest the confidence is, called the Brier score, where a lower is better, and it nearly halves from the small model to the big one. Bigger model, more trustworthy number. On

6:34 hardware, Palmer says the four and nine billion both fit on a 32 gigabyte Mac. In practice, the four billion is a sweet spot. A mid-range card runs it, and it's close enough to Jeff for real work. So

6:46 Kev's the drop-in. Same interface, public evals, runs on hardware you might already own. If you want one answer, that's it. But it's worth seeing the trick with no training at all. Semeth is

6:56 that version, from a developer called Theo Lee. His tagline is the whole spirit of this: semantic ifs from open models on a 3090 at home. Semeth doesn't ship a new model. It takes an open model

7:09 you already have, a four billion quan, and runs the logit scoring trick on it directly. No fine tuning, no training run, just the scoring. On the shared benchmark, it lands about 81% against

7:21 Jeff's 88. A few points back from a model you didn't train on a gaming card, and it's more than five times faster than making that same model write the answer as JSON. Semeth point is bigger

7:31 than one repo. If the trick is just reading logits, then a capable open model you already trust can become a decision engine today, without waiting on a special release. Lee is careful

7:42 about it, too. His comparison covers about 100 aligned questions, not the full set. And he says outright, you have to check calibration for your own workload. That's how you're

7:51 supposed to publish a number. So far, Kev and Semif both want a real GPU, though. So what if you don't have one? Then you want Vaughn, the small hardware answer. And it's built completely

8:02 differently. Instead of a chat model, Vaughn uses an encoder called modern BERT, which just means a model that reads text to judge it rather than to continue it. That's exactly what a

8:12 decision is. It's pre-trained on 2 trillion tokens, so it already knows language well. On top of that, Vaughn adds three decision heads. Choose an option, answer yes or no, or give a

8:23 rating. The whole model is 395 million parameters, about a gigabyte and a half on disk. That size means it runs anywhere. A CPU, an old laptop, a phone class chip, no graphics card required.

8:36 And each decision comes back in under 15 milliseconds. This is the one you embed in an app and forget about. The trade is accuracy, and it swings hard by task. On a clean

8:47 routing job, sorting into 20 buckets, Vaughn scores about 83%. Spread it over a messy suite of 49 different tasks, and it drops to 71. Sharp where it's focused, shaky where it isn't. So those

8:59 are your floor and your ceiling. Vaughn on a CPU for reach, Kev or Semif on a GPU for accuracy. Two more are worth a look because they add something the first three can't. The first is Open

9:10 Jev, from a developer who posts as Alex Wartega. It's trick is a sense the others don't have, sight. The 4 billion version reads images as well as text, so the state you hand it can be a

9:21 screenshot. Under the hood, Open Jev is trained on entailment. Does this evidence support this claim? Yes, no, or unrelated. Framed right, that covers a lot of decisions, and the image version

9:32 answers the same question about a picture. So you get decisions a text-only model can't make. Is this uploaded photo a receipt or a business card? Does this screenshot show an error

9:42 or a normal screen? It scores the options over the picture the same way it would over a sentence. The second is a warning as much as a tool.

9:50 Leia is tiny and speaks 100 languages, but its own docs are blunt. Out of the box, it's near random. On their benchmark, the base model scores 36% basically chance. Fine-tune it on your

10:01 actual task though, and it jumps to 77 past Dev on that same set. So, the real lesson under all of these, the good ones aren't oracles you plug in. They're fast bases you point at your problem. There

10:13 are more and they mostly show the same trick from a different angle. A 9 billion adapter for heavier context. A framework that turns any open model into a decision endpoint. A nano version

10:23 tuned for video games. Underneath, it's the same move. Read the logits, skip the writing. At this point, you've seen how it works and what to run. So, what do you actually point it

10:33 at? The sweet spot is high volume and low drama. Sorting support tickets by intent. Tagging feedback as bug, billing, or praise. Flagging spam before a human sees it. Scoring a lead from one

10:45 to five. All decisions, all cheap to double-check, all things you'd rather not send to a stranger server. Agents are the loud case. A browser agent that used to ask a big model what to do on

10:55 every page can ask a tiny local one instead. Which element, which action, one shot. Fewer round trips, lower cost, and it stays on your machine. Running one is mundane, which is the point.

11:07 Download the weights, start the server, point your code at localhost. On a model that fits your card, the first decision comes back in well under a second. Where it falls down is the opposite. One-off

11:18 calls you think about for a while anyway, or decisions with no fixed set of answers. If you can't write the menu, this isn't your tool. So, those are the tools. That leaves the two questions

11:29 that decide whether you'll actually do this. How much accuracy you give up and what it costs to run. On accuracy, be straight with yourself. The best local setups land a few points under Jev on

11:39 the task people have measured. Call it low to mid 80s against Jev's high 80s. That gap is tiny for sorting a support queue. It is not tiny for approving a loan. So, what can it reliably handle?

11:50 Close choices with clear options and plenty of examples, it's good at. Open-ended judgment, long messy context, or anything someone is trying to game. That's where the gap bites and you keep

12:01 a person in the loop. Running cost is close to nothing. No per token bill, no metered API. You pay for the GPU once or borrow one you already have and a million decisions cost you electricity.

12:14 On hardware, the latter is clear. No GPU, run Vonn on the CPU. A gaming card with eight or 12 gigs, run the 4 billion Kev or score a model you already trust with Semith. Want the 9 billion's

12:27 accuracy and you're into a workstation card or that 32 gig Mac. If you want my pick, it's the 4 billion Kev for most people and Semith when you'd rather score a model you already run. Kev

12:38 because it's a true drop-in. Keep your code, swap the address, read the public evals. Semith because it proves you don't even need a special model, just a technique.

12:48 The one caution I'll stand by, don't treat any of these numbers, Jev's included, as ground truth. They're mostly self-created. Before you wire one into anything that matters, run your own

12:58 100 examples and check that 87% confident really is right about that often. But, look at what you get for the work. That support ticket from the start, my

13:07 card got charged twice, gets read, sorted, and escalated in half a second on a card in your own rack, and the customer's words stay on your own machine. The hosted model is faster to

13:17 set up and a little sharper. The local one is yours and your data stays home. For a whole class of quiet, high-volume decisions, yours is the right trade. And once you've seen that a

13:27 decision is just reading scores off a model you already run, you start noticing how much you've been paying a cloud to decide for you. If you want the full build, the setup, and the models,

13:37 say so in the comments, and that's the next one. Subscribe for more local AI teardowns like this.

Frontier News · by Hyperjump Technology