Enjoying this issue?
Get tomorrow's AI & engineering digest in your inbox — hand-picked, summarized, and always spam-free.
TLDR
Jev, a decision-based model from Typesafe, challenges the assumption that general-purpose LLMs are the best fit for workflow automation by using structured inputs and probabilistic outputs to achieve 40-200x faster response times. Unlike autoregressive models, Jev supports parallel sampling and typed probabilistic decisions, making it suitable for tasks like sorting, RAG, and game playing. The real significance is a shift toward specialized, narrow models optimized for specific constraints rather than brute-forcing general-purpose foundation models.
Key points
Jev is a decision-based model that uses structured inputs and outputs with probability distributions.
Jev achieves end-to-end response times of 70-500 milliseconds, 40-200x faster than current LLMs.
Jev's primitive types are choice, score, and null, enabling categorical, scoring, and yes/no probability questions.
Typesafe has not disclosed details of its RLCD method, but similar ideas exist, like a Reddit model using bidirectional BERT.
Jev competes with flash or nano models like GPT-5.6 Luna, DeepSeek v4 Flash, and Sonnet 5, but only for workflow-specific tasks.
The 421 million parameter open-source model can run on consumer hardware.
Tools mentioned
Techniques
- Parallel sampling
- Typed probabilistic decisions
- Structured input/output formats
- Reinforcement learning with verifiable reward (RLCD, undisclosed)
Stop scrolling. Start reading smarter.
Receive the day's most important AI & engineering updates in one concise email. No spam.
Transcript (captions)
Jeff is a new type of model that makes us question how we should really think about optimizing LLMs. When we follow the orthodox path in AI from chachi that was released in 2022, models have been
optimized to assist humans in chat applications. Of course, as coding agents like clock code and codecs became mainstream around 2025, we also optimize the models for this as well. But when we
look at the AI stack, how AI actually gets end up being used in the application layer often puts a downward pressure on the layers below. And as a consequence, they tend to morph
themselves to optimize for the best outcome in the application layer. And models continue to optimize to be more helpful to humans and also optimize to be more helpful in coding agents. But
one area that always felt short in use cases was workflow automation. Even a highly intelligent model like Astra and Fable couldn't really cross a threshold here without a huge cost on the way. and
type safe AI that made Jev is making an assertion here that the reason why this threshold is notoriously difficult to cross even for highly intelligent models is because the state-of-the-art models
are optimized for different things. The way that they articulate this is a schism in the orthodox line where the optimizing the model for human preference and further optimizing them
for verifiable reward doesn't really carry over when decisions need to be quick given the uncertainty. Now, what we're talking about here isn't that one path is ultimately superior than the
other, but Jev is sort of this anti-thesis to our current trajectory in AI and arguing that we really have been neglecting the use cases that stand to make a huge difference in automation.
And types of AI is arguing that trying to force a model that's optimized for human interaction and agentic use cases into automation is fundamentally wrong way to go about it. When we look at all
the different projects that people are showing off on social media using Jev, a lot of them are actually more showing off the speed of execution rather than the depth of understanding. Sorting
emails, improving rag, playing games, model routing are all tasks that current LLMs can do, but certainly not at the speed of Jev that according to them are 40 to 200 times faster where end-to-end
response time is 70 to 500 milliseconds. Functionally, current LLMs can do everything that Jev offers, including mimicking the response so that it returns a result just like Jev would.
But matching the latency of 70 to 500 millisecond is not something that auto reggressive models can easily do, especially since tokens are generated one after another until completion. But
Jev is inherently designed to do parallel sampling and typed probabilistic decisions, which is something that LLMs don't do in the architecture. So it's less about what
Jev is functionally capable of, but what we're optimizing the model so that we can cover more ground in the application layer so that automation, game playing, mass sortings, these types of taxs are
practically possible. Now let's take a closer look at the model Jev and see how it actually works. But before we do that, we have to talk about coding agents which can be complicated to use.
And that's why Juny CLI, an AI coding agent from Jet Brains could be perfect for you. Juny is what scored near the top right here on SUI rebench leaderboard as you can see at 61.8%
resolved. So I can bring this very coding agent directly into my terminal right here and ask Juny to work on complex coding tasks. For example, here I'm asking Juny to keep my website up to
date. And since there's a lot of information to update, I can hit shift tab and put Juny into plan mode first. Instead of immediately changing code, Juny will then break the task down into
these multiple requirements. As you can see the design, the implementation stages and the tests that it wants to run. And now this plan is actually saved inside the folder.juny/plans.
So now I can review it before actually proceeding with Juny to start implementing the code. I can also bring my own API key and choose which model I want to use per task, which means I have
more control over the usage limits and pick and choose the models I want for the job. For a limited time, Gemini 3.8 Flash runs at 75% off base price in Juny. While your daily driver can still
be Gemini 3.7 as of default. You can open the ID plugin or Juny CLI, run a real task and see how much you might save. Link in the description below. Unlike traditional LLMs where we give
them an instruction and the model responds in text. Jev is different because it basically forces you away from raw text as input, but putting them in a structured format that looks
something like this. and Jeff responds back in a similar structured format but with a probability distribution. Okay, what do I do with this? And that's where a lot of people who are used to
interacting with AI through traditional LLMs sort of get stuck because our rules of engagement is totally different now with Jev. The basic primitive types of Jev are choice, score, and null. And you
can ask Jev a categorical question, score based on order choices, or a probability of yes or no question. And these are sort of these building blocks that software can now use to build on
top. It almost feels like we're back to logic gates and registers and now we have to build many abstractions on top to build more complicated application as an abstraction. For example, I can take
this long list of plants to sort through and running them through auto reggressive models like cloud opus 5 would inherently be different than asking Jeff to do the same task which is
done in a matter of a second or two. And of course you can also follow different design patterns given different permutations of these building blocks. So when we look at the parto frontier
Jev is competing closer to flash or nano models like GPT 5.6 Luna deepsec v4 flash or sonnet 5 but only on use cases like workflow specific tasks. So what I hope to see going forward is more
coverage in this paro that not only allows us to use models that are optimized to assist humans in solving complex and creative tasks and also daily driver models that helps us assist
in mundane tasks. But now with models like Jeff, an explosion of use cases in this section as well as more models continue to pop up and help us make actual automation possible from
companies that built on top of these primitives to create something useful and valuable. And as long as demand and use cases continue to exist, we'll likely see more models in this section
here. While Typesafe hasn't really disclosed their details around how RLCD method actually works, we've had many similar ideas in the past around the same idea. In fact, someone on Reddit
had already made a very similar model using birectional BERT. And this 421 million parameter model that's open- source can easily run on consumer hardware. So, how I process this entire
news around Jev is really more about our use cases growing horizontally than about the novelty of Jev itself. since narrow models that are specialized like this is something that we've had before
generative AI completely took over the public narrative and I do think that this is a net positive for the ecosystem to now appreciate a more wider landscape of models that doesn't necessarily brute
force a generalpurpose foundation model for all tasks and as we hit constraints that are made by the architecture and the training objectives we can now start building models that are optimized for
the right constraints instead Head.