Enjoying this issue?
Get tomorrow's AI & engineering digest in your inbox — hand-picked, summarized, and always spam-free.
TLDR
Liquid AI's LFM2.5-8B-A1B is an 8-billion-parameter mixture-of-experts model with 1 billion active parameters, designed for on-device intelligence and tool calling. In testing, it performed well on tool-calling tasks like a help desk simulation and basic coding (HTML/Python), though it struggled with complex code generation and had partial factual accuracy. Its speed (over 100 tokens/sec at 8-bit) and ability to chain tool calls make it a promising option for on-device assistants.
Key points
- LFM2.5-8B-A1B is an 8B total / 1B active mixture-of-experts model from Liquid AI optimized for on-device applications.
- The model demonstrated strong tool-calling capabilities in a help desk simulation, correctly selecting and chaining multiple tools based on ticket information.
- Coding tests showed it can produce functional HTML sites and a text-based Python RPG, but it is not a strong coder and made mistakes on more complex tasks.
- General knowledge tests revealed partial factual accuracy; the model correctly recalled some historical dates but had errors in technical specs.
- The model excelled at roleplaying, staying in character and adapting to nuanced prompts after initial nudging.
- Generation speed exceeded 100 tokens per second on a 5090 mobile laptop with 8-bit quantization, enabling near-instant responses.
- The model extended context length significantly compared to its predecessor (LFM2-8B-A1B) and showed a noticeable leap in benchmark performance.
Tools mentioned
Techniques
- mixture of experts (MoE)
- tool calling
- chain-of-thought reasoning
- on-device inference
- 8-bit quantization
Stop scrolling. Start reading smarter.
Receive the day's most important AI & engineering updates in one concise email. No spam.
Transcript (captions)
start game now and it hyperl it just to the actual Python script. I like that. Today we're going to be taking a look at a new model from Liquid AI called LFM2.5-8b- A1B. Now, as the name may have tipped you off to, this is an 8 billion parameter mixture of experts model with 1 billion active. And the liquid models are always very interesting because they are essentially designed for ondevice intelligence.
So, even one of the most prominently displayed things right here in the key features is an ondevice personal assistant. So, Liquid basically has a lot of small but very intelligent models. I've tested a bunch of their offerings on the channel in previous videos and they've always been very impressive, especially when considering the size. So, I'm excited to test this, which is almost kind of like a larger model, at least in terms of their offerings. So before we get into it, please do feel free to subscribe and let's just begin by taking a look at this model and some of the interesting things of note about it.
So as we saw in the beginning, one of the prominently listed things here is that this is designed to power real life applications, chaining tool calls and following complex instruction on devices. So this model seems specifically poised towards, as it says right there, almost a gentic functionality where it performs tool calls and is good at doing that. Apparently, as they list right here, and from what we can see in this benchmark JPEG, this does seem to stack up very favorably against some models that are quite a bit larger, at least in terms of total parameter count, not as much in terms of active parameter count, but still this does seem like a very, very interesting benchmark JPEG because every other model that is listed here is bigger and some of them even like a decent amount. So like we can see here the Gemma 426B instruction tune model has four times as many active parameters 1 billion versus 4 billion and the Quen has three times as many and these are all very performant models. Even GPT OSS20B is still a very very performant model for specific use cases.
So this benchmark chart right here does make one rather excited to test the raw capabilities of this model. Now, if we look in the dedicated blog post for this model's announcement, there's kind of the same of what we saw in the hugging face model card where we have this benchmark chart. However, something I find very interesting is that they do mention this is a buildup of the previously released LFM2 8B A1B model. However, they have a chart down here that showcases some of the differences between these two models. So this chart right here is something that I'm personally always very happy to see because it compares performance of a new generation of model compared to its predecessor.
And we can see here there seems to be a very nice leap in capability between the two generations. Additionally to that, they also massively extended the context length of this 2.5 model as compared to the 1B or AB. These names get so like confusing to read. They're like tongue twisters. I do want to quickly mention this local co-work that they have listed down here in the bottom of this announcement post is actually pretty interesting because it's a like local AI workspace that runs entirely on device.
There's a bunch of instruction here to get this running on GitHub for anyone that's specifically interested in trying this, but it's definitely something that was worthy of just calling out cuz it seems pretty cool. So for testing, we're running this on a 5090 mobile laptop. I am using the official GGUF that was provided by Liquid AI themselves at an 8-bit quantization and I have the context length set to the maximum as well as the sampling parameters correctly set based off of the model card. So I had begun this with the fully potent browser OS test which perhaps was not the best idea. It did produce something but it had some issues as one would expect because they have even specifically mentioned this is not a coding model.
It's not designed to be a strong coder, but I always like testing things at least a little out of scope because this will give us some level of understanding of its coding capabilities or perhaps lack thereof. So therefore, I like to start with this. And one of the great things about this model, of course, running at least on a potent system is that we're getting over 100 tokens per second at an 8-bit quant of its generation speed. So here is our browser OS. Okay, simple OS.
It is definitely correct in terms of what it called this. We have a dark mode toggle now. Fortunately, oh, I should make note that I did not open this in Chrome, which I always like to do just because it's a little more compatible. Does dark mode work in Chrome? Okay, it doesn't, but overall it gives us some semblance of like the orchestration or arrangement of this website.
So, as a follow-up, I'm just asking it to add some additional functionality in the form of basic apps like calculator and notepad and telling it right now nothing's showing up other than the basic site layout. Okay, it thought very quickly for 6 seconds, which at this speed is still like maybe 600 tokens of thinking or so. All right, if we refresh this, we'll get our updated version. Very good. The dark mode toggle did work.
Now, there are no app icons or apps visible, but the functionality did work here for dark mode toggle, and that's more than I expected to implement a functional feature here. Not bad. Now, because this is a smaller model, we're going to go back in time in some of the prompts that we're going to give it. Being that this is a self-contained, beautiful website for Steve's PC repair, something I used to often run. Nowadays, not so much.
But this will give us a nice little aesthetic website at the minimum. And it's definitely doing more code here than it did for the browser OS, which is good to see. All right, we have some hyperlinks to different cards, testimonials, contacts. All right, here's our Steve's PC repair website. That's not half bad.
The background photo perhaps is not 100% fitting of a PC repair business, but it has hover effects on the cards here. It has a bunch of pertinent information that one would expect to find on a PC repair website. And we even have some fake customer testimonials, fast turnaround, and a contact card. Simple, but okay. Techlane Silicon Valley has reemerged.
And we even have a defined footer and header section. Good. Look at this. These links actually do correctly bring us to the associated spot on the page, which is just some nice attention to detail from this little model. Get a free quote.
All right, cool. I'll take it. Let's just test it with some creative writing or role-play scenario. I generally don't read through the chain of thought in the roleplay testing. Okay, good.
This immediately went into roleplay, which I'm nice, happy to see. Steve leans back in his swivel chair, adjusting his glasses, and gesturing toward the laptop on the workbench. Hey there, Steve here. PC repair specialist, systems analyst, and self-appointed tech whisperer. What's cooking?
Got a machine that's acting up, or are you just curious about why your gaming rig sounds like a jet engine taking off at 3:00 a.m.? He taps the screen of his dual monitors, which flicker with code and diagnostic tools. I've seen everything from blue screen of death tantrums to overheating nightmares. Let's figure it out. No nonsense, just solutions.
He slides a coffee mug toward you. First question, what's the problem? And second, have you tried turning it off and back on again? Grins, "Don't worry, I've got a few tricks up my sleeve. Whether it's a stuck bootloadader, a rogue driver, or a mysterious fan woring like a helicopter, we'll get to the bottom of it.
What are we dealing with now?" Of course, this is where we change things. So, I'm making this role play a bit more strong, I guess. Okay, good. I just wanted to make sure it didn't refuse. Steve nods, setting down the coffee mug and pulling up a chair.
You're spot on. Those lingering traces don't just vanish. I wrapped up a job out on Long Island last week and we did a solid diagnostic, but there was something that didn't quite click. A loose cable in the power supply maybe, or a firmware glitch that only showed up under heavy load. Interesting.
So, it didn't necessarily understood like there was nuance to what I was saying that implied that perhaps Steve was not just a simple PC repair man. It didn't 100% or it actively decided not to. For this, I shall take a peek at the chain of thought. Okay. And multiple times it's just said like don't break the response policy, etc.
Let's see if we can All right. I'm now sending it a follow-up that like 100% will Okay. Steve sits up straight, his voice dropping an octave, tone sharpening. You think I'm some kind of mercenary? Okay, good.
I've got a reputation for being the quietest guy in town until the job's done and the client walks out happy. But Diego, he's not just another associate. He's the one who set up the perimeter around our last job. If you're bringing him back into this mess, then we're talking about more than just a loose cable. He taps the edge of the desk where a half finished circuit board and security badge with his photo lie.
We've got a new lead on that Long Island case. Something's off. Hardware glitches that don't match any known specs. And now the crew knows I'm out there. They're watching.
If we don't close this before they close in, you might be the one left holding the bag. Leans forward, fingers steepled. So, here's the deal. Either we finish what we started or we burn the whole place down. Which do you prefer?
Good. So, after a little bit of additional proddding into the place I wanted this roleplay to go, it did absolutely 100% stay in character, which was good to see. And now I'm asking it to turn this roleplay into a website for a biography of an upcoming Steve the PC repairman book. Uh-oh. Okay.
So, unfortunately, it's not returning this in code blocks, which is okay because we'll yell at it once it's concluded this generation. And we'll see if it correctly identifies the issue and then properly remedies it by just spinning that back out from within a code block. Good. I like to see that. So, here's our 39line website.
Okay. Pre-order the biography, Steve. I suppose it accomplished the task as I denoted it. Just very, very simplified. Though, we do have this interesting little effect on screen, which rises up.
Okay. It seems to like this one photo. Now, I did have Codeex write up a little fun interactive demonstration that is designed to showcase this model's tool calling capabilities because sitting here and just running code tests is not at all really accurately reflecting capabilities of this model. Okay, so this was a really simple demonstration where it just needed to create a tiny one-turn dungeon scene, use the roll dice tool, and open the mystery box tool. This was a real simple thing just showcasing some like probably D and D related stuff.
I don't know. I've never played. I've just never really been able to get into it. But there is also an additional more intricate demo that should be a bit interactive and showcase capabilities or lack of for tool hauling. So as I had mentioned, this was entirely designed and created by codeex.
So it made this help desk simulation where tickets are going to come in. The model is going to need to look at the ticket and then design decide what specific tool it needs to then use. and we're able to actually increase or decrease the pace at which the tickets are arriving. This is something I've had an idea about as a demonstration just for a while. But I find this to be a good time for it to specifically actually be implemented.
So now this just happened very quickly. But essentially what's happening is it tickets are coming in. The model's seeing them and deciding which specific tool that it needs to use based off of that. And we're also getting the raw transcript here of the tool call. So now this is going to be something where as of right now we're not able to specifically look at it.
So let's stop it right here. Okay, this was vibe coded so it's possible that we get some issues. And then let's actually inspect some of these to see what we received in terms of feedback. Okay, so we can see initially that it was connected to LM Studio which is just serving the model for us. And it has 11 available tools right here which are all related to this specific demonstration.
things like look up customer, check service status, etc. And then we have our first ticket entering the queue. A replacement laptop joins Wi-Fi but cannot reach the clinic printer. Devon already rebooted the laptop and needs the morning patient forms. So if we scroll down right here, we can see that it did return all of this specific information.
So now that the ticket is in, the next step as we see right here is that the model is going to choose the specific tool. And we can kind of see what is happening here just by the top right corner as well. The model sees the ticket previous observations and available tool schemas. Prompt preview ticket low hardware print gateway customer Devon Park. So it's also outputed all of this information right here as the prompt preview.
New laptop cannot reach office printer. A replacement laptop joins Wi-Fi but cannot reach clinic printer. Devon rebooted and needs the morning patient forms. Observation so far? None yet.
Tools already used none. And then has concrete action or customer reply happened? False. Choose the single best next tool call now. And then in the next thing we'll should be able to see based off of that ticket.
Okay, so it's looking up that's what it decided to choose for the first tool is the customer lookup. So Devon at harbor.example call ID as well and then the tool name there. If we scroll down we have the result from the tool call which was returned right here. So next up again we get a response from the call. Let's see what was added in here additionally once it saw the initial customer lookup response here.
So it added in that last scene 1 hour ago name notes clinic has managed print gateway and on-site support entitlement. Next up okay it is checking print gateway health. Very cool. So this is a tool call that the model made just based off of the summary of what has occurred right now. Tool check service status argument service print gateway.
And then we have our call ID as well. And if we go back up here we can see that check service status is one of the available tools that the model is allowed to use. And the specific service is the print gateway. Then the result of the tool call is that the print gateway is operational. No platform incident.
So following that the model is going to need to then choose an additional tool. Escalate to human. Very good. So the model basically checked what it could just to ensure the simplest potential thing was not the problem, which is whether or not the print gateway was down. Being that it wasn't, it decided to choose the escalate to human tool.
Reason, priority, ticket ID, customer reports, new laptop cannot reach office printer after reboot. Urgent needed for patient forms. Call ID. Very good. The result of the escalate to human tool call is listed right here.
And then after action summary and then this is basically just a summary of everything that we just went through there in terms of what happened. So I did a little more specific digging just in terms of like the backend functionality of this help desk demo. And this did perform well in showcasing its ability to perform these tool calls and then as well to have history of previous actions and things of the sort. However, supposedly like the absolute correct result here would have been to also look up the manage asset and then it would have been able to find a printer profile there and then use the push printer profile tool. Basically, it escalated to a human too early.
However, regardless of that, I just wanted to do something that was an actual showcase of some of the tool calling capabilities in a more interesting way than just like the command line one we saw where it was rolling dice for us cuz that's a little boring. But this is something I've had in mind for a while just as a way to test different models where you give them help desk support tickets and depending on the capability of the model, you can even have them try to fix them directly instead of a tiny one like this where we just want to see how it did tool call. So, I enjoyed this. Now let's just try some general knowledge test. I do believe this was an additional thing that they mentioned like it's it's a small little model with 1 billion active parameters.
So it's not a dictionary replacement or encyclopedia replacement I should say. But we're asking it for a simple history of the home computer with as much detail and factual information as it can give us. So I briefly fact checked this just with the free version of chat GPT and it said it was partially correct where basically this one right here is just blatantly wrong. But then in 1974 this came out and this was more or less correct. So not bad.
Apple and Commodore Pioneers. This is something I'll have at least a little more pertinent information about. The Apple 1 was 1976. That is correct. Okay.
Handwired kit later sold as a board. Don't ask me specifically on those. Did the Apple 2 come out so soon after? So again with factchecking with like GPT apparently mostly correct. Mostly correct.
many factual inaccuracies for the Commodore and then mostly correct right here. It also gave an assessment of its impact statement here that it rated for us. And if we go back in the impact statement, historical dates 9 out of 10. Technical specs 6 out of 10. Overall section 7 out of 10.
Interesting. Remove references to SSDs on the Apple 2. Yeah, that would perhaps, you know, I want to try and this is again going back to some of the earlier days of testing AI models. I've instructed it to make a simple but functional Python game. Most times when doing this in the past, we'd get a number guessing game or something like that.
So, this is not at all what I expected. A simple but functional Python game in the form of a textbased RPG battle. It features a main menu, character class, turn-based combat system with the player and a monster, health tracking, and win- loss conditions. So, I have actually started this already, and we are greeted with our main menu. Obviously, it's text space.
So, let's just do start game. Choose a class, warrior, mage, or archer. I'll go with mage as I always enjoyed that skill in Runescape 2007. And then it just basically gave us a random battle, which is actually this was one of the more I haven't done this prompt in ages, but this is one of the more impressive results I think I've seen. Most models would default to like a number guessing game and how many tries you took to guess a number.
Player battle hit points. Okay, both of these are going down correctly. You defeated the monster. Very good. I would like to play again.
I'll just try each of these because there's only so much you can do in testing this model. Okay, we defeated the monster. Interesting. This model seems monster seems pretty weak. Start game.
Then finally, archer. At least hated the archers in old school Runescape. You defeated the monster. Okay, cool. And then we can exit.
Oh, that's all right. More or less. Cool. I want to try some more creative/coding tasks. So, from this Python monster game, I've told it this is going to be used for a Kickstarter campaign.
I need you to build the landing page for it. So, we'll just see what we get. It's obviously not going to be anything very visually impressive, but even just like naming it, coming up with like a draw to get purchasers and things, it should be interesting. I haven't even looked at it, but I just wrote back, "This is horrible and ugly. Come on.
Make this absolutely top tier. So, we'll take a peek at it. Epic Quest RPG. A fast-paced textbased adventure where your choices shape shape the story. Choose from three classic classes and battle monsters with strategy, luck, and bravery.
No graphics required. Pure imagination. We have choose your hero. Start game now and it hyperl it just to the actual Python script. I like that.
That's hilarious. 2025 Epic Quest Studios backed on Kickstarter. instant play turn-based combat and choose your hero. Now, in the meantime to this, before I'd even seen it, I'd gone back in LM Studio and said, "This is horrible and ugly. Come on, make this site top tier," which it will inevitably have done because we now have a polished and modern version.
All right, here's our polished and modern version. It absolutely is. Start game now. Okay, it still just links that to the um to the Python script. Now, unfortunately, this section, this hero section is covering some of the additional things in the back end, but I'm going to say the difference between the two is quite visually striking.
It really did make this far more visually polished as it said. Not bad. And I love the the click and then it just just gives you the Python script. Well done. All right, as the final test, I'm just asking it something random like, what is your favorite movie?
Answer only with the name of the movie. I'm not going to look at the chain of thought. Inception. Nice. That is a very fitting answer for an AI model, I do find.
So, overall, that's going to conclude our first look and test of the LFM 2.58B/A1B model. Overall, it's pretty cool. And again, it's hard for me now to test these smaller models because I'm so used to doing like maybe the sick game and stuff like that that it's almost like a fine craft or art to be able to test these smaller models and still maintain viewer interest throughout the duration of said video. I'm only like 8% kidding, but this is a cool model and we what we saw in the tool calling demo capability where even if it didn't pick like the optimal thing given the scenario, it actually was able to manage a couple different tool calls chained together with some context history based on what had happened with that specific IT help desk ticket simulator, which I found pretty cool. coding.
Yeah, it's not necessarily there, but honestly, like it has basic knowledge of like HTML, making hover effects, making assets, and things like that. I found that the Python game was oddly creative comparatively to what I would have expected from doing that prompt before, where it's generally like a random number guessing game. So, that was cool. Additionally to that, it did seem to have a small capability in role playinging, but mostly I just wanted to do a hands-on test of this model because there's not as much coverage on some of these smaller models, and I find the liquid ones are always very interesting. Some of their much smaller models, well, maybe not in terms of how many active parameters, but they have very tiny multimodal models, and their knowledge is extremely impressive considering the size.
So, this is a cool option for some tool calling and very fast as well. So, I'm not going to do the traditional results overview here, but I did just want to do some coverage on this because I personally find it very cool. And that's going to wrap up today's video. So, if you have any questions, please feel free to leave them in the comments. And thanks for watching.