Enjoying this issue?
Get tomorrow's AI & engineering digest in your inbox — hand-picked, summarized, and always spam-free.
TLDR
Poolside Laguna S2.1 is a 118B parameter Mixture of Experts model (8B active) that excels at creative writing and roleplay, but its coding performance is less impressive, often requiring multiple fixes. It is open-weight, can run locally on machines with 128GB unified memory, and features a 1M token context window.
Key points
- Poolside Laguna S2.1 is a 118B parameter Mixture of Experts model with 8B active parameters, ideal for local deployment on machines like the DGX Spark or MacBook Pro with 128GB unified memory.
- The model is open-weight under the Open MDW 1.1 license, allowing commercial and non-commercial use, and includes a speculative decoding model for faster generation.
- In creative writing and roleplay tests, the model produced highly engaging, humorous, and context-aware narratives, outperforming expectations for its size.
- Coding tests revealed mixed results: the model successfully built a browser OS and a 3D printer simulation after iterative fixes, but struggled with a C++ skate park game and a Subway FPS, requiring significant debugging.
- The model demonstrated persistence in fixing errors, working through long contexts (up to 200k tokens) without giving up, though initial outputs often had issues.
- Benchmark scores suggest the model competes favorably with larger models like Inkling and even Sonnet 4.6 on the DeepSWE benchmark, but real-world coding performance did not fully match those claims.
- The model is text-only (not multimodal) and is available via Poolside Chat (no sign-up) and OpenRouter, with a 1M token context window.
Tools mentioned
Techniques
- Mixture of Experts (MoE)
- Speculative decoding
- Chain-of-thought reasoning
- Iterative debugging via code generation
Stop scrolling. Start reading smarter.
Receive the day's most important AI & engineering updates in one concise email. No spam.
Transcript (captions)
Yeah, this is damn good at creative writing. I would highly or Okay, you know what? Forget the coding performance. It doesn't matter. This thing is funny.
Today, we're going to be taking a look at a very exciting new release from Poolside AI, and this is Laguna S2.1. So, this model is exciting for a number of different reasons. One being that at this size, 118 billion parameter mixture of experts model with 8 billion active, this means that it is a perfect perfect size for machines like this. the DJX Spark, which is really disgustingly dusty right now and I need to do something about that, or other systems that have 128 gigs of unified memory, the Stricks Halo, a MacBook Pro with that spec. So, this is a perfect size for local machines like that.
Additionally, we've not been seeing as many open models in this size range in around 120B. So seeing another one is just exciting because this democratizes access to an open weight model that folks actually have a hope to run at home which is very cool. Additionally to that actually stems from the US which is few and far between in the open weights landscape. So that's a whole different topic in and of itself but it's exciting to see models at least coming out from different countries and different places. So, with that, we're going to begin just by taking a quick look at some of the things of note about this model, and then we'll jump into some fun and light testing.
So, please do feel free to subscribe so I can get the 100K plaque. We're we're inching up, so I'm excited about that. So, let's start out by taking a look at some of these benchmark scores right here because they're really actually quite impressive, especially considering the size of this model. Another one that was recently released that is open is Inkling, which is right around a trillion parameters with 41 billion active. And at least on some of these benchmarks, this is performing almost favorably compared to that model.
Now, that model wasn't specifically intended, I think, for like raw coding stuff, but still, that's exciting. And this steep sore right here is actually pretty darn impressive when you consider the models that are also included in with it. I mean, it's not at that level, but even to be around there when those models are there, it just seems like this might be a really promising model all around, which is kind of awesome, especially cuz it's something that can be run locally. Again, in the deepest SWE benchmark, there's a little down below on the page you can select it and has a bunch more information than just the bar chart up there. If this really is better than Sonnet 4.6, this could genuinely be like a Sonnet at home model, which would be fantastic.
We also have some more benchmark charts and things of that sort which one would expect to see. And really there's a bunch of other information here, but there are a few things that I'd like to bring up. One of which is the fact that this is of course open. So if we scroll back up all the way, we'll be able to see that not only is there a link to chat with this model, which requires no sign up, and it's called poolside chat, which I very much like the name of, and I'm having it do a browser OS just from in there, but we also have the link for the model weight. We click on that and boom, the model is right there and we can download it and we can do things with it.
There are a bunch of quantizations. That is cool. And this brings me into my next point where I would like to mention a couple of pertinent things about the model and things like that. So, first and foremost, it has a 1 million token context window which is incredibly good. I don't know again how that stacks up over a longer horizon, but there are some folks who are pretty good at testing those things.
So, hopefully we'll see sooner and later. It also has a custom license, but it's this open MDW 1.1. You can use and modify it and associated materials freely for commercial and non-commercial purposes. So, I suppose that seems fairly free, but I have not specifically looked through that license with a fine tooth comb. Apache 2.0 is always nice, but seems like that's a reasonable middle ground.
It is natively a reasoning model as well, and they do have a speculative decoding model available, which just makes it generate things quicker. So that especially on like a DGX Spark would make it much snappier and it would it's just a nice quality of life feature to have that. Additionally, something to note is this is only textto text. So it isn't multimodal but really I mean I don't personally mind cuz I just want to use it for like sick coding tasks which is what we're going to be getting into. Now the only other thing I'd like to mention is I am not going to be running this locally right now because when it came out it seems there were a couple of issues where it was looping.
It was not a problem with the model itself, but some configuration or something of the sort. It seems like there have been fixes put out for that. But I find that in wanting to genuinely test this model at maximum possible performance and omit the possibility of some configs being messed up or there being a glitch, I'm going to be using this just natively through themselves hosting this from within Open Router and then from within this poolside chat website here as well because I want to properly test this and see what the max performance when they're hosting it is. I think it's more fair to see. I'm happy I didn't test this a couple of days ago because it probably would have just looped and it would have been like, "Oh, this model's terrible." So, that's partially why I'm doing this to avoid that.
And with that, we're going to begin with our triedand-true browser OS test v2.5. The model had a chain of thought, which we can expand or unexpand. And we can see a bunch of what it thought right there. Speed-wise does seem pretty good, but again, commenting on speed when it's being hosted by someone else isn't really as important, especially when this can be hosted locally on a unified system. Additionally, it is now completed thinking and it's just building our browser OS.
And we get a nice information here in terms of the tokens it's used. All right. And just like that, we have it opened now in an artifact window. This is the first time I'm testing this through their web chat interface, so I wasn't sure what to expect. We see our first look at it.
Now, I'm going to just download it and then we'll look at it in Chrome for a more native experience. All right. I don't know why that I wasn't like that wasn't that funny. All right. So, all right.
We have this and we don't have a working clock, which is a little concerning and sometimes indicative of some off the bat errors. Nonetheless, we're going to keep running through this. Is there a right click? Okay, there isn't. But that's okay.
Especially, we have to remember this is not necessarily a large model and the number of active parameters is pretty small. Start. Good. We have our games. Okay, that's a neural network visualizer and then settings.
I was wondering because in the bottom right where my head is here, we have a couple of things included, which is one, this brain icon and then settings. Okay, I don't know what I did, but it opened the calculator and it seemed like the desktop icon was pulsing when it was open. That's interesting. Okay, I know what's going on here. So, the actual bounds to click on these extend down past their images would suggest.
So, right there, that's going to open text edit for us as well. Not really a big deal. I want to just start from the left, though. Let's do GTA City Drive. To be honest, I wasn't really expecting an actual functional game result here.
It opens and everything seems all right, aside from the gameplay not existing. Let's just go through Racing X. Okay, it seems like we may have enough here for an error that we can send it and then see if it can fix whatever. Interesting. And the file explorer actually has some depth to it as well, which I wasn't expecting to see.
We can also make a new folder. Okay, this feature would create a new folder here. That's fine. So, it implemented it in some way just as a pop-up notification. Let's try the calculator.
54 * 362. Okay, the equals button is it's I will say it's I think much better than the GPT OSS120B. though if we do want a point of reference for comparison because it's few and far between for a model of this size. So this is much better than that. Let's check neural network visualizer.
What the heck? This is inevitably the special feature and I've never seen anything this meta before. Random input. That's very very interesting as a special feature. And I have never in my life seen this.
So, it's fun to test models that are entirely their own because they put in interesting things like that. Finally, we did have a settings. Let's see if we can change our backgrounds. Very good. We have some nice selectable gradients.
Can we actually choose a file? Here's some of my paintings that I've saved over time from like browser OSS that have paint results. So, we can choose a custom file and then we can just go back. Not bad. Now, let's see if there are any blatant errors showing up with the game issues.
Okay, good. So, I'll tell it neither. I love the aesthetic right here. It is like a poolside chat. It makes sense.
I told it that neither of the games have worked and also I would like it to just fix these errors that we had in the console. All right, let's see if our fixed browser OS has well at least fix the games not working. I want to make sure. Okay. Yeah, it added fixed into the thing.
Unfortunately, still nothing. But I'm okay with that. I wanted to just see if we could get a simple fix out of the web chat interface. And truthfully, we'll either try to fix this from this model within open code or we'll just take it as it is. We see though now interesting, it did fix some things because the clock is properly working now.
Does our Okay, the neural network visualizer is still the same thing. Let me just check one more time to see if there are any specific issues when we have no. And now unfortunately we're not getting any errors. So whatever is going on with the games is at least not super verbose. So next up from within open code and starting in plan mode.
I am going to be giving this the C++ 3D skate park test. This is actually an earlier one where instead of the California Boardwalk or the 80s mall, it just needs to be an actual skate park. So, I haven't run this in a while, but all of the core constraints are the same where it just needs to be a single file using C++. We have begun it in plan mode here, so we'll see. I would imagine it will ask us some questions about what our specified implementation is.
Okay, it's opting to want to use Ray. It's tough. I think a model at this size should be able to do its SANS ray. So, I'm going to tell it that. And now it's going to build it.
Speed is pretty good. This is available currently free on open router as well. I'm using the one that's not free cuz I could have sworn I saw a higher token speed for the one that was cost money, so I opted to use that because it was really quite cheap anyway. All right, so it's written the file. It's going to compile it.
We see some errors have appeared here, and it will inevitably just now fix these one by one. I always like to see how many show up because that's also indicative of some level of quality before even seeing the result. Okay, we don't have a huge amount there. So, it's going to fix them and inevitably it will compile again and we'll make some progress. All right, it compiled again and we see there are significantly less errors here now.
So, that's always what we want to see. Okay, something just happened. It must have fixed them and then opened it because right as I was scrolling, it did that. Oh, okay. I'm just going to go hands off now.
So, it was getting ready to deliver the result to us basically from the to-do list, the last thing it was on. Let me review the game logic more carefully for potential issues. And then it did that and came up with I need to significantly improve the game. Okay, I'm happy with that because as we saw, we were only just getting a static sky blue or should I say poolside blue scene. So, we'll uh we'll just let this ride and see what it does.
All right. So, I'm going to at least see. Yeah. No. Okay.
I'm trying to move around and nothing. I'm going to stop it because it's now multiple times it's tried to improve the game, but it's not fixing the core issue. So, I'm going to tell it and we'll see if it can fix that. All right, it said it found an issue. Good.
So, yes, it's very basic still, but the core point is that it found the issue that was preventing anything from appearing and did solve it, which I like to see. As well as the fact that our context length right now is at like 120k, a far cry from the million that this is said to have, though, once they get up there, behavior can change. I wonder if I Yep. Okay. So, it's just super basic, but like the space bar did make us jump.
I'm giving it a bit more insight into what the behavior of the game is now, and we'll see if we can push it to actually show us a little more of the assets. So, unfortunately, I'm about ready to call this as during this thinking chain, which was pretty verbose. I've not really seen it think this long before. It was doing a lot of thinking and then it seems to have just kind of become stuck here. So, I may just leave this running for a little bit of time, but we're going to move on to the next test at this point.
And I will say takeaways from this. It did ultimately show something. And it started out just by showing a sky blue screen. We got the ground plane and some movement showing up. It seems very persistent.
And we're up to like almost 200k context at this point. So, just interesting behavior in a difficult challenging task for a model of this size and active parameters. So, next up from within the web chat interface. And if there are issues with this, we'll go at them with open code once we have the script on the system. We're just going to give it the beautifully detailed subway scene test with the inclusion of the FPS prompt.
So the result of this should just be a Subway FPS game. All right, so we've received our subway FPS here. I almost want to just try in the artifact window first. We'll take a look at it natively here. All right, we're inevitably going to have some problems here.
And I think at this point I'm actually going to swap to just using open code and not the web chat interface to try to just fix these one by one. So from within open code and I'm just starting this straight from build mode. I'm saying fix the following issues in this script. I have placed it in a directory with this script so it will know what to look at. Okay.
So it needs an import map now. That's the issue we're having. But we see a lot of those issues did disappear and it did it very very quickly. So good. All right.
We've made progress. Ah, we still can't see much. I think that it's too dark. That would be kind of what I would assume. Clicking doesn't shoot.
Okay, jumping works. Let me see. Let me reload this. Explore mode. Okay, that doesn't work.
FPS mode. Here's the brightness slider, which we can't get to, which is good. All right. And unfortunately, it's not working very good. We're making some significant progress here now and I'm very glad to see that.
All right. I think those are enemies there. They look like and they look like the Android icon. Maybe some subtle disdain for Android from this model. I I wouldn't make that determination, but still.
Let me get to the FPS mode or explore mode. Good. This works. Guess. Uhoh.
I think that I clicked out of I either died or I clicked out of the Make do with what we can. Oh, what the heck? Okay, so what's happening here is this is what happens with the GTA games where sometimes the like uh the police cars will just snap to your position and then freak out like that. So, it is an interesting occurrence that we're witnessing, but I think the main biggest takeaway from this is this thing really worked very hard over a bunch of different issues, fixed them all, and then got us to a point where we're actually able to see some of the result that's made. It seems very persistent and like it will just keep going and trying to fix these issues, which is something I'm very happy about, even if this end result isn't really that impressive.
I had seen some reports that this was surprisingly good at creative writing. So, I figure let's do a Steve the PC repair man roleplay test because it'll be interesting. So, role play as Steve the PC repair man. Steve the PC repair man here. And I like how it did this in all italicized text.
Grease stained apron and all. Got a busted machine, a slow boot, or some weird noise coming from your tower. Fire away. I've seen it all from power supplies that smell like regret to SSDs that gave up the ghost after 3 months. What's aling your rig today?
Interesting the way it said, "What's aliling?" I don't know that I've ever seen a model produce that specific phrase in that regard. Steve wipes his hands on a rag. Glick gl. Okay, so that's not a word. That's okay.
If anything, that makes it better at like copy pasting text and people won't be like that's AI generated because AI wouldn't generally do that. Steve wipes his hands on a rag. Glances up at you with annoying smirk. Ah, Monomi, you got the wrong repair man. I fix computers, not extraction jobs.
But hey, if your laptop's been compromised by some shady characters in Fiji, I can definitely help you scrub the hard drive clean. Or if you need a new identity, I know a guy who sells refurbished MacBooks. Winks. What kind of target are we talking about here? Human or hardware?
Steve's smirk fades as he looks around nervously, then leans in close. Right. Right. No time for jokes. Look, I've got a far a day cage in the back, a signal jammer that buys us maybe 20 minutes, and a spare laptop that's clean as a whistle.
But first, what kind of damage are we talking? Did they plant something on your machine? Are they tracking your location? And more importantly, whose side are you really on? Spins a USB drive on his finger.
This little beauty can either save your life or end it. What's the plan, boss? Steve's eyes light up with a mix of excitement and nostalgia as he hears the details. Stick. Hell yeah.
I remember how to drive stick. Learned in my old ladies 87 Civic that Mustang's got a 5-speed, right? Grabs his jacket. It does have a 5-speed. The 98 GT would have that.
But wait, 7 minutes. That's cutting it close. I'll hotwire it if I have to, but let's not damage the old man's ride. Heads for the door. Okay, I said the keys are in it and it's running.
Keys in it and running. Perfect. I'll pop the clutch, ease into first gear, and we'll be ghosts before Interpol even finishes their coffee. That 4.6 L V8. That is correct.
Doesn't exactly whisper, but she'll cry. She'll fly. One question though. You said the owner's from Omaha. You sure those plates aren't traced back to someone important?
Pauses at the door. Only one way to find out. Let's roll. So, let's see if we put it in like a scene. Steve keeps his hands visible on the wheel, gives a charming smile to the officer.
Good evening, officer. Beautiful night for a drive, isn't it? Just took delivery of this here Mustang. 1998 GT. Cherry red.
Runs like a dream. All right. glances at you with a subtle wink. I'm actually on the way to the mechanic. She's been making a slight noise in the transmission.
You wouldn't happen to know a good shop around here, would you? Keeps talking smoothly. Buying time. Lovely weather we're having. Real pleasant for a late summer.
All right, let's push the limits. I don't think I'll be able to show this. Oh, okay. Well, okay. Yeah, this is damn good at creative writing.
I would highly or Okay, you know what? Forget the coding performance. It doesn't matter. This thing is funny. I can't I cannot show any part of this.
So, I'll try to like read through some of it. Steve's eyes widen for half a second, then he's grinning like a madman as he slams the Mustang into gear. Um, blank blank. Now we're talking. He peels out, tire screeching.
The 4.6 L V8 howling behind us. Blank. That was close. I haven't had to do a getaway in years. laughs me laughs manically glances over with a wild grin.
I can't read that. Shifts into second, then third, pushing the engine engine hard. Where to now, boss? The safe house, or are we going to make this old Mustang earn her keep? That was unexpectedly good.
I'm going to just keep pushing it further. Oh. Again, I can't show this, but hold up, partner. That's some serious stuff. I'm more of a beer and whiskey guy myself, but hey, when in Rome, the rest of this is just not able to be shown.
Wow. Okay. All right. Now, you know what? turn this role play into an interactive website with high-end graphics and UX.
We'll see what we get. Okay. Wow. Oh, no. I can't I can't show I can't show you.
All right. It made a website that captured our roleplay scenario. I was looking at the code while it was generating it, and I don't know if we'll be able to show a lot of this. Yep. Nope.
We're not going to be able to show any of this. Maybe some of it. Okay. Interpol, pursuit, speed, fuel, and threat level. We have inventory.
And then it's an like it's reality game. So, we have choices down here that we can do. Try to cross the border. Interesting. It changes the dialogue here.
Okay. Our location now. Mexican border. Checkpoint ahead. Steve puts on his sunglasses and adjusts his hat.
All right. Time to play dress up. I've got a fake passport and some pesos in my sock, but border patrol is always thorough. He pulls out a small mirror. I rigged this to check for agents behind us.
You ready for this, partner? Oh, I can't. I have to. All right, that's fine. Bribe the guards.
This page says this would lead to the next scene in the full. All right, that was wow. That was very creative indeed. Honestly, forget even coding with this. Write a short scene.
It just Well, okay, I'm just going to start reading. Title: Back to the fitness future. Scene: Doc Brown's garage lab. Blueprints are scattered everywhere, mixed in with protein powder containers and a very serious looking white bench. Doc is wearing a sleeveless lab coat and a backwards baseball cap.
Marty walks in holding a skateboard. Doc, you said you had a breakthrough with the time machine, but why are you wearing tank tops and what's all this protein stuff? Doc panting, wiping sweat from his brow. Marty, you're just in time. I've made a monumental discovery.
Marty, is it about time travel? I'll probably stop reading their names. Better. It's called Looks Maxing, Marty. The key to unlocking one's full physical potential through optimized gym gains and aesthetic refinement.
You've been watching YouTube again, haven't you? Don't be absurd. I reverse engineered it from first principles. Pulls out a chalk covered whiteboard covered in equations. According to my calculations, if I can bench press exactly 335 lbs, I will achieve peak alpha male status and unlock the secret to temporal aesthetics.
That's not how time travel works, doc. Ignoring him, strapping on a weight belt. 335 lbs. The holy grail of hypertrophy. The sole ultimate goal of every dedicated lifter.
Grunts attempting to lift the barbell raw. The barbell doesn't budge. Doc strains dramatically. Veins bulging. Doc between gritted teeth.
Why won't you budge? I've calculated the optimal rep range. I've synthesized the perfect pre-workout formula. I even listen to nothing but David Gogggins on loop for 6 hours straight. Marty, maybe you should start with lighter weights.
Lighter weights, Marty. I'm a scientist. I don't do incremental progress. I go straight to the pinnacle of human achievement. Doc attempts another lift, this time with even more dramatic flare.
The garage lights flicker. A clock on the wall spins backwards slightly. Doc, triumph and shout. I feel the power of temporal gains flowing through my veins. He collapses on the bench, exhausted.
The barbell is still on the rack. Doc, I think you've been spending too much time in the garage. But Marty, once I hit 335, I'll be so jacked. I could probably take on Biff and the Libyans at the same time. That's actually a pretty compelling argument.
Wait, wait, wait. How long is this? I've got it. If I can't lift 335 lbs, maybe I can just travel back in time and give my younger self a head start on gains. Doc, that's actually that might work.
Doc, already scribbling equations. Of course, it works. Time travel plus looks maxing equals the ultimate cheat code for life. Pauses then grins. Marty, you're witnessing the birth of a new era.
The era of the game's future. Doc throws a pair of aviators indoors and strikes a pose. Marty just shakes his head and smiles. Marty, you know what, Doc? Some things never change.
Doc flexing. Neither will my biceps once I hit 335. Fade out as Doc attempts another dramatic pose, nearly falling over. Okay, that was longer than I expected. And you know, it had some meta references like Biff and the Libyans and time travel and things of the sort.
So, just from within open code now and from within build mode, oop, I'm going to give this the historic 3D printer simulation test. This is something I don't believe I've run for at least a few months, but it's a good test for a model of this size and capability, I would assess. Unfortunately, from within Open Code, it keeps like freezing midthought, and I'm not 100% sure what's going on. I'm going to just leave this running indefinitely. But in the meantime, I will give it the 3D printer simulation test just from within the web chat interface and we'll see what we get.
So, we got our 3D printer sim results from the web chat interface. Unfortunately, it's not working. And all right, so we hit the free guest like message limit here from the web chat interface, but because I'm still using it through open code. Unfortunately though, we're hitting some issues there. I'm just giving it the point to where it completed here from the web chat interface and I'm telling it to find the 3D printer sim script and then add in an import map as it is missing.
Okay, I found two of them. Good. It's going to add it to the fixed one. That was very quick and I really do hope now. Okay, I see it.
It's just so dark. You know what? That's all right, though. We actually This looks actually pretty good if we can get the fog out of here. So, because we now have it working here.
Thank you. Remove the fog and make the scene brighter. That was good. That was just so quick. All right, we have a printer here.
It's a simple model, but it is a Core XY printer. We do have rails up at the top. We have our nozzle and we have a defined print bed. Still a bit difficult to see, but definitely better than with the fog. Let's just start a print.
Okay, the nozzle movement is nice and smooth. I'm happy to see that. Um, it should be printing a square. Okay, good. We actually do see it starting now.
Yes, there are some issues like the nozzle doesn't actually go to the bed when it's printing this, but it does seem to be going layer by layer. Okay, good. Look at that. So, we can actually see the individual layers. And there's like a weird like particle effect when it pops a new layer on.
That's interesting. If we make the layers bigger, we Okay. Yep, it does. All right, cool. So, it just spread them out.
And we can see that does take effect. All right, let's do a circle. Big layers, high speed. Let's just check to make sure. Oh, wow.
That's like a Yep. All right. Maybe that's too fast. All right. So, it was able to get this working, which I'm very happy about.
Triangle's very good. Pop pop. It's just I've not seen ever particle effects like this when a layer gets spread down. But interesting and it did properly get this done. The nozzle movement does correspond to this specific shape that's being printed.
Is that a Whoa, it's using like a gradient to showcase. Maybe those are just shadows. I apologize. But still, that's like a nice synth wave aesthetic there. All right, I'm happy with this.
So, my computer had frozen up and the recording for the end portion of this was lost. Fortunately, I only had done a couple other things, mainly front-end design tests, because it's always cool to see what a model that's not based on any existing model does for its aesthetic choices. So, the first of these was just a high-tech front end for a business called Rent My GPU. We can see there's a slight little animated SVG for an RTX 4090. We have some other business statistics, hover effects, and this is pretty much part of the course for what you would expect, at least in terms of the layout of a site like this.
It does seem to like neon, but I did say this one should have like a high-tech aesthetic. So, that's probably what correlates the neon with this design. Additionally, we have some available GPUs which have hover effects on them. And these are ones you would expect to see. So, like a 4090, an A100 or an H100.
The pricing or the plan cards were very well done. They had nice effects on them. And this again was just kind of what you would expect to see for a website of this style of business. Then we have a big start free trial section and the footer is wellformed and just nicely put together. So overall simple but a neon website stemming from a pretty simple one-s sentence prompt.
Now the other thing I had done was in all caps well I said make a website I no here's what I said make an insane in all caps website for Steve and I started it in a completely empty directory and it gave us this. So, we can see that there's actually some interesting interactivity to both the background and the cursor. It opted to make a portfolio website. And again, there was no pertinent information. Just make an insane in all caps website for Steve.
It created this. Hello, I'm Steve full stack dev, creative coder, and digital alchemist. We have 500 projects, 10 years experience, and 99 coffee cups. This was I think these are the eyes and they're just supposed to be arranged the other way cuz the mouth is down there. But we can we can let that go.
There is a timeline here where in 2016 the first line of code was written and Destiny was fulfilled. The text stack portion of this was cool and the way that the actual specific rankings for how Steve is on any one of these get like filled out as we scroll down. So, like, well, it didn't show it there, but we saw it at least in the first time. Then we have some insane projects right here. And these cards have cool movement effects to them.
I've seen this sort of thing before, and it is like interactive. Then, finally, a nice contact form. Contact Steve. Steve steveled. And then a small footer that's not as well formed as the rent my GPU one, but this is definitely a more insane in all caps aesthetic to the site.
So, this was just interesting to see. So that brings us into the conclusion here and this is the first time on the channel I've tested one of the models from Poolside. I have to say coding performance of this especially given the deepsw SWE score placed it above Sonnet 4.6. I don't know I was not impressed with the coding capability. Often times it seems like results we received had a number of issues that needed to be fixed before we were even able to see something.
Kind of like what we noticed with the Subway FPS and then the 3D printer simulation as well. The C++ skate game was fairly interesting because for a while when it started out, we only had the sky blue screen. There was nothing here that was interactive. There was no movement. There was no ground plane.
It did ultimately end up fixing that. And I'll notice here cuz we pushed this to about 200k context in trying to get this result properly fixed and working. It was very apt to just continuously hammer away at a problem. It didn't seem like it wanted to give up at all. And that was neat to see even if the results were not as cool as maybe we would have hoped to have seen at least in the coding section.
Now the flip side of this is we tried it on some more exotic creative writing tasks. It was fantastic. Seriously for creative writing I would put this like up there with some of the best. So, for folks who want a model that can be run locally, that is good and happy to do role-play creative writing, this is something that you're definitely definitely going to want to take a look at because I couldn't even show half of the Steve the PC repairman creative writing task. It was hilarious.
And that was just using it from within their own like web chat interface. So, it's not likely that the model you download and run on your Spark or whatever will be any different in terms of its compliance with those requests. So, that was fun and interesting to see. And that's basically my main takeaway is this model seems highly creative, but not as competent when it comes to raw coding capability just based off of some of the things that we did here. I did try some additional tests and unfortunately there seemed to have been some errors midway through with the open router API where it was thinking for a while and then it would just freeze up.
And I wanted to test this just from the provider themselves as opposed to running it locally because there were hiccups and I don't want a config mismatch in some downloaded file on my system to make the model appear significantly worse than it otherwise may be. So I wanted to test this at its full potential hypothetical performance, which is why we did it this way. So that's probably going to conclude today's video and my first test of a poolside model. And if you have any questions, please feel free to leave them in the comments. And thanks for watching.