Cline Desktop Hands-On – Can OPEN Models Match Fable 5.1?

summarized

TLDR

DeepSeek V4.1 Flash replicated roughly 90% of a subway FPS game originally built by Claude Fable 5.1, using only the design markdown file Fable had generated — and did so through a $9.99/month Cline Pass subscription instead of a $100/month Fable plan. The Cline Desktop beta app made it possible to run this test across multiple open-weight models, and the result suggests that a cheap open model can closely match a frontier model's output when given a sufficiently detailed architectural specification.

Key points

DeepSeek V4.1 Flash achieved about 90% visual and gameplay replication of a Fable 5.1 game when given the same design markdown file.

Muse Spark 1.3 produced a less complete replication with missing sound, broken enemies, and no weapon switching.

The design markdown file for the subway FPS game was 2,600 lines and took Fable 5.1 two hours to generate.

Cline Desktop is a native beta app for the Cline coding agent, available on Mac and Windows with Linux support coming.

Cline Pass costs $9.99/month and provides access to multiple open-weight models including DeepSeek V4.1 Flash, Muse Spark 1.3, and GLM 5.3.

Tools mentioned

Techniques

  • Using design markdown files from a frontier model to instruct cheaper open-weight models
  • Agentic coding with Cline Desktop
  • Accessing multiple open-weight models through a single Cline Pass subscription
Transcript (captions)

0:00 because [laughter] Okay, so today we're going to be putting some openweight models to the test by giving them a design file that was authored by

0:12 Claude Fable 5.1 and seeing how well they can replicate the actual one:1 output that Fable had created when given the design markdown instructions. So, we're essentially giving them the

0:23 ingredients. And we're going to see how well they can prepare the dish compared to the model that both created the ingredients list and prepared the dish as well. Now, to do this, we are going

0:33 to be using a couple of different things. So, let's talk about that and then we'll jump into a quick overview of what exactly this design markdown file is, at least for the first task, and

0:42 what specific models we're going to be using to perform it. So for today's testing, we're going to be using the newly releasleased Klein for desktop app which is still in beta. However, this is

0:52 essentially the Klein coding agent as a native desktop experience or app as they say. Now what's cool about this is a couple of things. One is the source is available for this and there have been

1:03 some discussions recently about open- source versus closed source agentic coding tools and being able to actually look in and inspect and ensure it's not uploading the entirety of your file

1:13 somewhere is always nice to be able to have. Additionally to that, we're going to be using these models through Klein Pass, which is essentially a $10 a month subscription or $9.99 to be able to

1:23 access a bunch of Frontier openweight models. I use these all the time on the channel and I do have to be honest with you. I basically have separate accounts for every single one that is scrolling

1:33 through right here and it can become kind of cumbersome to have to deal with all of that. So a unified trustworthy place that you can access them all with is very very nice to have. So this is

1:42 how we're going to be testing our models today and we can see some of the ones that are available. Now truthfully I have not determined the exact ones we're going to be using to perform this test.

1:52 I have a general idea. Definitely probably GLM 53 flash, although the full 53 is also interesting. And then yeah, so I'm still on the fence about what specific ones we're going to use. But

2:03 before we get into it, there are a couple of additional things. I do want to say thank you to Klein for supporting the channel and allowing me to do this video and giving me some access to this.

2:12 Additionally to that, they do also have an introductory offer for this. So instead of $9.99, it is $1.99 and there will be a link in the description. Now really the next thing to do is start by

2:23 taking a look I suppose at the client desktop app which I do have downloaded right here. I am using Windows today as this is currently Mac and Windows only. However, they do mention Linux support

2:33 will be coming very soon. And I was just browsing through some of the marketplace for skills here adding in MCPS and things like that. And I was happy to see that Jev is already included. This has

2:43 been all of the rage very recently so you're very likely familiar with it. But I was happy to see that this seems to be pretty quickly updated because this just came out and became popular very

2:52 recently. Now to quickly look through a bit about this app, I am going to replay the introduction what you'll first see when you actually get this set up on your system. So you just see this and

3:01 the eyes do follow the cursor. So that's always nice to see. And you may be familiar with Klein as prior to this it was prominently like a VS code extension or something like that. So this is now

3:11 just a native experience and it moves out of the terminal also if you were familiar with using it that way. We also have the signin with client option which is what I am doing here being that I am

3:20 using this with client pass. I am going to skip the import sessions and things like that. As this is a fresh install on this system and clin is already set up for my account. So we're going to start

3:31 building and we can see we swap in and just select client pass and all the models that are currently available are listed right here that are included with that. Interestingly muse spark is as

3:40 well. So that's cool. And I do believe Meta had talked about open waiting that at some point. So hopefully that is something we'll see. They do also offer a few free models as well that change

3:50 variously depending on like what's hot things like solar pro. So that's pretty cool as well. That's a Frontier model from Korea. Laguna S2.1 and a bunch of other interesting models. So now for I

4:00 suppose the meat or what exactly we're going to be working with at least to begin here. This markdown file right here was created by Claude Fable 5.1. It took 2 hours when I first tested that

4:12 when it came out. We did a Subway FPS test in that video. It is live on the channel from I believe early September. I had run it using ultra code and I did not realize that it spent two hours

4:24 crafting a design markdown file and this is the culmination of that. So if we scroll down here, we're going to see that this is essentially a 2,600 line of text file denoting every single specific

4:37 thing that should go into this 3JS subway FPS game. Now, the reason that I actually had this idea was because I gave this to Quen 3.8 flash next at a Q4 quantization running locally on the

4:50 system behind me. It worked overnight and this was a couple of weeks ago and it really produced something that was probably like 80% of what Fable had produced once it went beyond this and

5:00 created the actual game. And it made me think this is actually pretty cool to see what you could get out of combining a Frontier model, using it for an agents.mmd or a design MD and then

5:13 instructing openweight models that are cheaper to use but still more or less capable up to a point comparatively and seeing if they can replicate it. So that's exactly what we're going to be

5:23 doing today. I have already placed this in a folder and really the thing I'm going to do once we decide what model to start with which I do think v4.1 flash is probably a prudent one though I am

5:33 also curious about muse spark because some folks were saying on max effort that's become pretty potent. So we'll have to see which one we choose but to start for now and we'll run it with a

5:43 couple. I want to go with v4.1 flash. I will put it on extra and I'm basically going to say build the design from this folder. And one of the things that I think will really just be the best way

5:55 to compare these is by looking at the actual source that was created by Fable 51. And then we'll get a good visual comparison of how a model does when given its instruction. A model that is

6:06 significantly cheaper to run in the case of V4.1 Flash than Fable 5.1. Our workflow is finished and I quickly want to just showcase some of what we saw. This is the Fable 5.1 result that we're

6:17 seeing right here. And there is some closed captioning going on so that's why we just saw that. But this is just to give ourselves a reference point of what specifically that design markdown file

6:26 had produced when it was done by Fable 5.1. So our next step now is to essentially go and see how Deepseek V4.1 Flash did with this. Keep in mind V4.1 Flash is a bit of a larger model than

6:40 the previous Flash. So it is more capable than you may expect if you are thinking of the V4 Flash model instead. So, this worked for quite a long time and I don't know that it's still like

6:50 100% polished, but I think it's more or less there being that I did take a nap and wake up and it was still working on some things. So, okay, we now have this all set. I had neglected to realize that

7:00 I needed to start the server myself. All right, we have the sound on right here. It said click to enable pointer lock. So, let's basically see how we did with this task when giving it to

7:15 Okay. So, um either the pointer is hyper sensitive or something's not 100% right here. However, I'm going to say just on first glance, aside from that, this is really almost like a pixel perfect

7:28 replication of it, this looks exactly like what we had just seen in the video, I'm actually a bit surprised here. Now, one of the issues I'm noticing is unfortunately there's something off with

7:39 the pointer like Yeah. So, um, is it the mouse? I don't think it's the mouse. Let me unplug this mouse I'm using. Yeah, it's okay. So, I think that we may perhaps have a bit

7:53 of troubleshooting to do. However, Okay. All right. Yeah, I think it's just like a hyper sensitive pointer. Oh, wow. And we have the right click actually allows us to aim just like that. We have

8:06 multiple different weapons which was also the case for the claw result. And then now the enemies are supposed to just start by coming down from the stairs. Keep in mind the mouse is still

8:15 very funky. So I'm having to like do some really weird things right here to keep it from going haywire. Now I do believe in the original the enemies came down from the stairs. I don't

8:25 necessarily see them coming right here. So I'm just telling it some of the issues that we're having and I'm saying fix these immediately. and it will hopefully immediately. All right, first

8:35 and foremost, the mouse movement is significantly better. Do we have sound? I don't actually know as of right now because unfortunately I can't shoot. All right, let's see. So,

8:48 in the original Hey, what the heck? I don't actually think the original had the ability to go up the stairs. And something else I'm going to note that I did want to verify is it actually did

9:01 draw Ashworth station which is what was named for this project or Ashworth Street. Maybe station street station would make more sense. I don't think the original had that. Now we're going to

9:11 notice here yes the enemies do seem to have some form of issue. I believe they don't actually have the proper tickets to get onto the subway track which would be keeping them there. Nonetheless, even

9:21 just the graffiti drawn here. This is very interesting. This really, I will say, forget about like the actual enemy logic and shooting right now in sound. I think the main takeaway right here at

9:31 least is the fact that this is really a proper proper graphical replication of the main result. Very interesting artifact there. I do believe there are some enemies over here as well. So, the

9:42 final thing I'm giving it is just the fact that the sound still isn't working. And also, when we click the left mouse button, we can't actually shoot any of the ammunition. All right.

9:50 Hypothetically, we now have some progress here. So, it's saying reload with Ctrl F5 and then shoot. So, let's see. Okay, we only have one error now. Good. Shooting works. I still don't hear

10:01 any sound. However, the fact is we are getting to shoot. Are there no bullet holes in the environment? Oh, there are. Look at that. They're clean, too. All right. All right.

10:11 So, I'm going to hide the dev console for now. We still don't have sound, but at least we know some of the enemies are definitely not very properly put together, but you know, that could be an

10:25 artistic decision and not necessarily a bug. And we can see, okay, zero left in wave 1. However, there were a bunch up there, though. This is the version where the train is supposed to deliver the

10:36 next enemy. So, all right, wave one cleared. So, now the train should come in. next train and we're getting a countdown timer. This is again where we'll be able to see how

10:47 well this managed to replicate the style. Okay, incoming train stand behind the yellow line. We do get the red blinking lights which were also present in the original. We just need to make

10:57 sure [clears throat] that is a very very very very similar replication of the train from the original. It even has the really nice interior. There was some smoke there and the way the doors open.

11:09 So overall the one thing that is kind of bad here I suppose I should say the two things the sound and then the enemies are kind of just as we can see here something is not quite right in the

11:20 geometrics of how these things are basically designed however we do have bullet holes we can't go on the train as with the original I know some folks had asked about that so still I find this

11:30 pretty interesting because we basically gave this a very very dense markdown file and it replicated all of the aesthetics of this pretty much like 95% of the way there. Seriously, the big

11:42 issues are sound, the enemies, and then we had to do a few more bug fixes, but really like the overall takeaway I think from right here, and they actually have similar behavior as well. The original,

11:52 these blue ones would actually run at you very quickly and attack you. So, I think the main takeaway right here, and I'm not going to try to get this perfectly 100% working because I wanted

12:01 to see what it realistically did. And I would say if we factor in that the sound didn't work and the enemies were a bit odd, overall I would say this is about 90% of what Fable 5.1 gave us, keeping

12:12 in mind this was all done with a $9.99 a month subscription. Whereas if this were to have been done with Fable, not even factoring in actually creating the game, just creating that markdown file, I

12:25 don't think you would have been able to do that on a $20 a month subscription to use Fable. I think you'd need like the $100 a month one. And even then, the limits are pretty low. So, this is

12:33 something that I've really been interested in doing where essentially giving these models that are significantly cheaper but pretty capable all things considered almost an

12:42 architectural diagram in the form of this design markdown file to see how well they can replicate what a big expensive closed source Frontier model would have made. And in this case, this

12:53 is a pretty darn good replication. Now, the thing I'd like to do next is try this with a model that I personally find to be a bit less capable, just in my personal experience, and that is Meta's

13:03 Muse Spark 1.3. So, we saw what Deepseek V4.1 Flash did, and that was a really fantastic replication of the Fable 5.1 result. Now, let's see what Muse Spark 1.3 can do. A model that I truthfully do

13:17 not believe will get us to that like 90% of the way there, though, I would be very pleased to be proven wrong. So, our next step is going to be to delete this session just because I'd like to have it

13:26 basically gone. So, we're starting fresh. I have selected Museep Spark 1.3 contributor. This is, of course, being used through the Clin Pass subscription. Keep in mind, Muse Spark 1.3 on the

13:38 contributor tier is very, very cheap. So, basically, you're going to get a ton of usage for this specific model. However, contributor means that you get it at that low cost in exchange for

13:48 allowing the data you send to it to be used to train and improve the model. Now, in the case of this where it would just basically produce a better Subway FPS, I'm okay with that because we could

13:58 all use better subway FPS's, I have also just basically gotten rid of every single file that was in this folder that DeepS had created. So, if we go right here, it is just stuck with the design

14:08 markdown file that we originally had. Now, the only thing I'd like to do is put this on extra. And we're now going to start this and see how Muse Spark 1.3 on its highest available setting here

14:19 does. Again, were I to take a guess on how this will perform, I have personally found Musepark 1.3 is not as strong of a coding model as some of the ones that we're going to be comparing it with. So,

14:30 keep that in mind. I would say it will probably get a 66% or so replication of the original, though, I'd be happy to be proven wrong. All right. Now, in 32 minutes, this has completed. This is

14:41 significantly quicker than what it took DeepSeek V4.1 Flash to produce. Now, that does not necessarily mean it's going to be less intricate. However, I would perhaps go out on a limb and

14:52 assume it may in fact be the case. Now, we can see a few of what was going on right here, such as teamwork happening, multiple agents being spawned to perform different tasks and things of the sort.

15:01 So, just cool to see a little bit of a TLDDR overview of what specifically it did to actually create this. However, I am quite eager to see what we actually get because again, I don't know how

15:13 bullish I am on this. Okay. So, we're I think we're definitely going to be seeing perhaps some of the limitations here of Muse Spark 1.3 compared to a bit of a more capable model. Now, we

15:26 definitely will have a couple of issues that we can maybe be uh send this. All right. So, I've given it the error log right there and we'll just take a peek and see. So, in 8 and 1/2 minutes, we

15:37 hypothetically have a fixed result here. It just said do a hard refresh, which I inevitably have forgotten the key mapping for on a Windows system. So, one second. So, we did get that issue fixed.

15:49 However, unfortunately, there is now a different issue and still none of the gameplay is loading in. So, I'm just going to tell it that. So, it's reporting both symptoms fixed in about 4

15:57 minutes. And that is one thing about this model is it is really quite fast. Very good. I don't believe we had issues, at least preliminarily. Okay, we're definitely seeing some progress

16:08 right now. It seems to be having some graphical issues. Good. Very good. So, one of the things I'm noticing is I think it didn't render because I had the development console open. So, let's hard

16:24 refresh and then we'll get a full screen look at this. Good. Good. This is kind of I think this is actually an interesting take here because it shows us that there

16:36 are elements of this that are pretty expected based off of what we've seen from both the original and then the very very high quality replication that Deep Seek had done. The enemies are

16:46 [laughter] quite something. Um now we should be able to swap weapons. Okay, I'm not noticing that ability. So unfortunately that seems a bit lost. Let's Okay, there

16:59 there are no bullet holes in the environment, which truthfully is something that I was kind of expecting would not exist. Let's just see if we can actually do any Wait a second. There

17:08 are bullet holes in the environment. At least over here. Well, in certain spots, I'm actually okay with that. Was that our health in the bottom left? Because if it was [laughter]

17:24 Reload, reload. All right. Well, [laughter] I don't know that we're going to be able to get to the point here where a train

17:35 actually shows up. I'll try one more hard refresh. All right. So, yeah, this one is definitely going to provide some insight just in terms of how Muse sparked it. Although I will say

17:47 it followed like the core of this, you know, the the skeleton is there and we do even have this upper section right here. So that's actually kind of cool to see. And

17:57 we get a different view here. And we do fall back down to the ground. So acceptable and kind of insightful, I think. So next up, we're going to be transitioning to something that I

18:06 believe is a bit more difficult. This is a C++ result that I just had Fable 5.1 make using Ultra Code. I stopped it when it was around six or seven hours in because all that I really wanted was the

18:18 design markdown file, of course. So, this is for a retro Formula 1 style racing game. I do believe it is allowing them to use Ray Lib. If you watch my videos, thank you. And you know that

18:29 sometimes I don't let them use RayBib for C++ results because it makes things more difficult. However, in this case, I just want to see how well these models will replicate it. We're going to be

18:39 starting off here just using GLM 5.3, the full-size version and not the flash variant on extra. So, the highest potential reasoning effort. Now, I would like to quickly just showcase what

18:50 exactly the end result that Fable had produced looked like for this specific task because it really was quite impressive, especially considering I had stopped it prematurely. And this will be

19:00 essentially what we want to see replicated by GLM53. So in addition to just a short video clip of how this game actually looks, here are just some static images that were taken during the

19:10 build process. So essentially this is basically what we are aiming to replicate here. It is pretty cool. There were multiple cameras denoted. We'll also take a quick look at the agent or

19:20 the design markdown file. But this is a pretty cool retro looking result and I think it'll be fun to see how well we can replicate this. Now, GLM53 is setting up the

19:42 dependencies on this system. Being that this is a Windows system, I don't generally use for a lot of this work. It doesn't have anything like Rail or all of that stuff on it. So, it's just doing

19:51 that right now. If we do see things flash up like PowerShell opening throughout going through this markdown file, that is going to be specifically why. But, let's take a look at this. So

20:01 once again, this is a pretty beefy design file that was created, of course, by Fable 5.1 on Ultra Code. We can see that this is much shorter than the Subway one. However, it is

20:12 1,573 lines, which is still a lot. And there's a lot of reference snippet, sample code here, and things like that. So hopefully this should allow us to get to a pretty

20:22 decently replicated point here. All right, I believe this took probably about an hour and a half. In the meantime, I've been troubleshooting an issue with the new iPhone, so I was kind

20:31 of just letting this run and not really checking in on it. But we can see right here we have our game delivered. I've not looked at this at all. I didn't see any of the potential screenshots it was

20:40 creating to try to debug or anything. So, I'm really very interested to see how this looks compared to our source that was just done with Fable 5.1. So, keep in mind as well, I did have Fable

20:51 5.1 at least finish this up to a point where I deemed it had gone enough to stop it. That was somewhere around five and a half to six hours where I stopped that. So keep that in mind when we look

21:00 at what this made in about an hour and a half to two hours max versus that. All right, let's take a look at it. My knowledge of how to run some of these things on Windows. Oh, is quite all

21:12 right because [laughter] Okay, so um All right, the car models look good. We're definitely noticing perhaps a couple of

21:28 issues. Good. We can sort through them. There's quite a decent amount of them. Let's just go. I liked the red and white car. Oh, okay. We have different camera angles.

21:43 No. Oh, okay. [laughter] All right. So, there's our cinematic camera and then Okay. There we also have um a restart. Okay. Okay. So, we're definitely going to

22:01 notice here it is perhaps [laughter] more of a differentiation between the two models. Now, keep in mind this model is significantly cheaper to run than Fable 51, and it got a lot of it more or

22:15 less similarly. The track looked good. Some of the car models themselves did not look bad. Things basically fell apart when we actually tried playing it, but like the skid marks and things

22:24 looked pretty good. So, this was definitely more of a difficult test. And I really can't help myself, but I also want to just try this with Muse Spark 1.3 because it's fast and I'm kind of

22:36 just curious to see if we can get some decent work out of that. This one had a lot more reference code in the documentation. So, I'm wondering maybe if Muse Spark can just follow the

22:45 instructions pretty closely, then maybe we'll be able to get something more impressive than the Subway game. At least I would hope. So, let's just look at this one more time. And then I will

22:55 essentially have to nuke this folder and we'll have Muse Spark go at it just to see because I am really curious to see what that model can do. All right,

23:06 I'll put it here. Like right here. This looks good. Well, okay. And then and then we kind of lose it a bit. But interesting to see what was actually maintained and then what

23:23 unfortunately went a bit ary. I don't actually really have much to say beyond that. It was definitely an interesting experiment though. Okay, so Muse Spark went for a very very

23:35 different approach here where it essentially turned this into a single file HTML game. Why HTML instead of C++ and Ray? Well, because it doesn't have those dependencies. So, let's just make

23:48 it something you can run. So, oh, I'm so used to using Ubuntu. Let's open with All right. Well, this is it's a

24:07 replication of the source. [laughter] [snorts] We'll see if the sound was included. Interesting track layout.

24:33 Well, that is [laughter] Were you incapable installing the dips? I am just going to ask it a follow-up question here. We may not go

24:48 through with the full rebuild of this in the specified language. However, I'm curious to see where its thought process was at when it decided this. Oh, okay. [laughter] I just addressing user

24:59 frustration and planning to follow the C++ rate. All right. We'll see what we get. All right. So, unfortunately, Muse Spark did not properly complete this with C++.

25:12 Though, as we saw, we did get that HTML variant of it, which was curious. But overall, this was really interesting because the thing I wanted to test with these various models was how well they

25:23 were able to replicate a given reference result when given the same exact agents or not agents, design markdown file that led to said result from a state-of-the-art model, in this case,

25:35 Fable 5.1. I will say Deepseek V4.1 Flash seems like an extremely extremely potent model. I almost have to wonder if that may be a bit better than GLM 5.3, the full-fledged one. And it's pretty

25:48 interesting to think that the Pro model has not yet been released for Deep Seek V4.1. So that's still the Flash model, even though it's bigger. I am very, very excited for that Pro model. Now,

25:58 overall, of course, what we also tested today was the Klein desktop app, and some of these sessions were going on for quite a long time. In this video, I filmed over basically like a day and a

26:07 half. Additionally to that, all of the models that we used were just through Klein Pass right here. So, having access to a bunch of these different models, which of themselves or by themselves

26:17 would require separate subscriptions, which I basically do have for most of these, and it is kind of a pain because sometimes you'll stop using one and you'll forget and then I'll see a charge

26:26 like, oh, I'm spending like $60 a month for this model that I have not used in a couple of months. And that does happen. So being able to just have one specific point where you can use all of them is

26:36 definitely very helpful. Additionally to that, you can use the API key for your client pass account and bring it into a different coding agent if you want. So you'd still be able to access these

26:47 various models from within the Klein Pass subscription, but bring it into somewhere else like almost a bring your own key, but you would bring this key into a different say PI or Open Code or

26:57 something like that and still have access to all of these models through Klein Pass. So, that's pretty cool as well. Additionally, the desktop app was pretty cool. I do like the eye movement

27:06 and things like that over some of these very long things that we had generated here. I didn't notice any hiccups with it and things seemed to run well. It didn't crash. Everything was smooth.

27:16 Keep in mind, this still is in beta and it will be released for Linux soon. So, if you're like me and that is exciting to you, don't worry. That will be coming soon. With that, I suppose that is

27:26 probably going to wrap up today's video. Again, thank you to Klein for allowing me to do this video and supporting the channel. Additionally to that, if you liked this video idea where basically we

27:37 see what these models can do with the agents markdown files, please let me know because I would think it'd be cool to have a repository of some really intricate markdown files done with like

27:47 Ultra Code on Fable 5.1 or GPT6 Astra on its highest setting just to see how folks and the models they run will be able to do that. It'd be cool to have like a benchmark chart of how various

27:58 models do when given the same static, very intricate markdown file to follow. So overall, that is going to conclude our first look and test of client desktop and the client pass

28:09 subscription, which allowed us to basically access every model that we use today. If you have any questions, please feel free to leave them in the comments. There will be a pinned comment and link

28:18 in the description for more information. And thanks for watching.

Frontier News · by Hyperjump Technology