Transcript (captions)
What do you get when Claude, Gemini, and GPT walk into a bar? The worst browser OS yet. Today we're going to be testing OpenRouter's Fusion. Now, this is something that a few folks have mentioned in recent video comments that I should test, and with the fact that Claude Fable was recently removed from access for anyone, this is something that is very interesting now because hypothetically, as we see in one of the benchmarks right here, fusing multiple models together, and we'll talk a little bit about the specifics of what that actually means, can hypothetically outperform Fable 5 just by itself. Now, the caveats to this are the specific benchmark that they used to come to these scores was more related to deep research tasks and things, not necessarily just raw coding capability, but the entire premise of this is something that is not 100% new.
It has been around for a while, the concept of using multiple models together, and it is very interesting though. I do wonder about perhaps some cost to do so because we're going to be using multiple state-of-the-art models in conjunction with one another. So, before we get into it, please do feel free to subscribe so I can hit that 100k figure, and let's start out by just seeing what exactly they mean when they talk about the concept of fusion, and it is listed in the introductory paragraph where they found that synthesizing the results of multiple models can significantly outperform what the individual models are capable of. So, fusion, as it works here, at least in specific for OpenRouter, this is a So, it you can choose a panel of participant models alongside a judge model responsible for fusing the individual results together. Now, essentially, this just takes the answers from multiple different models.
It will pick the best pieces from when specific models answer, and then the judge model will fuse the best possible pieces together to provide the user with the final answer. It's very interesting, and I do have to say, I did try something like this a while back, so this was posted 5 months ago. So, basically late 2025. It was a total disaster the way I did it here, but I did also have a feature in this video where they had to scrutinize each other's work as opposed to work together, which definitely made things a bit more entertaining, but less performant. So, they showcase here a bunch of different model combinations and their scores.
Unfortunately, Fable 5 is also listed in here. Rest in peace. Though, apparently, if we look down here to the first result where Fable 5 is not present, we see that this result will outperform Fable 5 by itself, and that is Opus 4.8 plus GPT 5.5 plus Gemini 3.1 Pro synthesized by Opus 4.8. And keep in mind, this score figure right here is specific to the benchmarks that they ran where they say they've performed 100 deep research tasks from the Draco benchmark. Some highlights of what they found.
So, this is not necessarily like, "Oh, it's SWE bench." or something. It is a specific benchmark, but they do showcase that there is some gains to be had having these models work together. And the whole point of fusion is to basically allow you to call this, as it were, a model in and of itself. So, right now, if I go into Open Code on the system, it may be a little glitchy here. I'm still having the specific settings worked out, but if I just type model slug, which will give us the name of the specific model that we're using, it will return here just saying Open Router Fusion.
And as we see right there. So, this is essentially able to be called as it were just a new or specific model. So, it's going to be very interesting to test these out. Now, beyond just showcasing, "Okay, if you bolt all of these state-of-the-art high-end expensive models together, they perform well." They also have different options here. So, some including open weights models like DeepSeek V4 Pro, and then mixing and matching.
So, in this case, we have Gemini 3 Flash along with two additional open weight models, and that's performing very well in this specific test, actually better than GPT 5.5 or Opus 4.8 by themselves. So, this is just conceptually very interesting. To begin, we're just going to initially run this from within the web chat interface where they do allow you to use this through this model fusion chat pane. Of course, we're going to start out with the tried and true browser OS test V2. And we can see right here, there are some presets depending on what specific model combination one would like.
So, quality is the latest state-of-the-art models across Anthropic, OpenAI, and Google. Okay, so our first go at this was a fail. Nonetheless, I'll just click retry. All right, I'm actually glad we ran this first from this web interface because it just gives us a bit more insight into what's going on underneath the hood. So, first step three for the analysis, and we can of course expand each of these, we're basically getting Okay, all three delivered the files.
Architecture centers on a global OS window. I'm looking here to see at least one game uses ray casting with DDA grid traversal to produce dependency-free pseudo-3D wall slices. Car control uses Euler physics. Okay. Key differences.
And this is actually kind of interesting because we see that there are differences in the special feature that each would have chosen to use for this specific test. If you watch the channel, thank you, and you're very familiar with like all of these things, at least in terms of what's included. So, this is very interesting. Partial coverage, unique insights, Z-buffer billboards, sprite occlusion, fill Okay. And then we have like almost citations essentially of which model specifically came up with each of these things.
And we do have some included from each of the models. Blind spots, none of the games were actually tested, just being the nature of running this through this web chat interface. And then step three is our final answer. So, this is the fused response. I've synthesized the best approaches into a complete working browser OS.
Here we have the complete fused answer. So, let's take a peek at this first, and then we'll talk about it a little. All right, so here is our final fused browser OS result answer. Okay, on first glance, this is pretty bad. Just aesthetically, I've seen better results individually from each of these models, but again, this is a first test.
There is no right click. There is a correct time in our locale. Let's just check our start menu. There's no search function or anything. All right, we'll just run through these sequentially one by one.
This is genuinely the quality that I would have received from like a 12 billion parameter Gemma. Let's just try random upload. Good. Okay, so that does work. We'll go back to something a little easier on the eyes.
Notepad. I'm going to assume it's okay. Simple. Type anything, it auto saves, and then Chronos can rewind it. Chronos inevitably being the special feature that was chosen to be implemented.
Okay. We don't even have like Neofetch. Calculator. All right, 54. I did this one times six, 300 24.
I'm just like speeding through these because my interest lies heavily in some of these games. Okay. So, Maze Raider 3D is sadly not 100% functional. Let's check our GTA game. Um oh, okay, very good.
It just uh the camera spawned in a building. Oh, wow, that's nice. So, the the patch that gets left when we accidentally collide with a pedestrian is a nice touch. We do have a wanted level. Can we get out of the car?
No. Okay, perhaps this was not the right test to run with this, but I did just want to Okay, restore this moment. So, the snapshot would essentially take a photo or a snapshot of everything that's currently open on the system, And then if we were to refresh this, it should save that in persistent storage. So, let's restore that moment. Okay.
Well, uh let's go back to open router here. Again, this was not touted as working well with code, but I did see so much coverage of this saying like this can beat Fable, and that is partially a side effect of like the way it was marketed, I would imagine. So, we're going to test it with code. All right. So, I've used this only just partially in open code ensuring that some of the configs worked, but that was very minimal aside from whatever system prompts and things like that open code will include.
So, more or less that browser OS result probably cost around $2.50, which may in fact be one of the worst um trade deals in history. So, again, it's possible this may be some form of out-of-scope testing, but I do like out-of-scope testing. All right, I have Fusion set up in open code. It's using the same exact models that we had just tested with the web chat interface. I have begun this from within plan mode, and I'm giving it the self-contained C++ skateboarding test.
All right, so we received a plan with some follow-up just asking some questions. The only thing I've told it is do not use raylib since the Fable 5 result didn't either. All other questions are up to you. Build it. And I've swapped it into plan mode.
Now, it's going to be doing a lot of research as this is kind of poised towards that. So, that's all right. We'll take a peek at whatever it generates for us. Okay, so I just realized that I don't still have the Claude Fable 5 skate game on this. I'll post like a picture of it or something, but we do have our C++ skate result right here.
So, I'm very excited to This is actually quite good. Well, aside from the Okay, I got to I have to not speak before like I'll say this is actually pretty clean and competently done. I have individually run this test with all of these models by themselves, though I do feel like I recall Opus, it was either 4.8 or 4.7 may have had a better one than this just in terms of the actual aesthetic. However, the walking looks good, our text looks good. Everything's actually just very clean.
This is a respectable result, I will say. Let's go over and take a peek at the water. And we'll try some of the other tricks as well. So, Q would be a kickflip. Okay, yeah, no.
See, that is something I don't want to see. That I would expect from maybe like a mid open weights model to just lock the actual player and the skateboard in the same motion for a trick. >> [laughter] >> The water looks good. All right, so here is where I want to kind of maybe touch upon the downside of that where if we go back into activity right here, we're going to see that up until running that we were at $2.98 spend. So, the entirety of what we just did will be whatever $2.98 to $12.70 equals.
So, really for coding, I'm going to make the determination here that even though I've only run two tests with this, it's not a good idea, at least with these state-of-the-art models to use them together in terms of coding. But, that is not necessarily the specifically listed thing that this is designed to excel at. That is deep research. So, I just so happen, completely coincidental, have a task that I performed the other day with Google Gemini 3.1 Pro in deep research mode, and we're going to now run that same exact task here with the Open Router Fusion model, and we're going to compare the results. So, I want to test this where it is listed as having a strength, and that is in the deep research capabilities.
Now, truly by genuine coin- coincidence, a few days ago I had used Gemini 3.1 Pro on extended thinking for a deep research task where I was looking for additional information about a car. It was a car that was hypothesized to have been perhaps shown at some car shows back in the day, and there was sparse information about the shop that built it and things of the sort. So, I asked it specifically to find information that was niche, not common things, information that was older than 2025. I'm looking for old photographs, passing mentions of this very identifiable vehicle, and anything that is older information. Unfortunately, when I did run this, it basically gave me a research report that almost was more in line with a creative piece of writing than what I had asked for.
I was essentially looking for, "Oh, here's like a really old archived page that actually mentions this specific vehicle." Instead, I got the genesis and provenance of the 04 Okay, that's just weird and not what I wanted. And then it basically gave us a bunch of very weird like the canvas, the architecture of this vehicle, things I had zero interest in that I could have found on Wikipedia. Following this, I basically scolded it. I didn't I said I just wanted to know if there were any pictures or mention of this exact specific car you could find from back in the day. It said, "I can't." I said, "Search Washington State because that's where the vehicle was built, that's where the shop was that built this vehicle for local archives, boards, or anything that may mention this older than 2012." And it just found a few very sparse things and forum posts relating to the name of the shop that built this vehicle.
So, this is a deep research task that I genuinely had done for me a couple of days ago because I needed it. So, I just went and ran this exact same thing through OpenRouter with Fusion, and I'm going to say this is 100% a strength of this, and I do actually believe that this is a pretty reasonable approach to deep research because what I'm going to find here is I sent this the exact same prompt that I sent to Gemini with deep research, and in the analysis here, there is a section on blind spots. Now, the downside is the blind spots are things that were not included in the actual results from the model, but everything listed here is a very deep deep combing of looking at the assessed results and identifying like, oh, this should have been brought up. This should have been brought up. Things that were not brought up with the Gemini result are listed here in this blind spot section.
So, even if they're not present in the actual end research report that we get here, this is fantastic. This seriously, because this is genuinely, and I know it's completely coincidental, but this is something I was doing the other day, cuz I bought this car, cuz I wanted it, but that's not important. The thing is, this actually identified that the shop that created this car had a show car Lamborghini, and it had the same exact rims that are on this car that I purchased as well. It's some old like cheap Range Rover. So, it's not like a big expensive purchase, but the fact that it identified those.
Now, I was only able to find that out by finding some old website on internet archive and being able to visually cross-reference, hey, that has the same exact rims that are the on this car. It was also mentioning internet archive to find the website for the shop that built it, which is now defunct. It mentioned like Dub Magazine, and it's just very interesting. It also basically talked about how Gemini, some of the claims fail verification, which could mislead a reader. Everything I'm seeing here, and I can 100% confirm that this is a very strong performance, because this is something I had to go and manually do the other day.
I spent like 4 hours researching this, and this surfaced a ton of stuff that this would have been significantly it would have saved me a lot of time. So, okay, yes, there's one hard record of the truck itself, which is fine, because it's not really recorded. Okay, it's talking more about the company, and what we're going to see right here is that this is far more isolated and focused on the actual question I asked, where the Gemini deep research one, and Gemini deep research is a very, very strong research model that kind of just went on like a creative writing thing and wrote me like a Wikipedia page just about the car in general, not specifically to that car, just that model. This is absolutely 100% correct. The car used the exact same formula as your truck.
This is very, very, very well done to find all of this information and I'm very happy that we did this because some of the coding results were kind of not the best, but this is absolutely fantastic and this does tie in directly to what they were saying that this benchmarks properly, claims to be skeptical of, okay, and it lists some things as well. Will the period photos most likely still exist? It's saying go to Wayback Machine reverse image search. Something else I want to notice, this is a succinct result, which in this case is actually better as opposed to what we had received just with Gemini by itself, where okay, we get like a bunch of random stuff here, but not necessarily what I wanted. Whereas what we did with Open Router Fusion was 100% on the ball for what I wanted and this I can definitively I will stand behind saying that this actually seems to be a very, very strong performing thing for deep research tasks in specific.
Now, I would imagine that's more of a niche thing than maybe people who use models to code, but it does have a place in terms of like I don't know, but this is impressive. I'm genuinely impressed with what we're seeing here. Well, at least for the deep research, not for um not for the coding. Let's see how much that See, okay, so we were at $15.44. And we've gone up to $18.93.
So, just that process of deep research through the web chat interface 15 This is $3.50, $3.49 basically. So, this is very expensive to use and we also do need to remember that I am using this on the specific quality setting. So, it's using the Frontier state-of-the-art models. I want to now, cuz I don't really have other deep research tasks that I can actually understand in depth to be able to look at and be like, this is right, this is wrong. I want to try it again, but with a front-end web design test.
So, for this next test, I'm giving these models a prompt that they need to create a beautiful front-end for a watch website, but they need to actually create 3D watches and then have a cinematic panning shot of said watch in the hero section. So, it's front-end design coupled with actually creating something attractive in terms of a 3D model. Oh, okay, I spelled that wrong, but I'm actually going to leave it the same because this is the same exact prompt that was run with Fable 5. I have also disabled web search here because there's no real need for it. This is entirely just a code test.
All right, so right here is the Claude Fable result that was produced for the same prompt that we have just run through the Openrouter Fusion. So, this created a really beautiful 3D model of a watch, and it did also have the panning animation for it. Just for some reference in terms of what was generated with Fable 5. So, now let's take a look at what we received with Openrouter Fusion. And before we take a look at it, let's just see how much it cost.
So, we were at $18.93, and now we're up to $21.46. So, basically, uh more money. I can't do that in my head right now. So, um I'll put it on screen how much. Okay?
Okay. >> [laughter] >> This is definitely a step down from what we received. Although, the similarities do exist, and that's Okay, that's actually kind of cool. That's due in part to the prompt being pretty static. This is actually a fairly good result.
Now, this funny enough, it made the same mistake that GLM 5.2 did with this prompt, where it keeps the glass too high and too small. So, that was actually noticed with the first run of GLM 5.2. Though, the crown is very well defined and some of the individual pieces of detail in this watch are actually quite nice. I will say, independent horology rendered for the digital I don't even know what that word is. I'm not going to pretend like I do.
Okay, and then we also have our summer collection with the two watches right here. Look at this. There is some actual luminosity to the number markers and I do believe the second hand is moving here as well. So, while this is not 100% in line with how well the Fable 5 model was, there are some really nice elements to this here as well. So, I did just want to do something that was basically a direct head-to-head comparison of the results side by side.
Now, truly, I do also have one more kind of interest and what I would like to do is run the same exact prompt, but in lieu of using it in the most expensive setting, I'd like to use it in the budget setting as this is a very, very interesting grouping of models. Now, it is just defaulting to fuse the results with Opus latest, but the models that will actually be producing the answer are Gemini Flash, DeepSeek V4 Flash, and then Moonshot Kimmy latest. All right, we have our completed result with the second test. So, we were at 21:46 and we only went up $1.40, so this was significantly cheaper. And just as a reminder, although I suppose it This was Gemini Flash, DeepSeek V4 Flash, and Moonshot Kimmy latest with Opus 4.8 as the model that fused them together and judged them.
So, let's take a look at our result when using the cheaper models. Okay, that >> [laughter] >> I think we can definitely see the difference. Although, if this were laid flat and perhaps some of the emissive properties of this chrome was not blinding or not blinding, the second hand here is actually quite nice. Let's take a look down. Everything in the mid-section is always very consistent.
Okay, and it is just static photos of the same watch. Unfortunately, it's not really animated to the same degree. This is a Oh, wow. So, we actually do have the ability to freely orbit around this. We see a strap there when we go like that.
That's kind of frustrating because there's definitely some promise here. I mean, the argument could be made that this is a more usable result, but it also does look pocket watch-ish. So, maybe the Swatch collab was used as a reference for this. So, that's probably going to conclude this because just from the short bit of tests that we've run here, I think I'm pretty confident in stating that for deep research tasks, this does actually seem to produce an improvement over just using one of the models in isolation. I'm fortunate that entirely by coincidence, I did have a deep research task that I ran a few days ago that is exactly in line with what this is saying it's good at.
And when we ran it with Openrouter Fusion with the state-of-the-art models, they did produce a significantly superior result to what I saw when trying it with Gemini 3 1 Pro extended in deep research. It gave me basically exactly what I had been looking for, so it was nice to have that side-by-side comparison in something I have some pertinent knowledge about. When it comes to coding, I I feel that the results here were what I had noticed in late December when I had run that models working together test video. I don't quite know that using models together for coding like this is going to produce better results than what one would get if they just used one of the models in isolation. It's definitely significantly more expensive.
I mean, the entirety of what we did today cost about $20, maybe a little more, which I mean, that's a one-month subscription for Anthropic or GPT, and you would get a lot more use than that. So, this definitely has a place perhaps, but I find it's a very niche place, and it's just interesting to see more evolution in terms of the way in which people go about using models. So, I wanted to test this because there was a lot of interest about it. If you have any questions, please feel free to leave them in the comments and thanks for watching.