Transcript (captions)
Claude just released Ultra Code and it is incredibly powerful but only if you know how to use it correctly. Ultra Code increases the speed, performance and accuracy by creating up to 10 parallel agents. This unlocks new capabilities and is an increase on whatever system you were using previously and nobody's talking about exactly why. And in this video, I'm going to show you exactly how to use Ultra Code correctly. And how to build your own Agentic operating system and the big watch out that you need to make sure you know so you don't make any mistakes.
All meaning you can make more money, save time, and get light years ahead of everybody else. And if you're new, I'm Jack. I built and saw my LEX startup with a gazillion customers. Now I build my own AI companies, and I share the stuff that works on this channel. So if you haven't already, grab that beautiful coffee and let's dive straight in.
So when we talk about ultra code, one of the best ways to sort of like understand or think about this is that we have one boss so to speak that is running 10 different specialists. Think of it like a factory floor and essentially this boss script is actually running multiple agents in parallel and then we actually have a judging function and this is going to be one of the biggest growth areas that we're going to see with AI model use in building anything. Now, when you're in claw code, you'll notice down on the bottom right hand corner you have these different settings, right? You have max and you can see it goes all the way down from low, medium, high, extra high, and then max level. Well, that has nothing to do with ultra code.
You actually activate ultra code by just using the word ultra code in the prompt. Think about the reasoning effort as how hard any given agent is thinking. The more that it's thinking, the more tokens that it actually spends. And then ultra code is more of this new strategy which kind of takes what we used to do previously to a whole new complete different level. Again ultra code you can do it for just one task or if you do ultra code on you can run the entire session in ultra mode.
Now as you know there's six things that we can do with include code. We can have the regular chat. We can have plan mode where we basically plan out the thing we want to build. We can have ultra think which extends its thinking over a period of time. sub agents where basically Claude just creates mini sub agents to do something fast mode which is the same multiple speeder but then ultra code which is the hero of the show here replaces the managing LLM with code and effectively it's called here a deterministic fan out what on earth does action mean well like sub agents we're still going to be spawning up loads of mini clots which is the same underlying model as you can see here to do things but instead of it being an agent that manages it and the performance can decay this guy in the middle can forget things or whatever it is, we actually have a script running it which means that it will not actually finish until we've achieved the specific output we want to and the general quality of the entire system will be a lot higher.
While there is so many interesting different approaches that we can use with ultra code like adversarial verify a judge panel pipelines what we call something called loop until dry. I've got a little bit more detail here on what they actually do. For example, adversarial verify will spawn a number of independent skeptics each told to refute a claim and kill it on a majority vote. Plausible but wrong findings do not survive the panel. So this is an example of how you can physically use it.
Then I'm going to show you how you can use this in the Gentic operating system which is going to freaking blow your mind. We have this thing called judge panel. This will generate several independent attempts, have judges score them, then synthesize the winner grafting the best ideas and running up. So, for example, why don't we show you exactly how that works? So, I could come over here and give it a prompt.
I might say, "Hey there, I'd like you to go ahead and use Ultra Code and basically spin up a panel of judges to debate what the best idea is for growing my LinkedIn. Number one is I learn everything myself. Number two is that I hire a top tier agency. Number three is that I hire an intern to do it and maybe give me a couple different ideas and then just show me what that output is." So, pretty much for a random question, but you'll see exactly how it works now. So, as you can see now, what's going on here?
It's got a perfect job for a judge panel, which is just one of the applications of Ultrakum. It's got a debate. So, each of your three options gets a sharp advocate, best possible case, then a ruthless skeptic, flaws, hidden costs, and failure mode. Then, we're going to have a judge, a fourperson panel is going to argue this out each through a different land, a growth strategist, a CFO, a time leverage coach, and a brand authenticity purist. Then the verdict.
The head judge calls a winner, scores all three, and throws in a couple of fresh hybrid ideas beyond your original three. If you're saying, "Jack, this sounds like an episode of Law and Order. I've seen courtrooms work like this." You would not be mistaken. So, you can see now it's got everything here. Now, what's really cool is we have 11 different agents.
And crucially, they're not managed by a master agents. We've actually created JavaScript or code to manage this. Now, check this out. I click on this. You can see, look at what's happening.
Our debate is happening before our very eyes. And we can see we got the judge. No agents have started yet. We got the verdict. But look at them go, guys.
I can see the tokens. I can see the tools. I can see how long they're doing it. Now, these agents are running in parallel. We call it parallelization, right?
Try to say that after a couple of cheeky ginonics. Very difficult to do. So, the point here is that we basically don't have to wait like an hour. Like things that would normally take 60 minutes, 2 hours, we can do now so much more quickly because they're all running together in parallel. Now, it's really cool.
And actually, you'll often find that agents when you're asking questions, not only do they hallucinate, they forget things, but they can become confidently incorrect. And what that means is you don't know when it's actually telling you something incorrect sometimes. And one of the ways that we can account for this is to have the agents ruthlessly and mercilessly critique each other in a loop. Because the more separate agents that we have challenging one another, the more likely it is to actually arrive at the best idea. As you can see now, we've got the judges, the growth strategist, the CFO, the time of leverage coast, and the brand authenticity purist.
We know from best practices in prompting that the best way to do this is to assign a specific lens, a specific role to actually then attack it. And this realistically is the best way for you to see exactly how this entire ultra code system actually works. And by the way, if this all sounds like I'm speaking Mandarin, I'll put a link down below for the full Claude code uh master class. It goes through foundation setup, building websites, power features, memory systems, Hermes agent stuff I have never shared on YouTube. It is the best thing that I've ever built terms of a course.
You get immediate access to the entire um basically cord code and Hermes operating system. I'll put a link down below so you can check that one out and get all of the juicy benefits cuz it is very detailed. And you'll understand guys when we're using these new strategies, we need to know when to apply it. And it's incredible stuff cuz you need to be slotting in Ultra Code. But you've got to be integrating at the right point.
And I'm going to show you exactly when it's perfect and when it isn't. And when I show you what you can build with those, we're actually together going to build out a brand new functionality in the Aentic operating system that is going to be incredible. And you'll see what I mean when we get to that point. One of the cool things I want to draw your attention to here is we've got the tool use, which is fantastic. How much time it took, but we can see the debate, the judge, and the verdict coming in as it comes through.
So, it's now currently just synthesizing all of the individual um aspects together. And it's wonderful. I mean all of these are effectively Claude Opus 4.8 and we can run this in any effort level we want to remember. Effort level is one axis and then the ultra code is just simply a different axis. Now it has essentially completed and now we're going to get the full output back from Claude.
And as you can see we get the full verdict back. Now this is just one particular example of how we might go ahead and use this ultra code feature. We can also use this to do something called perspective verify which is essentially basically getting a number of identical reviewers give each verifier a different lens. We can do pipeline versus parallel where it lets each item flow stage to stage with no barrier and a few others including completeness critic which essentially essentially ask the question what did we miss? So loads of really interesting ways that we can go ahead and use this.
Now we understand what this incredible technology does we need to apply it to a project so you can see exactly how this works. So we've got here the claude code operating system. This is the knowledge graph that I covered in my last video on Graphify. We're familiar with the operating system giving us the full view of our spend, how much we're spending in each platform, Codeex, anti-gravity CL code, dynamic suggestions, it dreams overnight based on usage, goals, memory systems, everything. So, what we're going to do here is add in a section to our wonderful cord code operating system, our ogentic OS.
And what I think we should do because one of the big areas in the future is going to be the ability to use the correct model for the correct thing which is why this OS is so important whether you're using Hermes agent open claw or whatever it is that you're physically doing we need to be using the correct model at the correct time and to do that we need to understand performance and cost and we call this the best model for the job in other words you are model agnostic okay so we have cost we have performance and we have all these different models claude openi Gemini Meta um you know as a company all their different models Mistl Deepseek and Grock and what's the idea here is that lockin is the enemy of high performance okay the winners are going to root by task so if we're using Homies agent or even using chord you if we're doing something with multimedia we want to tag in Gemini right if we're doing something like code review we may want to bring in chat GBT 5.5 but the model landscape can change so quickly and you having to think about it all the time is time that could otherwise be spent actually building things and so what we're going to do is actually build the system directly in iOS using this ultra code and I'll show you exactly what I mean. So the first thing I'm going to do is come over here. I'm going to give it the below which is hey though I would like you to use ultra code to create for me an improvement to my claude code operating system. Specifically, I would like a section, okay, probably somewhere in the homepage, maybe at the bottom, that basically covers models that I can toggle on and off. And effectively, what I want it to do is have a direct link into various different online sources.
Uh, that essentially shows me what are the best models that are running right now. Maybe we could get some data from open router and I would love to see all the different models. I want to see all the logos visible and essentially I would like it to be able to know what are the up andcoming models, what are the pros using right now, what is the performance, what is a benchmark, what is the sentimentality, maybe it's something we can refresh. The stated intention here is essentially so that I can look at this and use it as a knowledge base for my Hermes agent for my cla code to know exactly what model I can tag in for any specific task based on price, performance, speed, that kind of thing. Feel free to challenge my thinking on this, but essentially I want it maybe somewhere on the homepage near integrations, automations, that kind of place that doesn't clutter up the interface.
Before we start the ultra thing, you feel free to ask me any clarificatory questions. I do want it to be very up to date. Okay, so basically I kind of just explained all the things that I want to I would recommend that you do this before setting off Ultra Code cuz you saw the last example. I just gave it a random idea and it brought in a brand specialist. It brought in a you know a head of content or whatever it is.
Specifically speaking, the Ultra Code system is only as good as the thing it's doing. It's like, for example, putting a billion dollars into a project about trying to find sand in space. We could spend it. We could put our biggest efforts and warp power there, but are we actually putting it in the right place? You know, answer is probably not for some things, which is why you always want to have this context conversation ahead of time.
So, what's cool here is asking us a series of questions. So, what should the onoff toggle actually control? In other words, how does this feed Hermes and Claude code? So, essentially going to come down and give it this feedback here. don't need you to send anything to Hermes and claude code but effectively this is going to be a resource that we can call using Hermes or claude code to make informed decisions.
Okay, I'm going to come down and click on next. How fresh should the data be given the benchmark sentiment have no clean API? I think we should do something like on a weekly basis. So hybrid live open router plus dated snapshot. I think that's great.
Footage on the homepage. Cool. So I here's what I think about this. Awesome. for the homepage.
What I think would be great is I wonder if we had see these sessions per day at the bottom. Why don't we just make that something that's toggle like you can toggle on a couple of things like maybe we have something like sessions and we can toggle between sessions and models. Um so we can kind of like really think about the information density on the dashboard to make it a pleasure to use so we don't over complicate it. Little side hack by the way that is one of the biggest almost I'd say fumbles but one of the biggest side steps I see with any kind of dashboard or website in text density way too high. You'll see it on dashboards with like a 100 things on the left.
I would never go above five. What I would do as we build this out in the community I would probably either bring them into one section. Okay. Agents is slightly different. It's very clear.
The mind clusters things together. It groups things together. So it's okay looking at five things but I wouldn't have six. I'd probably merge two together or remove one or do something along those lines. And as you can see, Ultra Code is now working on the right hand side.
And we can see the different phases that we've got. We've got recon. So, this is going to do some reconnaissance on it. And they can see got multiple different agents going on. Then we'll have merge, verify, design, and finished by build, which is fantastic.
We've got all these working for us in the background. Beautiful. So, now it's gone ahead and designed it. We can come over to the operating system and take a look at what it has designed. And take a look at this guys.
We can search now by many different things. So, we've got down here default. We can search by smartest, uh what's le on the arena, the cheapest, the fastest, uh the most used based on the benchmarks. And this is super interesting. So, Deep Seek V4 Flash is the most used.
I can click into that. I can have a look at different things. It's really freaking cool actually. Level of basically configurability we get here. We can filter by Frontier models, filter by speed, filter by what's open source and what isn't.
This is really freaking handy actually. So I can scroll down and have a look and at the bottom here we've got a full breakdown. So as you can see the deepseek v4 flash this is the number one model in opinion by token throughput the cheapest credible agentic brain with $1 million context. I've done a full video breaking down deepseeek how we use that with Hermes and Claude and it's really cool. I just now get this beautiful breakdown for all of the models.
So if I'm ever thinking dude what model should I be using for this? I can actually say to Hermes or Claude. It can search the web obviously but even if I want context I'm coming in. And I'm like, dude, what are the latest models? If you're like, where's the puck moving?
What's new? What's interesting? And again, guys, if we wanted to build this out and do a lot of extra details, we can do like that's where we're at with this. We can literally now just get a nice short, sharp, crisp overview of everything that's going on by filtering by the models. I can search by different ideas.
There's miniax. It's just profoundly helpful. And we built this using ultra code. Now, there's one thing that we need to understand about Ultra Code. But before I do that, I just want to touch briefly on what we've done here.
So this is an example of an app, an agentic operating system. I share in my community of course, but you'll know many of you know that I'm building my current text to speech startup glider. One question I always get in my community and comment section is around compliance and risk. And I want to just call this out very quickly because it's really important to understand and that's it. If you're building any apps for yourself, for companies, whether it's an OS or whatever it is, you need to lock down your security and compliance.
People always ask me who I'm using. Company that I use is Vanta. We're using these with Glider because these guys offer over 35 different uh security compliance frameworks DDPR um sock um ISO 27,0001 and I've actually found honestly with a lot of like enterprise stuff if you don't have this level of like security they basically don't really it's almost like you don't exist for many of them unless you've got the level of security and detail obviously everyone needs different things. The reason we use these guys these guys basically help you automate your security and compliance. Obviously, there's like a million different companies.
We use these ones. I actually spoke to them and said, "Dude, I do a lot on apps. I should I'd actually like to talk about you on the channel." And they said, "Cool. They'd be happy to sponsor it." So, I want to give them a shout out in the video because we find it so important in building beautiful glider. And this itself does lead us very nicely onto what we're talking about with the product, right?
The fact that trust is the product. We've covered some of these different aspects here. And it leads us nicely onto this idea of when we're actually using ultra cut. Okay? So, essentially, we can do it surgically.
So, we use it for one turn or we can have it for an entire session. Now, bear in mind, Ultra Code is going to burn tokens like no other. So, you don't want to be using Ultra Code 24/7. You're basically going to burn through your entire credits. You need to be surgical and apply it at the correct moment.
And I'm going to explain what some of the use cases are in a sec. But on the surgical option, basically, you drop the keyword in and a single request, just say ultra code, and it will basically go ahead and do it. And after that, you're back to normal. Or you can do ultra code on which is their max mode. Standing default for the session.
Every substantial task authors and runs will work for by default. Maximum thoroughess token cost stops being the constraint. Session only. It resets when you exit. But till jar here is ultra is a power tool.
Leave it off for the conversation replies and trivial edits or you'll pay for orchestration that you quite simply didn't need. Set another way. It's like firing a bazooka when a hammer very easily would have done the job for us. So, if you're looking for a realistic split when you're doing big builds, you know, probably no more than 20% of the time will I be using ultra code. Um, essentially, you know, Ultra Code earns its keep on research audits, multi-perspective work, and separatable systems.
Anywhere where you don't know the answers shape. Good way to think about it, other 80% is back to normal using everything else. Um, people will tell you, and I've seen it online as well, that my wallet is crying watching the video. So, it's very powerful, but you don't want to be using it all the time for things where it's not required. And a few decision makers to make it easy for you.
First of all, you can get four to 7x the token build with this sort of stuff cuz parallel fan out burns tokens fast. You got multiple agents working on it. If I don't need a team to figure out the question, we don't use it. Spin up eats small tasks. It does for anything in the 5 minutes.
Setup plus verification costs more minute saves. And you can't parallelize a ch. In other words, if I have agent M that's got to pass something to agent B, then pass something to agent C, I could have a billion parallel agents, it doesn't actually make the process any faster because I still need A then B and C. So, really important to bear in mind. Now, when do we go and use it?
When we're doing research, audits, multi-perspective reviews. Let's say that you're facing a decision in your life or your business and you would quite like to have a broad variety of perspectives on something. Perfect for that. Like, hey, where should I live? Should I buy a Springer Spinal or Caucus Spaniel?
You know, life's very important questions. Adversarily verifying any claims. Building a separatable system you can reuse and discover where you don't know the shape of something. And we don't want to use it for single file refactors, normal feature builds, bug fixes, anything with strict step-by-step dependencies and quick sub fiveminute tasks. So use it for those tasks and you'll absolutely crush it.
And so if you look at what I developed here for example, if you using it strictly and again time was a constraint and cost was something you aware of, you'd probably use to think about hey spin up multiple agents to debate what would be the most, you know, useful set of functionalities to help a user make decisions about models based on different things and they could debate it. But for example, just building it, you'd probably do something more straightforward. So you can see it's got this band of incredible use cases where it is absolutely awesome. But the big problem here is that building with Ultra Code is just one of the key concepts that you need to understand when using Claude. So, the next thing that we have to do is learn all of the Claude superpowers to take our business to the next level, which we're going to do in this video right