Enjoying this issue?
Get tomorrow's AI & engineering digest in your inbox — hand-picked, summarized, and always spam-free.
TLDR
A master prompt for Claude, revealed in a YouTube video, audits your session for token waste and claims to find over 95% waste in most cases. The system then applies fixes and behavioral hacks to triple available usage, with the author claiming it effectively buys back 17 pages of context per session. The real value is the structured audit and specific habits—like /clear between jobs, one model per session, and editing over replacing messages—that reduce token burn without reducing output quality.
Key points
The master prompt audits token usage and identified 95.4% waste in the author's own Claude sessions.
The prompt has three phases: audit, fix, and apply, with key habits to keep usage low.
Using /clear between jobs prevents the context from growing unnecessarily.
Switching models mid-session forces a full reprocess of history, wasting tokens.
Edits to messages save tokens compared to sending corrections that add to conversation history.
Tools mentioned
Techniques
- Token auditing and waste identification
- One model per session to avoid reprocessing history
- Batching questions to reduce context re-reads
- Editing messages over replacing to avoid polluting context
- Using text instead of PDFs or images to reduce token consumption
- C100 communication style for clarity and fewer clarifying questions
Stop scrolling. Start reading smarter.
Receive the day's most important AI & engineering updates in one concise email. No spam.
Transcript (captions)
Imagine if you never run out of Claude tokens. I found a new system in Claude and it more than tripled my available usage. [music] And in this video, I'm going to show you exactly how to never
hit a rate limit in Claude again. So, you can build way more, spend less, and get light-years [music] ahead of everybody else. So, if you haven't already,
grab that beautiful coffee and let's dive straight in. Now, I've literally increased my usage by over 3x, enough in fact to read the entire Lord of the Rings to Claude 8,400
times. I'm going to give you the full possible system so you can do the exact same thing. Link will be the second link in the description. But before you do that, we need to understand how we
actually won't never run out of tokens again and I'm going to show you that across five very simple levels, all of them build on each other and by the end of this video, guys, you're going to be
the biggest token saver in the household. Now, let's jump in with level one, which essentially is the master prompt. What I want to do is click the link down below. It's going to take you
through to this page here. And this explains all the basics, how tokens burn. I'm sure you may already know that. Message turn, you know, basically we have a big conversation, you're
repacing the entire conversation, etc., etc. The key bit to understand is you're going to come down and you will literally, guys, going to come down and copy this prompt. Now, this has three
phases to it. The first thing it's going to do is audit. So, it's going to list everything that gets preloaded into a fresh session, your system tools, your MCP servers, how many tools each exposes
your Claude.md. It's going to read your session logs. It's going to measure every Claude.md in scope, the project, the parent folders. And essentially, it's going to flag any file over 5K or
if it totals over 10K, it's going to list skills. It's going to do lots of wonderful and beautiful things. So, let's head over to Claude right now. Now, all you're going to do, guys, is
come over and quite simply just paste that into the description. And the first thing it does is audits. Mine found a staggering 95.4% waste. I actually couldn't believe it
and I use Claude all the time. Of course, it's got three phases it will audit for you. It will fix and apply and then it's going to give you key habits to take so you can keep the your usage
even lower meaning you don't attack or get hit by the token monster which we're fighting. So, we're going to audit it to find out where are you right this very second wasting tokens. Once it's done
that, it's going to go ahead and fix. It's going to find out all the stuff that it can remove to make your entire usage way less. Then once it's done that, it will go ahead and show you
exactly how and it can get you a full breakdown just like this. But there is one thing I really want you to do first. So, if you take a look at this as an example, if I click on this little arrow
here, you can see it shows my contacts window and I'm using 384,000 and it tells you all the stuff that's taking up your precious tokens. Tokens that could otherwise be spent on
brilliant things like finding the best cheesecake in Budapest some would say. Now, what I want you to do is come down over here on the bottom left and you're going to go ahead and click on the plus
and click on plugins. And when you're on here, I want you to click on connectors. And then once you land in the connections section, I want you to grab the biggest axe you can get your hands
on and we're going to chop away any of these connectors that we're not using. We're very quick to add, we're very slow to remove. So, just literally grab them and remove them. And you can do that by
clicking on the thing itself, come down here and just click on disconnect. So, if you're not using Canva, we don't need Canva no more. Sorry Canva, nothing personal. We just don't need you right
now. And obviously in that document as well, I'll also include you this overview for skill. What is very cool about the skill is if you are working with clients, you can actually give them
this as like a bonus, a Claude audit to help you basically potentially triple your usage if not more, maybe less, who knows. Now, what was interesting to me is that 97.8% of every token came from
three sessions I left open which is crazy. It explains exactly what happened. These were three very big sessions obviously and it goes through the clients. It says, "Hey look, you've
got 9,000 tokens back from every single session. That buys back 17 pages of Lord of the Rings before you type a single word in every conversation. It tells you of every book what I'm saving on my book
log hook. It saves me there. And then it goes through with some other kind of very high-value habits. And those habits lead us nicely onto level two, which is now we've actually gone ahead and we've
clans. And Claude will show you how much money in tokens you can save, we now want to change our behavior because you are probably making at least one of these very costly mistakes. What are
those hacks and what do we physically need to change? Number one, here's the first hack I need to do. You're going to do {forward slash} clear in between jobs. If you go ahead and do compact,
it's actually re-reading the entire conversation. So, we're going to open up a new fresh tab, and when it's completed that task, we're going to come down literally {forward slash} clear,
badabing badaboom, like this. Clears everything, and now you can begin and carry on any conversation that you'd like. And the idea of this being that the longer the conversation goes on, the
greater the context window. And so, it's like 500,000 words, 600,000 words that's going to basically consistently and continuously has to be processed. And that's money
that you're paying when realistically, that first 500,000 words is not relevant any longer to what we're doing. Hack number two, this is a big one. This surprised me when I got under the skin
of this. This idea was one model, one session. So, what is a big no-no is if you're having a session in Claude, we do not change the model halfway through. So, for example, if we're using Fable 5
in this chat window to say, "Hey, how much protein in 100 g of Greek yogurt?" which I like to use Fable 5 extra for. I'm just joking, by the way. Never do this. This is for demonstration purposes
only. It's an ongoing joke I have. We would not then go ahead and switch to Opus. One session, one model. Because when you switch the model, you basically reprocess your entire history. That's
less tokens that we could be spending on design or learning about cheesecake, which is not a great thing. Hack number three is to batch your questions. So, basically, instead of sending a message,
then another message, then another message, hit like if you've got the dictation tool like Glider, for example, all you're going to do So, for example, you can do continuous
mode, and then I can ask a question, get my full stream of consciousness out, and just think about everything that you wanted. Because if not, every time you ask a question, the AI rereads the
entire chat history. By just this one small habit, you can save so many tokens. It is ridiculous. Hack four is about editing over replacing. Now, let's say that you are talking to Claude. Now,
most people would say something like, "Hey there, who is the tallest man that ever lived, and what is the best flavor cheesecake in the world?" Come over here, paste that one there, and let
Claude respond. Now, most people say, "Damn, I made a mistake. I meant who is the tallest woman." And they'll come down and say, "Oh, I didn't mean man, I meant woman." And they will just paste
that in. But, there's an issue. Is it it increases the length of the conversation. And not only that, we now have the correct response and the incorrect response. So, what you can
actually do instead is literally come back over here, click this pencil icon, and we can just literally change this to "woman." Hit save, and then it will replay the message without polluting the
context. This works in Cohere chat, but it doesn't always appear in code release notes for me, so just bear that in mind in where you are using this. And then, hack five is to use text instead of a
PDF if you can. It also applies for images. Sometimes, you just need to, but if you're going to be continually referencing it, just try to preprocess it first, and just get it into plain
text so you don't have to reuse it all the time. Again, I put this down below in the actual skill for you so you can see it clearly in between jobs. It's going to save you a lot of tokens.
Do not change the effort, either. So, even if you're using Fable five, and we want to find anything out, and you're thinking, "You know what? High is too high. I'm going to switch the effort
level." Don't do that. Keep both the effort level and the model selector exactly the same for the entire conversation. That will save you on your cash. Again, batch your questions, edit
the message, don't correct it where possible, and try to go for text over PDFs and images. And this leads on to level three. Now, level three is going to be an absolute godsend cuz have you
ever found yourself having a message back from Claude and thinking, "I don't understand a single thing it just said to me." Well, turns out that this is called the verbosity
problem, and it is very fixable by using this standard communication style, which is C100, and there is a perfect skill to use. And by the way, if this all sounds like I'm speaking Spanish, I'm going to
pull them down below for the full Claude code masterclass that will take you through foundations. It will take you through building beautiful websites, power features, memory systems, Hermes
agent apps, building anything, stuff I have never shared on the channel, as well as the full Claude code design system, as well as the dashboard that actually helps you save money,
dynamically dreams, does a bazillion other things, all down below so you can level up very quickly. Now, the skill, and I actually shared this with my community, funnily enough, about a week
or so ago, is essentially something you can give to Claude. It will basically break things down like it's 5 years old, and just look at how much easier this is to read. The problem, like you're five.
The fix, one sentence. Wait, what is that? Why is this so important? Well, the clearer Claude is, the less clarifying questions that you have to ask, so your conversations are shorter.
And also, the less brainpower that you spend. I'm thinking about your tokens over here, not just Claude's tokens. How about your token the token cost of using a human being? And all you're going to
literally do, guys, is copy this up, turn it into a skill by coming over here, you can paste it in, and then you can say, "Hey, turn this into a skill." And then once you've done that, you can
do forward slash, and however you named it, hit that text, and then it will speak to you like an actual human. And if you want to feel the difference, this document itself was written by Claude in
that style. Again, if you think about it, if you're landing a plane, the last thing you want is verbosity and a Bronte-esque style of communication. I want simple sentences, no unnecessary
super- basically superfluous language. Now, before we can understand the capabilities of level four, we have to understand what Claude cannot do. And this brings us nicely onto the sponsor
of today's video, GenSpark. Now, here's the thing. Claude is great, but it is one tool. One of the really cool things about GenSpark that you can do, and I've been playing around with this recently,
is you can use any model. Let me show you exactly what I mean. So, instead of having an image description, for example, what I can do is come down here to, I don't know, AI image chat, for
example, and I can say, "Hey, build me a um 16x9 photo of a beautiful English Springer Spaniel drinking a protein shake." So, a lot of people building vibe cutting things and up integrating
everything, GenSpark 6.0 has become this really cool place where it's got all the image gen models, the video gen models, and you can effectively tag around and play with multiple different models in
the same environment. And would you believe it, guys? I got that was actually pretty quick, to be fair. We got a beautiful Springer Spaniel rising ground. Believe it or not, that is a
true snapshot from my kitchen on a daily basis. But you can have a lot of really good fun with this. Like, one of the really cool use cases that I like, obviously, you've got AI sheets. So, for
example, if I'm trying to build out a tracker for my business, I can say, "Hey, create for me a bit of a detailed breakdown. Give me um income, ROI, profits, expenses, anything you think
you think might be relevant, and just give me a green headings and make it look beautiful." So, you can just open these things up effortlessly, and I find it really helpful to have the image on
the right-hand side. But what's really cool, for example, if you come back over here, let's take a quick look at AI slides. A lot of free users, they get free credits when they sign up, which is
awesome. But let's give it something. Let's say that you just finished the meeting, and you're thinking, "Damn, I need to get a presentation and a proposal up." You can literally drop the
transcript inside GenSpark. So, I've just uploaded a sample one. Of course, I can click on creative mode, and I might say, "Hey there, go ahead and build for me a beautiful pitch presentation off
the back of this meeting, just for example's sake. Just give me an example pricing, services, and then keep it to three to four pages maximum. Thank you. We can hit send on that, and you can
just see the quality of different graphics it's built. So, it's a kind of all-in-one platform that you can just play around with and essentially build anything. Now, let's come back to ask it
clarifying questions. So, I'm going to say, "That sounds good." I'll say, "Um go ahead and generate it. I just want to get a sense of the kind of thing that you design for this example." And then
Jasper comes back and gives us these beautiful slides, which are fantastic. And look at this. It's not bad. It's done all of those directly from it. Two modes, we can actually do this in
creative, in which case you get these beautiful images. You can then go ahead actually and export those, as you can see, inside Microsoft PowerPoint. So, I'm going to say you can download that.
If you want to edit it, you want to do it in professional mode, or you can get these great images. And then as you can see, I've now got this executive financial dashboard. So, if I was
working on something with my team, I could be like, "Hey guys, go and build this out." And my team now don't need to be experts at this stuff. They can see it visually. They can ask it questions
and do anything that they like. I'll put a link below, so go check it out and have a good time with them. And then speaking of having a good time, guys, we're going to talk a little bit about
level four. Now, level four itself is about routing, routing the work. Because for some of us, we only ever really use one type of model. In reality, we're leaving too much value on the table.
What do I mean by that? Well, I'm talking about using the right job um for the right brain. Now, the way I'd want you to think about this is Haiku is your grunt worker, Sonnet is the builder,
general questions, Opus can be your main session driver. And then the two really big interesting ones here actually are Fable is the artist. So, Fable is and still the most capable model, and we
want to use this where quality it it's really important that we get the decision correct. If you're on the $200 plan, I would use this much more liberally than on the $20 plan. But
Fable is the one that you want to use your big decisions. And if you want a creative premium, it is still better than Opus. It really is. And then if you actually do have a ChatGPT subscription,
the $20 ChatGPT subscription will get you very far with stuff and you can tag in Codex on the certain things like checking your work inside Claude. That is probably its biggest feature guys
because if you get it wrong, you may have to rebuild the entire thing. It goes down crazy rabbit holes. Have Codex just say, "Hey, I want you to use the Codex CLI to double-check this work." I
can tell you to no end the times I've had Claude tell me everything's fine, there's no issues. Check it with Codex. Codex, oh my gosh, found the biggest issue ever. Check it with Codex guys,
super important. So by using the correct model for the correct task, you're not using a bulldozer to open up a fridge door. That's the bottom line. Now once you've nailed down the model, we have to
nail down one of the biggest token burners in the world. This is going to put dollars in your pocket. And that's the ability when we're talking to codebases because codebases and GitHub
repos, fancy speak for place we put files online, burn tokens like the Joker out of The Dark Knight. And essentially, what happens with this is we want to use one really cool tool called GraphiPy.
Now GraphiPy has over 100,000 GitHub stars. You literally grab this link, you can come over and chat to Claude or if you've got an agentic operating system like me, all I do is I just add a
project, I throw it in, and then I can literally chat to it in my OS and have a great time. And effectively, what it does is it finds the relationship between the code. So it's like giving
your Claude a map. So rather than having to go through every single line of the map, imagine like a classic Command & Conquer style, it has a map and it understands the relationship between
things. So it makes it way easy to query and the cost to find out the information is significantly less. So you always want to be using GraphiPy when it's learning a new codebase, when it's
understanding or you're asking questions, this will save you so much in tokens. And by the way, if you are new here, I'm Jack. I built a small elastic startup with like a gazillion customers.
Now, I'm building my own startup. So, on this channel, I share with you the stuff that actually works. And one of the big things that works is what's in level five. And level five is around not just
talking to Claude, but designing in it. Now, I love designing in Claude. There's a couple of things that we can do. Number one, when you get a design style that you like, I want you to save it as
a skill so you don't burn loads of money exploring. Couple hacks that you need to know. Every screenshot is 5,000 tokens. So, try not to keep dropping screenshots in. Make it more like this. Make it more
like this. Try your best to actually explain what it physically looks like. And if you ask that 20 times, 100,000 tokens consistently burning. So, just be judicious with the screenshots that you
actually give it. And try to get it right on the first prompt just because performance slants down over time. Once you get that winner, I want you to go ahead and codify it. And that will only
take 2 kilobyte specs and obviously save you a ton on tokens. And crucially, I want you to spin up sub agent critics using the design loop skill. That does burn slightly more tokens, but I found
the output is incredible. And because once you get that beautiful output, you can then codify it and then just use that as many times as you like. And if you've never heard of a design loop,
I'll put a link to video on screen covering that in brilliant detail. Now, saving money on tokens is one thing, but you can't save your way to greatness. So, we need to know what we can actually
do with this incredible technology, which we're going to cover together in this beautiful video right here.