This makes Claude Free Forever

summarized

TLDR

A GitHub repo with over 50,000 stars aggregates free API keys from 50 providers, giving Claude Code users 1.3 billion tokens per month at no cost via Nvidia's free API. The system automatically reroutes to another free model when one tier runs dry, but the trade-off is lower performance and potential hallucinations. It's a useful hack for extending Claude Code usage on grunt work, but frontier models are still recommended for critical tasks. The real story is that this makes Claude Code effectively free for heavy users, provided they accept the performance trade-offs and monitor for model degradation.

Key points

The GitHub repo aggregates free API keys from 50 providers to offer 1.3 billion tokens per month.

Nvidia provides a free API key without requiring a credit card.

The system automatically reroutes to another free model when one tier runs dry.

Free models are less capable and may hallucinate, requiring extra verification.

The presenter recommends using frontier models for 20% of tasks and free models for 80%.

The repo has over 50,000 stars and is rapidly growing in popularity.

Techniques

  • API key aggregation
  • automatic model routing
  • MCP server integration
Transcript (captions)

0:00 Imagine if you could use Claude code 100% for free. The truth is that Claude is expensive and once you hit your limits, you are stopped dead in your tracks. However, a brand new repo just

0:12 came out that gives you 1.3 billion tokens to use inside the Claw Code harness every month for $0. And in this video, I'll show you exactly how to set it up in one prompt so you can use

0:26 unlimited claw code harness, build more, save time, and get light years ahead of everybody else. So, if you haven't already, grab that beautiful coffee and let's dive straight in. So, this is an

0:37 actual new system. I'm going to cover the repo and why this is a better way to do it. So one thing to understand with this is that effectively we can use claw code harness for free using Nvidia and

0:50 this incredible GitHub repo. But first of all you have to understand this if you're really going to get the benefits out of it. Even my dog agrees he's barking in the background. So when it

1:00 comes to performance it is a combination of two things. It there is the model. So that could be Claude Opus 4.8 Fable 5 insert the blank and then we have the harness. The harness is the sets of

1:12 rules. It's where you build. It's all the prompts around it. Classically speaking, it's this. It's using inside the Claude code harness with all of your skills and all of the other beautiful

1:22 things that Anthropic has put behind it. But you only get so much usage of these models. So, what happens when you run out or what if you want to spend over a billion tokens? Well, you can do that

1:33 for free with this system. And the way that I built this skill that I'm going to give you down below for free, it'll be the second link in the description. You just give it one command, you toggle

1:41 it off, and you go from paying all your subscriptions down to $0. And when you make it to the end of the video, you are going to have such a ease of life hack, and it is going to save you so many

1:50 dollars. Now, we're going to be using Nvidia, which is incredible. And you're going to get a 100% free API key. You do not need to put down any debit card information. It is free. I will explain

2:00 exactly the business model so you can understand how that works, but it's completely free. and you get access to 1.3 billion tokens from a combination of 50 different providers. Now you might be

2:10 wondering why are we doing this versus open router. Basically open router obviously you need your card which I think is fine. Obviously they should be making money. You get free models but

2:18 you can hit a daily wall and it'll cap you out and you can get rate limited and some have accused of just getting the leftovers sometimes but I do love open rout

2:28 start so it's very easy to get signed up. Um there is basically never a wall with you're going to hit with Nvidia and you get Frontier open models and it is 100 100% for free and they do this

2:39 because hopefully one day you go and actually rent from Nvidia because you love the server so much. So with that in mind, let's get started and get this downloaded. So now we've done that,

2:48 let's go ahead and get this whole thing running. So I want you to click the link down below. Once you've done that, you're going to see this here. This will explain exactly what it is, what the

2:55 problem is, how it works, and how we set up. And all you're going to do is come down, copy this, and then head over to Claude. Just drop that in, hit enter, and this will do the rest. So, we're

3:04 going to come down here and click yes, run it recommended. Now, this is using this GitHub repo, the one that's got over 50,000 stars and is blowing up right now. Then, once that is done, it

3:13 will open up this window on your own computer. And all we're going to do now is give it a connected account. And we're going to be doing it with Nvidia, which is going to be incredible. Then,

3:21 click on configure. It brings you down. Then, head over to this website. I'll put a link down below. It's build.envidia.com. You're going to go ahead and sign in and

3:30 create an account. Now, you're limited up to 40 requests per minute, which means you can hit the API essentially once every 1.5 seconds. That's a lot of requests, guys. I ask a lot of

3:40 questions. I don't ask that many questions. In other words, it's very, very generous. So, come over and click on API keys. And by the way, if you are new here, I'm Jack. I built and sold my

3:49 last tech startup with a gazillion customers. Now, I'm building my own AI startup. And on the channel, I share here the stuff that actually works. So come over and click on generate API key.

3:58 Give it a name. So I'm going to call this one free claw code expiration for 12 months. And then click on generate key. And then you're going to have the API key. You're just going to copy this

4:06 like so. Come back over here and just drop it in right there. Then once you've done that, click apply in the bottom right hand corner. And now that is fully saved. Now basically you can just go

4:14 ahead and enter in all the API keys from all of the different providers. So we're going to have open rout client pass. And obviously token limits change day by day and month by month sometimes. So once

4:25 you add in all the API keys, this does currently give you access to 1.3 billion tokens per month. So this is the idea. We've got all the tokens to all these different models. You can add them all

4:35 in and then it will pull through and pull a different model based on various different kind of like standards. So let's say I want to add in an open router API key. Awesome. Let's go ahead

4:44 and grab that. So I come over here, I click on new key for example, give it a name like claude free. And basically guys, we're just creating these free accounts from all these different

4:52 locations which is cool. Uh and effectively once you've done that you can just provide it to the model itself and use that and then once you've got it you can come down and throw it in here

4:59 if you want to configure throw it in and repeat the process. Now in reality you wouldn't do all 4 to2 think of this as like the key ring and then each of these API keys are the keys. So I would add in

5:10 personally Nvidia open routter is a classic Gemini Gro GitHub is a good one as well. And by the way, if this all sounds like I'm speaking design Spanish and you want to get even more out of

5:20 claw code AI systems, I'm going to put a link down below for the full Claude code masterass that will take you from foundation setups, building websites, skills and content I have never shared

5:29 on YouTube, including full access to the agentic operating systems, all of these crazy memory systems, and of course the design operating system. A gargantillion incredible things. I'll put a link below

5:41 if you want to level up. Don't tell your competitors though. it will give you a very unfair advantage. So now what have we done? We've effectively gone ahead and we've now given it all the keys we

5:50 need. So now we can head back over to Claude and actually use this really cool feature. So to go ahead and run this, all we're going to do is command spacebar and we're going to open up the

5:58 terminal like so. And then all we're going to do in the terminal guys is type in this command. And again, this will be available in the notion doc just FCC claude. And this right here is free

6:08 claude code which is awesome. Yes, I trust this folder. We can go ahead and now we are literally using our new free models. So to test it, we can say something like hey say pong. Now when

6:18 you ask this model what is it? It will tell you Opus 5. That's because it's running the tick. It will also show you you spent 48 cents. That's not true. It's basically actually connected to

6:30 Neatron via Nvidia and it's got the full cycle of different free models it can access. And just like in cord code you can do for/mcp and access all of your MCP servers. Now, before I show you one

6:41 of the really cool use cases with this, you need to understand why this is actually so cool. So, what's really cool is if a tier runs dry, it will automatically reroute it, which is great

6:51 cuz one of the biggest limitations with free models, so to speak, is that they can be on for a very short period of time. You'll be coding to your heart's content and then it's like game over.

7:02 Sorry, you can't use it. The cool thing about this system, we plug all the different free stuff in as and when it appears and then just automatically tracks you, meaning you get kind of

7:10 hassle-free coding on the freer models. And they've also included this really cool section that lets you connect any of your MCPs that you're already using inside Claude over to this new free

7:21 claude code. So, say for example, I wanted to go ahead and connect Zapia, I could basically copy this, come over here, type this in like so into the chat window, bring it over here, and all do

7:31 is enter in my URL. Now of all the ones that I would use because these models guys are not as powerful as the Frontier obviously you're trading on performance cost and speed right the thing about

7:43 Zapia is I can actually I it's my authentication layer I can use it across my Aentic operating system I can use it across cloud code I can use it everywhere to give all the specific

7:53 access I need and it has stuff that I just can't access in other places like the school API for example so what I like to do is come down and create a new MCP P server specifically for my new

8:04 fruit free cloud code sections. I'll put a link down below so you can grab this. I come over and I literally guys generate the token, copy it, and then I can use that inside of free code to do

8:15 things. So it has the ability to access different things. But one of the reasons I really like this with Zapia is because I can literally in the app say exactly what I do and do not want it to have

8:26 access to. For example, I can specifically remove certain requirements. Yes, you can also do this in claude, but it's very handy for me to have multiple different levels of

8:35 authentication that I like and I can use the same thing in multiple different apps and connections is one thing, but if you're not using it properly, you're never going to get the best results. So,

8:45 think about it like this. You do want the Frontier models for work. Like, you just genuinely do. And it is a what we'd call a false economy to think that you want to do free for everything. A tiny

8:57 bit of money can go a very long way. So, you never want to go all three. You want to have Frontier for your best best tasks. And if you're on the $200 plan, dude, rock on. You're kicking butt. It's

9:08 actually unlikely you're going to need this. If you watch this video on screen, which I highly, highly recommend that you do, it's going to blow your mind in terms of how it can actually increase

9:16 your token usage. But typically speaking, if you really want to crush the free model gauntlet, you want to do 80% uh on the cheaper models. So, that could be Haiku or this free tier. And

9:27 then really save the 20% for the front tier. Now, what's really important to bear in mind is the caveat. You are trading for cost. Since we're bringing the cost down, your performance is going

9:36 to come down. But also, sometimes models, guys, they will lie to us. They will hallucinate. So, just bear that in mind with your free models. And you need to be extra careful to doublech check,

9:46 spin up sub agents, and cross reference for any work you're doing on free. And where possible, have that reviewed by more intelligent frontier level model. Now, also, I want to equip you so you

9:56 walk into this with both eyes open. free tiers, they do shrink. So, for example, do you remember Ox Alpha? Fantastic, right? That's great. They open, some of the best labs open things up for free,

10:08 but they will take them off the market a time, which is one of the reasons why this kind of keychain repo is awesome because we have so many you you're less interrupted and things run a bit

10:16 smoother after like 10 minutes of setup. Also, the amount of free they give you can change from time to time, which is awesome. Free models perform, I'd say, the worst on taste. So, if you're

10:26 looking to design your next beautiful piece of art, the next Mona Lisa, maybe we don't bring in the free models for that, but it can be incredible for grunt work running on your local AI labs 24/7.

10:38 Just have a load of chucking on infinitely to do that, which is awesome. And bear in mind that it will sometimes say it is Opus 5. When we pull the Scooby-Doo mask off, we can tell it is

10:47 in fact something completely different. So, it's just reading from the surface level data it's got. It's not the model that's actually being used. And you can always check that by just asking Claude.

10:56 Now, using the Claude code harness is one thing, but I did one prompt and I actually tripled my usage inside Claude. It is one of the best systems I've seen. And so the next thing that we need to do

11:07 is get you running on that system in one prompt, which we can do in this video right

Frontier News · by Hyperjump Technology