Enjoying this issue?
Get tomorrow's AI & engineering digest in your inbox — hand-picked, summarized, and always spam-free.
TLDR
DeepSeek V4.1 is roughly 88x cheaper than GPT Astra and performs competitively on key coding and automation benchmarks, often matching or beating GPT 5.6. Its real value is as a workhorse model for structured tasks like data analysis and existing design amendments, while GPT Astra remains superior for frontier reasoning, complex PDFs, and design from scratch. The practical takeaway is to delegate routine or high-volume work to DeepSeek and reserve the expensive frontier model for genuinely hard problems.
Key points
DeepSeek V4.1 scored 54.8% on Automation Bench, up from 37.7% in V4.
DeepSeek V4.1 scored 74.2% on SWE-bench, beating GPT 5.6 by about 1%.
DeepSeek V4.1 scored 90.6% on Terminal Bench, outperforming both Opus 5 and GPT 5.6.
DeepSeek V4.1 costs $1.75/million input tokens versus GPT Astra at $60 and GPT 5.6 at $24.
The model uses more tokens per task than Astra, so total cost comparison must account for token count.
Tools mentioned
Techniques
- Model delegation via Codeex harness
- API key usage directly from DeepSeek to avoid rate limiting
- Connecting external services via MCP CLI for model tool use
- Using a workhorse model for high-volume, lower-complexity tasks
Stop scrolling. Start reading smarter.
Receive the day's most important AI & engineering updates in one concise email. No spam.
Transcript (captions)
The brand new Deep Seek model just dropped and it is 88 times cheaper than GPT Astra and as powerful as GPT 5.6 Soul. And in this video, I'm going to show you after days of testing this
model exactly when you need to be using this and what it's excellent at, what you shouldn't be using it for, and then most importantly, how we can actually combine this with GPT Astro, the world's
best model to get incredible performance for a lot lot cheaper. Grab that coffee. Let's go straight in and begin with Deep Astra. So, we can use the brand new Deepseek model actually inside of our
codeex framework. Now, many of you will know that you can use something called a Deepseek harness. This is just DeepSeek's version of Claude Code essentially where we can build and code
with them. But actually, if you come down and grab this free GitHub repo, I'll put a link down below. What we can do is literally come down here, copy this, and give it to Codeex. Then simply
say I would like you to open up deepseeek inside the terminal within the chat GPT codeex framework. Of course this explains everything inside it and you paste that link and it will do that
for you. Now this enables Astra to actually delegate work to this deepseek model using your existing files uh all your existing tools and everything that you have access to already uh within
chat GBT which is incredible. Now why do we actually care about this? cuz this is genuinely big moves and this is a model you can actually host locally as well. The bottom line is that Deep Seek got
way stronger like ridiculously. Look at this for example. This is automation bench. Automation Bench covers how well an agent completes multi-step automation tasks. That blue thing on the right is
DeepSeek. Look at how it compares to Opus 5. Look at how it compares to soil 5.6. And just to show you how much this has stepped up from V4. If I show you V4, look at this. We went from 37.7
to 54.8. It's crazy. Then I can show you Deep Suite V.1. This covers how well it solves software engineering tasks in the codebase. 74.2%. Literally um defeating Opus 5 and
defeating by about over 1% solve 5.6. And then we have Terminal Bench as well, Terminal Bench 2.1, which is how well it completes task using a terminal and its tools. And it's at 90.6%.
It is crushing it. It's multimodal. It can see it can do many different things. Now, what does it mean for you? For working web apps, it is exceptional. Again, better than 5.6 and its
performance on business workflows is pretty incredible actually. And then we've got long documents and I'm going to show you where this excels, where it falls behind, and how these benchmarks
hold up in practice. And I'm going to do that uh essentially across three levels. We're going to do design tasks, we're going to do real business life workflows, and then how it can actually
help with the agentic operating systems. And let's begin with level one, which is getting deepseek to design. If you're new here, by the way, I'm Jack. I built and sold my last tech startup like a
gazillion customers. Now, I'm building my own air startups and I share on this channel the stuff that really works. If you like that, feel free to drop a subscribe so you get more updates like
this on the cutting edge. Now, deep itself is very, very interesting. Two big ways that we can use this. One is with the deep astro, which I'll communicate on, and the other one that
we've got, as I explained previously here, is we have the deepsek harness, which you can grab from this repo here. Now, first thing that you need to do when you're testing this, I have found
is you want to head over to the Deepseek platform website. Another reason for this, guys, is I found when I was using Open Routter, I'd occasionally get rate limited. Once I got my API key directly
from Deepseek, which you can actually take that key over to Open anyway, I found that I just wasn't getting rate limited and it was just a way smoother experience. So, I just deposited like 20
whatever dollars you want to in here. You can easily create an API key and then just use that instead. And I actually run this across dozens of tests. But what I want to really test is
it ability to build beautiful presentations and images all orchestrated by GPT Astro. Now to do this we need to go ahead and connect Higfield. Higfield if you don't know
enables us to build and generate beautiful images and videos. As you can see many of the images I even used for this presentation I actually just built in Higsfield. You can use literally the
state-of-the-art models the brand new GPT image models. It is exceptionally userfriendly and I can turn everything into videos. It's incredible. So we want to give this skill to DeepSeek so it can
generate these images for us and that we can actually test it. Now to connect this just click on MCP at the top and then click on CLI on the right hand side and I'll put a link down below so you
can just see over here and effectively what we can do is give this to DeepSeek or Chat GBT and basically connect our computer to Hickfield and that means that DeepS will be able to
programmatically just connect to Hickfield. I'm into Deepseek. I'll just say hey there could you connect to the Hickfield CLI. going to come down here and enter that there. And what's really
interesting guys is if you aren't connected to this model yet, you can actually say to GPT Astra, hey, go ahead and connect me to this. And it will do that for you. And as you can see here,
guys, it tells me the CLI is already installed globally. So now I can go ahead and use this to generate images and videos. Now, I did so many tests, but the first thing I wanted to really
test out with here was his ability to go ahead and generate for me a presentation, which it went ahead and did. As you can see, this video here again, it just created with Higfield.
This was 100% by the way just deepseek deepseek v4.1 frontier reasoning open to everybody million token context the weights are open now I had to give it some reference designs to this I
basically said hey I like the look of this website kind of turned this into a design here and this is what it went ahead and actually delivered now let's go through and have a look we have slide
one we have twice the reasoning half the weight three things that we rebuilt I think it's a pretty gorgeous you know gorgeous animations to be fair again no astster involved in this whatsoever this
was just Deepseek. Very, very cool. As you can see, it's got some gorgeous animations. Uh, draft, check, revise. This is basically a presentation on DeepSeek, and it says it is now
available today. Now, I had to give this probably three to four revisions back. It didn't do the best job in the world on its first prompt on the design, but I think it shows that it is capable of
following instructions when you prompt it correctly. So it can probably do more than you think on a design point of view, but you just actually have to congeal it and spend a little bit of
time going back and forth, which is why combining it with Astra is actually very powerful and very interesting, which I'm going to cover in this video. The other thing that I did is I had this website
here that I built with GPT Astra. Now, one interesting thing that we can do here, guys, is we can have Astra design uh and leverage beautiful websites and then what we can do once those websites
are built and developed is pass over to Deepseek. So let's say that we had a template that we were building for our clients. I actually gave DeepSeek this individual website and I said, "Hey, go
ahead and change the text such that it's for a company called Bricks and they are kind of like an AI infra company and it came back with this." And it went through it basically changed all the
texts and you can see it went down here and it changed all the logos. Now look at this guys. It hasn't actually designed these logos itself. But can you see how it's managed to keep in line
with all of the design systems and design architecture that we had initially? So, it's been able to maintain that for me, which is really good. Now, if I'd have asked Astra to do
this, you know, I might have had to get a second mortgage just to kind of pay the bill. But with this, you can see it's been able to maintain that. It's changed the text. It's made amendments
to existing design systems. Come down here, and as you can see, it's even changed the images and everything else. The other cool out here, though, is it is not incredible at building straight
from scratch. It's still better than a lot of models were 6 months ago, but when you're using this, you want to essentially have Astra delegate to Deepseek if you're looking at doing
amendments or those existing design systems from scratch. It's still capable, but it's not the best model in the world yet. It's not the chief design engineer. So, very interesting. We can
have Astra leveraging Deepseek to go ahead and build some interesting things, but from scratch, I wouldn't necessarily just use Deepseek. It's good to support and build on, but it does lead us on to
a very interesting one here, which is essentially Deepseek's biggest strength, which is its capability to actually look at data. Now, I'm going to show you exactly how well this performed. And if
this sounds like I'm speaking Spanish, I'm going to put a link down below for my full GPT Astro Masterass, which is dropping, which will take you through AI foundations, productivity and live,
business and scaling systems, marketing, creation, automations, essentially things you can actually implement with GPT6 Astra that will move you ahead, that will save you time and help you
absolutely crush it. And you also get immediate access to my entire Agentic operating system, all my different memory systems, my design systems, my business systems, everything else. I'll
put a link down below so you can join that and get light years ahead. Now, I had both DeepSeek and Astra go through over 50,000 rows of data to answer basically 17 questions and both were
able to do so with 100% accuracy. It had it look at over 45,000 paid orders. Now, if I come down to the reports itself, you can actually have a look at this. We have one from Atlas, excuse me, one from
GPT and then one from uh Deepseek itself. And you can see we can actually see what all the revenue looks like, where the revenue lands, all the questions, and it's gone ahead and built
this. And then we have the same thing as well from Deepseek. Again, design's not quite at Astro level by itself, but you can see what it's been able to do. And this aligns very nicely with what the
benchmarks reporting in terms of its performance and ability to do these automation tasks. And so now we've actually tested there. I wanted to go ahead and do one final test, which
essentially was to actually solve a problem for me in my Aentic operating system. And I wanted to give it an idea here something a little bit difficult to see exactly how it would go ahead and
perform. Now recently I added into the agentic operating system in my community this basically this business OS that lets you look at your overview for all your accounts. You can ask questions.
You could say something like hey um how much money am I expected to receive this month. The idea of this is it helps you run your business better and it's like an operating system that brings all of
your models together. It tells you how much money you're spending. It tells you things like insights about how you can improve. It's really cool and it's got all this gorgeous stuff built into it.
So, this was a new design UI system. It can show you breakdowns and everything. You just plug in your account. It's like Mercury. It's all done locally for you and you get these incredible systems.
Now, what I wanted to do here, and look at this, it's even answered there with a model. I didn't even have to do anything. It's like my command center. What I wanted to do with this is take
this design system and I wanted to go ahead and say, "Look, can you redesign my dashboard in the exact same design system?" This is a large, I'd say, quite complicated product. I've been building
this for some time. I even built like an onboarding. It took me like eight hours to build to help people out. It's very interesting. You can even edit websites in here, edit the copy, do a bazillion
different things. Now, what did Deep Seek actually do? Well, it went ahead and it did it. Now, let me show you what the results were. If I come over to my dashboard, you can see that it's
actually taken the same typography and it's now breaking down this into different sections to make it easier. So instead of having one big scrolling section which was there previously, we
have the overview which is how much money I'm spending on AI every month, how much money it's saving me, what my activity is for improvements of how I can improve because it dreams every
night find that for me. It shows me detail on my usage and all the different models I have available and what their performances are. It shows me information on my workspace and my uh
memory systems, my knowledge graphs and how everything connects together. And then even shows me some system information. So, it was able to understand the existing design system
and build that for me. No, I can still get Astra to check it. But Deepseek is the workhorse model. I was super impressed with this. How much did this actually cost, guys? It costs less than
essentially 10 cents. This cost me 9, 9.5, which is crazy. It did it in 7 minutes and 18 seconds and was able to do everything for me. Okay, comparison. If you use a different model for that,
like a more frontier level model like, you know, GPT Astra, that would have been insane if you're doing it on API. And you can see here just some of the differences between this Deepseek model
and Astra 6 and Soul 5.6. Also, the token efficiency is one thing to bear in mind is that DeepS will use more tokens than Astra. 100% it will do. So, Astra, you can't just look at the headline
price because Astra's intelligence per token spend is higher. it. It's just it just gets more done per token. So tokens isn't the only lens that we can use here. Now, you can run this model
locally. To do that, we would estimate that if you're doing an RTX 5090 with this much RAM, you're looking at around about $13,000 plus. Just to be able to run this locally and you'd get around
five tokens per second. But this is a big breakthrough that we can run quantitized versions of these models locally is incredible. Now, what does it actually go ahead and cost? Well, if I
look at input and output, just to share this with you, GBT Astra is $60. GT 5.6 24 bucks and DeepSeek V4.175. That's an off peak and on peak it's about twice as expensive, which is
crazy. So, the idea here is not for us to not use Astra. My whole philosophy here is that Astra is frontier level technology. It is, but there are some tasks that DeepS v4.1 can do for us. And
if you're somebody who's hitting your limits in Astra, you're going beyond your subscription, having the ability for Astra to delegate to Deepseek V4.1 can make a big difference. I mean, I've
even connected to GL. For example, I can sometimes say something like, "Hey, dude. Could you go ahead and tell for me when was Deepseek V4.1 released? Go ahead and ask Deepseek this." So, this
is my transcription tool, right? And I've essentially set up a connection now where I can just ask Deep Seek questions. and I do this, it opens up the harness and I've just asked a
question immediately to Deep Seek V4.1. Such as the power of this harness and this connection. Now I've got it there to ask questions whenever I want to. Now, as you can see, it's now doing some
factchecking and it came out around September 10th, which is incredible. Then let's actually look at some of the drawbacks of this model in terms of where it is not the strongest. So, it is
not the strongest at doing designs from scratch. Number one, if you give it heavy reference materials and you're willing to go back and forth with it, it will get to a great output. But for your
time benefit, it is a false economy. Just use DeepSeek in that capacity. I wouldn't do it. It's okay to amend and build on existing systems. Have it as a workhorse model in the background
because it's pretty freaking incredible. And you can see here how it's borne out in terms of difficult coding. So your biggest coding tasks, I would still be leaning on Astra 6. After that, let's
take a look at complex PDFs. PDFs are famously a little bit of a wacky crazy situation. still capable, not as capable as Astra, so lean on to Astra for anything to do with PDFs expert
questions. Again, it's still up there, but again, it, you know, it's not at the level of SOL 5.6 or Astra 6. And then finally, factual reliability, not the world's most factual reliable model. So,
when it comes to doing fact checks, do not bring in Deepseek V4.1. And so, the way that I would be using Deepseek V4.1 going forward is my workhorse model. I have Astramm which is I use for my
frontier stuff my top level big brain thinking and I also have my claude systems but Deepseek V4.1 is what I would call the workhorse it's the one where basically essentially it's very
intelligent and it's very cheap right which is perfect so be certain number of tasks now where you can handily trust DeepSeek to do that for you and the fact that it's locally hosted is also
incredible so by using Deep Astra you can now actually have Astra say Hey, if you think this is a task for Deepseek, just delegate it to Deepseek and get the information back from it. But it does
take us on to a very important question, and that's the fact that we've got this AGI level intelligence. But what can we actually do with that to move our business forward, to move our life
forward? I show you exactly some of the technologies and capabilities you can do right now.