Agents are slower than LLMs?

summarized

TLDR

Agents are slower than LLMs primarily because they rely on external tool calls (e.g., fetching web pages, making API calls) that can take seconds to minutes, creating a lopsided time horizon where tool execution dominates. This latency propagates down the stack, causing expensive GPUs to idle while waiting for tool results, which motivates infrastructure innovations like disaggregated inference and external KV cache storage. The video also notes that model quality, harness design, and hardware budgeting (including CPU, storage, and networking) all affect agent speed.

Key points

  • Agents are slower than pure LLMs because they must execute external tool calls that take seconds to minutes, dwarfing the time for token generation.
  • Tool execution latency forces GPUs to hold cache and idle, wasting expensive compute resources (rental costs $2–$7 per GPU per hour).
  • Disaggregated inference separates batch processing (compute-heavy) from token generation (bandwidth-heavy), allowing external cache storage to reduce GPU idle time.
  • Chinese researchers' 'dual path' method optimizes data movement between compute and storage to mitigate overhead from cache transfers.
  • Anthropic's prompt caching (5-minute default, up to 1-hour with higher pricing) likely uses similar infrastructure techniques for agentic workloads.
  • Locally run agents don't benefit from disaggregated inference or cache offloading because a single user's GPU is fully dedicated.
  • Agent speed also depends on the harness (e.g., Claude Code vs. Codex) and the model's tool-calling accuracy, measured by benchmarks like BFCL.
  • Agentic workloads shift data center power budgets toward CPU, storage, and networking, not just GPU; home builders must balance PSU capacity across all components.

Tools mentioned

Techniques

  • disaggregated inference
  • external KV cache storage
  • dual path data movement optimization
  • prompt caching
  • ablation studies on agent harnesses
  • parallel agent orchestration
Transcript (captions)
Why is it that agents are much slower than LLM's? Clearly, LLM's can give answers at an incredible speed, like this 1,000 tokens per second chats from Cerebras on certain models. But, when we use an agent, we just can't seem to get that kind of speed and everything just seems to take forever to get done. But, why? Are agents just inherently slower than LLM's? When we give an agent a simple prompt, tokens are generated from bottom up from energy that powers the GPU chips and GPU chips that are run on data center infrastructure and data centers that run LLM's like Opus and finally LLM's that generate output tokens to our application that we see on our screen. But, unlike pure LLM's, agents are inherently different because agents typically use tools to actually engage with the external world. So, even if the layers below generate tokens at a decent speed of 300 tokens per second, to make a tool call request, the actual time it takes to execute the tool can take anywhere from few seconds to minutes. So, immediately, what we find is this lopsided time horizon where tool requests from the model only takes a small portion in comparison to actually executing the tool itself. For example, if I ask Claude Code to summarize this Wikipedia article on Rick Astley, it will eventually make a tool call to Claude Code to access the internet to fetch this article's content. And the agent is bound by how long it takes to actually reach the website and gather the content back to the Claude Code itself. And this same principle applies to every other tool calls like MCPs to make changes on Notion, making API calls to check the weather, building your code solution, making database calls, and transferring large files. All of these tool calls can easily take seconds if not minutes to execute them. Now, while this time horizon conundrum might seem like it only applies at the application layer, but it turns out the same conundrum applies practically at every layer underneath as well. The disparity that we just saw at the application layer on tool execution has a huge impact at the chip and infrastructure layer as well. Because this very time that it takes to execute the tool in a naive setup, it means that GPUs need to hold on to the cache until the tool finishes its job and sends back the result itself. And the GPU just sits there idling. And we all know just how expensive GPUs are, and not only how expensive they are to buy them, but also renting GPUs can range anywhere from two to seven dollars per hour per GPU. So, the time that it takes for tool execution to finish could mean wasting a large portion of the GPU time that's expensive. So, clearly, we need something more flexible to handle the agentic use cases at the chip and infrastructure layer as well. So, here's a thought. What if we come up with a system where we store the cache externally so that our GPUs can remain running so that during the time that the agent is doing its job to the external world, the cache can just exist outside of GPU that's expensive to a more abundant storage option elsewhere. This kind of method actually pairs up really nicely with what's called disaggregated inference, where that the infrastructure layer, we separate the compute in two sections where we have one section that handles batch processing of the input and tool results, and the other section that generates token after token after they have been processed. So, given this disaggregated inference, we let our cache that was processed to live somewhere else and accumulate through many agentic tool calls and only draw them up from the storage on demand as tool calls continue. Now, this might seem a bit counterintuitive at first, and you might be wondering, wouldn't it just add more overhead to keep transferring data like this back and forth? Which is a fair question to ask. And in fact, there's a huge overhead in moving cache around between systems. And recently, Chinese researchers have made a pretty clever step called dual path that optimizes the data movement between these two systems. And this aggregated inference certainly doesn't improve throughput. But, the idea behind it is to separate out the type of job by pairing them up with the right hardware for the job. For batch processing, that's typically compute heavy, we put a specialized GPU that can handle large compute. And for token generation, that's typically bandwidth heavy, we put a specialized GPU here as well for bandwidth and low latency. And now, with our agentic needs, we add an external storage option for the job for tasks that take long and need to be stored somewhere else so that the cache can live longer where we save cost on GPU, but sacrifice instead on cache being moved around as an overhead. In a similar way, Anthropic exposes prompt caching at the API level, as you can see here, where by default, we have a 5-minute prompt caching and optionally storing them up to 1 hour to live with higher pricing. And while Anthropic haven't released the exact method they use in the infrastructure layer, we can reasonably assume that there's some amount of what we just talked about to allow prompt caching in this way for agentic use cases. Now, you might look at this and think that this is just over-engineering for the sake of complicating things. And especially when we look at it from running inference locally at home using our own GPU definitely seems like an overkill. Locally run agents are typically used by one person, which is just you. And it makes very little sense to implement disaggregated inference and KV cache offloading because our compute is typically reserved entirely for our own needs. But at a mass scale, we need to have a dedicated GPU for various types of jobs and the mechanism to actually coordinate the movement of cache in this way. And for agentic use cases where you have constant tool calling in a loop, it requires a much more comprehensive system like this. Now, speaking of locally run agents at home, the perfect place to buy your computer parts to run agents at home is none other than Micro Center, who's sponsoring this video. Hardware prices are going up and up, and buying them during sales is always great. Micro Center is running sales on all kinds of laptops, and it's a perfect time to load up on your hardware, especially as AI demand continues to grow. And here's an example where RTX 5080 is offered at a huge discount while the deal lasts. If you have been thinking about building your own computer to try running agents at home, Micro Center is the perfect store to go buy your parts where you can select different kinds of CPUs to pick and choose from Intel to AMD. And for your GPU needs, they have Intel, Nvidia, and AMD, as well as all kinds of solid state drives to choose from. As AI demand continues to grow, shop at Micro Center for all your computer and computer parts needs. And right now, if you live near Austin, Texas, you can sign up for a free 128 GB flash drive to redeem it in store when they open later this year. And if you're near Columbus, Ohio, the store is also getting remodeled, and you can sign up for a free 128 GB flash drive on reopening later this year, as well. Micro Center also has valuable blogs around AI, so you can check out the articles as you experiment with your build. Link in the description below. Now, as we just saw how agentic use cases not only affects what happens in the model and application layers, it also affects what happens in the infrastructure and chip layer, as well. And following this trajectory, we can go one level deeper into the energy layer and find that agentic use cases can also impact how we portion our energy budget, as well. A typical data center that just generates token after token might heavily use racks and racks of GPUs in their power consumption. But agentic use cases at the application layer adds a pressure to increase the blend of other components in the data centers, like CPU for agentic tasks, storage options for keeping context and tool results, networking for accessing the web compared to a more traditional GPU hyperscaler. So, as agent workloads become more common, data centers also need to think about the ratio of GPU, CPU, storage, networking, and the power delivery. And the same principle applies to locally run agents as well. Even when we build our own PC at home, we have to think about how to budget our, let's say, 850-W PSU might handle a GPU like RTX 4090, which needs 450 W of power. So, the moment we start adding other parts as well, like CPU, memory, storage, cooling, and networking, which are all necessary for efficient agentic loop, we quickly realize that raw GPU power alone is not the only full constraint anymore. Now, when we zoom back out and look at the entire stack that we just covered, in a coding scenario like using a coding agents like Cloud Code to make changes to our repository, the agent might engage in a react loop, where Cloud Code will reason through the prompt and make action, observe the change, and repeat the cycle over and over again until it's done. And one interesting study to consider is that not all agents in this very harness yield the same result, meaning depending on the harness that you use, you might get a varying result on the speed of the agent as well. In fact, some people conduct ablation studies on various harnesses, like comparing Cloud Code against Codex, for example, to see the efficiency at the harnessing layer as well, which is a factor worth considering when it comes to agents. This is especially worth considering given the age-old debate whether the model's intelligence and speed derives from the model itself or the harnessing around it or both. What's also worth considering is the growing demand at the application side for parallel agents, whether that's sub-agents or even tasks like deep research, where you have multiple agents all working in parallel against a parent orchestrator. All of these use cases at the application layer also mixes up the entire stack underneath depending on the demand that we have from users. Looking back, there's one layer in the AI stack that I glossed over and that's the model layer. Some models are really good at calling the right tool at the right time while other models are notorious for being really bad at tool calling. And a lot of it comes down to how well the model reasons through the information provided and more importantly, the post training portion that the model goes through to actually learn when and how to call the right tool and how to structure the tool call so that in the real world, the agentic use cases work efficiently. There are benchmarks to consider like the BFCL that helps you measure this to compare which models are good at the accuracy part of the tool call. And a model that's inefficient at tool calling could easily mean that they're wasting time spinning in circles to do things by itself when it really should be relying on external services by calling tools whether it's MCP or custom application that it has access to.

Frontier News · by Hyperjump Technology