Enjoying this issue?
Get tomorrow's AI & engineering digest in your inbox — hand-picked, summarized, and always spam-free.
TLDR
AMD's ROCm 10 release ships a 'skills' folder that teaches Claude, Cursor, and Codex how to use its GPUs, effectively conceding that the moat in AI hardware is now what coding agents already know how to do. But the execution is thin: AMD published only 8 skills against Nvidia's 343, and the headline 3.3x performance gain was measured on a preview build that users cannot download. The move is strategically right, but the launch oversells a half-finished catalog.
Key points
AMD shipped ROCm 10, jumping from version 7.14 to 10 in 42 days.
AMD Skills are markdown files that instruct coding agents how to use AMD hardware.
Nvidia's skills catalog has 343 skills; AMD has 8, with 2 more in staging.
Hyperloom is an agent that autonomously optimizes inference workloads, achieving up to 193% improvement.
The 3.3x faster inference claim was measured on a preview build, not the released ROCm 10.
ROCm 10's known issues include a 9-25% training throughput drop on MI350X due to an attention kernel regressor.
Tools mentioned
Techniques
- Skill.md format: description loaded in agent context, body loaded only on task match.
- Hyperloom optimization loop: profile, analyze, plan, optimize, validate.
- Tree search and critic agent for autonomous optimization.
- ROCm unified build system: single pipeline for all accelerators and Windows/Linux.
- Evaluation harness and benchmark file per skill (noted as present in AMD's skills but no signatures).
Stop scrolling. Start reading smarter.
Receive the day's most important AI & engineering updates in one concise email. No spam.
Transcript (captions)
On the 15th of July, AMD shipped a release of its GPU software stack called ROCm 7.14. 42 days later, one ordinary release cycle, they shipped ROCm 10. The major
version went up by three inside 6 weeks. AMD shipped ROCm 1.0 back in April of 2016, so this one lands a decade in. The announcement subtitle says, "The rest of it out loud." Built for the age of
agentic AI. You don't skip two major versions for a library update, so it's worth asking what's in there. The headline feature isn't a compiler or a driver. It isn't a kernel, either. It's
a folder of markdown files that teaches your coding agent how to use an AMD graphics card. AMD calls them AMD skills, and they install into three coding agents: Claude, Cursor, and
CodeX. Which means a chip company is now shipping software whose only job is to make somebody else's agent smarter about its own hardware. I think that idea is right. I also think, on the receipts,
AMD is about 6 weeks late to it and a few hundred skills short. So, let's take both halves in turn, starting with what a skill actually is, because the word does a lot of work here. A skill is a
folder with a file called skill.md. In it, a name, one line describing when to use it, and the instructions underneath. The description stays loaded in the agent's context for almost
nothing. The body only loads when the task matches. That's the entire mechanism, and it's why 100 skills don't drown an agent. Anthropic wrote that format and released it as an open
standard. And other agent products picked it up, which is what makes a folder written by AMD portable to a tool AMD doesn't own. Now, AMD's own explanation of why they bothered is the
sharpest paragraph in the release, and I want to read it to you. Their words: "Documentation describes an API surface. Every flag, every option, neutral by design. A skill encodes the opinionated
path." That means which flags, which container image, which environment variables, in what order. The decisions a senior AMD engineer makes without thinking. And the
distinction is real because documentation gets written for a person who'll skim it once, while a skill gets written for a machine that follows it every single time. It all ships under
one name, roc-m.ai, and it has three parts. The first is a command line tool that folds a pile of separate install scripts into one binary. It can stand up a model
for inference or examine a broken driver and tell you which piece is wrong. AMD ships it as a tech preview, which is phrase for expect this to change. So, those are the skills and that's the
command line. The man who runs AMD's AI software group, Anush Elangovan, framed the whole thing as agents that profile, debug, and drive workloads toward peak performance. The third piece is where
that gets literal and it's called Hyperloom. Hyperloom is an agent that optimizes your inference workload without you in the room. It profiles the job, finds the
bottleneck, plans a change, writes the code, benchmarks it, and checks that answers still come out the same. It goes profile, analyze, plan, optimize, validate, on repeat until it hands you a
report of everything it changed and what each change bought. AMD says that turns weeks of manual tuning into hours. There's a real research paper underneath it.
On full stack inference optimization, the harness reaches up to 193% better on a combined throughput and latency curve against baselines the vendor had already tuned by hand. The number I find more
convincing is the control. One agent, no harness around it, plateaus at 33% and then crashes irrecoverably within hours. The tree search and the critic agent are what buy
the difference and that's a finding worth having. Now, look at what's driving all of it. Hyperloom's own supported features table lists one language model back end and it's Claude.
AMD's autonomous optimizer for AMD hardware is driven by somebody else's frontier model. The main instruction file in its own repository is named skill.md. I don't think that's
embarrassing. I think it's the argument of the whole release stated in a file path. So, we have the pitch, one command line, a folder of skills, and an agent that tunes your kernels for you.
Now, the number, because does a folder of markdown really move a hardware decision? Coverage of the release led with one number. 3.3 times faster inference and 2.4 times faster training
over ROCm 7. But, read AMD's own end note and it's a different sentence. The test ran on the 7th of July, 7 weeks before ROCm 10 existed. The baseline is ROCm 7.0, which
shipped back in September of last year. By the test date, nine newer releases had already shipped, and the fast side isn't ROCm 10, either. It's a preview build of ROCm.ai sitting on 7.22 with
hand-applied kernel, scheduling, and parallelism work on top. It ran on eight Instinct MI 355X accelerators in a rack. The measurement is real. It is not a
measurement of the software you can download today. And ROCm 10's own release notes carry a known issues list, which is specific. Hugging Face training throughput can fall 9 to 25% on Instinct
MI 350X, because the attention kernel picker regressed to a slower path. The workaround is to rebuild PyTorch and pin the old version. PyTorch fine-tuning can reset the GPU outright on some Radeon
cards, and on a different set of three, inference can fail to start. Each of those ships with an environment variable as the fix, and one of them warns that the fix costs you performance. Over on
the local model subreddit, the release thread ran past 260 votes, and the replies come from people who'd actually installed it. One reads, "Installed it today and built llama.cpp.
No change in speed or any difference for me." Another, from a 7900XTX owner, "Literal no difference." Which tracks, and it isn't a contradiction.
3.3 times was measured in a data center rack on a preview build. A desktop card was not in that test. So, that is the performance story and it's oversold. The skill story is the one carrying the
weight and there's something in it neither AMD announcement mentions. Nvidia got there first. Nvidia's official skills catalog repository was created on the 25th of February.
AMD's was created 42 days later on the 8th of April. I counted both repository trees on the 29th of August. Nvidia publishes 343 skills. AMD publishes eight with two more sitting in a staging
folder. In fairness, most of Nvidia's aren't kernel work. 60 of them cover its networking chip alone, but everyone ships with a detached signature you can verify after download and a benchmark
file beside it. AMD's ship a skill card and an evaluation harness and not one signature in the tree. And there's a smaller thing that tells you where AMD actually is. Their launch post names a
skill for quantizing models on epic processors. I read the entire catalog and it is not there. The diagnostic line tool sits in a staging folder marked planned
in AMD's own table. That's a launch blog describing a catalog that's still being written, which is normal and it is not the same thing as shipped. And underneath the marketing, there's real
plumbing, which deserves saying. Every part of ROCm now comes out of one automated build system called the ROC primitives, libraries, and framework wheels from a single pipeline validated
across Instinct accelerators, Radeon cards, and Ryzen integrated graphics in the same pass. On Windows, the old separate SDK is retired. Windows and Linux draw from
the same source tree and the same 6-week cadence now. Although Windows still ships as a tarball you extract yourself and native installers are promised later this year. So, here's where I land. The
move is right and the execution is thin. For a decade, the argument was that CUDA's moat is the compiler and the libraries and AMD spent that decade narrowing it. This release concedes,
without saying so, that the moat has moved somewhere else. The moat is now what your coding agent already knows how to do with the card in your machine, and whoever writes those instructions owns
the default. Defaults are how hardware gets bought. If you rent Instinct racks, ROCm 10 is the best ROCm there's been, known issues and all, and I'd take it at twice the migration cost. One built
system, one SDK across Windows and Linux, a 6-week cadence, and Hyperloom is a new kind of tool. If you own a Radeon card, this is a packaging release with a known issues
list, and you should read that list before you upgrade. The villain in this story isn't AMD. It's the launch day number, a 3.3 times measured on software you cannot install, repeated in the
headlines while the end note that undoes it sat in AMD's own newsroom the whole time. And my question is the one AMD's own file path already asked. If what sells a GPU is how well an agent knows
it, who should be writing those instructions? The vendor whose blog name skills it hasn't shipped, or the people who already got the card working?