AMD Shipped Skills for Claude, Cursor and Codex (All 8 of Them)

summarized

TLDR

AMD's ROCm 10 release ships a 'skills' folder that teaches Claude, Cursor, and Codex how to use its GPUs, effectively conceding that the moat in AI hardware is now what coding agents already know how to do. But the execution is thin: AMD published only 8 skills against Nvidia's 343, and the headline 3.3x performance gain was measured on a preview build that users cannot download. The move is strategically right, but the launch oversells a half-finished catalog.

Key points

AMD shipped ROCm 10, jumping from version 7.14 to 10 in 42 days.

AMD Skills are markdown files that instruct coding agents how to use AMD hardware.

Nvidia's skills catalog has 343 skills; AMD has 8, with 2 more in staging.

Hyperloom is an agent that autonomously optimizes inference workloads, achieving up to 193% improvement.

The 3.3x faster inference claim was measured on a preview build, not the released ROCm 10.

ROCm 10's known issues include a 9-25% training throughput drop on MI350X due to an attention kernel regressor.

Tools mentioned

Techniques

  • Skill.md format: description loaded in agent context, body loaded only on task match.
  • Hyperloom optimization loop: profile, analyze, plan, optimize, validate.
  • Tree search and critic agent for autonomous optimization.
  • ROCm unified build system: single pipeline for all accelerators and Windows/Linux.
  • Evaluation harness and benchmark file per skill (noted as present in AMD's skills but no signatures).
Transcript (captions)

0:00 On the 15th of July, AMD shipped a release of its GPU software stack called ROCm 7.14. 42 days later, one ordinary release cycle, they shipped ROCm 10. The major

0:11 version went up by three inside 6 weeks. AMD shipped ROCm 1.0 back in April of 2016, so this one lands a decade in. The announcement subtitle says, "The rest of it out loud." Built for the age of

0:24 agentic AI. You don't skip two major versions for a library update, so it's worth asking what's in there. The headline feature isn't a compiler or a driver. It isn't a kernel, either. It's

0:34 a folder of markdown files that teaches your coding agent how to use an AMD graphics card. AMD calls them AMD skills, and they install into three coding agents: Claude, Cursor, and

0:45 CodeX. Which means a chip company is now shipping software whose only job is to make somebody else's agent smarter about its own hardware. I think that idea is right. I also think, on the receipts,

0:56 AMD is about 6 weeks late to it and a few hundred skills short. So, let's take both halves in turn, starting with what a skill actually is, because the word does a lot of work here. A skill is a

1:07 folder with a file called skill.md. In it, a name, one line describing when to use it, and the instructions underneath. The description stays loaded in the agent's context for almost

1:17 nothing. The body only loads when the task matches. That's the entire mechanism, and it's why 100 skills don't drown an agent. Anthropic wrote that format and released it as an open

1:28 standard. And other agent products picked it up, which is what makes a folder written by AMD portable to a tool AMD doesn't own. Now, AMD's own explanation of why they bothered is the

1:38 sharpest paragraph in the release, and I want to read it to you. Their words: "Documentation describes an API surface. Every flag, every option, neutral by design. A skill encodes the opinionated

1:50 path." That means which flags, which container image, which environment variables, in what order. The decisions a senior AMD engineer makes without thinking. And the

2:00 distinction is real because documentation gets written for a person who'll skim it once, while a skill gets written for a machine that follows it every single time. It all ships under

2:09 one name, roc-m.ai, and it has three parts. The first is a command line tool that folds a pile of separate install scripts into one binary. It can stand up a model

2:19 for inference or examine a broken driver and tell you which piece is wrong. AMD ships it as a tech preview, which is phrase for expect this to change. So, those are the skills and that's the

2:30 command line. The man who runs AMD's AI software group, Anush Elangovan, framed the whole thing as agents that profile, debug, and drive workloads toward peak performance. The third piece is where

2:41 that gets literal and it's called Hyperloom. Hyperloom is an agent that optimizes your inference workload without you in the room. It profiles the job, finds the

2:50 bottleneck, plans a change, writes the code, benchmarks it, and checks that answers still come out the same. It goes profile, analyze, plan, optimize, validate, on repeat until it hands you a

3:02 report of everything it changed and what each change bought. AMD says that turns weeks of manual tuning into hours. There's a real research paper underneath it.

3:11 On full stack inference optimization, the harness reaches up to 193% better on a combined throughput and latency curve against baselines the vendor had already tuned by hand. The number I find more

3:23 convincing is the control. One agent, no harness around it, plateaus at 33% and then crashes irrecoverably within hours. The tree search and the critic agent are what buy

3:33 the difference and that's a finding worth having. Now, look at what's driving all of it. Hyperloom's own supported features table lists one language model back end and it's Claude.

3:43 AMD's autonomous optimizer for AMD hardware is driven by somebody else's frontier model. The main instruction file in its own repository is named skill.md. I don't think that's

3:53 embarrassing. I think it's the argument of the whole release stated in a file path. So, we have the pitch, one command line, a folder of skills, and an agent that tunes your kernels for you.

4:04 Now, the number, because does a folder of markdown really move a hardware decision? Coverage of the release led with one number. 3.3 times faster inference and 2.4 times faster training

4:15 over ROCm 7. But, read AMD's own end note and it's a different sentence. The test ran on the 7th of July, 7 weeks before ROCm 10 existed. The baseline is ROCm 7.0, which

4:27 shipped back in September of last year. By the test date, nine newer releases had already shipped, and the fast side isn't ROCm 10, either. It's a preview build of ROCm.ai sitting on 7.22 with

4:39 hand-applied kernel, scheduling, and parallelism work on top. It ran on eight Instinct MI 355X accelerators in a rack. The measurement is real. It is not a

4:50 measurement of the software you can download today. And ROCm 10's own release notes carry a known issues list, which is specific. Hugging Face training throughput can fall 9 to 25% on Instinct

5:01 MI 350X, because the attention kernel picker regressed to a slower path. The workaround is to rebuild PyTorch and pin the old version. PyTorch fine-tuning can reset the GPU outright on some Radeon

5:12 cards, and on a different set of three, inference can fail to start. Each of those ships with an environment variable as the fix, and one of them warns that the fix costs you performance. Over on

5:22 the local model subreddit, the release thread ran past 260 votes, and the replies come from people who'd actually installed it. One reads, "Installed it today and built llama.cpp.

5:33 No change in speed or any difference for me." Another, from a 7900XTX owner, "Literal no difference." Which tracks, and it isn't a contradiction.

5:43 3.3 times was measured in a data center rack on a preview build. A desktop card was not in that test. So, that is the performance story and it's oversold. The skill story is the one carrying the

5:54 weight and there's something in it neither AMD announcement mentions. Nvidia got there first. Nvidia's official skills catalog repository was created on the 25th of February.

6:05 AMD's was created 42 days later on the 8th of April. I counted both repository trees on the 29th of August. Nvidia publishes 343 skills. AMD publishes eight with two more sitting in a staging

6:18 folder. In fairness, most of Nvidia's aren't kernel work. 60 of them cover its networking chip alone, but everyone ships with a detached signature you can verify after download and a benchmark

6:29 file beside it. AMD's ship a skill card and an evaluation harness and not one signature in the tree. And there's a smaller thing that tells you where AMD actually is. Their launch post names a

6:39 skill for quantizing models on epic processors. I read the entire catalog and it is not there. The diagnostic line tool sits in a staging folder marked planned

6:51 in AMD's own table. That's a launch blog describing a catalog that's still being written, which is normal and it is not the same thing as shipped. And underneath the marketing, there's real

7:01 plumbing, which deserves saying. Every part of ROCm now comes out of one automated build system called the ROC primitives, libraries, and framework wheels from a single pipeline validated

7:12 across Instinct accelerators, Radeon cards, and Ryzen integrated graphics in the same pass. On Windows, the old separate SDK is retired. Windows and Linux draw from

7:23 the same source tree and the same 6-week cadence now. Although Windows still ships as a tarball you extract yourself and native installers are promised later this year. So, here's where I land. The

7:33 move is right and the execution is thin. For a decade, the argument was that CUDA's moat is the compiler and the libraries and AMD spent that decade narrowing it. This release concedes,

7:44 without saying so, that the moat has moved somewhere else. The moat is now what your coding agent already knows how to do with the card in your machine, and whoever writes those instructions owns

7:54 the default. Defaults are how hardware gets bought. If you rent Instinct racks, ROCm 10 is the best ROCm there's been, known issues and all, and I'd take it at twice the migration cost. One built

8:04 system, one SDK across Windows and Linux, a 6-week cadence, and Hyperloom is a new kind of tool. If you own a Radeon card, this is a packaging release with a known issues

8:15 list, and you should read that list before you upgrade. The villain in this story isn't AMD. It's the launch day number, a 3.3 times measured on software you cannot install, repeated in the

8:26 headlines while the end note that undoes it sat in AMD's own newsroom the whole time. And my question is the one AMD's own file path already asked. If what sells a GPU is how well an agent knows

8:37 it, who should be writing those instructions? The vendor whose blog name skills it hasn't shipped, or the people who already got the card working?

Frontier News · by Hyperjump Technology