Videos list
Thumb Title Channel Status Published
Weight Folding, CUDA Streams, and the Bug That Made My Model Speak Backwards — Filip Makraduli

A paper co-authored by Filip Makraduli proposes two algebraic tricks—weight folding and deferred normalization—that...

AI Engineer summarized 2026-09-19 19:00
Run a $10,000 AI Model at Home, Here’s How

Open source models have narrowed the capability gap with frontier labs to 3-6 months, with recent releases like Kimi...

David Ondrej summarized 2026-09-09 10:27
AMD Shipped Skills for Claude, Cursor and Codex (All 8 of Them)

AMD's ROCm 10 release ships a 'skills' folder that teaches Claude, Cursor, and Codex how to use its GPUs,...

Cloud Codes summarized 2026-08-31 09:00
The KV Cache Layer That Makes LLMs 10x Faster? (LMCache)

The KV cache is the dominant cost in long-context LLM serving, and the built-in prefix caching in frameworks like...

Cloud Codes summarized 2026-08-27 19:30
Speculative Decoding: The ONLY Video You Need to Speed Up Inference

Speculative decoding is a lossless inference acceleration technique that uses a small draft model to propose...

Cloud Codes summarized 2026-08-26 19:30

Frontier News · by Hyperjump Technology