Videos list
Thumb Title Channel Status Published
Qwen3.8-27B & How to Serve it Fast

Qwen released the 27B parameter Qwen3.8-27B model, which significantly outperforms its predecessor Qwen3.6-27B and...

Sam Witteveen summarized 2026-08-18 13:00
Qwen 3.8 27B is 3X Faster With ONE Setting

Qwen's 27B model ships with a multi-token prediction head that can nearly triple throughput for free—no extra...

Cloud Codes summarized 2026-08-17 12:00
Nemotron Lightning - NVIDIA's Super Fast Agent MoE

NVIDIA quietly released Nemotron Lightning, a small open-weights MoE model (30B total, 3B active) built specifically...

Sam Witteveen summarized 2026-08-11 13:30
The Inference Frontier: 10x Faster Models to Self-Optimizing AI — Philip Kiely & Ali Taha, Baseten

Inference engineering for large language models involves a stack of optimizations including KV cache-aware routing,...

Latent Space summarized 2026-08-03 21:35
The 100,000 Sandbox Problem — Akshat Bubna, Modal CTO

Modal is a cloud platform built for AI workloads, focusing on elastic inference, sandboxes, and training. They've...

Latent Space summarized 2026-07-08 22:42

Frontier News · by Hyperjump Technology