Videos
5 total
| Thumb | Title | Channel | Status | Published |
|---|---|---|---|---|
|
|
Qwen3.8-27B & How to Serve it Fast
Qwen released the 27B parameter Qwen3.8-27B model, which significantly outperforms its predecessor Qwen3.6-27B and... |
Sam Witteveen | summarized | 2026-08-18 13:00 |
|
|
Qwen 3.8 27B is 3X Faster With ONE Setting
Qwen's 27B model ships with a multi-token prediction head that can nearly triple throughput for free—no extra... |
Cloud Codes | summarized | 2026-08-17 12:00 |
|
|
Nemotron Lightning - NVIDIA's Super Fast Agent MoE
NVIDIA quietly released Nemotron Lightning, a small open-weights MoE model (30B total, 3B active) built specifically... |
Sam Witteveen | summarized | 2026-08-11 13:30 |
|
|
The Inference Frontier: 10x Faster Models to Self-Optimizing AI — Philip Kiely & Ali Taha, Baseten
Inference engineering for large language models involves a stack of optimizations including KV cache-aware routing,... |
Latent Space | summarized | 2026-08-03 21:35 |
|
|
The 100,000 Sandbox Problem — Akshat Bubna, Modal CTO
Modal is a cloud platform built for AI workloads, focusing on elastic inference, sandboxes, and training. They've... |
Latent Space | summarized | 2026-07-08 22:42 |
Frontier News · by Hyperjump Technology