Videos
4 total
| Thumb | Title | Channel | Status | Published |
|---|---|---|---|---|
|
|
What's New in Inference Engineering — Philip Kiely, Baseten
Turboquant, which uses polar coordinates to quantize the KV cache to four bits, has a fatal flaw for data center... |
AI Engineer | summarized | 2026-09-19 17:00 |
|
|
China Just Open-Sourced 6 Ways to Speed Up AI Inference (Tencent AngelSpec)
Tencent's AngelSpec is a training framework for speculative decoding that lets you compare six drafting... |
Cloud Codes | summarized | 2026-08-30 19:30 |
|
|
Speculative Decoding: The ONLY Video You Need to Speed Up Inference
Speculative decoding is a lossless inference acceleration technique that uses a small draft model to propose... |
Cloud Codes | summarized | 2026-08-26 19:30 |
|
|
1,000 Tokens/Sec on One RTX 3090 (Here's the Config)
A configuration of nine software changes on an RTX 3090 achieves 1,000 tokens/sec for 64 concurrent users by... |
Cloud Codes | summarized | 2026-08-25 17:00 |
Frontier News · by Hyperjump Technology