Videos list
Thumb Title Channel Status Published
What's New in Inference Engineering — Philip Kiely, Baseten

Turboquant, which uses polar coordinates to quantize the KV cache to four bits, has a fatal flaw for data center...

AI Engineer summarized 2026-09-19 17:00
China Just Open-Sourced 6 Ways to Speed Up AI Inference (Tencent AngelSpec)

Tencent's AngelSpec is a training framework for speculative decoding that lets you compare six drafting...

Cloud Codes summarized 2026-08-30 19:30
Speculative Decoding: The ONLY Video You Need to Speed Up Inference

Speculative decoding is a lossless inference acceleration technique that uses a small draft model to propose...

Cloud Codes summarized 2026-08-26 19:30
1,000 Tokens/Sec on One RTX 3090 (Here's the Config)

A configuration of nine software changes on an RTX 3090 achieves 1,000 tokens/sec for 64 concurrent users by...

Cloud Codes summarized 2026-08-25 17:00

Frontier News · by Hyperjump Technology