The Inference Frontier: 10x Faster Models to Self-Optimizing AI — Philip Kiely & Ali Taha, Baseten
Inference engineering for large language models involves a stack of optimizations including KV cache-aware routing, disaggregated prefill/decode, speculative decoding, and quantization to achieve 10x speedups. The field is converging with training, as models like GLM52 write their own GPU kernels, and future gains will come from faster interconnects and continual learning via KV cache compaction.