Videos
4 total
| Thumb | Title | Channel | Status | Published |
|---|---|---|---|---|
|
|
One GPU. 30 People. What Runs Out First? (vLLM)
For shared GPU serving with vLLM, memory space from the KV cache — not compute — is what usually limits how many... |
Cloud Codes | summarized | 2026-09-29 12:00 |
|
|
Two Bugs That Hid in Plain Sight: A vLLM Debugging Detective Story — Asaf Gardin & Yuval Belfer
vLLM's Mamba state cache had two silent bugs—decode-before-prefill scheduling and uint32 index overflow—that... |
AI Engineer | summarized | 2026-09-19 18:30 |
|
|
Qwen 3.8 27B and Hermes Agent built a vLLM Monitoring App for Local AI
Qwen 3.8 27B at FP16, despite being slower, outperforms Qwen 3.8 Flash Next at int4 for agentic code tasks because... |
Digital Spaceport | summarized | 2026-09-06 04:50 |
|
|
Qwen 3.8 Flash Next + HERMES AGENT = AWESOME LOCAL AI AGENTS!
Qwen 3.8 Flash Next running in a Hermes agent on a quad-3090 setup achieves 55-60 tokens per second with a Q4... |
Digital Spaceport | summarized | 2026-09-02 02:10 |
Frontier News · by Hyperjump Technology