Figure 8
Where one step's time goes
Gemma 4 31B (NVFP4) on one RTX PRO 6000, one request at a time, a fresh screenshot. Writing pays one forward pass per output token; reading stops after the prefill. The first slot reader lost most of its time in the server's request handling, not on the GPU.
Segments are approximate, built from vLLM's per-request metrics (prefill 79–87 ms; decode ≈ 19 ms per token) and our measured end-to-end medians: vanilla native JSON 472–546 ms for ~20 tokens; slots 237 ms before the fix and 146 ms after (8 API-server processes, JPEG screenshots at model size, 256 label ids per request); vanilla with thinking 2.6–3.1 s.