Figure 8

Where one step's time goes

Gemma 4 31B (NVFP4) on one RTX PRO 6000, one request at a time, a fresh screenshot. Writing pays one forward pass per output token; reading stops after the prefill. The first slot reader lost most of its time in the server's request handling, not on the GPU.

request handling (HTTP, template, image)prefill on the GPU one decode step per output token8 slot requests waiting on one API process

Segments are approximate, built from vLLM's per-request metrics (prefill 79–87 ms; decode ≈ 19 ms per token) and our measured end-to-end medians: vanilla native JSON 472–546 ms for ~20 tokens; slots 237 ms before the fix and 146 ms after (8 API-server processes, JPEG screenshots at model size, 256 label ids per request); vanilla with thinking 2.6–3.1 s.