Figure 2

LLMs as Decision Model

Each question is just a short tail on the same prompt. The prefix (instructions, the state, an image) is processed once and cached; every extra question costs a few tokens and one read.

system + questions + state (+ image) ≈ hundreds of tokens · computed once (prefix cache)

This is what systemone sends to vLLM: one batched request; prefix caching processes the shared prompt once. Slot reading (later) uses the same idea for an action and its coordinates.