Figure 1

Writing a decision vs reading it

Writing

the model generates its step as JSON, one token per forward pass
…screenshot, task, history…
<assistant> ▌
0 ms

Reading

the prompt ends where each answer goes; one pass gives every distribution
…screenshot, task, history…
<assistant> action: ▮ x: ▮ y: ▮
0 ms

Measured: slot read 146 ms per step; vanilla Gemma 4 writing its native {"action", "point": [y, x]} JSON 472–546 ms (about 20 tokens). Both start with the same prefill of the screenshot and prompt (shown here as ~110 ms); writing then spends one forward pass per token.