Figure 6
Which tokens should the 256 bins be?
Every bin needs its own single-token label, and the model has to learn that each label means a position. Two properties of the labels' embeddings go with the sets that learned: they are distinct (no near-duplicates), and they have some built-in order (neighbouring bins already a little alike).
Measured on Gemma 4's embeddings (input and output are tied; label_geometry.py). Near-duplicate pairs: share of label pairs with cosine above 0.99. Built-in order: mean cosine between neighbouring bins minus that of random pairs (bar scale up to +0.15). Mean pairwise cosine alone does not separate the sets: words have the lowest (0.034) because unrelated words point in near-perpendicular directions in the high-dimensional embedding space, yet they failed. These properties go with what learned; we did not test them causally. Accuracy: ScreenSpot clicks (1,272) and held-out AGUVIS steps (200), joint training on identical data; the words-only result is from phase 1 (clicks only).