Figure 6

Which tokens should the 256 bins be?

Every bin needs its own single-token label, and the model has to learn that each label means a position. Two properties of the labels' embeddings go with the sets that learned: they are distinct (no near-duplicates), and they have some built-in order (neighbouring bins already a little alike).

Ordinary words

256 common words, one token each
appleriverstonecloudpencil…
near-duplicate pairs0%
built-in order+0.007
Unrelated words point in near-perpendicular directions (mean cosine 0.03): perfectly distinct, but with no order at all. Nothing in the embeddings says bin 117 sits next to 118, and each label has a meaning to unlearn.
alone: never learned (0.6% of clicks)
with a coarse slot (arm A): 0.835

101 byte-like tokens

84 single characters + 17 two-letter tokens · arm B+ (the GPC-1 grid size)
!"#$%…z{~
near-duplicate pairs0%
built-in order+0.099
Little meaning to unlearn. But 101 bins is coarse: ~13 px per bin on a 1,280-px screen.
clicks 0.867 · agent steps 0.605

256 byte-like tokens

84 characters + 172 two-letter tokens; the first 101 are B+'s · arm E
!"…~abac…gs
near-duplicate pairs0%
built-in order+0.115
Distinct (no near-duplicates; the closest pair is at cosine 0.83), some built-in order, little meaning to unlearn, and 256 bins: ~5 px per bin.
clicks 0.858 · agent steps 0.700 (best)

Byte-fallback tokens

<0x00> … <0x7F>: the tokenizer's raw bytes
<0x00><0x01>…<0x7F>
near-duplicate pairs74%
built-in order+0.084
Some order, but 74% of pairs are near-identical (cosine above 0.99): the model can hardly tell those labels apart.
rejected before training

Measured on Gemma 4's embeddings (input and output are tied; label_geometry.py). Near-duplicate pairs: share of label pairs with cosine above 0.99. Built-in order: mean cosine between neighbouring bins minus that of random pairs (bar scale up to +0.15). Mean pairwise cosine alone does not separate the sets: words have the lowest (0.034) because unrelated words point in near-perpendicular directions in the high-dimensional embedding space, yet they failed. These properties go with what learned; we did not test them causally. Accuracy: ScreenSpot clicks (1,272) and held-out AGUVIS steps (200), joint training on identical data; the words-only result is from phase 1 (clicks only).