How much data the wider alphabet needs

Parity task, fixed counting range, model size, and optimizer budget. Exact match on the longest split against the training-set size, for macro orders k=1, 2, 4. Hover the points.

k=1 (vocab 9) k=2 (vocab 11) k=4 (vocab 15)

At the smallest budget of 60 examples the widest-vocabulary order k=4 reaches only 0.44, while all orders reach 1.00 once the training set is large enough. This is the sample-complexity side of the positions-versus-vocabulary trade.