The blind spot
The shuffle test
Two kinds of “where”
Learned absolute positions
Token vector + position vector
xi = E[idi] + P[i]
The position table
What the learned table ends up looking like
Where learned positions break
The maximum-length wall
Same relation, different numbers
Three answers, three injection points
Exceptions and sharp edges
Left padding shifts positions
compute position ids from the attention mask, not from the raw index.Packed documents
pick one, and use matching attention masks.Under-trained tail rows
include enough long examples, or use a scheme without a table.No positions is not zero information
a curiosity, not a replacement for real positions.Swapping position scheme
decide at the start, or plan a fine-tuning phase.