The door learned to turn with context.
A normal language-model head is one fixed door through which every context must leave—and through which its training signal must return. This design makes several doors from the current hidden state. Each context still has a finite opening, but the opening does not face the same direction every time.
- Everyone assumed
- A language model should send every context through one fixed output projection.
- I asked
- What if the important bottleneck is the orientation of the training signal returning to the layer stack?
- I changed
- Replace one fixed vocabulary-space aperture with several context-conditioned directional carveouts whose surviving backward subspace rotates with the hidden state.
- Now there is
- A same-budget decoder design, invariant probes, trained demonstrations awaiting import, and a broader model laboratory that also contains exact circular positional geometry and LELU.