The decoder head stopped being one door.

The output head becomes a family of backward conduits rather than a singleton, without pretending that the goal is merely a higher-rank forward distribution.

The door learned to turn with context.

A normal language-model head is one fixed door through which every context must leave—and through which its training signal must return. This design makes several doors from the current hidden state. Each context still has a finite opening, but the opening does not face the same direction every time.

Everyone assumed
A language model should send every context through one fixed output projection.
I asked
What if the important bottleneck is the orientation of the training signal returning to the layer stack?
I changed
Replace one fixed vocabulary-space aperture with several context-conditioned directional carveouts whose surviving backward subspace rotates with the hidden state.
Now there is
A same-budget decoder design, invariant probes, trained demonstrations awaiting import, and a broader model laboratory that also contains exact circular positional geometry and LELU.