One fixed aperture
A standard low-rank output head sends every token through the same narrow backward subspace, regardless of what the current hidden state is trying to express.
Active research · language-model architecture
Do not merely increase the forward rank of the decoder. Change which directions in vocabulary space can teach the layer stack at this context.
Can a context-rotating backward aperture improve useful vocabulary coverage and specialization without sacrificing the stability of an ordinary tied head?
A standard low-rank output head sends every token through the same narrow backward subspace, regardless of what the current hidden state is trying to express.
The hidden state selects a context-dependent backward subspace while the total parameter budget remains comparable to the ordinary construction.
Coverage-preserving vocabulary dropout, separation, redundancy, and residual controls keep the moving subspace from collapsing into a small recurring set of favored directions.
The research includes invariant probes and an empirical demonstrator so any gain can be distinguished from a disguised increase in capacity or an unstable change of coordinates.
A PyTorch implementation, coverage-preserving vocabulary dropout, separation and redundancy controls, residual paths, invariant probes, and an empirical demonstrator are public. The object remains a research system because the decisive comparison is not whether it can emit logits. It is whether the moving backward geometry changes what the rest of the model can learn.