Active research · language-model architecture

The unembedding
system

Do not merely increase the forward rank of the decoder. Change which directions in vocabulary space can teach the layer stack at this context.

Current question

Can a context-rotating backward aperture improve useful vocabulary coverage and specialization without sacrificing the stability of an ordinary tied head?

The mechanism

Ordinary head

One fixed aperture

A standard low-rank output head sends every token through the same narrow backward subspace, regardless of what the current hidden state is trying to express.

Proposed head

The aperture rotates

The hidden state selects a context-dependent backward subspace while the total parameter budget remains comparable to the ordinary construction.

Coverage

Vocabulary directions take turns

Coverage-preserving vocabulary dropout, separation, redundancy, and residual controls keep the moving subspace from collapsing into a small recurring set of favored directions.

Evidence

Probe the invariants

The research includes invariant probes and an empirical demonstrator so any gain can be distinguished from a disguised increase in capacity or an unstable change of coordinates.

What exists

A PyTorch implementation, coverage-preserving vocabulary dropout, separation and redundancy controls, residual paths, invariant probes, and an empirical demonstrator are public. The object remains a research system because the decisive comparison is not whether it can emit logits. It is whether the moving backward geometry changes what the rest of the model can learn.