Observe the boundary
Build a self-context hidden path inside an otherwise ordinary network, declare the first matrix maps that consume each nonlinear representation, and let Residual Polar Transport adapt geometry only at those boundaries.
- Form
- Implementation guide
- Primary example
- Natural self-context MLP
- Extension
- Pre-norm transformer block
- Default
- Matrix Transport everywhere
The model chooses a chart. The optimizer decides how to move through it.
A self-context block lets an activation construct a local metric, allocate responsibility across fixed projection frames, and reread itself through that allocation. Its forward pass can therefore change the useful coordinate structure from one example to the next. The optimizer has a different job: carry accumulated update intent across changing gradient frames without allowing old momentum to reverse the live descent decision.
The two mechanisms meet at a nonlinear boundary. Before that boundary, observing several consecutive affine maps mostly repeats one residual relation in different coordinates. After it, the activation Jacobian is data-dependent: the same incoming direction may be stretched, folded, suppressed, or divided among different self-context charts. That is new geometry worth measuring.
Observation is consequently selective. It does not enable the optimizer, and it does not decide whether a parameter learns. It only permits a matrix to replace part of its transported request with a norm-matched polar request when the local training residual demonstrates persistent spectral concentration.
This division matters. Self-context changes what the network can represent in one forward pass. Residual Polar Transport changes how the parameters approach a solution over many backward passes. The optimizer never substitutes its state for the activation’s chart, and the model never edits momentum. They exchange exactly one local fact: whether the post-nonlinearity residual operator is still directionally narrow.
One nonlinear boundary, two observed branches
In the natural MLP below, only down.base and down.metric are observed. They are parallel contractions of the representation produced by the main LELU, so each receives the same source activation but a different output cotangent during backpropagation.
Construct a representation
- EmbedMatrix Transport
- Up self-contextMatrix Transport
- LELUNew tangent geometry
- down.baseObserved down.metricObserved
- OutputMatrix Transport
Measure the unresolved operator
Xsaved post-LELU activation(X, Δbase)base-branch cotangent(X, Δmetric)metric-branch cotangentρbase, ρmetriclocal normalized entropy ranksWhy these matrices carry the observation
Write the hidden state entering the second self-context block as a = LELU(up(embed(x))). The base branch computes an ordinary affine reading Wba. In parallel, the metric branch computes the factor from which that block constructs its activation-conditioned positive-semidefinite metric. Both branches are therefore immediate matrix contractions of the same newly nonlinear state, but they do different work and receive different cotangents. Combining them into one observation would erase that distinction; observing them separately preserves it.
The embedding and the matrices inside the first block are upstream of the main LELU boundary. Observing all of them would repeatedly describe residual geometry before the representation has undergone another material nonlinear change. They still receive the complete Matrix Transport update—two-sided matrix scaling, transported momentum, live-axis closure, and the Euclidean descent certificate—but their polar weight remains zero.
The scalar response heads and scalar output head are ineligible for a different reason. Their residual operators have only one singular value, so normalized entropy rank is identically one whenever the operator is nonzero. A one-dimensional spectrum can report magnitude, which the optimizer already sees, but it cannot report anisotropy. A multidimensional output head could be observed in another architecture if it were the first contraction after a new nonlinearity; scalarity, not the word “output,” is the exclusion.
Derive sites from the graph, not from layer names
- Mark operations whose Jacobian depends on the current example: activations, normalization with live statistics, softmax mixing, multiplicative gates, self-context allocation, routing, or recurrence.
- Follow each resulting branch until its first trainable matrix contraction. That matrix is eligible because it is the first parameterized reader of the new representation.
- If the representation fans into parallel readers, retain each reader separately. Stop following that branch until another nonlinear transformation creates a new local geometry.
- Discard candidates whose smaller matrix dimension is one. If a fused functional call hides an otherwise valid matrix, expose it as a module or call the optimizer’s explicit observer.
The observer measures shape, not loss
For a selected linear map, the observer regresses the output cotangent against the activation before the minibatch is contracted into a single weight gradient. Centered batches and a relative ridge make the statistic insensitive to offsets and stable near weak directions.
Consider one observed linear module with input width n, output width m, and B samples after flattening any leading token or spatial dimensions. The forward hook retains the centered activation matrix X ∈ ℝB×n. During backward, the module receives the centered output cotangent Δ ∈ ℝB×m. Ordinary backpropagation will soon contract them into the weight gradient G = ΔTX / B. That contraction is sufficient to update the weight, but it no longer says whether the residual relation was broad across directions or concentrated in a small operator tail.
O = (XTX + λsI)−1XTΔThe observer pauses before that information is destroyed. It solves a small regularized regression from activation coordinates to output-cotangent coordinates. The relative scale s = tr(XTX)/n makes the ridge follow activation energy instead of imposing an absolute unit. Fixed deterministic sketches bound both feature dimensions before this solve when either exceeds the configured observer width. They change the resolution of the estimate, not the training examples or objective.
The singular vectors of O locate coupled activation/residual directions; its singular values say how much energy each coupled direction carries. Absolute singular magnitude is deliberately discarded. If the entire loss is multiplied by a constant—as under AMP scaling—the singular values scale together and the resulting distribution is unchanged.
pi = σi(O)2 / Σjσj(O)2ρ = exp(−Σipi log pi) / min(n,m)When singular energy is evenly spread, the entropy effective rank approaches the full smaller dimension and ρ approaches one. When a few singular directions dominate, effective rank falls and so does ρ. The statistic therefore distinguishes isotropic unresolved work from a residual that remains trapped in a narrow set of coupled directions. It does not say whether the loss is large, whether the prediction is correct, or whether training should stop.
α* = 1 − ρThe complement α* is a target, not a hard switch. Each observed parameter low-pass filters it into its own state. The same smoothing path operates when anisotropy rises and when it releases, so a matrix may approach polar shaping and later return to Matrix Transport. No schedule, task label, or validation metric can force that transition.
Transport remains the default update
Residual observation controls only the last shaping decision. The underlying request is built for every parameter, observed or not, by a non-elementwise matrix metric and transported momentum. The polar branch cannot silently enlarge the step because the two requests are norm matched before interpolation.
For a matrix gradient G, the optimizer maintains exponentially averaged row and column second moments rather than one second moment per element. Their inverse fourth roots whiten the gradient from both sides. This produces a live request in a matrix-aware coordinate system while retaining the original parameter shape.
Momentum cannot simply be reused in those coordinates. As the live request rotates, yesterday’s stored vector contains both motion along the current route and motion transverse to it. Minimum-rotation transport moves the history into the new frame; frame agreement determines how much transverse history remains trustworthy. A live Nesterov component then closes the identifiable lag along the current axis. Only after those operations does residual observation decide whether an observed matrix should equalize some singular-value pressure.
The polar factor is computed from the transported request, not directly from the raw gradient. It is rescaled to the transported request’s Frobenius norm before mixing, and the blend is normalized again. Thus α changes how the step distributes effort across singular directions, not how large the step is. A final inner-product certificate removes any component that would oppose the current gradient.
- 1
Read the gradient as a matrix
Row and column covariance factors form
L−1/4 G R−1/4. This replaces an elementwise second moment with a two-sided matrix metric. - 2
Carry momentum into the live frame
The previous request is moved by the minimum rotation between consecutive normalized live requests. Transverse history is retained only in proportion to frame agreement.
- 3
Close the live axial discrepancy
A Nesterov-style live component corrects identifiable lag along the present request axis without allowing stored momentum to reverse the current decision.
- 4
Shape only an observed anisotropic matrix
The transported request and its Newton–Schulz polar factor are norm matched and mixed by the smoothed weight α. Unobserved matrices have α = 0.
- 5
Certify Euclidean descent
If numerical composition produces a request with negative gradient inner product, the opposing component is removed before the parameter update.
Remain on Matrix Transport.
Approach norm-matched polar shape.
Put self-context inside an ordinary MLP
The base path remains affine. A learned local metric allocates a fixed bank of unprivileged projections, lifts their weighted response back into activation coordinates, and uses that bounded context to reread the same activation. The correction is residual: the layer can decline to use it.
The layer begins with a conventional affine map base(x). That path is never removed, so the module can behave like an ordinary linear layer when its structural correction is not useful. Beside it sits a fixed bank of low-rank projection frames. The frames are seeded without semantic labels; they are possible views, not a supplied Fourier basis, rotation generator, or task coordinate system.
The activation generates a factor A(x), and the layer forms M(x)=A(x)A(x)T/r. This guarantees a positive-semidefinite local metric while allowing its orientation and condition to vary with the example. Each fixed frame projects the same authentic activation. Metric cost, projection energy, signed mean, and absolute response become evidence for a soft allocation over those views.
The allocated projections are lifted back into activation space and normalized to the authentic activation’s RMS. A strength of 0.25 means the layer proposes a chart point one quarter of that normalized context away from x. Crucially, repeated context is anchored to x; the proposal does not recursively accumulate and drift. The allocation is recomputed at the proposed chart point, pooled through a shared low-rank map, and added as a learned residual correction to base(x).
Backpropagation remains exact in the setup below. The optimizer does not detach the self-context path or temper its Jacobian. Its observer merely records the activation and the cotangent already passing through the selected readers. Model expressivity and optimizer geometry therefore remain separable experimental variables.
def forward(self, x):
# A learned positive-semidefinite metric reads this activation.
factor = self.metric(x).view(batch, rank, rank)
metric = factor @ factor.transpose(1, 2) / rank
# Fixed projection frames have no task-specific semantic names.
projected = einsum("dri,bi->bdr", self.primitive, x)
cost = einsum("bdr,brs,bds->bd", projected, metric, projected)
weight = softmax(self.response(projected, cost) - normalize(cost), dim=1)
# Lift the allocated views, bound them to the authentic activation scale,
# and reread the chart from an anchored proposal rather than accumulating drift.
context = einsum("bd,bdr,dri->bi", weight, projected, self.primitive) / rank
context = context * rms(x).detach() / rms(context)
chart_x = x + self_context_strength * context
_, projected, weight = self.allocate(chart_x)
pooled = einsum("bd,bdr->br", weight, projected)
# The ordinary affine map survives; context contributes a residual correction.
return self.base(x) + softplus(self.scale) * (pooled @ self.shared)
import torch
from ML_experiment.models import LELU, SoftEikonalLinear
class ContextMLP(torch.nn.Module):
def __init__(self, input_dim, output_dim, width=128):
super().__init__()
self.embed = torch.nn.Linear(input_dim, width)
self.up = SoftEikonalLinear(
width, 2 * width,
directions=12,
rank=4,
self_context_strength=0.25,
)
self.activation = LELU()
self.down = SoftEikonalLinear(
2 * width, width,
directions=12,
rank=4,
self_context_strength=0.25,
)
self.output = torch.nn.Linear(width, output_dim)
def forward(self, x):
x = self.embed(x)
x = self.activation(self.up(x))
x = self.down(x)
return self.output(x)
def optimizer_observation_modules(self):
# These parallel maps first consume the post-LELU representation.
return (self.down.base, self.down.metric)
from ML_experiment.optimizer import ResidualPolarTransport
model = ContextMLP(input_dim=n_features, output_dim=n_targets)
optimizer = ResidualPolarTransport(
model.parameters(),
lr=3e-3,
weight_decay=1e-4,
)
observed = frozenset(model.optimizer_observation_modules())
optimizer.attach_model(
model,
module_filter=lambda module: module in observed,
)
try:
for x, target in loader:
optimizer.zero_grad(set_to_none=True)
prediction = model(x)
loss = criterion(prediction, target)
loss.backward()
optimizer.step()
finally:
optimizer.remove_observers()
The runtime contract
- Evaluation
- Forwards inside
torch.no_grad()are ignored. A grad-enabled forward abandoned before backward should be followed byclear_observations(), otherwise its retained activation could be consumed by a later backward pass. - Checkpoints
- Optimizer moments, covariance factors, cached roots, entropy ranks, and polar weights serialize normally. Python hook handles do not. Attach the restored optimizer to the restored model before the next observed forward pass.
- Gradient accumulation
- Several backward passes before one
step()are supported. Each selected parameter accumulates its local rank observations, and the optimizer averages them when that update is applied. - AMP and DDP
- Uniform AMP loss scaling leaves entropy rank unchanged. Distributed workers observe local minibatches; synchronize an explicit rank when every replica must maintain exactly identical optimizer state.
- Compilation
- Full-backward hooks can create graph breaks under
torch.compile. A compiled or fused model can callobserve_residualexplicitly with the same activation, cotangent, and target parameter instead. - State cost
- The reference keeps full row and column covariance factors, requiring
O(rows² + columns²)state and periodic dense symmetric eigendecompositions. Large transformer matrices require blocked or low-rank roots before this is a credible production optimizer.
The same boundary rule extends to a transformer
Attention and a gated feed-forward path both create nonlinear, data-dependent representations. The first explicit matrix after attention mixing and the two matrix readers after the self-context activation are natural observer sites. This is an architectural translation of the rule, not yet a transformer benchmark claim.
The query, key, and value projections are upstream of attention’s nonlinear mixing. Softmax converts their pairwise scores into example- and token-dependent routing weights; the attended value field is the new representation. An explicit attention_output matrix is therefore the first parameterized reader after that boundary. Its observer sees token rows as additional samples because the hook flattens all leading dimensions.
The residual addition has no trainable matrix and creates no observer site. After the second normalization, the self-context feed-forward path expands the token through context_up, applies LELU, and sends the result to two parallel readers inside context_down. Those readers—not every matrix inside the block—receive local residual-rank state. The QKV projection, the upstream self-context matrices, normalization parameters, biases, and residual paths continue to use the transport update without polar shaping.
This explicit declaration is preferable to optimizer-side graph guessing. Real transformer implementations fuse attention, use functional linears, share weights, or route tokens conditionally. The model author knows which operations create new representations and can expose those boundaries as modules. The optimizer only needs a set of module identities; it does not need transformer-specific code.
import torch.nn.functional as F
class ContextTransformerBlock(torch.nn.Module):
def __init__(self, d_model, heads, dropout=0.0):
super().__init__()
assert d_model % heads == 0
self.heads = heads
self.head_dim = d_model // heads
self.dropout_p = dropout
self.norm1 = torch.nn.RMSNorm(d_model)
self.qkv = torch.nn.Linear(d_model, 3 * d_model, bias=False)
# Kept explicit so the optimizer can observe the post-attention map.
self.attention_output = torch.nn.Linear(d_model, d_model, bias=False)
self.norm2 = torch.nn.RMSNorm(d_model)
self.context_up = SoftEikonalLinear(
d_model, 4 * d_model, self_context_strength=0.25
)
self.activation = LELU()
self.context_down = SoftEikonalLinear(
4 * d_model, d_model, self_context_strength=0.25
)
self.dropout = torch.nn.Dropout(dropout)
def forward(self, x):
batch, tokens, width = x.shape
h = self.norm1(x)
qkv = self.qkv(h).view(
batch, tokens, 3, self.heads, self.head_dim
).permute(2, 0, 3, 1, 4)
q, k, v = qkv.unbind(0)
mixed = F.scaled_dot_product_attention(
q, k, v,
dropout_p=self.dropout_p if self.training else 0.0,
)
mixed = mixed.transpose(1, 2).reshape(batch, tokens, width)
x = x + self.dropout(self.attention_output(mixed))
h = self.norm2(x).reshape(batch * tokens, width)
h = self.context_down(self.activation(self.context_up(h)))
return x + self.dropout(h.view(batch, tokens, width))
def optimizer_observation_modules(self):
return (
self.attention_output,
self.context_down.base,
self.context_down.metric,
)
# A transformer owns the boundary declaration; the optimizer consumes it.
observed = frozenset(
module
for block in model.blocks
for module in block.optimizer_observation_modules()
)
optimizer = ResidualPolarTransport(model.parameters(), lr=3e-3)
optimizer.attach_model(
model,
module_filter=lambda module: module in observed,
)
What changes in a real transformer implementation
Some attention kernels use an output weight through a functional call instead of invoking an nn.Linear module. Automatic module hooks cannot see that call. Keep the post-attention projection explicit, wrap the functional contraction in a small module, or supply the activation/cotangent relation through observe_residual.
For a stack, collect the declared modules from every block into one identity set and pass that set to attach_model. This produces several independent local controllers, not a single transformer-wide polar switch. A narrow residual in one late attention block cannot force an unrelated early feed-forward matrix toward polar geometry.
The example is appropriate for architectural experiments at small and medium width. It is not a claim that the reference root backend is ready for a billion-parameter model. The observer already uses bounded sketches, but the optimizer’s full matrix covariance state remains the dominant scaling problem.
What remains unresolved
The placement rule has been tested on the standing self-context MLP architecture, not yet on every model family. A residual observer should be attached to the first eligible matrix after each material nonlinear representation change; architectures with recurrent state, attention, shared modules, or several simultaneous charts may expose more than one such boundary.
The next engineering experiment is temporal rather than architectural. The selected entropy ranks stabilize quickly, so measuring them every four to eight updates—and forcing an early refresh only after substantial gradient-frame rotation—may remove most remaining observer cost without changing the optimizer’s geometry.
Research record
The complete layer implementation, optimizer, tests, and research record remain in BFFT, the authoritative repository.