Blind acceleration rings
Old motion is reused after the problem seen by the next update has changed.
FROM BFFT TRANSPORT RESEARCH · EXPERIMENTAL PYTORCH OPTIMIZER
Acceleration by refusing the motion that makes the trajectory spend its next hundred steps undoing the present one.
Ordinary momentum remembers a vector. Anchor remembers a request and the directional frame in which that request was learned.
When the gradient turns, reusing the old request unchanged confuses persistence with correctness: part of the stored motion still follows the valley, while another part now points across it. Anchor rotates the stored request into the new frame, combines it with the live gradient, identifies the component created by their disagreement, and removes only that component.
The resulting step can be smaller than the unrestrained proposal and still arrive sooner. The claim is about the trajectory’s relaxation time—not about making any single update larger.
BFFT’s Meyer experiment alternates two convex image subproblems: one requests cartoon structure and the other requests bounded oscillatory texture. The published-style nested solve can spend hundreds or thousands of inner iterations repeatedly rebuilding nearly the same answer. Naive Anderson or heavy-ball acceleration is unreliable because the object being accelerated changes meaning when the other block moves.
The decisive improvement came from treating the inner Bregman fields as carried transport state: keep the useful state alive, re-enter it through the current subproblem, interleave the two relaxations, and constrain proposals that do not belong to the present geometry. On the motivating workload, a process normally requiring more than 400 iterations fell to roughly 30; fusion pushed it lower again.
The experiment exposed the reusable idea: useful progress is a request with a history, and that history has to be transported when the local frame turns. Anchor applies that principle to gradient optimization.
Old motion is reused after the problem seen by the next update has changed.
State must be moved into a common frame before it can condition current motion.
Measured geometry may suppress an incompatible request; it may not invent unsupported thrust.
Carry the slow request, read it against the live gradient, remove its turn-crossing component, then apply it locally.
A slow state stores the update request that persisted across prior gradients. A separate signature records the directional frame in which that state currently lives.
The current unit gradient and carried signature define a spherical midpoint. Anchor applies the minimum rotation that moves the slow state into that midpoint frame. At a 180° ambiguity it resets rather than fabricating a rotation.
The transported slow request and raw current gradient form a fast proposal. Anchor projects that proposal onto u − r—the axis separating the current gradient from the carried signature—and attenuates only that projection according to the measured turn energy.
The remaining request is scaled so its proposed displacement is at most 2% of the row’s parameter scale. Braking is immediate; recovery is deliberately slow. Other rows keep their own stability budgets.
Transferred: state has a frame; compare only after transport; restraint may condition a requested motion; iteration count is the target.
Not transferred: no image prior, G-norm, Bregman divergence, Hessian, loss lookahead, or second-moment estimator appears in Anchor.
Move the turn control. The carried signature is blue, the current unit gradient is red, transported memory is cyan, the unrestrained lead request is amber, and the applied request is white. The gap between amber and white is the component removed along the new turn.
Each matrix row is a local cell. Normalize its gradient to u, combine it with the carried frame r, and set the next signature to the unit-ball projection of r + u. Except at degeneracy, its direction is the spherical midpoint q.
The minimum spherical rotation moves the stored slow request from r to q. If r and u are antipodal, the midpoint and shortest transport are not unique, so that cell resets instead of inventing a turn.
A fixed two-timescale recurrence mixes transported memory with the current gradient. The fast state reaches the parameters; the slow state remains the carried request.
Anchor projects the fast request onto the current-gradient minus carried-signature axis and subtracts that component in proportion to the turn energy. It does not damp every coordinate.
If the proposed update would move a cell by more than 2% of its parameter scale, the local multiplier drops immediately. It recovers only 5% toward the currently admissible rate per update.
INVARIANT Directional restraint cannot increase the norm of the fast proposal. The stored slow request is transported, but never clipped by the restraint operator.
Let r be the carried signature frame, u the current normalized gradient, and q their normalized midpoint. The constants are fixed: c = e−1 and a = 1 − c.
s̄t = T[r → q](st)Carry the slow request by the minimum rotation.
ft = a s̄t + c gtRead the transported state against the present gradient.
st+1 = a s̄t + c ftStore a slower request; do not store the restrained output.
d = u − r, τ = ½(1 − u·r)The turn axis and its half-angle energy come from consecutive signatures.
vt = ft − τPdftSubtract only the component that conflicts with the live turn.
ht = min(1, .02 bt / ηmax‖vt‖)Here bt = max(‖pt‖, 1) for the local cell.
λt = ht if ht < λt−1; otherwise λt−1 + .05(ht − λt−1)The active multiplier never exceeds the current admissible multiplier.
ηmaxλt‖vt‖ / max(‖pt‖,1) ≤ 0.02A troubled cell brakes without spending another cell’s stability budget.
No extra loss evaluation or closure is required.
Anchor does not estimate coordinatewise gradient variance.
The battery uses one ceiling and one trust radius everywhere.
Rows brake independently; vectors form one cell.
Twenty-four tasks, vanilla and self-context networks, three paired seeds, width 24, 500 updates, batch 256, and validation every five updates. AdamW uses 0.003, SGD uses 0.03, and Anchor uses the unchanged 1.0 ceiling. Initialization and minibatch schedules are identical inside every comparison.
Area under the validation-acquisition curve. Higher means the optimizer acquired useful behavior in fewer updates.
Positive mean AUC margin over the stronger baseline and paired-seed wins in at least two of three seeds.
Updates required to reach the best validation score attained by either baseline on the same task, model, and seed.
Held-out continuation is reported independently. It is never folded into the acquisition claim.
Across the unfiltered battery, Anchor finishes second in mean acquisition AUC: 0.7280, above SGD’s 0.6343 and below AdamW’s 0.7712. It satisfies the paired winner rule on 9 of 48 task × architecture pairs with no numerical failures. Seven wins occur in the vanilla MLP and two in self-context; five of those acquisition winners also lead held-out score. The distribution supports a specialist method with identifiable favorable geometries.
Means across every task, architecture, and seed show broad central tendency. The paired winner register below answers the narrower question of whether a specific problem is learned faster.
| Optimizer | Runs | Acquisition AUC | Best validation | Held-out | Failures | Mean fit time |
|---|---|---|---|---|---|---|
| Waiting for the frozen battery. | ||||||
Measured curves populate after every paired fit completes.
The interactive chart is limited to the six largest AUC margins for legibility. This register is complete.
Anchor’s first contact is 40× earlier and sustained residence begins 13.3× earlier. A 1.5 ceiling touched the floor at update 50 but rang until 2,700; 1.0 remains the standing optimizer because the target is fewer useful updates, not a transient crossing.
Interpretation boundary. A task that no optimizer learns is not evidence of optimization superiority. A validation win is not an extrapolation win. A fast crossing that immediately leaves the basin is not a speed win. Conversely, poor held-out continuation cannot retroactively erase faster acquisition on the sampled domain. The page keeps those questions separate.
The public file is copied byte-for-byte from the research repository. Gradient clipping remains outside the optimizer so every comparator receives the same training-loop treatment.
from anchor import Anchor
optimizer = Anchor(
model.parameters(),
lr=1.0,
trust_radius=0.02,
)
optimizer.zero_grad(set_to_none=True)
loss.backward()
optimizer.step()
Loading anchor.py…