Skip to main content
Back to timeline
arXivSource publication:

Weakening the skip connection causes a depth-induced rank collapse that no loss term repairs, while restoring the skip reopens the gradient path

Synopsis

The authors collapsed small transformers into depth-induced rank collapse by weakening the skip connection and then treated copies with two rank-based loss terms differing only in whether their gradient vanishes; neither repaired the collapse even though the bounded force reached about a tenth of the task gradient, because the task gradient no longer reached the query and key weights that decide where attention looks, whereas restoring the skip connection, a runtime action changing no weight, reopened that path at once, after which the rank recovered but only far above the scale of collapse and with the loss still above a healthy network.

Source-provided article image: Force without transmission: a depth-induced rank collapse that no loss on the representation reopens
Figure 1 ·

Figure 1: Stable rank of the last block after the treatment is switched on, all treatments and collapse events of the main study (50-step median, as in the recovery rule; dotted line: recovery threshold). Only the reversed ramp recovers.

arXiv

Interpretation

In depth-induced rank collapse driven by a weakened skip connection, neither rank-based loss term recovered the rank: penalty 0/8 and hinge 0/8 in the main study, none in the replication, and untreated copies did not recover either. A bounded-force loss term had repaired an established attention-entropy collapse from saved checkpoints; this work transfers that idea to rank collapse and tests whether it holds. Ten collapse events across two depths and five seeds, plus a replication with five new seeds per depth; collapse was defined as a last-block stable rank below 1.5 after settling, and recovery as the 50-step median rank reaching half the healthy value.

The bounded force was genuinely large yet ineffective: the hinge's force at the parameters had a median of 0.0887 of the task gradient, about 28 times the penalty's 0.0032, and 0.3071 at four times the strength, with the replication at 0.0782 and 0.0037. Separates the size of a corrective force from whether it arrives, showing that a term's magnitude at the loss does not represent its effect. Force was measured at the parameters as the ratio of added to task gradient, with medians and ranges per collapse event; even with the added gradient exempt from clipping and strength raised up to ten times, 0 of 32 branches recovered.

Block-by-block measurement showed the task gradient barely reached the query and key weights: in 17 of 18 networks it was below threshold in every block, against a per-block level in healthy networks, falling by more than three orders of magnitude from the last block to the first and sitting almost entirely on the value projection. Extends the result Noci et al. derive at initialisation to trained networks and provides a per-block arrival profile inside the collapsed network. Measured without any training on all 18 saved collapsed networks, block by block, for both the task gradient and the hinge gradient, with the same-seed network trained without the driver as the healthy reference.

Restoring the skip connection changed no weight yet reopened the path at once: at the first step after the scale was set to 1, before any update, the task gradient on query and key rose by about an order of magnitude even in the block where it rose least and reached 1.2 times the healthy value in the median block; rank recovery required the skip scale to return to 0.25–0.97 at depth 12 and higher at depth 24, at least 12 times above the point of collapse, a hysteresis. Reframes repair from how much force is added to whether the gradient path is restored, and specifies the scale and timing conditions for recovery. Every network recovered on the reversed ramp, 8/8 and 10/10; in immediate-switch branches 7/8 and 8/8 recovered at 0.3 and 0.6 while only 2/8 did at 0.1; switching after delays of 500, 1500 or 3000 steps recovered later or not at all, and no delayed depth-24 branch recovered.

Perspective

This work addresses settings where a depth-induced rank collapse has already occurred during training and the architecture cannot be changed, showing that the available lever is restoring the gradient path rather than increasing a loss term. For practitioners, measuring block by block where the gradient arrives is cheap and needs no training, and can distinguish a collapse with a closed path from a rank drop with the path still open; a runtime action such as restoring the skip connection reopens the path without changing any weight, but the scale must be returned well above the point of collapse and earlier treatment works better. The authors also note that under a learning-rate burst with the skip intact the path stayed open and the rank recovered untreated, providing a contrast for choosing an intervention.

Several open questions remain for a careful reader: the gradient path was measured on one probe batch per network, the registered tests rest on ten and eight held-out networks of one system, and the design was fixed before the main runs but not independently timestamped. What changes with time in the collapsed state was not identified, and the attention pattern did not change. No settled collapse was produced with the skip intact, so whether such a collapse would close the path is untested. In addition, networks that recovered by rank still ended with a loss above their control, so the rank reported recovery while the function had not recovered, leaving how to judge a genuinely recovered run an open question.

Sources