Adding a paragraph coordinate to attention makes real paragraphs compress deeper than random labels, but depth varies by corpus and no corpus-only statistic fully reproduces it
Synopsis
Using a hierarchical rotary positional encoding (hRoPE) that splits paragraph, sentence, and token indices into separate channels, the authors intervene on the paragraph coordinate while holding the token sequence fixed and measure cross-paragraph attention with a token-distance-exact estimator, finding that attention is compressed in all three corpora but that a density-matched random-label channel is compressed too; what distinguishes real structure is the depth of compression, which is greater and corpus-dependent while the control's is not, and across eight corpus-only quantities spanning lexical persistence, paragraph length, and embedding-based coherence, none fully reproduces the cross-corpus ordering of depth, though embedding-based coherence comes closest.
Interpretation
With the token sequence fixed and only the paragraph coordinate changed, attention changes, showing that hierarchical paragraph position is a causal factor in attention rather than an artifact of token content. Prior positional-encoding work treats position as a one-dimensional reading-order coordinate, or uses hierarchical coordinates for generation and efficiency; here the paragraph coordinate is isolated for counterfactual intervention, paired with a density-matched random-label control and a periodic mirror control. Three corpora (WikiText-2, OpenWebText, Python code) with three seeds each; merging paragraphs compresses attention and splitting dilates it; flat and sent_axial show no effect by definition; rand_axial responds markedly more weakly. The single-example demonstration is explicitly labeled by the authors as an illustrative example, not a statistical result.
Compression alone is not diagnostic of genuine structure: a density-matched random-label channel is compressed in every corpus too, only more shallowly; depth, not location, distinguishes real from random structure. Separates the phenomenon of attention being compressed near paragraph boundaries from the question of whether that compression indicates genuine hierarchical structure, and reports the depth gap under a paired document-cluster bootstrap. The paired bootstrap interval for the hrope_axial versus rand_axial depth gap excludes zero in Code and WikiText-2, and includes zero in OpenWebText, which the authors describe as not resolvable at their seed count and treat as evidence in neither direction.
Compression depth is corpus-dependent while the random control's depth tracks the corpus less strongly, yet none of eight corpus-only quantities fully reproduces the cross-corpus ordering of depth. Compares the model-derived depth against candidates computed entirely from the corpus without any model representations or attention values, and reports that an earlier apparent correspondence, based on non-comparable runs and an unsound distance-bin pooling procedure, does not survive correction. Eight quantities across three constructs: five lexical-persistence variants (one a rescaling that shares ordering by construction and is not an independent data point), median paragraph length, and two summaries of one embedding-coherence curve; the authors count constructs rather than rows and conclude lexical persistence and paragraph length fail while embedding coherence comes closest but leaves one of three corpus pairs unresolved.
In the two corpora with an identifiable interior minimum, the response strengthens before reversing, a shape that cannot be produced by an additive superposition of a monotone-decaying attraction and a monotone-growing suppression. Uses a proposition to rule out the simplest additive competing-process family, locating the turnover in some interaction rather than treating it as a trivial sum of two independent monotone processes. Proposition 5.1 proves the discrete first difference is strictly positive; the authors state this rules out only the additive family, leaving softmax-competitive and non-monotonic alternatives untested and flagged as the main open mechanistic question.
Perspective
This work is aimed at readers studying positional encodings and long-document modeling, and applies to an 8-layer, 512-dimensional, context-1024, 5000-step small-scale setting across WikiText-2, OpenWebText, and a Python source corpus, with paragraphs defined corpus-specifically (delimited prose blocks, blank-line-delimited source blocks). It enables follow-up work to treat the paragraph coordinate as an independently intervenable variable when testing whether a model uses hierarchical position, to use density-matched random labels and periodic mirrors to separate channel presence from channel content, and to track compression depth rather than compression location as the reproducible quantity. The authors also note that an analogous sentence-level signature is not addressed, that periodic coordinates are not testable under this protocol, and that production scale and downstream metrics are not evaluated.
The cross-corpus ordering is descriptive: the document-level bootstrap quantifies within-corpus uncertainty only, and the three corpora are fixed by choice rather than drawn from a common population, so Code deeper than WikiText-2 deeper than OpenWebText establishes that the corpora differ in depth, not that the specific order generalizes. On OpenWebText the real-versus-random depth gap is unresolvable at the current seed count, and the authors explicitly treat it as evidence in neither direction, unable to separate a true null from a power limitation. The origin of compression depth remains open: the three tested constructs do not reproduce the ordering, but the authors rule out only those specific accounts, not other corpus properties or the architecture and training dynamics themselves. Mechanistically, only the additive competing-process family is ruled out, leaving softmax-competitive and non-monotonic alternatives untested. In addition, WikiText-2's depth location is structurally non-identifiable and its depth shifts by about one gate's confidence-interval width across sampling thresholds, and the masking probe uses a single checkpoint per corpus and is exploratory. A reader who sees only the abstract without figures and appendices cannot check the threshold sensitivity, the seed-level directional evidence, or the probe details, which is an open point to keep in mind.
