Skip to main content
Back to timeline
arXivSource publication:

A survey recasts attention evolution as contextual-memory organization, using 59 release records and 11 open-weight endpoints

Synopsis

Treating model-internal contextual memory as the shared unit of analysis, this survey introduces five dimensions—Memory Representation, Memory Update, Access, Readout, and Integration—and reviews five research lines (Softmax Attention, Sparse Attention, Linear Attention, State Space Models, and Hybrid Architectures), using 59 release-level records across 14 major model lineages and a frozen comparison of 11 high-performing open-weight endpoints through September 22, 2026 to argue that attention design is diversifying rather than converging, with hybridization and cross-layer reuse making network depth a dimension along which contextual memory is organized.

AI-generated editorial illustration: The Evolution of Attention in Large Language Models: Mechanisms, Trade-offs, and Emerging Trends

Interpretation

The survey proposes a five-dimensional lens—Memory Representation, Memory Update, Access, Readout, and Integration—asking what historical information remains represented, how it changes, what is eligible for the current query, how eligible memory is read, and how one or more readouts form the module output. Where prior surveys organize the literature by architecture family, computational complexity, long-context strategy, state dynamics, or KV-cache management, this one uses functional roles as a common vocabulary so that Softmax Attention, Sparse Attention, Linear Attention, State Space Models, and Hybrid Architectures can be compared without reducing them to one computational model. The lens is defined through abstract operators (memory schema, update, eligible view, readout, integration) with a terminology and symbol table; the authors state that the five dimensions are functionally distinguishable but not necessarily orthogonal or separately implemented, presenting it as an analytical vocabulary rather than a universal computational graph.

At the mechanism level, the explicit-memory route and the state-based route start from different memory forms yet expand their control scopes until they overlap: the explicit route centers on Representation and Access and extends into Update for bounded or multiresolution memories, into Readout when routing scores or summary estimates affect aggregation, and into Integration through head coordination and gating; the state-based route centers on Representation and Update and expands through additional rows, slots, segments, and state groups, making Access, Readout, and Integration more explicit. The survey decomposes linear-attention updates into retain/decay, erase/correct, and write/commit roles, tabulating progressions from no decay to fixed to input-dependent decay, from no explicit erase to coupled corrective to decoupled erase, and from direct to corrective to erase–write-decoupled writing, while a per-token objective view reinterprets the same recurrences as online optimization. The conclusion rests on unified update equations and comparison tables covering Linear Transformer, DeltaNet, RetNet, GLA, RWKV-5/6/7, HGRN2, GDN, Lightning Attention-2, DeltaProduct, KDA, GDN2, and EDA, plus capacity-expansion (SSE, SDM) and temporal-expansion (Log-Linear Attention, DLA, MHLA) comparisons.

At the architecture level, publicly documented LLM attention design diversifies rather than converges: of 59 release-level records, 58 are classifiable, 32 retain one principal attention or memory family, and 26 are classified as Hybrid; Hybrid records rise from none in 2022–2023 to 9 of 25 in 2024–2025 and 17 of 27 in 2026, dominated by local/global explicit-attention schedules and recurrent-state/explicit-attention schedules. The survey connects the mechanism review to two architecture evidence sets: a longitudinal release inventory and a frozen cross-section of 11 high-performing open-weight endpoints captured on September 22, 2026, which shows full GQA, latent-compressed sparse attention, block-sparse attention, local/global or compressed dense/sparse hybrids, and recurrent-state layers combined with full, latent, or sparse explicit retrieval coexisting near the frontier. The inventory draws primarily on official papers, technical reports, model cards, and released configurations, marking records with insufficient architectural disclosure as undisclosed rather than inferred; the authors state it is purposively curated rather than exhaustive or market-share weighted, and that leaderboard scores depend on scale, training, post-training, reasoning budget, and systems implementation, so the evidence characterizes adoption and coexistence rather than causal attribution.

Cross-layer reuse extends coordination from module placement to the depth-wise lifecycle of memory and routing artifacts: GLM-5.2 and GLM-5.3 reuse retrieval indices or top-k candidate decisions through IndexShare/IndexCache, DeepSeek-V4.1-Flash reuses selected global KV representations and indexer keys through CSA2's Full, Reindex, and Reuse modes, and LongCat-2.0 and LongCat-Flash-Lite-Sparse reuse one layer's selected token set via LSA's cross-layer indexing. This introduces a distinction between capability placement—which memory functions operate at different depths—and artifact-lifecycle coordination—which memory or routing artifacts persist across layers and when they are reused or refreshed; five of the 27 records in 2026 carry the cross-layer attribute, corresponding to three architectural patterns across GLM, DeepSeek, and LongCat systems. The conclusion comes from the five cross-layer release-level records in the inventory and their public technical documentation; the authors caution that the cross-layer attribute overlaps the mutually exclusive single-family/hybrid classification and should not be added to those columns.

Perspective

The survey targets model-level mechanisms in autoregressive LLMs and closely related causal sequence mixers, covering publicly available work through September 22, 2026; test-time learning, retrieval from external databases, multimodal-specific memory designs, and implementation-only optimizations remain outside the core taxonomy unless they directly alter contextual-memory semantics. The longitudinal inventory is purposively curated rather than exhaustive or market-share weighted, and records with insufficient architectural disclosure are marked undisclosed; the frontier comparison is a frozen September 22, 2026 snapshot used to characterize adoption and coexistence rather than causal attribution. The five-dimensional lens is an analytical vocabulary whose dimensions the authors explicitly note need not be orthogonal or separately implemented, so it works best as a comparison and design checklist rather than a unified computational graph.

The stateful multidimensional memory-routing hypothesis remains forward-looking, and its feasibility depends on jointly coordinating write and read routing over one address space and on allocating sparse-write and sparse-read budgets; the text offers no empirical validation of the hypothesis. Cross-layer reuse currently appears in only five release-level records corresponding to three architectural patterns, so its benefits and applicability conditions need more public cases. The inventory and frontier comparison are descriptive evidence at a specific date, and leaderboard scores depend on model scale, training, post-training, reasoning budget, and systems implementation, so no attention mechanism can be credited with higher model quality on this basis. In addition, this reading is full-text but figures appear as placeholders, so node fills, context-length tiers, and visualization details in Figures 14 and 15 cannot be verified and judgments about those figures should defer to the original.

Sources