DepthBench sweeps width–depth ratio across 10 architectures at fixed parameter budget: HC and Full AttnRes keep lowering validation loss at extreme deep-narrow shapes while Pre-LN degrades
Synopsis
The work introduces DepthBench, a controlled benchmark that varies the width–depth aspect ratio (d_model/n_layer) from shallow-wide to deep-narrow under a fixed parameter budget and pre-training recipe, comparing 10 representative residual and normalization architectures, and finds that the benefit of depth is strongly architecture-dependent: Pre-LN and most of its norm- and scaling-based variants worsen as models get deeper and narrower (Pre-LN rises from 2.759 to 2.782), whereas HC and Full AttnRes improve consistently (HC from 2.729 to 2.699, Full AttnRes from 2.751 to 2.718), accompanied by stronger cross-layer causal dependence and order sensitivity, though deeper models also incur compute and memory overhead.
Interpretation
DepthBench treats depth as a capacity-allocation choice: with total parameters approximately fixed (400M total in the main benchmark, 7 shapes, aspect ratios from 76.0 to 9.1) and the pre-training recipe held fixed, it systematically varies the width–depth ratio and reports each configuration at its best-performing learning rate after a sweep, separating the effect of depth from model size and suboptimal learning rates. Previously LNS, KEEL, AttnRes, HC and mHC were evaluated with different architectures, training recipes and system configurations, so their gains could not be cleanly attributed; this work compares 10 residual designs on a shared LLaMA-like backbone, the same data (FineWeb-Edu, 20 tokens per parameter) and the same optimization settings. Main benchmark at 400M total parameters, 7 shapes, 10 architectures; plus an iso-backbone control fixing the backbone at 300M, a 200M–500M multi-scale suite, and a 1.6B validation suite; architecture-specific learning-rate sweeps at every aspect ratio.
Depth benefits depend strongly on residual connection design: Pre-LN and most normalization/scaling variants (Sandwich-LN, LNS, DeepNorm, KEEL, MoDA) are insensitive or even unfavorable to deeper-narrower shapes, with optima at wider or intermediate shapes, whereas HC and Full AttnRes improve consistently as depth grows, up to 70 layers and an extreme aspect ratio of 9.1. This provides a controlled ranking result: not every method that improves normalization or residuals converts architectural depth into effective computational depth; the residual topology itself is the key variable. Validation-loss numbers: Pre-LN rises monotonically from 2.759 to 2.782; Full AttnRes falls from 2.751 to 2.718; HC falls from 2.729 to 2.699; with backbone size fixed, deep-narrow models improve even as total size decreases, indicating the gain comes from the width–depth allocation rather than a larger backbone.
The gains extend beyond pre-training loss: deeper HC and Full AttnRes generally achieve lower teacher-forced NLL on coding, STEM and math tasks, with the trend especially clear for STEM and math, indicating that the width–depth advantage translates into domain-specific predictive capability. It moves the criterion for whether depth is effective from a single validation loss to domain-task NLL, and shows the trend persists across 200M–500M and 1.6B scales (Full AttnRes more stable; HC less stable at larger scale, with gradient explosion at the largest aspect ratio in the 500M run). Multi-scale suite covering 200M, 300M, 400M and 500M, plus a 1.6B suite with three shapes; evaluation uses an NLL-based protocol rather than training loss alone.
Layer-level analyses link the gains to layers actually being used: HC and AttnRes show larger angular distances with markedly non-smooth structure across depth, and higher fractions of late-layer causal and permutation scores above threshold (roughly for Full AttnRes and for HC, versus about for Pre-LN); meanwhile Block AttnRes has an early dead segment where the depth-mixing softmax gives essentially zero weight to the running sum, and mHC's composed residual maps have lower effective rank (– versus – for HC), suggesting stabilization constraints reduce long-range route diversity. Beyond reporting performance differences, it offers a mechanism-level contrast: Full AttnRes directly retrieves earlier sub-layer outputs while HC propagates and recombines them through interacting residual streams, whereas Pre-LN-like architectures become smooth and partly interchangeable in late layers. Three controlled layer-level analyses (angular distance, causal score, permutation score) focused on the last three quarters of the network to avoid the universally strong early layers; the mHC and Block AttnRes explanations are supported by routing-weight and residual-map spectral analyses.
Perspective
The results are aimed at researchers and engineering teams choosing model shapes under a fixed parameter budget: when the goal is for more layers to contribute effective computation, multi-stream or cross-layer-access residual designs such as HC and Full AttnRes are the natural candidates, while Pre-LN and its normalization variants suit wider, shallower shapes. The setting is controlled pre-training comparison, with the main benchmark at 400M total parameters, 20 tokens per parameter and sequence length 2048, and trend validation at 200M–500M and 1.6B; the authors also release code and checkpoints spanning a wide range of aspect ratios for further exploration of the accuracy–efficiency trade-off.
Open questions remain: whether these trends hold far beyond 1.6B, under different data recipes and tokenizers; HC is more sensitive to optimization hyperparameters, with gradient explosion at the largest aspect ratio in the 500M run and no clear depth gain at 1.6B, so its stability boundary needs finer learning-rate and initialization tuning; the overhead of deep-narrow models in prefill FLOPs (roughly 3.0 to 4.4 TFLOPs per sequence), KV cache (more than doubling), peak memory and GPU-hours means practical gains depend on kernels and parallel infrastructure; and although this is a full-text parse, some figure values appear as ranges or omissions, so exact thresholds and complete heatmaps require the appendix.
