模型生长、递归与边界算子如何影响缩放指数
核心概要
该工作以 prelude–core–coda 循环 Transformer 为统一框架,在 FineWeb 上为每种架构单独拟合计算最优配方并训练到约 10^20 FLOPs 的缩放阶梯,发现模型生长(训练中途增加核心遍数)与边界算子(归一化残差流并重新注入 prelude 输出)能改变预训练的缩放指数,使计算效率优势随规模扩大,而权重解绑只改变常数;在数据受限的多轮训练中,最优循环次数随计算量增长,扩大循环次数比扩大模型规模更省算力。
Figure 5: In multi-epoch training, the optimal loop count grows with compute, scaling the loop count beats scaling model size even against tuned weight decay, and looping leaves the optimal weight decay nearly unchanged. Models are trained on 100M unique FineWeb tokens for 10 epochs, and all losses shown are validation losses on FineWeb. First: matched-compute cuts. The loss of each loop count is interpolated in compute between neighbouring depths, curves are quadratics in log loss against log loop count, and stars mark their minima. The optimal loop count rises from 1.4 to 6.7 across budgets, the gap to Operator-1 grows from 0.00 to 0.11. Second: marginal gain from looping. For each depth, we plot the loss at loop count K K minus the Operator-1 loss at the same depth. Every curve decreases monotonically, and from d 8 d8 upward the curves lie within 0.006 of one another at every loop count, so the gain from looping appears nearly independent of depth. Third: weight decay when scaling loops versus model size. Scaling model size at fixed weight decay overfits, and retuning weight decay at every model size only reaches 3.40 loss at 7.5 × 10 18 7.5\times 10^{18} FLOPs. Scaling the loop count at fixed model size and fixed weight decay matches that loss with 2.2 × 2.2\times less compute, and tuning weight decay on top gains at most 0.02, so loop-count scaling is more compute-efficient and nearly free of tuning. Fourth: optimal weight decay by depth and loop count. The optimum rises with depth but is nearly flat in loop count, so overfitting tracks stored parameters rather than executed depth.
arXiv深度剖析
模型生长(训练中途把核心遍数从 2 增加到 4)改善缩放指数,而非仅改善常数。 以往的生长与循环研究多报告单一或少数目标规模上的收益,未在计算最优、架构专属超参缩放的设定下考察缩放指数;本文显示 Untied-Grow 相对 Vanilla 的计算乘数从 10^18 FLOPs 的 1.30× 扩大到 10^20 FLOPs 的 1.55×,相对 Untied-2 也从 1.09× 增至 1.16×。 八条计算最优阶梯、每架构独立拟合配方,log–log 斜率回归标准误低于 10^-3;生长时机扫描在固定规模与预算下呈内部极小值,且生长模型偏好更小的初始模型与更多 token(最优 tokens per parameter 从 6 升到 7–8)。
仅加入边界算子(Norm(h)+αe,并在 coda 前同样施加)即可改善缩放指数。 此前深度相关干预只在固定模型规模下改善损失,未被证明能改变缩放指数;本文中 Operator-1 与 Vanilla 参数量和 FLOPs 相同,计算乘数从 10^18 FLOPs 的 1.12× 升至 10^20 FLOPs 的 1.25×,且消融显示去掉归一化、注入或 coda 注入都会削弱增益。 参数与 FLOPs 匹配的对照比较,加上边界算子各组成部分的消融(Table 4、Figure 15(a));KL 有效深度上 Untied-2 达 24 而 Deep Vanilla 为 20(10^20 FLOPs),且仅缩放宽度时指数优势退化为常数优势。
权重解绑只带来固定倍数的常数改善,不改变缩放指数。 把深度效应与权重共享效应分离:Untied-2 与 Loop-2 计算图与 FLOPs 相同,Untied-2 在 10^18 与 10^20 FLOPs 分别高出 1.08× 与 1.06×,该倍数不随规模扩大;因此绑定权重并生长仍保留生长的指数改善,只损失一个常数因子。 FLOPs 与深度匹配的配对比较,并在 Figure 3 右图以拟合指数 γ 对常数 log A 的位移方向加以区分。
在数据受限的多轮训练中,最优循环次数随计算量增长,扩大循环次数比扩大模型规模更省算力。 此前工作表明重复数据下最优权重衰减随参数量增长、需逐规模重调;本文在 100M 唯一 FineWeb token 上训练 10 轮,发现最优循环次数从约 1 升到约 7,固定模型规模与固定权重衰减下扩大循环次数以 2.2× 更少算力达到逐规模调权重衰减的最佳损失,且在其上再调权重衰减最多只再降 0.02。 Operator-1 到 Loop-12、120M–1.4B 参数的网格与权重衰减扫描;最优权重衰减随深度上升但对循环次数近乎平坦,解绑循环未胜过调好的 Operator-1。
启示与展望
该结果面向单轮、数据不受限的计算最优预训练,以及固定 100M 唯一 token 上多轮重复的数据受限设定;作者建议在小规模校准块分配、循环次数与生长比例后按比例放大复用,并强调必须为每种架构单独调参与缩放学习率,否则指数改善会被掩盖。它也为后续探索提供了方向:上下文长度、专家数量与宽度同样随算力增长却在训练前固定,按网络所需的时间表生长它们可能进一步改变指数,深度上分阶段(2→4→6 遍)亦被列为可能路径。
读者仍会关注:指数改善在更大算力与更多语料上的持续性,作者以 8× 拟合算力的单次留出运行与到 10^25 FLOPs 的投影给出初步支持;与 GPT-3 13B 的约 20× 算力差距因训练数据不同、参考分数来自另一套评测流程而被作者称为指示性而非受控比较;KL 有效深度只是计算深度的代理,落在阈值内并不代表后续块不再改变预测,网络中部未使用的块也无法被检测;随机循环主要降低额外测试时遍数的惩罚,并未带来持续的测试时扩展;Muon 与 Adam 在宽度单独缩放与宽度深度联合缩放下的表现差异,被作者列为需要进一步研究的方向。
