37k-Parameter Learned Optimizer Trained in 0.87 GPU-Hours Transfers Zero-Shot, Cutting BERT-Tiny Validation Loss by 9.1%
Synopsis
This paper presents a lightweight and versatile learned optimizer that dynamically recombines gradient history represented as averages over disjoint time spans: the prediction space is reduced to one scalar coefficient per gradient average shared by multiple parameters, and progressively averaging older gradients minimizes the memory cost of long history while keeping their contributions independently accessible; a 37k-parameter network trained in 0.87 GPU-hours generalizes zero-shot to unseen tasks, lowering validation loss by 9.1% and 0.4% on BERT-Tiny and GPT-Tiny, improving test accuracy over Adam by 3.5 percentage points on a Vision Transformer and by 2.7 percentage points on average across nine graph models, with FLOPs overhead as low as 0.3%.
Figure 2: Dynamic recombination of multi-resolution gradient history. For each parameter group, the current gradient enters a multi-resolution history. A group feature encoder converts each stored average into a fixed-dimensional slot feature, and a joint temporal predictor assigns the coefficients used to recombine the history into the group update. Step index t t is omitted in the figure as it represents a single optimization step.
arXivInterpretation
Introduces a learned optimizer whose core mechanism dynamically recombines gradient history, with history represented as averages over disjoint time spans. Relative to prior learned optimizers, it reduces the prediction space to one scalar coefficient per gradient average, shared by multiple parameters, lowering the complexity of learning and prediction while retaining historical information. The summary provides the mechanism plus quantitative results: a 37k-parameter network trained in 0.87 GPU-hours, with zero-shot transfer reported across several model families.
Progressively averaging older gradients minimizes the memory cost of long history while keeping their contributions independently accessible. This design avoids storing long-range gradient history in full while still allowing the optimizer to call on each span separately, balancing memory efficiency with flexible use of history. The summary states the mechanism; it does not give specific memory figures, only that memory cost of long history is minimized.
The trained optimizer generalizes zero-shot to unseen tasks and outperforms Adam across multiple architectures. It lowers validation loss by 9.1% and 0.4% on BERT-Tiny and GPT-Tiny, improves test accuracy over Adam by 3.5 percentage points on a Vision Transformer and by 2.7 percentage points on average across nine graph models. Results span language, vision, and graph models, and report FLOPs overhead as low as 0.3%, indicating gains come with low computational cost.
Perspective
The result is aimed at researchers and practitioners who want a general-purpose optimizer at low training and compute cost, and applies to the language models (BERT-Tiny, GPT-Tiny), Vision Transformer, and graph-model training settings described in the summary; its value lies in zero-shot transfer to unseen tasks and in validation-loss and test-accuracy gains over Adam at FLOPs overhead as low as 0.3%.
The summary does not specify how the disjoint time spans are partitioned, what objective learns the scalar coefficients, or the training task distribution, nor does it give concrete memory figures or statistical significance; these open questions need confirmation in the original. Also, based only on the summary, performance at larger model scales and longer training cannot be judged.
