Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data
Synopsis
The work proposes an "infinite-parameter LLM" architecture in which a compact hypernetwork encodes run-time data (facts, instructions, demonstrations) into a low-dimensional latent code, that code generates a low-rank additive modulation of a shared base feed-forward network so each token's "expert" is generated rather than drawn from a stored bank, and a Bayesian belief over the latent code is carried and updated online so the effective weight keeps evolving within a session; the authors also specify an evaluation protocol that pits carrying knowledge in the weights against carrying it in the prompt at matched budget.
Figure 1: Three regimes for turning data into model capability, ordered by how often the model’s weights change and how recent the data they learn from is. Pre-training and post-training (SFT, RLHF) both update the weights offline, in batch, and leave them frozen thereafter; they differ mainly in the data they use and how often they run. Live-learning, the regime this work targets, updates a generated low-rank code continuously, at inference, on the interaction data — facts, corrections, outcomes — that the others cannot reach in the loop, and keeps adapting rather than freezing. In-context learning and retrieval (bottom) also act at inference, but they leave the weights unchanged and carry the data in the prompt, where it is re-read every request and then discarded. The regimes are complementary, not competing: live-learning does not replace pretraining (§ 4), it reaches the data pretraining and prompting leave on the table.
arXiv · Page 3Interpretation
A generative expert architecture and its design space: a shared base FFN's weights are made dynamic through a generated low-rank delta driven by a latent code, with no stored expert bank. Relative to stored-bank MoE (e.g. DeepSeek-V3's 671B parameters with 37B activated per token) and to bank-free but frozen μMoE and ∞-MoE, the expert is generated rather than selected; the authors position ∞-MoE, μMoE, DFC, HyperMoE and MoEGen along two axes of where the weight comes from and whether it keeps changing after generation. An architectural and methodological design argument; the text gives the form W(z)=W0+B(z)A(z)ᵀ and contrasts the per-token cost O(r(d+h)) with the base FFN's O(hd), but reports no training or inference results.
A precise statement of the infinite-parameter view: the expert space is an unbounded continuous generated family, one expert per token, extended over time by adaptation, explicitly distinguished from unbounded knowledge. The authors stress that what a network can store is bounded by its parameters (citing roughly two bits per parameter), so "infinite" names the unbounded set of realisable effective weights and behaviours, not unbounded storage, with Bayesian-nonparametric mixtures of experts as a guiding analogy. A conceptual argument resting on cited capacity results and on the authors' own scoping of the claim, with no new empirical measurement.
Online adaptation formulated as a belief over the latent code, the element that most sharply separates this design from one-shot weight generators. One-shot generators such as Text-to-LoRA and SHINE read the context once and freeze the adapter, turn-level and memoryless; here the code is a latent variable with a prior, updated by an amortized recursive Bayesian filter at three cadences (contextual, per-turn, per-token) with uncertainty-gated retention, locating in-context learning, one-shot hypernetworks, continual-learning posteriors and fast weights as points within one probabilistic formulation. A formalisation and positioning argument at the method level; the authors state that "Bayesian" is an empirical claim about a calibrated posterior and that exact-filter recovery, the amortization gap and calibration of posterior precision are left to future work.
An evaluation protocol aimed at the comparison the design actually faces: whether carrying knowledge and behaviour in the weights beats carrying them in the prompt at matched budget. The headline baselines are the prompt family (in-context learning and retrieval), with one-shot weight generators and point-estimate test-time training as adaptation baselines and stored-bank MoE as a reference point; the documented failure mode of weight generators, memorisation rather than generalisation, is treated as a first-class evaluation concern. A protocol-design contribution; the text reports no results from executing the protocol.
Perspective
The design targets within-session adaptation at deployment: adaptation is framed as bounded, low-dimensional and reversible, and the authors state it does not repeal the capacity law or substitute for pretraining, but reaches the run-time data that pretraining and prompting leave on the table. It applies to the FFN sub-layer of chosen layers in a standard decoder-only transformer, with attention and base weights frozen throughout; the encoder spans a footprint-versus-reading-fidelity axis from a light readout head, through backbone reuse, to a fully separate hypernetwork, with the specific point left as an empirical question.
A careful reader would still watch: whether the amortized filter recovers exact-filter behaviour, how large the amortization gap is, and whether posterior precision is calibrated; whether uncertainty-gated retention and process-noise forgetting remain stable over long sessions; where the optimum lies between encoder size and reading fidelity; and whether the known memorisation-rather-than-generalisation tendency of weight generators is mitigated here. In addition, the available text is truncated within the method section, so Figure 4, the concrete architectural choices of Section 3.5 and the details of the Section 4 evaluation protocol could not be read, and descriptions of specific hyperparameters, base choices and experimental design are therefore limited to what appears in the text.
