Google Deploys a Dual-Loop LLM Agent System on YouTube That Autonomously Found RMSprop, Gated Path, and a Multi-Objective Reward, Lifting Multiple A/B Metrics
Synopsis
The work proposes and deploys a self-evolving recommendation system in which Gemini-family LLM agents act as machine learning engineers: an Offline Agent (Fast Loop, waking every 5 minutes and running daily) generates and offline-scores candidate configurations, while an Online Agent (Slow Loop, running daily) ranks candidates and drives A/B experiments; across several YouTube recommendation surfaces the agents autonomously discovered optimizer switches (Adagrad to RMSprop, FTRL), a Gated Path (GLU) architecture, activation refinement, multi-objective reward synthesis, and reward hyperparameter tuning, with multiple online metrics improving at the 95% confidence level and experiment throughput rising from Θ(1)-Θ(10) per week to Θ(100) per week.
Figure 1: The Self-Evolving System Architecture. The framework operates as a dual-loop, self-evolving system centered around a shared Experiment Journal containing a persistent knowledge base, historical trials and their resulting metrics. The Offline Agent (Fast Loop) serves as the high-frequency nomination engine, where LLMs are invoked to instantiate specialized reasoning personas that generate and refine hypotheses into executable code. Tool calls are made to assign offline scores to candidates. The Online Agent (Slow Loop) is the low-frequency ranking engine, selecting high-potential candidates and promoting them to online experiment. It manages the entire experiment lifecycle including fetching online north star metrics.
· Page 4Interpretation
An autonomous MLE framework for industrial-scale recommender systems: centered on a shared Experiment Journal, the Offline Agent (Fast Loop) handles high-frequency candidate generation and offline scoring, while the Online Agent (Slow Loop) ranks by north star metrics, promotes candidates to training and live experiments, and reclaims resources; humans only supply the initial research idea and review final metrics. Prior AutoML (HPO, NAS, optimizer search) can only select or remix operations inside a predefined search space and cannot read production code, invent new modules, or author new reward logic; this work moves the "AI scientist" paradigm into a live production recommender, handling noisy feedback, safety guardrails, and rigorous A/B protocols. The paper presents the full system design and prompt templates (Figures 1-4) and reports deployment across several YouTube surfaces, with human engineers involved only at the first and last high-level steps.
The agents made structural discoveries in optimizer and architecture: replacing legacy Adagrad with RMSprop (learning_rate=0.005, rho=0.95, etc.), proposing a GLU-like Gated Path multiplicative-gate architecture, and later refining sigmoid gates into GELU activations with layer normalization; by adjusting batch sizes, epochs, and optimizer hyperparameters the agent also cut training time by 8x in total (first 4x, then 2x) without degrading business metrics. These changes were not picked from a fixed operator menu but written as new code; the authors describe Gated Path as yielding some of the most robust gains in the deployment and show the agent both exploring new structures and exploiting and fine-tuning structures it judged superior. Offline loss drops and live-traffic gains are reported as statistically significant; Table 1 gives RMSprop at +0.06% [+0.03%, +0.09%] YouTube-level and +0.12% [+0.05%, +0.19%] surface-level, and Gated Path at +0.06% [+0.02%, +0.11%] and +0.14% [+0.08%, +0.21%], all marked significant at the 95% confidence level.
On reward engineering, the Reward Persona first runs large-scale open-ended analysis over user logs with run_sql_query to form hypotheses, then uses compute_eval to compute loss-independent proxy metrics (such as long-watch correlation), synthesizing a multi-objective reward that incorporates a novel signal for whether the user is actively engaging with content on the site, significantly outperforming the human-engineered baseline; it then tuned four reward hyperparameters in two weeks using only the online slow loop and no offline metrics. The authors stress that changing the reward definition changes the optimization landscape, so comparing Lproxy across reward definitions is ill-defined and this persona cannot use compute_loss; this distinguishes it from prior work such as Eureka and LEARN-Opt evaluated in robotics or simulations with immediate feedback and a clear oracle. Multi-objective reward synthesis appears in Table 1 at +0.05% [+0.02%, +0.08%] YouTube-level and +0.17% [+0.13%, +0.22%] surface-level; reward hyperparameter tuning at +0.05% [+0.01%, +0.08%] and +0.21% [+0.13%, +0.29%], both significant; the authors note prior manual attempts over several months failed to improve both metric levels simultaneously.
Ablations show discovery depends on model reasoning power and context engineering: Gemini 2.5 Pro (opt_2p5, normalized loss -0.84 [-1.70, -0.01]) beats Gemini 2.5 Flash (opt_flash, 0.85 [0.66, 1.05]); removing the expert MLE persona (opt_no_role, -0.52), ordering history by timestamp instead of loss (opt_no_sort, 0.06), limiting to top-1 (opt_top_1, 0.11), or providing no history (opt_no_context, 1.05) all degrade performance, while the full loss-sorted history (opt_top_5, -0.72) approaches the baseline. This set of controls decomposes "can the agent discover" into measurable factors, indicating that expert persona framing and a complete, ranked history are key to iterative discovery, not model scale alone. Results are averaged over 6 independent runs exploring 70 ideas each on the optimizer-component task, reported as normalized z-scores of the loss with intervals.
Perspective
The results target engineering teams that run large-scale production recommender systems, can afford daily training and online A/B traffic, and have persistent experiment logs and safety guardrails; for such teams it frees humans from repetitive code generation, compilation, and experiment orchestration toward setting strategic guardrails, ethical constraints, and long-term vision. The method is described as model-agnostic, but the deployment environment is a value-based RL ranking model on YouTube's video watch page, with training typically requiring Θ(hours). What readers can directly reuse are the dual-loop division of labor, the shared Experiment Journal, delta-based proposals, and expert-persona prompting.
Readers should still note: the core online gains come from a single company platform, so cross-company and cross-domain generality is unverified; some Table 1 entries (e.g., the 4x training-efficiency YouTube-level -0.01% [-0.05%, +0.03%] and activation-refinement -0.02% [-0.05%, +0.01%]) have intervals crossing zero, so not every change is robust on every metric; ablations cover only the optimizer component, with no separate sensitivity analysis for the architecture and reward roles; the paper does not disclose the rejection or failure rate of candidates, nor long-term stability and safety-incident statistics for agent proposals. In addition, this is a full-text parse in which Figures 1-4 are architecture and prompt-template illustrations, so the complete prompt text and tool implementation details require the original paper and code. Worth watching next: whether this framework works equally well on non-recommendation industrial ML systems, and how to define and audit the agent's decision boundaries when humans retain only the first and last steps.
