Skip to main content
Back to timeline
发表出处待核验Source publication:

SAGE optimizer lifts accuracy, cold-start recall, and diversity together on Amazon and RecIF-Bench generative recommendation

Synopsis

The work identifies a "Symmetric Conservatism" failure mode of GBPO in generative recommendation (symmetric update bounds suppress rare positive signals such as cold-start items, static negative-sample constraints fail to prevent diversity collapse, and group-normalized multi-objective rewards yield low-resolution training signals) and proposes SAGE, which uses a geometric-mean sequence-level importance ratio and a decoupled multi-objective advantage estimator to reduce token-level variance and mitigate reward collapse, plus asymmetric adaptive bounding that applies a positive Boost to successful slates and an entropy-aware penalty to low-diversity failures; on three Amazon Product Reviews categories and the large-scale RecIF-Bench, the native-text TextRec optimized with SAGE outperforms O

AI-generated editorial illustration: SAGE: Sequence-level Adaptive Gradient Evolution for Generative Recommendation

Interpretation

The paper identifies and names the "Symmetric Conservatism" failure mode of GBPO: symmetric update bounds suppress rare positive signals such as cold-start items, static negative-sample constraints fail to prevent diversity collapse under rejection-dominated feedback, and group-normalized multi-objective rewards produce low-resolution training signals. Where GBPO was previously framed as a stability fix for negative-sample gradient explosion, this work reframes its symmetric conservative design as a structural bottleneck for recommendation, reporting observations such as a 44.7% drop in cold-start video views and an 11.7% increase in cluster density. Based on analysis of the OneRec-V2/GBPO mechanism and the observation values reported in the text, this is a diagnostic argument rather than an independent controlled experiment.

SAGE replaces token-level importance ratios with a geometric-mean sequence-level ratio and pairs it with a decoupled multi-objective advantage estimator (normalize each objective independently within the sampling group, then aggregate, then apply batch-level z-score normalization) to reduce token-level variance and mitigate multi-objective reward collapse. Relative to GRPO's group normalization and GBPO's conservative bounds, SAGE elevates the importance ratio to the slate level (fixed slate length L=10) and draws on GDPO's decoupled normalization so that objectives such as clicks, duration, and comments retain resolution before aggregation. The method is presented with equations and an algorithm; ablations show that replacing signal decoupling with a simple weighted reward sum reduces NDCG@10 and other metrics.

SAGE introduces asymmetric adaptive bounding: for positively rewarded slates a Boost factor (ε_boost in 0.2–0.5, set to 0.3 in experiments) allows super-linear updates, while for negatively rewarded slates a list-entropy-based regulation factor λ imposes stronger penalties on low-diversity failures. Relative to GBPO's symmetric design that caps positive updates at 1.0, SAGE draws on DAPO's Clip-Higher and BAPO's entropy-clip ideas to make the boundary dynamic with respect to sample and entropy. Ablations show that removing Positive Boost drastically reduces cold-start recall and removing the entropy penalty substantially decreases Entropy@10, indicating each component acts in its targeted area.

On Amazon Product Reviews (Beauty, Sports, Toys) and RecIF-Bench, TextRec using the native text vocabulary and optimized with SAGE achieves the best top-K accuracy, improving NDCG by roughly 6.27%–7.8% over OneRec-GBPO while substantially improving cold-start recall and diversity. Relative to OneRec variants that depend on a separate Semantic-ID vocabulary, the results indicate that directly reusing the LLM's native vocabulary with preference optimization can outperform Semantic-ID architectures, with clearer advantages on instruction-conditioned tasks (Interactive Rec, Label-Conditional Rec). Covers three Amazon categories and six RecIF-Bench tasks, including roughly 120M interactions from 200K users under a strict user-based split; cold-start recall improves by roughly 89.9%–101% and entropy by roughly 11.17%–11.64%.

Perspective

The results target list-wise generative recommendation, applying to systems with a Qwen3-8B backbone, slate length L=10, and a two-stage SFT-then-RLHF pipeline, validated in both Semantic-ID and native-text action spaces. For teams seeking to relieve cold-start suppression and diversity collapse while keeping training numerically stable, SAGE offers an optimizer option that can replace GBPO; for multi-objective feedback (clicks, duration, comments) where signal resolution matters, the decoupled advantage estimator offers a reusable practice. The text also notes that native-vocabulary recommendation requires language-level item rendering, which constrains latency-sensitive systems, so this route fits environments where compute and deployment tooling continue to improve.

Readers will still watch: accuracy gains under the Semantic-ID action space are not stable (OneRec-SAGE is competitive but does not consistently lead), suggesting the interaction between strong exploration and a constrained discrete action space needs further characterization; how to trade off the accuracy–diversity balance that accompanies cold-start and diversity gains across different business objectives is treated as an empirical observation favoring the native-text setting and awaits validation in more scenarios; additionally, although this is a full-text parse, the algorithm pseudocode and some figure references in the body are incomplete (e.g., Algorithm ?? and Figure 5 presented only via normalized descriptions), so reproducing exact numbers and hyperparameter sensitivity would still require the original figures and code.

Sources