Skip to main content
Back to timeline
arXivSource publication:

Tencent team proposes Adaptive Reward Routing, taking best result on nine of ten metrics across LTX-2 and LTX-2.3

Synopsis

The work proposes Adaptive Reward Routing, which jointly adapts where reward-driven updates act and how competing rewards are coordinated during forward-process RL (DiffusionNFT) of joint audio-video diffusion models: it uses bidirectional cross-attention responses as a proxy for evolving cross-modal influence to dynamically reweight token-aware losses and scale gradients across cross-modal layers, and after warm-up uses branch-specific reward-gradient interactions as residual corrections to predefined preference weights, training on 19,487 audio-video prompts over the LTX-2 (19B) and LTX-2.3 (22B) backbones and achieving the best result on nine of ten metrics on the 10,140 prompts of JavisBench.

AI-generated editorial illustration: Adaptive Reward Routing: Dynamic Multi-Reward Optimization for Joint Audio-Video Diffusion via Forward-Process RL

Interpretation

The paper reframes reward-guided post-training of joint audio-video diffusion models as dynamic multi-reward optimization along two coupled dimensions: where reward-driven updates should act, and how multiple rewards should be coordinated. Prior methods such as OmniNFT fix layer routing based on the base model, GDPO aggregates rewards with fixed weights, and MARBLE adjusts weights by gradient geometry without encoding preference or modality responsibility; this work unifies both dimensions into an adaptive process that evolves with training. The paper probes base and OmniNFT-trained checkpoints to show that cross-modal functions and gradient flows change during fine-tuning, making static routing progressively stale; this is motivation-level evidence rather than final performance evidence.

Cross-Modal Influence-Guided Routing uses bidirectional cross-attention responses as a proxy for cross-modal influence, with layer aggregation yielding token weights and token aggregation yielding layer scales, localizing updates without additional model interventions. Unlike fixed routing, token and layer routes are recomputed from the current model during training; layer scales convert to a soft detachment coefficient that scales backward gradients while leaving forward values unchanged, so strongly influential layers retain more gradient and weakly coupled layers are increasingly detached. Disabling each of the 48 A2V or V2A blocks separately at four training checkpoints shows the proxy closely recovers the intervention-based layer ranking with Spearman correlations near one for A2V and V2A; the proxy recomputed from the current model stays high throughout fine-tuning while the proxy frozen at initialization falls to low values; blocking the top-scoring 10% of target tokens changes the final prediction more than blocking an equally sized random group.

Preference-Preserving Modality-Aware Reweighting estimates reward conflicts within the branch responsible for each objective and, after warm-up, uses the resulting coefficients as residual corrections to predefined reward weights. GDPO preserves explicit preferences but cannot adapt to conflicts, and MARBLE adapts to local gradient geometry but does not encode preference or modality responsibility; this method lets the prior set a nonzero floor while the residual term adapts to current conflicts, avoiding dominant rewards suppressing weak but essential objectives. Ablations show branch-aware balancing, residual mixing, and warm-up each contribute improvements, and the routing chain keeps GDPO weighting fixed while the weighting chain inherits no routing, isolating each axis; the complete model gives the strongest overall balance across quality, semantic consistency, and synchronization.

Across two joint audio-video diffusion backbones, the method achieves the best result on nine of ten JavisBench metrics and maintains favorable trajectories across all five component rewards during training. Compared with GDPO, MARBLE, OmniNFT, and the authors' released OmniNFT* checkpoint, it improves modality quality, semantic consistency, cross-modal consistency, and synchronization consistently, with the same trend on both backbones. On LTX-2: VQ 3.336, AQ 5.868, AV-IB 0.235, AVHScore 0.234, JavisScore 0.206, DeSync 0.341; on LTX-2.3: VQ 3.599, AQ 5.979, AV-IB 0.267, AVHScore 0.266, JavisScore 0.236, DeSync 0.302; GDPO, MARBLE, OmniNFT, and the proposed method are each trained independently with three random seeds under the same data, LoRA, and optimization budgets, with arithmetic means reported, while the base model and OmniNFT* are fixed checkpoints.

Perspective

The result targets multi-reward post-training of joint audio-video diffusion models and applies when modality tokens are identifiable and directional interaction responses can be isolated; the paper validates on the dual-stream LTX-2 (19B) and LTX-2.3 (22B) backbones and states that routing can extend to the unified single-stream JavisDiT++ backbone, while reward coordination is architecture-independent. Training uses 19,487 audio-video prompts from a VGGSound-derived corpus and optimizes five reward signals, with evaluation on 10,140 JavisBench prompts under a four-second, 24 FPS, 16 kHz protocol. For researchers using other diffusion objectives, the paper notes that routing can extend through their token losses and gradient paths, and reward coordination requires only reward-wise gradients.

The paper states that no established unified reward jointly captures modality quality, semantic consistency, and temporal synchronization, and that human-preference models provide a complementary overall signal but do not replace fine-grained modality-specific supervision; for models with inseparable modality representations or inaccessible interaction responses, routing remains future work. In ablations, individual components may favor different objectives, so intermediate configurations do not necessarily improve every metric monotonically, and the overall-balance conclusion depends on the complete configuration. In addition, this evidence bundle is full text without the figures themselves, so the qualitative examples in Fig. 3, the training-dynamics curves in Fig. 4, and the proxy-validation curves in Fig. 5 can only be understood from the prose descriptions.

Sources