Skip to main content
Back to timeline
arXivSource publication:

SoL-Refiner refines low-resolution generated video to 4K in one denoising step, cutting 2K refinement latency from 57.461 to 6.447 seconds

Synopsis

The work presents SoL-Refiner, a one-step video refiner that turns low-resolution generator outputs into 4K video through a three-stage recipe of high-resolution continual training, reinforcement-learning post-training with frame-based reward models, and final one-step distillation, and introduces Refiner-Bench, a video refinement benchmark using a shared-input protocol to compare refiners at roughly 2K output resolution; at 2K the one-step model outperforms all evaluated external refiners on VBench and UniPercept averages, at 3840×2176 it improves both metrics over the three-step LTX-2.3 Refiner, and with the complete acceleration stack it achieves an 8.91× speedup in refinement latency over the same baseline in the 2K latency setting.

AI-generated editorial illustration: SoL-Refiner: Speed-of-Light One-Step Refinement for High-Resolution Video

Interpretation

SoL-Refiner is a one-step video refiner that transforms low-resolution model outputs into 4K videos with a single denoising step, and it can refine outputs from different base generators without modifying or retraining the base models. Existing refiners often require multiple target-resolution denoising steps, creating a second sampling bottleneck, and many are developed for a particular base generator with transfer to other generators rarely evaluated; this work compresses refinement to one step and explicitly assesses cross-generator transfer. The paper reports that on Refiner-Bench the one-step SoL-Refiner outperforms all evaluated external refiners on VBench AQ, IQ, and AVG and on UniPercept AVG, including SEEDVR2 under the same step budget, with VBench and UniPercept averages of 0.81048 and 60.4150 respectively.

A three-stage training recipe links high-resolution refinement capability, perceptual quality, and one-step inference: continual training learns the low-to-high-fidelity mapping, RL post-training improves perceptual quality with frame-based reward models, and DMD-GAN distillation compresses the model to one step. Rather than supervising the student directly with perceptual rewards during one-step distillation, this work applies frame-based rewards to the multi-step refiner before distilling it through a few-step intermediate student; ablations show RL post-training gives the highest VBench and UniPercept averages, and after distillation the one-step model remains above the continual-training baseline. The training-recipe ablation reports three-stage numbers: continual (multi-step) VBench AVG 0.80888 and UniPercept AVG 56.6620; post-training (multi-step) 0.81691 and 61.3170; DMD-GAN (one-step) 0.81048 and 60.4150. The reward-learning ablation shows that removing either the regularization recipe or the multiple reward models lowers the VBench average.

Refiner-Bench is a video refinement benchmark constructed from the outputs of different video generators, using a shared-input protocol to compare refiners at approximately 2K output resolution. The benchmark separates refinement performance from the choice of upstream generator, containing 150 videos — 50 each from WAN 2.1 1.3B, SANA-Video 2B, and LTX-2 Stage 1 — spanning 12 content groups, three motion levels, and six camera-motion types, with WAN and SANA-Video providing out-of-distribution inputs and LTX-2 Stage 1 providing in-distribution inputs. The paper states that all methods receive the same 150 aligned inputs resized to a common size, that LingBot Stage-2 Refiner produces one output resolution while the other refiners produce another, and that UniPercept uses eight uniformly sampled frames per video.

The complete acceleration stack (one-step distillation, TAE, and Sol-Engine) reduces refinement latency from 57.461 to 6.447 seconds in the 2K latency setting, an 8.91× speedup over the three-step LTX-2.3 Refiner; cross-base-generator experiments likewise show that low-resolution generation plus one-step refinement lowers end-to-end latency while improving quality averages. The result stacks one-step distillation with system-level acceleration (TAE replacing the full VAE, and Sol-Engine's kernel fusion and sparse attention), indicating that the two provide complementary gains rather than step-count compression alone. The latency analysis is measured on one H100 GPU, adding one-step distillation, TAE, and Sol-Engine in sequence to obtain adjacent speedups and the 57.461-to-6.447-second end-to-end figures; the WAN-5B, WAN-1.3B, and Cosmos-Nano pipelines reduce latency by 54.7%, 71.1%, and 64.4% respectively while improving mean VBench and UniPercept.

Perspective

The result targets generative video systems architected as a two-stage pipeline of low-resolution base generation plus high-resolution refinement, and suits developers who want to improve local detail and texture without modifying or retraining the base model; the paper notes that refinement can improve a generated video at its existing resolution or accompany upsampling, so it serves both 2K/4K enlargement and same-resolution artifact correction. Refiner-Bench's shared-input protocol lets different refiners be compared on the same 150 aligned inputs, helping separate refinement performance from the choice of upstream generator. The cross-base-generator experiments and the SoL-H3 deployment on DGX Spark indicate the refiner can serve as the refinement stage in a cascade.

The limitations section notes that on Refiner-Bench the 23-step model achieves higher VBench and UniPercept averages than the one-step model, so one-step conversion involves a quality-versus-speed trade-off; the refinement objective targets local visual detail while preserving the base video's content and motion, and correcting large semantic, geometric, or motion errors is not an explicit training objective; Stage II evaluates three sampled frames per video, and frame-based rewards do not directly measure long-range temporal consistency; Refiner-Bench covers AI-generated videos from several base generators, while camera-captured degradations and broader video editing tasks remain outside the study. In addition, some values in the loaded text appear in placeholder form (for example the resolution notation at 3840×2176 and the way several speedup and latency numbers are rendered in the body versus the appendix), so exact figures should still be checked against the original tables and figure captions.

Sources