Skip to main content
Back to timeline
arXivSource publication:

COFLOW adaptively selects step counts from prompt features, achieving over 2.5x speedup in image and video generation while preserving perceptual and semantic quality

Related research and updates

Synopsis

The work proposes COFLOW, an inference-time method that adaptively selects the step count for each generation based on prompt features and is trained online with an unsupervised reward balancing inference efficiency and generation fidelity; it is plug-and-play, requires no retraining of the underlying generative model, generalizes to image and video generation with over 2.5x speedup while preserving perceptual and semantic quality, and is accompanied by a theoretical analysis establishing an O(1/K) forward-Euler discretization error bound under standard regularity conditions.

Source-provided article image: Contextual Flow Matching: Adaptive Step Selection in Flow Models for Efficient Visual Generation
Figure 1 ·

Figure 1 : Generation across integration step budgets. Different inputs can exhibit different behavior as the integration budget changes. The rightmost plots illustrate relative velocity variation along the corresponding trajectories.

arXiv

Interpretation

COFLOW adaptively selects the number of steps for each generation at inference time based on prompt features, rather than using a uniform number of function evaluations for all inputs. Existing acceleration methods reduce the number of function evaluations but often introduce additional training overhead, degrade quality, or fail to account for input-dependent variability; COFLOW makes step selection input-dependent and requires no retraining of the underlying generative model. Summary-level evidence: the method is described as a plug-and-play inference-time approach and is reported to achieve over 2.5x speedup on image and video generation while preserving perceptual and semantic quality.

COFLOW is trained online with an unsupervised reward that balances inference efficiency and generation fidelity. Relative to acceleration routes that rely on additional training overhead or quality trade-offs, this places the efficiency-fidelity trade-off inside an online-learned unsupervised reward. Summary-level evidence: the abstract states that the context-aware method is trained online with an unsupervised reward, but gives no reward form, training scale, or ablation details.

The authors provide a theoretical analysis establishing an O(1/K) forward-Euler discretization error bound under standard regularity conditions. This offers a convergence characterization of the error as a function of step count K for forward-Euler discretization, complementing an acceleration landscape dominated by empirical results. Summary-level evidence: the abstract states the bound holds, but does not list the specific regularity conditions or proof details.

Perspective

The result targets readers who use flow matching for visual generation and want to reduce function evaluations at inference time, in image and video generation settings; its selling point is that it can be plugged in without modifying the underlying generative model, making it suitable for serving and deployment of existing flow-matching models. The O(1/K) error bound stated in the abstract is given for forward-Euler discretization under standard regularity conditions, indicating that the method's scope is tied to this discretization framework.

The abstract does not specify the form of the unsupervised reward, the scale and stability of online training, or the datasets, comparison baselines, and evaluation metrics used, so how the over 2.5x speedup holds across resolutions, models, and prompt distributions needs confirmation in the main text. The regularity conditions, constant factors, and correspondence to the actual step-selection policy behind the O(1/K) error bound also require the main text to judge its applicable scope. This assessment is based on the abstract only, without reading figures or experimental details, and is limited to what the abstract states.

Sources