Public articles linked to the same research event.
arXiv The work proposes COFLOW, an inference-time method that adaptively selects the step count for each generation based on prompt features and is trained online with an unsupervised reward balancing inference efficiency and generation fidelity; it is plug-and-play, requires no retraining of the underlying generative model, generalizes to image and video generation with over 2.5x speedup while preserving perceptual and semantic quality, and is accompanied by a theoretical analysis establishing an O(1/K) forward-Euler discretization error bound under standard regularity conditions.
The work proposes COFLOW, an inference-time method that adaptively selects the step count for each generation based on prompt features and is trained online with an unsupervised reward balancing inference efficiency and generation fidelity; it is plug-and-play, requires no retraining of the underlying generative model, generalizes to image and video generation with over 2.5x speedup while preserving perceptual and semantic quality, and is accompanied by a theoretical analysis establishing an O(1/K) forward-Euler discretization error bound under standard regularity conditions.
The work proposes COFLOW, an inference-time method that adaptively selects the step count for each generation based on prompt features and is trained online with an unsupervised reward balancing inference efficiency and generation fidelity; it is plug-and-play, requires no retraining of the underlying generative model, generalizes to image and video generation with over 2.5x speedup while preserving perceptual and semantic quality, and is accompanied by a theoretical analysis establishing an O(1/K) forward-Euler discretization error bound under standard regularity conditions.
The work proposes COFLOW, an inference-time method that adaptively selects the step count for each generation based on prompt features and is trained online with an unsupervised reward balancing inference efficiency and generation fidelity; it is plug-and-play, requires no retraining of the underlying generative model, generalizes to image and video generation with over 2.5x speedup while preserving perceptual and semantic quality, and is accompanied by a theoretical analysis establishing an O(1/K) forward-Euler discretization error bound under standard regularity conditions.