Skip to main content
Back to timeline
arXivSource publication:

Adobe Research and KAIST introduce FlowTool: treating retouching tool parameters as conditional generation, cutting L1 by up to 28.8% and L2 by up to 47.4% on reference metrics with at least 50x lower latency

Synopsis

Researchers at Adobe Research and KAIST introduce FlowTool, which uses conditional rectified flow to directly model the distribution of high-quality retouching tool parameters conditioned on the input image and user instruction, combining a vision-language model backbone with a Diffusion Transformer parameter generator and a tool-presence head instead of autoregressive MLLM reasoning and token-by-token numeric generation; across MMArt-Bench, FlowTool-Eval, ArtEdit-Bench, and MIT-Adobe5K it achieves stronger reference-based performance than specialized MLLM editing agents and proprietary MLLMs, remains competitive with proprietary models under reference-free evaluation, and reduces inference latency by at least 50x while requiring nearly 2x less memory.

AI-generated editorial illustration: FlowTool: Controlling Tool Parameter in Image Retouching via Flow Matching

Interpretation

It recasts tool-based image retouching from autoregressive language generation into conditional generative modeling over a structured continuous tool-parameter space, directly modeling p(a | x, c). Prior work such as JarvisArt and RetouchIQ has MLLMs generate reasoning, tool selections, and discretized parameter values token by token; this work instead performs conditional flow matching in continuous parameter space, preserving the continuous structure of parameters and modeling the conditional joint distribution over tool parameters. The paper provides a formal problem setup: parameters are normalized to [-1,1], a tool-presence mask m distinguishes active tools, and training is defined by a rectified-flow linear path and velocity-matching loss; the introduction offers three motivations (continuous numerical representation, cascading errors, inference cost), which are argumentative rather than ablation evidence.

FlowTool combines a VLM backbone, a DiT-based tool parameter generator, and a tool-presence head, trained with a two-stage supervised flow-matching curriculum followed by reward-based post-training. Unlike purely imitating expert parameters in supervised training, the second phase applies DiffusionNFT-style forward-process reward post-training, with a reward that is a weighted sum of a reference-based negative L1 reward against the expert edit and a VLM-as-a-judge reference-free reward, so the model need not exactly reproduce a single expert solution. The method section gives full formulations (flow-matching loss, presence BCE loss, reward normalization, and positive/negative branch objectives); training data comes from internally curated expert editing records with instructions synthesized by Gemma4-31B, totaling 496,693 training examples; ablations show the two-stage curriculum beats stage 1 only (L1 9.63, L2 18.92, SC 7.73, O 8.50) and stage 2 only (L1 9.90, L2 20.10, SC 5.99, O 7.42), with the full curriculum at L1 9.37, L2 17.78, SC 7.96, O 8.69.

Across four benchmarks FlowTool substantially outperforms specialized MLLM agents and proprietary MLLMs on reference-based metrics, and is competitive with proprietary models on reference-free metrics. Against the strongest specialized baseline it reduces L1 by up to 28.8% and L2 by up to 47.4%; against frontier proprietary models it reduces L1 and L2 by up to 13.0% and 21.7% respectively, while achieving stronger perceptual quality and remaining competitive in semantic consistency; on MIT-Adobe5K it outperforms all baselines. Results come from Table 1 comparing on MMArt-Bench, FlowTool-Eval, ArtEdit-Bench, and MIT-Adobe5K, reporting L1, L2, SC, PQ, O plus PSNR, SSIM, and LPIPS on MIT-Adobe5K; FlowTool-Eval contains 300 held-out internal examples with no overlap with training data; reference-free metrics use GPT-5 as the judge.

Directly generating editing parameters yields substantial efficiency gains: at least 50x lower latency and nearly 2x lower peak GPU memory. Autoregressive agents may generate lengthy reasoning and tool-call trajectories that are not the system's final objective; FlowTool requires no natural-language reasoning or numeric generation at inference and uses three ODE steps by default. The efficiency comparison samples 200 examples total (50 from each benchmark) for single-sample inference on an NVIDIA A100; FlowTool takes 0.19s versus 9.66s to 32.11s for baselines; locally deployed baselines use 9.89 GB and 17.09 GB peak GPU memory, while proprietary MLLMs are accessed via API so their memory is unavailable; ablations show one step is insufficient, two steps improve substantially, three steps give the best balance, and more steps add little.

Perspective

The work targets standardized tool sets controlled by continuous parameters normalized to a common range; the authors state the current design is for "standardized linear parameter spaces, such as parameters normalized to [−100, 100]". It applies to retouching scenarios that must translate natural-language editing intent into concrete tools and parameter values, especially those sensitive to interactive latency or deployed on resource-constrained devices; evaluation covers MMArt-Bench, ArtEdit-Bench, MIT-Adobe5K, and the internal held-out FlowTool-Eval, with training data from internal expert editing records and instructions synthesized by Gemma4-31B. The authors identify extending to cyclic or polar-valued parameters such as hue, and to discrete or categorical parameters such as textual options, as future directions, which defines the current reach of the method across editing tools.

Evaluation relies on internal expert editing data and the internal held-out set FlowTool-Eval, whose composition external readers cannot directly verify; reference-free metrics use GPT-5 as the judge, whose scoring behavior is itself a variable to watch. In the efficiency comparison proprietary MLLMs are accessed via API with memory unavailable, so the memory conclusion mainly concerns locally deployed specialized agents. Ablations show that scaling the backbone from 4B to 9B brings no additional improvement, a phenomenon the paper does not elaborate on. In addition, although this evidence bundle is the full text, tables and figures are rendered as text, so exact full comparisons for individual numbers still require the original tables.

Sources