MT-OPSD uses on-policy self-distillation on self-generated states to lift open-source image editors from 0.03–0.15 to 0.38–0.52 ten-turn success
Synopsis
The work finds that instruction-based image editing models degrade rapidly under recursive multi-turn editing and attributes this to a train–test mismatch in the conditioning distribution, since models are trained on clean source images but must repeatedly condition on their own imperfect outputs at inference; it proposes MT-OPSD, which rolls out the current student under an identity instruction to produce self-generated states containing its own errors, then uses the same model conditioned on the clean source as a teacher to transfer clean-condition editing behavior onto those states via sparse velocity matching, without multi-turn annotations; it also builds LME-Bench of 100 ten-turn editing sessions, raising SR@10 to 0.38–0.52 and cutting CR@10 to at most 0.
Interpretation
The paper identifies multi-turn collapse as a phenomenon across editing models and attributes it to a train–test mismatch in the conditioning distribution: training conditions only on clean source images, whereas inference repeatedly conditions on the model's own previous outputs, so errors accumulate turn by turn. Prior work often treated multi-turn degradation as a model-specific issue or mitigated it with image-space or latent-space corrections after the fact; this work offers a unified diagnostic view and uses identity rollouts, where the model is asked to reproduce its input unchanged, to separate model-induced errors from the semantic changes requested by instructions. The paper reports observing this behavior across several editing models and provides the identity-reproduction diagnostic; on LME-Bench the three open-source backbones drop to SR@10 of 0.03–0.15 and rise to CR@10 of 0.25–0.61.
It proposes MT-OPSD, an on-policy self-distillation framework in which the student is conditioned on self-generated rollout states and the teacher is the same pretrained model conditioned on the clean source, transferring clean-condition editing behavior onto degraded states through sparse query-based velocity matching, with no multi-turn annotations or external teacher. Unlike reinforcement-learning routes that rely on external rewards (such as MT-EditFlow) and visual self-distillation that needs paired target images, this approach treats the clean source image itself as the teacher's privileged context while the student context is its own rollout state; training alternates an identity branch that suppresses further drift and an editing branch that preserves editability on self-generated states. Ablations show that removing the editing branch drops SR@10 to 0.00 and raises CR@10 to 0.80; replacing on-policy velocity matching with standard flow matching on teacher-generated pseudo-targets lowers SR@5 from 0.91 to 0.59, lowers SR@10 from 0.44 to 0.03, and raises CR@10 from 0.02 to 0.67.
It introduces LME-Bench: 100 ten-turn editing sessions with 1,000 instructions in total, each session containing six local and four global edits, evaluated turn by turn for prompt following, consistency with the previous state, and visual quality, with SR@t and CR@t as metrics. Existing benchmarks contain at most five consecutive turns and cannot characterize long-horizon degradation; this benchmark places global edits non-adjacently and not in the final turns, so later local edits operate on images that have already undergone global transformations, and it uses six families of validity rules to avoid invalid or ambiguous instructions. Source images are generated by Z-Image-Turbo and evenly distributed across 10 semantic categories; instructions are first checked by a VLM and then manually reviewed; each turn is scored 0–10 by GPT-4o, a turn succeeds only when both prompt following and consistency reach at least 7, and a session counts as collapsed only after two consecutive degraded turns.
Across three editing backbones, MT-OPSD substantially improves long-horizon editing success and reduces collapse while largely preserving single-turn editing quality. The rollout curriculum typically saturates at around four turns, yet the gains persist through ten-turn evaluation, indicating that robustness learned from self-generated states transfers beyond the rollout depths seen in training; training-free methods show backbone-dependent gains. On LME-Bench, SR@10 rises to 0.38–0.52 and CR@10 stays at most 0.04; on MSE-Bench, SR@5 improves from 0.24 to 0.51 on Qwen-Image-Edit-2511 and from 0.40 to 0.69 on FireRed-Image-Edit, while matching FLUX.2-klein-base at 0.49; ImgEdit single-turn scores change only slightly.
Perspective
The result targets instruction-based image editing models built on flow matching that can be adapted with LoRA, in a setting where each turn receives only the previous turn's output and the original image is never reintroduced; the method depends on the model already having strong single-turn clean-condition editing ability, trains on roughly 2,000 source-image/instruction pairs from OmniEdit without using ground-truth edited images, and selects teachers asynchronously via a VLM judge on a held-out gate set of 23 ten-turn sessions. LME-Bench and the gate set use source images generated by Z-Image-Turbo and are scored by GPT-4o, so the benchmark and metrics can be used directly to compare other multi-turn editing methods and the same idea can be extended to video or other recursive generation tasks.
The rollout curriculum typically saturates at around four turns while gains hold at ten-turn evaluation, and the boundary of that extrapolation deserves continued watching; the identity-versus-editing branch trade-off appears in ablations as a balance between success rate and controllability, and different applications may prefer different points; teacher promotion relies on a VLM judge and a held-out gate set, and the influence of judge behavior is not developed in the main text; compared with proprietary systems, GPT-Image-2 and Nano Banana Pro remain stronger at long horizons, and how far open-source models close the gap varies by backbone; in addition, this reading covers the main text and appendices, and per-cell numbers in the tables were not verified cell by cell, so exact figures should be taken from the original tables.
