Skip to main content
Back to timeline
arXivSource publication:

ImIR replaces the text prompt with an embedding read off the degraded image, letting one frozen Qwen-Image-Edit plus a single low-rank adapter beat text conditioning on all six restoration tasks and still work without a degradation label.

Synopsis

The work recasts all-in-one image restoration as editing guided by an instruction derived from the image itself: a lightweight token mapper predicts the clean-image instruction embedding from the degraded image's vision-language embedding in place of a text prompt, and on a frozen Qwen-Image-Edit a single low-rank adapter trained on about 688 pairs for roughly three hours on one GPU covers low-light enhancement, deraining, dehazing, deblurring, denoising, and JPEG artifact removal, outperforming text conditioning on all six tasks under a matched comparison while also supporting task-agnostic restoration without a degradation label and a family of valid restorations obtained by scaling the instruction.

AI-generated editorial illustration: ImIR: Image-Instruction Tuning for All-in-One Image Restoration

Interpretation

The conditioning for all-in-one restoration is changed from a discrete text prompt to a continuous instruction vector derived from the degraded image, removing per-task prompt engineering. Prior work in this line (such as Edit2Restore) drives the editor with task-specific text prompts; here the degraded image's own vision-language embedding is corrected by a token mapper into the clean target's instruction embedding, with the text prompt left empty. In a matched comparison on the same backbone, test inputs, preprocessing, and metric code, the image instruction leads the text adapter on all six tasks: 21.3 dB versus 16.3 dB on low-light and 33.2 dB versus 32.5 dB on deraining.

A task-agnostic variant performs on par with the task-aware model without any degradation label, whereas text conditioning degrades markedly once the task label is removed. The text adapter collapses with a single neutral prompt (deraining drops from 32.5 dB to 17.5 dB), and per-image prompts stay within about 1 dB of the neutral prompt; task-agnostic ImIR exceeds the task-aware text adapter in PSNR on all six tasks. Task-agnostic ImIR versus task-aware ImIR is 21.2 dB versus 21.3 dB on low-light and 33.0 dB versus 33.2 dB on deraining, with the same pattern on the other tasks; the authors attribute the difference to image conditioning rather than to the editor or adapter.

Because the instruction is a continuous vector, scaling its residual shift yields a family of valid restorations for tasks whose target is not unique. A fixed text prompt offers no comparable continuous control; on low-light enhancement the scale parameter produces a spectrum of exposures from near the dark input to over-brightening, and on dehazing the output moves progressively closer to the ground truth. The trends are reflected by mean-luma and mean Color Attenuation Prior readouts; the authors note the control axis is meaningful for tasks whose target is a range and close to inert for tasks whose target is a point.

Ablations in image space attribute the gain to the predicted instruction itself rather than to degradation cues the degraded image already carries. Relative to the identity baseline with no mapper, the mapper gains 1 to 5 dB; an oracle using the clean embedding confirms that the clean embedding is a useful instruction and bounds the mapper's remaining headroom. ImIR closes at least 82 percent of the identity-to-oracle gap on every task; adding the global-context branch lifts the reduction in embedding-space MSE over the identity baseline from 25.1 percent to 30.3 percent and raises cosine similarity from 0.724 to 0.742, worth 1 to 5 dB in image space.

Perspective

The result targets full-reference restoration evaluated with paired data: one benchmark per task (LOLv2-real, Rain100L, RESIDE SOTS, GoPro, SIDD, Kodak), a training set balanced across tasks with 688 pairs in total, a frozen Qwen-Image-Edit backbone (the 2511 release reported as primary), a rank-16 low-rank adapter, and 326M learned parameters in the mapper and adapter (235M in the adapter, 91M in the mapper) trained in about three hours on one H100. It lets researchers and practitioners cover several degradation types with a single model without writing prompts or supplying a degradation label, and choose an operating point such as exposure by scaling the instruction on tasks whose target is not unique. The authors explicitly restrict the efficiency claim to data and training time and make no compute-normalized comparison; from-scratch full-data specialists still lead on full-reference fidelity for several tasks, and the work does not claim to beat them there.

The authors run no dedicated generalization study and do not claim it as a result; blind and real-world benchmarks are left to future work, so how the task-agnostic variant behaves on out-of-distribution degradations remains open. Inference is iterative sampling, taking 14 seconds per image at about one megapixel with batch size one and 30 steps, slower than feed-forward specialists, and the generative prior can synthesize plausible but unfaithful detail. On deraining the image instruction and the text adapter are close to a statistical tie on PSNR and SSIM, and the denoising LPIPS win rate of 69 percent rests on a small test set of 32 images. In addition, this document is a full-text parse in which tables and figures are given as text, so checking exact numbers and the visual comparisons of Fig. 3 and Fig. 4 still calls for the original figures.

Sources