AutoRef lets a coding agent rewrite the harness automatically, lifting frozen FLUX.2 [klein] 4B from 5.72 to 7.37 on four-reference MultiBanana and matching Nano Banana Pro and GPT-Image-1.5
Synopsis
AutoRef has a coding agent iteratively rewrite harness code while keeping both the image generator and the reasoning model frozen, separating the tasks whose feedback informs proposals from those used to select candidates and continuing the search from a beam of top-ranked harnesses; the resulting AutoRef-Harness raises open-weight FLUX.2 [klein] 4B from 5.72 to 7.37 (out of 10) on the held-out four-reference MultiBanana test split, matching or exceeding proprietary models such as Nano Banana Pro and GPT-Image-1.5, and transfers without re-optimization across reference counts, benchmarks and evaluators, generators, and reasoning models.
Interpretation
AutoRef treats harness optimization itself as the search object: the image generator and reasoning model stay frozen, and only the executable program that specifies how references are interpreted, prompts are constructed, candidates are generated and evaluated, and the final image is selected gets rewritten. Earlier harness optimization methods such as Meta-Harness mainly target tasks verifiable with discrete labels or executable tests, whereas image generation relies on noisy visual evaluation where scalar scores give little diagnostic information; AutoRef redesigns the search loop for that setting. The paper formalizes the harness optimization objective (Equations 2 and 3) and states that optimization acts only on the executable code surrounding the models, allowing changes to prompting, generation, evaluation, selection, and control flow.
AutoRef separates the tasks used to propose harness updates from those used to select among candidates, and uses beam search that keeps the top-ranked harnesses on the selection tasks as parents for the next iteration. Meta-Harness lets the proposer read the full history and ranks candidates on the same tasks whose feedback it reads, while Greedy Search keeps only the single best harness; AutoRef combines train-validation separation with a multi-parent beam to reduce overfitting to both the search tasks and evaluator noise. Under the same generator, reasoning model, and evaluator, AutoRef yields the strongest final harness on the held-out test split, followed by Greedy Search and Meta-Harness; Meta-Harness and Greedy Search draw 4.2 and 5 images per task versus 3 for AutoRef-Harness, and AutoRef's second-best harness still outperforms both methods' final harnesses.
The discovered AutoRef-Harness consists of four separable components: reference-grounded prompting, structurally diverse drafts, complaint-directed revision, and failure-aware selection that counts hard failures before pairwise judgment. Compared with human-written harnesses, it specializes and chains each step for multiple references: every prompt assigns each requested element to its reference, the two drafts differ in prompt structure rather than only sampling, complaints name the reference they concern, and selection first checks for hard failures such as a missing reference. Leave-one-out ablation shows that removing structurally diverse drafts (7.37 to 7.01), complaint-directed revision (to 6.99), or failure-aware selection (to 6.93) lowers the average; grounding only with a single image reaches 6.91, close to IPR at 7.02 with three images, selection only reaches 6.15, close to Best-of-3 at 6.01, and no partial combination matches the full harness at 7.37.
The gains of AutoRef-Harness do not depend on the exact configuration used during search: it still improves results when the number of references, the benchmark and evaluator, the generator, or the reasoning model changes. The search runs only on the four-reference MultiBanana training and validation splits, with the test split unseen during search; transfer tests cover three- and five-reference settings, the never-used OmniContext benchmark with its official GPT-4.1 evaluator, generators such as FLUX.2 [klein] 9B and Qwen-Image-Edit-2511, and the open-weight Qwen3-VL-32B reasoning model. Four-reference held-out test split improves from 5.72 to 7.37; three references from 6.94 to 7.76 and five references from 5.19 to 6.36; on OmniContext FLUX.2 [klein] 4B rises from 8.30 to 8.85; with Qwen3-VL-32B it still improves from 5.72 to 6.90; in human evaluation it wins 70% against its base generator (17% losses), 64% versus 24% against FLUX.2 [klein] 9B, 63% versus 27% against Seedream 4.5, and is competitive with Nano Banana Pro at 46% versus 40%.
Perspective
This work targets multi-reference image generation specifically, and applies to settings where one wants to improve how existing models are combined without retraining the generator, such as advertising, virtual try-on, and content creation where people, objects, clothing, backgrounds, and styles are specified through separate images. It enables follow-up work to keep optimizing prompt construction, candidate generation, diagnosis, and selection around frozen models, and indicates that harness optimization and model-weight optimization such as DyRef fine-tuning can be combined. The authors release the AutoRef implementation and AutoRef-Harness together with data splits and pinned model revisions, so the protocol can be reproduced and extended; they also note that the evaluator and proposer are replaceable, making stronger vision-language models and coding agents, and richer evaluators such as ones combining a VLM with segmentation models, possible directions.
The search maximizes a single vision-language evaluator's score; although the paper reports that the harness transfers to OmniContext and its GPT-4.1 evaluator, how the evaluator's own preferences shape the final harness structure remains worth watching. The harness only changes how a frozen generator is used, so when the generator itself rarely produces an acceptable image at a given reference count, as with Qwen-Image-Edit-2511 at five references where the hard-failure check flagged all three drafts on 87 of 96 tasks, better selection cannot compensate, and the paper notes that overcoming this limit likely requires reducing how many references the generator must compose at once. In addition, the human evaluation uses four raters and 50 tasks per baseline, and the qualitative examples were selected among tasks with the largest score gap between the base generator and AutoRef-Harness rather than sampled randomly, so they illustrate the failures removed rather than the overall distribution. The search runs five iterations with fixed beam and candidate counts, and what larger budgets or different proposers would yield is left as an open direction.
