Iris-3B pretrains a 3B pixel-space text-to-image model from scratch and finds no pixel-space prior advantage on depth or restoration
Synopsis
The authors pretrain Iris-3B, a 3B-parameter pixel-space text-to-image transformer trained from scratch through a 256→512→1024 curriculum with v-prediction and DINOv2 REPA, and convert the latent FLUX.2 Klein base 4B to pixel space; after fine-tuning both pixel backbones for monocular depth and image restoration, they find no significant improvement from a pixel-space generative prior: on depth Iris-3B is level with latent FLUX.2 Klein while the converted model trails it, and on DIV2K restoration neither pixel model beats a latent FLUX.2 Klein fine-tune.
Fig. 2: x x - versus v v -prediction at 256 2 256^{2} , absolute scores at matched steps (EMA, 25-step FlowDPM++, CFG 4.5). Following JiT [ 5 ] , the x x -prediction arm also uses its logit-normal timestep distribution ( 0.8 , 0.8 ) (0.8,0.8) instead of ( 0.0 , 1.0 ) (0.0,1.0) , so the contrast does not isolate the prediction target.
arXivInterpretation
Iris-3B shows pixel-space text-to-image pretraining scales to 3B parameters: under the official evaluators it reaches GenEval 0.798, DPG 86.52, LongText 0.857 and OneIG 0.540, matching Qwen-Image on OneIG while trailing i1, Z-Image and Qwen-Image on GenEval (the latter two use prompt rewriting), with English long-text rendering its clearest gap. Prior pixel-space generation was mostly validated at smaller scale or through converted models; this provides a from-scratch 3B pixel-space text-to-image model with released weights and training code, plus the 256px ablation that selected its recipe. Scored with the official, unmodified GenEval, DPG-Bench, LongText-Bench and OneIG-Bench evaluators at 1024 resolution with CFG 3, 100 sampling steps and 4 samples per prompt, every sample scored (no best-of), on the released SFT step-665K checkpoint.
A 256px ablation campaign on a 1.3B PixelDiT reproduction finds v-prediction tied with x-prediction on GenEval at every milestone while x-prediction is worse on DPG and FID, and among alignment objectives REPA with DINOv2 is strongest overall, a DINOv3 teacher lowers GenEval at every milestone and worsens FID, iREPA gains only on FID at 150K and 200K with that gain nearly gone by 250K, and Self-Flow trails on all four metrics. It places the prediction target and four alignment objectives (REPA, iREPA, a DINOv3 teacher, Self-Flow) under one data subset, global batch 256 and one scoring protocol, giving a controlled basis for choosing a pixel-space pretraining recipe. Each arm is compared with its own control at matched steps under one protocol: EMA weights, 25-step FlowDPM++ sampling, CFG 4.5, GenEval, DPG-Bench, FID-10K against a held-out split and CLIP score; the authors note the x-prediction arm also switched to JiT's logit-normal timestep distribution, so that contrast does not isolate the prediction target alone.
Using a pixel-space generative prior for dense prediction did not deliver the expected gain: on depth Iris-3B is level with latent FLUX.2 Klein on the mean of AbsRel and δ1 (mean AbsRel 0.071 vs 0.072) while the converted pixel model trails on every benchmark except KITTI, and on DIV2K restoration the converted pixel model is worse on most metrics including SSIM and NIQE, with Iris-3B best in the table on MUSIQ but slightly below both FLUX.2 Klein models on fidelity and DeQA. It directly tests the hypothesis that removing the lossy VAE should help detail-critical tasks, covering both routes to a pixel backbone (from-scratch pretraining and latent-to-pixel conversion) and documenting reproducible depth and restoration recipes. Depth uses one matched direct-regression recipe and budget, scored on NYUv2, KITTI, ETH3D, ScanNet and DIODE with the unmodified Marigold V2 evaluation code; restoration uses the one-step HYPIR recipe on DIV2K with PSNR, SSIM, LPIPS, NIQE, MUSIQ and DeQA. The authors state each arm is a single run at a short budget with no significance test, so small differences should be read as no advantage.
Pixel-space models leave a measurable patch grid in dense predictions: the mean curvature of predicted depth at patch borders divided by its mean elsewhere is about one for the latent model (0.97–0.99), 1.31–1.44 for the converted pixel model, and highest for Iris-3B at 1.61–2.25. It turns a faint artifact into a direct quantitative measure and attributes it to non-overlapping patch tokenization with largely per-patch decoding, in contrast to the latent model's VAE decoder with overlapping receptive fields. The ratio is reported on Hypersim, KITTI, ETH3D and DIODE; the authors also note PixelUMM independently reports the same grid and finds convolutional output heads suppress it.
Perspective
This work is aimed at researchers and engineers weighing pixel-space backbones: it provides a downloadable 3B pixel-space text-to-image model with training code, a pretraining recipe selected by 256px ablations (v-prediction with DINOv2 REPA), and directly reusable recipes for direct-regression depth and one-step HYPIR restoration. The latent-to-pixel conversion contributes two concrete engineering steps: a gain-matched ridge initialization of the new input projection, and scaling the trunk learning rate by the ratio of the parent's weight RMS to a reference model's. The authors report that together these turn patch noise into photorealistic, prompt-aligned images within 1K steps, while neither alone produces coherent images. The results apply to text-to-image generation at roughly a one-megapixel budget and to monocular depth and image restoration with a generative backbone.
The authors state that the downstream comparisons are confounded: Iris-3B has no latent twin, so its depth comparison with latent FLUX.2 Klein mixes representation space with differences in size, pretraining data and compute; the two matched comparisons share one converted parent whose generative prior is likely weaker than that of the latent model it came from, so they compare a briefly converted model with a fully pretrained one; each depth and restoration arm is a single short-budget run with no significance test; and the Iris-3B restoration arm is not matched to the FLUX.2 Klein arms. Generation quality (GenEval, DPG, FID) of the converted models was not measured, so the attribution of the converted model's deficit to conversion rather than to pixel space remains a hypothesis. The ridge initialization and LR scaling were validated only in probes at global batch 8 and 1K steps, and it is untested whether the benefit survives at the global batch of 128 used by Jiang et al. In addition, this is a fast parse in which formulas and some numeric values appear as placeholders, so exact hyperparameters and loss weights should be checked against the original text before reproduction.
