Image quality scoring drops human labels: one Jacobian from a frozen encoder, and correlation barely moves when resolution doubles
Lead
A perceptual distance derived from the output sensitivity of a frozen vision encoder, fitted once on 100 unlabeled images, reaches the highest mean correlation across four image databases and stays nearly unchanged when resolution doubles.
Story
Image compression, restoration and generation all need to measure how different two images look to a person, while pixel error treats every pixel change equally and disagrees with human vision. The most accurate perceptual distances were typically fitted to human judgements, tying them to a fixed data and resolution. On TID2013, doubling the image resolution from 256 to 512 pixels drops the correlation of DISTS with human scores from 0.815 to 0.717 and LPIPS-VGG from 0.757 to 0.661.
The new distance combines the locality of early patch features with the perceptual sensitivity carried by later representations, weighting directions by the encoder output's Jacobian with respect to an early patch token. Earlier choices either used early features directly, weighting all directions alike, or used late representations that have already pooled away spatial detail. On a frozen DINOv2-S, unweighted early features reach a correlation of 0.784 and output tokens 0.742, while the weighted features reach 0.850.
The Jacobian lens is one fixed metric tensor in feature space, keeping the 64 leading eigendirections of 384 that account for 76.5% of its trace. No perceptual distance had previously transferred a network's output sensitivity back to its early features in this way. The lens is fitted once from 100 unlabeled DIV2K images in about 35 seconds on one A100 GPU, using no image pairs, distortions or human labels.
The same construction extends to video by swapping the encoder, using a 16-direction lens to capture temporal effects such as flicker that no single frame reveals. Frame-wise image distances could not cover such temporal effects. On the 240 Waterloo IVC 4K pairs it reaches 0.786 against 0.611 for VMAF, and on held-out AVT-VQDB-UHD-1 it reaches 0.912 against 0.932 for VMAF.
What to watch
Next steps can fit a lens for other differentiable encoders such as audio, medical imaging or 3D by the same procedure, needing roughly 100 representative unlabeled samples disjoint from evaluation data and storing the corpus manifest and random seed. Applications that need a strict pseudometric can use the lens term alone, which satisfies the triangle inequality; using the full distance as a training loss requires care about the null space formed by the 320 discarded feature directions.
The lens term is invisible along the 320 discarded directions, so direct optimisation against it can exploit them, and the full distance is not guaranteed to satisfy the triangle inequality, with at most 107 violations per million audited triplets. It is weaker on PIPAL restoration outputs and on TID2013 global contrast changes, and on held-out video it still trails VMAF, which is trained on human video scores. The human study has only ten participants, and the method assumes spatially aligned inputs.
