A perceptual distance from a network that only learned to see

Measure an image change with early features, and weight those features by how much the encoder's output responds to them. On four standard databases the result matches human judgements better than distances trained on them.

Thesis. A frozen self-supervised encoder already contains a perceptual distance. It takes one matrix to read it out, and that matrix needs no human labels.

Main technical point. Early patch features see every local change but weight all directions alike. The encoder's output Jacobian tells us which of those directions matter. Keeping the top 64 of 384 raises agreement with human scores on TID2013 from 0.784 to 0.850.

Practical implication. JLD reaches a mean Spearman correlation of 0.926 on four standard databases, above LPIPS, DISTS, PieAPP and DreamSim, and holds its accuracy when image resolution doubles. The fit takes 35 seconds on 100 unlabelled images.

The problem

Every codec, every restoration network and every image generator needs a number that says how different two images look to a person. The oldest answer is pixel error. PSNR is cheap and differentiable, and it treats a change to any pixel like a change to any other. People do not.

The JLD pipeline, and five distortions of one TID2013 image at nearly equal PSNR with their human scores, per-patch maps, and the rank each measure assigns.
Five TID2013 distortions of one image, all between 26.8 and 27.2 dB. Human scores run from 5.58 to 2.64. The table at the bottom gives the rank each measure assigns; only JLD reproduces the human order. The top row is the whole method.

Across every pair of TID2013 images with equal PSNR, PSNR picks the one people prefer 51.1% of the time, which is a coin flip.

The standard fix is to compare images in the feature space of a pretrained network and then fit that comparison to human ratings. LPIPS learns a weight per channel from human choices. DISTS fits weights on feature statistics. DreamSim fine-tunes the encoder itself on human triplets. These work, but the fit ties the distance to the data and the resolution it was fitted on. Rescore TID2013 at 512 pixels instead of 256 and DISTS falls from 0.815 to 0.717 Spearman correlation with human scores. LPIPS-VGG falls from 0.757 to 0.661. The ratings are the same. Only the pixel grid changed.

So we asked a narrower question. Does a network that has only learned to see, with no human quality label anywhere in its training, already contain a useful perceptual distance?

Two layers, two halves of the answer

We looked inside a frozen DINOv2-S, a 12-block vision transformer trained by self-supervision. No single depth gives a good distance, and the two ends fail for opposite reasons.

Patch tokens after the first block are local. Each one has looked at its own 14 × 14 patch and interacted once with its neighbours, so a distortion anywhere in the image moves some token. But a token has 384 dimensions, and plain Euclidean distance counts movement along all of them equally. Many of those directions barely affect anything the network computes later. Raw block-1 distance reaches 0.784 on TID2013.

The output of the encoder has the opposite problem. It has decided what matters, and in doing so it has pooled away the spatial detail that blur, blocking and noise act on. Distance between output tokens reaches 0.742.

One end has the locality and the other has the weighting. JLD takes the features from block 1 and the weights from the output. The object that connects a layer to the output is the Jacobian.

The Jacobian lens

Write \(\mathbf h_t\) for the block-1 token at patch \(t\) and \(\mathbf z\) for the encoder's final class token. The Jacobian \(J_t=\partial\mathbf z/\partial\mathbf h_t\) answers a simple question: if we nudge this one patch token by \(\delta\), how far does the output move? To first order, by \(J_t\delta\). Square that and average over natural images and patch positions:

\[ \mathbf M=\mathbb E_x\Big[\frac1N\sum_{t=1}^N J_t(x)^{\top}J_t(x)\Big]. \]

\(\mathbf M\) is a 384 × 384 matrix, and \(\delta^{\top}\mathbf M\,\delta\) is the expected squared output change caused by moving a patch by \(\delta\). Its eigenvectors sort the feature directions by how strongly the output responds to them. A push along the first eigenvector changes the output the most. A push along the last leaves it almost unchanged.

For DINOv2-S this spectrum is steep. The top 64 eigenvectors hold 76.5% of the trace. We keep those 64 as the columns of a matrix \(\mathbf U\) and call it the Jacobian lens. It is one fixed linear projection, shared by every image, every patch position and every resolution.

Left: the fitting procedure. Middle: a sketch of a feature change split into the part the lens retains. Right: a scatter plot showing lens displacement tracks the encoder output change better than raw displacement.
Left, the fit. Middle, what the projection does to one feature change. Right, a check on 153 perturbations at 40 dB: lens displacement tracks the change in the encoder output (ρ = 0.643), raw block-1 displacement hardly does (ρ = 0.147).

Fitting it in 35 seconds

We never form a Jacobian. Draw a random Gaussian vector \(v\), take the scalar \(v^{\top}\mathbf z\), and backpropagate once. The gradient that lands on each patch token is \(g_t=J_t^{\top}v\), and averaging \(g_tg_t^{\top}\) over probes gives \(J_t^{\top}J_t\) in expectation. One backward pass yields an estimate at every patch at once.

The shipped lens uses 100 unlabelled DIV2K images, four 224 × 224 crops from each, and eight probes per crop. That is 35 seconds on one A100. There are no image pairs, no distortions and no ratings in the fit. Refitting with fresh crops and probes recovers the same subspace, with an overlap of 0.962.

Human data enters in one place, and we want to be exact about it. We used CSIQ and a development split of KADID-10k to choose among 245 configurations: which block, how many directions, how to weight the global term. TID2013 informed the chroma pre-filter. LIVE, the KADID-10k test references and PIPAL were held out.

The distance

With the lens in hand, the distance has three parts. A fixed pre-filter \(R\) blurs the two chroma channels with a 2-pixel Gaussian and leaves luminance alone, because human vision resolves colour more coarsely than brightness. Then

\[ \begin{aligned} D(x,y)={}&\Big[\frac1N\sum_{t=1}^N \big\|\mathbf U^{\top}\big(\mathbf h_t(Rx)-\mathbf h_t(Ry)\big)\big\|^2\Big]^{1/2}\\ &+\tfrac12\big[1-\cos\big(\mathbf z(Rx),\mathbf z(Ry)\big)\big]. \end{aligned} \]

The first term is the lens term: the root mean square, over patches, of the token change as seen through the lens. Its per-patch values make a map that shows where the distance comes from. The second is a small cosine term on the output. It catches whole-image changes such as a global colour shift, which are spread too thinly over patches to register in the first term.

What the lens throws away

The clearest way to see what the lens does is to split each token's movement into the part it keeps and the part it discards. A change with no preferred direction would keep 64/384, or 16.7%, of its energy.

Per-patch scatter of kept against discarded block-1 displacement for JPEG, Gaussian blur and colour noise at three distortion levels of one TID2013 image.
Three distortions of one TID2013 image at levels 1, 3 and 5. Each point is a patch: horizontal is the displacement the lens keeps, vertical is what it discards. Colour noise moves tokens a long way, mostly upward.

Take JPEG and colour noise at the strongest level on this image. Their PSNR is nearly identical, 21.6 and 21.3 dB. People rate the JPEG image 1.66 and the noisy one 4.47. The lens keeps 14.3% of the JPEG displacement and 6.4% of the noise displacement, so JLD comes out four times larger for JPEG (0.516 against 0.129), which is the human verdict. Colour noise moves the tokens plenty, but along directions the output ignores. JPEG block edges move them along directions the output uses.

Over all 24 distortion types the lens keeps a median of 10.6% of the displacement energy.

Because the lens is a fixed linear map, the lens term is a true pseudometric and has a local form we can write down: near an image it is a quadratic form in pixel space whose iso-distance contours are ellipses. We checked those ellipses against measured contours on real images, and the radii agree to within 9 to 11%. Around one image the contours are 2.33 times longer along chroma than along blur, so a blur costs far more distance than a colour change of equal pixel energy.

Against LPIPS, DISTS and DreamSim

We evaluated 17 full-reference distances under one protocol. A few rows from the main table:

DistanceWhere the weighting comes fromMean SRCCms per pair
PSNRNone0.7601.4
LPIPS-VGGChannel weights learned from human choices0.80210.9
DreamSimEncoder fine-tuned on human triplets0.87151.3
DISTSWeights on feature statistics, fitted to human scores0.9009.9
VSIHand-built saliency and gradient model0.9159.9
JLD-fastThe encoder's own output sensitivity0.9112.5
JLDThe encoder's own output sensitivity0.92612.2

Mean Spearman correlation with human ratings over TID2013, CSIQ, LIVE and KADID-10k test. Times are on an A100 at 512 × 384.

JLD has the highest mean of the 17 and beats every learned deep distance on each of the four databases.

The ablation says where the gain comes from. All 384 raw directions give 0.772. So do 64 random directions. 64 PCA directions give 0.851, and the 64 lens directions give 0.905. All three subsets have 64 directions, so the gain comes from which ones are kept. The chroma filter and the output term add the rest.

SRCC against resolution on TID2013 and CSIQ, mean SRCC against time per pair, and video SRCC on two databases.
(a, b) Correlation with human scores as the short side grows from 256 pixels to native size. (c) Accuracy against time per pair. (d) Video.

The resolution result is the one we care about most. From 256 to 512 pixels, the lens term reads 0.850, 0.860, 0.845 on TID2013. The same block-1 features without the lens fall from 0.784 to 0.656.

JLD-fast stops the encoder after block 1 and drops the output term. It keeps a mean of 0.911 at 2.5 ms per pair, 4.3 times faster than LPIPS-VGG. Only PSNR is faster.

The construction also carries to video by swapping the encoder. A 16-direction lens on block 2 of a frozen VideoMAE-Base, fitted on 50 unlabelled clips, reaches 0.786 on Waterloo IVC 4K, where VMAF reaches 0.611. And it is not special to DINOv2: fitting a lens to each of 14 frozen encoders, including CLIP and VGG16, improved on raw features in all 14.

Honest edges

How to use it

from jld import JLD

metric = JLD.pretrained("full")                       # or "fast"
distance = metric("reference.png", "distorted.png")   # lower means more similar

Clone the repository and run uv sync. The fitted lens ships inside the package and the DINOv2-S weights download on first use. A few things worth knowing:

Citation

@misc{saini2026jld,
  title         = {{JLD}: Perceptual Distance Through A Jacobian Lens},
  author        = {Saini, Shreshth and Adsumilli, Balu and Bovik, Alan C.},
  year          = {2026},
  eprint        = {2610.05967},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CV},
  doi           = {10.48550/arXiv.2610.05967},
  url           = {https://arxiv.org/abs/2610.05967}
}