Research release · 27 September 2026
Data and release report
This release collects the preprint, the image-JLD implementation, fitted lens and selected recorded image and video results. It adds a browser presentation without rerunning metric inference or changing measured scores.
What you can download
| Record | Scope |
|---|---|
| Paper PDF | Preprint, 35 pages |
| Data records ZIP | JSON and CSV evidence, selection rules and reports |
| Image comparison summary | Six cases, twelve distortions, JLD, DISTS, LPIPS-VGG, PSNR and MOS |
| Main qualitative records / gallery records | Full local map arrays, source hashes and selection criteria |
| Image benchmark table | Recorded dataset correlations, including compared baselines |
| Measured runtime table | Original hardware/protocol-dependent measurements |
| Resolution study | Recorded per-resolution evaluations |
| Video study | Recorded aggregate comparisons |
| Waterloo displayed scores | Four VideoMAE comparison pairs |
| AVT displayed scores | Four separate framewise image-JLD pairs |
| Video map explorer | Four scenes, twelve displayed frames per scene, map scales, timestamps, source-array hashes and movie provenance |
| Paper video profile | Two 200-frame distance curves and original high/low-frame figure evidence |
| Demo distances | Eight Kodak pairs, four metric variants |
| SHA-256 manifest | All published site assets except the manifest itself |
Image selection and maps
The first two cases maximize human MOS gaps among same-reference comparisons with PSNR differences below 0.02 dB, where JLD agrees with MOS and DISTS and LPIPS-VGG reverse it. The gallery adds two agreements and two reversals using its archived exclusion and PSNR rules. A has higher MOS in all six cases. KADID examples use the held-out reference split. These are selected demonstrations, not a random sample or an accuracy estimate.
All 18 original source image files were matched to their archived SHA-256 values before the recorded 504 × 378 center crops were exported. Map values come directly from the recorded 27 × 36 local displacement arrays. They use nearest-neighbor rendering and a shared viridis scale from zero to the largest value within each pair. The full scalar metric includes the global CLS term; the spatial maps do not. Source paths in exported JSON use portable descriptive roots.
Video protocols
The four Waterloo examples use the VideoMAE-B/K400 block-2, 16-dimensional temporal lens. Human MOS favors A in all four. The lens agrees for railway, football and dance; it reverses the game example, where VMAF agrees. These clips were chosen for metric disagreements. The full source identities, recorded score rows and display frames are in the Waterloo evidence record.
The four AVT examples belong to a separate framewise image-JLD pilot with no temporal encoder and no global term. Twelve 448 × 252 frames score each complete clip; human ratings were collected at 4K. The new map movies contain synchronized images and actual recorded local response maps. The original comparison movies contain images and score cards. The orange example is a shared failure of the compared measures. The AVT protocol and selection records specify the study.
The eight original comparison MP4 files retain the exact hashes in the supplement manifest. The four additional map MP4s are copied unchanged from the research archive; browser-compatible WebM derivatives are supplied too. The excerpts are display media, not the complete evaluated sequences. See the video README and video manifest.
The frame explorer uses twelve evenly spaced frames from each 90-frame map
movie. Reference and distorted panels are lossless crops of the decoded
display movie. Maps are nearest-neighbor renderings of the square roots of
the archived squared-displacement arrays, with a single linear magma scale
from zero to the absolute maximum over A, B and all 90 display frames.
Nothing is clipped. Timestamps come from the original encoding rate and
source-frame indices. The original arrays and their hash receipts are in
video-map-arrays/. Clip scores pool twelve separately sampled
frames; displayed frames can have a different ordering.
The paper's separate full-1080p profile scores every third frame of two ten-second Big Buck Bunny encodes. The unchanged paper figure shows local maps at the highest and lowest image-JLD frames. Its overlays share a 99.5th-percentile display limit. That figure's display normalization and full-image distance curves differ from the spatial pilot above; see the original profile measurements.
Direct media: railway, football, dance, game, animation, orange, vegetables, water.
Verification and scope
The release check validates local page links, published asset hashes, all eight video hashes against the archived supplement manifest, image dimensions, score-table identity and the bundled fitted-lens hash. The new video maps also pass pixel-exact checks against their archived arrays; all four map WebMs decode and play in Chromium. The Python implementation and its original recorded demo scores are carried forward from the tested release package. Model inference was not rerun for the website. Current lightweight checks do not establish fresh end-to-end benchmark reproduction.
The compact package implements image JLD. The temporal-video pipeline and bulk benchmark datasets are not included. The paper reports selected negative controls and limits of transfer. The lens term is a pseudometric; the complete score need not satisfy the triangle inequality. Runtime values depend on the reported hardware and evaluation protocol.
Credits and reuse
The code's MIT grant applies to the implementation. It does not grant rights to third-party benchmark media, model weights or the manuscript. Dataset and source-content terms continue to apply.
- KADID-10k: Hanhe Lin, Vlad Hosu and Dietmar Saupe, QoMEX 2019. The official dataset is freely available to the research community and requests citation.
- TID2013: Nikolay Ponomarenko and collaborators, 2013. Refer to the dataset's official documentation and publication when using its images or ratings.
- Kodak Lossless True Color Image Suite: Eastman Kodak images hosted by Rich Franzen. This release uses kodim23, kodim03 and kodim05. The mirror describes its understanding of unrestricted use; this release does not independently grant image rights.
- Waterloo IVC 4K: compression-video database. Source-content credits and dataset terms remain with the original providers.
- AVT-VQDB-UHD-1: Rao and collaborators, IEEE ISM 2019. Animation source: Big Buck Bunny, Blender Foundation. Other source identities are preserved in the AVT manifest.
- DINOv2: frozen encoder; cite its paper and follow the upstream model terms.
- LPIPS: Zhang and collaborators, CVPR 2018. DISTS: Ding and collaborators, IEEE TPAMI. Cite these methods when reporting their measurements.
The project page links to the preprint PDF. An arXiv identifier has not been assigned yet.