← JLD project page

Research release · 27 September 2026

Data and release report

This release collects the preprint, the image-JLD implementation, fitted lens and selected recorded image and video results. It adds a browser presentation without rerunning metric inference or changing measured scores.

What you can download

RecordScope
Paper PDFPreprint, 35 pages
Data records ZIPJSON and CSV evidence, selection rules and reports
Image comparison summarySix cases, twelve distortions, JLD, DISTS, LPIPS-VGG, PSNR and MOS
Main qualitative records / gallery recordsFull local map arrays, source hashes and selection criteria
Image benchmark tableRecorded dataset correlations, including compared baselines
Measured runtime tableOriginal hardware/protocol-dependent measurements
Resolution studyRecorded per-resolution evaluations
Video studyRecorded aggregate comparisons
Waterloo displayed scoresFour VideoMAE comparison pairs
AVT displayed scoresFour separate framewise image-JLD pairs
Video map explorerFour scenes, twelve displayed frames per scene, map scales, timestamps, source-array hashes and movie provenance
Paper video profileTwo 200-frame distance curves and original high/low-frame figure evidence
Demo distancesEight Kodak pairs, four metric variants
SHA-256 manifestAll published site assets except the manifest itself

Image selection and maps

The first two cases maximize human MOS gaps among same-reference comparisons with PSNR differences below 0.02 dB, where JLD agrees with MOS and DISTS and LPIPS-VGG reverse it. The gallery adds two agreements and two reversals using its archived exclusion and PSNR rules. A has higher MOS in all six cases. KADID examples use the held-out reference split. These are selected demonstrations, not a random sample or an accuracy estimate.

All 18 original source image files were matched to their archived SHA-256 values before the recorded 504 × 378 center crops were exported. Map values come directly from the recorded 27 × 36 local displacement arrays. They use nearest-neighbor rendering and a shared viridis scale from zero to the largest value within each pair. The full scalar metric includes the global CLS term; the spatial maps do not. Source paths in exported JSON use portable descriptive roots.

Video protocols

The four Waterloo examples use the VideoMAE-B/K400 block-2, 16-dimensional temporal lens. Human MOS favors A in all four. The lens agrees for railway, football and dance; it reverses the game example, where VMAF agrees. These clips were chosen for metric disagreements. The full source identities, recorded score rows and display frames are in the Waterloo evidence record.

The four AVT examples belong to a separate framewise image-JLD pilot with no temporal encoder and no global term. Twelve 448 × 252 frames score each complete clip; human ratings were collected at 4K. The new map movies contain synchronized images and actual recorded local response maps. The original comparison movies contain images and score cards. The orange example is a shared failure of the compared measures. The AVT protocol and selection records specify the study.

The eight original comparison MP4 files retain the exact hashes in the supplement manifest. The four additional map MP4s are copied unchanged from the research archive; browser-compatible WebM derivatives are supplied too. The excerpts are display media, not the complete evaluated sequences. See the video README and video manifest.

The frame explorer uses twelve evenly spaced frames from each 90-frame map movie. Reference and distorted panels are lossless crops of the decoded display movie. Maps are nearest-neighbor renderings of the square roots of the archived squared-displacement arrays, with a single linear magma scale from zero to the absolute maximum over A, B and all 90 display frames. Nothing is clipped. Timestamps come from the original encoding rate and source-frame indices. The original arrays and their hash receipts are in video-map-arrays/. Clip scores pool twelve separately sampled frames; displayed frames can have a different ordering.

The paper's separate full-1080p profile scores every third frame of two ten-second Big Buck Bunny encodes. The unchanged paper figure shows local maps at the highest and lowest image-JLD frames. Its overlays share a 99.5th-percentile display limit. That figure's display normalization and full-image distance curves differ from the spatial pilot above; see the original profile measurements.

Direct media: railway, football, dance, game, animation, orange, vegetables, water.

Verification and scope

The release check validates local page links, published asset hashes, all eight video hashes against the archived supplement manifest, image dimensions, score-table identity and the bundled fitted-lens hash. The new video maps also pass pixel-exact checks against their archived arrays; all four map WebMs decode and play in Chromium. The Python implementation and its original recorded demo scores are carried forward from the tested release package. Model inference was not rerun for the website. Current lightweight checks do not establish fresh end-to-end benchmark reproduction.

The compact package implements image JLD. The temporal-video pipeline and bulk benchmark datasets are not included. The paper reports selected negative controls and limits of transfer. The lens term is a pseudometric; the complete score need not satisfy the triangle inequality. Runtime values depend on the reported hardware and evaluation protocol.

Credits and reuse

The code's MIT grant applies to the implementation. It does not grant rights to third-party benchmark media, model weights or the manuscript. Dataset and source-content terms continue to apply.

The project page links to the preprint PDF. An arXiv identifier has not been assigned yet.