Challenges of HDR

An abstract particle flow from the smaller BT.709 gamut to the larger BT.2020 gamut.

From capture and SDR-to-HDR reconstruction to collecting human judgments at scale.

HDR gives us more room to represent light, from a dim room to a bright window in the same shot. Using that room well is harder. A recording, a reconstructed highlight, and a human quality score all depend on what happens between the scene and the screen. These are the problems behind our work on Beyond8Bits, HDR quality assessment, and SDR-to-HDR conversion.

What HDR changes

A recording of a shaded room and a sunlit window must fit both into the range its pipeline can capture, encode, and display. HDR widens that range, preserving differences among bright surfaces and detail in shadows. The result still depends on the recording and the viewer’s screen.

Start with this Beyond8Bits recording. The same sky, autumn trees and reflections appear in both views. The SDR version compresses the bright parts to fit a 100-nit reference display; the original keeps its native HDR signal. The matching crops and plots show what changes before your screen applies its own display mapping.

Tone-mapped SDR · BT.709
SDR clip loads here
Sky, shoreline and reflections
Matched detail loads with the clip
Native HDR · PQ / BT.2020
Native HDR clip loads here
Sky, shoreline and reflections
Matched detail loads with the clip

The paired clips load together and loop when playback is supported.

Color and luminance in the marked regionWaiting for decoded samples
Decoded color and luminance samples at matching pixel locations and frame times.
SDR · 100-nit reference whiteSamples unavailable
HDR · decoded PQ luminanceSamples unavailable

Dots: fixed sample grid · Diamond: brightest pixel. Near-black dots omitted; colors use an SDR preview mapping.

A Beyond8Bits source and its BT.709 conversion, using fixed Reinhard tone mapping, luminance-preserving gamut desaturation, and a 100-nit BT.1886 reference.[1] The plots measure decoded signal values; the videos use native playback and depend on your display.
Where the light goes 0.0 s
Measured pixel pairs and luminance distributions from the same six-second river crop.
The same sampled pixels, before and after conversion. Left: each dot pairs HDR luminance with its decoded SDR value; the line is the intended tone curve. The two axes have different ranges. Right: each curve gives the share of samples at or below a luminance, on a common logarithmic axis. Actual decoded samples replay at 5 fps; these are signal measurements, not screen measurements. Measurement details.

Dynamic range compares a high luminance with a low one. In stops, a useful way to write it is

Dstops = log2(Lhigh / Llow)

Each extra stop doubles the ratio. The low end needs a clear definition, such as a measured display black level or a usable shadow level. Choosing zero luminance for the denominator makes the ratio undefined. The luminance range of a scene, the range represented in a file, and the range emitted by a display are different measurements.

Color gamut and precision matter alongside luminance. BT.2020’s primaries enclose a wider chromaticity region than BT.709’s; the opening animation illustrates that difference. Ten-bit coding gives 1,024 possible codes per channel, compared with 256 for eight bits, before accounting for the permitted signal range. Neither more codes nor wider primaries alone guarantees a good HDR picture.[2], [3], [4]∗

∗A chromaticity triangle has no brightness axis. The animation’s particle colors are schematic; an ordinary screen cannot reproduce every color inside the BT.2020 boundary.

HDR at scale

As a phone moves past sunlit leaves and shaded plants, fine edges, saturated flowers, and dark gaps must survive capture, compression, and display. Improving the picture starts with finding where that path changed it.

This is already a large practical problem. In July 2023, Meta reported millions of HDR uploads each day to Facebook and Instagram. In March 2025, Netflix reported more than 300% growth in HDR streaming over five years, more than twice as many HDR-configured devices watching Netflix, and over 11,000 hours of HDR titles.[5], [6]

Five-year HDR growth at Netflix Relative to each measure's own starting level of one, HDR streaming grew to more than four times its earlier level, and HDR-configured devices watching Netflix to more than twice. Open endpoints show reported lower bounds, not exact values. The comparison ends in March 2025. Five-year HDR growth at NetflixReported March 2025 · each measure normalized to its own starting level HDR streamingHDR viewing devices >4×>2× 1×2×3×4×5×Multiple of the starting level Open circles mark lower bounds; arrows mean the actual values exceed them.
Netflix reported over 300% growth in HDR streaming and more than a doubling of HDR-configured devices watching Netflix over the five years preceding March 2025.[6] These are platform-specific comparisons, not annual observations or global adoption rates.

YouTube introduced HDR uploads and playback in 2016, then HDR live streaming in 2020. It supplies HDR to compatible devices and SDR versions to others.[7], [8], [9] Phones added both ends of the pipeline: iPhone X supported HDR playback in 2017, iPhone 12 added Dolby Vision recording in 2020, and Pixel 7 added 10-bit HDR recording in 2022.[10], [11], [12] These are capability milestones, not a count of active HDR viewers. Public evidence does not give us a consistent annual series of worldwide HDR uploads or phones.

Wider adoption requires preserving HDR through capture and compression, converting SDR footage, and judging results on suitable screens.

Reading the signal

Start with the signal. PQ maps code values to absolute display luminance, with a defined range reaching 10,000 cd/m². That is the signal’s ceiling, not the brightness of every HDR screen. HLG represents relative scene light and uses a display mapping that depends on viewing conditions. Bit depth describes precision; color primaries describe gamut. A 10-bit file or a BT.2020 tag alone cannot establish that the intended HDR picture reaches the viewer.[2]

The container, codec, and color description answer separate questions. MP4 packages tracks and timing. HEVC specifies a compression format, with profiles that support different precision and sampling choices. Transfer characteristics, color primaries, matrix coefficients, and signal range tell the decoder how to interpret the values. An .mp4 extension tells us none of those color properties by itself.

PackagingContainer, tracks, timestamps
CompressionCodec, profile, precision, chroma sampling
Color meaningTransfer function, primaries, matrix, range
PresentationDisplay mapping, HDR metadata, screen limits

A player can decode the codec yet apply an unintended color conversion. Incorrect tags can send PQ samples through the wrong transfer function. Resizing, trimming, and transcoding can change metadata or precision, so check each output.

PQ decoding curveThe ST 2084 inverse transfer curve maps a normalized code value of 0.5 to about 92 cd/m², 0.75 to 983 cd/m², and 1 to 10,000 cd/m². The vertical luminance axis is logarithmic. This is the mathematical signal definition, not measurements from a video or display.A code value is not a brightness percentagePQ signal decoded to absolute luminance0.010.111010010001000000.250.50.751Normalized PQ code valuecd/m² · logarithmic axis0.50 → 92 cd/m²0.75 → 983 cd/m²1.00 → 10,000 cd/m²
PQ decoding for a neutral pixel with equal RGB components, evaluated from the BT.2100 constants. A code value of 0.75 represents about 983 cd/m², not 75% of a screen’s peak brightness; actual screens apply their own output limits.[2]

A common failure happens before any model runs. Decoding into ordinary 8-bit RGB can discard precision and change the brightness representation. For YCbCr input, reconstruct nonlinear RGB before applying the transfer function to each channel; applying PQ directly to luma Y′ is not the same calculation. Treating PQ values as linear light gives arithmetic a different physical meaning. Normalizing every frame independently can hide exposure changes that a viewer would notice. Preserve the original signal, record transfer function and color metadata, and test the decoding path with known inputs.

What reaches the viewer

Peak brightness, black level, tone mapping, screen size, and room light affect visible detail. Even eligible HDR devices differ in brightness settings and display behavior. A controlled lab can measure these conditions; a browser study has less control.

The browser’s dynamic-range: high query does not establish that HDR mode is currently active. Codec and transfer-function support also do not measure the light leaving the screen.[13] Asking participants to disable automatic brightness and follow viewing instructions can reduce variation; it cannot calibrate thousands of displays remotely.†

†Qualification asks whether a setup is suitable. Calibration measures and adjusts its output against a target. A successful browser check supplies only part of that evidence.

We should measure these effects rather than assume their size. LIVE-HDR compared 5 and 200 lux and found no statistically significant difference in its resolution-and-bitrate group comparisons.[14] For reference-based evaluation, ColorVideoVDP explicitly includes display and viewing parameters. PU21 provides a perceptual encoding for adapting familiar metrics to HDR display light.[15], [16] The useful question is which viewing conditions a result covers.

Comparison figures have the same problem. An SDR export needs a stated tone mapper, target luminance, gamut mapping, and output transfer function. Brightness filters or an unexamined browser canvas conversion do not establish a matched SDR master. Keep the source frames and crop coordinates aligned so a visible difference comes from the conversion being studied.

A useful color analysis separates chromaticity from light. After decoding the transfer function into linear light and converting the appropriate RGB primaries to XYZ, chromaticity is

x = X / (X + Y + Z),   y = Y / (X + Y + Z)

Plotting x, y, and luminance Y gives a three-dimensional view of a patch’s colors and brightness.[17] For SDR, state the reference white and black luminance assumptions so both plots use comparable cd/m² units. Distance in xyY is not a perceptual quality score. For a moving patch, use the same frame times, pixel locations, sampling rule, and axes for HDR and SDR. Report the fraction outside BT.709 separately from luminance percentiles. Black pixels need special handling because their chromaticity is undefined. These plots should come from the decoded signals, before the page’s own display conversion.

The river clip at the beginning is a native Beyond8Bits recording. Its bright shoreline, autumn foliage and water reflections share the frame. I converted the first six seconds to BT.709 SDR with a fixed luminance curve, LSDR = 100L / (L + 100), then reduced saturation where needed to fit the SDR gamut. The curve compresses highlight contrast while preserving its ordering; it is one rendering choice, not a universal SDR appearance.

The opening comparison’s marked crop includes the bright sky, shoreline and reflections. In its first frame, median luminance falls from 161 to 62 cd/m²; the 95th percentile falls from 621 to 86 cd/m², a roughly sevenfold reduction. About 72% of the HDR crop lies above 100 cd/m². The HDR point cloud stretches upward while SDR is confined near the floor. Here the larger volume comes mainly from luminance, not a dramatic shift in chromaticity. The diamond marks each frame’s brightest pixel.

The plots follow 384 fixed pixel locations in the marked patch at five samples per second. Both sides come from decoding the actual files, including the encoded SDR output. PQ supplies absolute reference luminance; the SDR calculation assumes a 100-nit white, zero black and the BT.1886 display response.[18] The same axes make the reduced luminance range visible without changing scales between panels. Dot colors use a separate SDR preview mapping, so they do not reproduce the native HDR light.

Building Beyond8Bits

HDR datasets already existed. The gap was a large collection of ordinary user-generated recordings, their processing variants, and human judgments collected through an HDR-capable viewing path. Earlier controlled studies supplied useful evidence, but did not cover the combination of capture defects, content diversity, and scale we needed.[14], [19], [20] Building that collection became part of the research.

A source can already contain blur, noise, camera shake, or a difficult exposure. “Reference” identifies the source of our transcodes without claiming it is pristine. Measuring compression damage and predicting overall perceived quality therefore require different judgments.

Our collections grew from BrightVQ, introduced with BrightRate, and CHUG into Beyond8Bits. BrightVQ has 300 sources and 2,100 clips; CHUG has 856 sources and 5,992 clips. Beyond8Bits incorporates and extends that work, so these totals should not be added as independent datasets.[21], [19], [22]

Beyond8Bits public release, checked October 2026
CollectionCount
Source/reference clips5,917
Transcodes, six per source35,502
Total clips41,419
Human ratings, release documentationAbout 1.46 million

The sources include 2,153 crowd contributions and 3,764 Vimeo recordings. Collection involved checking HDR signaling, removing duplicates and unsuitable content, and making short clips while preserving PQ, 10-bit HEVC, and BT.2020.[20], [1] Those checks catch some malformed or incorrectly signaled inputs. They cannot prove that every camera used HDR well. Coverage also needs inspection. Portrait and landscape footage, dark scenes, bright lights, motion, and capture devices can be unevenly represented.

Near-duplicate scenes can inflate apparent diversity. Short clips can lose transitions that expose an artifact, while aggressive filtering can remove difficult recordings a quality model needs. Scene coverage, clip boundaries, and eligibility all affect the resulting task.

Redistribution is a separate constraint. Publicly viewable footage is not automatically available for a dataset release. The Beyond8Bits paper reports 6,861 sources, approximately 44,276 clips, and over 1.5 million ratings. The public release is smaller while license clearing continues.[22], [20] The plots here use the released CSV; the paper’s model results use its reported experimental collection.

Scenes across Beyond8Bits

Eleven native HDR recordings load and loop together.

Eleven other Beyond8Bits recordings: a suspension bridge, fireworks, open roads, snow, an aerial view, concert lighting a wet street at night, an arena, an outdoor performance, fire and evening architecture.[1] Each tile preserves the original landscape or portrait frame and opens its full 10-bit PQ/BT.2020 recording. HDR output depends on your viewing setup. Source metadata.
Beyond8Bits source quality distributions and conditional quantilesExact empirical cumulative distributions for 2153 Crowd and 3764 Vimeo reference source clips. 27.7 percent of Crowd and 16.1 percent of Vimeo scores are at most 60. The lower panel shows 10th, 25th, 50th, 75th and 90th percentiles within four origin and orientation groups, containing 1725 Crowd portrait, 428 Crowd landscape, 482 Vimeo portrait and 3282 Vimeo landscape clips. Every source appears once; transcodes are excluded. Quantile intervals describe clip variation, not confidence intervals.How source quality is distributedCumulative share of source clips at or below each score.Crowd · 2,153 sourcesVimeo · 3,764 sources0%25%50%75%100%020406080100Source video MOSAt MOS 60 or below: 27.7% of Crowd, 16.1% of Vimeo.The same origins, split by orientationThin line 10th–90th percentile · thick line middle 50% · dot median020406080100MedianCrowd · portrait1,725 source clips64.8Crowd · landscape428 source clips62.2Vimeo · portrait482 source clips69.1Vimeo · landscape3,282 source clips67.7Source video MOS
Exact cumulative source-score distributions and origin-by-orientation quantiles from 5,917 reference clips, each counted once. Intervals show variation across source clips; capture origin and orientation are not randomized, so differences do not establish a causal effect on quality.[1]

The source pools also differ in ways that a pooled score can hide. About 80.1% of crowd sources are portrait, compared with 12.8% of Vimeo sources. Their median source MOS values are 64.44 and 67.88. The cumulative curves show that 27.7% of crowd sources have MOS at or below 60, compared with 16.1% of Vimeo sources. Origin, orientation, scene choice, and capture pipelines vary together here; this comparison cannot tell us that one orientation causes better quality.[1]

Each released source has six encoding conditions: 360p at 0.2 Mbps, 720p at 0.5 and 2 Mbps, and 1080p at 0.5, 1, and 3 Mbps. Keeping their source identities lets us compare the same content at different settings. It also prevents a common evaluation leak: training on one encode of a scene and testing on another.‡

‡Seven released clips share each source: its reference and six transcodes. Count the 5,917 source families when measuring content diversity, and keep each family in one split.
Small lights against a dark background
Source / reference
Same region, enlarged
360p · 0.2 Mbps target
Video preview loads with the clip
Same region, enlarged
Video preview loads with the clip

Source preview. The paired clips loop when playback is supported.

The Christmas tree recording and its 360p, 0.2 Mbps transcode, with matched detail views. Both files retain HDR; watch the lights and dark background for compression artifacts.[23] Source · Transcode.

An HDR study on MTurk

MTurk reaches the larger participant pool needed to rate thousands of recordings. Device eligibility, playback, instructions, and quality control become part of the measurement system. Our CHUG paper introduced, to our knowledge, the first large-scale Amazon Mechanical Turk study of user-generated HDR video quality. The claim concerns that setting; HDR subjective studies already existed.[19]

211,848retained ratings
700+participants
35.4ratings per clip on average

Those CHUG totals cover 5,992 clips from 856 sources. Ratings per clip vary, and these counts exclude recruiting effort and playback troubleshooting.[19]

Finding eligible viewers

Participants need a suitable display, operating-system configuration, browser, decoder, and connection. Excluding unsupported setups reduces the recruiting pool and changes which viewers the data represents. Report these requirements and exclusions; qualified participants do not form a random sample of everyone who watches video.

CHUG screened device/browser capability, bit depth, HEVC support, resolution, and network conditions, and monitored changes in HDR settings.[19] These safeguards reduce failures without establishing uniform luminance or calibration. Recording device families and settings would let us test differences across groups.

Checking playback

A low score could describe compression, buffering, a decode failure, or an unintended SDR fallback. Those causes require different responses. We preloaded clips, tracked playback completion, and allowed participants to report technical problems. CHUG used progressive checks at 25%, 50%, and 75% of a session.[19] I would pilot playback on each supported device class and retain its logs alongside ratings; a completed download cannot confirm correct presentation.

Training and screening participants

CHUG sessions included six training videos followed by 94 test presentations. Ten test presentations were controls: five repeated clips and five golden-set clips. Training helps participants learn what the quality scale means; repeats test consistency, and golden items provide a known comparison. Screening also used technical-issue flags and observer filtering.[19] Design these checks before collection and test thresholds in a pilot.

People still use scales differently, even when they watch carefully. Some avoid the ends; others spread scores widely. We used SUREAL to estimate quality while accounting for observer bias and inconsistency.[19] Genuine disagreement can remain, especially for difficult scenes. For a new study, retaining individual ratings lets us inspect disagreement around the estimated quality.§

§Variation across individual judgments and uncertainty in an estimated mean are different quantities. A widely disputed clip can still have a precisely estimated average.

Reproducing the study requires versioned manifests, source identities, exclusions, and admitted viewing configurations. Models trained on these ratings still need tests on new devices and content.

What the ratings show

The released data supports comparisons that a single average can hide. At the same target 0.5 Mbps, 720p has a mean opinion score of 50.30 and 1080p has 46.89. More pixels compete for the same bit budget. Pairing the encodes by source shows that 720p scores higher in 5,033 of 5,917 cases, or 85.1%. The remaining pairs matter too; this is a pattern in these recordings and settings, not a rule for every video.[23]

Beyond8Bits scores overlap across encoding conditionsSeven groups of 5917 clips each. Marks show empirical 10th, 25th, 50th, 75th and 90th percentiles of per-clip opinion scores. These are distributions across clips, not confidence intervals. The two 0.5 Mbps groups are highlighted.The spread behind each meanThin line 10th–90th percentile · thick line middle 50% · dot median020406080100360p · 0.2 Mbps720p · 0.5 Mbps1080p · 0.5 Mbps1080p · 1 Mbps720p · 2 Mbps1080p · 3 MbpsSource / referencePer-clip opinion score
Same-source encoding scores at target 0.5 MbpsAll 5917 paired source clips. Horizontal axis is the 1080p opinion score and vertical axis is the 720p score, with identical scales from 0 to 80. The dashed diagonal marks equal scores. In 5033 pairs, the 720p score is higher. Marks show measured clip scores, not model predictions; this descriptive count is not a significance test.Same source, same 0.5 Mbps targetEach point compares the two encodes of one source clip.002020404060608080720p score1080p scoreAbove line720p scoredhigherBelow line1080p scoredhigher720p scores higher in 5,033 of 5,917 pairs (85.1%).
Public Beyond8Bits scores, with 5,917 clips per condition; the upper quantiles show variation across clips, and the lower plot uses source identities from the expanded manifest to pair two encodes at the same 0.5 Mbps target. These are descriptive score comparisons, without uncertainty intervals or a significance claim.[1][23] Mean values (CSV).

The rate curves give another view of the same experiment. At 1080p, mean MOS rises from 46.89 at 0.5 Mbps to 61.54 at 1 Mbps, then to 64.06 at 3 Mbps. The gains differ across those intervals. The source mean is 65.38, which also reminds us that an encoder cannot repair every capture defect by spending more bits.[1]

Release quality at six resolution and target bitrate settingsThe 1080p mean MOS rises from 46.89 at 0.5 Mbps to 61.54 at 1 Mbps and 64.06 at 3 Mbps. 720p rises from 50.30 at 0.5 Mbps to 61.04 at 2 Mbps. 360p at 0.2 Mbps scores 37.34. Source mean is 65.38. Lines connect measured condition means and do not predict untested rates.What a larger bit budget buysMean video MOS · the same 5,917 sources per condition30405060700.20.5123Target bitrate, Mbps · logarithmic axisSource mean 65.3837.34360p50.3061.04720p46.8961.5464.061080p
Release means at the six configured encoding settings, with lines connecting settings at the same resolution. These are target rates, not measured delivered bitrates; the connecting lines do not predict untested settings.[1]
Paired quality differences at the same target bitrateEach dot is one source. Horizontal coordinate is the mean of its two scores, vertical coordinate is 720p minus 1080p. Zero means equal scores. The gold line connects descriptive within-bin medians, not a fitted predictor or uncertainty interval.Where 720p gains at the same bitrateAll 5,917 matched sources at target 0.5 Mbps-30-20-100102030020406080Positive: 720p scores higherMean of the two clip scoresScore difference · 720p − 1080pGold line: median difference within 10-point mean-score bins.
Discrete rate quality Pareto frontierFour tested settings are non-dominated: 360p at .2 Mbps, 720p at .5, 1080p at 1, and 1080p at 3. The two remaining settings are dominated in mean quality versus configured target bitrate. Connecting lines do not estimate unmeasured settings.The observed rate–quality frontierLower target rate and higher mean quality are preferred.30405060700.20.5123360p720p1080p1080p720p1080pTarget bitrate, Mbps · logarithmic scaleFilled: not dominated among these six measured settings.Open: another setting has no higher target rate and higher mean MOS.Mean opinion score
Paired differences show how content changes the resolution tradeoff; bin medians describe these clips, not statistical significance.[23] The frontier uses the six measured condition means and configured target rates. It excludes the source reference, which has no common rate target; dashed connections do not predict untested settings.[1]

This supports evaluating resolution and rate jointly, using matched content. A model should follow changes across encodes and recognize low-quality sources. Whole-collection correlation can hide failures on difficult subsets.

For reproducible experiments, identify the manifest and split explicitly. The public Beyond8Bits split is 70/10/20 for training, validation, and testing; the paper describes 70/20/10. The expanded manifest supplies source-family identities for grouping.[22], [20], [23] Hold out complete source families, check overlap with earlier collections, and report results across scene types and viewing conditions. The small release CSV contains quality scores and metadata, without human-written defect explanations or calibrated display measurements.

Predicting and explaining quality

Beyond8Bits also supports learning a quality predictor. In our HDR-Q work, an HDR-adapted encoder processes native PQ information, and HDR-Aware Policy Optimization trains the model to use that evidence when predicting quality and producing an explanation. The paper reports rank correlation of 0.9206 for the full model versus 0.8914 for its SDR variant.[20]

Two HDR-Q correlation metricsPaper Table 1 reports SDR variant SRCC 0.8914 and full-model SRCC 0.9206. PLCC increases from 0.8895 to 0.9118. The plotted axis runs from 0.85 to 0.95. No per-video prediction scatter or uncertainty interval is available here.HDR-Q, SDR variant and full modelPublished correlations on Beyond8Bits; higher is better.0.850.8750.90.9250.95Rank (SRCC)0.89140.9206Linear (PLCC)0.88950.9118SDR variantFull modelAxis shown from 0.85 to 0.95; correlations are not accuracy percentages.
HDR-Q results reported in Table 1 of the paper, with the same two variants shown for rank and linear correlation. These aggregate values are not new measurements on the public release and do not supply confidence intervals.[20]

That comparison tests score prediction. A score can agree with viewers while the accompanying explanation names the wrong defect.[20], [24] Keep the exact video, the model’s unedited output, its predicted score, and the corresponding human score together, then locate each claimed defect in the footage. A paired SDR/HDR comparison should also use the same checkpoint and sampling procedure unless the difference being tested is explicitly model adaptation.

A botanical scene with mist, from the HDR-Q presentation
ModelPredicted quality, 0–100Published explanation, paraphrased
Ovis 2.585Describes uneven exposure, dull plant colors, and color spreading around the purple flowers.
HDR-Q with HAPO82Describes gradual highlight transitions in the mist and stable hues. Attributes softer detail to the mist and reports no obvious motion artifacts.
Published qualitative excerpts from slide 14 of the HDR-Q presentation, paraphrased here.[25] The slide supplies these two model predictions but no human MOS or matched SDR/HDR inference pair.

HDR-Q samples eight frames. A brief flash, exposure jump, or dropped frame can occur between them. That makes temporal stress tests useful even when overall correlations are high. Evaluating separate capture defects, compression changes, highlights, shadows, and device groups would tell us more about where a predictor can be trusted.

Reconstructing HDR from SDR

Existing SDR libraries make inverse tone mapping (ITM), or SDR-to-HDR conversion, useful. The difficulty is that the forward process can discard information. Several bright values may become the same clipped white. Colors outside the SDR gamut can collapse together. A creative grade, camera processing, and compression can all be part of the input. There is no single inverse that can undo every combination.

What disappears in a white sky
Hard-clipped SDR · 100-nit limit
Waterfall video loads here
The same clouds, enlarged
Matched sky detail loads here
Tone-mapped SDR · 100-nit limit
Waterfall video loads here
The same clouds, enlarged
Matched sky detail loads here

The matching six-second clips load and loop together.

Two SDR conversions of the same Beyond8Bits waterfall recording.[1] Left: luminance above 100 nits is deliberately clipped. Right: a smooth tone curve compresses it, preserving cloud variation. Both use the same gamut reduction and 100-nit reference display. This isolates clipping; it is not a comparison of all SDR and HDR content. Native HDR source · Conversion details.

In the waterfall example, bright clouds collapse into a flat white area after clipping. Increasing that white value cannot bring back their shape. A model could generate plausible clouds, but they need not match the recording. In video, those details must also remain consistent as the camera moves; a good still frame can hide flickering textures.

Bright signs against a dark room make the problem easier to see. In this Beyond8Bits clip, a fixed tone mapper compresses the lights and reduces colors that do not fit BT.709. The pair below shows the encoded SDR input beside the recorded HDR target, with matching crops around the signs.

Bright signs, deep shadows
SDR conversion · BT.709
Neon-lit video loads here
The same signs, enlarged
Matched signs load here
Recorded HDR · PQ / BT.2020
Neon-lit video loads here
The same signs, enlarged
Matched signs load here

The native HDR and converted SDR clips load and loop together.

A neon-lit Beyond8Bits recording and its encoded SDR conversion.[1] Both use the same frames; the SDR version uses the fixed luminance curve and gamut reduction described above. The right side is the recorded target, not a reconstructed prediction. Signal and conversion details.

This is one way to build aligned training pairs for SDR-to-HDR: start with the recorded HDR and generate an SDR input whose processing is known. Expanding luminance cannot, by itself, undo color desaturation, clipping, quantization or compression. Testing on pairs from the same conversion pipeline is easier than restoring arbitrary SDR footage with an unknown history.

This example shows an SDR input and its recorded HDR target, not a model prediction. A reconstruction should be compared with the target using both signal measurements and temporal checks.

Our LumaFlux work approaches this with an eight-step rectified-flow bridge from SDR to PQ/BT.2020 HDR. It uses luminance and image features to condition a frozen generative backbone, then a monotone tone-field decoder to map the result into display-referred light. The released video path shares noise across frames and smooths tone-curve parameters; it has no explicit long-range motion model. Evaluation therefore needs a known HDR target where available, color and luminance measurements, and temporal checks. A brighter output alone does not establish a faithful reconstruction.[26]

Training pairs bring their own bias. Synthesizing SDR from an HDR master gives an aligned target, but a model may learn the chosen tone mapper rather than handle real camera and grading pipelines. LumaFlux’s released data recipe combines expert grades where available with multiple tone mappers and codec settings.[27] Source-level holdouts and unseen conversion pipelines help test what actually transfers. HDR generation adds another question: a desired brightness distribution must fit the scene. Our LumaGuide work explores luminance-distribution guidance for generation; matching a histogram alone cannot tell us whether the highlights belong on the right objects.[28]

Using a shared dataset

A public collection lets another lab repeat a comparison without rebuilding the study. Processing variants and human scores support research on compression, representations, and perceived quality. Industry teams can investigate delivery settings, while academic work can test uncertainty and transfer across viewing conditions. A changed deployment setting may still require new data.

Start with an end-to-end pilot that checks decoding, playback, instructions, controls, and individual ratings. Hold out source scenes and conversion pipelines for SDR-to-HDR; hold out source families for quality assessment. Choose metrics according to whether the experiment measures fidelity, perceived quality, or plausible reconstruction.

Public repositories and what each provides
RepositoryUseful for
Beyond8BitsClips, scores, source identities, splits, and HDR-Q research materials. The public tree currently lacks HDR-Q training and inference code.
CHUG · BrightVQ / BrightRateEarlier datasets, study papers, and BrightRate model files. Check overlap before combining collections.
LumaFlux · ComfyUI nodesSDR-to-HDR inference, data preparation, and evaluation protocols.
LumaGuideLuminance guidance for HDR generation. Public FLUX implementation; other integrations have separate release status.
My P.910 forkMicrosoft’s crowd-testing toolkit, including task templates and rating analysis. A starting point, not a verified release of our custom HDR study backend.[29]
HDR-Shorts_Buffer_VideosAdditional clip files stored through Git LFS, not a study platform.

For the crowd toolkit, inspect the task configuration as well as the code. Available checks are not necessarily enabled, a visibility test is not HDR calibration, and technical failures need their own recorded outcome. Credit for the underlying toolkit belongs to Microsoft’s P.910 project. The study papers above document our published protocols.

References

  1. Beyond8Bits. Public metadata manifest, Git blob 4ff3304. CC BY 4.0.
  2. ITU-R. BT.2100-3, Image parameter values for high dynamic range television. February 2025.
  3. ITU-R. BT.709-6, HDTV image parameters and color primaries. June 2015.
  4. ITU-R. BT.2020-2, UHDTV image parameters and color primaries. October 2015.
  5. Meta Engineering. Bringing HDR video to Reels. July 2023.
  6. Netflix TechBlog. HDR10+ Now Streaming on Netflix. March 2025.
  7. YouTube. True colors: adding support for HDR videos. November 2016.
  8. YouTube. Seeing is believing: launching HDR live streams. December 2020.
  9. YouTube Help. Upload HDR videos.
  10. Apple. The future is here: iPhone X. September 2017.
  11. Apple. iPhone 12 and iPhone 12 mini announcement. October 2020.
  12. Google. Pixel 7 and Pixel 7 Pro announcement. October 2022.
  13. W3C. Media Queries Level 5, dynamic-range capability and HDR mode.
  14. Shang et al. Subjective Quality Assessment of High Dynamic Range Videos Under Different Ambient Conditions. ICIP 2022.
  15. Mantiuk et al. ColorVideoVDP: A visible difference predictor for colour images and videos. SIGGRAPH 2024.
  16. Mantiuk and Azimi. PU21: A novel perceptually uniform encoding for adapting existing quality metrics for HDR. 2021.
  17. CIE S 017:2020. Chromaticity coordinates.
  18. ITU-R. BT.1886, Reference electro-optical transfer function for HDTV production displays. March 2011.
  19. Saini et al. CHUG: Crowdsourced User-Generated HDR Video Quality Dataset. ICIP 2025.
  20. Saini et al. Seeing Beyond 8bits: Subjective and Objective Quality Assessment of HDR-UGC Videos. 2026.
  21. Saini et al. BrightRate: Quality Assessment for User-Generated HDR Videos. WACV 2026.
  22. Beyond8Bits. Public release documentation and licensing. Checked October 2026.
  23. Beyond8Bits. Expanded release manifest with source-family identities. Git blob 75afce1.
  24. BrightRate-LM. Multi-exposure inference, scoring, and explanation generation.
  25. Saini. HDR-Q presentation, slide 14: qualitative model outputs.
  26. Saini et al. LumaFlux. Public implementation and current evaluation protocol. Checked October 2026.
  27. LumaFlux. Data preparation, source collections, and pairing.
  28. Saini et al. LumaGuide: Guiding Image and Video Generation into the High Dynamic Range. 2026.
  29. Microsoft. P.910 subjective video quality crowd-testing toolkit.

Citation

@misc{saini2025hdrquality,
  author = {Shreshth Saini},
  title = {Challenges of HDR},
  year = {2025},
  url = {https://shreshthsaini.github.io/blogs/hdr-video-quality-assessment.html}
}
@article{saini2026seeing,
  title = {Seeing Beyond 8bits: Subjective and Objective Quality Assessment of HDR-UGC Videos},
  author = {Saini, Shreshth and Chen, Bowen and Birkbeck, Neil and Wang, Yilin and Adsumilli, Balu and Bovik, Alan C.},
  journal = {arXiv preprint arXiv:2603.00938},
  year = {2026}
}
@inproceedings{saini2026brightrate,
  author = {Saini, Shreshth and Chen, Bowen and Wang, Yilin and Birkbeck, Neil and Adsumilli, Balu and Bovik, Alan C.},
  title = {BrightRate: Quality Assessment for User-Generated HDR Videos},
  booktitle = {Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)},
  year = {2026},
  pages = {1522--1532}
}
@inproceedings{saini2025chug,
  author = {Saini, Shreshth and Bovik, Alan C. and Birkbeck, Neil and Wang, Yilin and Adsumilli, Balu},
  title = {CHUG: Crowdsourced User-Generated HDR Video Quality Dataset},
  booktitle = {2025 IEEE International Conference on Image Processing (ICIP)},
  year = {2025},
  pages = {2504--2509},
  doi = {10.1109/ICIP55913.2025.11084488}
}