Challenges of HDR
From capture and SDR-to-HDR reconstruction to collecting human judgments at scale.
HDR gives us more room to represent light, from a dim room to a bright window in the same shot. Using that room well is harder. A recording, a reconstructed highlight, and a human quality score all depend on what happens between the scene and the screen. These are the problems behind our work on Beyond8Bits, HDR quality assessment, and SDR-to-HDR conversion.
What HDR changes
A recording of a shaded room and a sunlit window must fit both into the range its pipeline can capture, encode, and display. HDR widens that range, preserving differences among bright surfaces and detail in shadows. The result still depends on the recording and the viewer’s screen.
Start with this Beyond8Bits recording. The same sky, autumn trees and reflections appear in both views. The SDR version compresses the bright parts to fit a 100-nit reference display; the original keeps its native HDR signal. The matching crops and plots show what changes before your screen applies its own display mapping.
The paired clips load together and loop when playback is supported.
Dots: fixed sample grid · Diamond: brightest pixel. Near-black dots omitted; colors use an SDR preview mapping.
Dynamic range compares a high luminance with a low one. In stops, a useful way to write it is
Each extra stop doubles the ratio. The low end needs a clear definition, such as a measured display black level or a usable shadow level. Choosing zero luminance for the denominator makes the ratio undefined. The luminance range of a scene, the range represented in a file, and the range emitted by a display are different measurements.
Color gamut and precision matter alongside luminance. BT.2020’s primaries enclose a wider chromaticity region than BT.709’s; the opening animation illustrates that difference. Ten-bit coding gives 1,024 possible codes per channel, compared with 256 for eight bits, before accounting for the permitted signal range. Neither more codes nor wider primaries alone guarantees a good HDR picture.[2], [3], [4]∗
∗A chromaticity triangle has no brightness axis. The animation’s particle colors are schematic; an ordinary screen cannot reproduce every color inside the BT.2020 boundary.HDR at scale
As a phone moves past sunlit leaves and shaded plants, fine edges, saturated flowers, and dark gaps must survive capture, compression, and display. Improving the picture starts with finding where that path changed it.
This is already a large practical problem. In July 2023, Meta reported millions of HDR uploads each day to Facebook and Instagram. In March 2025, Netflix reported more than 300% growth in HDR streaming over five years, more than twice as many HDR-configured devices watching Netflix, and over 11,000 hours of HDR titles.[5], [6]
YouTube introduced HDR uploads and playback in 2016, then HDR live streaming in 2020. It supplies HDR to compatible devices and SDR versions to others.[7], [8], [9] Phones added both ends of the pipeline: iPhone X supported HDR playback in 2017, iPhone 12 added Dolby Vision recording in 2020, and Pixel 7 added 10-bit HDR recording in 2022.[10], [11], [12] These are capability milestones, not a count of active HDR viewers. Public evidence does not give us a consistent annual series of worldwide HDR uploads or phones.
Wider adoption requires preserving HDR through capture and compression, converting SDR footage, and judging results on suitable screens.
Reading the signal
Start with the signal. PQ maps code values to absolute display luminance, with a defined range reaching 10,000 cd/m². That is the signal’s ceiling, not the brightness of every HDR screen. HLG represents relative scene light and uses a display mapping that depends on viewing conditions. Bit depth describes precision; color primaries describe gamut. A 10-bit file or a BT.2020 tag alone cannot establish that the intended HDR picture reaches the viewer.[2]
The container, codec, and color description answer separate questions. MP4 packages tracks and timing. HEVC specifies a compression format, with profiles that support different precision and sampling choices. Transfer characteristics, color primaries, matrix coefficients, and signal range tell the decoder how to interpret the values. An .mp4 extension tells us none of those color properties by itself.
A player can decode the codec yet apply an unintended color conversion. Incorrect tags can send PQ samples through the wrong transfer function. Resizing, trimming, and transcoding can change metadata or precision, so check each output.
A common failure happens before any model runs. Decoding into ordinary 8-bit RGB can discard precision and change the brightness representation. For YCbCr input, reconstruct nonlinear RGB before applying the transfer function to each channel; applying PQ directly to luma Y′ is not the same calculation. Treating PQ values as linear light gives arithmetic a different physical meaning. Normalizing every frame independently can hide exposure changes that a viewer would notice. Preserve the original signal, record transfer function and color metadata, and test the decoding path with known inputs.
What reaches the viewer
Peak brightness, black level, tone mapping, screen size, and room light affect visible detail. Even eligible HDR devices differ in brightness settings and display behavior. A controlled lab can measure these conditions; a browser study has less control.
The browser’s dynamic-range: high query does not establish that HDR mode is currently active. Codec and transfer-function support also do not measure the light leaving the screen.[13] Asking participants to disable automatic brightness and follow viewing instructions can reduce variation; it cannot calibrate thousands of displays remotely.†
We should measure these effects rather than assume their size. LIVE-HDR compared 5 and 200 lux and found no statistically significant difference in its resolution-and-bitrate group comparisons.[14] For reference-based evaluation, ColorVideoVDP explicitly includes display and viewing parameters. PU21 provides a perceptual encoding for adapting familiar metrics to HDR display light.[15], [16] The useful question is which viewing conditions a result covers.
Comparison figures have the same problem. An SDR export needs a stated tone mapper, target luminance, gamut mapping, and output transfer function. Brightness filters or an unexamined browser canvas conversion do not establish a matched SDR master. Keep the source frames and crop coordinates aligned so a visible difference comes from the conversion being studied.
A useful color analysis separates chromaticity from light. After decoding the transfer function into linear light and converting the appropriate RGB primaries to XYZ, chromaticity is
Plotting x, y, and luminance Y gives a three-dimensional view of a patch’s colors and brightness.[17] For SDR, state the reference white and black luminance assumptions so both plots use comparable cd/m² units. Distance in xyY is not a perceptual quality score. For a moving patch, use the same frame times, pixel locations, sampling rule, and axes for HDR and SDR. Report the fraction outside BT.709 separately from luminance percentiles. Black pixels need special handling because their chromaticity is undefined. These plots should come from the decoded signals, before the page’s own display conversion.
The river clip at the beginning is a native Beyond8Bits recording. Its bright shoreline, autumn foliage and water reflections share the frame. I converted the first six seconds to BT.709 SDR with a fixed luminance curve, LSDR = 100L / (L + 100), then reduced saturation where needed to fit the SDR gamut. The curve compresses highlight contrast while preserving its ordering; it is one rendering choice, not a universal SDR appearance.
The opening comparison’s marked crop includes the bright sky, shoreline and reflections. In its first frame, median luminance falls from 161 to 62 cd/m²; the 95th percentile falls from 621 to 86 cd/m², a roughly sevenfold reduction. About 72% of the HDR crop lies above 100 cd/m². The HDR point cloud stretches upward while SDR is confined near the floor. Here the larger volume comes mainly from luminance, not a dramatic shift in chromaticity. The diamond marks each frame’s brightest pixel.
The plots follow 384 fixed pixel locations in the marked patch at five samples per second. Both sides come from decoding the actual files, including the encoded SDR output. PQ supplies absolute reference luminance; the SDR calculation assumes a 100-nit white, zero black and the BT.1886 display response.[18] The same axes make the reduced luminance range visible without changing scales between panels. Dot colors use a separate SDR preview mapping, so they do not reproduce the native HDR light.
Building Beyond8Bits
HDR datasets already existed. The gap was a large collection of ordinary user-generated recordings, their processing variants, and human judgments collected through an HDR-capable viewing path. Earlier controlled studies supplied useful evidence, but did not cover the combination of capture defects, content diversity, and scale we needed.[14], [19], [20] Building that collection became part of the research.
A source can already contain blur, noise, camera shake, or a difficult exposure. “Reference” identifies the source of our transcodes without claiming it is pristine. Measuring compression damage and predicting overall perceived quality therefore require different judgments.
Our collections grew from BrightVQ, introduced with BrightRate, and CHUG into Beyond8Bits. BrightVQ has 300 sources and 2,100 clips; CHUG has 856 sources and 5,992 clips. Beyond8Bits incorporates and extends that work, so these totals should not be added as independent datasets.[21], [19], [22]
| Collection | Count |
|---|---|
| Source/reference clips | 5,917 |
| Transcodes, six per source | 35,502 |
| Total clips | 41,419 |
| Human ratings, release documentation | About 1.46 million |
The sources include 2,153 crowd contributions and 3,764 Vimeo recordings. Collection involved checking HDR signaling, removing duplicates and unsuitable content, and making short clips while preserving PQ, 10-bit HEVC, and BT.2020.[20], [1] Those checks catch some malformed or incorrectly signaled inputs. They cannot prove that every camera used HDR well. Coverage also needs inspection. Portrait and landscape footage, dark scenes, bright lights, motion, and capture devices can be unevenly represented.
Near-duplicate scenes can inflate apparent diversity. Short clips can lose transitions that expose an artifact, while aggressive filtering can remove difficult recordings a quality model needs. Scene coverage, clip boundaries, and eligibility all affect the resulting task.
Redistribution is a separate constraint. Publicly viewable footage is not automatically available for a dataset release. The Beyond8Bits paper reports 6,861 sources, approximately 44,276 clips, and over 1.5 million ratings. The public release is smaller while license clearing continues.[22], [20] The plots here use the released CSV; the paper’s model results use its reported experimental collection.
Eleven native HDR recordings load and loop together.
The source pools also differ in ways that a pooled score can hide. About 80.1% of crowd sources are portrait, compared with 12.8% of Vimeo sources. Their median source MOS values are 64.44 and 67.88. The cumulative curves show that 27.7% of crowd sources have MOS at or below 60, compared with 16.1% of Vimeo sources. Origin, orientation, scene choice, and capture pipelines vary together here; this comparison cannot tell us that one orientation causes better quality.[1]
Each released source has six encoding conditions: 360p at 0.2 Mbps, 720p at 0.5 and 2 Mbps, and 1080p at 0.5, 1, and 3 Mbps. Keeping their source identities lets us compare the same content at different settings. It also prevents a common evaluation leak: training on one encode of a scene and testing on another.‡
‡Seven released clips share each source: its reference and six transcodes. Count the 5,917 source families when measuring content diversity, and keep each family in one split.Source preview. The paired clips loop when playback is supported.
An HDR study on MTurk
MTurk reaches the larger participant pool needed to rate thousands of recordings. Device eligibility, playback, instructions, and quality control become part of the measurement system. Our CHUG paper introduced, to our knowledge, the first large-scale Amazon Mechanical Turk study of user-generated HDR video quality. The claim concerns that setting; HDR subjective studies already existed.[19]
Those CHUG totals cover 5,992 clips from 856 sources. Ratings per clip vary, and these counts exclude recruiting effort and playback troubleshooting.[19]
Finding eligible viewers
Participants need a suitable display, operating-system configuration, browser, decoder, and connection. Excluding unsupported setups reduces the recruiting pool and changes which viewers the data represents. Report these requirements and exclusions; qualified participants do not form a random sample of everyone who watches video.
CHUG screened device/browser capability, bit depth, HEVC support, resolution, and network conditions, and monitored changes in HDR settings.[19] These safeguards reduce failures without establishing uniform luminance or calibration. Recording device families and settings would let us test differences across groups.
Checking playback
A low score could describe compression, buffering, a decode failure, or an unintended SDR fallback. Those causes require different responses. We preloaded clips, tracked playback completion, and allowed participants to report technical problems. CHUG used progressive checks at 25%, 50%, and 75% of a session.[19] I would pilot playback on each supported device class and retain its logs alongside ratings; a completed download cannot confirm correct presentation.
Training and screening participants
CHUG sessions included six training videos followed by 94 test presentations. Ten test presentations were controls: five repeated clips and five golden-set clips. Training helps participants learn what the quality scale means; repeats test consistency, and golden items provide a known comparison. Screening also used technical-issue flags and observer filtering.[19] Design these checks before collection and test thresholds in a pilot.
People still use scales differently, even when they watch carefully. Some avoid the ends; others spread scores widely. We used SUREAL to estimate quality while accounting for observer bias and inconsistency.[19] Genuine disagreement can remain, especially for difficult scenes. For a new study, retaining individual ratings lets us inspect disagreement around the estimated quality.§
§Variation across individual judgments and uncertainty in an estimated mean are different quantities. A widely disputed clip can still have a precisely estimated average.Reproducing the study requires versioned manifests, source identities, exclusions, and admitted viewing configurations. Models trained on these ratings still need tests on new devices and content.
What the ratings show
The released data supports comparisons that a single average can hide. At the same target 0.5 Mbps, 720p has a mean opinion score of 50.30 and 1080p has 46.89. More pixels compete for the same bit budget. Pairing the encodes by source shows that 720p scores higher in 5,033 of 5,917 cases, or 85.1%. The remaining pairs matter too; this is a pattern in these recordings and settings, not a rule for every video.[23]
The rate curves give another view of the same experiment. At 1080p, mean MOS rises from 46.89 at 0.5 Mbps to 61.54 at 1 Mbps, then to 64.06 at 3 Mbps. The gains differ across those intervals. The source mean is 65.38, which also reminds us that an encoder cannot repair every capture defect by spending more bits.[1]
This supports evaluating resolution and rate jointly, using matched content. A model should follow changes across encodes and recognize low-quality sources. Whole-collection correlation can hide failures on difficult subsets.
For reproducible experiments, identify the manifest and split explicitly. The public Beyond8Bits split is 70/10/20 for training, validation, and testing; the paper describes 70/20/10. The expanded manifest supplies source-family identities for grouping.[22], [20], [23] Hold out complete source families, check overlap with earlier collections, and report results across scene types and viewing conditions. The small release CSV contains quality scores and metadata, without human-written defect explanations or calibrated display measurements.
Predicting and explaining quality
Beyond8Bits also supports learning a quality predictor. In our HDR-Q work, an HDR-adapted encoder processes native PQ information, and HDR-Aware Policy Optimization trains the model to use that evidence when predicting quality and producing an explanation. The paper reports rank correlation of 0.9206 for the full model versus 0.8914 for its SDR variant.[20]
That comparison tests score prediction. A score can agree with viewers while the accompanying explanation names the wrong defect.[20], [24] Keep the exact video, the model’s unedited output, its predicted score, and the corresponding human score together, then locate each claimed defect in the footage. A paired SDR/HDR comparison should also use the same checkpoint and sampling procedure unless the difference being tested is explicitly model adaptation.
| Model | Predicted quality, 0–100 | Published explanation, paraphrased |
|---|---|---|
| Ovis 2.5 | 85 | Describes uneven exposure, dull plant colors, and color spreading around the purple flowers. |
| HDR-Q with HAPO | 82 | Describes gradual highlight transitions in the mist and stable hues. Attributes softer detail to the mist and reports no obvious motion artifacts. |
HDR-Q samples eight frames. A brief flash, exposure jump, or dropped frame can occur between them. That makes temporal stress tests useful even when overall correlations are high. Evaluating separate capture defects, compression changes, highlights, shadows, and device groups would tell us more about where a predictor can be trusted.
Reconstructing HDR from SDR
Existing SDR libraries make inverse tone mapping (ITM), or SDR-to-HDR conversion, useful. The difficulty is that the forward process can discard information. Several bright values may become the same clipped white. Colors outside the SDR gamut can collapse together. A creative grade, camera processing, and compression can all be part of the input. There is no single inverse that can undo every combination.
The matching six-second clips load and loop together.
In the waterfall example, bright clouds collapse into a flat white area after clipping. Increasing that white value cannot bring back their shape. A model could generate plausible clouds, but they need not match the recording. In video, those details must also remain consistent as the camera moves; a good still frame can hide flickering textures.
Bright signs against a dark room make the problem easier to see. In this Beyond8Bits clip, a fixed tone mapper compresses the lights and reduces colors that do not fit BT.709. The pair below shows the encoded SDR input beside the recorded HDR target, with matching crops around the signs.
The native HDR and converted SDR clips load and loop together.
This is one way to build aligned training pairs for SDR-to-HDR: start with the recorded HDR and generate an SDR input whose processing is known. Expanding luminance cannot, by itself, undo color desaturation, clipping, quantization or compression.¶ Testing on pairs from the same conversion pipeline is easier than restoring arbitrary SDR footage with an unknown history.
¶This example shows an SDR input and its recorded HDR target, not a model prediction. A reconstruction should be compared with the target using both signal measurements and temporal checks.Our LumaFlux work approaches this with an eight-step rectified-flow bridge from SDR to PQ/BT.2020 HDR. It uses luminance and image features to condition a frozen generative backbone, then a monotone tone-field decoder to map the result into display-referred light. The released video path shares noise across frames and smooths tone-curve parameters; it has no explicit long-range motion model. Evaluation therefore needs a known HDR target where available, color and luminance measurements, and temporal checks. A brighter output alone does not establish a faithful reconstruction.[26]
Training pairs bring their own bias. Synthesizing SDR from an HDR master gives an aligned target, but a model may learn the chosen tone mapper rather than handle real camera and grading pipelines. LumaFlux’s released data recipe combines expert grades where available with multiple tone mappers and codec settings.[27] Source-level holdouts and unseen conversion pipelines help test what actually transfers. HDR generation adds another question: a desired brightness distribution must fit the scene. Our LumaGuide work explores luminance-distribution guidance for generation; matching a histogram alone cannot tell us whether the highlights belong on the right objects.[28]
Using a shared dataset
A public collection lets another lab repeat a comparison without rebuilding the study. Processing variants and human scores support research on compression, representations, and perceived quality. Industry teams can investigate delivery settings, while academic work can test uncertainty and transfer across viewing conditions. A changed deployment setting may still require new data.
Start with an end-to-end pilot that checks decoding, playback, instructions, controls, and individual ratings. Hold out source scenes and conversion pipelines for SDR-to-HDR; hold out source families for quality assessment. Choose metrics according to whether the experiment measures fidelity, perceived quality, or plausible reconstruction.
| Repository | Useful for |
|---|---|
| Beyond8Bits | Clips, scores, source identities, splits, and HDR-Q research materials. The public tree currently lacks HDR-Q training and inference code. |
| CHUG · BrightVQ / BrightRate | Earlier datasets, study papers, and BrightRate model files. Check overlap before combining collections. |
| LumaFlux · ComfyUI nodes | SDR-to-HDR inference, data preparation, and evaluation protocols. |
| LumaGuide | Luminance guidance for HDR generation. Public FLUX implementation; other integrations have separate release status. |
| My P.910 fork | Microsoft’s crowd-testing toolkit, including task templates and rating analysis. A starting point, not a verified release of our custom HDR study backend.[29] |
| HDR-Shorts_Buffer_Videos | Additional clip files stored through Git LFS, not a study platform. |
For the crowd toolkit, inspect the task configuration as well as the code. Available checks are not necessarily enabled, a visibility test is not HDR calibration, and technical failures need their own recorded outcome. Credit for the underlying toolkit belongs to Microsoft’s P.910 project. The study papers above document our published protocols.
References
- Beyond8Bits. Public metadata manifest, Git blob 4ff3304. CC BY 4.0.
- ITU-R. BT.2100-3, Image parameter values for high dynamic range television. February 2025.
- ITU-R. BT.709-6, HDTV image parameters and color primaries. June 2015.
- ITU-R. BT.2020-2, UHDTV image parameters and color primaries. October 2015.
- Meta Engineering. Bringing HDR video to Reels. July 2023.
- Netflix TechBlog. HDR10+ Now Streaming on Netflix. March 2025.
- YouTube. True colors: adding support for HDR videos. November 2016.
- YouTube. Seeing is believing: launching HDR live streams. December 2020.
- YouTube Help. Upload HDR videos.
- Apple. The future is here: iPhone X. September 2017.
- Apple. iPhone 12 and iPhone 12 mini announcement. October 2020.
- Google. Pixel 7 and Pixel 7 Pro announcement. October 2022.
- W3C. Media Queries Level 5, dynamic-range capability and HDR mode.
- Shang et al. Subjective Quality Assessment of High Dynamic Range Videos Under Different Ambient Conditions. ICIP 2022.
- Mantiuk et al. ColorVideoVDP: A visible difference predictor for colour images and videos. SIGGRAPH 2024.
- Mantiuk and Azimi. PU21: A novel perceptually uniform encoding for adapting existing quality metrics for HDR. 2021.
- CIE S 017:2020. Chromaticity coordinates.
- ITU-R. BT.1886, Reference electro-optical transfer function for HDTV production displays. March 2011.
- Saini et al. CHUG: Crowdsourced User-Generated HDR Video Quality Dataset. ICIP 2025.
- Saini et al. Seeing Beyond 8bits: Subjective and Objective Quality Assessment of HDR-UGC Videos. 2026.
- Saini et al. BrightRate: Quality Assessment for User-Generated HDR Videos. WACV 2026.
- Beyond8Bits. Public release documentation and licensing. Checked October 2026.
- Beyond8Bits. Expanded release manifest with source-family identities. Git blob 75afce1.
- BrightRate-LM. Multi-exposure inference, scoring, and explanation generation.
- Saini. HDR-Q presentation, slide 14: qualitative model outputs.
- Saini et al. LumaFlux. Public implementation and current evaluation protocol. Checked October 2026.
- LumaFlux. Data preparation, source collections, and pairing.
- Saini et al. LumaGuide: Guiding Image and Video Generation into the High Dynamic Range. 2026.
- Microsoft. P.910 subjective video quality crowd-testing toolkit.
Citation
@misc{saini2025hdrquality,
author = {Shreshth Saini},
title = {Challenges of HDR},
year = {2025},
url = {https://shreshthsaini.github.io/blogs/hdr-video-quality-assessment.html}
}@article{saini2026seeing,
title = {Seeing Beyond 8bits: Subjective and Objective Quality Assessment of HDR-UGC Videos},
author = {Saini, Shreshth and Chen, Bowen and Birkbeck, Neil and Wang, Yilin and Adsumilli, Balu and Bovik, Alan C.},
journal = {arXiv preprint arXiv:2603.00938},
year = {2026}
}@inproceedings{saini2026brightrate,
author = {Saini, Shreshth and Chen, Bowen and Wang, Yilin and Birkbeck, Neil and Adsumilli, Balu and Bovik, Alan C.},
title = {BrightRate: Quality Assessment for User-Generated HDR Videos},
booktitle = {Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)},
year = {2026},
pages = {1522--1532}
}@inproceedings{saini2025chug,
author = {Saini, Shreshth and Bovik, Alan C. and Birkbeck, Neil and Wang, Yilin and Adsumilli, Balu},
title = {CHUG: Crowdsourced User-Generated HDR Video Quality Dataset},
booktitle = {2025 IEEE International Conference on Image Processing (ICIP)},
year = {2025},
pages = {2504--2509},
doi = {10.1109/ICIP55913.2025.11084488}
}