SDR to HDR inverse tone mapping

LumaFlux: Lifting 8-Bit Worlds to HDR Reality with Physically-Guided Diffusion Transformers

Shreshth Saini1, Hakan Gedik1, Neil Birkbeck2, Yilin Wang2, Balu Adsumilli2, Alan C. Bovik1

1The University of Texas at Austin, 2Google, Inc.

LumaFlux teaser showing SDR inputs and corresponding HDR results
LumaFlux converts 8-bit SDR (BT.709) images and video into 10-bit HDR (PQ, BT.2020).

Abstract

LumaFlux converts 8-bit SDR content in BT.709 into 10-bit HDR in PQ and BT.2020. It adapts a frozen FLUX.1-dev MM-DiT using physical image cues and frozen perceptual features, while the backbone, VAE, and SigLIP encoder remain frozen. The resulting model is prompt-free and uses 71.3M trainable parameters, 0.57% of the backbone.

Physically-Guided Adaptation injects luminance, gradient, saturation, and spectral-band cues through gated low-rank attention residuals. Perceptual Cross-Modulation supplies FiLM conditioning from SigLIP, and the HDR Residual Coupler combines both paths under timestep-and-layer modulation. A monotone RQS tone-field decoder calibrates the frozen VAE output into display-referred HDR luminance. Inference follows an 8-step rectified-flow bridge with no text prompt or sampling-strength knob. For video, the same image model uses shared bridge noise and an EMA over applied RQS parameters, with no temporal training.

8solver steps
71.3Mtrainable parameters
0.57%of the backbone
Prompt-freeSDR to HDR bridge

Method

LumaFlux keeps the FLUX.1-dev MM-DiT and VAE frozen, and also keeps the SigLIP encoder frozen. It learns a compact physical and perceptual adaptation around that backbone, then calibrates the decoded result into HDR luminance.

Overview of the LumaFlux architecture, including physical and perceptual adaptation paths and the HDR decoder
LumaFlux combines physical cues, frozen SigLIP features, a frozen FLUX.1-dev backbone, and an RQS tone-field decoder.
PGA

Physically-Guided Adaptation

Gated low-rank attention residuals receive luminance, gradient, saturation, and spectral-band cues.

PCM

Perceptual Cross-Modulation

FiLM conditioning carries perceptual features from a frozen SigLIP encoder into the adaptation.

COUPLER

HDR Residual Coupler

The coupler fuses the physical and perceptual paths under timestep-and-layer modulation.

RQS

RQS tone-field decoder

A monotone rational-quadratic spline calibrates the frozen VAE decode into display-referred HDR luminance.

Comparison of SDR to HDR adaptation paradigms and the LumaFlux physically-guided design
The LumaFlux paradigm adapts a frozen generative backbone through physically-guided and perceptual paths.

Prompt-free bridge. Inference starts at z1 = VAEenc(SDR) + 0.05ยทฮต and integrates from t = 1 to t = 0 in K = 8 steps. There is no text prompt and no sampling-strength knob.

Training-free video stabilization. Video uses the same image model with one shared bridge-noise realization across frames and an EMA on the applied RQS parameters. It adds no temporal layers, optical flow, video fine-tuning, or temporal loss.

Results

Luma-Eval covers synthetic tone-mapping and codec degradations, expert-graded SDR from LIVE-TMHDR, and native paired SDR/HDR. Every method is scored on identical inputs, output encoding, and metric implementations.

Luma-Eval

Run-by-us synthetic-track protocol on 100 held-out frames. Higher is better for PU21-PSNR, PU21-SSIM, and HDR-VDP-3. Lower is better for dE-ITP, HDR-LPIPS, and FR-HIDRO.

Luma-Eval comparison
Method PU21-PSNR โ†‘ PU21-SSIM โ†‘ dE-ITP โ†“ HDR-LPIPS โ†“ HDR-VDP-3 โ†‘ FR-HIDRO โ†“
BT.2446c inverse23.740.829161.600.3075.6110.753
Reinhard inverse23.620.828661.760.2955.7260.753
HDRTVNet++23.030.819266.670.3285.4170.703
FMNet22.860.817667.920.3305.4120.724
KUNet22.270.811170.400.3255.4450.693
ITM-LUT22.920.814267.410.3155.4960.761
VAE+RQS (no-diffusion control)18.540.7647111.670.3813.6730.924
LumaFlux24.230.829456.540.2695.8120.631

LumaFlux is best in every column. Its margin over the strongest baseline is +0.49 dB in PU21-PSNR and -5.06 in dE-ITP. FR-HIDROVQA falls from 0.703 for HDRTVNet++ to 0.631.

HDRTV1K

Results on 117 published test pairs using the authors' protocol. The in-domain checkpoint and zero-shot main checkpoint are distinct.

In-domain versus zero-shot HDRTV1K evaluation
Model PSNR SSIM SR-SIM dE-ITP HDR-VDP-3
LumaFlux in-domain (lumaflux-hdrtv1k)33.340.94270.994114.697.962
LumaFlux zero-shot (lumaflux-main)27.520.93070.983729.867.618

The 33.34 dB result comes from the HDRTV1K-trained checkpoint. The mixed-corpus main model was never trained on HDRTV1K, so its 27.52 dB result is zero-shot. Published baseline numbers on HDRTV1K retain their authors' own protocols and are not directly rankable against these results.

Generative-class comparison

Like-for-like native-track comparison on 117 pairs.

Generative inverse tone-mapping methods
Method Steps PU21-PSNR dE-ITP s/megapixel Trainable params
LEDiff5010.99198.5817.0860 M
X2HDR3018.6592.2016.2149 M
LumaFlux823.5647.592.571 M

Against the closest diffusion competitor, LumaFlux gains +4.9 dB, costs about 6x less per output pixel, uses 2x fewer trainable parameters, and requires fewer steps.

53.3%

Less signed excess flicker

Shared noise and RQS-parameter EMA reduce signed excess PU21 flicker by 53.3%. Pooled raw temporal energy remains within 3.3% of the reference, with zero temporal training.

0.55 dB

Compression robustness

Across the HDRTV1K fixed QP 27 to 42 sweep, the mixed-corpus LumaFlux model loses 0.55 dB. HDRTVNet++ loses 2.78 dB and FMNet loses 2.71 dB.

Getting the weights

Download the adapter checkpoints from the LumaFlux Hugging Face repository. Each file contains adapters only. The frozen FLUX.1-dev backbone and SigLIP are downloaded separately.

lumaflux-main.safetensors

Main checkpoint

Trained for 100k steps on the mixed UGC+PGC corpus of 314,396 pairs. Use it by default, for video, and for Luma-Eval.

lumaflux-hdrtv1k.safetensors

HDRTV1K checkpoint

Trained for 50k steps on the HDRTV1K train split. Use it for in-domain HDRTV1K comparisons.

Two-command quickstart
huggingface-cli download shreshthsaini/LumaFlux lumaflux-main.safetensors --local-dir .
python -m lumaflux.inference.cli --config configs/model/flux_dev.yaml --adapters lumaflux-main.safetensors --input input_sdr.mp4 --output output_hdr.mp4 --num-steps 8

Gated backbone: FLUX.1-dev is gated on Hugging Face. Accept its license separately before running inference.

Limitations

BibTeX

@article{saini2026lumaflux,
  title   = {LumaFlux: Lifting 8-Bit Worlds to HDR Reality with
             Physically-Guided Diffusion Transformers},
  author  = {Saini, Shreshth and Gedik, Hakan and Birkbeck, Neil and
             Wang, Yilin and Adsumilli, Balu and Bovik, Alan C.},
  journal = {arXiv preprint arXiv:2604.02787},
  year    = {2026}
}