Rectified-CFG++: Fixing Guidance for Flow Models
Rectified-CFG++ evaluates guidance at a predicted intermediate state. The reported image-quality gains vary by model and metric; its local perturbation bound has a narrower scope than exact manifold preservation.
Thesis. Strong classifier-free guidance can introduce artifacts in both diffusion and flow sampling. Rectified-CFG++ uses a conditional predictor and an intermediate guidance correction to change how that guidance enters each step.
Main technical point. The update anchors to the current conditional velocity, then evaluates the conditional-minus-unconditional difference at a provisional midpoint. This is a scheduled correction, not generally a convex interpolation.
Practical implication. The method needs no retraining, but adds a conditional evaluation per step. Compare quality at matched cost and check the checkpoint’s time and guidance conventions.
Guidance in a flow sampler
Classifier-free guidance∗[3] computes a guided velocity by linearly combining the unconditional and conditional predictions:
∗The scale convention matters. In the displayed formula, omega = 1 gives the conditional branch and omega = 0 the unconditional branch. A pipeline may use zero to mean a disabled guidance option.\[ \hat{v} = (1 - \omega)\, v^u + \omega \, v^c = v^u + \omega\,(v^c - v^u) \]
For \(\omega>1\), this extrapolates beyond the conditional velocity. Strong guidance can improve prompt adherence while introducing artifacts. A stochastic diffusion sampler adds noise, but that noise is not a projection onto a data manifold and does not guarantee correction of guidance errors. CFG also works with deterministic diffusion samplers.[3]
This post uses a flow ODE with noise at \(t=1\) and data at \(t=0\):
\[ \frac{dx_t}{dt} = v_\theta(x_t, t), \qquad t: 1 \to 0. \]
The ODE has no stochastic term. Errors may grow or contract according to the field’s dynamics; deterministic integration does not imply that every error must accumulate. The practical concern here is how guidance changes the trajectory and the resulting image quality.
These Flux examples illustrate changes in color and structure under different guidance settings. They are qualitative examples from the project, not evidence that a trajectory stayed on a particular manifold.[5]
Predictor and correction
Rectified-CFG++† uses a conditional predictor followed by a guidance correction evaluated at the predicted state.[5]
†The name "predictor-corrector" comes from numerical ODE methods (e.g., Heun's method). The idea is the same: use a provisional step to gather information, then correct the final update. Here the "correction" is guidance-aware rather than purely numerical.Each step uses the following sequence:
- Follow the conditional flow to a predicted midpoint
- Evaluate the guidance signal (conditional minus unconditional) at that midpoint
- Apply a time-dependent correction and complete the update
Rectified-CFG++ step with a signed time increment
- Predictor (conditional half-step).
\(h=t_{\mathrm{next}}-t<0,\qquad \tilde{x}_{t+h/2}=x_t+\frac{h}{2}v^c_\theta(x_t,t)\)Here \(h\) is a signed time increment. This Euler predictor approximates the conditional trajectory; it does not remain exactly on it. - Evaluate at the predicted midpoint.
\(v^c_{\mathrm{mid}}=v^c_\theta(\tilde{x}_{t+h/2},t+h/2),\quad v^u_{\mathrm{mid}}=v^u_\theta(\tilde{x}_{t+h/2},t+h/2)\)Both evaluations use the same predicted state. - Corrector (anchored guidance):
\(\hat{v} = v^c_\theta(x_t, t) + \alpha(t) \cdot \bigl(v^c_{\text{mid}} - v^u_{\text{mid}}\bigr)\) The anchor is the conditional velocity at the current point. The correction term is the guidance direction evaluated at the midpoint, scaled by \(\alpha(t)\). - Euler update.
\(x_{t+h}=x_t+h\hat v\)A descending time grid requires a negative increment when the network predicts \(dx_t/dt\). Check whether a particular implementation instead predicts a denoising direction.
The conditional velocity is the anchor, and \(\alpha(t)\) scales the midpoint correction. If both branches were evaluated at the current state, this would be ordinary CFG with \(\omega=1+\alpha\). The midpoint changes the evaluation location; it does not turn the algebra into a convex combination.
A schedule illustration
The paper’s Algorithm 1 writes \(\alpha(t)=\lambda_{\max}(1-t)^\gamma\). Under its stated noise-at-one convention, that expression grows toward the data endpoint. Its accompanying discussion of late-time decay is not consistent with that formula. The public SD3 implementation inspected on October 4, 2026 uses a constant true_cfg; its proposed ramp is commented out.[9]‡ The following two-ended bump is an illustration for this post, not the paper’s stated schedule:
\[ \alpha(t) \;=\; \lambda_{\max}\,\kappa_{\gamma,\delta}\; t^{\gamma}\,(1-t)^{\delta}, \qquad \kappa_{\gamma,\delta} = \Bigl(t_\star^{\gamma}(1-t_\star)^{\delta}\Bigr)^{-1}, \quad t_\star = \frac{\gamma}{\gamma+\delta}, \]
For positive \(\gamma\) and \(\delta\), this illustrative curve vanishes at both endpoints and peaks at \(t_\star\). The constant \(\kappa_{\gamma,\delta}\) normalizes its peak to \(\lambda_{\max}\). It is related in shape to limited-interval guidance, which has been studied separately for diffusion models.[8] The shape alone establishes no image-quality benefit.
- At \(t=1\). The illustrative correction is zero. The conditional predictor still uses the prompt.
- At interior times. The correction reaches its chosen peak. The exponents move and reshape this peak.
- Near \(t=0\). The added correction tends to zero. Earlier trajectory differences can remain.
Examining text rendering
Text rendering makes small structural errors easy to notice.§ The examples below show selected SD3 outputs, but they do not isolate which sampling phase caused a letter to improve or fail.
§Inspect the actual characters as well as edge sharpness. A crisp pseudo-letter can look convincing at a glance.When assessing these examples, separate three questions:
- Does the output contain the requested text, with the correct spelling and layout? A sharp-looking inscription can still contain the wrong characters.
- Does the comparison hold the prompt, seed, checkpoint, sampler, and compute budget fixed? A guidance change alone does not identify a late-step causal mechanism.
- Do the gains persist across prompts and guidance scales? Selected images support a qualitative comparison, while an OCR or human study is needed to measure legibility more broadly.
The following project examples compare CFG and Rectified-CFG++ on SD3. These selected images show lettering in several kinds of scenes; they are not a systematic text-accuracy test.
Stop Sign
Neon Sign
Carved Text
Newspaper Headline
Postage Stamp
These examples suggest differences in lettering and edge structure. They do not establish that constant guidance always harms text, or that the conditional model already predicts the correct letters before guidance is applied.
What the local bounds establish
The paper analyzes midpoint guidance and a one-step perturbation.¶[5] The conditions and the distinction between local and accumulated error matter. The formulation below makes time regularity explicit.
¶These are local bounds under stated regularity assumptions. Their constants are not measured for the neural networks in the examples.Lemma (Midpoint Guidance Consistency)
Suppose each velocity is \(L_x\)-Lipschitz in space and \(L_t\)-Lipschitz in time, and the conditional predictor has magnitude at most \(V_{\max}\). For a step of magnitude \(|h|\), the triangle inequality gives
\[\bigl\|(v^c_{\mathrm{mid}}-v^u_{\mathrm{mid}})-(v^c_t-v^u_t)\bigr\|\leq (L_xV_{\max}+L_t)|h|.\]
The spatial displacement and the change in time both contribute. Spatial Lipschitz continuity alone does not give the stated linear-in-step bound across two different times. A numerical step count, such as 20 or 50, does not by itself show that the bound is small.
One-step perturbation
Compare guided and conditional Euler updates from the same state. If \(\|v^c_{\mathrm{mid}}-v^u_{\mathrm{mid}}\|\leq B\), then
\[\|x_{t+h}^{\mathrm{guided}}-x_{t+h}^{\mathrm{cond}}\|\leq |h|\,|\alpha(t)|B.\]
This is a local comparison with a conditional Euler step from the same input. It is not a bound on distance to the exact conditional path, nor a claim that accumulated error disappears when guidance is switched off.
For two Euler trajectories starting together, a spatial Lipschitz bound gives a recurrence of the form \(e_{n+1}\leq(1+L|h_n|)e_n+|h_n|\,|\alpha_n|B\). Earlier deviations can persist or grow after \(\alpha_n\) becomes zero. A vanishing endpoint schedule alone therefore does not imply a collapsing tube or exact manifold preservation.
Results
The first table transcribes the selected columns of the paper’s MS-COCO 10K Table 1. FID is lower-is-better; the other scores are higher-is-better. Improvements are mixed: CLIP decreases for Lumina and SD3.5, for example.[5]
MS-COCO 10K: Across Architectures
| Model | Method | FID ↓ | CLIP ↑ | PickScore ↑ | HPSv2 ↑ |
|---|---|---|---|---|---|
| Lumina | CFG | 26.9321 | 0.3511 | 0.5867 | 0.2797 |
| Rect-CFG++ | 22.4899 | 0.3464 | 0.6133 | 0.3004 | |
| SD3 | CFG | 23.8898 | 0.3439 | 0.4408 | 0.2751 |
| Rect-CFG++ | 23.3945 | 0.3471 | 0.5591 | 0.2897 | |
| SD3.5 | CFG | 20.2945 | 0.3506 | 0.4923 | 0.2933 |
| Rect-CFG++ | 20.2169 | 0.3497 | 0.5077 | 0.2946 | |
| Flux-dev | CFG | 37.8625 | 0.3351 | 0.3248 | 0.2621 |
| Rect-CFG++ | 32.2262 | 0.3493 | 0.6752 | 0.2996 |
Guidance methods on SD3.5, MS-COCO 1K
| Method | FID ↓ | CLIP ↑ | ImageReward ↑ | HPSv2 ↑ |
|---|---|---|---|---|
| No guidance | 77.3049 | 0.3260 | 0.3852 | 0.2421 |
| CFG | 67.7133 | 0.3515 | 1.0530 | 0.2941 |
| CFG-Zero* | 68.3909 | 0.3458 | 0.9947 | 0.2879 |
| APG | 67.2311 | 0.3513 | 1.0748 | 0.2935 |
| Rect-CFG++ | 67.1495 | 0.3506 | 1.0845 | 0.2959 |
Values from Table 3 of the paper. Its 1K-sample FID values are not directly comparable with the 10K-sample values above.[5]
User Study
The paper reports a four-way study comparing CFG, APG, CFG-Zero*, and Rectified-CFG++. Thirty expert participants judged detail, naturalness and color, text legibility, and overall preference across 32 prompts and four backbones. These are shares of four-way selections, not a paired win rate against CFG.[5]
Implementation
The pseudocode below shows the paper’s midpoint predictor and Euler correction with a signed time increment. It accepts a schedule explicitly. The public SD3 pipeline differs: it makes a full-step conditional prediction, optionally adds noise, then applies a correction scaled by constant true_cfg. It batches both branches at both evaluations. This pseudocode is therefore an explanation of the paper’s update, not a transcription of that release.[9]
def rectified_cfgpp_sample(model, x, prompt, timesteps, alpha_fn):
"""x starts as noise; model returns dx/dt; times descend 1 to 0."""
for t, t_next in zip(timesteps[:-1], timesteps[1:]):
h = t_next - t # negative for this time convention
v_cond = model(x, t, prompt=prompt)
x_mid = x + 0.5 * h * v_cond
t_mid = t + 0.5 * h
v_cond_mid = model(x_mid, t_mid, prompt=prompt)
v_uncond_mid = model(x_mid, t_mid, prompt=None)
v_hat = v_cond + alpha_fn(t) * (v_cond_mid - v_uncond_mid)
x = x + h * v_hat
return x
A few implementation notes:
- Evaluation cost. This step uses one conditional evaluation at the current state and two branch evaluations at the midpoint. Standard true CFG uses two branch evaluations. This gives a 3-to-2 branch-evaluation ratio for the pseudocode at equal steps. The released SD3 pipeline evaluates four branches in two batched calls. Neither count guarantees a wall-clock ratio or fewer required steps.
- Schedule. Treat the bump plot as illustrative. The inspected SD3 release uses a constant correction coefficient. Peak values from the paper’s formula or the bump illustration do not transfer directly to that configuration.
- Compatibility. The paper evaluates Flux, SD3, SD3.5, and Lumina variants. Applying the formula elsewhere requires meaningful conditional and unconditional branches and the correct velocity parameterization.
- Batching the midpoint evaluations: The two midpoint passes share the same input \(\tilde{x}\), so run them as one batched forward with a doubled batch (the standard CFG trick). This makes two sequential calls with unequal batch sizes. Measure wall-clock time and peak memory on the target hardware.
- Solver choice. The pseudocode uses Euler. A Heun or midpoint wrapper needs fresh stage evaluations of the state-dependent guided field; replacing the final line alone does not produce a second-order method.
- Guidance-distilled checkpoints. A guidance embedding is not necessarily an explicit unconditional branch. Check the checkpoint and pipeline rather than assuming all Flux variants expose the same true-CFG interface.
- Text-heavy evaluation. Check requested spelling, long strings, and dense typography across seeds. Do not treat the selected examples or a high preference-model score as a guarantee of correct text.
- Evaluation. Report step count, branch evaluations, batching, wall-clock time, and memory. Sweep guidance for each method under the same prompt and seed protocol.
References
- Liu et al., Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow, 2022.
- Lipman et al., Flow Matching for Generative Modeling, 2022.
- Ho and Salimans, Classifier-Free Diffusion Guidance, 2022.
- Chung et al., CFG++: Manifold-constrained Classifier Free Guidance for Diffusion Models, 2024.
- Saini et al., Rectified-CFG++ for Flow-Based Models, NeurIPS 2025.
- Esser et al., Scaling Rectified Flow Transformers for High-Resolution Image Synthesis (Stable Diffusion 3), 2024.
- Black Forest Labs, Flux.1, 2024.
- Kynkäänniemi et al., Applying Guidance in a Limited Interval Improves Sample and Distribution Quality in Diffusion Models, 2024.
- Rectified-CFG++ public SD3 pipeline, master branch inspected October 4, 2026.
Citation
The paper
@inproceedings{saini2025rectifiedcfgpp,
title = {Rectified-CFG++ for Flow Based Models},
author = {Shreshth Saini and Shashank Gupta and Alan C. Bovik},
booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
year = {2025}
}
This post
@misc{saini2025rectifiedcfgpp_blog,
author = {Saini, Shreshth},
title = {Rectified-CFG++: Fixing Guidance for Flow Models},
year = {2025},
month = {December},
howpublished = {\url{https://shreshthsaini.github.io/blogs/rectified-cfgpp.html}},
note = {Blog post}
}