Rectified-CFG++: Fixing Guidance for Flow Models

A toy particle field illustrates how an extrapolation scale changes a target shape. It uses 520 particles and includes decorative noise and swirl. It does not simulate a trained generator or test Rectified-CFG++.

Rectified-CFG++ evaluates guidance at a predicted intermediate state. The reported image-quality gains vary by model and metric; its local perturbation bound has a narrower scope than exact manifold preservation.

Thesis. Strong classifier-free guidance can introduce artifacts in both diffusion and flow sampling. Rectified-CFG++ uses a conditional predictor and an intermediate guidance correction to change how that guidance enters each step.

Main technical point. The update anchors to the current conditional velocity, then evaluates the conditional-minus-unconditional difference at a provisional midpoint. This is a scheduled correction, not generally a convex interpolation.

Practical implication. The method needs no retraining, but adds a conditional evaluation per step. Compare quality at matched cost and check the checkpoint’s time and guidance conventions.

Guidance in a flow sampler

Classifier-free guidance∗[3] computes a guided velocity by linearly combining the unconditional and conditional predictions:

∗The scale convention matters. In the displayed formula, omega = 1 gives the conditional branch and omega = 0 the unconditional branch. A pipeline may use zero to mean a disabled guidance option.

\[ \hat{v} = (1 - \omega)\, v^u + \omega \, v^c = v^u + \omega\,(v^c - v^u) \]

For \(\omega>1\), this extrapolates beyond the conditional velocity. Strong guidance can improve prompt adherence while introducing artifacts. A stochastic diffusion sampler adds noise, but that noise is not a projection onto a data manifold and does not guarantee correction of guidance errors. CFG also works with deterministic diffusion samplers.[3]

This post uses a flow ODE with noise at \(t=1\) and data at \(t=0\):

\[ \frac{dx_t}{dt} = v_\theta(x_t, t), \qquad t: 1 \to 0. \]

The ODE has no stochastic term. Errors may grow or contract according to the field’s dynamics; deterministic integration does not imply that every error must accumulate. The practical concern here is how guidance changes the trajectory and the resulting image quality.

Diffusion SDE + CFG (stochastic increments) reference path noise guided drift noise increment Flow ODE + CFG (deterministic updates) reference path noise deviation reference path
Schematic paths under stochastic and deterministic updates. The arrows illustrate possible deviations, not a theorem that noise restores a manifold or that an ODE cannot correct an error.

These Flux examples illustrate changes in color and structure under different guidance settings. They are qualitative examples from the project, not evidence that a trajectory stayed on a particular manifold.[5]

No Guidance Flux output with no guidance
Guidance disabled
Standard CFG Flux output with standard CFG
CFG (w=3.5)
Rect-CFG++ Flux output with Rectified-CFG++
Ours (w=3.5)

Predictor and correction

Rectified-CFG++† uses a conditional predictor followed by a guidance correction evaluated at the predicted state.[5]

†The name "predictor-corrector" comes from numerical ODE methods (e.g., Heun's method). The idea is the same: use a provisional step to gather information, then correct the final update. Here the "correction" is guidance-aware rather than purely numerical.

Each step uses the following sequence:

  1. Follow the conditional flow to a predicted midpoint
  2. Evaluate the guidance signal (conditional minus unconditional) at that midpoint
  3. Apply a time-dependent correction and complete the update

Rectified-CFG++ step with a signed time increment

  1. Predictor (conditional half-step).
    \(h=t_{\mathrm{next}}-t<0,\qquad \tilde{x}_{t+h/2}=x_t+\frac{h}{2}v^c_\theta(x_t,t)\)Here \(h\) is a signed time increment. This Euler predictor approximates the conditional trajectory; it does not remain exactly on it.
  2. Evaluate at the predicted midpoint.
    \(v^c_{\mathrm{mid}}=v^c_\theta(\tilde{x}_{t+h/2},t+h/2),\quad v^u_{\mathrm{mid}}=v^u_\theta(\tilde{x}_{t+h/2},t+h/2)\)Both evaluations use the same predicted state.
  3. Corrector (anchored guidance):
    \(\hat{v} = v^c_\theta(x_t, t) + \alpha(t) \cdot \bigl(v^c_{\text{mid}} - v^u_{\text{mid}}\bigr)\) The anchor is the conditional velocity at the current point. The correction term is the guidance direction evaluated at the midpoint, scaled by \(\alpha(t)\).
  4. Euler update.
    \(x_{t+h}=x_t+h\hat v\)A descending time grid requires a negative increment when the network predicts \(dx_t/dt\). Check whether a particular implementation instead predicts a denoising direction.

The conditional velocity is the anchor, and \(\alpha(t)\) scales the midpoint correction. If both branches were evaluated at the current state, this would be ordinary CFG with \(\omega=1+\alpha\). The midpoint changes the evaluation location; it does not turn the algebra into a convex combination.

Geometric view of Rectified-CFG++ predictor-corrector on the flow manifold
The paper’s geometric illustration of the conditional predictor and guidance correction. A local update bound does not establish exact manifold preservation.[5]

A schedule illustration

The paper’s Algorithm 1 writes \(\alpha(t)=\lambda_{\max}(1-t)^\gamma\). Under its stated noise-at-one convention, that expression grows toward the data endpoint. Its accompanying discussion of late-time decay is not consistent with that formula. The public SD3 implementation inspected on October 4, 2026 uses a constant true_cfg; its proposed ramp is commented out.[9]‡ The following two-ended bump is an illustration for this post, not the paper’s stated schedule:

‡The plot is a separate teaching example. Neither zero correction nor a zero unconditional weight means that text conditioning has been removed.

\[ \alpha(t) \;=\; \lambda_{\max}\,\kappa_{\gamma,\delta}\; t^{\gamma}\,(1-t)^{\delta}, \qquad \kappa_{\gamma,\delta} = \Bigl(t_\star^{\gamma}(1-t_\star)^{\delta}\Bigr)^{-1}, \quad t_\star = \frac{\gamma}{\gamma+\delta}, \]

For positive \(\gamma\) and \(\delta\), this illustrative curve vanishes at both endpoints and peaks at \(t_\star\). The constant \(\kappa_{\gamma,\delta}\) normalizes its peak to \(\lambda_{\max}\). It is related in shape to limited-interval guidance, which has been studied separately for diffusion models.[8] The shape alone establishes no image-quality benefit.

  • At \(t=1\). The illustrative correction is zero. The conditional predictor still uses the prompt.
  • At interior times. The correction reaches its chosen peak. The exponents move and reshape this peak.
  • Near \(t=0\). The added correction tends to zero. Earlier trajectory differences can remain.
γ release = 2.0 · δ warmup = 1.5
An illustrative normalized bump compared with a constant correction. This is not the schedule stated in Algorithm 1, and the figure does not measure image quality. Sampling runs from right to left.

Examining text rendering

Text rendering makes small structural errors easy to notice.§ The examples below show selected SD3 outputs, but they do not isolate which sampling phase caused a letter to improve or fail.

§Inspect the actual characters as well as edge sharpness. A crisp pseudo-letter can look convincing at a glance.

When assessing these examples, separate three questions:

  1. Does the output contain the requested text, with the correct spelling and layout? A sharp-looking inscription can still contain the wrong characters.
  2. Does the comparison hold the prompt, seed, checkpoint, sampler, and compute budget fixed? A guidance change alone does not identify a late-step causal mechanism.
  3. Do the gains persist across prompts and guidance scales? Selected images support a qualitative comparison, while an OCR or human study is needed to measure legibility more broadly.

The following project examples compare CFG and Rectified-CFG++ on SD3. These selected images show lettering in several kinds of scenes; they are not a systematic text-accuracy test.

Stop Sign

Standard CFG (w=3.0) Stop sign with standard CFG showing text artifacts
CFG stop-sign detail
Rect-CFG++ (w=2.5) Stop sign with Rectified-CFG++ showing clean text
Rectified-CFG++ stop-sign detail

Neon Sign

Standard CFG (w=2.5) Neon sign with standard CFG
CFG neon-sign detail
Rect-CFG++ (w=2.5) Neon sign with Rectified-CFG++
Rectified-CFG++ neon-sign detail

Carved Text

Standard CFG (w=3.5) Carved text with standard CFG
CFG carved-text detail
Rect-CFG++ (w=2.5) Carved text with Rectified-CFG++
Rectified-CFG++ carved-text detail

Newspaper Headline

Standard CFG (w=3.5) Headline with standard CFG
CFG newspaper detail
Rect-CFG++ (w=3.0) Headline with Rectified-CFG++
Rectified-CFG++ newspaper detail

Postage Stamp

Standard CFG (w=3.5) Stamp with standard CFG
CFG stamp detail
Rect-CFG++ (w=3.5) Stamp with Rectified-CFG++
Rectified-CFG++ stamp detail

These examples suggest differences in lettering and edge structure. They do not establish that constant guidance always harms text, or that the conditional model already predicts the correct letters before guidance is applied.

What the local bounds establish

The paper analyzes midpoint guidance and a one-step perturbation.[5] The conditions and the distinction between local and accumulated error matter. The formulation below makes time regularity explicit.

These are local bounds under stated regularity assumptions. Their constants are not measured for the neural networks in the examples.

Lemma (Midpoint Guidance Consistency)

Suppose each velocity is \(L_x\)-Lipschitz in space and \(L_t\)-Lipschitz in time, and the conditional predictor has magnitude at most \(V_{\max}\). For a step of magnitude \(|h|\), the triangle inequality gives

\[\bigl\|(v^c_{\mathrm{mid}}-v^u_{\mathrm{mid}})-(v^c_t-v^u_t)\bigr\|\leq (L_xV_{\max}+L_t)|h|.\]

The spatial displacement and the change in time both contribute. Spatial Lipschitz continuity alone does not give the stated linear-in-step bound across two different times. A numerical step count, such as 20 or 50, does not by itself show that the bound is small.

One-step perturbation

Compare guided and conditional Euler updates from the same state. If \(\|v^c_{\mathrm{mid}}-v^u_{\mathrm{mid}}\|\leq B\), then

\[\|x_{t+h}^{\mathrm{guided}}-x_{t+h}^{\mathrm{cond}}\|\leq |h|\,|\alpha(t)|B.\]

This is a local comparison with a conditional Euler step from the same input. It is not a bound on distance to the exact conditional path, nor a claim that accumulated error disappears when guidance is switched off.

For two Euler trajectories starting together, a spatial Lipschitz bound gives a recurrence of the form \(e_{n+1}\leq(1+L|h_n|)e_n+|h_n|\,|\alpha_n|B\). Earlier deviations can persist or grow after \(\alpha_n\) becomes zero. A vanishing endpoint schedule alone therefore does not imply a collapsing tube or exact manifold preservation.

Results

The first table transcribes the selected columns of the paper’s MS-COCO 10K Table 1. FID is lower-is-better; the other scores are higher-is-better. Improvements are mixed: CLIP decreases for Lumina and SD3.5, for example.[5]

MS-COCO 10K: Across Architectures

ModelMethodFID ↓CLIP ↑PickScore ↑HPSv2 ↑
LuminaCFG26.93210.35110.58670.2797
Rect-CFG++22.48990.34640.61330.3004
SD3CFG23.88980.34390.44080.2751
Rect-CFG++23.39450.34710.55910.2897
SD3.5CFG20.29450.35060.49230.2933
Rect-CFG++20.21690.34970.50770.2946
Flux-devCFG37.86250.33510.32480.2621
Rect-CFG++32.22620.34930.67520.2996

Guidance methods on SD3.5, MS-COCO 1K

MethodFID ↓CLIP ↑ImageReward ↑HPSv2 ↑
No guidance77.30490.32600.38520.2421
CFG67.71330.35151.05300.2941
CFG-Zero*68.39090.34580.99470.2879
APG67.23110.35131.07480.2935
Rect-CFG++67.14950.35061.08450.2959

Values from Table 3 of the paper. Its 1K-sample FID values are not directly comparable with the 10K-sample values above.[5]

User Study

The paper reports a four-way study comparing CFG, APG, CFG-Zero*, and Rectified-CFG++. Thirty expert participants judged detail, naturalness and color, text legibility, and overall preference across 32 prompts and four backbones. These are shares of four-way selections, not a paired win rate against CFG.[5]

Four-way user study results by model and criterion
The paper’s four-way human-preference results, shown by model and criterion. Do not interpret a share here as a head-to-head win rate.[5]
Comparison across methods on multiple prompts
Selected qualitative comparisons from the project across guidance methods.[5]

Implementation

The pseudocode below shows the paper’s midpoint predictor and Euler correction with a signed time increment. It accepts a schedule explicitly. The public SD3 pipeline differs: it makes a full-step conditional prediction, optionally adds noise, then applies a correction scaled by constant true_cfg. It batches both branches at both evaluations. This pseudocode is therefore an explanation of the paper’s update, not a transcription of that release.[9]

def rectified_cfgpp_sample(model, x, prompt, timesteps, alpha_fn):
    """x starts as noise; model returns dx/dt; times descend 1 to 0."""
    for t, t_next in zip(timesteps[:-1], timesteps[1:]):
        h = t_next - t  # negative for this time convention
        v_cond = model(x, t, prompt=prompt)
        x_mid = x + 0.5 * h * v_cond
        t_mid = t + 0.5 * h

        v_cond_mid = model(x_mid, t_mid, prompt=prompt)
        v_uncond_mid = model(x_mid, t_mid, prompt=None)
        v_hat = v_cond + alpha_fn(t) * (v_cond_mid - v_uncond_mid)
        x = x + h * v_hat
    return x

A few implementation notes:

  • Evaluation cost. This step uses one conditional evaluation at the current state and two branch evaluations at the midpoint. Standard true CFG uses two branch evaluations. This gives a 3-to-2 branch-evaluation ratio for the pseudocode at equal steps. The released SD3 pipeline evaluates four branches in two batched calls. Neither count guarantees a wall-clock ratio or fewer required steps.
  • Schedule. Treat the bump plot as illustrative. The inspected SD3 release uses a constant correction coefficient. Peak values from the paper’s formula or the bump illustration do not transfer directly to that configuration.
  • Compatibility. The paper evaluates Flux, SD3, SD3.5, and Lumina variants. Applying the formula elsewhere requires meaningful conditional and unconditional branches and the correct velocity parameterization.
  • Batching the midpoint evaluations: The two midpoint passes share the same input \(\tilde{x}\), so run them as one batched forward with a doubled batch (the standard CFG trick). This makes two sequential calls with unequal batch sizes. Measure wall-clock time and peak memory on the target hardware.
  • Solver choice. The pseudocode uses Euler. A Heun or midpoint wrapper needs fresh stage evaluations of the state-dependent guided field; replacing the final line alone does not produce a second-order method.
  • Guidance-distilled checkpoints. A guidance embedding is not necessarily an explicit unconditional branch. Check the checkpoint and pipeline rather than assuming all Flux variants expose the same true-CFG interface.
  • Text-heavy evaluation. Check requested spelling, long strings, and dense typography across seeds. Do not treat the selected examples or a high preference-model score as a guarantee of correct text.
  • Evaluation. Report step count, branch evaluations, batching, wall-clock time, and memory. Sweep guidance for each method under the same prompt and seed protocol.

References

  1. Liu et al., Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow, 2022.
  2. Lipman et al., Flow Matching for Generative Modeling, 2022.
  3. Ho and Salimans, Classifier-Free Diffusion Guidance, 2022.
  4. Chung et al., CFG++: Manifold-constrained Classifier Free Guidance for Diffusion Models, 2024.
  5. Saini et al., Rectified-CFG++ for Flow-Based Models, NeurIPS 2025.
  6. Esser et al., Scaling Rectified Flow Transformers for High-Resolution Image Synthesis (Stable Diffusion 3), 2024.
  7. Black Forest Labs, Flux.1, 2024.
  8. Kynkäänniemi et al., Applying Guidance in a Limited Interval Improves Sample and Distribution Quality in Diffusion Models, 2024.
  9. Rectified-CFG++ public SD3 pipeline, master branch inspected October 4, 2026.

Citation

The paper

@inproceedings{saini2025rectifiedcfgpp,
  title     = {Rectified-CFG++ for Flow Based Models},
  author    = {Shreshth Saini and Shashank Gupta and Alan C. Bovik},
  booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
  year      = {2025}
}

This post

@misc{saini2025rectifiedcfgpp_blog,
  author       = {Saini, Shreshth},
  title        = {Rectified-CFG++: Fixing Guidance for Flow Models},
  year         = {2025},
  month        = {December},
  howpublished = {\url{https://shreshthsaini.github.io/blogs/rectified-cfgpp.html}},
  note         = {Blog post}
}