Rectified Flow: A Technical Note on Objectives, Geometry, and Training

The curved paths are numerical solutions of an analytic Gaussian-mixture field. The animation blends them into straight segments with the same computed endpoints. This visual morph illustrates straightening; it does not train or simulate successive reflow models.

Rectified Flow as regression on a velocity field: why straight paths sample fast, where RF sits next to DDPM and flow matching, and how reflow buys few-step generation.

Thesis. Rectified Flow is best understood as supervised regression of a time-indexed velocity field under a chosen coupling between noise and data distributions.

Main technical point. Linear training paths can produce curved sampling trajectories. Their acceleration, the learned-field error, and the solver all affect few-step quality.

Practical implication. Check time sampling, solver error, and coupling choice before attributing poor few-step samples to the architecture.

Intuition: Straight Lines Are Cheap

Flow and diffusion generators move a simple distribution, usually Gaussian noise, toward the data distribution. Sampling repeatedly evaluates a learned field. Count those evaluations, or NFE, when comparing cost: an Euler step uses one field evaluation, while a higher-order step may use several. Guidance can add conditional and unconditional evaluations too.

A DDPM learns to reverse a noising process. Its associated deterministic samplers follow paths whose geometry depends on the schedule and the data. Rectified Flow trains on linear interpolations between paired endpoints. A path traversed at constant velocity can be integrated exactly with one Euler step. The aim of rectification is to approach that case, although a linear training interpolation alone does not make the learned flow straight.[1]

This note develops the coupling, regression objective, and sampling ODE, then checks their geometry in a linear example. The distinction between a training segment and a sampled trajectory will matter throughout.

Setup: Coupling, Interpolation, and Targets

Let \(x_0 \sim p_0\) be a base sample (usually Gaussian), and \(x_1 \sim p_1\) be a data sample. A coupling \(\pi(x_0, x_1)\) defines how these endpoints are paired∗[1]. Given \(t \in [0,1]\), define the linear interpolation

∗In this note, "coupling" means a joint distribution \(\pi(x_0, x_1)\) with marginals \(p_0\) and \(p_1\). Independent noise-data pairing is a common first-pass choice. Other couplings change the regression problem and need separate evaluation.

\[ x_t = (1-t)x_0 + t x_1. \]

The instantaneous displacement target is

\[ u(x_0,x_1) = x_1 - x_0. \]

Rectified Flow trains a neural vector field \(v_\theta(x,t)\) to predict this displacement from \((x_t,t)\). With the exact conditional-mean field and the regularity assumptions needed for a well-defined ODE, the flow matches the interpolation marginals and reaches \(p_1\). Finite training and numerical integration introduce approximation error.[1][2]

Curved Transport Rectified Transport noise data noise data
Geometric view: low-step ODE solvers approximate straight trajectories much better than curved ones.

One subtlety deserves emphasis early, because it explains most of what follows. Each pair \((x_0,x_1)\) is connected by a perfectly straight segment, but the learned field is not automatically straight. Many different pairs pass through the same point \((x,t)\), and the network can only output one velocity there, the average over all pairs passing through:

\[ v^\star(x,t)=\mathbb{E}\bigl[x_1-x_0 \,\big|\, x_t=x\bigr]. \]

When several endpoint pairs are compatible with the same \((x,t)\), their velocities are averaged. That conditional mean can bend the sampled trajectories even though every training segment is straight. Reflow replaces the original coupling with endpoints paired by the learned ODE. This can reduce the ambiguity; it does not guarantee that one retraining round removes all curvature.[1]

Loss Function and the Regression Optimum

The basic objective is

\[ \mathcal{L}_{\mathrm{RF}}(\theta) = \mathbb{E}_{(x_0,x_1)\sim\pi,\; t\sim\rho} \bigl[\|v_\theta(x_t,t) - (x_1-x_0)\|_2^2\bigr]. \]

Here \(\rho(t)\) is the time-sampling distribution (uniform is the default; non-uniform choices are often better in practice). The minimizer at each \((x,t)\) is the conditional expectation

\[ v^\star(x,t) = \mathbb{E}[x_1-x_0 \mid x_t=x,\; t]. \]

This follows from the projection theorem in \(L_2\):

\[ \mathbb{E}\|v-u\|^2 = \mathbb{E}\|v-v^\star\|^2 + \mathbb{E}\|u-v^\star\|^2, \]

so optimization can only reduce the first term; the second is irreducible variance from ambiguous pairings. This decomposition is useful when debugging: if training loss plateaus early with large data variance, the bottleneck may be coupling noise, not model capacity.

In matrix form for a batch of size \(B\), define \(X_t \in \mathbb{R}^{B\times d}\), \(U \in \mathbb{R}^{B\times d}\), and \(V_\theta(X_t,t)\in\mathbb{R}^{B\times d}\). Then

\[ \mathcal{L}_B(\theta)=\frac{1}{B}\|V_\theta(X_t,t)-U\|_F^2. \]

This Frobenius form makes clear that RF training is a standard least-squares regression over features \((x_t,t)\), which is why most optimization tricks transfer directly from diffusion training pipelines[5].

Where RF Sits: DDPM, Score SDE, Flow Matching

RF is easiest to place once you see that the whole family shares one template. Pick a path that interpolates between a noise sample and a data sample,

\[ x_t = \alpha_t\, x_1 + \sigma_t\, x_0, \qquad \dot{x}_t = \dot{\alpha}_t\, x_1 + \dot{\sigma}_t\, x_0, \]

then predict noise, a clean endpoint, a score, or a velocity. For Gaussian corruption, their population-optimal predictions can be converted into one another at interior times where the relevant coefficients are nonzero. At finite capacity, the parameterization and loss weighting affect optimization. Endpoint singularities and the sampler also matter.

ModelSchedule / interpolantNetwork targetTypical samplerPath geometry
DDPM[6] VP schedule, with time reversed for noise-to-data sampling noise \(\epsilon\) ancestral (stochastic), 50 to 1000 steps schedule- and data-dependent
Score SDE / EDM[5] VP or VE schedules, continuous time score \(\nabla_x\log p_t\) SDE or probability-flow ODE, 20 to 80 steps curved, schedule-dependent
Flow matching[2] any \((\alpha_t,\sigma_t)\), often linear velocity \(v\) ODE solver, 10 to 50 steps depends on path choice
Rectified Flow[1] \(\alpha_t=t\), \(\sigma_t=1-t\) (linear) velocity \(v\) Euler / midpoint, 1 to 8 steps after reflow straight training paths; curved flow possible

RF uses the conditional flow-matching objective with a linear interpolant and adds reflow as a way to change the coupling. Diffusion models also admit deterministic probability-flow ODEs. Their required step count depends on the learned field, schedule, solver, and error tolerance; it is not fixed by the choice of an MSE target.[1][2][3]

Matrix Sanity Check: Deterministic Linear Map

Consider a deterministic linear transport \(x_1 = A x_0\)†, with \(x_0 \sim \mathcal{N}(0, I)\). Then

†This section is intentionally "toy" but useful: if your implementation fails this case, the issue is usually target construction, time conditioning, or solver parameterization.

\[ x_t = \bigl((1-t)I + tA\bigr)x_0 = M_t x_0, \quad M_t := (1-t)I+tA. \]

Assume \(M_t\) is invertible for every time of interest. The target displacement is \(u=(A-I)x_0\), so expressed in terms of \(x_t\):

\[ u = (A-I)M_t^{-1}x_t = K_t x_t, \quad K_t := (A-I)M_t^{-1}. \]

Therefore the optimal velocity is exactly linear in \(x_t\) at each \(t\). If you fit a linear model class

\[ v_W(x,t)=W_t x, \]

then \(W_t=K_t\) is the minimizer. With samples stored as rows, the batch prediction is \(X_tW_t^\top\), so the normal equations give

\[ (W_t^{\star})^\top = (X_t^\top X_t)^{-1}X_t^\top U,\qquad \operatorname{rank}(X_t)=d. \]

This is an exact diagnostic for a deterministic linear coupling. It does not imply that a nonlinear image flow has the same closed-form regression solution. It also requires invertibility: for example, \(A=-I\) makes \(M_{1/2}\) singular.

A concrete instance makes the geometry vivid. Take \(A\) to be the 90° rotation \(A=\begin{pmatrix}0 & -1\\ 1 & 0\end{pmatrix}\). Then

\[ M_t=\begin{pmatrix}1-t & -t\\ t & 1-t\end{pmatrix}, \qquad \det M_t = (1-t)^2 + t^2 \;\ge\; \tfrac{1}{2}, \]

so the interpolation never collapses. The ODE solution is exactly \(x_t=M_tx_0\), with constant velocity \((A-I)x_0\) along each path. These trajectories are straight, despite the time dependence of \(K_t\). This deterministic rotation is already a reflow fixed point. Both endpoint distributions are standard Gaussians, so the identity map has zero quadratic transport cost, while the rotation has positive cost. Straightness therefore does not imply optimal transport in multiple dimensions.[1][12]

Sampling Dynamics and Why Curvature Matters

Generation solves the ODE

\[ \frac{dx_t}{dt} = v_\theta(x_t,t), \qquad x_{t=0}\sim p_0. \]

Euler integration with \(N\) steps uses

\[ x_{k+1} = x_k + \Delta t\, v_\theta(x_k,t_k), \qquad \Delta t = 1/N. \]

The second time derivative along the trajectory is

\[ \ddot{x}_t = \partial_t v_\theta(x_t,t) + J_{v_\theta}(x_t,t)\,v_\theta(x_t,t), \]

and this directly controls local truncation error:

\[ e_{\text{local}} = \frac{\Delta t^2}{2}\ddot{x}_t + \mathcal{O}(\Delta t^3). \]

For a fixed Euler step size, small acceleration reduces the leading local discretization error. Here “curvature” includes changes in speed along the path. Reflow can improve this geometry, but discretization error is only one part of the final sampling error.[1]

Straightness can be made quantitative. For a flow with trajectories \(x_t\), define

\[ S \;=\; \int_0^1 \mathbb{E}\,\bigl\|\, (x_1 - x_0) - \dot{x}_t \,\bigr\|^2 \, dt . \]

\(S=0\) means almost every trajectory is a straight line traversed at constant speed, so one Euler step is exact. The reflow theorem gives an \(O(1/K)\) bound on the best straightness value among the first \(K\) rectified flows, under its assumptions. It does not establish that every consecutive value of \(S\) decreases. Track \(S\) as a geometric diagnostic alongside sample-quality measurements.[1]

Euler has first-order global error under standard smoothness assumptions; Heun has second-order error and normally uses two field evaluations per step. Compare solvers at matched NFE, not just matched step count. EDM studies schedules and solvers for diffusion ODEs, while DPM-Solver exploits their particular semilinear structure.[5][11] Neither result makes one solver universally best for an arbitrary learned RF field. Check coefficient singularities when converting prediction types; the endpoint estimates below themselves contain no divisions.

Changing the coupling, the time parameterization, or the solver can change sampling error in different ways. Reflow changes which endpoints are paired. A solver only approximates the field it is given.

Flow explorer: watch the paths change A one-dimensional toy using analytic fields and numerical integration, in the spirit of the interactive figures in Fjelde et al.’s flow-matching introduction[13]. The straight-path mode joins numerically computed endpoints directly and has no training stage. The finite-noise VP example starts from an approximate Gaussian prior. This is a geometric illustration, not a matched-quality speed benchmark.
Diffusion SDE · 8 steps

Reflow, Distillation, and Few-Step Generation

Before reflow: crossings Idealized straight coupling average bends here noise data same marginals, no crossings
At the population optimum, rectification preserves endpoint distributions while changing their coupling. In practice, reflow learns from numerically generated teacher pairs, so its target distribution inherits the teacher and solver errors.[1]

A common rectification loop is:

  1. Train base \(v_{\theta_0}\) with the standard RF objective.
  2. Sample \(x_0\sim p_0\), then integrate \(\dot{x}=v_{\theta_0}(x,t)\) to obtain \(\hat{x}_1=\Phi_{\theta_0}^{0\to1}(x_0)\).
  3. Use pairs \((x_0,\hat{x}_1)\) to retrain a new field \(v_{\theta_1}\).

This process is closely related to reflow/trajectory-straightening ideas in RF literature and to practical guidance adaptations in conditional flow sampling[1][4].

The retrained objective is

\[ \mathcal{L}_{\mathrm{reflow}}(\theta) = \mathbb{E}_{x_0\sim p_0,\; t\sim\rho} \left[\left\|v_\theta\bigl((1-t)x_0+t\hat{x}_1,t\bigr) - (\hat{x}_1-x_0)\right\|^2\right]. \]

Time-dependent weights change the emphasis of the regression. One illustrative boundary-heavy choice is

\[ \mathcal{L}_{\lambda}(\theta) = \mathbb{E}\left[\lambda(t)\,\|v_\theta(x_t,t)-u\|_2^2\right], \qquad \lambda(t)=\frac{1}{t(1-t)+\tau}. \]

This example upweights both endpoints. \(\tau>0\) bounds the weight by \(1/\tau\); neither this formula nor a particular value of \(\tau\) is a generally established stability recipe. Validate it against unweighted training and account for the time-sampling density as well.

From a broader viewpoint, this sits inside the stochastic-interpolant family: change the interpolation law and you change both the conditional targets and the geometry of the learned field[3].

Reflow retrains a velocity field on pairs generated by a teacher flow. Distillation trains a student to reproduce a teacher mapping or a shorter sequence of steps. These can be combined, as in InstaFlow.[8] Distillation need not be limited to one step, and reflow does not guarantee graceful quality degradation at every lower step count. Measure quality and diversity at the actual deployment budget.

For exact population fields with well-defined ODE solutions, the rectification results distinguish three properties.[1][12]

  • Marginal preservation. Rectification preserves the prescribed endpoint distributions. A fitted neural model and a discretized solver need not preserve them exactly.
  • Transport-cost monotonicity. The expected displacement cost does not increase for convex \(c\), assuming the required expectations exist. This does not guarantee a globally optimal coupling for an arbitrary convex cost.
  • Best-iterate straightness. The theorem bounds the best \(S\) among the first \(K\) flows by a quantity of order \(1/K\). It does not assert strict monotonicity each round or identify every fixed point with an optimal transport map.

Progressive distillation halves a teacher sampler’s step count through successive student training.[9] Consistency models learn an endpoint map that agrees along a probability-flow trajectory and can be sampled in one or several steps.[10] Reflow retains a velocity field and changes its training coupling. Each approach needs its own quality and cost evaluation; teacher errors can be inherited, but teacher quality is not a universal mathematical ceiling on a retrained student.

Implementation Corner

Parameterization identities (useful in implementation)

If \(v\) is predicted, then endpoint estimates follow directly:

\[ \hat{x}_1 = x_t + (1-t)\,v_\theta(x_t,t), \qquad \hat{x}_0 = x_t - t\,v_\theta(x_t,t). \]

These formulas are often used for auxiliary losses: e.g., reconstruction losses on \(\hat{x}_1\), or consistency regularizers between endpoint predictions at neighboring times.

Illustrative weighted training loss

def rf_weighted_loss(model, x1, tau=1e-2):
    x0 = torch.randn_like(x1)
    # Beta sampling can emphasize boundaries without singular weights
    t = torch.distributions.Beta(0.9, 0.9).sample((x1.size(0),)).to(x1.device)
    t = t[:, None, None, None]

    xt = (1.0 - t) * x0 + t * x1
    target_v = x1 - x0
    pred_v = model(xt, t)

    w = 1.0 / (t * (1.0 - t) + tau)
    per_example = ((pred_v - target_v) ** 2).flatten(1).mean(1, keepdim=True)
    return (w.flatten(1) * per_example).mean()

Time sampling: where the gradient signal goes

The sampling density \(\rho(t)\) controls which times contribute most often to training. Uniform sampling is a baseline. Symmetric Beta distributions with parameters below one emphasize endpoints. A logit-normal distribution can emphasize the middle, depending on its mean and variance. SD3 selected a logit-normal recipe after ablation.[7] These density choices alone do not guarantee sharper textures or better global structure.

0 1 t ρ(t) Uniform Beta(0.9, 0.9), endpoint focus Logit-normal, mid-path focus
Schematic time-sampling densities. Uniform weights times equally; Beta(0.9, 0.9) emphasizes the boundaries; a suitable logit-normal emphasizes the middle. Their effect on image quality requires an experiment.[7]

Resolution changes the clock: timestep shift

SD3 motivates a resolution-dependent time shift using a signal-to-noise argument and then tests it empirically.[7] Its convention runs from data at zero to noise at one. With this post’s opposite convention, \(x_0\) noise and \(x_1\) data, the corresponding shift is

\[ t' \;=\; \frac{t}{s-(s-1)t},\qquad s>0. \]

For \(s>1\), this maps interior times toward zero, the noisy end in this note. SD3 uses a shift of 3.0 for its 1024px experiments. A training distribution over times and an inference step grid serve different purposes, so they need not be identical. Preserve the model’s time convention and compare candidate grids at a fixed evaluation budget.[7]

Implementation checks

  • Time conditioning. Preserve the checkpoint’s time scale and embedding convention. A model trained on a scaled discrete index should receive that scale; a continuous-time architecture should receive its expected continuous input.
  • Numerical guards. The endpoint estimates above have no divisions by time. Guard any additional score or noise conversion that divides by \(t\) or \(1-t\), and test its endpoint limits. Do not truncate the integration interval without checking the resulting bias.
  • Latent statistics. Use the VAE scaling and offset expected by the checkpoint. RF does not require the data distribution to have unit variance, but changing latent scaling changes the model’s input distribution.
  • Loss weighting is schedule design. A uniform-\(t\) velocity MSE is implicitly a particular SNR weighting of the equivalent \(\epsilon\)-loss; if you port a weighting recipe from a diffusion codebase, re-derive what it does to the velocity view before trusting it[5][7].
  • Pairing. A condition belongs to the data sample. Preserve that association when constructing batches. Re-pairing noise and data changes the coupling; check its effect instead of assuming a caption or class bucket reduces variance.
  • EMA evaluation. Compare EMA and current weights under the same sampler. Choose a teacher checkpoint using measured quality rather than assuming one averaging decay is always best.

Practical Notes (Detailed Checklist)

  • Time sampling. Start with uniform sampling or the pretrained model’s recipe. Test alternative densities while keeping the effective loss weighting explicit.
  • Loss scale. Inspect losses by channel and resolution. Changing target normalization changes the predicted velocity unless that transform is inverted at sampling time.
  • Checkpoint choice. Report whether results use EMA weights and how the averaging schedule was chosen.
  • Solver choices: Compare Euler and higher-order methods at matched evaluations; their image-quality ranking is empirical. If the model is trained for 8-16 step inference, verify quality with the exact solver used in deployment, not only with high-step validation.
  • Guidance scale. Sweep guidance at the intended sampling budget. Strong guidance can change the field’s derivatives and the solver error; a universal late-time increase is not established.
  • Reflow scheduling: One reflow pass often gives a large gain; additional passes may have diminishing returns. Measure both FID-like scores and user-facing prompt adherence before paying extra retraining cost.
  • Batch construction. Independent Gaussian noise paired with each data sample is a valid baseline for both unconditional and conditional RF. Keep captions aligned with their data samples.
  • Matrix diagnostics: Track \(\|J_v\|_F\) proxies (finite differences) and trajectory curvature statistics during validation. If curvature grows while training loss decreases, expect low-step sampling regressions.
  • Guidance window. Limited-interval CFG has empirical support in diffusion models. Treat its transfer to another flow model as an experiment, with the same prompt set and cost budget.[14][4]
  • Stochastic variants. EDM’s noise-injection recipe is derived for a diffusion sampler. Adding arbitrary noise to an RF ODE can change its marginals; use a justified stochastic formulation and evaluate it separately.[5]
  • NFE sweep. Evaluate several budgets, such as 1, 2, 4, and 8 field evaluations, plus a well-resolved reference. Report the solver, guidance passes, and wall-clock cost.
  • Step grid. Translate any schedule into the checkpoint’s time convention and validate it. Training-time sampling and inference-time discretization need not use the same density.[7]
  • Reflow diagnostics. Track \(S\), sample quality, and diversity by round. The population theorem is a best-iterate bound, so a non-monotone empirical curve alone does not identify a bug.[1]

What Ships Today

SD3 is a large-scale example of rectified-flow training with logit-normal time sampling and an MM-DiT backbone.[7] Its weighting recipe is not the illustrative boundary weight above. InstaFlow combines reflow and distillation for one-step text-to-image generation.[8]

Guidance changes the sampling field, so its effect must be evaluated with the solver and checkpoint in use. Our Rectified-CFG++ work studies a conditional predictor and a guidance correction evaluated at a provisional state.[4]

References

  1. Liu et al., Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow, 2022.
  2. Lipman et al., Flow Matching for Generative Modeling, 2022.
  3. Albergo, Boffi, and Vanden-Eijnden, Stochastic Interpolants: A Unifying Framework for Flows and Diffusions, 2023.
  4. Saini et al., Rectified CFG++ for Flow Based Models, NeurIPS 2025.
  5. Karras et al., Elucidating the Design Space of Diffusion-Based Generative Models, 2022.
  6. Ho et al., Denoising Diffusion Probabilistic Models, 2020.
  7. Esser et al., Scaling Rectified Flow Transformers for High-Resolution Image Synthesis (SD3), 2024.
  8. Liu et al., InstaFlow: One Step is Enough for High-Quality Diffusion-Based Text-to-Image Generation, 2023.
  9. Salimans and Ho, Progressive Distillation for Fast Sampling of Diffusion Models, 2022.
  10. Song et al., Consistency Models, 2023.
  11. Lu et al., DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling, 2022.
  12. Liu, Rectified Flow: A Marginal Preserving Approach to Optimal Transport, 2022.
  13. Fjelde, Mathieu, and Dutordoir, An Introduction to Flow Matching, 2024.
  14. Kynkäänniemi et al., Applying Guidance in a Limited Interval Improves Sample and Distribution Quality in Diffusion Models, 2024.

Citation

@misc{saini2025rectifiedflow,
  author       = {Saini, Shreshth},
  title        = {Rectified Flow: A Technical Note on Objectives, Geometry, and Training},
  year         = {2025},
  month        = {September},
  howpublished = {\url{https://shreshthsaini.github.io/blogs/rectified-flow.html}},
  note         = {Blog post}
}