MEND: RL for Flow Models via Proximal Velocity Matching

Shreshth Saini1,2 Neil Birkbeck2 Yilin Wang2 Balu Adsumilli2 Alan C. Bovik1,3

1The University of Texas at Austin 2Google 3University of Colorado Boulder

arXiv:2610.05954 · 2026

Five prompts, each shown as an SD3.5-M image above and a MEND image below from the same initial noise, with PickScore on each tile. MEND scores higher on all five.
Training reward against updates for MEND, ReFL and DiffusionNFT under the equal-budget protocol. MEND is highest at every evaluated update. Held-out PickScore against updates. MEND reaches 24.03 at update 100, above ReFL and DiffusionNFT, and passes the Flow-GRPO level before update 30. Trained HPSv2.1 reward against updates for the MEND run trained on HPSv2.1, with ReFL and DiffusionNFT.
MEND stays ahead of ReFL and DiffusionNFT at every evaluated update. Top: SD3.5-M above, MEND below, same prompt and initial noise. Bottom, equal budget: (a) training reward and (b) held-out PickScore against ReFL and DiffusionNFT under the same protocol, with Flow-GRPO (about 4k updates) at its own setting as a level; (c) the HPSv2.1 run's trained reward.

Abstract

Reward post-training of flow models either reweights the model's own samples under a KL penalty or a frozen reference, often for thousands of updates, or backpropagates the reward and moves every sample without checking that the move is worth its size. We introduce MEND, a reinforcement learning method built on proximal velocity matching. MEND caps rewards within each prompt group, so samples that already score well receive no move. Below the cap, it proposes moves along the reward gradient and accepts one only when its capped reward gain exceeds a quadratic displacement price. The model then regresses onto the resulting velocity targets, with no KL term, frozen reference model, or advantage weights. In 100 updates, MEND outperforms Flow-GRPO (about 4k updates) on five of six evaluators at the same distance to base-model images. Under an equal-budget protocol, it surpasses ReFL and DiffusionNFT at every evaluated update across four training rewards, reaching PickScore 24.03 versus 23.92 and 23.43, respectively. A 300-update three-reward run also surpasses the five-reward DiffusionNFT model on all three rewards it trains on. MEND is general and easy to adopt: it applies to any flow backbone with a differentiable reward.

Motivation and background

Learned reward models such as PickScore, HPSv2.1 and ImageReward score an image for human preference. Reward post-training uses such a score to update a pretrained text-to-image flow model. The open problem is how to turn a scalar score into an update that raises reward quickly without eroding what the base model already does well. A score alone does not say which samples should change or where they should go.

Existing methods fall into three families.

  • Policy gradient. Flow-GRPO reweights stochastic trajectories by group-relative advantages under a KL penalty.
  • Weighted regression. DiffusionNFT reweights the model's own samples inside a regression loss.
  • Reward backpropagation. ReFL and DRaFT differentiate the reward through sampling, which supplies a direction.

The first two adjust how much existing samples count, and the Flow-GRPO and DiffusionNFT models we compare against use about 4k and 1.7k updates. Reward backpropagation moves every sample, including those that already score well, and nothing checks that a move earns its size. A KL penalty or reference term limits drift of the distribution as a whole. None of them decides, sample by sample, whether a change is worth making.

A one-dimensional flow with three modes in four panels: the pre-trained flow, reward reweighting, a gradient target, and MEND. Reweighting loses the hard mode, the gradient target pushes samples off the data, and MEND keeps all three modes.
(a) The base flow. (b) Reward reweighting concentrates on the peaks and loses the hard mode. (c) Gradient ascent moves every sample and pushes some of the mass off the data. (d) MEND: capped samples stay, certified moves are short, and the hard mode keeps its mass.

Objective components. The target check compares reward gain with displacement cost. MEND uses a lagged behavior adapter, not a frozen reference. Updates refer to the compared runs.

MethodKL / frozen referenceReward weightsReward gradientTarget checkUpdates
Flow-GRPOyesyesnono~4k
ReFLnonoyesno100
DiffusionNFTyesyesnono1.7k
MENDnonoyesyes100

Method

MEND rests on one principle: a sample's target moves only when the reward it gains pays for the distance it travels. Each round applies this principle sample by sample to build targets, then fits them with a plain regression loss.

MEND overview in three panels. Panel 1: a group of rollouts for one prompt with a reward cap, the top sample kept. Panel 2: three proposals along the reward gradient for a sample below the cap, the shortest accepted and two rejected after the distance price. Panel 3: the trained adapter regresses onto the behavior prediction plus the accepted displacement.
MEND overview. (1) A group for prompt \(c\); samples at or above the cap \(\kappa\) are kept. (2) Each sample below the cap gets 3 proposals along its reward gradient; the verdict subtracts a distance price from the capped reward and keeps the best candidate, here the shortest move \(y_1\). (3) The trained adapter regresses at the stored state onto the behavior prediction plus the accepted displacement \(d\); the behavior model then follows it by EMA.

Two adapters share the base model: the trained adapter \(\theta\) and a behavior adapter \(\theta_{\mathrm{old}}\). For each prompt \(c\), the behavior adapter generates a group of \(G\) deterministic 10-step rollouts with endpoints \(x^{(1)},\ldots,x^{(G)}\) and a stored intermediate state \(z_q\) near \(t_q = 0.278\). \(R(x,c)\) scores the decoded latent and \(g=\nabla_x R(x,c)\) is its gradient.

  1. Cap. Pushing samples that already score well buys little reward and costs diversity. Within each prompt group, rewards are capped at the quantile \(q = 0.75\), subject to a global floor \(\kappa_{\mathrm{glob}}\) that rises slowly across rounds. Samples at or above the cap receive zero displacement.
    \[ \kappa(c)=\max\Big\{Q_q\big(\{R(x^{(k)},c)\}_{k=1}^{G}\big),\;\kappa_{\mathrm{glob}}\Big\},\qquad R_\kappa(x,c)=\min\{R(x,c),\,\kappa(c)\} \]
  2. Propose. Each sample below the cap receives \(K=3\) proposals along its normalized reward gradient. Each is decoded and scored once, so the decision rests on the reward attained at \(y_j\), not on a first-order prediction of it.
    \[ y_j = x + \eta_j\,\frac{g}{\lVert g\rVert},\qquad \eta_j\in\{0.1,\,0.2,\,0.4\} \]
  3. Verify. A proposal must pay for its length. The verdict selects the candidate with the highest capped reward minus a quadratic displacement price, and the unchanged sample is always among the candidates.
    \[ y^{\star}=\arg\max_{y\in\{x,\,y_1,\ldots,y_K\}} J(y),\qquad J(y)=R_\kappa(y,c)-\frac{\lVert y-x\rVert^{2}}{2\tau} \]
    This is the proximal-point objective, maximized over a finite candidate set. The price \(\tau\) is fixed within a round and adjusted between rounds to target an acceptance rate of 30% to 60% among proposed samples.
  4. Regress. Let \(d=y^{\star}-x\) for accepted samples and \(d=0\) otherwise. At the stored state, the trained adapter matches the behavior prediction plus \(d\), with \(\lambda_{\mathrm{keep}}=10\).
    \[ \mathcal L(\theta)=\mathbb E_{\mathrm{moved}}\big\lVert \hat x_\theta(z_q)-\operatorname{sg}[\hat x_{\mathrm{old}}(z_q)+d]\big\rVert^{2}+\lambda_{\mathrm{keep}}\,\mathbb E_{\mathrm{kept}}\big\lVert \hat x_\theta(z_q)-\operatorname{sg}[\hat x_{\mathrm{old}}(z_q)]\big\rVert^{2} \]
    A kept sample asks for no change and a moved sample asks for exactly the correction \(d\).

Since \(\hat x_\theta(z_q)=z_q-t_q\,v_\theta(z_q,t_q,c)\), each term of the loss is a velocity-matching loss:

\[ \big\lVert \hat x_\theta(z_q)-\operatorname{sg}[\hat x_{\mathrm{old}}(z_q)+d]\big\rVert^{2}=t_q^{2}\,\big\lVert v_\theta(z_q,t_q,c)-v^{\star}\big\rVert^{2},\qquad v^{\star}=\operatorname{sg}\big[v_{\mathrm{old}}(z_q,t_q,c)-d/t_q\big] \]

Kept samples have \(v^{\star}=v_{\mathrm{old}}\). MEND is therefore velocity matching onto a target selected by the proximal rule, which is why we call it proximal velocity matching. The behavior adapter follows \(\theta\) by EMA, so the keep term anchors velocities to a lagged adapter that follows training, not to a frozen reference. Each round takes one AdamW step, and MEND never backpropagates through the sampler.

Theory

Verified target selection. Every accepted target improves capped reward by more than its quadratic price, and its displacement bound shrinks to zero as the sample approaches the cap:

\[ \frac{\lVert y^{\star}-x\rVert^{2}}{2\tau}\;\le\;R_\kappa(y^{\star})-R_\kappa(x)\;\le\;\kappa-R_\kappa(x) \]

First-order realization. At the behavior parameters, the descent direction of the regression loss backpropagates normalized reward gradients through the clean predictions of accepted samples only, scaled by their selected steps, so MEND keeps the direction of reward backpropagation and adds the two decisions it lacks: which samples move, and how far.

Results

What is compared. We train SD3.5-M on PickScore under the equal-budget protocol of Zhou et al. (2026): 48 Pick-a-Pic prompts × 24 images per update, rank-32 LoRA, 10-step rollouts at 512 pixels, and 100 updates. ReFL and DiffusionNFT are the baselines under this protocol. We also compare with the Flow-GRPO PickScore adapter (about 4k updates) and the five-reward DiffusionNFT model (1.7k updates). Evaluation uses DrawBench, 200 prompts × 5 seeds and 40 Euler steps, scored with six evaluators. Distance to base is the DreamSim distance to the base image generated with the same prompt, noise and sampling settings.

Main comparison

MEND outperforms Flow-GRPO on five of six evaluators at the same base distance. Its three-reward run surpasses DiffusionNFT on those three rewards. DrawBench 200 × 5. Bold: best; underline: second best.

MethodRewardsUpdatesPickScoreHPSv2.1HPSv3ImageRewardCLIPScoreAestheticDist. to base ↓
SD3.5-Mnone022.350.2802.720.830.2835.390
Flow-GRPOPickScore~4k23.520.3167.041.270.2805.900.313
DiffusionNFTfive1.7k23.820.3317.501.490.2926.020.538
MEND (ours)PickScore10023.700.3197.151.320.2915.880.313
MEND (ours)three30023.890.3445.871.280.3005.920.488

MEND outperforms Flow-GRPO on five of six evaluators at the same base distance (0.313), with roughly 40-fold fewer updates. Aesthetic score is the one evaluator on which Flow-GRPO is higher (5.90 versus 5.88). The run trained on PickScore, HPSv2.1 and CLIPScore surpasses DiffusionNFT (1.7k updates on five rewards) on those three evaluators at a smaller base distance. It is lower on HPSv3, ImageReward and aesthetic score, which it does not train on.

Equal budget, four training rewards

Each MEND run uses SD3.5-M, 100 updates and the equal-budget protocol; the base uses the same evaluation setting. The last two columns give ReFL and DiffusionNFT under this protocol on the row's training reward. Bold: MEND's trained metric; underline: second best among the three methods on that metric.

Training rewardPickScoreHPSv2.1ImageRewardCLIPScoreReFLDiffusionNFT
none (SD3.5-M)20.580.207−0.520.239
PickScore24.030.3011.130.27323.9223.43
ImageReward22.330.3031.510.2661.281.46
HPSv2.123.060.3601.210.2680.3580.336
CLIPScore22.270.2680.980.3140.3080.298

At equal budget, MEND is ahead of both baselines at every evaluated update. It reaches PickScore 24.03 at update 100, versus 23.92 for ReFL and 23.43 for DiffusionNFT. Separate runs trained on ImageReward, HPSv2.1 and CLIPScore are each above both baselines on their training reward at update 100 and at every evaluated update. All four runs finish above the base on every held-out evaluator.

Trained reward against updates for MEND runs trained on CLIPScore and on ImageReward, with ReFL and DiffusionNFT under the same protocol. MEND is highest at every evaluated update in both panels.
Trained reward against updates under the equal-budget protocol, one MEND run per reward, against ReFL and DiffusionNFT under the same protocol; markers give their 100-update values.

Update efficiency and fidelity

On 64 unseen Pick-a-Pic prompts, MEND is above DiffusionNFT under the same protocol at every evaluated update. By update 25 it also surpasses the levels of Flow-GRPO and DiffusionNFT, which use about 4k and 1.7k updates. Reward keeps improving while base distance barely changes late in training: between updates 25 and 100, DrawBench PickScore rises from 23.36 to 23.70, while base distance moves only from 0.294 to 0.313. Training costs 10.0 GPU-hours on 3 GB200 GPUs, excluding evaluation.

Three panels against updates: PickScore on unseen prompts for MEND and DiffusionNFT with Flow-GRPO and DiffusionNFT levels; percent change over SD3.5-M on ImageReward, HPSv2.1 and CLIPScore; and distance to base, which flattens near the Flow-GRPO level after about 30 updates.
(a) PickScore on 64 held-out Pick-a-Pic prompts at the protocol's setting, against DiffusionNFT under the same protocol, with Flow-GRPO and DiffusionNFT (about 4k and 1.7k updates) as levels; (b) held-out evaluators and (c) distance to the base on DrawBench at guidance 4.5.
Two prompts followed through training: SD3.5-M and MEND after 25, 50, 75 and 100 updates from the same initial noise, with PickScore on each tile.
SD3.5-M and MEND after 25, 50, 75 and 100 updates, with the same initial noise along each row and PickScore on each tile. Top: the grasshopper appears by update 50 and the moustache by update 75. Bottom: the wifi symbol appears at update 50 and is sharp by update 75.

Qualitative comparisons

In matched comparisons, MEND corrects prompt errors of the base while keeping its composition. All grids use matched prompts and initial noise, with PickScore on each tile. Many MEND examples have higher saturation, which reflects a preference of the reward.

Grid of matched prompts, one per row, with columns SD3.5-M, Flow-GRPO, DiffusionNFT and MEND, and PickScore on each tile.
Rows share a prompt and seed across SD3.5-M, Flow-GRPO, DiffusionNFT, and MEND. PickScore on each tile.
DrawBench prompts on counts, attributes and spatial relations, SD3.5-M above and MEND below from the same initial noise, with PickScore on each tile.
DrawBench prompts (text under each column), SD3.5-M above and MEND below from the same initial noise, PickScore on each tile. MEND renders counts, attributes or spatial relations that SD3.5-M misses.
Additional matched pairs, SD3.5-M above and MEND below, with PickScore on each tile.
More SD3.5-M and MEND results.

Generality

MEND is not tied to one backbone. It applies to any flow model with a differentiable reward, and it improves every backbone we test at the same update budget. On SD3-M, PickScore training raises PickScore from 20.36 to 23.70, above Linear-DPO at 20.96, and held-out HPSv2.1 from 0.215 to 0.298. On Z-Image-Turbo, a distilled model sampled in nine steps at 1024 pixels, PickScore training raises PickScore from 22.86 to 24.10 and held-out HPSv2.1 from 0.295 to 0.315. HPSv2.1 training reaches 0.357 with held-out PickScore 23.27.

SD3-M. DrawBench 200 × 5, update 100.

ModelPickScoreHPSv2.1
SD3-M20.360.215
Linear-DPO20.960.256
MEND, PickScore23.700.298

Z-Image-Turbo. DrawBench 200 × 5, 1024 pixels, 9 steps.

ModelPickScoreHPSv2.1
Z-Image-Turbo22.860.295
MEND, PickScore24.100.315
MEND, HPSv2.123.270.357

Limitations. MEND requires a differentiable reward. Its guarantees hold for the targets of each round. Longer training can over-optimize a reward, and some single-reward runs lose held-out score within the 100-update budget.

BibTeX

@misc{saini2026mend,
  title         = {{MEND}: {RL} For Flow Models via Proximal Velocity Matching},
  author        = {Saini, Shreshth and Birkbeck, Neil and Wang, Yilin and Adsumilli, Balu and Bovik, Alan C.},
  year          = {2026},
  eprint        = {2610.05954},
  archivePrefix = {arXiv},
  primaryClass  = {cs.LG},
  doi           = {10.48550/arXiv.2610.05954},
  url           = {https://arxiv.org/abs/2610.05954}
}