Preprint · under review

ReCASTReward Credit Assignment across Timesteps for Online Diffusion Reinforcement

Give each reward influence at the denoising steps where its feedback is most informative.

1UCLA    2Arena AI    *Equal contribution

The idea in one minute

Preference says how much. ReCAST learns when.

Multi-reward diffusion fine-tuning needs to answer two different questions. The budget λ should encode the user's trade-off across rewards; a Rényi discriminability curve should decide where each reward spends that budget across denoising time.

2four-reward
settings
30matched
model pairs
8/8held-out judge
means improved
57.2%OCR preference
vs. static
Full abstract

Training diffusion models with multiple rewards requires distinguishing user preference from reward informativeness. User preference determines how much each reward should contribute to the overall objective; reward informativeness determines when its feedback is useful during denoising. Some rewards can meaningfully evaluate a sample as soon as global structure emerges, but others become informative only when the sample is nearly clean. To address both questions jointly, we propose ReCAST (Reward Credit ASsignment across Timesteps), the first method, to our knowledge, for per-reward, timestep-dependent credit assignment in diffusion reward fine-tuning. ReCAST separates user preferences from temporal allocation through a reward-by-timestep weight matrix W, whose row sums match the user-specified reward budgets λ, while its column sums are equal, assigning the same total weight to each denoising step. Under these marginal constraints, ReCAST allocates weight according to each reward's informativeness, quantified by its Rényi discriminability gain at each step. These gains telescope to the total discriminability between the reward-induced positive policy and the current policy, providing a basis for temporal credit assignment.

We evaluate ReCAST by training SD3.5-Medium under two distinct four-reward settings, each across five reward budgets λ. ReCAST improves the training rewards in one setting and matches them in the other, improves every held-out judge in both, and is preferred by an independent LLM-as-a-Judge. Together, these results show that ReCAST yields improvements that generalize beyond the training rewards and support its core principle: assigning each reward greater weight at the denoising timesteps where its feedback is most informative.

How it works

Method

Replace the static weights by a reward-by-timestep matrix W, constrained on both margins:

rt(x0, c) = ∑i T · Wi,t · ri(x0, c)
∑t Wi,t = λi reward i's budget, set by the user ∑i Wi,t = 1/T every step gets the same total

The margins fix the totals but not the shape. How row i spreads its budget across t is exactly when reward i counts — and that is what we estimate instead of hand-setting.

1. Measure

The Rényi divergence Di,t between reward i's positive policy and the current one, per step. Its gain ΔDi,t is the new discriminability at that step, and the gains telescope to the data-level total — a decomposition across t with no double-counting. A de Bruijn-type theorem gives ΔDi,t ≥ 0.

2. Project

Gains measure utility but are not valid weights. Mean-normalize each curve, so only its shape survives, exponentiate into a kernel, and take its KL projection onto the two margins — a unique Sinkhorn solution, <100 log-domain iterations. A flat curve returns exactly the static baseline.

3. Train

Only the posterior expected reward needs estimating; the marginal cancels in the gain. Self-rollouts are too weak a proposal, so one strong external generator (GPT Image 1.5) stands in for every reward. The matrix is built once per λ, then reused unmodified.

Three panels: per-step discriminability gain for seven rewards, the mean-normalized gain, and the Sinkhorn weight matrix it projects to.
Left: per-step gain ΔDi,t. Center: mean-normalized, the kernel's exponent. Right: the Sinkhorn matrix at uniform λ, over the four rewards of the OCR setting; both margins hold exactly. Axis is t/T, cleanest step left to noisiest right. The curves are strongly reward-specific — GenEval, OCR, and ImageReward peak sharply toward the noisy end, the rest stay near uniform — and under uniform λ the aggregate demand varies by 5.4× across steps.

What changed

Results

Generalization is the headline. ReCAST improves all four held-out judges in both settings—even when the training rewards are effectively tied in Stage-GenEval.
Model
LoRA on SD3.5-Medium, DiffusionNFT, β = 0.1, T = 25
Settings
Stage-OCR and stage-GenEval, four rewards each
Pairs
5 budgets × 3 seeds = 15 matched pairs per setting
Baseline
Static weighting at the same λ, same seeds

Stage-OCR training rewards and every held-out judge improve

RewardbasestaticReCASTwin
Training rewards
Aggregate r1.6422.8672.92011/15
ClipScore0.2710.3130.32013/15
HPSv20.1950.2880.30312/15
PickScore0.7730.8880.89915/15
OCR target0.1290.8930.9009/15
Held-out judges
ImageReward1.1401.4221.48213/15
Aesthetic5.3885.4865.57812/15
HPSv312.6912.7413.1411/15
UnifiedReward-23.0153.1543.17611/15

Stage-GenEval training rewards tie; every held-out judge improves

RewardbasestaticReCASTwin
Training rewards — tied overall
Aggregate r1.7842.8812.8787/15
ClipScore0.2400.2990.2996/15
HPSv20.2130.3130.3125/15
PickScore0.7900.9110.9086/15
GenEval target0.2440.8730.8749/15
Held-out judges
ImageReward1.0241.1981.2899/15
Aesthetic5.9495.8645.96810/15
HPSv39.198.028.5810/15
UnifiedReward-22.6822.6712.7059/15

In Stage-OCR, all four training rewards and all four held-out judges improve. In Stage-GenEval, whose training rewards are tied in the paper, every held-out judge also improves.

Mean over 5 budgets × 3 seeds at the final checkpoint; base is untuned SD3.5-Medium, win counts matched pairs won. Bold marks the better of static and ReCAST. Standard deviations are in the paper.

Independent LLM-as-a-Judge MMRBv2 · 1,000 held-out OCR prompts · positions swapped

57.18%overall preference
58.00%faithfulness preference
54.97%aesthetics preference
4/5budgets prefer ReCAST

A multimodal judge (gemini-3.5-flash, temperature 0) scores every pair twice with image positions reversed to cancel position bias. A rate of 50% is parity; higher values favor ReCAST. Per-budget rates are in the paper.

Citation

BibTeX

@misc{chen2026recast,
  title  = {ReCAST: Reward Credit Assignment across Timesteps for Online
            Diffusion Reinforcement},
  author = {Chen, Yihang and Ban, Yuanhao and Kao, Kuei-Chun and Hsieh, Cho-Jui},
  year   = {2026},
  eprint = {2609.13425},
  archivePrefix = {arXiv}
}