Off-policy reinforcement learning

Reshaped Sequence Policy Optimization

Gradient signal should follow what a response needs—not how stale it has become.

ReSPO reshapes sequence-level importance weights by advantage sign. It preserves learning signal for under-generated positive responses and suppresses severely over-generated negative responses.

Gradient allocation sequence level
\(\hat{A} > 0\) Positive response

Preserve signal when the response is becoming less likely.

\(\hat{A} < 0\) Negative response

Suppress both already-suppressed and over-generated tails.

\(W\) is the sequence likelihood ratio between the current and rollout policies.

00 / Abstract

Learning from reused rollouts without starving useful gradients.

Reusing rollouts across policy updates saves generation cost, but the current policy gradually diverges from the policy that produced the data. Standard clipped objectives then underweight positive responses that have become unlikely, while large-ratio negative responses can dominate an update.

ReSPO replaces clipping with a smooth, two-branch sequence-level kernel derived from an \(\alpha\)-divergence objective and an exponential variance-control tilt. The result is faster early optimization, stronger late-stage training scores, and higher held-out benchmark performance when rollout reuse is high.

01 / The idea

Clipping leaves two tails untreated.

Positive and negative group-relative advantages need different behavior at the extremes of policy drift.

01

Starved positive update

For \(\hat{A} > 0\) and \(W \to 0\), clipping makes the recovery signal vanish. A useful response that the current policy is forgetting receives almost no gradient.

ReSPO keeps a nonzero limit
02

Dominating negative update

For \(\hat{A} < 0\) and \(W \to \infty\), the unclipped gradient weight grows without bound. A negative response can overwhelm the rest of a reused rollout batch.

ReSPO drives the tail to zero

Required tail limits

One kernel family, two sign-specific branches.

Branch \(W \to 0\) \(W \to \infty\)
Positive \(\hat{A} \geq 0\) preserve suppress
Negative \(\hat{A} < 0\) suppress suppress

02 / Method

Reshape the sequence weight, then control its variance.

ReSPO turns tail requirements into a smooth per-sequence gradient coefficient.

Positive and negative ReSPO kernels compared with clipping and VESPO across the sequence importance ratio
ReSPO preserves positive-branch signal near \(W=0\) while suppressing both negative-branch tails.
01

Change the measure

Interpret replacing \(W\) with \(\phi(W)\) as defining an unnormalized measure over responses.

02

Shape each tail

Use an \(\alpha\)-divergence power mean, selecting \(\alpha\) by the sign of the group-relative advantage.

03

Apply a smooth tilt

Project under a moment constraint to obtain an exponential factor that controls large weights.

Positive branch

\(\displaystyle \phi^{(+)}(W)=\frac{1+W}{2}\exp(1-W)\)

Maximum signal at \(W=0\); smooth decay as the response becomes over-generated.

Negative branch

\(\displaystyle \phi^{(-)}(W)=\sqrt{W}\exp\!\left(2(1-\sqrt{W})\right)\)

Zero weight at both extremes; value and slope match the positive branch at \(W=1\).

03 / Results

The advantage grows when rollout reuse is hardest.

Experiments cover Qwen3-1.7B and Qwen3-30B-A3B, three rollout-reuse ratios, and six held-out math benchmarks.

\(\frac{5}{6}\)

Highest late score

Highest observed mean over the final 128 updates across model–reuse settings.

\(\frac{17}{24}\)

High-reuse benchmark wins

Highest observed value across model–benchmark comparisons at \(N=16\) and \(N=32\).

\(\frac{4}{4}\)

High-reuse evaluations

Highest benchmark average at \(N=16\) and \(N=32\) on both models.

Training dynamics

ReSPO separates early and stays ahead.

Normalized training score
Training curves for GRPO, GSPO, VESPO, and ReSPO on Qwen3 1.7B and 30B-A3B across rollout reuse ratios 8, 16, and 32
ReSPO reaches the highest observed early peak in five of six settings and exceeds the strongest clipped baseline in late mean across all six.

Training-log summary

ReSPO leads in \(10\) of \(12\) settings.

Figure 2 summaries from the paper. The early score is the maximum through policy update \(256\); the late score is the policy-step-weighted mean over updates \(896\text{--}1024\). Bold marks the largest observed value in each row.

Early optimization

First \(256\) max

Maximum normalized training score through policy update 256.
Model\(N\)GRPOGSPOVESPOReSPO
Qwen3-1.7B\(8\)\(0.169\)\(0.158\)\(0.161\)\(0.192\)
\(16\)\(0.156\)\(0.127\)\(0.151\)\(0.166\)
\(32\)\(0.118\)\(0.106\)\(0.127\)\(0.157\)
Qwen3-30B-A3B\(8\)\(0.478\)\(0.416\)\(0.414\)\(0.468\)
\(16\)\(0.403\)\(0.369\)\(0.386\)\(0.445\)
\(32\)\(0.362\)\(0.341\)\(0.364\)\(0.418\)
Late optimization

Last \(128\) mean

Mean normalized training score over policy updates 896 through 1024.
Model\(N\)GRPOGSPOVESPOReSPO
Qwen3-1.7B\(8\)\(0.213\)\(0.209\)\(0.275\)\(0.286\)
\(16\)\(0.210\)\(0.199\)\(0.268\)\(0.264\)
\(32\)\(0.197\)\(0.184\)\(0.230\)\(0.276\)
Qwen3-30B-A3B\(8\)\(0.530\)\(0.538\)\(0.538\)\(0.595\)
\(16\)\(0.498\)\(0.521\)\(0.531\)\(0.563\)
\(32\)\(0.504\)\(0.493\)\(0.523\)\(0.588\)

Held-out evaluation

ReSPO gains ground as rollout reuse increases.

At \(N=8\), ReSPO remains close to the strongest macro average. At \(N=16\) and \(N=32\), ReSPO has the highest observed average on both model scales. Results are final-checkpoint pass@1 percentages; the final column is the six-benchmark macro mean.

Rollout reuse

\(N=8\)

Comparable averagesCompetition gains are offset on the broader suite.
1.7B \(-0.5\,\mathrm{pp}\) 30B \(-0.2\,\mathrm{pp}\)
Final-checkpoint evaluation at rollout reuse N equals 8.
ModelMethodAIME25AIME24AMC23Olymp.MinervaMATH500Avg.
Qwen3-1.7BGRPO\(4.2\)\(7.5\)\(35.3\)\(28.4\)\(28.6\)\(66.5\)\(28.4\)
GSPO\(6.2\)\(7.7\)\(40.2\)\(30.4\)\(27.2\)\(65.8\)\(29.6\)
VESPO\(7.3\)\(12.7\)\(47.8\)\(38.8\)\(30.7\)\(74.2\)\(35.3\)
ReSPO\(9.0\)\(13.3\)\(48.4\)\(36.8\)\(28.9\)\(72.6\)\(34.8\)
Qwen3-30B-A3BGRPO\(15.2\)\(24.8\)\(73.8\)\(51.1\)\(43.6\)\(85.2\)\(49.0\)
GSPO\(18.1\)\(26.7\)\(71.3\)\(52.1\)\(43.9\)\(86.1\)\(49.7\)
VESPO\(15.4\)\(27.3\)\(70.0\)\(51.7\)\(45.4\)\(85.3\)\(49.2\)
ReSPO\(25.2\)\(30.6\)\(79.1\)\(42.7\)\(41.1\)\(78.3\)\(49.5\)
Rollout reuse

\(N=16\)

ReSPO leads both scalesThe advantage emerges as off-policy drift grows.
1.7B \(+0.8\,\mathrm{pp}\) 30B \(+3.2\,\mathrm{pp}\)
Final-checkpoint evaluation at rollout reuse N equals 16.
ModelMethodAIME25AIME24AMC23Olymp.MinervaMATH500Avg.
Qwen3-1.7BGRPO\(7.1\)\(8.3\)\(43.1\)\(30.3\)\(27.8\)\(66.3\)\(30.5\)
GSPO\(4.6\)\(7.5\)\(41.9\)\(31.9\)\(28.3\)\(67.2\)\(30.2\)
VESPO\(6.5\)\(10.6\)\(46.4\)\(36.4\)\(30.1\)\(71.2\)\(33.5\)
ReSPO\(9.4\)\(12.3\)\(46.7\)\(35.8\)\(29.9\)\(71.7\)\(34.3\)
Qwen3-30B-A3BGRPO\(13.8\)\(24.6\)\(73.3\)\(49.0\)\(42.4\)\(83.1\)\(47.7\)
GSPO\(15.6\)\(26.5\)\(70.0\)\(49.3\)\(45.2\)\(84.8\)\(48.6\)
VESPO\(15.2\)\(30.0\)\(70.0\)\(51.5\)\(43.6\)\(85.4\)\(49.3\)
ReSPO\(21.3\)\(32.1\)\(79.2\)\(51.2\)\(45.7\)\(85.7\)\(52.5\)
Rollout reuse

\(N=32\)

ReSPO leads at maximum reuseThe high-reuse advantage remains on both model scales.
1.7B \(+1.3\,\mathrm{pp}\) 30B \(+1.9\,\mathrm{pp}\)
Final-checkpoint evaluation at rollout reuse N equals 32.
ModelMethodAIME25AIME24AMC23Olymp.MinervaMATH500Avg.
Qwen3-1.7BGRPO\(5.2\)\(4.0\)\(38.3\)\(28.9\)\(27.8\)\(65.3\)\(28.3\)
GSPO\(4.6\)\(5.4\)\(38.1\)\(29.9\)\(29.9\)\(66.3\)\(29.0\)
VESPO\(8.1\)\(13.1\)\(45.6\)\(34.1\)\(28.2\)\(69.8\)\(33.2\)
ReSPO\(8.8\)\(11.7\)\(48.8\)\(36.7\)\(28.0\)\(72.9\)\(34.5\)
Qwen3-30B-A3BGRPO\(12.3\)\(31.3\)\(71.7\)\(50.3\)\(44.4\)\(85.0\)\(49.2\)
GSPO\(14.6\)\(23.1\)\(66.4\)\(47.5\)\(42.3\)\(83.5\)\(46.2\)
VESPO\(17.7\)\(27.7\)\(70.9\)\(50.0\)\(44.4\)\(85.2\)\(49.3\)
ReSPO\(20.2\)\(26.9\)\(79.5\)\(51.0\)\(45.1\)\(84.4\)\(51.2\)

04 / Citation

Build on ReSPO.

Read the full derivation, implementation details, and complete evaluation protocol in the paper.

BibTeX
@misc{chen2026respo,
  title  = {ReSPO: Reshaped Sequence Policy Optimization
            for Gradient Starvation in Off-Policy Learning},
  author = {Chen, Yihang and Ban, Yuanhao and Hsieh, Cho-Jui},
  year   = {2026},
  note   = {Preprint}
}