Sequential Beats Joint: On the Interplay between On-Policy Distillation and RLVR
2026-09-07 12:00Models🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
0mentions
SummaryAI generated
On September 7, 2026, arXiv cs.AI published the paper “Sequential Beats Joint”, proposing a two-stage model training method: first, policy-based distillation (OPD), and then reinforcement learning (RL). Experiments show that the ‘OPD-then-RL’ approach consistently outperforms pure OPD, pure RLVR, and all existing joint baselines in logical and mathematical reasoning benchmarks. Behavioral analysis revealed that OPD expands students’ coverage of teacher-supported solutions, while RL performs fine-tuning within this range; optimizing both signals simultaneously leads to interference. Additionally, experiments confirmed that OPD validation scores are the key signal for switching to RL, and OPD is more suitable than supervised fine-tuning (SFT) as a cold start for RL. This study established an effective approach to transform two entangled training signals into complementary stages.