AuraTracer智迹闻
中文

EVENT DOSSIER

Sequential Beats Joint: On the Interplay between On-Policy Distillation and RLVR

2026-09-07 12:00 Models 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
0mentions
SummaryAI generated

On September 7, 2026, arXiv cs.AI published the paper “Sequential Beats Joint”, proposing a two-stage model training method: first, policy-based distillation (OPD), and then reinforcement learning (RL). Experiments show that the ‘OPD-then-RL’ approach consistently outperforms pure OPD, pure RLVR, and all existing joint baselines in logical and mathematical reasoning benchmarks. Behavioral analysis revealed that OPD expands students’ coverage of teacher-supported solutions, while RL performs fine-tuning within this range; optimizing both signals simultaneously leads to interference. Additionally, experiments confirmed that OPD validation scores are the key signal for switching to RL, and OPD is more suitable than supervised fine-tuning (SFT) as a cold start for RL. This study established an effective approach to transform two entangled training signals into complementary stages.

Related eventsRELATED EVENTS

All reports (1)SOURCES

A arXiv cs.AI en 2026-09-07 12:00

Sequential Beats Joint: On the Interplay between On-Policy Distillation and RLVR

本文提出一种两阶段训练方案,即先进行基于策略蒸馏(OPD)再进行强化学习(RL),该方案在逻辑与数学推理基准上持续优于纯 OPD、纯 RLVR 及所有现有联合基线。研究通过 pass@$k$ 行为、学习动态和参数更新提供了系统性解释:OPD 扩展了学生对教师支持解的覆盖范围,而 RL 在此范围内进行精细优化;联合优化两种信号则导致相互干扰。实验发现,OPD 验证分数是切换至 RL 的关键信号,且 OPD 比监督微调(SFT)更适合作为 RL 的冷启动。这些结果确立了"OPD-then-RL"作为一种简单而强大的方法,将两种纠缠的信号转化为互补的训练阶段。