AuraTracer智迹闻
中文

EVENT DOSSIER

Extremely Sparse Supervision Incentivizes Reasoning Ability

2026-09-07 12:00 Models 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
3mentions
SummaryAI generated

Based on the traditional assumption that post-training large language model training relies on massive amounts of dense supervised tokens, the research team found in the on-policy distillation (OPD) setup based on the Qwen3 family that a very small proportion of generated tokens (only one or two per inference trajectory, accounting for 0.05% of the total) can effectively stimulate the model’s reasoning ability. Experiments show that this sparse supervision approach often matches or even surpasses the performance of full-token training in most cases. This phenomenon was observed in teacher - student configurations with nine different model sizes and was further confirmed in mathematical reasoning tasks, code-based reasoning, and reinforcement learning based on verifiable rewards (RLVR). The research results challenge the traditional view that post-training must rely on dense tokens, suggesting that extremely sparse supervision may be closer to the natural learning process.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
Llama modelsProximal Policy OptimizationQwen3 family

Coverage · reports per dayLANGUAGE SPLIT

Entity relations
Llama models × Proximal…1Llama models × Qwen3 fa…1Proximal Policy Optimiz…1

SignalsSIGNALS

Keyword heat
  • Qwen3 family1
  • Llama models1
  • Proximal Policy Optimization1

All reports (1)SOURCES

A arXiv cs.AI en 2026-09-07 12:00

Extremely Sparse Supervision Incentivizes Reasoning Ability

大型语言模型通过有效后训练展现出日益增强的推理能力,但现有方法优化海量 Token 并隐含假设学习必须依赖 Token 密集。研究团队在基于 Qwen3 家族的 on-policy distillation (OPD) 设置中,发现极小比例的生成 Token(每推理轨迹仅一两个,占总量 0.05%)即可有效激励推理能力,该稀疏监督在多数情况下匹配或超越全 Token 训练效果。这一现象在九个不同模型规模的教师 - 学生配置及数学推理任务上一致观察到,并在编码推理、Llama 模型及基于可验证奖励 (RLVR) 的强化学习中得到进一步验证。研究挑战了后训练必须依赖 Token 密集的传统假设,并指出极稀疏监督可能更接近自然学习过程。