AuraTracer智迹闻
中文

EVENT DOSSIER

Towards Understanding Pause Token Fine-Tuning Dynamics: A Mode Retention Perspective

2026-09-07 12:00 Models 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
2mentions
SummaryAI generated

The study proposes the Masked Boundary Pause (MBP) rule, which intervenes in the training process by placing pause tokens with masked loss masks at the boundaries of inference steps. This strategy was tested on 1B-8B models such as Qwen and Llama, resulting in a 6-point improvement in mathematical task performance and a 2.5-point improvement in code tasks, while maintaining general language understanding capabilities. Experiments show that MBP optimizes training dynamics through a pattern retention mechanism: in synthetic continuous learning tasks, masked pause tokens only rewrite approximately 4 times the previously learned distribution when matching the final fitness; in mathematical inference tasks, boundary-neighboring tokens encode more information about downstream steps. Additionally, this pattern retention strategy extends benefits to GRPO, indicating that pause tokens are tools for intervening in training dynamics rather than merely computational tools during inference.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
LlamaQwen

Coverage · reports per dayLANGUAGE SPLIT

Entity relations
Llama × Qwen1

SignalsSIGNALS

Keyword heat
  • Qwen1
  • Llama1

All reports (1)SOURCES

A arXiv cs.AI en 2026-09-07 12:00

Towards Understanding Pause Token Fine-Tuning Dynamics: A Mode Retention Perspective

一项关于暂停令牌微调动态的研究提出 Masked Boundary Pause(MBP)规则,通过在推理步骤边界放置且损失掩码的暂停令牌来干预训练过程。该策略在 1B-8B Qwen 和 Llama 模型上测试,数学任务提升达 6 分,代码任务提升达 2.5 分,同时保持通用语言理解能力。实验表明,MBP 通过模式保留机制优化了训练动态:在合成持续学习任务中,掩码暂停令牌以匹配最终适应度时仅重写约 4 倍于前已学分布;在数学推理探测任务中,边界相邻令牌编码了更多下游步骤信息。此外,该模式保持策略将收益扩展至 GRPO,表明暂停令牌是训练动态干预而非仅推理时的计算工具。