AuraTracer智迹闻
中文

EVENT DOSSIER

Optimizer Memory Schedules for Outscaling the Overtraining Axis

2026-09-07 12:00 Science 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
4mentions
SummaryAI generated

For the scenario of long-term over-training of large models, the matrix conditional method (Muon, SOAP) and momentum scheduling method (ADANA) were compared. Experiments covered models with 51M to 253M parameters and over-training factors of 1x to 256x. The results showed that as the training duration increased, the optimal learning rate scheduling could reverse, the optimal weight decay coefficient scaled approximately with the square root, and long-term training generally preferred a fixed memory strategy. ADANA lagged behind Muon and SOAP in the initial stage after adjusting the fixed memory of AdamW, but as the training progressed, the benefits brought by log-time weight decay and momentum cooling continued to accumulate, eventually surpassing Muon and competing with SOAP under the highest over-training factor, leading to a significant advantage close to the theoretical prediction, far ahead of AdamW.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
ADANAAdamWMuonSOAP

Coverage · reports per dayLANGUAGE SPLIT

Entity relations
ADANA × AdamW1ADANA × Muon1ADANA × SOAP1AdamW × Muon1AdamW × SOAP1Muon × SOAP1

SignalsSIGNALS

Keyword heat
  • ADANA1
  • AdamW1
  • Muon1
  • SOAP1

All reports (1)SOURCES

A arXiv cs.LG en 2026-09-07 12:00

Optimizer Memory Schedules for Outscaling the Overtraining Axis

研究团队发现优化器性能与最优超参数随训练时长显著变化。对比矩阵预条件方法(Muon、SOAP)及动量调度方法(ADANA)在 51M 至 253M 参数模型及 1x 至 256x 过训练因子下的表现,结果显示:首选学习率调度可能随过训练轴反转,最佳权重衰减系数约按 sqrt(OT) 缩放,长训练时长通常偏好长固定内存。ADANA 在调整 AdamW 固定内存后仍保持优势,其日志时间权重衰减与动量冷却带来的收益随训练增加而累积;经此处理,ADANA 以接近 DANA 理论预测的指数优势超越 AdamW。Muon 和 SOAP 在大部分测量范围内提供恒定的 token 效率优势,但 ADANA 初期落后于两者后差距缩小,并在最高过训练因子下超越 Muon 并与 SOAP 竞争。这些结果确立了训练时长作为优化器评估…