Optimizer Memory Schedules for Outscaling the Overtraining Axis
2026-09-07 12:00Science🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
4mentions
SummaryAI generated
For the scenario of long-term over-training of large models, the matrix conditional method (Muon, SOAP) and momentum scheduling method (ADANA) were compared. Experiments covered models with 51M to 253M parameters and over-training factors of 1x to 256x. The results showed that as the training duration increased, the optimal learning rate scheduling could reverse, the optimal weight decay coefficient scaled approximately with the square root, and long-term training generally preferred a fixed memory strategy. ADANA lagged behind Muon and SOAP in the initial stage after adjusting the fixed memory of AdamW, but as the training progressed, the benefits brought by log-time weight decay and momentum cooling continued to accumulate, eventually surpassing Muon and competing with SOAP under the highest over-training factor, leading to a significant advantage close to the theoretical prediction, far ahead of AdamW.