AuraTracer智迹闻
中文

EVENT DOSSIER

Less Data, Faster Training: repeating smaller datasets speeds up learning via sampling biases

2026-09-07 12:00 Science 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
0mentions
SummaryAI generated

A recent AI study suggests that reusing smaller datasets can significantly accelerate model training. The study found that in terms of algorithm tasks, network architectures, and optimizer settings, reusing a small number of samples is more effective at saving computational resources than using a large single dataset. This phenomenon cannot be explained by existing theories; the speed improvement stems from the inter-layer growth effect caused by sampling bias, and it is even more pronounced in environments with small datasets. The study confirmed through theoretical analysis and multiple experimental interventions that in scenarios where data is scarce, the “small dataset plus reusing” strategy is not only a passive coping mechanism but also an advantageous inductive bias that can be actively utilized, especially in reasoning tasks.

Related eventsRELATED EVENTS

All reports (1)SOURCES

A arXiv cs.AI en 2026-09-07 12:00

Less Data, Faster Training: repeating smaller datasets speeds up learning via sampling biases

一项新研究提出,重复使用较小数据集可加速模型训练。该工作发现,在算法任务、架构及优化器上,重复少量样本比使用大样本能节省计算资源,且此现象无法用既有理论解释。研究认为,速度提升源于采样偏差带来的合适层间增长,该效应在小数据集下更为显著。作者提供了理论分析与多项干预的实证证据,表明在小数据稀缺时采用“小数据集加重复”不仅是被动策略,更可主动利用其作为优化中的有利归纳偏置,尤其在推理任务中效果明显。