How far can we go with ImageNet for Text-to-Image generation?
2026-09-07 12:00Models🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
4mentions
SummaryAI generated
A recent study challenges the traditional paradigm in text generation that suggests “the larger the data, the better”. It demonstrates that using only the enhanced ImageNet dataset can achieve performance comparable to that of the FLUX model. This approach scored 5 points higher than SD3 in the GenEval evaluation and 12 points higher than SDXL in the DPGBench evaluation. At the same time, the amount of images required for training is only one-thousandth of the original approach, and the number of parameters has been reduced to one-thirtieth to one-tenth of the original amount. Since ImageNet data is widely available and standardized training requires only 500 hours of H100 computing power, this study provides a new way to improve reproducibility.