AuraTracer智迹闻
中文

EVENT DOSSIER

How far can we go with ImageNet for Text-to-Image generation?

2026-09-07 12:00 Models 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
4mentions
SummaryAI generated

A recent study challenges the traditional paradigm in text generation that suggests “the larger the data, the better”. It demonstrates that using only the enhanced ImageNet dataset can achieve performance comparable to that of the FLUX model. This approach scored 5 points higher than SD3 in the GenEval evaluation and 12 points higher than SDXL in the DPGBench evaluation. At the same time, the amount of images required for training is only one-thousandth of the original approach, and the number of parameters has been reduced to one-thirtieth to one-tenth of the original amount. Since ImageNet data is widely available and standardized training requires only 500 hours of H100 computing power, this study provides a new way to improve reproducibility.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
FLUXImageNetSD3SDXL

Coverage · reports per dayLANGUAGE SPLIT

Entity relations
FLUX × ImageNet1FLUX × SD31FLUX × SDXL1ImageNet × SD31ImageNet × SDXL1SD3 × SDXL1

SignalsSIGNALS

Keyword heat
  • ImageNet1
  • FLUX1
  • SD31
  • SDXL1

All reports (1)SOURCES

A arXiv cs.CV en 2026-09-07 12:00

How far can we go with ImageNet for Text-to-Image generation?

一项新研究挑战了“数据量越大越好”的文本生成范式,证明仅使用增强后的 ImageNet 数据集即可达到 FLUX 模型性能。该方案在 GenEval 上比 SD3 高 5 分,在 DPGBench 上比 SDXL 高 12 分,同时训练图像用量仅为原来的千分之一,参数量减少 3 至 10 倍。由于 ImageNet 广泛可用且标准化训练仅需 500 小时 H100,该研究为可复现性研究提供了新路径。