SCAPES: Semantically Conditioned Autoregressive Prior for Environmental Sounds
2026-09-07 12:00Models🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
1mentions
SummaryAI generated
On September 7, 2026, arXiv released the SCAPES model, a semantic conditional autoregressive prior model for environmental sound generation. This model operates based on the continuous latent manifold of neural audio encoders, using a segmented strategy combined with continuous normalization and flow matching techniques. Experiments show that training high-fidelity, long-term stable, and semantically consistent sound generation instances (with approximately 36 million parameters) on a limited unselected dataset can be achieved with just a consumer-grade graphics card, and it supports smooth semantic interpolation. Currently, the code, pre-trained weights, and interactive demonstrations of this model are available publicly.