AuraTracer智迹闻
中文

EVENT DOSSIER

SCAPES: Semantically Conditioned Autoregressive Prior for Environmental Sounds

2026-09-07 12:00 Models 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
1mentions
SummaryAI generated

On September 7, 2026, arXiv released the SCAPES model, a semantic conditional autoregressive prior model for environmental sound generation. This model operates based on the continuous latent manifold of neural audio encoders, using a segmented strategy combined with continuous normalization and flow matching techniques. Experiments show that training high-fidelity, long-term stable, and semantically consistent sound generation instances (with approximately 36 million parameters) on a limited unselected dataset can be achieved with just a consumer-grade graphics card, and it supports smooth semantic interpolation. Currently, the code, pre-trained weights, and interactive demonstrations of this model are available publicly.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
SCAPES

Event frameEVENT FRAME

Launch

arXiv:2609.04634v1 SCAPES 发布语义条件自回归环境声音生成模型

SignalsSIGNALS

Keyword heat
  • SCAPES1

All reports (1)SOURCES

A arXiv cs.AI en 2026-09-07 12:00

SCAPES: Semantically Conditioned Autoregressive Prior for Environmental Sounds

arXiv:2609.04634v1 发布 SCAPES,一种用于环境声音生成的语义条件自回归先验模型。该轻量级模型基于神经音频编码器的连续潜在流形运行,通过分段策略结合连续归一化流与流匹配技术,在有限未筛选数据集上仅需单张消费级显卡即可训练。实验表明,3600 万参数实例收敛于约两倍源音频时长后,能生成高保真、长期稳定且语义一致的声音,并支持平滑语义插值。代码、预训练权重及交互式演示已公开。