AuraTracer智迹闻
中文

EVENT DOSSIER

X-VC: Zero-shot Streaming Voice Conversion in Codec Space

2026-09-07 12:00 Models 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
2mentions
SummaryAI generated

On September 7, 2026, arXiv cs.AI released the X-VC system, which achieved zero-sample streaming speech conversion in the latent space of pre-trained neural codecs. The system uses a dual-condition acoustic converter to jointly model the acoustic conditions at the frame level of the source codecs’ latent frames and the target reference speech, and injects speaker information through adaptive normalization. To reduce the mismatch between training and inference, the research team used generated paired data and a role assignment strategy combining standard, reconstruction, and reverse modes for training; streaming inference employed a block-based reasoning and overlapping smoothing scheme aligned with the segmented training paradigm of codecs. In the Seed-TTS-Eval experiment, X-VC achieved the best streaming word error rate (WER) in both English and Chinese environments, exhibited high speaker similarity across language settings, and had significantly lower offline real-time performance compared to the baseline system.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
Seed-TTS-EvalX-VC

Event frameEVENT FRAME

Launch

arXiv:2604.12456v3 X-VC 提出零样本流式语音转换系统

Coverage · reports per dayLANGUAGE SPLIT

Entity relations
Seed-TTS-Eval × X-VC1

SignalsSIGNALS

Keyword heat
  • X-VC1
  • Seed-TTS-Eval1

All reports (1)SOURCES

A arXiv cs.AI en 2026-09-07 12:00

X-VC: Zero-shot Streaming Voice Conversion in Codec Space

X-VC 系统实现了在预训练神经编解码器潜空间中进行的一步式零样本流式语音转换。该系统采用双条件声学转换器,联合建模源编解码潜帧与源自目标参考语音的帧级声学条件,并通过自适应归一化注入目标说话人信息;为降低训练与推理不匹配,使用生成配对数据及结合标准、重建和反向模式的角色分配策略进行训练;流式推理则采用与编解码器分段训练范式对齐的分块推理及重叠平滑方案。在 Seed-TTS-Eval 实验上,X-VC 在英语和中文中均取得最佳流式 WER,在同语言与跨语言设置下表现出强说话人相似度,且离线实时因子显著低于对比基线。