X-VC: Zero-shot Streaming Voice Conversion in Codec Space
2026-09-07 12:00Models🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
2mentions
SummaryAI generated
On September 7, 2026, arXiv cs.AI released the X-VC system, which achieved zero-sample streaming speech conversion in the latent space of pre-trained neural codecs. The system uses a dual-condition acoustic converter to jointly model the acoustic conditions at the frame level of the source codecs’ latent frames and the target reference speech, and injects speaker information through adaptive normalization. To reduce the mismatch between training and inference, the research team used generated paired data and a role assignment strategy combining standard, reconstruction, and reverse modes for training; streaming inference employed a block-based reasoning and overlapping smoothing scheme aligned with the segmented training paradigm of codecs. In the Seed-TTS-Eval experiment, X-VC achieved the best streaming word error rate (WER) in both English and Chinese environments, exhibited high speaker similarity across language settings, and had significantly lower offline real-time performance compared to the baseline system.