Training-Free Speech-Centric Omni Understanding with Frozen VLMs
On September 7, 2026, the TFO framework paper was published on arXiv and Hugging Face. It can improve audio quality without training…
2026-09-07 12:00Models🔥 47.2 heat score
2sources
1days unfolding
47.2heat score
7mentions
SummaryAI generated
The researchers proposed a training-free framework called TFO (Training-Free Omni), aimed at converting frozen video-language models (VLM) into full-modal models for speech centers. This framework was published by Ankan Deria et al. in August 2026, without the need to modify the architecture or perform multi-modal re-alignment. TFO utilizes Whisper for confidence filtering and timestamped transcription, routing audio information through the existing language interfaces of the VLM while maintaining the visual path unchanged. In 56 benchmark tests and comparisons across 21 languages, this framework was competitive with native full-modal models in audio and video understanding, improving audio performance across all five model settings and achieving significant cross-language speech gains. Additionally, frozen VLM generally retained stronger image/video understanding, visual localization, encoding, mathematical reasoning, and medical question-answering capabilities than corresponding native full-modal checkpoints. The results indicate that powerful full-modal speech center understanding can usually be achieved through modular audio-to-language routing…
Ankan DeriaHanoona RasheedHugging FaceMohamed Bin Zayed University of Artificial IntelligenceTFOWhisperXilin He
Event frameEVENT FRAME
Research
arXiv:2609.04242v1Training-Free Omni (TFO)提出无需训练即可增强语音能力的框架
Status
On September 7, 2026, the TFO framework paper was published on arXiv and Hugging Face. It can improve audio quality without training…
Coverage · reports per dayLANGUAGE SPLIT
Entity relations
Integrated timelineUNIFIED TIMELINE
2026-09-07
Publication of the TFO framework paper
Ankan Deria et al. published the TFO training-free framework on arXiv and Hugging Face. It uses Whisper to extract information and routes audio through the VLM language interface. The effectiveness was verified in 56 benchmark tests and 21 languages.
2 reports
SignalsSIGNALS
Keyword heat
TFO1
Whisper1
Hugging Face1
Mohamed Bin Zayed University of Artificial Intelligence1