AuraTracer智迹闻
中文

EVENT DOSSIER

Training-Free Speech-Centric Omni Understanding with Frozen VLMs

On September 7, 2026, the TFO framework paper was published on arXiv and Hugging Face. It can improve audio quality without training…

2026-09-07 12:00 Models 🔥 47.2 heat score
2sources
1days unfolding
47.2heat score
7mentions
SummaryAI generated

The researchers proposed a training-free framework called TFO (Training-Free Omni), aimed at converting frozen video-language models (VLM) into full-modal models for speech centers. This framework was published by Ankan Deria et al. in August 2026, without the need to modify the architecture or perform multi-modal re-alignment. TFO utilizes Whisper for confidence filtering and timestamped transcription, routing audio information through the existing language interfaces of the VLM while maintaining the visual path unchanged. In 56 benchmark tests and comparisons across 21 languages, this framework was competitive with native full-modal models in audio and video understanding, improving audio performance across all five model settings and achieving significant cross-language speech gains. Additionally, frozen VLM generally retained stronger image/video understanding, visual localization, encoding, mathematical reasoning, and medical question-answering capabilities than corresponding native full-modal checkpoints. The results indicate that powerful full-modal speech center understanding can usually be achieved through modular audio-to-language routing…

Related eventsRELATED EVENTS
Quick factsQUICK FACTS
56Number of benchmark tests
21Number of languages
Key entitiesKEY ENTITIES
Ankan DeriaHanoona RasheedHugging FaceMohamed Bin Zayed University of Artificial IntelligenceTFOWhisperXilin He

Event frameEVENT FRAME

Research

arXiv:2609.04242v1 Training-Free Omni (TFO) 提出无需训练即可增强语音能力的框架

Status

On September 7, 2026, the TFO framework paper was published on arXiv and Hugging Face. It can improve audio quality without training…

Coverage · reports per dayLANGUAGE SPLIT

Entity relations
TFO × Whisper1Ankan Deria × Hanoona R…1Ankan Deria × Hugging F…1Ankan Deria × Mohamed B…1Ankan Deria × Xilin He1Hanoona Rasheed × Huggi…1

Integrated timelineUNIFIED TIMELINE

  1. 2026-09-07

    Publication of the TFO framework paper

    Ankan Deria et al. published the TFO training-free framework on arXiv and Hugging Face. It uses Whisper to extract information and routes audio through the VLM language interface. The effectiveness was verified in 56 benchmark tests and 21 languages.

    2 reports

SignalsSIGNALS

Keyword heat
  • TFO1
  • Whisper1
  • Hugging Face1
  • Mohamed Bin Zayed University of Artificial Intelligence1
  • Ankan Deria1
  • Hanoona Rasheed1
  • Xilin He1

All reports (2)SOURCES

H Hugging Face Papers en 2026-09-07 08:00

Paper page - Training-Free Speech-Centric Omni Understanding with Frozen VLMs

Ankan Deria、Hanoona Rasheed、Xilin He 及 Fahad Shahbaz Khan、Salman Khan 提出了一种名为 Training-Free Omni(TFO)的框架,该框架无需架构修改或模态重对齐即可将冻结的多模态大模型(VLM)转化为语音中心的全模态理解模型。现有全模态模型通常依赖专用音频编码器及昂贵的音视频文本联合训练,导致能力耦合特定 VLM 主干并削弱原有视觉与推理能力。TFO 利用 Whisper 提取置信度过滤和时间戳转录本以解决此问题。该研究发表于 Hugging Face Daily Papers,论文 ID 为 2609.04242,于 2026 年 8 月 7 日发布。

A arXiv cs.CV en 2026-09-07 12:00

Training-Free Speech-Centric Omni Understanding with Frozen VLMs

研究人员提出名为 TFO 的无训练框架,将冻结的视频 - 语言模型(VLM)转化为语音中心的全模态模型,无需架构修改或多模态重对齐。该框架利用 Whisper 提取置信度过滤和时间戳转录本,通过 VLM 现有语言接口路由,同时保留视觉路径不变。在 56 个基准测试和 21 种语言的对比中,TFO 在音视频理解上与原生全模态模型具有竞争力,提升了所有五种模型设置下的音频单独性能,并实现了显著的跨语言语音增益。冻结 VLM 还普遍保留了比对应原生全模态检查点更强的图像/视频理解、视觉定位、编码、数学推理及医疗问答能力。这些结果表明,强大的语音中心全模态理解通常可通过模块化音频到语言路由获得,而非昂贵的特定骨干训练。