AuraTracer智迹闻
中文

EVENT DOSSIER

Motion-Omni: End-to-End Joint Speech and Full-Body Motion for Spoken Dialogue

The model was simultaneously released on arXiv and Hugging Face Papers on September 7, 2026, and the core metrics have been announced.

2026-09-07 12:00 Models 🔥 47.2 heat score
2sources
1days unfolding
47.2heat score
8mentions
SummaryAI generated

On September 7, 2026, the Motion-Omni model was simultaneously released by arXiv and Hugging Face Papers. This research proposes a end-to-end framework for jointly generating speech and full-body movements (including facial, hand, and upper and lower body parts), aiming to address the issue of lack of multimodal synergy in traditional cascading approaches. The model uses Qwen2.5-7B-Instruct as its backbone network, directly extracting motion features from the hidden state of speech, and forces joint training of the LLM, speech generator, and action generator by freezing the speech path. Its supervised data comes from a scalable pipeline containing 422,856 pairs of high-quality pairs (with a total of 1,402 hours). In public evaluations, the Motion-Omni-Q7 instance achieved a gap of less than 2% in reference-independent motion metrics compared to teacher-cascaded systems, with a 5.4-fold increase in inference speed (RTF=0.78), and surpassed all non-teacher-cascaded systems in beat relevance and diversity, with a word error rate of only 2.62%.…

Related eventsRELATED EVENTS
Quick factsQUICK FACTS
Qwen2.5-7B-InstructBackbone network
0.78 tokens per second (RTF)Inference speed
2.62%WER
Only 2% off the teacher cascade levelError margin
Key entitiesKEY ENTITIES
Chengqian MaHaoyu ZhangHugging FaceMotion-OmniQwen2.5-7B-InstructSwDA-500Wei TaoYiwen Guo

Event frameEVENT FRAME

Launch

政府 · 科研机构across 1 days

Status

The model was simultaneously released on arXiv and Hugging Face Papers on September 7, 2026, and the core metrics have been announced.

Coverage · reports per dayLANGUAGE SPLIT

Entity relations
Chengqian Ma × Haoyu Zh…1Chengqian Ma × Hugging …1Chengqian Ma × Wei Tao1Chengqian Ma × Yiwen Guo1Haoyu Zhang × Hugging F…1Haoyu Zhang × Wei Tao1

Integrated timelineUNIFIED TIMELINE

  1. 2026-09-07

    The Motion-Omni model was also released simultaneously.

    Motion-Omni was released on arXiv cs.CV and Hugging Face Papers, proposing an end-to-end framework for generating voice and full-body movements.

    2 reports

SignalsSIGNALS

Keyword heat
  • Hugging Face1
  • Chengqian Ma1
  • Wei Tao1
  • Haoyu Zhang1
  • Yiwen Guo1
  • Motion-Omni1
  • SwDA-5001
  • Qwen2.5-7B-Instruct1

All reports (2)SOURCES

H Hugging Face Papers en 2026-09-07 08:00

Paper page - Motion-Omni: End-to-End Joint Speech and Full-Body Motion for Spoken Dialogue

Motion-Omni proposes an end-to-end model that natively generates dialogue voices as well as facial, hand, upper and lower body movements. This model performs only 2% worse than teacher-cascades in terms of reference-independent motion metrics, with a reasoning speed of 0.78 tokens per second (RTF) and a WER of 2.62%. Its core mechanism involves jointly optimizing the LLM, voice generator, and motion generator to address the issue of inability to perform joint optimization in traditional cascade solutions. The supervised data comes from a model-independent pipeline containing 422,856 pairs of high-quality paired data, and the first public evaluation protocol has been released.

A arXiv cs.CV en 2026-09-07 12:00

Motion-Omni: End-to-End Joint Speech and Full-Body Motion for Spoken Dialogue

arXiv:2609.04250v1 发布 Motion-Omni,提出一种端到端联合生成语音与全身动作的框架。该框架利用 Qwen2.5-7B-Instruct 骨干网络,直接从产生语音的隐藏状态中输出面部表情及手、上半身和下半身的动作,并通过冻结语音路径强制 LLM、语音生成器和动作生成器协同训练以恢复对齐。监督数据来自包含 422,856 对高质量配对(1,402 小时)的可扩展管道,并发布了首个针对随机开放式全身口语对话的公开评估协议。Motion-Omni-Q7 实例在参考无关动作指标上与教师级联差距小于 2%,响应速度提升 5.4 倍(RTF=0.78),在节拍相关性和多样性上超越所有非教师级联系统,且词错误率低至 2.62%。