AuraTracer智迹闻
中文

EVENT DOSSIER

MCPO: Modality-Contrastive Preference Optimization for Multimodal Chain-of-Thought Compression

2026-09-07 12:00 Models 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
1mentions
SummaryAI generated

On September 7, 2026, arXiv published the paper MCPO (Modality-Contrastive Preference Optimization), proposing an efficient multi-modal thought chain compression method that requires fewer than 900 training samples. This method uses two-stage optimization: in the compression stage, a pruning algorithm based on normalized cross-modal mutual information (NCMI) is introduced to automatically eliminate visually irrelevant steps by comparing differences in graph-based and graph-free context reasoning; in the alignment stage, supervised fine-tuning is performed first, followed by asymmetric multi-modal length control preference loss optimization, using a non-linear odds ratio formula to strengthen the length constraints on preferred trajectories and maintain modality consistency. Experiments show that this method can reduce the length of thought chains in mainstream base models such as Qwen3-VL-Thinking by up to 69.5%, achieving an end-to-end reasoning acceleration of up to 3.34 times, while maintaining the original accuracy.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
Qwen3-VL-Thinking

SignalsSIGNALS

Keyword heat
  • Qwen3-VL-Thinking1

All reports (1)SOURCES

A arXiv cs.AI en 2026-09-07 12:00

MCPO: Modality-Contrastive Preference Optimization for Multimodal Chain-of-Thought Compression

提出 Modality-Contrastive Preference Optimization (MCPO),一种仅需少于 900 个训练样本的高效两阶段多模态思维链压缩方法。该方法在压缩阶段引入基于归一化跨模态互信息 (NCMI) 的剪枝算法,通过对比有图与无图上下文推理差异自动剔除视觉无关步骤;在对齐阶段先进行监督微调,再采用非对称多模态长度控制偏好损失优化,利用非线性几率比公式强化优选轨迹的长度约束并维持模态一致性。实验表明,该方法可将 Qwen3-VL-Thinking 等主流基模型的思维链长度减少最多 69.5%,实现最高 3.34 倍端到端推理加速,同时保持原始准确率。