MCPO: Modality-Contrastive Preference Optimization for Multimodal Chain-of-Thought Compression
2026-09-07 12:00Models🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
1mentions
SummaryAI generated
On September 7, 2026, arXiv published the paper MCPO (Modality-Contrastive Preference Optimization), proposing an efficient multi-modal thought chain compression method that requires fewer than 900 training samples. This method uses two-stage optimization: in the compression stage, a pruning algorithm based on normalized cross-modal mutual information (NCMI) is introduced to automatically eliminate visually irrelevant steps by comparing differences in graph-based and graph-free context reasoning; in the alignment stage, supervised fine-tuning is performed first, followed by asymmetric multi-modal length control preference loss optimization, using a non-linear odds ratio formula to strengthen the length constraints on preferred trajectories and maintain modality consistency. Experiments show that this method can reduce the length of thought chains in mainstream base models such as Qwen3-VL-Thinking by up to 69.5%, achieving an end-to-end reasoning acceleration of up to 3.34 times, while maintaining the original accuracy.