Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning
The multimodal large language model uses Visual CoT in reasoning across space, time, and embodied environments. While it provides an intuitive forward-looking mechanism, it introduces significant reasoning overhead, which is particularly unfavorable for active video reasoning. The study proposes the Internalized Visual Thinking (IVT) framework, which aims to enable the model to learn visual thinking during training and perform reasoning directly during the reasoning phase. This framework is a post-training method that jointly optimizes text prediction and……