AuraTracer智迹闻
中文

EVENT DOSSIER

Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning

2026-08-24 08:00 Models 🔥 28.9 heat score
1sources
1days unfolding
28.9heat score
1mentions
SummaryAI generated

Apple’s machine learning research team released new research findings titled “Beyond Visual CoT” on August 24, 2026. This research proposes an internalized visual thinking mechanism aimed at achieving active video reasoning. Unlike traditional methods, this mechanism does not rely on external visual cues but drives the model’s ability to actively analyze and reason about video content through internalized visual thinking processes.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
Multimodal large language models

SignalsSIGNALS

Keyword heat
  • Multimodal large language models1

All reports (1)SOURCES

A Apple ML Research en 2026-08-24 08:00

Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning

The multimodal large language model uses Visual CoT in reasoning across space, time, and embodied environments. While it provides an intuitive forward-looking mechanism, it introduces significant reasoning overhead, which is particularly unfavorable for active video reasoning. The study proposes the Internalized Visual Thinking (IVT) framework, which aims to enable the model to learn visual thinking during training and perform reasoning directly during the reasoning phase. This framework is a post-training method that jointly optimizes text prediction and……