Breaking the Black Box Conjecture: The Path to True Spatial Intelligence in Large Models
# Breaking the Black Box Hypothesis: The Path for Large Models to Achieve True “Spatial Intelligence” Perform a brain CT scan on the model to cost-effectively reshape real spatial intelligence. In their paper “SpatialSV: Internalizing Interpretable 3D Spatial Awareness in MLLMs via Task-Oriented Visual Supervision”, the team from Sun Yat-sen University revealed the blind spots in the model’s spatial cognition by using “CT scans” of its internal representations. ## Key Discovery: The Hidden Geometric Structures within the Model The team extracted intermediate hidden layer features from models like LLaVA during reasoning and generation, and rendered them as 3D spatial point clouds. The truth became clear instantly: in LLaVA’s “mind,” key references (such as the small table in the example) were not modeled at all. ## Technical Approach: 2D-to-3D Geometric Reconstruction SpatialSV uses “intermediate layer extraction…”