MultihopSpatial: Multi-hop Compositional Spatial Reasoning Benchmark for Vision-Language Model
2026-09-07 12:00Models🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
3mentions
SummaryAI generated
On September 7, 2026, arXiv published the benchmark for visual language models (VLM) called MultihopSpatial. This research aims to address the issue that existing VLMs ignore multi-step combined reasoning and precise visual localization. The main contributions include: constructing a comprehensive multi-step combined spatial reasoning benchmark covering complex queries of one to three steps; proposing the Acc@50IoU metric for synchronously evaluating reasoning and visual localization capabilities; and releasing the dedicated large-scale training corpus for MultihopSpatial-Train. Extensive evaluations show that combined spatial reasoning remains a significant challenge, but by training the model through reinforcement learning on the corpus, its intrinsic spatial reasoning abilities and downstream embodied operation performance can be improved.