On September 7, 2026, arXiv released RoboSPA, a large-scale robotic operation dataset and benchmark designed to diagnose the embodied reasoning capabilities of visual-linguistic-action (VLA) models. The benchmark covers 10 task categories and 56 basic tasks, with each task having five difficulty levels, resulting in 280 variants involving 527,000 trajectories and various scenarios. RoboSPA focuses on two core aspects: fine spatial reasoning and long-term process planning, and introduces diagnostic metrics beyond binary success rates to evaluate model performance. Experimental results show that existing VLA models still face significant challenges in understanding complex spatial relationships, achieving precise low-level execution, and planning with high memory consumption.