GSM8K-V: Can Vision Language Models Solve Grade School Math Word Problems in Visual Contexts
2026-09-07 12:00Science🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
1mentions
SummaryAI generated
On September 7, 2026, researchers released the GSM8K-V benchmark, aimed at evaluating the ability of visual language models to solve elementary math problems in image contexts. The benchmark included 1,319 high-quality samples with automated pipeline mapping and manual verification, requiring the model to extract quantities through visual perception and integrate implicit clues across scenarios for reasoning. Assessments of 34 visual language models revealed significant modal differences: text tasks had an accuracy of over 90%, while GSM8K-V achieved only 59%, far below the 91% accuracy of humans; models with enhanced visual math reasoning abilities showed no improvement on this benchmark, proving that their capabilities were independent. Error analysis indicated that the main bottleneck was implicit visual inference errors (IVIE), where the model failed to recover visually semantic information that was not explicitly stated. The relevant code and data have been made open source.