AuraTracer智迹闻
中文

EVENT DOSSIER

GSM8K-V: Can Vision Language Models Solve Grade School Math Word Problems in Visual Contexts

2026-09-07 12:00 Science 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
1mentions
SummaryAI generated

On September 7, 2026, researchers released the GSM8K-V benchmark, aimed at evaluating the ability of visual language models to solve elementary math problems in image contexts. The benchmark included 1,319 high-quality samples with automated pipeline mapping and manual verification, requiring the model to extract quantities through visual perception and integrate implicit clues across scenarios for reasoning. Assessments of 34 visual language models revealed significant modal differences: text tasks had an accuracy of over 90%, while GSM8K-V achieved only 59%, far below the 91% accuracy of humans; models with enhanced visual math reasoning abilities showed no improvement on this benchmark, proving that their capabilities were independent. Error analysis indicated that the main bottleneck was implicit visual inference errors (IVIE), where the model failed to recover visually semantic information that was not explicitly stated. The relevant code and data have been made open source.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
GSM8K-V

SignalsSIGNALS

Keyword heat
  • GSM8K-V1

All reports (1)SOURCES

A arXiv cs.AI en 2026-09-07 12:00

GSM8K-V: Can Vision Language Models Solve Grade School Math Word Problems in Visual Contexts

研究人员发布 GSM8K-V 基准测试,将 GSM8K 转化为多图像序列以评估视觉语言模型在图像语境中解决小学数学问题的能力。该基准包含 1,319 个经自动化管道映射及人工验证的高质量样本,要求模型通过视觉感知提取数量并整合跨场景的隐含线索进行推理。对 34 个视觉语言模型的评估显示显著模态差距:文本任务准确率超 90%,而 GSM8K-V 最高仅达 59%,远低于人类 91% 的准确率;增强视觉数学推理能力的模型在此基准上未见提升,证明其评估了独立能力。错误分析表明主要瓶颈在于隐含视觉推断错误(IVIE),即模型无法恢复未明确陈述的视觉语义。相关代码与数据已开源。