AuraTracer智迹闻
中文

EVENT DOSSIER

MedProb: Probing Internal Representations of Vision-Language Models for Medical Question Answering

2026-09-07 12:00 Science 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
1mentions
SummaryAI generated

The study proposes the MedProb lightweight detection framework, which can extract signals from the frozen visual-language model (VLM) internal representation to predict answers in multi-choice medical visual question-answering (Med-VQA) tasks without fine-tuning. This framework performs better than prompt methods and existing medical VLM and agent systems on the PATH-VQA, SLAKE, and VQA-RAD datasets. Experiments show that the detection method reduces the performance gap between small and large models; however, in the matching test of 14 pairs of general-purpose and medical VLM, medical adaptation did not consistently improve linear decodability. Additionally, free text generation results in an answer position deviation of up to 10 percentage points, while MedProb performs differently despite this effect. The study mainly focuses on multi-choice/multi-class settings and demonstrates that the detection method can be extended to open-ended generation tasks through rejection sampling scoring.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
MedProb

SignalsSIGNALS

Keyword heat
  • MedProb1

All reports (1)SOURCES

A arXiv cs.CL en 2026-09-07 12:00

MedProb: Probing Internal Representations of Vision-Language Models for Medical Question Answering

MedProb is a lightweight detection framework that can predict multi-choice medical visual question-answering (Med-VQA) answers from frozen visual-language model (VLM) representations without fine-tuning the model. On PATH-VQA, SLAKE, and VQA-RAD datasets, MedProb restored more answer-related signals than other prompt methods, performing better than medical VLMs and agent systems. The study also found that detection reduced the gap between small and large models compared to prompts; in a matching test of 14 pairs of general-purpose and medical VLMs, medical adaptation did not consistently improve linear decodability. Additionally, free text generation exhibited an answer position deviation of up to 10 percentage points, while MedProb also showed this deviation but in a different manner. The main results were applicable to multi-choice/multi-class Med-VQA settings and demonstrated that detection can be extended to open-ended generation through rejection sampling scoring.