VISTA: Dense Multi-Label Classroom Coding with Vision-Language Models
The VISTA model improved the restricted macroaccuracy to 80.1% in three non-participating chemistry lecture tests, outperforming the zero-sample variant’s 74.9%. The study proposed a VISTA baseline, which was implemented using MiniCPM-V-4.5 on dense sliding windows and fine-tuned by using a lightweight Multi-Layer Perception (MLP) to freeze the backbone. The prediction results were then pooled into COPUS grids every two minutes. The evaluation dataset was based on the Classroom Observation Protocol for Undergraduate STEM (COPUS) constructed by a consensus of five experts, including external validation vocabulary and code reliability metrics based on human evaluators. VISTA exhibited maximum residuals in teacher codes with visually similar visual features and rare audio-dependent codes. The study also identified three systematic failure modes (audible observability, fine-grained group work differentiation, long-tail recall), and released benchmark tools, prompts, and baseline code.