📄 Research Article
Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy
Multimodal large language models (MLLMs) are increasingly used to interpret visualizations, yet most evaluations remain chart-centric and offer limited insight into scientific visualization (SciVis) literacy. This study benchmarks six MLLMs (three closed-source, three open-source) on a standardized SciVis literacy assessment — 49 items spanning 18 scientific visualizations, 8 techniques, and 11 task types — and compares model performance against data from 485 human participants.
Results show MLLMs do not exhibit uniform SciVis literacy. Gemini is the strongest model overall, exceeding the human mean on several evaluated subsets, while all open-source models fall below the human baseline. Performance is highly uneven across techniques and tasks: models do best on scientific illustration, search, and spatial understanding, but struggle on texture-based and integration-based visualizations and on fine-grained quantitative estimation. For STEM Education, this delineates where AI can responsibly support interpretation of scientific figures versus where it remains unreliable. The work contributes a reusable benchmark methodology and underscores that current Generative AI multimodal systems should not be treated as substitutes for human AI Literacy in reading scientific visualizations — a finding relevant to assessment design in Higher Ed and to Formative Assessment of visualization competence.
Connected Concepts
Connected Articles
Citation
Patrick Phuoc Do, Chau M. Ta, Chaoli Wang (2026). Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy. arXiv:2607.15176.