Patrick Phuoc Do, Chau M. Ta, Chaoli Wang (2026) โ arXiv preprint. arXiv:2607.15176.
๐ Full text (arXiv)
Multimodal large language models (MLLMs) are increasingly used to interpret visualizations, yet most evaluations remain chart-centric and offer limited insight into scientific visualization (SciVis) literacy. This study benchmarks six MLLMs (three closed-source, three open-source) on a standardized SciVis literacy assessment โ 49 items spanning 18 scientific visualizations, 8 techniques, and 11 task types โ and compares model performance against data from 485 human participants.
Results show MLLMs do not exhibit uniform SciVis literacy. Gemini is the strongest model overall, exceeding the human mean on several evaluated subsets, while all open-source models fall below the human baseline. Performance is highly uneven across techniques and tasks: models do best on scientific illustration, search, and spatial understanding, but struggle on texture-based and integration-based visualizations and on fine-grained quantitative estimation. For stem-education, this delineates where AI can responsibly support interpretation of scientific figures versus where it remains unreliable. The work contributes a reusable benchmark methodology and underscores that current generative-ai multimodal systems should not be treated as substitutes for human ai-literacy in reading scientific visualizations โ a finding relevant to assessment design in higher-ed and to formative-assessment of visualization competence.
Related Pages
- ai-literacy โ benchmarks MLLMs on scientific visualization literacy vs 485 humans
- stem-education โ implications for AI support in interpreting scientific figures
- benchmark โ reusable SciVis-literacy benchmark methodology (49 items, 18 viz)
- generative-ai โ evaluates multimodal generative-AI systems on literacy tasks
- higher-ed โ undergraduate science-literacy assessment context
- formative-assessment โ assessment instrument for visualization competence