🧠 AI Ed Wiki

Multimodal large language models (MLLMs) are increasingly used to interpret visualizations, yet most evaluations remain chart-centric and offer limited insight into scientific visualization (SciVis) literacy. This study benchmarks six MLLMs (three closed-source, three open-source) on a standardized SciVis literacy assessment — 49 items spanning 18 scientific visualizations, 8 techniques, and 11 task types — and compares model performance against data from 485 human participants.

Results show MLLMs do not exhibit uniform SciVis literacy. Gemini is the strongest model overall, exceeding the human mean on several evaluated subsets, while all open-source models fall below the human baseline. Performance is highly uneven across techniques and tasks: models do best on scientific illustration, search, and spatial understanding, but struggle on texture-based and integration-based visualizations and on fine-grained quantitative estimation. For STEM Education, this delineates where AI can responsibly support interpretation of scientific figures versus where it remains unreliable. The work contributes a reusable benchmark methodology and underscores that current Generative AI multimodal systems should not be treated as substitutes for human AI Literacy in reading scientific visualizations — a finding relevant to assessment design in Higher Ed and to Formative Assessment of visualization competence.

Connected Concepts

  • STEM Education
  • Generative AI
  • AI Literacy
  • Higher Ed
  • Formative Assessment
  • Connected Articles

  • Lata Ferpa Compliant Local LLM Autograder — LaTA: A Drop-in, FERPA-Compliant Local-LLM Autograder for Upper-Division STEM Coursework
  • Learning Engagement Assistant Lea — Learning Engagement Assistant (LEA): Cross-Course Scalability and Classroom Evaluation of an Agentic AI Tutoring System
  • LLM Misconception Difficulty Easy Trap — The Easy Trap: Why LLMs Underestimate Misconception-Driven Difficulty
  • AI Generated Feedback Higher Ed — Artificial intelligence and feedback in university education: effectiveness and student perceptions
  • LLM Psychometric Calibration Cdp — Aligning LLM-Simulated and Human Examinees for Psychometric Calibration: A Cognitive Diagnostic Profiling Approach
  • Authentic Products Authenticated Processes 2026 — From authentic products to authenticated processes: authentic assessment in AI-rich higher education
  • Citation

    Patrick Phuoc Do, Chau M. Ta, Chaoli Wang (2026). Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy. arXiv:2607.15176.