Research Article
OmniPhys: A Unified Multimodal Benchmark for Physics Understanding and Generation from Chinese Educational Corpora
Synthesis: Chen et al. (2026) introduce OmniPhys, a large-scale multimodal benchmark for physics understanding and reasoning built from Chinese educational corpora, spanning middle-school through university-level problems. The benchmark comprises 15,246 questions and 19,850 images with fine-grained annotations supporting analysis of reasoning processes and knowledge usage, and — unusually — systematically evaluates multimodal output generation, including models' ability to synthesize structured physics diagrams. Extensive evaluations reveal critical gaps in current multimodal LLMs, especially in complex reasoning and visual generation, positioning OmniPhys as a foundational resource for advancing multimodal intelligence in physics and for informing automated physics Assessment.
Why physics needs a unified multimodal benchmark
Multimodal large language models (MLLMs) have demonstrated strong abilities across visual and textual reasoning tasks, but their development in the physics domain is hindered by the lack of a comprehensive benchmark. Physics is a prototypical multimodal arena, demanding rigorous integration of textual descriptions, visual diagrams, and symbolic logic for accurate reasoning. Existing physics datasets rarely satisfy three critical criteria simultaneously: (1) cross-stage knowledge fusion spanning middle school to university; (2) multimodal input comprehension requiring interpretation of complex textual and visual cues; and (3) multimodal output generation assessing the model's ability to actively synthesize diagrams rather than merely select options. OmniPhys addresses this gap by unifying all three in a single benchmark.
Benchmark design
OmniPhys is a Chinese benchmark designed to assess physics mastery from secondary education to university levels, covering five major physical disciplines including mechanics, electromagnetism, and optics. All questions are curated from contemporary examination papers and authoritative textbooks, undergoing a strict multi-stage filtering process to guarantee difficulty and pedagogical validity. The benchmark includes:
- 15,246 questions and 19,850 images, with detailed annotations supporting fine-grained analysis of reasoning processes and knowledge usage.
- A multimodal output generation subset that assesses MLLMs' capabilities in physics diagram understanding and editing — a fundamental component of authentic physics problem solving that most benchmarks omit.
- Coverage spanning question types from multiple-choice to open-ended problem solving, grounded in authentic assessment material.
Evaluation findings
Comprehensive baseline evaluations reveal that despite recent advances, current multimodal LLMs exhibit significant capability gaps in the physics domain, especially in complex reasoning and visual generation. The multimodal output tasks — where models must synthesize or edit structured physics diagrams — proved particularly challenging, underscoring that generating authentic physics representations remains an open problem for MLLMs. These findings have direct implications for whether generative AI systems can serve as reliable partners in physics learning and automated assessment, connecting to broader knowledge base evidence that AI systems still struggle with the specialized, multimodal, and diagram-heavy tasks characteristic of authentic STEM assessment.
Implications for physics education and AI
OmniPhys matters to the knowledge base for three reasons. First, it extends the physics-education evidence base on AI capability with a large, authentic, Chinese-educational-corpus benchmark — complementing studies of AI performance on physics problems and LLM support for computational thinking in physics. Second, its emphasis on diagram generation connects to multimodal learning and authentic physics problem solving, where the ability to construct representations is as important as selecting answers. Third, its finding that MLLMs struggle on complex reasoning and visual generation informs realistic expectations for tutoring and automated assessment in physics, supporting the knowledge base's recurring theme that AI excels at routine tasks but underperforms on the authentic, high-level reasoning that defines deep disciplinary learning.
Connected Concepts
- Physics Education
- Multimodal
- LLM
- Generative AI
- Benchmark
- Assessment
- Automated Assessment
- STEM Education
- Intelligent Tutoring
- Learning Gains
- Computational Thinking
- AI Ed Evaluation
Connected Articles
- Probing AI Generated Physics Solutions 2026 — Probing AI-generated physics solutions and preparing students to critique them
- LLM Computational Thinking Physics 2026 — LLM support for computational thinking in physics
- Hashmi Socratic Physics Chatbot 2025 — Socratic physics chatbot
- Physics Chatbot Epistemological Beliefs 2026 — Physics chatbot and epistemological beliefs
- AI Grading Handwritten Physics 2026 — Large-scale AI grading of handwritten physics assessments
- GenAI Oop Programming Assessments 2026 — GenAI performance on authentic introductory OOP assessments
- LLM Formative Feedback Systematic Review 2026 — Systematic review of LLM-based formative feedback
- Assessment Latent Structure Human LLM 2026 — Assessment instruments for humans and LLMs
- Syal Multimodal Dialogue STEM 2026 — Multimodal dialogue in STEM
- Evaluation Age AI Output Evidence 2026 — Evaluation in the age of AI: output as evidence of learning
Citation
Chen, H., Lin, Y., Yushanjiang, N., Lin, X., & Zhang, M. (2026). OmniPhys: A unified multimodal benchmark for physics understanding and generation from Chinese educational corpora. arXiv:2608.25398.