---
source_url: https://arxiv.org/abs/2603.00925
ingested: 2026-05-07
sha256: 5136f56aef80b903b3bb01c7ed5abb9a44abaf1e91701f66a36b711ec98accd5
---
# The Aftermath of DrawEduMath: Vision Language Models Underperform with Struggling Students and Misdiagnose Errors

**Authors:** Li Lucy, Albert Zhang, Nathan Anderson, Ryan Knight, Kyle Lo  
**arXiv:** 2603.00925  
**Submitted:** 1 March 2026  
**Length:** 15 pages, 10 figures  
**Subjects:** Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Computers and Society (cs.CY)

## Abstract

Effective mathematics education requires identifying and responding to students' mistakes. For AI to support pedagogical applications, models must perform well across different levels of student proficiency. Our work provides an extensive, year-long snapshot of how 11 vision-language models (VLMs) perform on DrawEduMath, a QA benchmark involving real students' handwritten, hand-drawn responses to math problems. We find that models' weaknesses concentrate on a core component of math education: student error. All evaluated VLMs underperform when describing work from students who require more pedagogical help, and across all QA, they struggle the most on questions related to assessing student error. Thus, while VLMs may be optimized to be math problem solving experts, our results suggest that they require alternative development incentives to adequately support educational use cases.

## Core Findings

- **Benchmark:** DrawEduMath — a QA dataset using real students' handwritten, hand-drawn math responses
- **Scope:** Year-long evaluation of 11 vision-language models (VLMs)
- **Critical Gap:** All VLMs perform worse on work from struggling students (those requiring more pedagogical help)
- **Weakest Area:** Models struggle most on questions assessing student error — the core of effective math pedagogy
- **Implication:** Current VLM optimization for math problem-solving expertise is insufficient for educational applications; alternative development incentives are needed to support pedagogical use cases.
