Research Article
Benchmarking Large Language Models for Diagnosing Students' Cognitive Skills from Handwritten Math Work
Synthesis: MathCog Benchmark (3,036 teacher-annotated diagnostic verdicts, 639 handwritten responses, 18 LLMs): all models severely underperform (macro F1 < 0.5) — over-attributing evidence, overthinking minimal cues, hallucinating nonexistent evidence (Hallucination Risk) — calling for evidence-aware architectures and teacher-in-the-loop designs (Knowledge Tracing, Multimodal AI, benchmark).
Summary
This paper introduces MathCog, a benchmark dataset of 3,036 teacher-annotated diagnostic verdicts across 639 student handwritten math responses to 110 problems. Evaluating 18 LLMs, the authors find that all models severely underperform (macro F1 < 0.5), with systematic failure modes: over-attributing evidential strength, overthinking minimal cues, and hallucinating nonexistent evidence. Performance degrades sharply when student evidence is vague or implicit. The study calls for evidence-aware architectures and teacher-in-the-loop designs.
Key Contributions
- MathCog Benchmark — First benchmark for cognitive skill diagnosis from handwritten math, grounded in TIMSS 2019 cognitive framework with Evident/Vague evidential strength labels
- Systematic Large Language Models (LLMs) Evaluation — 18 models spanning reasoning, multimodal, and text-only architectures, all showing F1 < 0.5
- Error Taxonomy — Five systematic error patterns identified: evidence misidentification, rubric misinterpretation, over-inference, inconsistency, and hallucination
- Evidential Calibration Metrics — Introduces OverAttr and FalseAttr to quantify models' tendency to over-claim evidential confidence
Core Findings
Universal Underperformance
No model achieves F1 ≥ 0.5. The best performers (GPT-4o-img at 0.448, DeepSeek-R1 at 0.442) still fail on nearly half of diagnostic decisions. Accuracy (mean 0.680) is misleading due to class imbalance — most student responses provide Evident Yes evidence, inflating accuracy.
Evidence Sensitivity Gap
All models perform worse when student evidence is Vague (implicit, incomplete, or context-dependent). Multimodal models show a larger Evident-to-Vague performance drop than text-only models, suggesting visual inputs may amplify over-interpretation rather than improve evidential calibration.
Systematic Error Patterns
- Evidence Over-Attribution (M = .580, SD = .120): Models frequently assign "Evident" to cases the teachers had labeled "Vague"
- Evidence False-Attribution (M = .585, SD = .145): Incorrect diagnoses are often accompanied by false claims of evidential confidence
- Hallucination: Models fabricate evidence quotes not present in student handwriting
- Over-inference: Drawing strong diagnostic conclusions from minimal or ambiguous cues
What this means for practice
- Software developers. Detect insufficient evidence inside the system instead of trusting the model's stated confidence: models over-attributed evidential strength (OverAttr M = .580, SD = .120) and did the same on incorrect diagnoses (FalseAttr M = .585, SD = .145).
- Software developers. Ship diagnosis as teacher-in-the-loop support rather than automatic verdicts: no model reached macro F1 ≥ 0.5, and the mean accuracy of 0.680 hides that failure behind class imbalance.
- Software developers. Do not treat multimodality or reasoning architectures as a calibration fix: multimodal models showed a larger Evident-to-Vague drop than text-only models, and the best performers (GPT-4o-img at 0.448, DeepSeek-R1 at 0.442) still carried high OverAttr and FalseAttr.
- Software developers. Route vague or implicit student work to human review, since performance degrades sharply when the evidence is Vague and models fabricated evidence quotes absent from the handwriting.
- Researchers. Anchor diagnosis in a published cognitive framework — MathCog's TIMSS 2019 grounding shifts handwritten math assessment from answer correctness to cognitive skill diagnosis — and report evidential calibration metrics like OverAttr and FalseAttr alongside accuracy, because the hallucinated-evidence and over-attribution patterns here extend The LLM Fallacy: Misattribution in AI-Assisted Cognitive Workflows.
Limitations
- Korean middle-school math only, with Korean-to-English machine translation that may introduce artifacts; generalizability to other languages, grade levels, and subjects is unknown
- 3,036 verdicts across 639 responses — moderate dataset size
- Only TIMSS "Knowing" and "Applying" domains covered; "Reasoning" skills excluded due to problem set characteristics
- Static benchmark; does not capture iterative diagnostic processes teachers use in practice
Citation
Kim, Y., Jin, H., Doh, H., Kim, E., Jung, D., Kim, S., Choi, K., Son, J., & Kim, J. (2025). Benchmarking Large Language Models for Diagnosing Students' Cognitive Skills from Handwritten Math Work.