Research Article
Generative AI and Mathematical Reasoning: Evaluating the Accuracy, Justification, and Proof Quality of AI-Generated Mathematics Solutions
Synthesis: Boafo, Antobre, and Appiah (2026) argue that a correct answer is not evidence of sound mathematical reasoning, and that generative AI evaluation in mathematics education has been organized around final-answer accuracy at the expense of the reasoning that produced it. Reviewing benchmarks, mathematics education studies, and error-detection work, they propose the Accuracy–Justification–Proof (AJP) framework, which scores a final answer (0–2), the validity of the reasoning steps (0–4), and — where a proof is required — the rigor of the argument (0–5) as three independent dimensions. They pair it with a twelve-code error taxonomy (calculation, algebraic, conceptual, theorem misuse, unsupported assertion, logical gap, circular reasoning, and others) and a classification matrix that separates a correct outcome with unreliable reasoning from a sound strategy with a computational slip. The review's central claim is that critical engagement with AI output is the educational goal: students must validate solutions against explicit warrants rather than accept fluency as proof.
Key Findings
- Accuracy is necessary but insufficient. A benchmark that scores only the final answer conflates an arithmetic slip with an invalid theorem, so a correct conclusion reached by invalid reasoning can pass as competent.
- Reasoning quality is a distinct construct. ReasonEval's reasoning-step indicators show that gains in final-answer accuracy do not necessarily correspond to gains in reasoning quality.
- Proof is not a longer explanation. A valid proof establishes a claim from accepted statements by valid methods; a fluent argument can still be circular, gap-ridden, or built on unstated assumptions.
- The AJP framework separates the three dimensions. Accuracy is scored 0–2, justification 0–4, and proof 0–5, and the three must be assessed independently because a correct answer can appear without the reasoning to support it.
- A finer error taxonomy is proposed. Twelve codes (E1 calculation through E12 false verification) distinguish computational, algebraic, conceptual, logical, and evidential failures that a single correct/incorrect label hides.
- The classification matrix exposes error type. Pairing accuracy, justification, and proof quality reveals whether a system's failures are primarily computational, conceptual, logical, or evidential.
- Educational value depends on student validation. Studies of proof construction show students may accept AI-generated proofs uncritically; validating a solution means connecting it to a theorem, definition, or condition rather than to its fluency.
Beyond Accuracy: Why Reasoning Needs Its Own Score
The review opens from a simple asymmetry: mathematical correctness has at least two dimensions, the correctness of the conclusion and the correctness of the reasoning that leads to it. The first can hold while the second fails — an answer may be numerically right yet rest on an invalid inference, a misapplied theorem, or a justification that does not establish the claim. This is sharpest in proof tasks, where a true statement is not a proof. Benchmarks such as MathEval cover a wide range of datasets, disciplines, and difficulty levels, but a single accuracy percentage collapses qualitatively different failures into one number. An answer with a sound overall method and one arithmetic mistake is scored like an answer built on an invalid theorem. The authors read the field's trajectory — from answer-only benchmarks toward reasoning-aware and error-detection evaluation — as evidence that process-level measurement is becoming the standard, and they build the review around organizing what those strands measure into one framework.
The AJP Framework and an Error Taxonomy
The proposed framework scores three dimensions separately. Accuracy asks whether the final answer is correct (0 incorrect, 1 partially correct, 2 correct). Justification asks whether the reasoning steps validly support the conclusion (0 absent through 4 fully valid). Proof quality, applied where a proof is required, asks whether the argument establishes the claim rigorously (0 absent through 5 rigorous and well structured). The point of separating them is that a correct result can be produced without the reasoning that would earn the justification score. The companion taxonomy is deliberately finer than correct/incorrect: it runs from calculation and algebraic errors through conceptual error, theorem misuse, unsupported assertion, logical gap, circular reasoning, incomplete proof, case omission, interpretation error, internal contradiction, and false verification — a response claiming a step was verified when its stated work does not support the claim. A classification matrix then maps the combinations: correct/valid/valid is a defensible solution; correct with invalid or incomplete reasoning is a correct outcome with unreliable reasoning; incorrect with mostly valid reasoning is a sound strategy carrying a substantive error.
What Students and Teachers Do with AI Proofs
The review draws a deliberate line between evidence about a model's capabilities and evidence about how learners use the model's outputs, warning that synthesizing the two as one object would be an unwarranted conflation. On the capability side, error-detection research treats diagnosing and correcting faulty reasoning as a skill distinct from producing correct answers. On the learner side, mathematics-education studies find that students' conceptions of proof and of AI shape whether they accept AI-generated proofs, that learners use AI feedback to revise claims for clarity and additional justification, and that AI-generated arguments can still contain faulty reasoning requiring critical examination. Validation, in this account, is what happens when a student connects their judgment about an AI solution to an explicit warrant — a theorem, definition, condition, or principle — rather than to the answer's surface plausibility. The framework is offered as a tool for evaluating systems and for designing classroom tasks in which students explain and justify the validity of AI-generated solutions.
What this means for practice
- Instructors. Grade mathematical work on reasoning as well as answers: ask students to justify each step and, on proof tasks, to name the warrant that licenses it, so a correct answer with a broken argument does not pass.
- Instructors. Turn AI-generated solutions into objects of critique rather than answer keys — present a fluent solution and have students locate the error type (misapplied theorem, logical gap, circular reasoning) before trusting it.
- Assessment designers. Do not let a single accuracy score stand in for mathematical competence; score final-answer correctness and reasoning validity separately, as the AJP framework does.
- Researchers. Use the error taxonomy and classification matrix to report error type, not just error rate, when evaluating AI systems on mathematics.
Limitations
- The paper is a critical narrative review, not a systematic one; the authors state it does not claim to identify and extract all existing research on the topic, so its evidence base is illustrative rather than exhaustive.
- The AJP framework and the twelve-code taxonomy are proposed for future empirical work and are not themselves validated against a corpus of scored solutions in this paper.
- The review explicitly separates capability evidence from learner evidence, so it does not report a single integrated effect of AI on mathematics learning; the claims about students rest on the cited proof-construction studies rather than on a pooled estimate.
- The sources span 2024–2026 and include benchmarks whose model versions and evaluation conditions differ, so scores are not directly comparable across the reviewed studies.
Citation
Boafo, R. O., Antobre, B. O., & Appiah, G. (2026). Generative AI and Mathematical Reasoning: Evaluating the Accuracy, Justification, and Proof Quality of AI-Generated Mathematics Solutions. EdArXiv Preprints.