Research Article
ChatGPT-generated help produces learning gains equivalent to human tutor-authored help on mathematics skills
Synthesis: Pardos & Bhandari (2024) report a randomized efficacy study (N=274) comparing ChatGPT-generated hints to human tutor-authored hints and a no-help control across four mathematics subject areas. Only the ChatGPT condition produced statistically significant learning gains versus control, with no significant difference between ChatGPT and human-authored hints — and ChatGPT's 32% raw hint-error rate was reducible to near zero (algebra) or 13% (statistics) using the self-consistency hallucination-mitigation technique.
Authoring help content for educational technologies is labor-intensive and costly — a full-time employee may take a year to produce a textbook's worth of material. If LLMs can generate hints with sufficiently low error and sufficient learning efficacy, they could alleviate the most time- and cost-intensive component of tutoring-system development and enable scaling to many domains. This study evaluates whether ChatGPT-generated worked-solution hints can do so.
Study design
A 3 × 4 between-subjects design with 274 participants (Mechanical Turk workers, all MTurk Masters with at least a high-school degree). Participants were randomly assigned to one of three hint conditions — control (no hints), human tutor, or ChatGPT — paired with one of four mathematics subjects drawn from OpenStax CC BY textbooks: Elementary Algebra, Intermediate Algebra, College Algebra, or Statistics. Each participant completed a 3-item pre-test, a 5-item acquisition phase in the OATutor platform, and a 3-item post-test (same questions as pre-test). ChatGPT hints were single worked-solution hints generated by prompting ChatGPT 3.5 with each problem's text; human hints came from the existing OATutor content library authored by UC Berkeley undergraduates.
Key findings
RQ1 — Hint quality and hallucination mitigation
- 32% of ChatGPT-generated hints failed quality checks (75 problems; 24 disqualified) due to containing incorrect answers and/or incorrect solution steps. No hints contained inappropriate language or grammatical errors. Disqualification ranged from 25% (Elementary Algebra) to 47% (Intermediate Algebra).
- Self-consistency drastically reduced errors: prompting 10 times per problem and returning the modal answer reduced the error rate to near 0% for the three algebra subjects and 13% for statistics.
- Human inter-rater agreement was high (Fleiss' κ = 0.857–0.929, "almost perfect"). Manual quality checking averaged 37.6 seconds per hint.
RQ2 — Learning efficacy
- ChatGPT hints produced the largest learning gain: 17.00% (pre 43.51% → post 60.52%, p<0.001) — statistically significant versus the control's 1.85% gain (p=0.011).
- Human tutor hints produced an 11.62% gain (p=0.001), not statistically separable from the ChatGPT condition (p=0.416).
- The control produced a non-significant 1.85% gain (p=0.192), establishing that gains were attributable to the hint conditions rather than mere test-retest recall.
- ChatGPT gains were higher than human-authored gains in all four subjects, though not significantly separable.
- No significant time-on-task difference between the ChatGPT and human conditions (both higher than control).
What this means for practice
- Designers. Generate first-draft worked-solution hints with an Large Language Models (LLMs), then run self-consistency — prompt each problem repeatedly and serve the modal answer — before students see them: this took the raw 32% hint-error rate to near zero in the three algebra subjects and 13% in statistics.
- Designers. Keep a human quality-check step wherever near-zero error is not achievable; manual checking averaged 37.6 seconds per hint, cheap enough to run across a full content library.
- Instructors. Where verified error rates are not near zero, present AI help as coming from an "imperfect robot" or peer-like source so students evaluate it critically instead of treating it as authoritative.
- Designers. Treat LLM-authored hints as a candidate replacement for costly human authoring only after re-testing error rates in the target domain: disqualification ranged from 25% in Elementary Algebra to 47% in Intermediate Algebra even though generation cost roughly 1/20th of human authoring.
Limitations
- Participants were 274 crowdsourced Mechanical Turk workers (all MTurk Masters with at least a high-school degree), not in-situ secondary or post-secondary students.
- Problems containing graphical figures were excluded because the ChatGPT version available at the time could not accept image input, so the corpus covers text-only problems only.
- The study relied on a closed-source model (ChatGPT 3.5) whose weights are not public, leaving the generation pipeline unreproducible with open alternatives.
- Attrition was high at 30%, though roughly even across conditions (36–43 excluded participants per condition), and the scope was limited to secondary and early post-secondary mathematics.
Citation
Pardos, Z. A., & Bhandari, S. (2024). ChatGPT-generated help produces learning gains equivalent to human tutor-authored help on mathematics skills. PLOS ONE, 19(5), e0304013. https://doi.org/10.1371/journal.pone.0304013