๐Ÿง  AI Ed Wiki

Pardos & Bhandari (2024) report a randomized efficacy study (N=274) comparing ChatGPT-generated hints to human tutor-authored hints and a no-help control across four mathematics subject areas. Only the ChatGPT condition produced statistically significant learning gains versus control, with no significant difference between ChatGPT and human-authored hints โ€” and ChatGPT's 32% raw hint-error rate was reducible to near zero (algebra) or 13% (statistics) using the self-consistency hallucination-mitigation technique.

Authoring help content for educational technologies is labor-intensive and costly โ€” a full-time employee may take a year to produce a textbook's worth of material. If LLMs can generate hints with sufficiently low error and sufficient learning efficacy, they could alleviate the most time- and cost-intensive component of tutoring-system development and enable scaling to many domains. This study evaluates whether ChatGPT-generated worked-solution hints can do so.

Study design

A 3 ร— 4 between-subjects design with 274 participants (Mechanical Turk workers, all MTurk Masters with at least a high-school degree). Participants were randomly assigned to one of three hint conditions โ€” control (no hints), human tutor, or ChatGPT โ€” paired with one of four mathematics subjects drawn from OpenStax CC BY textbooks: Elementary Algebra, Intermediate Algebra, College Algebra, or Statistics. Each participant completed a 3-item pre-test, a 5-item acquisition phase in the OATutor platform, and a 3-item post-test (same questions as pre-test). ChatGPT hints were single worked-solution hints generated by prompting ChatGPT 3.5 with each problem's text; human hints came from the existing OATutor content library authored by UC Berkeley undergraduates.

Key findings

RQ1 โ€” Hint quality and hallucination mitigation

  • 32% of ChatGPT-generated hints failed quality checks (75 problems; 24 disqualified) due to containing incorrect answers and/or incorrect solution steps. No hints contained inappropriate language or grammatical errors. Disqualification ranged from 25% (Elementary Algebra) to 47% (Intermediate Algebra).
  • Self-consistency drastically reduced errors: prompting 10 times per problem and returning the modal answer reduced the error rate to near 0% for the three algebra subjects and 13% for statistics.
  • Human inter-rater agreement was high (Fleiss' ฮบ = 0.857โ€“0.929, "almost perfect"). Manual quality checking averaged 37.6 seconds per hint.
  • RQ2 โ€” Learning efficacy

  • ChatGPT hints produced the largest learning gain: 17.00% (pre 43.51% โ†’ post 60.52%, p<0.001) โ€” statistically significant versus the control's 1.85% gain (p=0.011).
  • Human tutor hints produced an 11.62% gain (p=0.001), not statistically separable from the ChatGPT condition (p=0.416).
  • The control produced a non-significant 1.85% gain (p=0.192), establishing that gains were attributable to the hint conditions rather than mere test-retest recall.
  • ChatGPT gains were higher than human-authored gains in all four subjects, though not significantly separable.
  • No significant time-on-task difference between the ChatGPT and human conditions (both higher than control).
  • Implications for AI in education

    The findings suggest that LLM-generated worked solutions can be as effective as human-authored tutoring content while being produced in a fraction of the time (roughly 1/20th), opening the door to autonomous generation of effective mathematics tutoring content from arbitrary educational resources. However, the authors are explicit about caution: at a 32% raw error rate, ChatGPT should not be used to give feedback the way a teacher or TA would unless in a domain verified to have near-zero error. Where error mitigation cannot achieve near-zero rates, designers should frame LLM feedback as coming from an "imperfect robot" or peer-like source so students consider it critically. The error-reduction via self-consistency is central, and the findings ground the performance-vs-learning distinction by demonstrating genuine learning gains, not just performance.

    Connected Concepts

  • Generative AI
  • LLM
  • AI Tutoring
  • Intelligent Tutoring
  • Scaffolding
  • Math Education
  • Learning Gains
  • Hallucination Risk
  • Adaptive Learning
  • Feedback Loop
  • Connected Articles

  • Oatutor Open Source Adaptive Tutor 2023 โ€” OATutor: Open-Source Adaptive Tutoring System
  • AI Tutor Effectiveness Review โ€” AI Tutor Effectiveness Review
  • GenAI Performance Vs Learning โ€” Distinguishing Performance Gains from Learning
  • AI Generated Feedback Higher Ed โ€” AI Feedback in University Education
  • Generative AI Guardrails Harm Learning โ€” GenAI Without Guardrails Can Harm Learning
  • From Answer Generators To Reasoning Facilitators AI Tutors โ€” From Answer Generators to Reasoning Facilitators
  • Access Not Enough AI Tutoring 2026 โ€” Access Is Not Enough: AI Tutoring
  • Adaptive Pretesting Retention โ€” Adaptive Pretesting and Retention
  • Citation

    Pardos, Z. A., & Bhandari, S. (2024). ChatGPT-generated help produces learning gains equivalent to human tutor-authored help on mathematics skills. PLOS ONE, 19(5), e0304013. https://doi.org/10.1371/journal.pone.0304013