๐ Research Article
ChatGPT-generated help produces learning gains equivalent to human tutor-authored help on mathematics skills
Pardos & Bhandari (2024) report a randomized efficacy study (N=274) comparing ChatGPT-generated hints to human tutor-authored hints and a no-help control across four mathematics subject areas. Only the ChatGPT condition produced statistically significant learning gains versus control, with no significant difference between ChatGPT and human-authored hints โ and ChatGPT's 32% raw hint-error rate was reducible to near zero (algebra) or 13% (statistics) using the self-consistency hallucination-mitigation technique.
Authoring help content for educational technologies is labor-intensive and costly โ a full-time employee may take a year to produce a textbook's worth of material. If LLMs can generate hints with sufficiently low error and sufficient learning efficacy, they could alleviate the most time- and cost-intensive component of tutoring-system development and enable scaling to many domains. This study evaluates whether ChatGPT-generated worked-solution hints can do so.
Study design
A 3 ร 4 between-subjects design with 274 participants (Mechanical Turk workers, all MTurk Masters with at least a high-school degree). Participants were randomly assigned to one of three hint conditions โ control (no hints), human tutor, or ChatGPT โ paired with one of four mathematics subjects drawn from OpenStax CC BY textbooks: Elementary Algebra, Intermediate Algebra, College Algebra, or Statistics. Each participant completed a 3-item pre-test, a 5-item acquisition phase in the OATutor platform, and a 3-item post-test (same questions as pre-test). ChatGPT hints were single worked-solution hints generated by prompting ChatGPT 3.5 with each problem's text; human hints came from the existing OATutor content library authored by UC Berkeley undergraduates.
Key findings
RQ1 โ Hint quality and hallucination mitigation
RQ2 โ Learning efficacy
Implications for AI in education
The findings suggest that LLM-generated worked solutions can be as effective as human-authored tutoring content while being produced in a fraction of the time (roughly 1/20th), opening the door to autonomous generation of effective mathematics tutoring content from arbitrary educational resources. However, the authors are explicit about caution: at a 32% raw error rate, ChatGPT should not be used to give feedback the way a teacher or TA would unless in a domain verified to have near-zero error. Where error mitigation cannot achieve near-zero rates, designers should frame LLM feedback as coming from an "imperfect robot" or peer-like source so students consider it critically. The error-reduction via self-consistency is central, and the findings ground the performance-vs-learning distinction by demonstrating genuine learning gains, not just performance.
Connected Concepts
Connected Articles
Citation
Pardos, Z. A., & Bhandari, S. (2024). ChatGPT-generated help produces learning gains equivalent to human tutor-authored help on mathematics skills. PLOS ONE, 19(5), e0304013. https://doi.org/10.1371/journal.pone.0304013