On this page

Synthesis: This paper explores how an embodied AI agent can act as a coach that accelerates human motor-skill development using reinforcement learning. The authors argue that effective coaching requires dynamically balancing guidance with learner autonomy — too much assistance leads to Over-Reliance and skill atrophy, while too little leaves learners struggling.

Key findings:

  • An RL-based coaching policy that adapts its level of intervention to the learner's current skill level significantly accelerates skill acquisition compared to static assistance levels.
  • The AI coach that gradually fades scaffolding (consistent with Scaffolding theory in Intelligent Tutoring) produced the best long-term retention and transfer performance.
  • Over-reliance emerged when the coach provided excessive intervention, confirming the Over-Reliance concern documented in Generative AI tutoring contexts.

Implications:

What this means for practice

  • Instructional designers. Optimize coaching for the learner's independent competence rather than for task performance: the competence-trained coach cut lap time by 27.9% (p = 0.005, dz = -1.08) and failures by 3.52 per lap (p < 0.001, dz = -2.13) across one coached session.
  • Instructional designers. Do not fade assistance on a fixed schedule: rule-based fading produced no reliable lap-time change (11.3%, p = 0.21) and a failure-count gain roughly one-third the size of the adaptive coach's, so modulate support from an estimate of current skill.
  • Designers. Calibrate intervention to estimated skill within the task: the coach gave a novice more assistance than a more skilled learner and varied the assistance sharply around the gates that decide whether a learner collides or passes.
  • Researchers. Carry the objective beyond motor skills: the authors argue that coding agents are optimized for task performance with no explicit incentive for what the human retains, which makes independent-competence objectives a transferable design target for embodied and academic skill training alike.

Limitations

  • The randomized user study ran with N = 33 participants, about 11 per coaching condition, so the between-group contrasts favoring the adaptive coach (Hedges' g between 0.60 and 0.76) carried p-values of 0.09 to 0.16; only the within-subject tests reached significance.
  • All measurement happened in a high-fidelity first-person-view drone-racing simulator over a single 15-lap session of roughly 40 minutes, with human control limited to yaw and roll rate while pitch and thrust were automated.
  • The coaching policy was trained against simulated learners whose skill evolution followed a probabilistic automaton, and the authors acknowledge that simulated learners may not capture the full variability of real human learning in more complex motor-skill training settings.
  • Only the assistance modulation was learned: verbal instructions and visual cues were predefined and triggered by fixed rules, so the results say nothing about jointly optimized multimodal coaching.

Citation

Wang, W., Gu, E., Loquercio, A., Hu, H., & Mangharam, R. (2026). AI Coaching for Accelerating Human Skill Development with Reinforcement Learning.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.