Special-R1: Reinforcement Learning for Special Education โ€” Aligning LLM Tutors to Diverse Learners through Disability-Adaptive Training

Created: 2026-06-01 | Tags: intelligent-tutoringllmspecial-educationpersonalized-learningreinforcement-learningk-12scaffolding

Authors: Unggi Lee, Jihoi Na, Yeil Jeong, Haeun Park, Yeonju Jang (2026)

What It Is

Special-R1 is a framework that extends pedagogical reinforcement learning (RL) to special education. While prior RL-based tutor alignment methods targeted only generic math learners, Special-R1 explicitly models cognitive and communicative diversity across five disability profiles.

How It Works

The framework has two core components:

1. Two-dimensional adaptive system prompt: Couples a difficulty-based support level (scaffolding) with a disability-specific teaching style, forming a persona-aware prompt that guides the LLM tutor during multi-turn dialogue.

2. Persona-aware Thinking Reward: The judge rubric used to compute the training reward is conditioned on the learner's disability profile rather than a generic student. This shapes the tutor to produce responses that are helpful, safe, and appropriately challenging for each specific persona.

Key Results

Critical Insight

Students with specific learning disabilities in mathematics remain underserved, suggesting a need for multimodal extensions (visual aids, interactive diagrams) in future work.

Why It Matters

This is the first multi-turn pedagogical RL framework specifically targeting special education. It demonstrates that LLM tutors can be systematically aligned to support students with disabilities, improving both perceived helpfulness and pedagogical fit. The persona-conditioned reward rubric provides a replicable recipe for adapting RLHF-based tutor fine-tuning to diverse learner profiles.

Open Questions

Related Pages

Citation

APA: Lee, U., Na, J., Jeong, Y., Park, H., & Jang, Y. (2026). Special-R1: Reinforcement learning for special education: Aligning LLM tutors to diverse learners through disability-adaptive training. arXiv:2605.30670.