Research Article
Application-Driven Pedagogical Knowledge Optimization of Open-Source LLMs via Reinforcement Learning and Supervised
Synthesis: EduQwen: Pedagogical RL — A multi-stage optimization strategy combining reinforcement learning (DAPO) and supervised fine-tuning (SFT) to enhance the pedagogical knowledge of open-source LLMs, producing a family of dense 32B-parameter models that achieve state-of-the-art performance on the Cross-Domain Pedagogical Knowledge (CDPK) Benchmark, surpassing even much larger proprietary systems such as Gemini-3 Pro. Demonstrates that domain-specialized optimization can transform mid-sized open-source LLMs into true pedagogical domain experts, prioritizing guided learning over answer-giving.
Key Findings
The EduQwen project addresses a fundamental misalignment in Large Language Models (LLMs) behavior for education: general-purpose models are optimized for immediate helpfulness — providing answers directly — while effective pedagogy requires guiding learners to discover answers themselves. This gap, labeled the Catching The Correct Answer Trap: Characterising AI Tutor Blind Spots When Analysing Student Reasoning, drives the core research question.
Three-stage optimization pipeline. The team used a dense Qwen3-32B backbone (chosen over MoE architectures for superior responsiveness to iterative optimization) and applied:
- Stage 1 — RL with DAPO: Decoupled Advantage Policy Optimization was selected over GRPO for its stable gradients on complex pedagogical reasoning tasks, using asymmetric clipping to prevent catastrophic divergence. Hard-negative mining identified 440 questions the base model could not answer perfectly across 30 attempts, then sorted them by error frequency into a difficulty-ordered curriculum. Extended rollouts (5→8 steps) enabled multi-step pedagogical decision-making. Result: 94.13% on CDPK, already SOTA.
- Stage 2 — Synthetic SFT: The RL1 model generated 40,000 synthetic responses; only correct responses with gradient-based selection were retained, yielding 1,050 high-quality difficulty-ordered data points. Difficulty-weighted sampling kept all hard examples while sampling easy ones sparsely. Result: 96.20%.
- Stage 3 — Final RL (RL2): A second DAPO round on the SFT checkpoint reused the original hard-negative dataset, allowing the further refined model to tackle originally challenging problems. Result: 96.52% — definitive SOTA.
Benchmark dominance. EduQwen 32B-SFT-RL2 established new SOTA results across the Interactive Pedagogy Benchmark Leaderboard, surpassing Gemini-3 Pro (90.55%) — a system that is orders of magnitude larger. This proves that dense, mid-sized open-source models can become pedagogical domain experts through specialized optimization.
What this means for practice
- Developers. Use DAPO rather than GRPO for pedagogical reinforcement learning, and mine hard negatives first: 440 questions the base model failed across 30 attempts became a difficulty-ordered training curriculum.
- Developers. Insert a synthetic SFT stage between the two RL rounds, filtering 40,000 generated responses down to 1,050 difficulty-ordered correct ones, instead of training on all available data indiscriminately.
- Administrators. Budget for locally hosted 32B open-source models where governance forbids sending student interactions to proprietary endpoints; the optimized checkpoint is reported to beat a far larger proprietary system on the same benchmark.
- Researchers. Evaluate pedagogical quality by guidance behavior rather than answer accuracy alone, and treat exam-style multiple-choice scores as a screening result until free-form tutoring dialogue is tested.
Limitations
- The primary result (96.52% on CDPK) comes entirely from teacher-exam-style multiple-choice questions; the authors name free-form tutoring dialogs and longitudinal learning gains as untested, and classroom or platform validation as a prerequisite for deployment claims.
- The training signal is small and self-distilled: hard-negative mining drew on 440 questions, and the SFT stage retained 1,050 data points from 40,000 synthetic responses generated by the model being trained.
- Only one backbone was optimized (Qwen3-32B, plus a single Qwen3-30B-VL run); the authors state they could not test larger open-source models such as DeepSeek-R1 or Qwen3-235B-A22B-Thinking because of compute and time costs.
- The TutorBench transfer result of 61.64% came from one SFT run with 596 filtered synthetic responses and used a different judge (Claude-4.5-Sonnet) than the official leaderboard (Claude-4-Sonnet), so it is not directly comparable with published leaderboard numbers.
Citation
Singh, N. P., Wang, X., Garikipati, A., Ciobanu, M., Mao, Q., & Das, R. (2026). Application-Driven Pedagogical Knowledge Optimization of Open-Source LLMs via RL and SFT.