Research Article
The Tutoring Effectiveness Index: Predicting LLM Math Tutor Quality from Four Conversation Signals
Synthesis: TEI is a training-free, judge-free index that selects the best tutoring response from multiple LLM candidate outputs using only four internal conversation signals — no RL training, no external judge model, no reward model required.
What It Is
How It Works
TEI combines four signals computed during decoding with fixed weights:
- V (Verify ratio): Regex over thinking trace counting Schoenfeld Verification keywords ("let me check", "verify", "double-check")
- M̃ (Math-step density): Regex on visible output, min-max normalized within candidate pool
- Q (Ends-question rate): Regex detecting if tutor turn ends with a question
- D (Deep-reasoning gate): Binary, fires if ≥40% of tokens have JSD below threshold
Formula: TEI(y) = 1.0·V + 0.75·M̃ - 1.0·Q + 0.5·1[DTR ≥ 0.4]
The signs matter more than magnitudes: reward verification and math content, penalize ending with questions, small bonus for deep reasoning.
Key Results
- TEI@8 raises improvement rate on pre-incorrect scenarios from 59.0% to 81.9% (+22.9 pp) on frozen DeepSeek-R1-8B, with no training
- TEI@4 achieves 75.7%, beating both Random@4 (58.6%) and DTR@4 (61.2%)
- Feature ablation: dropping Verify costs -0.054 AUC, dropping Math-steps costs -0.036, dropping Deep-reasoning gate costs only -0.009
- TEI@8 costs 4.1× tokens of greedy (16,334 vs 3,984), about half of self-consistency
The Alignment Tax
The paper quantifies severe degradation from pedagogical GRPO fine-tuning:
- Thinking length drops from 1,764 to 119 words/turn (−93%)
- Content-Knowledge accuracy falls by −71% relative
- Pedagogical-Knowledge accuracy falls by −80% relative
- Student Δ Solve Rate crosses from +0.180 to −0.012 — the aligned tutor becomes detrimental
Why It Matters
TEI demonstrates that simple lexical and structural signals can effectively steer a frozen LLM to be a much better math tutor without any training. This is especially valuable when RL fine-tuning is shown to catastrophically degrade tutoring quality. The approach is cost-effective and immediately deployable on frozen models.
Open Questions
- Does TEI generalize to non-math tutoring domains (writing, science, language)?
- Can the fixed weights be optimized per-domain without losing the training-free property?
- How does TEI interact with different base model architectures and sizes?
es and sizes?
What this means for practice
- Software developers. Rerank candidate responses at inference time before reaching for training: TEI@8 raised the improvement rate on pre-incorrect scenarios from 59.0% to 81.9% on a frozen model, with no RL.
- Software developers. Budget the rerank explicitly: TEI@8 costs 4.1× the tokens of greedy decoding (16,334 vs 3,984) and roughly half the cost of Cons@8 (31,868).
- Software developers. Keep verification language in the signal set and stop rewarding question-ending turns: dropping the verification term removes −.054 AUC and math-step density −.036, while the question-rate term carries a 1.0 penalty.
- Software developers. Audit any pedagogically fine-tuned tutor for reasoning collapse: the GRPO run behind this study's alignment tax cut thinking from 1,764 to 119 words per turn (−93%), with content-knowledge accuracy down 71% relative and pedagogical knowledge down 80%.
- Software developers. Gate deployment on a student-outcome check rather than a judge score: the Student Δ Solve Rate crossed from +0.180 to −0.012, so an aligned tutor that passes rubric-style evaluation can still be detrimental.
Limitations
- Student outcomes use GPT-4o-mini as a synthetic learner rather than real students.
- All experiments are in mathematics, so the four weights and the signals behind them may need re-derivation for other subjects.
- The Schoenfeld verification signal is regex-derived and surface-level: a GPT-4o-mini paragraph classifier disagrees with it on more than half of paragraphs.
- The four TEI weights are fixed a priori from theory, and the Helpful and leak rates are LLM-judge metrics that inherit same-model bias; the primary outcome (Δ Solve Rate) does not.
Citation
Shim, J., & Lee, U. (2026). The Tutoring Effectiveness Index: Predicting LLM Math Tutor Quality from Four Conversation Signals.