On this page

Synthesis: TEI is a training-free, judge-free index that selects the best tutoring response from multiple LLM candidate outputs using only four internal conversation signals — no RL training, no external judge model, no reward model required.

What It Is

How It Works

TEI combines four signals computed during decoding with fixed weights:

  • V (Verify ratio): Regex over thinking trace counting Schoenfeld Verification keywords ("let me check", "verify", "double-check")
  • M̃ (Math-step density): Regex on visible output, min-max normalized within candidate pool
  • Q (Ends-question rate): Regex detecting if tutor turn ends with a question
  • D (Deep-reasoning gate): Binary, fires if ≥40% of tokens have JSD below threshold

Formula: TEI(y) = 1.0·V + 0.75·M̃ - 1.0·Q + 0.5·1[DTR ≥ 0.4]

The signs matter more than magnitudes: reward verification and math content, penalize ending with questions, small bonus for deep reasoning.

Key Results

  • TEI@8 raises improvement rate on pre-incorrect scenarios from 59.0% to 81.9% (+22.9 pp) on frozen DeepSeek-R1-8B, with no training
  • TEI@4 achieves 75.7%, beating both Random@4 (58.6%) and DTR@4 (61.2%)
  • Feature ablation: dropping Verify costs -0.054 AUC, dropping Math-steps costs -0.036, dropping Deep-reasoning gate costs only -0.009
  • TEI@8 costs 4.1× tokens of greedy (16,334 vs 3,984), about half of self-consistency

The Alignment Tax

The paper quantifies severe degradation from pedagogical GRPO fine-tuning:

  • Thinking length drops from 1,764 to 119 words/turn (−93%)
  • Content-Knowledge accuracy falls by −71% relative
  • Pedagogical-Knowledge accuracy falls by −80% relative
  • Student Δ Solve Rate crosses from +0.180 to −0.012 — the aligned tutor becomes detrimental

Why It Matters

TEI demonstrates that simple lexical and structural signals can effectively steer a frozen LLM to be a much better math tutor without any training. This is especially valuable when RL fine-tuning is shown to catastrophically degrade tutoring quality. The approach is cost-effective and immediately deployable on frozen models.

Open Questions

  • Does TEI generalize to non-math tutoring domains (writing, science, language)?
  • Can the fixed weights be optimized per-domain without losing the training-free property?
  • How does TEI interact with different base model architectures and sizes?

es and sizes?

What this means for practice

  • Software developers. Rerank candidate responses at inference time before reaching for training: TEI@8 raised the improvement rate on pre-incorrect scenarios from 59.0% to 81.9% on a frozen model, with no RL.
  • Software developers. Budget the rerank explicitly: TEI@8 costs 4.1× the tokens of greedy decoding (16,334 vs 3,984) and roughly half the cost of Cons@8 (31,868).
  • Software developers. Keep verification language in the signal set and stop rewarding question-ending turns: dropping the verification term removes −.054 AUC and math-step density −.036, while the question-rate term carries a 1.0 penalty.
  • Software developers. Audit any pedagogically fine-tuned tutor for reasoning collapse: the GRPO run behind this study's alignment tax cut thinking from 1,764 to 119 words per turn (−93%), with content-knowledge accuracy down 71% relative and pedagogical knowledge down 80%.
  • Software developers. Gate deployment on a student-outcome check rather than a judge score: the Student Δ Solve Rate crossed from +0.180 to −0.012, so an aligned tutor that passes rubric-style evaluation can still be detrimental.

Limitations

  • Student outcomes use GPT-4o-mini as a synthetic learner rather than real students.
  • All experiments are in mathematics, so the four weights and the signals behind them may need re-derivation for other subjects.
  • The Schoenfeld verification signal is regex-derived and surface-level: a GPT-4o-mini paragraph classifier disagrees with it on more than half of paragraphs.
  • The four TEI weights are fixed a priori from theory, and the Helpful and leak rates are LLM-judge metrics that inherit same-model bias; the primary outcome (Δ Solve Rate) does not.

Citation

Shim, J., & Lee, U. (2026). The Tutoring Effectiveness Index: Predicting LLM Math Tutor Quality from Four Conversation Signals.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.