Research Article
Knowledge without Wisdom: Measuring Misalignment between LLMs and Intended Impact
Synthesis: The gap between what LLMs are capable of and what actually benefits learners — benchmark performance, downstream task quality, and intended educational impact are three distinct and often-misaligned levels.
The Three-Layer Alignment Problem
Hardy & Kim (2026) identify a cascading proxy problem in AI-for-education evaluation:
- Benchmark alignment (MMLU, pedagogical knowledge tests) — what models are typically evaluated on.
- Downstream task alignment (expert human ratings of teaching quality) — what models are asked to do.
- Intended impact alignment (student learning gains / VAMs) — what actually matters.
The paper demonstrates these three layers are not just loosely coupled — they can be negatively correlated.
Empirical Evidence
Study Design
- Dataset: NCTE Main Study — ~350 4th/5th-grade math teachers, US; lesson transcripts.
- Tasks: 7 classroom observation dimensions from MQI (explanations, error remediation, student questioning, language precision) and CLASS (behavior management, instructional dialogue, positive climate).
- Models: 16 leading LLMs (GPT-3.5 through Llama 4) with 3 prompting strategies each.
- Metrics: Bias-corrected distance correlation (dCor²) for dependence; Kendall's τ for directional alignment with expert ratings and student VAMs.
Finding 1: LLMs Share a Homogeneous "Pedagogy Heuristic"
Large Language Models (LLMs)-LLM agreement is substantially higher than LLM-human agreement. Models converge on a shared latent heuristic of "good teaching" that doesn't match expert human distinctions. This is attributed to shared pretraining on Internet text lacking authentic classroom discourse.
Finding 2: Benchmark Alignment ≠ Student Impact
Some models align moderately with expert ratings, but alignment with student learning gains is often near zero or negative. Human raters show a real (τ ≈ 0.11–0.14) signal with VAMs; LLMs largely don't. Reasoning-enhanced variants (o1, DeepSeek-R1) showed no improvement.
Finding 3: Ensembles Amplify Misalignment
Both benchmark-weighted aggregation and unanimous-voting ensembles worsened alignment with learning. Aggregating multiple misaligned models compounds the problem rather than averaging it out.
Finding 4: Model/Prompt Selection = 15% of Error
Choice of LLM and prompting strategy accounts for only ~15% of misalignment; the rest is shared across all models — common pretraining data and objectives are the dominant driver. Prompt engineering and model selection are weak levers.
Broader Implications
- Stop benchmarking alone — High scores on MMLU or even pedagogy-specific benchmarks do not predict beneficial educational impact. See TeachBench - Evaluating LLM Teaching Ability for syllabus-grounded alternatives.
- Ensembles are not a safety net — When models share the same flawed pretraining priors, voting and weighting make things worse.
- Pretraining is the intervention point — The field's focus on post-hoc alignment (RLHF, prompting) misses that shared pretraining corpora embed the core misalignment. See Training Pedagogical LLMs for Tutoring for training approaches.
- Measure impact directly — Practitioners must evaluate against intended student outcomes, not proxy task accuracy. Connects to The Evidence Base on AI in K-12: A 2026 Review demands for causal evidence.
This finding is a deep challenge to the ITS effectiveness literature: if even the best models can't align with student learning, what does "effective" tutoring AI look like? It also reinforces the The Evidence Base on AI in K-12: A 2026 Review finding that general-purpose AI underperforms pedagogically-designed systems.
Open Questions
- Can pretraining on authentic classroom data (not just Internet text) close the alignment gap?
- Are there tasks where the alignment gap is smaller (e.g., factual tutoring vs. qualitative judgment)?
- How does this interact with The LLM Fallacy: Misattribution in AI-Assisted Cognitive Workflows — do students over-trust misaligned AI outputs?
What this means for practice
- Instructors. Do not read a model's agreement with expert ratings as evidence that it understands your classroom: some models align moderately with expert human raters while their alignment with student learning gains is often near zero or negative.
- Researchers. Report dependence against intended outcomes, not just expert agreement. The human raters in this study show a real signal with VAMs (Kendall τb 0.11/0.03 and 0.14/0.06) that the 16 LLMs largely fail to reproduce, even though LLM-LLM agreement was substantially higher than LLM-human agreement.
- Researchers. Stop treating multi-model benchmarking ensembles as a safety net: both benchmark-weighted aggregation and unanimous-voting ensembles worsened alignment with learning rather than averaging the error out.
- Instructors. Expect model and prompt choice to be weak levers. Selection of LLM or prompting strategy accounts for only 15% of all measured misalignment error, and reasoning-enhanced variants showed no measurable improvement, so upgrading models is unlikely to repair pedagogical validity.
- Software developers. Plan evaluation around how the model is used, not only what it predicts: the dominant variance shares sit in classroom-text-conditioned interactions (LLM×ITEM×OBS 0.19, LLM×PROMPT×OBS 0.14), meaning failures attach to particular kinds of instructional evidence rather than to a deficient model.
Limitations
- Teaching-quality measures come from the NCTE Main Study, which comprises observations of roughly 350 4th and 5th-grade mathematics teachers across four U.S. school districts, so generalization from U.S. primary mathematics classrooms to all classrooms is not demonstrable here.
- Expert ratings pertain solely to a subset of rating items on a specific rubric — 7 observation dimensions drawn from MQI and CLASS — which may limit conclusions about other tasks of classroom instructional support.
- Value-added measures are imperfect, high-variance estimates of causal impact and transcript segments are partial, lossy views of instruction; the authors present their variance decomposition as a statement about where misalignment concentrates under this measurement system, not as a definitive census of all sources of pedagogical effectiveness.
- Estimates are conditional on the sampled items, segments, models, and prompt families (16 LLMs, 3 zero-shot prompt techniques), and the authors flag that the test set may have unobserved confounding factors in its construction.
Citation
Hardy, M., & Kim, Y. (2026). Knowledge without Wisdom: Measuring Misalignment between LLMs and Intended Impact.