Concept
Large Language Models (LLMs)
Large Language Models (LLMs) — neural network models trained on vast text corpora that generate human-like text, powering most modern AI in education applications. LLMs are the computational backbone of generative AI tutoring, assessment, and content generation in education.
Questions to Consider
- What do you believe an AI chatbot 'knows' when it answers you? The page frames LLMs as generating probable text rather than retrieving verified facts — how does that distinction change how much you would trust a model's explanations?
- LLMs are described as the engine behind most modern AI education tools — tutoring, grading, content generation, and even diagnosing what students know. Of these uses, which do you think is most and least appropriate for a probabilistic text generator, and why?
- The page reports that three different LLMs produced sharply divergent support plans for the same learning-analytics input, each with different demographic assumptions. If models aren't interchangeable as advisors, what does that mean for an institution that adopts one?
- Because LLM output is sensitive to prompts and settings, two people can get very different results from the same model. How should this influence how you — as a learner or designer — phrase requests, and how much you trust a single output?
- A key limitation is hallucination — plausible-sounding but ungrounded content. In a tutoring or grading context, what would it take for you to feel confident the model wasn't inventing something, and what safeguards would you demand before letting it assess a real student?
Introduction
LLMs as the engine of AIED
LLMs are the most-referenced concept in the knowledge base (60+ articles) because they underpin nearly every AI education application:
- Tutoring: AI tutors use LLMs for dialogue, explanation, and Problem Solving guidance. Pedagogical training adapts general LLMs for educational use.
- Assessment: Grading systems, essay scoring, and item difficulty prediction leverage LLM capabilities. Razavi and Powers (2026) show GPT-4o can estimate the difficulty of K-5 math and reading items (N = 5170) calibrated under the Rasch IRT model: zero-shot ratings correlated moderately-to-strongly with true difficulties (r = 0.83 math, r = 0.81 reading) but varied by grade, while a feature-based strategy in which the LLM extracts cognitive and linguistic features for tree-based models reached correlations up to r = 0.87 — evidence that structured feature extraction can outperform a single holistic LLM judgment. Across the aggregate grading literature, a PRISMA-guided systematic review of 42 empirical studies (2023–2025) concludes that LLMs match human raters on short, well-structured tasks with detailed rubrics yet cannot fully replace human judgment on complex, open-ended, or subjective work, and that model version is a dominant determinant of grading quality (Can ChatGPT Replace the Teacher in Assessment? A Review of Research on the Use of Large Language Models in Grading and Providing Feedback). Reliability also varies sharply by item type: Falahat et al. (2026) found ChatGPT-5 matched faculty closely on objective pharmacy-exam items (CCC 0.935–1.000) but was unreliable on short-answer (CCC ≈0) and essay (0.341–0.854) items, and that providing a rubric did not consistently improve agreement.
- Multimodal AI reasoning LLMs as graders: when a multimodal, reasoning-capable LLM (GPT-o4-mini) graded a 296-student handwritten general-chemistry exam page-by-page against rubric images, single-run total scores were highly reproducible (ICC(A,1) = 0.967; averaging five runs reached 0.993) and agreed strongly with TA totals (R² = 0.91), yet item-level reliability was sharply format-dependent — textual and reaction-equation answers graded well while drawing and graphing were worse than random (background grids distract AI vision). This shows an LLM grader's trustworthiness is a function of response format and task, not just raw model capability, and that selective deferral via confidence filters is needed for high-stakes use (Assisting the grading of a handwritten general chemistry exam with artificial intelligence).
- Content: Generative AI content creation relies on LLMs. Question generation and video generation are LLM-driven.
- Safety: Pedagogical Safety, Hallucination Risk, and SafeTutors: Benchmarking Pedagogical Safety in AI Tutoring Systems research examine LLM-specific risks.
- Diagnosis: Knowledge Tracing and Cognitive Diagnosis increasingly incorporate LLMs for richer student modeling. Grounding matters enormously for error diagnosis: Reddig, Arora & MacLellan (2025) showed that supplying GPT-4 the tutor interface structure plus Bayesian Knowledge Tracing skill estimates raised logical-error identification from 40% to 81% on factoring (overall error diagnosis ~87.8%), while multi-step problems and responses with several errors remained weak cases and hallucinated "common-misconception" diagnoses persisted — evidence that an LLM's diagnostic value is as much a function of the structured context and learner-model signals it receives as of the model itself.
- Assessment model shift (2017–2024): Morley et al.'s scoping review of auto-marking short-answer science questions traces the field's move from fine-tuning smaller BERT models (dominant through 2021) toward prompting larger LLMs (GPT-1/2/3.5/4) from roughly 2022 — adopted via Prompt Engineering rather than fine-tuning — with domain-augmented models, rubric-aware prompting, and chain-of-thought lifting accuracy. Yet GPT models were rarely benchmarked against BERT on standard corpora, few auto-markers could explain their marks, and bias was seldom examined, cautions that apply to LLM assessment generally (Auto-marking short answer questions in science: The foundational years of transformer-based models from BERT to GPT-4).
Model-specific research
The knowledge base covers both general-purpose LLMs (GPT-4, Claude) and education-specific adaptations. Small language model benchmarks compare SLM performance for tutoring. Educational alignment research addresses how to make LLMs pedagogically appropriate. A classroom study across three frontier families — Oppenheimer, Cash & Connell Pensky (2025) — found that ChatGPT, Gemini, or Claude could act as collaborative critique partners for argumentative writing: over a semester of iterative essays, students improved on argument quality, prompt engineering, and response-to-AI feedback by roughly a full standard deviation each (all p < .001) and engaged deeply (87.8% rebutting LLM claims), positioning general-purpose LLMs as viable collaborative learning partners rather than mere answer generators.
A complementary line of work reframes LLMs from static graders into emulators of pedagogical reasoning. Yaşar et al. (2026) showed that GPT-4, scaffolded with a semantically precise, iteratively co-refined rubric, could approximate human evaluative judgment in design-based learning: initial LLM–human agreement was poor (Cronbach's Alpha = 0.393; Kappa −0.06 to 0.18), but iterative rubric refinement raised mean agreement from 54.75% to 81.25% (final Alpha = 0.798, Kappa 0.40–0.55), and K-means clustering of human and LLM score matrices showed highly correlated centroids (r = 0.89). The study positions the rubric as a mediating interface between human pedagogical intent and machine inference — evidence that off-the-shelf LLMs are not interchangeable as evaluators either, and that their assessment behavior is a design outcome shaped by the rubric and prompts they are given. Raw model capability differentiates grading too: benchmarking eleven GenAI and sentence-embedding models on 1,885 open-ended responses, Pecuchova, Benko & Drlik (2025) found only GPTo1 reached almost-perfect agreement with expert human graders (Fleiss' Kappa 0.82), with Claude3 and PaLM2 slightly behind, while reference-aligned models such as BERT fell far short — showing that frontier-model context-sensitivity matters for reliable open-ended assessment. Model differences also matter for high-stakes downstream uses. López-Pernas et al. (2026) showed that three LLMs produced sharply divergent student-support prescriptions for the same learning-analytic input, and each imposed distinct demographic priors on the learner profiles they generated — evidence that off-the-shelf LLMs are not interchangeable as prescriptive advisors. Similarly, Olvet et al. (2026) found that GPT-4's scoring of pre-clerkship medical open-ended questions climbed to substantial-to-almost-perfect inter-rater agreement with faculty (weighted kappa up to 0.94) only after three rounds of iterative rubric refinement, and fell to moderate (κw = 0.54) on a holistic-rubric item — reinforcing that rubric design, not raw capability alone, is the decisive lever for LLM scoring reliability. Model-specific behavior also shows up in how LLMs respond to skeptical users: an algorithmic audit queried ten frontier LLMs 500 times each with a rural-Montana K-12 AI-skeptic persona to test whether AI systems consulted by skeptical users are predisposed to encourage adoption. Eight of ten acknowledged user concerns then redirected to AI-engagement framings; composite scores spanned from 3.85 (Claude Sonnet) to 7.52 (Gemini 3.1 Pro Preview), with a cross-family AI scorer panel clearing Cohen's kappa >= 0.70. The pattern was a model-dependent design outcome. Model capability also depends on how models are combined: Bird (2026) fine-tuned eight state-of-the-art transformers (BERT, ELECTRA, RoBERTa, XLNet, ERNIE, ALBERT, DistilBERT, Longformer) to classify English literature by UK Key Stage, finding the best unimodal transformer (BERT) reached only an F1 of 0.75 — while fusing a fine-tuned ELECTRA with a computational-linguistics neural network lifted F1 to 0.996, showing that transformer text classification alone is limited and that fusion with complementary features is where the gains lie.
Connected Concepts
- Generative AI
- Prompt Engineering
- RAG (Retrieval-Augmented Generation)
- Hallucination Risk
- Pedagogical Safety
- Intelligent Tutoring
- Automated Assessment
- AI Literacy
- Knowledge Tracing
- Higher Education
- Scaffolding
- Training Pedagogical LLMs for Tutoring
- Learning by Teaching
- Technologies — Umbrella: AI technologies and techniques (models, LLM training, robotics, RAG, agentic)
Connected Articles
- What Students Ask Matters: LLM Interaction Depth, Task Quality, and Immediate Recall in Higher Education — What students ask matters: LLM interaction depth, task quality, and immediate recall (Tsiligkiris 2026)
- One Click Away: AI Tutoring with Khanmigo in a Two-Year School Experiment — One Click Away: Khanmigo in a two-year school experiment
- Assessing the Quality of AI-Generated Exams: A Large-Scale Field Study — Assessing the quality of AI-generated exams: a large-scale field study
- Neuro-symbolic pedagogical alignment (NSPA) for long-horizon classroom discourse analysis: Mitigating dialect bias via counterfactual preference optimization — Neuro-symbolic pedagogical alignment (NSPA)
- LLMs Do Not Grade Essays Like Humans — LLMs do not grade essays like humans (Mathew et al. 2026)
- Knowledge without Wisdom: Measuring Misalignment between LLMs and Intended Impact
- CSTutorBench: Benchmarking Small Language Models as Tutors for Block-Based Programming
- SafeTutors: Benchmarking Pedagogical Safety in AI Tutoring Systems
- Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction
- EduGuard: A Safe RAG-Based LLM Tutor for Programming Education
- From Evaluated Models to Evaluation Aids: A Multi-Evidence Study of LLM-Based Difficulty Calibration for Programming Examinations
- ELBench: A Multi-Dimensional Benchmark for Education-Facing Large Language Models
- Toward Convergence in Student-LLM Interactions: A Rapid Scoping Review and Taxonomy for Learning-Oriented Use
- LearnLM: Improving Gemini for Learning — LearnLM: pedagogical instruction following
- TeachLM: Post-Training LLMs for Education Using Authentic Learning Data — TeachLM: post-training with authentic learning data
- Conversational AI agents in education: an umbrella review of current utilization, challenges, and future directions — Umbrella review of conversational AI agents in education
- Can AI deliver appropriate support for diverse student profiles? A large-scale evaluation — Can AI deliver appropriate support for diverse student profiles? A large-scale evaluation
- Not for People Like Me: How Frontier AI Models Redirect Skeptical Rural School Staff — Algorithmic audit: how frontier LLMs redirect skeptical rural K-12 staff
- From evaluation to emulation: LLMs as agents of iterative pedagogical design — LLMs as agents of iterative pedagogical design
- Estimating Item Difficulty Using Large Language Models and Tree-Based Machine Learning Algorithms — Estimating item difficulty using LLMs and tree-based ML
- Auto-marking short answer questions in science: The foundational years of transformer-based models from BERT to GPT-4
- Generating In-Context, Personalized Feedback for Intelligent Tutors with Large Language Models
- You've Got AI Friend in Me: LLMs as Collaborative Learning Partners
- Automated Grading of Open-Ended Questions in Higher Education Using GenAI Models
- Assisting the grading of a handwritten general chemistry exam with artificial intelligence
- Bridging technology and education: The use of ChatGPT in grading pharmacy student exams
- Can Generative Artificial Intelligence Reliably Score Open-Ended Question Assessments in Undergraduate Medical Education?
- Can ChatGPT Replace the Teacher in Assessment? A Review of Research on the Use of Large Language Models in Grading and Providing Feedback
- StudentBench: AI and human tutoring yield equivalent GRE learning gains — StudentBench: AI and human tutoring yield equivalent GRE learning gains
- Evaluating Feedback Focus and Pedagogical Adaptivity in LLM-Generated Feedback on Student Writing — Evaluating Feedback Focus and Pedagogical Adaptivity in LLM-Generated Feedback on Student Writing
- Student Use of LLMs and the Limits of AI-Generated Question Difficulty in Data Science Courses — Student Use of LLMs and the Limits of AI-Generated Question Difficulty in Data Science Courses
- EduBehaviors: Assertion-based Schemas for Auditable Coding of Educational Dialogues — EduBehaviors: Assertion-based Schemas for Auditable Coding of Educational Dialogues
- From Sentiment Classification to Actionable and Responsible Feedback: A Scoping Review and Evidence Map of NLP in Student Evaluation of Teaching, 2015–2026 — From Sentiment Classification to Actionable and Responsible Feedback: A Scoping Review and Evidence Map of NLP in Student Evaluation of Teaching, 2015–2026
- Evaluation of pre-trained models for pedagogical assessment of novel AI-assisted educational questions — Evaluation of pre-trained models for pedagogical assessment of novel AI-assisted educational questions
- Open Questions Towards Skill-Sustaining Reliance in Reflective AI Engagement — Open Questions Towards Skill-Sustaining Reliance in Reflective AI Engagement