On this page

LLM Training and Fine-Tuning — how educational AI models are made: the pipeline from pretraining through post-training to adaptation, what each stage costs, and what the evidence says it buys. The organizing problem is an incentive mismatch: general-purpose models are optimized to answer, while teaching requires withholding the answer. Post-training and fine-tuning are the two levers that change that behavior, and the knowledge base's clearest result is that they are a third choice, not the first — retrieval and prompting come earlier in the decision, and one study here found supervised fine-tuning failing at a task where plain prompting won (WrAFT). Written for educational software developers and teaching practitioners deciding what to build with.

Questions to Consider

  • If you have a tutoring task where a general model answers too readily, is your first move a better prompt, retrieval over your own materials, or training? What would you need to measure to tell which one helped?
  • Fine-tuning needs data. Where would your project's training examples come from, who owns them, and what privacy review would they need before they could be used?
  • A study here found that fine-tuning a model on assessment data improved formatting and similarity while a separate fine-tune failed outright at generating feedback. What does that suggest about matching the training method to the task?
  • Post-training with reinforcement learning can teach a model to guide rather than answer. What reward would you write to capture "guides well", and how would a model game it?
  • The knowledge base reports that model and prompt choice account for only about 15% of the gap between an LLM and student learning gains. If that is right, how should it change your build-versus-buy decision?
  • A fine-tuned model that is accurate can still be unsafe across a long conversation. What would you test before letting one talk to students unsupervised?

Introduction

Educational AI development runs into the same wall from two directions. General-purpose language models are post-trained on human preference for helpfulness, which in practice means answering promptly and completely; tutoring requires the opposite, because the pedagogical goal is to help a student reach the answer rather than to hand it over. At the same time, a general model knows nothing about your curriculum, your rubric, or your institution's voice, and no amount of prompt engineering reliably installs those.

This page covers the techniques that change a model rather than the text you send it. They sit on a ladder of increasing commitment: prompting, then retrieval, then parameter-efficient adaptation, then full fine-tuning, then post-training with preference or reward signals. Each rung costs more data, more compute, and more evaluation discipline than the last, and each is a worse first choice than the rung below it unless a specific measurement justifies the climb. Much of the research literature reports only the top of that ladder, which makes it easy for a developer to reach for training when retrieval would have been sufficient.

The page is organized around the decisions a builder actually faces: which lever to pull (this section and the next), what adaptation buys and costs in practice, what post-training can shape that adaptation cannot, and how these systems fail. Terminology is used in its standard sense: pretraining is the from-scratch stage on a general corpus, post-training is everything after it that shapes behavior (supervised fine-tuning, preference optimization, reinforcement learning), and fine-tuning covers both the supervised stage of post-training and the later task- or domain-adaptation work, including parameter-efficient methods.

The training pipeline, stage by stage

Three stages, with very different accessibility. Knowing which one a paper is describing prevents most misreadings of this literature.

Pretraining builds a base model from a general text corpus. It is the stage that produces the model's broad capability, and it is out of reach for essentially every educational project: the cost is measured in millions of dollars and the data is a web-scale crawl. Nothing in this knowledge base does it. Its relevance to practitioners is diagnostic rather than actionable — pretraining data is the dominant lever on how a model behaves, and it is the one lever you cannot pull. That asymmetry is why the field's remaining techniques all operate on an already-trained model.

Post-training shapes behavior on top of the base model. Supervised fine-tuning (SFT) teaches the model to imitate demonstrations; preference optimization and reinforcement learning then push it toward outputs a reward model or a set of human judgments rates higher. This is the stage where a model can be taught to guide instead of answer, and it is where the most striking educational results live. It is also the stage most sensitive to reward design, because a reward is a compressed specification of what you want and models optimize whatever you actually wrote.

Adaptation fits an existing model to your task or domain, usually with far less data and compute. Parameter-efficient fine-tuning (PEFT), of which LoRA is the common form, trains a small number of added parameters and leaves the base weights frozen. Full fine-tuning updates everything. Distillation and unlearning sit alongside them. The sections below take each in turn. For most educational developers this is the practical stage — the one where a few hundred to a few thousand examples and a single GPU can produce a deployable model.

Prompt, retrieve, or train? The decision that comes first

The strongest single result in this knowledge base on that question is a retrieval study, not a training study. Building a course knowledge-base assistant, Shen et al. (2026) evaluated local models in three configurations and found that retrieval, not the model, is the first-order design decision: their local LLM with no retrieval reached only 52.3% accuracy, below a TF-IDF baseline of 55.4%, while adding retrieval-augmented generation with no fine-tuning at all lifted it to 66.6% (Towards sustainable AI knowledge-base assistants in computer science education: on-premise deployment and optimization with open educational resources). The fine-tuning configurations they also tested are reported alongside retrieval ablations, quantization and energy per query, which makes the comparison unusually honest: a base model with no grounding can be worse than a decades-old lexical method, and the cheapest fix is not training.

What the field actually builds reflects a similar ordering. A PRISMA-guided review of 23 empirical studies (2020–2025) on customizing AI for writing instruction found prompt engineering dominant (N = 13), ahead of fine-tuning (N = 7) and hybrid architectures (N = 3) (Customizing AI for writing pedagogy: a systematic review of pedagogical goals, theoretical principles, and technical design). The review's sharper finding is a structural misalignment rather than a technique ranking: stated pedagogical goals have moved toward writing processes, feedback literacy and higher-order academic skills, while the dominant implementations still pursue product-focused goals through prompt design or fine-tuning.

Where prompting and training have been compared head-to-head, the result depends on the task, which is exactly why the decision should be measured rather than assumed:

What fine-tuning buys: controllability over scale

When a general model cannot be made to hold a property you need, fine-tuning on a modest dataset often can — and the recurring result is that targeting beats size. Three 8B models fine-tuned on an expert-designed children's reading curriculum outperformed zero-shot GPT-4o and Llama 3.3 70B on difficulty-related metrics, with negligible safety issues. The authors frame this as controllability over scale, since a compact model can be tuned to a specific reading level and error pattern in a way a general model cannot be asked to (Children's English Reading Story Generation via Supervised Fine-Tuning of Compact LLMs with Controllable Difficulty and Safety).

The pattern repeats across very different tasks:

  • Curriculum grounding. A 24,795-example multilingual instruction dataset grounded in Indian Knowledge Systems produced a 7B fine-tune scoring 6.39 on a five-judge external panel (median over 1,201 stratified items). That is within 0.15 of a strong general reference model at a fraction of the deployment cost — while the same base model scored near zero on IKS-specific dimensions without the fine-tune (IKS-Instruct: A 24,000-Example Multilingual Dataset for Teaching Language Models Indian Knowledge Systems). That gap between "competent generally" and "competent here" is the whole case for domain adaptation. Note also the counterintuitive detail the authors report: quality did not rise monotonically with data curation.
  • Assessment reliability, cheaply. A single LoRA adapter trained on roughly 3,900 pooled graded examples brought five small open models (4B–30B) to parity or better with a human grader across two computer-science exams, and nearly erased persona sensitivity (drift ≤ 0.32 MAE) (Where LLM Graders Succeed and Break: Evidence from Two Computer-Science Exams). One adapter, one dataset, several base models — this is the shape of a practical deployment.
  • Selective automation. Confidence was a reliable predictor of scoring error (β = −0.602, p < .001). Routing the least confident 20% of responses to human review moved a fine-tuned GPT-3.5 model from r = 0.781 to r = 0.822 (RMSE 0.5990 → 0.5544) while cutting manual scoring work by roughly 80% (Know When to Trust: Making AI Scoring More Reliable for Educational Assessment). Fine-tuning here is not replacing the human; it is making the human's attention affordable.
  • Measuring constructs you would otherwise hand-score. Fine-tuned Hungarian transformers (hubert-base-cc, PULI-BERT-Large) were benchmarked against TF-IDF features and Qwen3 embeddings for scoring reflective writing in teacher training (Automatic Reflection Level Classification in Hungarian Student Essays), and a fine-tuned multimodal model (Qwen3.5-based) was shown to reconstruct item characteristic curves for multiple-choice items, learning the response patterns encoded in 3PL and MCM curves rather than being told them (Multimodal Item Parameter Estimation using Simulated Response Probabilities).

Parameter-efficient adaptation in practice

LoRA and its relatives are where most educational teams will actually work, and the literature contains unusually specific guidance about how they behave.

Rank is a trade-off, not a dial to maximize. Lu et al. (2026) built 360 system–user–assistant dialogues from a Linear Control Systems course, restructured answers into a Solution–Method–Teaching-Points format, and applied LoRA to Qwen2.5-3B and 7B at ranks 4, 8 and 16. Structured-output coverage moved from near zero at base to roughly 1.00, and the best configuration (7B, r = 16) reached ROUGE-L 0.4093 with bootstrap confidence intervals for the gain entirely above zero. But gain per million adapter parameters fell monotonically as rank rose, so course-level alignment is a scale-and-rank trade-off rather than a free upgrade (LoRA Fine-Tuned Models for Control Systems Course Q&A: A Multidimensional Evaluation of Model Scale and Rank Effects). Their metrics measure similarity and formatting, not derivational accuracy, which is the standard caveat for this whole family of evaluations.

Scale does not predict success, and identical settings behave differently across architectures. AiAWE, an open-source automated writing evaluation system built on a LoRA-adapted Gemma-3-27B-it, reached RMSE 0.474, QWK 0.828, and agreement within ±0.5 of the human score on 90.56% of 360 evaluation essays, outperforming LLaMA-3.3-70B and a fine-tuned GPT-3.5 baseline — while running on a consumer-grade server. Three broader findings matter more than the score: model scale was not a reliable predictor of downstream performance under LoRA adaptation, identical LoRA hyperparameters produced qualitatively different adaptation behaviors across architectures, and a well-tuned mid-size open model can be competitive with proprietary systems (AiAWE: An Open-Source LLM Automated Writing Evaluation System Using LoRA-Adapted Instruction-Tuned Models).

Which layers to adapt is a real decision. In a discourse-analysis system for English teaching, fine-tuning a truncated BERT by updating only the last four Transformer layers beat both alternatives: adapting only the top layer hit a lower ceiling, and full 12-layer fine-tuning overfitted and oscillated (Automatic discourse relation classification and feedback optimization in English teaching based on transformer BERT model).

The base architecture can matter more than the parameter count. Fine-tuning three open text-to-image models on 1,000 captioned nuclear-engineering images substantially improved Stable Diffusion XL, gave limited gains for SD-v3.5-Medium, and produced no measurable improvement for the flow-matching Flux.1 model at all (NuclearDiffusion: Text-to-Image Foundation Models for Learning Nuclear Energy Concepts). A fine-tuning recipe is not portable across architectures, which is the same lesson AiAWE reports from the other direction.

Two adjacent techniques round out the adaptation stage. Distillation compresses a large or black-box system into a small deployable one. A two-stage pipeline distilled a fitted black-box ML estimator and its post-hoc interpretation into a small open-weight LLM, with a 2B-parameter "mentee" achieving near-lossless recovery of the oracle effect surface (r > .90). That result is reported under a faithfulness-first evaluation that audits every narration against the attribution it claims to describe (Distilling Black-Box Machine Learning into a Small, Self-Explaining Language Model for Learning Analytics). Unlearning removes targeted content after training: gradient-based unlearning was applied to three models to strip PII and harmful content, tested in two removal orders (PII-first and harmful-content-first) (Balancing AI responsibility with privacy, safety, and utility: Unlearning in large language models for mathematics education) — the relevant tool when a model has memorized something it should not carry.

Post-training: shaping behavior, not just format

Adaptation teaches a model what to produce; post-training teaches it how to behave. This is where the pedagogical results are most striking, and where the design of the reward or preference signal decides everything.

The pipeline result. Singh et al. (2026) transformed Qwen3-32B into EduQwen through three stages — initial RL, synthetic SFT, final RL — reaching 96.52% on the CDPK benchmark and surpassing Gemini-3 Pro's 90.55%. The intermediate results are the instructive part: the first RL stage alone reached 94.13%, SFT on 40,000 self-generated responses took it to 96.20%, and the final RL round added the last fraction. The reward model prioritized guiding responses over direct answers, hard-negative mining excluded questions the base model already solved, and rollouts were extended from 5 to 8 steps to capture multi-step pedagogical decisions (Application-Driven Pedagogical Knowledge Optimization of Open-Source LLMs via Reinforcement Learning and Supervised). Against that, the Pedagogy Benchmark — drawn from real teacher professional-development exams across 97 models — found accuracy ranging from 28% to 89%, evidence that pedagogical knowledge is not acquired incidentally during general pretraining (Lelièvre et al., 2025).

Training the decision rather than the utterance. TACT post-trained a tutor on a 13-strategy taxonomy plus a two-axis student-move taxonomy and gained 20.30 points over its Qwen3.5-4B backbone, with a diagnostic benchmark that withholds the learner-state labels available during training so the model must infer state from dialogue (TACT: Taxonomy-Aligned Post-Training for Pedagogically Adaptive English Tutoring). The same logic appears in Special-R1 for special education alignment (Special-R1: Reinforcement Learning for Special Education — Aligning LLM Tutors to Diverse Learners through Disability-Adaptive Training) and in heuristic-RL work that aligns models as Socratic guides rather than answerers (Beyond Direct Answering: Aligning Educational LLMs as Socratic Guides via Heuristic Reinforcement Learning). Post-training need not only shape what the model says: one platform pairs a guarded tutoring chatbot with a reinforcement-learning agent that chooses the next practice problem, so the trained policy decides what the learner does next (Effective Personalized AI Tutors via LLM-Guided Reinforcement Learning).

Distilled reasoning models. A cheaper route to pedagogical behavior is to teach a small model to imitate a larger one: Pedagogy-R1 (1.5B and 7B) was instruction-tuned on pedagogically filtered outputs distilled from a QwQ-32B teacher, paired with Chain-of-Pedagogy prompting (Pedagogy-R1: Pedagogical Large Reasoning Model and Well-balanced Educational Benchmark).

Instruction-conditioned post-training versus authentic data. LearnLM frames education-model training as pedagogical instruction following, carrying system-level instructions that let developers and teachers specify tutor behavior without committing to one definition of pedagogy. It is mixed into Gemini's post-training stages by co-training, and experts preferred it over GPT-4o (+31%), Claude 3.5 Sonnet (+11%) and base Gemini 1.5 Pro (+13%). The key finding is that RL is substantially more effective than SFT alone for following nuanced pedagogical instructions in long conversations (LearnLM: Improving Gemini for Learning). TeachLM takes the opposite bet: that prompt engineering is a stopgap and the scarce ingredient is authentic learner–tutor interaction data. Trained on 100,000 hours of one-on-one sessions under rigorous anonymization, it doubles student talk time, improves questioning style, and increases dialogue turns by 50% (TeachLM: Post-Training LLMs for Education Using Authentic Learning Data). Together they define the design space: instruction-conditioned post-training when your data is scarce, fine-tuning on real tutoring interactions when you have it.

Supervision quality beats supervision quantity. SWIM's progression for a writing simulator is the cleanest demonstration. Rubric-grounded prompting gave limited proficiency control (best average trait QWK 0.577 for Claude Sonnet, 0.422 for GPT-5.4, near zero for an open 7B model); supervised fine-tuning lifted that 7B model to 0.474 ± 0.023. GRPO against an automated-essay-scoring-derived reward lifted it further to 0.618 ± 0.005 across every trait and prompt, with the reward designed as a dense trait-normalized accuracy because exact-match rewards are too sparse in the multi-trait setting (SWIM: Student Writing Simulation via Proficiency-Conditioned Generation). The misconception-modeling study reaches a sharper conclusion about what the supervision must contain: instruction-tuned models learned algebra misconceptions only when trained on step-level solution traces, with accuracy staying below 30% at every data size when trained on final answers alone. Its student role overgeneralized the learned error until correct examples were explicitly mixed in at ratios as low as one in four, while the tutor role showed no such cost, holding correct accuracy from 93% to 98% across ten jointly trained misconceptions (Misconception Acquisition Dynamics in Large Language Models).

The training data is itself something models can curate: Edu-QuRating adapts preference distillation to educational data curation, replacing a single "is this educational?" score with 20 rubric dimensions covering factual accuracy, pedagogical structure and level suitability (Edu-QuRating: Multi-Dimensional Educational Data Curation with Distilled Pairwise Judgements).

Theory can be a training scaffold. For agents automating instructional systems design, a hybrid of classical ADDIE and Dick & Carey frameworks with ReAct-style reasoning outperformed both pure theory (structured but inflexible) and technique-only agents (flexible but ungrounded), across 25,795 scenarios from a 51-variable context matrix with a multi-judge protocol to reduce LLM-as-judge bias (ISD-Agent-Bench: A Comprehensive Benchmark for Evaluating LLM-based Instructional Design Agents). And a study of supervision labels found that assigning every training example one target behavior — subject competence, curriculum grounding, diagnostic reasoning, or scaffolding — lifted every model scale tested, with the largest gains in scaffolding and in using a learner's history, while knowledge-state diagnosis remained weakest at 54.04% (OmniEdu: Open Foundation Models for Learning and Teaching).

Training the learner side, not only the tutor

The same machinery is increasingly pointed at simulating students, which changes the economics of evaluating a tutor: instead of recruiting learners, you generate them.

The caution is that simulated students must be validated like any other instrument. Benchmarking fine-tuned and prompted models on 382 held-out dialogues from the largest public corpus of real student–tutor mathematics dialogues — across seven metrics spanning linguistic, behavioral and cognitive aspects — produced the field's most direct test of whether simulated students behave like students (Simulated Students in Tutoring Dialogues: Substance or Illusion?). Two further approaches push on fidelity: INSIDE fine-tunes LLMs to both act and think like students, generating internal dialogue grounded in Bloom's Taxonomy across cognitive, affective and action dimensions and training on paired think-traces and actions (INSIDE the Student's Mind: Jointly Modeling Latent Reasoning and Action in LLM Student Simulators), and history-aware profiles condition simulation on a student's prior trajectory rather than a static persona (Who Am I? History-Aware Profiles for Student Simulation in Tutoring Dialogues). For a developer, the practical lesson is that a simulator is a measurement instrument and inherits every validity question that implies.

What goes wrong

Sycophancy is a training objective, not a usability setting. Tutoring requires corrective friction, and models resist it. EduFrameTrap shows that models which withstand context-switch attacks still capitulate under authority or social-affective pressure and withhold corrective feedback, which is why its authors argue "kind-but-correct" behavior should be an explicit training requirement rather than a preference (Sycophancy is an Educational Safety Risk: Why LLM Tutors Need Sycophancy Benchmarks). Training that rewards guiding over answering — as EduQwen's DAPO reward model does — is one structural lever against it, but contextual sycophancy persists after prompting and alignment, with learners' errors still propagating into AI advice (The Hidden Cost of Contextual Sycophancy: an AI Literacy Intervention in Human-AI Collaboration).

Training does not make a model safe over a long conversation. SafeTutors shows that even specialized pedagogical models degrade across sustained dialogue and can commit answer over-disclosure harms (SafeTutors: Benchmarking Pedagogical Safety in AI Tutoring Systems), so pedagogical safety has to be tested at conversation length rather than at the single turn.

Validation can substitute for training — or be required instead of it. A frontier untrained GPT-4 produced roughly 35% too-general, incorrect or answer-revealing hints when authoring feedback for an intelligent tutoring system, and its own automated quality checks misaligned with human judgment; the authors conclude that LLMs lack an internal model of instruction and that robust validation or domain-specific training is needed before unsupervised learner-facing use (Generating In-Context, Personalized Feedback for Intelligent Tutors with Large Language Models). The practical implication cuts both ways: sometimes the right answer is a validation layer rather than a training run, and sometimes validation is what tells you a training run failed.

Your evaluation may be measuring the wrong thing. This is the most common trap in the fine-tuning literature here. ROUGE-L and QWK measure similarity and ranking, not derivational correctness (LoRA Fine-Tuned Models for Control Systems Course Q&A: A Multidimensional Evaluation of Model Scale and Rank Effects); a fine-tune can reach excellent QWK while the same system's feedback is unparseable (WrAFT: a Modularized Automated Writing Evaluation System for Argumentative Essays); and a model can rank students correctly while being wrong about how likely each is to need help. Where the target is a construct, fine-tuning should be paired with an instrument that measures the construct — the confidence-routing result is a good template, since it converts an accuracy number into an operating policy with a human in the loop (Know When to Trust: Making AI Scoring More Reliable for Educational Assessment).

A practical sequence for educational developers

Distilled from the evidence above, in the order that avoids wasted compute:

  1. Establish a baseline you can beat, including a trivial one. The course-assistant study's no-retrieval model scored below TF-IDF. Log your prompt-only accuracy before you consider training.
  2. Add retrieval before parameters. Grounding the model in your own materials is the cheapest large gain reported here (52.3% → 66.6%), and it is reversible.
  3. Engineer the prompt, and re-engineer it iteratively. Rubric-guided refinement moved agreement from 54.75% to 81.25% with no training at all, and item generation improved only after prompt refinement had plateaued — that plateau is your signal that training might add something.
  4. Fine-tune when the target is a stable format, a controlled property, or a domain the base model lacks. LoRA on a few thousand examples can put small open models at human-grader parity, and an instruction dataset can take a 7B model from near zero to within 0.15 of a much larger reference on domain-specific dimensions.
  5. Choose rank and layers deliberately, and expect architecture-specific behavior. Gain per adapter parameter fell as rank rose, four-layer adaptation beat both shallower and full-depth fine-tuning, and one model family improved substantially while another did not move at all.
  6. Use post-training when you need to change how the model teaches. Reward guiding over answering, and expect the reward to be gamed — write it against a benchmark, not an intuition.
  7. Supervise at the step level. Final-answer-only training produced sub-30% misconception accuracy at every data size.
  8. Keep a human in the loop where confidence is low. Routing the least confident fifth of responses removed about 80% of manual work while improving agreement.
  9. Test safety across whole conversations. Long dialogues are where specialized models degrade.
  10. Calibrate expectations. Model and prompt choice account for only about 15% of the misalignment with learning gains; some of what you want from training is not available from training.

Open Questions

  1. Does pedagogical post-training generalize across subjects, or is subject-specific tuning always needed?
  2. Can the RL–SFT–RL pipeline be combined with longitudinal memory for personalization across terms?
  3. What is the smallest supervision set that still teaches step-level pedagogical reasoning, and can it be shared across institutions without sharing student data?
  4. How should evaluation of fine-tuned educational models be standardized, given that similarity metrics can be excellent while the model is unusable?

This page is the how a model is made companion to Large Language Models (LLMs), which covers what large language models are and how they behave. It sits under Technologies as the training-and-adaptation node, beside RAG (Retrieval-Augmented Generation) (the retrieval alternative that usually comes first), Prompt Engineering (the cheapest lever), and Reinforcement Learning (the RL half of post-training). Its outputs feed Intelligent Tutoring and Adaptive Learning; its most common educational applications are Automated Assessment, Automated Essay Scoring and AI Feedback Quality; and its closest conceptual relatives are Educational NLP, Simulating Students, Open Source and Learner Modeling and Adaptive Instruction. The risks it creates are held by Pedagogical Safety, AI Sycophancy, Hallucination Risk and Privacy.

Connected Concepts

Connected Articles

Connected FAQs

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.