On this page

Synthesis: This PRISMA-guided systematic review synthesizes 23 empirical studies (2020–2025) on customizing AI systems for writing instruction — any deliberate change to system behavior (prompt design, fine-tuning, data adaptation, interface configuration) aligning outputs with pedagogical intent. Its key result is structural misalignment: goals have moved beyond surface error correction toward writing processes, Feedback Literacy, and higher-order academic skills, yet the dominant route remains product-focused goals justified by performance-oriented theories and implemented through Prompt Engineering or fine-tuning. Automated scoring still outnumbers dialogic mediation; the limiting factor is missing alignment across intent, theory, and technique.

Key Findings

  1. Five databases (Web of Science 53, Scopus 50, ProQuest 17, ERIC 4, LLBA 2) yielded 126 records, 82 unique after duplicate removal; screening excluded 57 (25 generic applications, 8 benchmarking-only, 15 perception-only, 9 unrelated disciplinary contexts), leaving 23 studies (January 1, 2020 – September 30, 2025).
  2. Multi-dimensional feedback and scoring was the most common goal (N = 9), then writing process and metacognition (N = 6) and higher-order academic skills (N = 6); alternative assessment models and learner feedback literacy (N = 1 each) were marginal.
  3. Prompt engineering dominated methods (N = 13), ahead of fine-tuning (N = 7) and hybrid or integrated architectures (N = 3).
  4. Implementation modes were standalone automated writing evaluation (N = 8), specialized academic task tools (N = 5), conversational agents or tutors (N = 3), framework integration (N = 3), integrated writing environments (N = 2), and AI-based research metrics (N = 2).
  5. Technical rationales outpaced learning theory: NLP performance optimization guided 6 studies, against process-oriented pedagogy (N = 4), cognitive and affective learning theories (N = 4), sociocultural and interactionist theories (N = 4), applied and computational linguistics (N = 3), and writing feedback theory (N = 2).
  6. Coding recorded 25 positive outcomes and 17 challenges: benefits on writing quality and learning outcomes, learning processes, validation of AI-supported assessment, and feedback quality; challenges were technical (unstable or overgeneralized feedback, cost) and pedagogical (overreliance, reduced critical engagement, limited transparency).
  7. Cross-tabulation showed product-oriented goals clustering with performance-driven theories and weak links between learning-process, dialogic, or feedback-literacy goals and sociocultural frameworks; hybrid architectures embedding learning theory in adaptive, developmentally sequenced behavior were rare.

How the review was assembled

The review follows Keele (2007), Xiao and Watson (2019), and PRISMA. Studies qualified only if they customized model or data for pedagogical purposes, tested hybrid architectures integrating goals such as formative feedback or genre awareness, or treated Prompt Engineering as a research construct. Coding covered five dimensions (goals, customization methods, theoretical principles, implementation methods, outcomes/challenges); reliability rested on cross-checking and re-coding a random 25% subset after one week. The corpus concentrates on university-level second-language academic writing in EAP/EFL settings, though K-12 and distance-learning cases appear.

What the included studies optimized for

Zheng and Zhang's fine-tuned feedback system cut grammatical errors by 78.1% versus 31.0% in the control over 12 weeks with 162 learners. Shi compared ChatGPT with five human teachers on 110 CET-4 compositions: more feedback but lower precision (83% versus 99%) and higher recall (86% versus 55%). Tang et al. found prompt design drove reliability, the "criteria and sample-referenced justification" prompt agreeing most with human scores (GPT-4, QWK = 0.5677). Link et al.'s ELECTRA classifier hit 96.4% exact match and 77.4% precision / 79.6% recall. AES-style Automated Assessment answers the feedback bottleneck but not whether learners act on it.

CoachGPT scaffolded an 11-stage writing process, a GPT-4-turbo suite guided six Korean ninth-grade EFL students through drafting and revision, and EvaluMate judged peer review comments against five quality features. In Guo et al.'s five-week quasi-experiment with 124 Chinese undergraduates, the AI-supported group's peer feedback improved significantly on four of five dimensions and their writing more than the control's — though the chatbot never saw the essay under review and students risked copying revisions passively. A three-week WILLM study (19 participants) improved grammar and vocabulary scores and scored 84.21 on the System Usability Scale, though gains were non-significant and engagement low.

The theory–design mismatch

Customization rarely encodes the principles a study names. Process-oriented pedagogy aligned support to planning, drafting, and revising; cognitive and affective theories brought in Self-Regulated Learning, cognitive load, and Motivation; sociocultural work positioned AI as a dialogic partner. But theory lived at the interface or task-framing level, not in the model architecture, leaving behavior "largely theory-agnostic": no customization encoded contingent mediation, diagnostic sensitivity, or implicit-to-explicit progression.

Overreliance, superficial revision, and unstable or overgeneralized responses follow from optimizing output quality rather than pedagogically specified interaction. Its reframing treats customization as theoretically constrained system specification, implicating researchers, designers, and educators who judge tools by theoretical alignment and learner Learner Agency, not output quality. Four gaps define the agenda: short-term interventions with small samples; underexplored uptake of feedback; narrow genre coverage centered on argumentative writing; and closed-source models that inhibit transparency and sustainability.

What this means for practice

  • Instructors. Ask what a tool's feedback was optimized for; AI feedback matches raters on content and organization but is unstable on context-sensitive features, so learners need support to use it critically.
  • Instructional designers and curriculum teams. Treat theory as an architecture decision, not a wrapper: sequencing stages, withholding answers, and revision loops change behavior.
  • Assessment designers and administrators. Keep a human in the loop for high-stakes AI scoring: precision gaps (83% versus 99% against teachers) argue for complementary use, not replacement.
  • Educational technology developers. Build hybrid architectures that embed pedagogical rules, domain knowledge, and adaptive control rather than prompt constraints.
  • Researchers. Pursue longitudinal, multi-cycle designs, fine-grained analysis of dialogue logs and revision histories, and how learners take up AI guidance.

Limitations

  • No backward or forward citation searching: coverage rests on the five-database search (126 records), so work outside Web of Science, Scopus, ProQuest, ERIC, and LLBA may be missed.
  • Reliability rests on intra-coder consistency (a random 25% subset re-coded after one week); no second coder or inter-coder agreement statistic is reported.
  • The 23-study corpus reflects strict inclusion criteria and concentrates on higher education L2 academic writing, describing a bounded corpus.
  • Then-current models (GPT-3.5 Turbo, GPT-4, GPT-4o, GPT-4-turbo, GPT-3, BERT-family) and a search closed September 30, 2025 are superseded generations, so capability constraints behind reported challenges may no longer hold.

Citation

Luo, Y. (2026). Customizing AI for writing pedagogy: a systematic review of pedagogical goals, theoretical principles, and technical design. Educational Technology Research and Development.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.