On this page

Synthesis: Auto-marking short answer questions in science: The foundational years of transformer-based models from BERT to GPT-4 — Morley, Walland, and Vidal Rodeiro scoping-review the transformer-based auto-marking of short-answer science questions during the foundational period 2017–early 2024, coding 21 articles under PRISMA-ScR guidelines. BERT models (and variants) dominated, peaking in 2021 before GPT-based approaches arrived, while models augmented with domain-specific data such as textbooks and marking rubrics consistently outperformed those without. The review surfaces enduring threats to reliability, explainability, and fairness and argues auto-markers should support rather than replace human examiners.

Key Findings

  • Model landscape and temporal shift. BERT (base and variants such as RoBERTa, DistilBERT, SciBERT) was the most common transformer approach, used in 20 of 21 studies and peaking in 2021; GPT models (GPT-1/2, GPT-3.5, GPT-4) appeared from roughly 2022, adopted via Prompt Engineering rather than fine-tuning, reflecting a field-wide move from fine-tuning smaller models toward prompting larger LLMs.
  • Datasets and performance. The SciEntsBank corpus dominated (2-way, 3-way, and 5-way items with unseen-answers/questions/domains tests), alongside ASAP SAS, Beetle, BEAR, TIMMS and others — all US-collected. BERT models generally outperformed earlier auto-markers on SciEntsBank, though no study yet benchmarked GPT models there; one direct comparison found GPT-3.5 outperformed BERT base across items.
  • Domain data and data augmentation help. Models incorporating additional domain knowledge — textbooks, marking rubrics, science journals, further pre-training, or meta-learning — consistently outperformed models without it; GPT-generated synthetic data also improved fine-tuned BERT performance, while rubric-aware and chain-of-thought prompting lifted GPT accuracy.
  • Reliability and validity concerns. Auto-markers can learn "spurious correlations" with surface features (punctuation, grammar, wording) rather than construct-relevant scientific understanding, and small input changes can produce large output differences, threatening reliability and construct validity.
  • Explainability and bias gaps. Few models could justify marks in human-comprehensible terms, GPT "rationale" generation and chain-of-thought prompting only partially address this, and bias across demographic and linguistic groups was rarely examined — undermining Trust and raising ethical stakes, especially in high-stakes settings.
  • Recommendations for the field. The authors call for more diverse and shared datasets, comprehensive evaluation frameworks (reliability, validity, fairness, explainability, robustness, practical utility), transparent and hybrid human-machine scoring models, rigorous bias analyses, and movement beyond simple classification toward multi-mark-point and partial-credit scenarios.

Connected Concepts

Connected Articles

Citation

Morley, F., Walland, E., & Vidal Rodeiro, C. (2026). Auto-marking short answer questions in science: The foundational years of transformer-based models from BERT to GPT-4. International Journal of Artificial Intelligence in Education, 36, Article 100005.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.