On this page

Synthesis: This quasi-experimental mixed-methods longitudinal study (N=142) deploys a fully automated marking pipeline for handwritten mock examinations in A-Level sciences, removing the human-marking bottleneck that normally caps the frequency of formative mocks. High-frequency, automatically-marked Formative Assessment cycles were associated with improved student outcomes, providing field evidence for the pedagogical payoff of Automated Grading at upper-secondary level. The handwritten-work pipeline links to Automated Grading of Handwritten Mathematics Using Vision-Capable LLMs and Confidence Estimation in Automatic Short Answer Grading with LLMs, while the human-in-the-loop trade-offs echo Hybrid E-Assessment in Higher Education: Semi-Automated Grading of Paper-Based Written Examinations.

What this means for practice

  • Instructors. Replace two termly human-graded mocks with weekly or bi-weekly automatically marked practice: the intervention cohort completed 4.1 times the practice volume and scored an adjusted 81.2% on the final mock against 64.4% for the control group.
  • Act on the granular annotations within the two-hour return window, since the study attributes "error-type decay" — students correcting procedural misconceptions weeks before their peers — to step-by-step marking at the precise point of failure.
  • Keep the assessments handwritten and blind-graded even though feedback is automated, as high-stakes conditions and external marking were what preserved Assessment Validity in this design.
  • Designers. Prioritize error-carried-forward logic and marginalia that locate the specific failing step over blanket comments, because intervention students favored granular annotations to general remarks like "check your working."
  • Log turnaround time, submission frequency, and practice volume as first-class outputs; the reported gains are inseparable from evidence that higher-frequency assessment is actually reaching students.

Limitations

  • Intact classes were assigned to conditions rather than individual students (control N = 70, intervention N = 72, aged 16–18) at one international secondary school in Europe, so the quasi-experimental design cannot rule out cohort or timetable effects.
  • The design cannot separate the quality of automated feedback from the 4.1-fold increase in practice volume; the authors state it remains unknown whether high-frequency human feedback would produce identical results.
  • Exam confidence, task-specific self-efficacy, and trust in automated feedback were measured by 5-point Likert self-report surveys rather than behavioral measures.
  • The automated marking pipeline is a proprietary system supplied free of charge by its developer (Cortex Global) for the study period, so the marking method is not reproducible from the reported details alone.

Citation

Matey Yordanov, Mikhail Bychkov, Andrei Kuchma (2026). The Effect of High-Frequency, Automatically-marked Formative Assessments on Student Outcomes in A-Level Sciences.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.