On this page

Synthesis: Nourashrafi, Alavinia and Darvishi (2026) use a sequential-explanatory mixed-methods design to compare the pedagogical and content quality of AI-generated versus human-developed Assessment tasks in English language teaching. Twenty experienced Language Learning teachers rated 52 assessment tasks drawn from six human-developed lesson plans (from American English File 3) and six equivalent AI-generated plans using a TPCK-aligned rubric, followed by semi-structured interviews analyzed with Braun & Clarke's thematic analysis. Quantitative results (Chi-square tests) show no statistically significant differences across the eight quality criteria: teachers preferred AI for grammar and vocabulary tasks but favored human development for reading, listening, Writing and speaking. Qualitative analysis locates the split in TPCK integration — AI shows strong technological–content (TK–CK) integration for rule-based content but limited pedagogical (PK) delivery for communicative, context-dependent skills. AI tools are positioned as useful initial resources requiring teacher mediation rather than complete replacements.

Key Findings

No significant overall difference. Across eight quality criteria (Chi-square tests of independence, all p > .05), ratings showed no statistically significant difference between AI-generated and human-developed assessment tasks — AI is already at near-parity for routine EFL task generation.

A clear division of labor by skill. Teachers preferred AI-generated grammar and vocabulary tasks (69% positive for each) but favored human-developed reading, writing, listening and speaking tasks (81%, 78%, 75%, 83% respectively). The qualitative analysis links this to TPCK integration: AI robustly integrates technological and content knowledge (TK–CK) for rule-based, objective content, but delivers limited pedagogical knowledge (PK) for communicative, context-dependent skills.

TPCK as the evaluative lens. The rubric operationalizes Technological Pedagogical Content Knowledge (TPACK) across technological, content and pedagogical knowledge domains. For AI tasks the three domains appeared dissociated rather than integrated — strong TK+CK but weak PK. Human tasks showed greater CK–PK integration, reflecting the "person-oriented" nature of communicative competence.

AI strengths. Qualitative analysis found AI's advantages lie in generating tasks quickly, producing a diverse range, and managing rule-based content (drills, substitutions, closed questions) via pattern recognition/NLP-based automatic item generation — efficiency gains relevant to Automated Question Generation and Automated Assessment.

The "complexity continuum". AI's performance gap widens with task complexity: it performs comparably to humans on closed-structure, rule-governed content, but weakens on open-ended, person-oriented, context-dependent, culturally specific skills (consistent with DeKeyser's complexity continuum). The authors suggest that more detailed, pedagogically-specified prompting may narrow this gap.

Teacher mediation essential. Personalization, cultural awareness, emotional depth and real-life applicability require human intervention — teachers viewed AI as lacking empathy and situational awareness — positioning AI as an initial resource rather than a full replacement: a human-AI complementarity model for EFL assessment design.

What this means for practice

  • Instructors. Spend AI on the tasks it wins: teacher approval was highest for AI-generated grammar and vocabulary items (69% each, against 61% and 59% for human-developed versions), so use it to produce drill, substitution, and closed-question practice, and keep your own drafting time for reading, writing, listening, and speaking tasks (where human-developed ratings reached 81%, 78%, 75%, and 83%).
  • Instructors. Treat every AI-generated task as a first draft that needs mediation. Add the personalization, cultural reference, and real-life applicability the teachers in the interviews said the tool lacked — they described AI as an assistant, not a decision maker, and themselves as the final Assessment authority.
  • Designers. Write pedagogical specifications into prompts rather than content requests alone. The TK–CK/PK dissociation behind the split suggests AI handles rule-governed, closed-structure content but weakens on open-ended, context-dependent skills, and more detailed pedagogical prompting is the authors' proposed lever for closing that gap.
  • Designers. Review AI items against all three Technological Pedagogical Content Knowledge (TPACK) domains instead of technological and content quality alone; AI tasks looked technically and content-strong but pedagogically thin, which a content-level review will not catch.
  • Administrators. Keep teacher judgment in the assessment loop: interview participants saw AI as workload relief but rejected replacement, and ratings show no significant quality advantage for AI on the communicative skills your curriculum is built around.

Limitations

  • The rating panel is 20 experienced EFL teachers (8 men, 12 women; ages 24–43), purposefully sampled because they already used AI for task design and had taught American English File 3 — a small, self-selected expert group.
  • The task corpus is 52 assessment activities drawn from just 12 lesson plans: 6 human-designed plans from one B2-level coursebook plus 6 ChatGPT-4 equivalents, so no other proficiency level, coursebook, or subject area was sampled.
  • Quality was judged from teacher ratings on a TPCK rubric, not from student performance with either set of tasks; the chi-square tests of independence reported no significant differences (p > .05) rather than demonstrating equivalence.
  • Only 11 of the 20 raters were interviewed, in sessions of roughly 30 minutes, so the qualitative account of why AI trails on communicative skills rests on just over half the sample.

Citation

Nourashrafi, F. K., Alavinia, P., & Darvishi, S. (2026). AI-Generated versus Human-Developed Assessment Tasks in EFL Context: Insights from TPCK Model. Computers and Education Open, 100415. https://doi.org/10.1016/j.caeo.2026.100415

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.