AI Ed Wiki logoAI Ed WikiUse with AI

Synthesis: Nourashrafi, Alavinia and Darvishi (2026) use a sequential-explanatory mixed-methods design to compare the pedagogical and content quality of AI-generated versus human-developed Assessment tasks in English language teaching. Twenty experienced Language Learning teachers rated 52 assessment tasks drawn from six human-developed lesson plans (from American English File 3) and six equivalent AI-generated plans using a TPCK-aligned rubric, followed by semi-structured interviews analysed with Braun & Clarke's thematic analysis. Quantitative results (Chi-square tests) show no statistically significant differences across the eight quality criteria: teachers preferred AI for grammar and vocabulary tasks but favoured human development for reading, listening, Writing Education and speaking. Qualitative analysis locates the split in TPCK integration — AI shows strong technological–content (TK–CK) integration for rule-based content but limited pedagogical (PK) delivery for communicative, context-dependent skills. AI tools are positioned as useful initial resources requiring teacher mediation rather than complete replacements.

Key Findings

No significant overall difference. Across eight quality criteria (Chi-square tests of independence, all p > .05), ratings showed no statistically significant difference between AI-generated and human-developed assessment tasks — AI is already at near-parity for routine EFL task generation.

A clear division of labour by skill. Teachers preferred AI-generated grammar and vocabulary tasks (69% positive for each) but favoured human-developed reading, writing, listening and speaking tasks (81%, 78%, 75%, 83% respectively). The qualitative analysis links this to TPCK integration: AI robustly integrates technological and content knowledge (TK–CK) for rule-based, objective content, but delivers limited pedagogical knowledge (PK) for communicative, context-dependent skills.

TPCK as the evaluative lens. The rubric operationalises TPACK across technological, content and pedagogical knowledge domains. For AI tasks the three domains appeared dissociated rather than integrated — strong TK+CK but weak PK. Human tasks showed greater CK–PK integration, reflecting the "person-oriented" nature of communicative competence.

AI strengths. Qualitative analysis found AI's advantages lie in generating tasks quickly, producing a diverse range, and managing rule-based content (drills, substitutions, closed questions) via pattern recognition/NLP-based automatic item generation — efficiency gains relevant to Automated Question Generation and Automated Assessment.

The "complexity continuum". AI's performance gap widens with task complexity: it performs comparably to humans on closed-structure, rule-governed content, but weakens on open-ended, person-oriented, context-dependent, culturally specific skills (consistent with DeKeyser's complexity continuum). The authors suggest that more detailed, pedagogically-specified prompting may narrow this gap.

Teacher mediation essential. Personalization, cultural awareness, emotional depth and real-life applicability require human intervention — teachers viewed AI as lacking empathy and situational awareness — positioning AI as an initial resource rather than a full replacement: a human-AI complementarity model for EFL assessment design.

Connected Concepts

Connected Articles

Citation

Nourashrafi, F. K., Alavinia, P., & Darvishi, S. (2026). AI-Generated versus Human-Developed Assessment Tasks in EFL Context: Insights from TPCK Model. Computers and Education Open, 100415. https://doi.org/10.1016/j.caeo.2026.100415