On this page

Synthesis: Lee, Atif, and Kang (2026) reframe learner-generated questions as diagnostic signals of epistemic engagement in AI-mediated learning, classifying 434 authentic student queries from 12 information technology courses into three constructivist instructional roles: knowledge transmitter, facilitator, and co-learner. After human-in-the-loop to reach consensus (Fleiss' kappa rising from 0.60 to 0.83) and augmentation to 582 balanced questions, four transformer models of the NLP family — baseline BERT, fine-tuned BERT, DeBERTa, and RoBERTa — were compared. DeBERTa achieved the highest overall accuracy (86.36%) and excelled on factual knowledge-transmitter items (96.67% precision), but all models struggled to separate higher-order facilitator and co-learner questions, where subtle contextual nuance and implicit intent proved decisive. The authors ground this gap in Constructivism theory (Piaget, Vygotsky, Freire) and cognitive load theory: generative AI handles low-complexity transmission reliably while faltering at the Scaffolding that the zone of proximal development demands. The study's contribution is practical as much as technical — an LMS-embedded classification-and-feedback loop in which AI handles initial categorization while instructors validate outputs, so that personalized feedback, curriculum diagnostics, and human pedagogical judgment remain aligned.

Key Findings

  • DeBERTa led overall classification: DeBERTa reached 86.36% accuracy, 87.18% precision, 86.36% recall, and an 86.52% F1 score, outperforming fine-tuned BERT (84.09% accuracy, 84.14% F1), RoBERTa (82.95% accuracy, 83.06% F1), and baseline BERT (81.82% accuracy, 81.95% F1).
  • Factual questions are easy; higher-order questions are not: DeBERTa classified knowledge-transmitter (level 1) questions with 96.67% precision, but its level 2 (facilitator) precision fell to 78.79% — facilitator questions were the hardest category for every model.
  • Fine-tuned BERT excelled at co-learner recall, at a precision cost: Fine-tuned BERT achieved 92.00% recall on level 3 (co-learner) questions but only 74.19% precision, indicating semantic overlap between the co-learner and facilitator categories.
  • Label quality improved through iteration: Expert labeling of 434 questions yielded moderate reliability (Fleiss' kappa = 0.60; complete agreement on 264 questions), improving to kappa = 0.83 (complete agreement on 361) after a second consensus round.
  • Class imbalance was corrected by augmentation: The original distribution (196 knowledge transmitter, 135 facilitator, 103 co-learner) was balanced to 582 questions (196 / 193 / 192) using back-translation and paraphrasing.
  • Three recurring error patterns emerged: Conceptual similarity between roles, ambiguous learner intent (co-learner questions read as facilitator prompts), and technical complexity — where domain-specific phrasing was misread as cognitive depth.
  • Augmentation may have introduced lexical shortcuts: The authors note that paraphrasing and back-translation may have pushed models to rely on lexical patterns rather than semantic depth.
  • Human oversight remains indispensable: Misclassification between facilitator and co-learner roles means a fully automated system risks propagating errors, supporting a hybrid model with expert validation.

Study Design & Method

  • Data source: 434 authentic learner-generated questions collected as assignment appendices across 12 undergraduate and graduate information technology courses at an Australian university, spanning programming, data analytics, and software engineering; students used generative AI tools (e.g., ChatGPT) for brainstorming, idea refinement, and solution verification.
  • Participants: 11 students enrolled in undergraduate and master's degree programs in IT, contributing questions that ranged from simple informational queries to complex Problem Solving prompts.
  • Expert labeling: Three external doctoral-level experts with more than 5 years of Constructivism research experience conducted multi-phase focus group labeling, refining operational definitions (Phase 1), labeling independently with Fleiss' kappa checks (Phase 2), and validating the augmented set (Phase 3).
  • Augmentation: Back-translation and paraphrasing expanded underrepresented categories, producing a balanced 582-question corpus reviewed by experts for alignment with role definitions.
  • Model comparison: Baseline BERT (no fine-tuning), fine-tuned BERT (last three layers updated), DeBERTa, and RoBERTa, trained with a 70/15/15 train-validation-test split, 128-token maximum length, batch size 8, learning rate 2e-5, ReduceLROnPlateau scheduling, and early stopping (patience = 20 epochs) on Python 3.8, PyTorch 1.12.0, CUDA 11.2, and an NVIDIA RTX3070 GPU.
  • Evaluation and error analysis: Accuracy, precision, recall, F1, and confusion matrices for quantitative comparison, followed by manual qualitative review of misclassified questions to derive NLP failure patterns.
  • Ethics: Anonymised, non-identifiable student data with informed oral consent; participation was voluntary and unlinked to assessment, and institutional guidance classified the study as low-risk, requiring no formal ethics review.

Implications for AI in Education

  • Embed question classification in the LMS, with humans in the loop: A Formative Assessment system built into the LMS can classify learner questions in real time, but instructors should validate outputs — a collaborative model that combines AI throughput with pedagogical judgment.
  • Use inquiry depth as a diagnostic signal: Distinguishing knowledge-transmitter, facilitator, and co-learner questions gives educators a structured view of epistemic engagement and lets personalized responses escalate when a learner keeps asking only low-complexity factual questions.
  • Prompt deeper questioning deliberately: If a student consistently asks factual questions, the system can nudge them toward reflective, critical inquiries, developing Metacognition and self-regulation.
  • Aggregate question data for curriculum decisions: Patterns of questions reveal topics where students struggle, letting instructors adjust content, sequence, or support before end-of-term assessment exposes gaps.
  • Plan for what NLP cannot yet do: Because models confuse higher-order facilitator and co-learner questions, deployments should pair classification with explainable outputs, bias checks, and ethical guidelines rather than relying on automation alone.
  • Broaden the evidence base beyond IT: The corpus is confined to IT courses, so generalizing to humanities, social sciences, and STEM requires cross-disciplinary and longitudinal validation before institutional-scale adoption.

Connected Concepts

Connected Articles

Citation

Lee, H., Atif, A., & Kang, K. (2026). Analysing AI utilisation in education through learner question types: A constructivist approach. Australasian Journal of Educational Technology, 42(2), 77–94.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.