Abdolali Faraji, Mohammadreza Molavi, Zohreh Rasoulkhani, Mohammadreza Tavakoli, Gábor Kismihók (2026) — AIED 2026 📄 Full text (arXiv)
Evaluates cross-dataset generalization of ML/DL methods and LLMs for automatic Bloom's taxonomy classification of assessment questions across five datasets. Supervised ML/DL models degraded substantially on unseen datasets, while LLMs with tailored prompting (in-context examples + course-specific action verbs) showed stable performance. A lightweight UI was developed for instructors to classify large question banks, with usability study indicating low workload and high usability.
Key Contributions
- LLMs with tailored prompting generalize better than supervised models for cross-dataset Bloom's taxonomy classification of assessment questions.
Related Pages
- ai-k12-evidence-base — Empirical evidence on AI in K-12 education outcomes
- intelligent-tutoring — Automated tutoring systems and their evaluation
- student-experience — Student perspectives on AI in education
- learning-analytics — Data-driven approaches to understanding learning
- llm-feedback-programming-classroom — LLM feedback in classroom settings
Citation
APA: Abdolali Faraji, Mohammadreza Molavi, Zohreh Rasoulkhani, Mohammadreza Tavakoli, Gábor Kismihók (2026). Cross-Dataset Bloom Question Classification: Supervised Models and Prompted LLMs. arXiv:2606.13684. AIED 2026.