📄 Research Article
Cross-Dataset Bloom Question Classification: Supervised Models and Prompted LLMs
Evaluates cross-dataset generalization of ML/DL methods and LLMs for automatic Bloom's taxonomy classification of assessment questions across five datasets. Supervised ML/DL models degraded substantially on unseen datasets, while LLMs with tailored prompting (in-context examples + course-specific action verbs) showed stable performance. A lightweight UI was developed for instructors to classify large question banks, with usability study indicating low workload and high usability.
Key Findings
Study Design & Method
The motivation is practical: Bloom's taxonomy supports the systematic design, analysis, and alignment of instructional activities and assessments, but manually classifying assessment questions is time-consuming, especially for large item banks or repeated course offerings. The study compares two families of approaches — supervised ML/DL models and prompted LLMs — under cross-dataset conditions, moving beyond the within-dataset evaluations that dominated prior work. Because labeling is subjective and teacher-dependent, the authors also assessed how prompting strategies could be tailored (in-context examples, course-specific action verbs), and they validated the instructor-facing tooling with a usability study.
Implications for AI in Education
For instructors and institutions, the results suggest that LLM-based classification with tailored prompting is a more portable approach than training supervised models for Bloom's taxonomy labeling, reducing the burden of maintaining dataset-specific models. The lightweight UI demonstrates a realistic deployment path for classifying large question banks, supporting Formative Assessment and Automated Assessment workflows while keeping the instructor in control. The finding that supervised models do not transfer across datasets is also a cautionary lesson for Educational NLP generally: strong within-dataset results should not be assumed to generalize, and evaluation designs should include cross-dataset conditions. The work connects to Teacher Role discussions about how AI can shoulder routine classification labor so that instructors focus on higher-level design and feedback.
Connected Concepts
Connected Articles
Citation
Abdolali Faraji, Mohammadreza Molavi, Zohreh Rasoulkhani, Mohammadreza Tavakoli, Gábor Kismihók (2026). Cross-Dataset Bloom Question Classification: Supervised Models and Prompted LLMs. arXiv:2606.13684. AIED 2026.