On this page

Synthesis: Wang et al. (2026) evaluate the cognitive depth of LLM-generated educational questions through a Bloom's Taxonomy lens. Across six widely-used LLMs, they find that models excel at factual recall but struggle to generate questions that stimulate higher-order thinking — a key limitation for Automated Question Generation and Formative Assessment in education, with implications for how AI is used in assessment.

LLM-generated educational questions show varying cognitive depth; models excel at factual recall but struggle with higher-order thinking questions per Bloom's taxonomy. While LLMs show promise in automating educational content creation, their ability to generate questions that stimulate higher-order thinking remains understudied. This work evaluates six widely-used LLMs through a Bloom's Taxonomy lens, focusing on their capacity to transcend rote memorization. The findings inform Automated Question Generation and assessment design, connecting to Critical Thinking and Cognitive Diagnosis.

What this means for practice

  • Instructors. Do not assume that a model which can classify a Bloom level can generate at that level: classification accuracy and Bloom-level consistency were negatively correlated, with most models performing best on "Apply" while misclassifying "Analyze" and "Create".
  • Instructors. Use fine-grained prompting rather than chain-of-thought when you need questions at a specified cognitive level: fine-grained prompting raised question usability and Bloom-level consistency and coverage across all six models, while chain-of-thought left more redundancy.
  • Instructors. Screen for upward drift before assigning generated questions: CogLeap — questions generated above the intended level — exceeded 44% of mismatched instances under chain-of-thought for every model, reaching 54.87% for LLaMA-3-8B, an abstraction bias the authors warn may overload learners.
  • Software developers. Automate screening but not the final judgment: the automated rubric agreed with expert annotations on uniqueness (90.75%), readability (98.08%), and answerability (95.83%), while Bloom-level agreement was only 46.58%.
  • Researchers. Track knowledge coverage alongside cognitive level: knowledge-identification accuracy was consistently associated with coverage of the intended knowledge units, which makes identification quality part of the question-generation pipeline rather than a side task.

Limitations

  • Cognitive levels are largely machine-assigned: the computer-science set (1,406 questions across 12 knowledge units) and the K–12 math set were annotated with a Bloom-aligned verb list rather than expert judgment, and the framework's own Bloom-level agreement with experts reached only 46.58%, so the cognitive-shift metrics inherit measurement error in the construct they report.
  • Human validation covers only a 100-item social-science subset (1,200 generated questions); the main experiment of 20,700 questions is scored by the same automated rubric the paper proposes.
  • The comparison covers six LLMs (GLM4-9B-Chat, Qwen2.5-7B-Instruct, Baichuan2-7B-Chat, InternLM3-8B-Instruct, LLaMA-3-8B-Instruct-Chinese, Spark3.5-Max), most of them Chinese-oriented models at a single point in time.
  • The K–12 math dataset contains only "Apply"-level questions, and no classroom or learning-outcome evidence is reported; the authors place classroom-based studies of whether cognitive-shift metrics correlate with learning outcomes in future work.

Citation

Xiaolong Wang, Zhe Zhao, Song Lai, Chaoli Zhang, Zijie Geng, Yu Tong, Ye Wei, Qingsong Wen (2026). From Memorization to Creation: Evaluating the Cognitive Depth of LLM-Generated Educational Questions.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.