Research Article
From Execution to Education: A Bloom-Aligned Framework for Measuring Educational Control in LLMs
Synthesis: Introduces a Bloom-aligned framework for measuring 'educational control' in LLMs: the ability to preserve a task's instructional intent while shifting its cognitive demand toward higher-order Bloom levels, offering a metric for evaluating whether AI assistance scaffolds or shortcuts learning. The work connects to broader debates about how Generative AI systems reshape Student Experience and the conditions under which AI support scaffolds rather than undermines learning. It has direct implications for The Pedagogy of AI Mistakes: Fostering Higher-Order Thinking and the risk of Over-Reliance when assistants absorb too much of the cognitive load. Findings also bear on AI Literacy and Self-Regulated Learning, and on how institutions should govern Student Experience and Academic Integrity. Practitioners in Higher Education and teachers can use the evidence to calibrate when to deploy Large Language Models (LLMs)-based help and how to pair it with Feedback that preserves learning gains.
What this means for practice
- Instructors. Do not infer educational control from execution benchmarks: the framework's premise is that a model can be a proficient problem solver while remaining a poor task designer, so evaluate candidate models with cognitive-shift and target-accuracy checks rather than pass rates alone.
- Instructors. Verify simplification requests instead of trusting them: under the general "Easier" prompt both models still produced positive cognitive shift (0.715 for the general model, 0.194 for the coder model), and accuracy against "Lower" Bloom targets fell below 30% for both.
- Instructors. Prompt for higher-order demand with reasonable confidence — the general model hit "Higher" targets 79.2% of the time and the coder model 63.4% — but inspect the generated task, since upward mutation is the models' strong direction and downward mutation their weak one.
- Learners. Treat a rewritten "easier" task as unverified: the same prompt family that reliably raises the Bloom level does not reliably lower it, so check the level yourself before studying from the regenerated task.
Limitations
- The evaluation covers two models only — Qwen3-Next-80B-A3B-Instruct (general) and Qwen3-Coder-Next (coder) — chosen because they share a 48-layer architecture and tokenizer; the authors state that validation across additional model families, scales, and training recipes remains an important next step.
- Bloom's Taxonomy serves as a structural proxy, so the framework measures task-level cognitive demand and says nothing about multi-turn tutoring dynamics or learner-specific adaptation.
- The benchmark suite is English-language and Python-centric (2,520 tasks drawn from three code benchmarks), leaving other programming languages and non-programming learning tasks unexamined.
- Augmented tasks were produced with a zero-shot prompting protocol and no human audit of the generated mutations, and the Bloom judge, even with its reported agreement against a 150-question human-validated subset, may not hold on mutations that shift away from that validation distribution.
Citation
S. Bekkouch, T. Constantinou, M. Ovaere, et al. (2026). From Execution to Education: A Bloom-Aligned Framework for Measuring Educational Control in LLMs.