🧠 AI Ed Wiki

Introduces 'design problems' (DPs): concise, scenario-based prompts that require applying knowledge in transfer contexts, generated with LLMs to assess higher-order thinking (HOT) in project-based learning. Traditional PjBL assessments often fail to capture HOT, especially transfer; DPs target that gap.

Bridges Generative AI generation with Formative Assessment and Active Learning, linking to Scaffolding of complex tasks and Higher Ed/CS Education contexts. It contributes a concrete method for scaling HOT assessment and informs Generative AI used for evaluation rather than just content delivery.

Key Findings

  • Design problems (DPs) are concise, scenario-based prompts that require applying project concepts in new situations, targeting higher-order thinking (HOT) that traditional PjBL assessments often fail to capture, especially in transfer contexts.
  • Surveys of 31 instructors showed that instructors value DPs for assessing HOT, but creation effort is a barrier to adoption.
  • An evaluation of 80 LLM-generated DPs showed LLMs can produce high-quality prompts with strong expert agreement, effectively lowering the creation barrier.
  • Students rated DPs from different LLMs similarly, and their performance on DP tasks showed negligible correlation with traditional project grades, suggesting DPs capture distinct aspects of higher-order thinking rather than duplicating existing measures.
  • Keystroke data suggested deeper cognitive engagement through planning and revision behaviors while students worked on DP tasks.
  • Study Design & Method

    The study triangulates three perspectives: instructor perceptions (surveys with 31 instructors), LLM generation capability (80 generated DPs evaluated for quality and expert agreement), and student experience (performance data plus keystroke logs). The negligible correlation between DP performance and traditional project grades is the key psychometric signal: it indicates the assessment captures a different construct — transfer-oriented higher-order thinking — than project artifacts alone.

    Implications for AI in Education

    DPs appear to be a useful complement to traditional assessments, particularly in situations where AI use or collaboration may undermine individual learning: because DPs demand application in novel scenarios, they are harder to complete by simply reusing project artifacts or generated code. The strong expert agreement on LLM-generated prompts makes scalable HOT assessment feasible, and the keystroke evidence connects DP work to deeper engagement. This positions LLM-generated DPs as a practical instrument for Formative Assessment and for evaluating transfer in project-based computing education.

    Connected Concepts

  • Generative AI
  • Formative Assessment
  • Active Learning
  • Scaffolding
  • Higher Ed
  • CS Education
  • Connected Articles

  • AI Generated Feedback Higher Ed — Artificial intelligence and feedback in university education: effectiveness and student perceptions
  • Hybrid E Assessment Semi Automated Grading — Hybrid E-Assessment in Higher Education: Semi-Automated Grading of Paper-Based Written Examinations
  • Student Misconceptions Conditionals Loops Taxonomy — How Students (Mis)understand Conditionals and Loops -- A Taxonomy
  • Slidesqaqa Pedagogical Question Generation — Slide Deck Q&A Quality Assurance App: A Multi-Stage Pipeline for Pedagogical Question Generation
  • Mllm Scientific Visualization Literacy — Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy
  • Lata Ferpa Compliant Local LLM Autograder — LaTA: A Drop-in, FERPA-Compliant Local-LLM Autograder for Upper-Division STEM Coursework
  • Citation

    Ahmad D. Suleiman, Daqing Hou, Maliha Noushin Raida (2026). LLM-Generated Design Problems for Assessing Higher-Order Thinking in Project-Based Learning. arXiv:2607.11032. arXiv preprint.