🏷️ assessment
51 pages tagged with assessment(40 articles, 11 concepts)
📄 Knowledge, Skills, Attitudes, Production: Competency-Based Education After Generative AI
> **Synthesis:** This conceptual paper proposes adding *production* — the capability to deliver professional-standard work by directing tools and other people — as a fourth attribute of competency-bas…
2026-08-12 · assessment-validity, academic-integrity, generative-ai, higher-ed, automated-assessment
🏷️ AI Misuse and Learning Harm
> **AI misuse and learning harm** — the causal relationship between students offloading cognitive work to generative AI and reduced durable learning, even when immediate task performance rises. The de…
2026-08-12 · over-reliance, cognitive-offloading, academic-integrity, self-regulated-learning, motivation
🏷️ Cognitive Diagnosis
> **Cognitive diagnosis** — the inference of a learner's latent knowledge state — the specific concepts, skills, and misconceptions they have or lack — from their responses or behavior. It is the asse…
🏷️ Reducing AI Misuse
> **Reducing AI misuse** — the design, pedagogical, and policy levers that prevent students from substituting generative AI for their own cognitive work and instead steer them toward ethical, producti…
📄 Rethinking Elementary Education's Writing Instruction in The Age of Generative AI: A Systematic Review
> **Synthesis:** This systematic literature review synthesizes 8 peer-reviewed studies (2019–2025) on AI literacy for elementary writing instruction, finding that AI integration efficiently supports w…
📄 Generative AI in Higher Education: A Systematic Review of Opportunities, Challenges, and Pedagogical Innovations (2022–2025)
> **Synthesis:** This PRISMA-guided systematic review synthesizes 125 peer-reviewed studies (2022–2025) on generative AI in higher education, documenting exponential adoption (92% student usage by 202…
📄 Can Large Language Models Foster Critical Thinking, Teamwork, and Problem-Solving Skills in Higher Education?: A Literature Review
> **Synthesis:** Can Large Language Models Foster Critical Thinking, Teamwork, and Problem-Solving Skills in Higher Education?: A Literature Review…
📄 From Evaluated Models to Evaluation Aids: A Multi-Evidence Study of LLM-Based Difficulty Calibration for Programming Examinations
> **Synthesis:** Yan, Xiong, Li & Chen (2026) reposition LLMs from benchmark targets to auxiliary evidence sources for interpreting programming-exam difficulty, showing that AI difficulty estimates co…
📄 AI literacy alone is not enough: Student AI readiness and career adaptability in business and management education
> **Synthesis:** Testa, Apuzzo, and Pittaway (2026) investigate how AI-related competencies contribute to career adaptability in business and management education. Surveying 339 university students in…
📄 Coauthorship integrity: Reconceptualising assessment validity for the age of generative artificial intelligence
> **Synthesis:** This paper addresses concerns that students use GenAI to submit texts they do not understand, adopting an assessment validity lens. It proposes Coauthorship Integrity as a new concept…
2026-08-10 · generative-ai, conversational-agents, assessment-validity, academic-integrity, ai-education
📄 Generative AI interactive textbook in electrotechnics: A four-year comparative study on student performance and inclusion
> **Synthesis:** This four-year comparative study presents results of implementing a Generative-AI Interactive Textbook built on GPT-4, integrated into an Electrical Engineering course. With a sample …
📄 Learning with machines: Toward a theory of epistemic co-agency
> **Synthesis:** This paper introduces the Epistemic Entanglement Framework, a theory-informed model capturing how learners engage with AI systems. Drawing on distributed cognition, sociomaterialism, …
📄 Reimagining feedback through generative AI in engineering education
> **Synthesis:** This study investigates the capacity of a large language model to generate formative feedback for student-created UML diagrams in a university software engineering course. Across two …
📄 Students' engagement with generative AI in academic learning: A self-determination theory and epistemic network analysis study
> **Synthesis:** This qualitative case study examines undergraduate students' engagement with GenAI in academic learning using self-determination theory and epistemic network analysis. Data from 23 se…
📄 Teaching Intro AI When the Tools Can Do the Homework: A Course Redesign and a Student Bill of Rights
> **Synthesis:** This experience report describes the redesign of an introductory AI course at the University of Washington Bothell in response to LLMs being able to complete most assignments. The red…
📄 Anchor Is the Key: Toward Accessible Automated Essay Scoring with Large Language Models Through Prompting
> **Synthesis:** Choi, Tate, Ritchie, Nixon & Warschauer (2025) investigate the most practical approach to LLM-based automated essay scoring — prompting — and find that providing anchor papers (exampl…
🏷️ Automated Essay Scoring
> **Automated Essay Scoring (AES)** — the use of AI to evaluate and score written essays, spanning traditional statistical approaches, fine-tuned language models, and increasingly accessible LLM-based…
🏷️ Benchmark
> **Benchmark** — standardized test suites and evaluation frameworks used to measure AI model performance on educational tasks. Benchmarks enable reproducible comparison across models and approaches, …
🏷️ Learning Gains
> **Learning gains** — measurable improvements in student knowledge, skills, or competencies resulting from educational interventions, including AI-assisted instruction. In AI in education research, l…
📄 Using LLMs to Detect Growth in Computational Thinking in Introductory Physics
> **Synthesis:** Savage, Shanker, Michlitsch & Rebello (2026) investigate using LLMs to evaluate students' written explanations of computational physics problems at scale. Establishing a human-coded b…
📄 What out-of-the-box LLMs can(t) do in law? A Turing test in Italian exams for lawyers, judges and notaries
> **Synthesis:** This paper reports a blind Turing Test evaluating leading LLMs on three Italian professional legal examinations: the Bar exam, Judges exam, and Notary exam. LLMs generated full writte…
📄 Behaviorally Adaptive Visual Diversion for Inclusive and Resilient Digital Assessment Delivery
> **Behaviorally Adaptive Visual Diversion for Inclusive and Resilient Digital Assessment Delivery** — Proposes BAVD, a theoretical framework for adaptive visual diversion in digital assessment that r…
📄 CODE-GEN: A Human-in-the-Loop RAG-Based Agentic AI System for Multiple-Choice Question Generation
> **A dual-agent RAG-based system for generating and validating coding comprehension MCQs**, evaluated by 6 SMEs across 7 pedagogical dimensions (N=288 questions, 2,016 rating pairs). AI excels at cri…
📄 OECD Digital Education Outlook 2026
> **OECD flagship report** synthesising empirical evidence and expert insights on generative AI in education. Central finding: general-purpose AI chatbots improve task performance but produce no durab…
📄 From authentic products to authenticated processes: authentic assessment in AI-rich higher education
> Generative AI has not created the need for authentic assessment — it has made weaknesses in assessment design harder to ignore. Polished products can now be generated or substantially mediated by to…
📄 Beyond Detection: redesigning authentic assessment in an AI-mediated world
> Detection-led responses face well-documented limits: validity and fairness failures (bias against non-native writers), notable error rates, erosion of trust, and distraction from assessment design. …
📄 Responsible Assessment in the AI Era: Key Insights from a Future-Focused Conference
> **Responsible assessment in the AI era** — assessment grounded in learners' sociocultural contexts and designed to generate valid, trustworthy, context-specific inferences from accumulated evidence,…
📄 The Theoretical Foundation of Socratic Tests: Dynamic, Multimodal, Conversational Examinations
> **Ilya Mikhelson** — Submitted to Computers and Education: Artificial Intelligence (2026).…
📄 Confidence-Aware Automatic Short Answer Grading
> **Confidence-Aware ASAG** — A hybrid confidence estimation framework for Automatic Short Answer Grading with LLMs that fuses model-based confidence signals (verbalized, latent, consistency-based) wi…
📄 The Effect of High-Frequency, Automatically-marked Formative Assessments on Student Outcomes in A-Level Sciences
This quasi-experimental mixed-methods longitudinal study (N=142) deploys a fully automated marking pipeline for handwritten mock examinations in A-Level sciences, removing the human-marking bottleneck…
📄 Make or Take: How Students Navigate Self-Created and Instructor-Provided Cheat Sheets
Chen, Sakhnini and Istead run a three-wave longitudinal study in a senior software-requirements course where students could use instructor-provided or self-created cheat sheets in exams. Choices were …
📄 Navigating the moral panic: encouraging appropriate use of GenAI in the classroom rather than condemning innovation as disruption
> **Jennifer M. Krebsbach & Victoria L. Cross (University of California, Davis)** — *Assessment & Evaluation in Higher Education* (Taylor & Francis). Open Access, CC BY 4.0. doi:10.1080/02602938.2026.…
📄 A Tool-Invariant Framework for Teaching and Assessing Computational Methods in the Age of Agentic AI
> **Larry Engelhardt (Francis Marion University)** — *arXiv:2607.15518* [physics.ed-ph], submitted 17 Jul 2026. CC BY 4.0. doi:10.48550/arXiv.2607.15518. > **Note on type:** This is a *framework / pos…
📄 Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction
This paper introduces Epi2Diff (Episode to Difficulty), a framework that maps LLM reasoning traces into cognitively grounded episode sequences for predicting human item difficulty in [[assessment|educ…
📄 A bit of chaos and madness: The AI Assessment Scale and the work of assessment reform
📄 [PDF](https://arxiv.org/pdf/2606.26729) This study examines the implementation of the Artificial Intelligence Assessment Scale (AIAS), a structured framework for redesigning [[assessment|university…
📄 Why Machines Misread Pedagogical Quality: Human-Machine Alignment in LLM-Based Pretest Question Evaluation
Tseng et al. (2026) investigate human-machine alignment in LLM-based pretest question evaluation — a critical bottleneck for scalable AI-assisted assessment. Their AI-assisted workflow combines automa…
📄 Test-Driven, AI-Assisted Learning: Replacing Lectures with Weekly Closed-Book Tests
Liu et al. (2026) report on a 13-week Test-Driven, AI-Assisted (TDAA) redesign of a Theory of Computation course at HKUST (Guangzhou). The course replaced all lectures with self-directed, AI-assisted …
📄 From Memorization to Creation: Evaluating the Cognitive Depth of LLM-Generated Educational Questions
LLM-generated educational questions show varying cognitive depth; models excel at factual recall but struggle with higher-order thinking questions per Bloom's taxonomy. From Memorization to Creation: …
📄 LearnOpt: Recovering the Latent Cognitive Structure of Standardized Examinations via Knowledge Graphs and Constrained Optimization
Standardized examinations are typically treated as uniform syllabus coverage problems. LearnOpt recovers stable latent cognitive structures diverging systematically from official syllabi, using LLM-ta…
📄 Hybrid E-Assessment in Higher Education: Semi-Automated Grading of Paper-Based Written Examinations
**Hartwig Grabowski, Michael Canz** — cs.AI, cs.CV, cs.CY This paper identifies the didactic narrowing caused by fully digital e-assessment (overuse of closed question formats) and proposes a hybrid a…
🏷️ AI Ed Evaluation
> **AI-ed evaluation** — the body of methods, benchmarks, and criteria used to assess whether AI education tools (LLM-based tutors, automated graders, feedback systems, agents) actually work — not jus…
🏷️ AI Plagiarism Detection
Technologies and methods for detecting AI-generated content in academic submissions, including classifier-based approaches, watermarking, and stylistic analysis. The effectiveness and reliability of t…
📄 Reinforcement Learning Measurement Model
Interactive assessments generate sequential process data that conventional item response models (IRT) cannot adequately handle. This paper proposes a **reinforcement learning measurement model** that …
📄 Understanding Student Effort Using Response-Time Propensities During Problem Solving
Adaptive learning systems produce substantial learning gains, yet many students engage too briefly or superficially to benefit. This paper addresses the central challenge of **measuring student effort…
📄 AI Literacy Assessment: Self-Reported vs Performance Misalignment
>Highlights critical misalignment between self-reported AI literacy and actual performance. Teachers overestimate their AI skills by 40% on average. Performance-based assessments correlate better (r=0…
🏷️ Automated Question Generation
Automated question generation leverages NLP and LLMs to create educational assessments at scale. Wei & Stamper (2025) introduced the **generate-then-validate** paradigm, reducing hallucination by 62% …
📄 Authentic Assessment
> Wiggins (1990) proposed AA as a counterbalance to standardised tests: direct examination of "student performance on worthy intellectual tasks." > Authentic assessment (AA) has evolved from workplace…
📄 Educational VLM Evaluation
> Benchmarking vision-language models (VLMs) not on their ability to solve problems, but on their ability to *support learners* — particularly struggling learners and those making errors. Traditional …
📄 Multimodal Learning with Generative AI
> The guide adopts a middle way between "techno-fixing" and rejecting AI as an existential threat. It argues that: > A comprehensive educator's guide to integrating Generative AI into multimodal teach…
🏷️ Formative Assessment in AI Education
Assessment designed to inform ongoing instruction and learning, as opposed to summative evaluation. AI systems can generate, validate, and adapt formative assessment items at scale, though quality var…
🏷️ Human-in-the-Loop AI for Education
Educational AI systems that strategically interleave automated generation with human expert judgment, preserving pedagogical quality while scaling production. Two recent implementations illustrate dis…