🏷️ assessment-validity
26 pages tagged with assessment-validity(22 articles, 4 concepts)
📄 The AI Literacy Heptagon: A Structured Approach to AI Literacy in Higher Education
> **Synthesis:** Hackl, Müller, and Sailer (2026) present the AI Literacy Heptagon, a structured seven-dimensional framework for AI literacy (AIL) in higher education, developed through an integrative…
📄 The Competence Paradox: Negotiating Ease, Risk, and Creative Identity in Text-to-Image Generative AI Use Among Art and Design Students
> **Synthesis:** Liu, Meng, and Zhang (2026) examined technology acceptance of text-to-image (T2I) generative AI in art and design education from both educators' and students' perspectives, using a mo…
📄 Knowledge, Skills, Attitudes, Production: Competency-Based Education After Generative AI
> **Synthesis:** This conceptual paper proposes adding *production* — the capability to deliver professional-standard work by directing tools and other people — as a fourth attribute of competency-bas…
📄 Coauthorship integrity: Reconceptualising assessment validity for the age of generative artificial intelligence
> **Synthesis:** This paper addresses concerns that students use GenAI to submit texts they do not understand, adopting an assessment validity lens. It proposes Coauthorship Integrity as a new concept…
🏷️ Academic Integrity
> **Academic integrity** — the ethical framework governing honest academic work in the age of AI. The wiki documents how the concept has been reframed by generative AI: from a problem of detecting dis…
2026-08-09 · ai-literacy, plagiarism-detection, authentic-assessment, educational-policy-ai, regulation
🏷️ Automated Assessment
> **Automated assessment** — the use of AI to evaluate student work, from formative quizzes to high-stakes exams. Automated assessment spans multiple modalities — multiple-choice, short answer, essay,…
🏷️ Automated Grading
> **Automated grading** — AI systems that evaluate student work, from multiple-choice scoring to essay assessment and code review. Automated grading is one of the most mature and widely-deployed AI in…
📄 Comparative Validation of GPT-4o-mini and Teacher Mean Scores for Automated Scoring of Music Analysis Responses: Single-Pass Deployment, Repeatability, and Strategy-Specific Bias
> **GPT-4o-mini can produce stable rubric-based scores for open-ended music analysis responses, with few-shot chain-of-thought prompting agreeing most strongly with teacher means while RAG systematica…
📄 From authentic products to authenticated processes: authentic assessment in AI-rich higher education
> Generative AI has not created the need for authentic assessment — it has made weaknesses in assessment design harder to ignore. Polished products can now be generated or substantially mediated by to…
📄 Beyond Detection: redesigning authentic assessment in an AI-mediated world
> Detection-led responses face well-documented limits: validity and fairness failures (bias against non-native writers), notable error rates, erosion of trust, and distraction from assessment design. …
📄 CoTAL: Human-in-the-Loop Prompt Engineering for Generalizable Formative Assessment Scoring and Feedback
> 1. **Evidence-Centered Design (ECD)** — assessments and rubrics aligned to curriculum goals from the start 2. **Human-in-the-loop prompt engineering** — labelled examples and prompts refined iterati…
2026-08-03 · formative-assessment, automated-grading, human-in-the-loop, prompt-engineering, benchmark
📄 Responsible Assessment in the AI Era: Key Insights from a Future-Focused Conference
> **Responsible assessment in the AI era** — assessment grounded in learners' sociocultural contexts and designed to generate valid, trustworthy, context-specific inferences from accumulated evidence,…
📄 The Theoretical Foundation of Socratic Tests: Dynamic, Multimodal, Conversational Examinations
> **Ilya Mikhelson** — Submitted to Computers and Education: Artificial Intelligence (2026).…
📄 AI-based scoring systematically underestimates conceptual understanding of linguistically weak students' explanations in physics
> **Authors:** Markus S. Feser, Paul L. Tschisgale (Leibniz Institute for Science and Mathematics Education, Kiel, Germany)…
📄 LLM Fallacy Misattribution in Education
> **The LLM Fallacy** is a cognitive attribution error in which individuals misinterpret LLM-assisted outputs as evidence of their own independent competence — producing a systematic gap (∆C) between …
2026-07-29 · llm, misinformation, ai-literacy, ai-assistance-reduces-persistence, cognitive-load-theory
📄 Representation Robustness under Executable Reasoning Constraints in Large Language Models for Mathematical Problem Solving
This study probes how sensitive [[llm]] mathematical problem solving is to the surface representation of an item — a question with direct bearing on [[assessment-validity]] when LLMs are used for scor…
📄 What Does the Credential Still Certify? Cognitive Stewardship for AI-Mediated Education
Generative AI undermines a basic premise of educational assessment: that submitted work reliably evidences the human capacities a credential certifies. This paper proposes *cognitive stewardship*, a f…
📄 Confidence-Aware Automated Assessment of Student-Drawn Scientific Models
> **Luyang Fang, Yingchuan Zhang, Jongchan Park, Zhaoji Wang, Ping Ma, Xiaoming Zhai** (2026). arXiv cs.AI preprint…
🏷️ AI Ed Evaluation
> **AI-ed evaluation** — the body of methods, benchmarks, and criteria used to assess whether AI education tools (LLM-based tutors, automated graders, feedback systems, agents) actually work — not jus…
📄 REC-CBM: Rubric-Aware Error-Correction Concept Bottleneck Models for Trustworthy Open-Ended Grading
**REC-CBM: Rubric-Aware Error-Correction Concept Bottleneck Models** advances the [[automated-grading]] frontier by solving a fundamental trust problem: even accurate AI graders are unusable if educat…
📄 LaTA: A Drop-in, FERPA-Compliant Local-LLM Autograder for Upper-Division STEM Coursework
> LaTA: A Drop-in, FERPA-Compliant Local-LLM Autograder for Upper-Division STEM Coursework **Rodríguez (2026)** — Oregon State University. Submitted to Computers & Education.…
📄 AI Literacy Assessment: Self-Reported vs Performance Misalignment
>Highlights critical misalignment between self-reported AI literacy and actual performance. Teachers overestimate their AI skills by 40% on average. Performance-based assessments correlate better (r=0…
📄 Quality-Conditioned Agreement in Automated Short Answer Scoring: Mid-Range Degradation and the Impact of Task-Specific Adaptation
> Schleifer, Ariely & Klebanov (2026) investigate a critical gap in [[automated-grading]]: **how scoring quality degrades for mid-range student responses**. Most ASAS evaluations focus on clearly corr…
📄 The University AI Didn''t Replace: Rethinking Universities in the AI Era
> **Synthesis:** Rather than replacing universities, generative AI **redefines their essential functions** — this paper proposes a four-level framework of institutional AI adoption and argues that the…
📄 A meta-analysis of the effect of generative AI on productivity and learning in programming
> Maier, Gunzenhäuser & Schweisthal (2026) conduct a **meta-analysis synthesizing evidence** on how generative AI tools affect both programming productivity and learning outcomes. This is a **confiden…
📄 Faculty Readiness for AI-Supported Teaching and Scalable Online Program Delivery in Higher Education: The EPIQ-AI Framework for Epistemic Integrity
> **Synthesis:** Sangwa, Ndahayo & Dusengumuremyi (2026) develop the EPIQ-AI Readiness Framework synthesizing data from 2020-2025 to explain how institutions can align faculty capacity, governance, an…