Stub โ pending source ingestion. AI-powered automated grading and assessment systems for educational contexts.
Connections
Related Pages
- cotal-formative-assessment-scoring-2026 โ CoTAL human-in-the-loop prompt engineering
- structrag-diagram-reasoning-ai-tutoring โ Diagram-based assessment and structural verification
- llm-cognitive-diagnosis-handwritten-math โ MathCog benchmark: 18 LLMs evaluated on cognitive skill diagnosis from handwritten math; all F1 < 0.5; systematic over-attribution and hallucination of evidence (2025)
- correct-answer-trap-ai-tutor โ 8 of 8 papers in May 28 scan
- rubric-aware-grading-rec-cbm โ 2 of 8 papers in May 28 scan
- student-misconceptions-conditionals-loops-taxonomy โ error taxonomy enabling explanatory feedback beyond correct/incorrect
- llm-automated-assessment-student-self-explanations โ LLM vs. semantic similarity for scoring student self-explanations in programming (2026)
- lata-ferpa-compliant-local-llm-autograder โ FERPA-compliant local LLM autograder with 0.02% error rate
- academiclaw-student-agent-benchmark โ AcademiClaw: 6-technique multi-dimensional rubric approach connects to automated assessment methodology
- self-referential-l2-writing-llm-assessment โ Paradigm shift from inter-learner ranking to intra-learner profiling
- short-answer-scoring-quality-degradation โ Reveals significant quality degradation for mid-range student responses
- multimodal-ai-feedback-learning โ Zhao et al.: multimodal feedback improves student experience without changing grading accuracy
- ground-truth-reliability-aied โ Thomas et al.: ground truth quality framework for the labels that train and evaluate grading systems
- llm-feedback-programming-classroom โ LLM-generated feedback as automated grading in intro programming- aiawe-automated-writing-evaluation โ Open-source automated writing evaluation with LoRA-adapted LLMs
- learnopt-exam-cognitive-structure -- Standardized exams have stable latent cognitive structures recoverable via LLM-tagged question analysis and knapsack optimization
- cross-dataset-bloom-question-classification -- LLMs with tailored prompting generalize better than supervised models for cross-dataset Bloom taxonomy classification
- llm-chatbots-cs-multiple-choice -- ChatGPT answers with explanations do not improve student MCQ performance; GPT-4o/5 outperform smaller models
- confidence-aware-student-drawing-assessment -- Confidence-Aware Automated Assessment of Student-Drawn Scientific Models
- psyscore-essay-scoring-zpd-feedback -- PsyScore: A Psychometrically-Aware Framework for Trait-Adaptive Essay Scoring and ZPD-Scaffolded Feedback
- test-driven-ai-assisted-learning -- A lecture-free CS course with AI-assisted weekly closed-book tests maintained accountability and was scalable with a version-controlled AI agent workspace.
- machines-misread-pedagogical-quality -- Human-machine disagreements in AI pretest evaluation are systematic; rubric revision has a larger alignment effect than rationale-first evaluation, and the two are complementary.
- automated-grading-linux-bash-examinations-large-language-models โ Automated Grading of Linux/Bash Examinations Using Large Language Models
- teaching-feedback-classification-benchmark โ Automated classification (2026-07-14)
- knowledge-distillation-ai-tutor-evaluation โ Scalable tutor grading (2026-07-14)
- automated-formative-assessments-a-level-sciences โ Automating the marking of handwritten mock exams enables much higher formative-assessment frequency
- aicode-collaborative-feedback-system โ Multi-LLM approach to automated feedback quality
- ai-assistance-discretionary-feedback โ AI-generated feedback scaffolds in grading workflows
- responsible-assessment-ai-era-stanford-2026 โ Report on AI scoring as the most common AI application in assessment, with validity costs
- gpt4o-mini-music-analysis-scoring โ GPT-4o-mini vs teacher mean scores for automated scoring of music analysis responses
๐ 41 other pages tagged automated-grading
- A Durability and Cross-Language Transfer Benchmark for a Validated Teaching-Feedback Classification Protocol
- Advancing diagram-based reasoning in AI tutoring systems: a structural approach for STEM education
- AI Assistance for Discretionary Work: Increasing Feedback Provision in Higher Education
- AI-based scoring systematically underestimates conceptual understanding of linguistically weak students' explanations in physics
- AI-Driven Assessment of Human Tutors: Linking Training Performance to Real-Life Practice
- AiAWE: An Open-Source LLM Automated Writing Evaluation System Using LoRA-Adapted Instruction-Tuned Models
- AICoFe: Implementation and Deployment of an AI-Based Collaborative Feedback System for Higher Education
- AISSA: AI-based Student Slides Analysis Tool for Academic Presentations
- Are LLM-based Chatbots Good Enough to Support Computer Science Students in Multiple-Choice Exercises?
- Automated Grading of Handwritten Mathematics Using Vision-Capable LLMs
- Automated Grading of Linux/Bash Examinations Using Large Language Models
- Automated Question Generation
- Automatic Short Answer Grading with LLMs
- Catching The Correct Answer Trap: Characterising AI Tutor Blind Spots When Analysing Student Reasoning
- Comparative Validation of GPT-4o-mini and Teacher Mean Scores for Automated Scoring of Music Analysis Responses: Single-Pass Deployment, Repeatability, and Strategy-Specific Bias
- Comparing human and LLM ordered coding of qualitative data: How coding differences cascade through temporal analysis
- Confidence-Aware Automated Assessment of Student-Drawn Scientific Models
- Confidence-Aware Automatic Short Answer Grading
- CoTAL: Human-in-the-Loop Prompt Engineering for Generalizable Formative Assessment Scoring and Feedback
- Cross-Dataset Bloom Question Classification: Supervised Models and Prompted LLMs
- Ensuring Reliability in Programming Knowledge Tracing: A Re-evaluation of Attention-augmented Models and Experimental Protocols
- Exploring the Effectiveness of Using LLMs for Automated Assessment of Student Self Explanations in Programming Education
- From Memorization to Creation: Evaluating the Cognitive Depth of LLM-Generated Educational Questions
- Generate-Then-Validate: Question Generation for Education
- Generative AI as a Design Variable: An Evidence-Centered Framework for Principled Governance in STEM Assessment
- Hybrid E-Assessment in Higher Education: Semi-Automated Grading of Paper-Based Written Examinations
- ICLE++: Modeling Fine-Grained Traits for Holistic Essay Scoring
- Knowledge Distillation for Automated AI Tutor Evaluation
- KT4EQG: Personalized Exercise Question Generation via Knowledge Tracing
- LaTA: A Drop-in, FERPA-Compliant Local-LLM Autograder for Upper-Division STEM Coursework
- LLM-Generated Feedback in Introductory Programming: A Classroom Study
- Modernizing Ground Truth: Four Shifts Toward Improving Reliability and Validity in AI in Education
- PsyScore: A Psychometrically-Aware Framework for Trait-Adaptive Essay Scoring and ZPD-Scaffolded Feedback
- Quality-Conditioned Agreement in Automated Short Answer Scoring: Mid-Range Degradation and the Impact of Task-Specific Adaptation
- REC-CBM: Rubric-Aware Error-Correction Concept Bottleneck Models for Trustworthy Open-Ended Grading
- Rethinking LLM-Judged Helpfulness as a Pedagogy Signal: A Pre-Registered Audit Across Tutor Models
- The Correct Answer Trap: Pedagogically-Grounded Detection and Feedback for Hidden Misconceptions
- The Effect of High-Frequency, Automatically-marked Formative Assessments on Student Outcomes in A-Level Sciences
- The Theoretical Foundation of Socratic Tests: Dynamic, Multimodal, Conversational Examinations
- The Tutoring Effectiveness Index: Predicting LLM Math Tutor Quality from Four Conversation Signals
- Towards Self-Referential Analytic Assessment: A Profile-Based Approach to L2 Writing Evaluation with LLMs