Concept
Educational Measurement
Educational measurement — the psychometric theory and methods for quantifying and validating learning and its constructs — runs through the knowledge base's Item Response Theory, Knowledge Tracing, and Assessment Validity pages. The Large Language Models (LLMs) era forces measurement to reconcile classical psychometrics with new AI-generated response streams: automated scoring, AI-predicted difficulty, and Multimodal AI traces must be validated against established measurement principles to preserve reliability and validity.
Questions to Consider
- An LLM predicts that an exam item is easy, but empirical test data says students find it hard. Which do you Trust, and what would convince you to trust the machine's estimate?
- Educational measurement is about turning observations of learning into defensible quantitative claims. When an AI grades an essay or scores a response, is a 'score' automatically a measurement — or does something have to be validated first? What?
- Some research suggests assessment instruments may not measure the same thing for humans as they do for LLMs — the latent structure diverges. If the constructs genuinely differ across humans and AI, what does that imply about AI-generated grades or difficulty ratings?
- AI can score, generate items, and predict difficulty at unprecedented scale. Is 'more measurement' the same as 'better measurement'? What makes a score reliable and valid, and can those standards be preserved when the measurement is done by a generative model?
- Benchmarks and AI-generated scores are everywhere now. What would you need to see before you'd treat an AI-based assessment as evidence about a learner's actual understanding rather than just a number?
Introduction
Educational measurement is the discipline of turning observations about learning — responses, behaviors, scores — into defensible quantitative claims. It encompasses construct definition, item/test design, scaling, reliability, and validity. In AI in education, measurement questions are everywhere: does a benchmark score measure what we think? Is an AI-generated grade reliable and valid? Do AI-predicted item difficulties agree with empirically estimated ones?
How educational measurement appears in the research
-
AI's decade-long reshaping of the field: Xiong and Li (2026) map AI's impact across three eras (Formative 2015–2018, Expansion 2019–2022, Generative 2023–present) via an Efficiency–Enhancement–Transformation framework, spanning AI on scoring and item generation, psychometric modeling, assessment innovation and process data, and fairness/ethics/equity. They argue for a new paradigm integrating measurement theory with AI methods and for reconceptualizing constructs in the context of human–AI interaction — the same boundary that latent-structure comparison probes empirically.
-
AI-predicted difficulty and calibration: LLM difficulty calibration and item-difficulty prediction use LLMs to estimate item difficulty, which must be validated against psychometric estimates (see Item Response Theory). Razavi and Powers (2026) add a large-scale K-5 test of this premise: across 5,170 math and reading items calibrated under the Rasch IRT model, GPT-4o's zero-shot difficulty ratings correlated moderately-to-strongly with true difficulties (r = 0.83 math, r = 0.81 reading) but were uneven across grades and no better than a grade-mean dummy regressor for grades K and 1. A feature-based approach — LLM-extracted cognitive and linguistic features fed into tree-based models — reached correlations up to r = 0.87, with grade level and word count the top predictors. The study illustrates both the promise and the limits of AI-predicted difficulty as a measurement input, and offers a practical seven-step workflow for testing professionals.
-
Psychometric awareness in AI Assessment: psychometrically aware AI is the standard that AI-based assessment be aligned with measurement theory — calibrated, uncertainty-aware, and validity-preserving (see Confidence Aware AI Assessment).
-
Automated scoring and validity: AI scoring and language bias and multimodal item-parameter estimation examine how automated scoring and multimodal data affect measurement quality.
-
Validity frameworks: Assessment Validity and Educational NLP supply the standards and tools for validating LLM-based measurement.
-
Latent-structure comparison: Strugatski et al. (2026) extend educational measurement to the LLM setting by testing whether assessment instruments show the same factor structure for humans and LLMs. Using EFA, factor congruence, and resampling, they show LLM–human latent structures systematically diverge across chemistry and quantitative-reasoning instruments, implying the constructs measured differ across populations — a necessary check before human validity evidence is assumed to transfer to AI.
-
AI as a rater with measurable error: Cvengros & Kortemeyer treat AI-assigned points as fallible observations with multiple error sources (items, runs, tasks) under generalizability theory and Kane's argument-based validity. Reliability analysis of a 296-student handwritten general-chemistry exam showed high stability of total scores across five AI runs (ICC(A,1) = 0.967, Kendall's W = 0.959, 95% repeatability coefficient 5.33 of 60 points) with lower item-level stability (ICC(A,1) = 0.836) — unsystematic errors partially cancel when summed (a Spearman–Brown aggregation effect that lifts total-score agreement to R² = 0.91), while a small positive intercept with slope < 1 revealed a "timid grader" score-compression bias. This frames run-averaging and aggregation as measurement design decisions, and motivates confidence filters calibrated to IRT risk as a validity safeguard before automated measures are trusted.
-
Calibration that scales with new data rather than the whole history: Jewsbury et al. (2026) confront the operational consequence of AI item generation — banks larger, sparser and more frequently updated than conventional ones, where refitting all accumulated response data at each update is costly and can exceed available memory. Their consensus calibration combines independently calibrated quarterly periods by mapping each period's posterior draws onto a common metric with a per-draw robust Haebara link (so linking uncertainty propagates rather than being treated as a fixed transformation), then aggregating as a product of Gaussian posteriors from which each period's estimated, hierarchical prior is subtracted and a consensus prior reinstated. On operational Duolingo English Test data across four periods the aggregate matched a pooled single-run benchmark on posterior means (r = .998 for difficulty, .991 for log-discrimination) and standard deviations (r = .970 and .920), with residual under-dispersion of 0.91–0.98 in the consensus/pooled SD ratio across exposure tertiles — concentrated in the sparsest item tertile, where prior over-counting bites hardest. The measurement lesson is twofold: an estimated prior cannot be handled by the standard subset-prior devices of parallel Bayesian computation, and posterior dispersion — the uncertainty adaptive selection and scoring consume — is as much a target of the correction as the posterior mean.
-
Identifying a system's dispersion from two published numbers: Restrepo Morales et al. (2026) show that an achievement distribution's standard deviation can be recovered from published summary statistics alone — the mean and the share of students at or above a fixed proficiency cut score — because under a distributional assumption σ = (c − μ) / z(1 − p). Applied to PISA 2025 El Salvador (mathematics mean 346, 12.28% at or above the level-2 cut of 420.07), the implied σ is 63.8 (reading 76.8, science 63.8), well below the international benchmark of 100, as expected for a distribution pressed against the floor of the scale; the three estimates were obtained from independent pairs of figures and lie within about 13 points of one another, a mild internal consistency check. The share at Level 2 is reported for every PISA system and underpins the SDG 4.1.1 indicator, so the identification needs no microdata, and the same paper demonstrates the denominator problem in effect sizes by reporting one mathematics gap three ways — 129 points, 1.29 international standard deviations, or 2.02 standard deviations of the Salvadoran distribution itself — on the argument that standardized effect sizes divided by their own sample's dispersion are not legitimately comparable across studies.
Measurement instruments in the knowledge base
A central function of educational measurement is the development, validation, and use of instruments — the concrete scales, tests, and coding schemes that operationalize constructs. The knowledge base's articles document a wide range of instruments for AI-in-education constructs, which can be categorized by what they measure and by their measurement approach.
AI / GenAI literacy instruments
AI literacy is the construct with the richest instrument coverage in the knowledge base. Two broad families exist: performance-based tests (objective, less susceptible to self-report bias) and self-report scales (subjective, capturing perceived competence). The trade-offs of that second family are the subject of Self-Report Measures: self-report reaches attitudes and perceptions cheaply, but it cannot carry a claim about competence or behavior, and the knowledge base documents a 40% overestimation gap when parallel self-report and performance measures are compared.
- Performance-based (objective) measures. The flagship is GLAT (Generative AI Literacy Assessment Test), a 20-item multiple-choice instrument built on a 25-concept blueprint across four dimensions (Know & Understand, Use & Apply, Evaluate & Create, Ethics) and validated with CTT + 2PL IRT on 355 students (RMSEA = 0.03, CFI = 0.97, α = 0.80, ω = 0.81). Critically, GLAT scores predicted AI-assisted task performance where self-report did not — evidence that performance-based measurement outperforms self-report for AI literacy. Related work in How to Assess AI Literacy: Misalignment Between Self-Reported and Objective-Based Measures quantifies the gap between self-reported and performance-based AI literacy (teachers overestimate by ~40%), and Tracing GenAI Literacy: Student-AI Interaction Patterns in Academic Writing traces actual student–AI interaction patterns rather than relying on reported use.
- Self-report scales. SAIL operationalizes AI literacy across three domains (AI Concepts; Application and Technical Skills; AI Digital Citizenship) and four scaffolded levels; the AI Literacy Heptagon structures seven dimensions (technical, application, critical thinking, ethics, social impact, integration, legal/regulatory) with four Bloom-aligned proficiency levels. The GenAI Skill Bypass: Mapping Divergent Pathways of University Students and Staff AI Literacy maps divergent AI-literacy pathways for students vs. staff, and Towards AI literacy: A proposal of a framework based on the Episodes of Situated Learning grounds literacy assessment in situated learning episodes. Instrument development for younger learners is also advancing: the AI Literacy Self-Assessment Questionnaire (AIL-SAQ) of Thianwan and Srikoon (2025) is a 15-item self-report scale for students in Grades 4 to 6, organized around Learning About AI, Learning About How AI Works, and Learning for Life with AI, validated with exploratory factor analysis on 335 students and confirmatory factor analysis on 579 more, with an overall Cronbach's alpha of .934.
- Domain-specific AI literacy. Teacher education for artificial intelligence literacy through a self-determination theory perspective develops teacher AI-literacy measures within a Self-Determination Theory framework (382 teachers, factor-validated); Conceptualizing pre-service teachers' readiness for AI integration into teaching practices: An intelligent-TPACK approach measures pre-service teacher AI readiness via intelligent-Technological Pedagogical Content Knowledge (TPACK); AI literacy alone is not enough: Student AI readiness and career adaptability in business and management education assesses student AI readiness and career adaptability in business education; and Can Large Language Models Foster Critical Thinking, Teamwork, and Problem-Solving Skills in Higher Education?: A Literature Review reviews instruments for LLM-supported critical-thinking and teamwork outcomes. Extending into educator measurement, the Teachers' AI Literacy Scale (TAILS) operationalizes the ED-AI framework's six dimensions to measure AI literacy specifically within language teacher education (validated with factor analysis), filling a gap in assessments that target students or general users.
Attitudes, acceptance, and motivation instruments
- Technology acceptance. Instruments grounded in TAM/UTAUT measure perceived usefulness, ease of use, and behavioral intention to use AI. See Technology Adoption Models and its application in Acceptance of AI-Assisted English Language Learning Tools in Higher Education: Psychological Correlates Across Disciplinary and Proficiency Groups (AI-assisted English learning tools, psychometric validation across disciplinary/proficiency groups) and GenAI adoption pathways.
- Self-efficacy and motivation. Self-Efficacy instruments and motivation scales (e.g., SDT-based measures of autonomy/competence/relatedness in Teacher education for artificial intelligence literacy through a self-determination theory perspective) capture the motivational antecedents and consequences of AI use. These connect to Student Engagement and Prior Knowledge measurement.
- Project-based learning perceptions. Zhu and Kong (2026) develop and validate a context-grounded AI project-based learning scale (AI-PBLS) measuring students' perceptions of PBL when using AI for problem solving in AI literacy courses (EFA and CFA on 1,027 secondary and university students, 446 complete), and their SEM application demonstrates how empowerment and ethical awareness mediate PBL-to-satisfaction relationships — a robust instrument for assessing perceived PBL experiences in AI applications.
Assessment-quality and validity instruments
- Automated scoring and rubric instruments. HARMOGEN-R generates assessment rubrics; AI-assisted, instructor-supervised grading and feedback in higher education: Design and evaluation of an end-to-end pipeline evaluates AI-grading quality against Elaborated-Feedback criteria; A bit of chaos and madness: The AI Assessment Scale and the work of assessment reform addresses how AI disrupts traditional assessment scales.
- Validity-strengthening designs. assessment twins pair a GenAI-vulnerable task with a less-vulnerable equivalent assessing the same outcomes, mapping threats across Messick's six strands of validity evidence.
- Discourse and engagement coding. Measuring Cognitive Engagement in Collaborative Discourse with an Extended ICAP Framework: Comparing Human Annotation, In-Context Learning, and Reflective LLM Agents extends the ICAP framework into a 7-point cognitive-engagement coding scheme, comparing human annotation (κ = 0.906–0.998) with LLM-based labeling (κ = 0.541–0.609) — a measurement-instrument study showing automated coding still trails trained humans. Automated coding is also advancing via context-aware prompting, which models contextual dependencies and fuses cognitive and social abilities to code collaborative problem-solving skills from process data at superior performance over strong baselines — enabling large-scale, real-time assessment while addressing the labor-intensity of manual coding.
- Skills extraction. Principal Trait Analysis: Towards Deriving 'Skills' in Human-AI Collaboration derives "skills" in human–AI collaboration via principal-trait analysis — a data-driven measurement of collaboration competency.
Measurement approach matters
The knowledge base's evidence repeatedly shows that how a construct is measured changes the conclusions. Self-reported AI literacy diverges sharply from performance-based measures (How to Assess AI Literacy: Misalignment Between Self-Reported and Objective-Based Measures); LLM annotation of engagement diverges from trained human coding (Measuring Cognitive Engagement in Collaborative Discourse with an Extended ICAP Framework: Comparing Human Annotation, In-Context Learning, and Reflective LLM Agents); and latent structures differ between humans and LLMs (Do Assessment Instruments Measure the Same Thing for Humans and LLMs? A Latent Structure Analysis). Rigorous instrument validation — reliability, structural validity, external/predictive validity — is therefore not a formality but the foundation of trustworthy AI-in-education evidence, connecting to Assessment Validity and Psychometrically Aware AI. A recent material-level study shows how much validation work sits behind a simple acceptance claim: Wang, Chuang and Wu (2026) used a split-half design (EFA on a development sample, CFA on a held-out half) for two age-tiered AI literacy guidebooks, reporting KMO = .820, 67.3% variance explained, acceptable student fit (CFI = .943, RMSEA = .059, SRMR = .059), composite reliability of .84 to .90, and measurement invariance across editions — then documenting where the instrument strains, with HTMT values up to .950 and a constrained playfulness-intention correlation that had to be tested for unity. Their supplementary checks for differential item functioning are a model of the honesty this page argues for: invariance held, but low-variation response patterns among younger students meant one group difference weakened once those respondents were excluded, and the paper reports that alongside the headline result rather than instead of it.
- Agreement coefficients are corpus properties, and scores are joint products. An automated-scoring validation of 60 marketing posts reported absolute agreement ICC(2,1) of .435 for an LLM, .266 for a rules-plus-LLM hybrid and .091 for deterministic rules, with uncertainty estimated from 2,000 writer-cluster bootstrap samples and MAE of 6.28-17.22 points; adding researcher-authored anchors moved inter-rater agreement from .338 to .902, showing how strongly these estimates depend on the reference set (Agreement and error in automated scoring of student marketing posts). An expert re-grading audit of physics benchmarks makes the complementary point at the instrument level: 95.20% of audited rejections were benchmark or grader errors rather than model failures, so a measured score must be reported as a property of the item bank, the rubric and the grader together (How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks).
Reliability is partly recoverable from the model itself. Organisciak and Acar (2026) evaluate three cheap upgrades to LLM-based scoring on more than 20,000 responses to the Alternative Uses Test: reading the model's own token-level confidence, taking a probability-weighted mean over its top-n predicted scores instead of a single best guess, and ensembling across models. Each improved agreement with human scores (correlation rising from r = 0.781 to 0.823 and RMSE falling from 0.599 to 0.498 in the best configuration), and their diagnostics locate a specific failure of single-pass scoring: the model's expressed confidence was negatively related to its accuracy (β = −0.602). This is a measurement-approach result rather than a model result — the same model, scored differently, is a more reliable instrument — and it sits alongside the corpus-level agreement evidence above.
A 2026 appraisal of 33 teacher AI literacy instruments shows where instrument development is mature and where it is not. Zainal, Mohd Matore and Maat (2026) graded the instruments against a decision matrix adapted from COSMIN and Terwee et al. (2007) and found internal consistency the strongest domain (28 of 33 at Grade A, 84.8%) and fairness the weakest, with only five instruments (15.2%) reporting measurement invariance or differential item functioning evidence. Structural validity was strong, with 24 instruments (72.7%) at Grade A through CFA, PLS-SEM or IRT modeling, yet content validity rested mostly on qualitative review, with 21 instruments (63.6%) at Grade B for lacking quantitative expert agreement statistics.
Issues and limitations: what measurement can miss or get wrong
Educational measurement is powerful but fallible. Understanding its failure modes is essential to reading AI-in-education evidence critically — and to recognizing where an apparent learning gain or construct claim may be an artifact of measurement rather than a real effect.
- Reliability limits. Measurement is never perfectly reliable; error variance is always present. When instruments have low internal consistency or test–retest stability, observed differences may be noise. In AI contexts, new failure modes compound this: LLM-generated responses can be scored with high machine agreement yet diverge from human scoring (Measuring Cognitive Engagement in Collaborative Discourse with an Extended ICAP Framework: Comparing Human Annotation, In-Context Learning, and Reflective LLM Agents), and automated scoring can be internally consistent while systematically wrong — precision without validity. The measurement limitations of AIEd research document unreliable instruments as a cross-cutting weakness.
- Rater agreement sets the practical ceiling. A large-scale validation on Uruguay's Acredita EB found that human rater agreement itself sets the practical ceiling for AI scoring: among ten experts independently scoring 50 texts, no rubric item reached unanimous agreement with consensus, the most divergent rater typically fell below 80% agreement while the best exceeded 90%, and Cohen's Kappa was only slight or fair for several skewed items that almost all responses satisfy. The authors therefore judge automated scores against a reference standard that is itself imperfect — ten raters, a single operational score for most responses, and several items in the conventional 70% acceptability zone — a caution for any measurement claim built on single-rater operational labels.
- Validity — measuring the wrong thing. Validity asks whether an instrument measures the construct it claims to. Common failures include construct under-representation (an AI-literacy test that samples only technical knowledge, missing ethics) and construct-irrelevant variance (an item that rewards reading fluency rather than the target skill). AI scoring and language bias shows how surface features — language, phrasing, style — can drive automated scores in ways unrelated to the intended construct. Assessment Validity is the guardrail against these threats.
- The self-report gap. Self-report measures capture perceived competence, not actual competence. The knowledge base repeatedly shows self-reported AI literacy diverging sharply from performance-based measures (How to Assess AI Literacy: Misalignment Between Self-Reported and Objective-Based Measures, ~40% overestimation by teachers) and that self-report fails to predict real AI-assisted performance where performance tests succeed (GLAT). Measures that rely on self-report can systematically overstate constructs and conceal true skill gaps.
- Constructs that don't transfer across populations. Strugatski et al. show assessment instruments can have a different factor structure for humans and LLMs — meaning the same items may not measure the same latent construct across populations. Even within humans, instruments validated on one group (e.g., Western, resourced Higher Education) may not generalize to others (Global South), a concern for the generalizability of AI-in-education measures.
- What measurement can miss. Some of the constructs most central to AI-in-education learning are the hardest to measure well — and therefore the most easily missed or distorted by instruments:
- Process and strategy. Standard outcome measures capture the product of learning, not the process. They can miss how students actually engage with AI — over-reliance, unreflective acceptance, or critical verification. Tracing GenAI Literacy: Student-AI Interaction Patterns in Academic Writing uses interaction traces precisely because self-report and outcome tests miss these dynamics.
- Longitudinal and durable learning. A single post-test may show inflated performance from AI assistance while missing the erosion of durable, unassisted knowledge (the performance–learning gap). Measurement at one time point can be actively misleading about learning.
- Equity and access. Instruments that assume uniform device/connectivity access, or that are normed on privileged samples, can miss — or systematically under-measure — the capabilities of under-resourced Learners, misattributing access gaps to ability gaps.
- Affective and motivational states. Engagement, motivation, and self-efficacy are often measured by self-report, inheriting the self-report gap above; they may not capture the situated, momentary dynamics that drive learning.
- Automated measurement can be confidently wrong. The combination of high machine confidence and opaque scoring is a distinctive AI-era risk: an LLM grader or annotator can produce highly self-consistent scores that are systematically biased, and the appearance of rigor (large N, high inter-LLM agreement) can mask invalidity. The AI Ed Evaluation and Psychometrically Aware AI frameworks are the antidote — requiring calibration, uncertainty awareness, and validity evidence before automated measures are trusted.
In short, educational measurement can miss what it does not sample (process, durability, access, affect) and can get wrong what it samples poorly (self-perception, surface features, cross-population constructs). Reading AI-in-education findings therefore requires asking not just what was measured but how — and what the instrument may have failed to capture.
- Auditable coding separates definitional error from model error. EduBehaviors decomposes a construct into separately annotated observable behaviors plus an explicit aggregation rule, so a reviewer can see which evidence produced a label and re-derive labels after a rule change without new model calls; it reached macro-F1 0.673 on Teacher TalkMoves, competitive with direct LLM prompting (Bernado et al., 2026).
Connections
Educational measurement is the foundation for Item Response Theory, Assessment Validity, Knowledge Tracing, and Learner Modeling and Adaptive Instruction. It connects to Learning Analytics (measurement of learning data), Educational NLP (measuring language), and Psychometrically Aware AI (AI aligned with measurement theory). Its validity and reliability concerns underpin AI Ed Evaluation and the measurement limitations of the field. For the constructs it measures, it intersects with AI Literacy, Technology Adoption Models, Self-Efficacy, Motivation, and Student Engagement.
- Causal modeling of support interventions (2026): a structural causal modeling protocol moves educational assessment beyond associative item-response-theory belief updating toward interventional and counterfactual reasoning (e.g., the effect of hints), with structural equations elicited from experts using purely logical information — illustrated on compulsory-school algorithmic-skills tasks (Causal Modelling of Support Interventions for Student Competency Assessment).
Connected Concepts
- Interpreting and Applying AIEd Research
- Item Response Theory
- Assessment Validity
- Psychometrically Aware AI
- Knowledge Tracing
- Learner Modeling and Adaptive Instruction
- Educational NLP
- Learning Analytics
- AI Ed Evaluation
- Automated Assessment
- Limitations in AIEd Research
- AI Literacy
- Technology Adoption Models
- Self-Efficacy
- Benchmark
- Self-Report Measures
- Motivation
- Student Engagement
Connected Articles
- A Human-in-the-Loop Framework for AI-Assisted Scoring in Large-Scale Writing Assessment — A Human-in-the-Loop Framework for AI-Assisted Scoring in Large-Scale Writing Assessment
- Semantic Variability of LLM-Generated Replies Across LLMs: Implications for Designing Conversation-Based Assessment
- Causal Modelling of Support Interventions for Student Competency Assessment — Causal Modeling of Support Interventions for Student Competency Assessment
- Do Assessment Instruments Measure the Same Thing for Humans and LLMs? A Latent Structure Analysis — Do assessment instruments measure the same thing for humans and LLMs? (Strugatski et al. 2026)
- GLAT: The Generative AI Literacy Assessment Test — GLAT: IRT-validated GenAI literacy test (Jin et al. 2025)
- Benchmarking the Pedagogical Knowledge of Large Language Models — LLM pedagogical-knowledge benchmark (CDPK + SEND)
- Validating AI-generated classroom observations: Reliability, accuracy, and limits of LLM-based pedagogical judgment — LLM classroom observation reliability and accuracy (Melo et al. 2026)
- Measuring Cognitive Engagement in Collaborative Discourse with an Extended ICAP Framework: Comparing Human Annotation, In-Context Learning, and Reflective LLM Agents — Measuring cognitive engagement with an extended ICAP framework
- Assessment twins: An approach for strengthening assessment validity in the age of generative AI — Assessment twins for strengthening assessment validity under GenAI
- The AI Literacy Heptagon: A Structured Approach to AI Literacy in Higher Education — The AI Literacy Heptagon
- How to Assess AI Literacy: Misalignment Between Self-Reported and Objective-Based Measures — Self-reported vs performance AI literacy misalignment
- Teacher education for artificial intelligence literacy through a self-determination theory perspective — Teacher AI literacy through self-determination theory
- Acceptance of AI-Assisted English Language Learning Tools in Higher Education: Psychological Correlates Across Disciplinary and Proficiency Groups — AI acceptance measures for English learning tools
- From Evaluated Models to Evaluation Aids: A Multi-Evidence Study of LLM-Based Difficulty Calibration for Programming Examinations — From evaluated models to evaluation aids
- Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction — Cognitive evaluation of LLM item-difficulty prediction
- Multimodal Item Parameter Estimation using Simulated Response Probabilities — Multimodal item-parameter estimation
- AI-based scoring systematically underestimates conceptual understanding of linguistically weak students' explanations in physics — AI scoring and language bias in physics
- Analyzing Undergraduate Problem-Solving in Physics Through Interaction With an AI Chatbot — Socratic physics chatbot
- From Student Risk Prediction to SC2R: Semantics-Constrained Counterfactual Recourse for Educational Decision Support — From Student Risk Prediction to SC2R: Counterfactual Recourse
- The End of Assessment? Disruption and Transformation in the Age of AI
- Automating Learner Assessment: Benchmarking Machine Learning and Deep Learning Models for EEG-Based Familiarity Prediction — Automating Learner Assessment: EEG-Based Familiarity Prediction
- The AI Writes the Code, the Student Writes the Model: A Theory and Measurement Programme for Learning by Construction with Generative AI — Model authorship: theory & measurement for learning-by-construction with GenAI
- Assessing students' DRIVE: A framework to evaluate learning through interactions with generative AI — DRIVE: assessing learning through GenAI interaction (DRI + Visible Expertise)
- A Decade of Reflection and Thematic Review on Artificial Intelligence's Impact on Educational Measurement — Decade thematic review of AI in educational measurement
- Design and Validation of a Questionnaire on Teachers' Uses of Generative Artificial Intelligence — Questionnaire on teachers' uses of generative AI (Pérez-Montesdeoca et al. 2026)
- Enhancing AI Literacy Course Satisfaction Through Empowerment in AI Problem-Solving and Ethical Awareness: Development and Validation of an AI Project-Based Learning Scale — AI-PBLS scale; empowerment and ethical awareness mediating PBL-to-satisfaction in AI literacy courses (Zhu & Kong 2026)
- Context-aware prompting for collaborative problem solving skill identification — Context-aware prompting for automated collaborative problem-solving skill coding
- An Exploratory Machine Learning Approach to Understanding Determinants of Future ChatGPT Use in Higher Education — ML/SHAP determinants of future ChatGPT use in higher education
- Personalized neural cognitive architecture search — AutoML personalized neural cognitive architecture search for learner profiles
- Language teachers’ AI literacy: A psychometric study based on the ED-AI framework — Teachers' AI Literacy Scale (TAILS) psychometric study (ED-AI framework)
- Estimating Item Difficulty Using Large Language Models and Tree-Based Machine Learning Algorithms — Estimating item difficulty using LLMs and tree-based ML
- Assisting the grading of a handwritten general chemistry exam with artificial intelligence
- Measuring Acceptance of Age-Tiered AI Literacy Guidebooks: A Developmentally Informed Study of K-12 Students and Teachers — Split-half EFA/CFA validation with invariance, HTMT and DIF checks, including an honest account of playfulness-intention construct overlap
- ProIQA: A Process-Based Framework for Fine-Grained Math Item Quality Assessment — ProIQA: Process-Based Math Item Quality Assessment
- Beyond AI Literacy: A Structured Review and Exploratory Meta-Analysis of Measures for Competent Generative-AI Use — Beyond AI Literacy: A Structured Review and Exploratory Meta-Analysis of Measures for Competent Generative-AI Use
- Towards Scalable Measurement of Durable Skills — Toward Scalable Measurement of Durable Skills
- Bayesian Consensus Calibration of Continuously Evolving IRT Item Banks — Bayesian consensus calibration: divide-and-conquer recalibration of a continuously evolving IRT item bank (Jewsbury et al. 2026)
- How much selection would be enough? Bounding the learning claim of El Salvador's artificial intelligence tutoring pilot — Bounding the learning claim of El Salvador's AI tutoring pilot (Restrepo Morales et al. 2026)
- Know When to Trust: Making AI Scoring More Reliable for Educational Assessment — Know When to Trust: model self-confidence, probabilistic scoring and ensembles as scoring reliability levers
- The AIR Scale: Development and Validation of a Measure of Motivations for Using AI During Reading — The AIR Scale: development and validation of an instrument for motivations to use AI while reading
- Can LLMs Model Incorrect Student Reasoning? A Case Study on Distractor Generation for Multiple-Choice Questions — distractor quality criteria and the reasoning strategies behind them
- Development of an AI literacy self-assessment questionnaire in upper primary school students — A validated 15-item self-assessment questionnaire for upper-primary AI literacy (Thianwan & Srikoon 2025)
- Assessing teachers' AI literacy: a systematic review of measurement tools — Systematic appraisal of 33 teacher AI literacy instruments across COSMIN-style quality domains (Zainal et al. 2026)
- Can large language models reproduce higher education grade bands? Cross-model study of calibration and grading bias in authentic student writing — Can large language models reproduce higher education grade bands? Cross-model study of calibration and grading bias in authentic student writing
- Mental Health Literacy Across Psychology Students and Large Language Models — Mental Health Literacy Across Psychology Students and Large Language Models
- StudentBench: AI and human tutoring yield equivalent GRE learning gains — StudentBench: AI and human tutoring yield equivalent GRE learning gains
- Student Use of LLMs and the Limits of AI-Generated Question Difficulty in Data Science Courses — Student Use of LLMs and the Limits of AI-Generated Question Difficulty in Data Science Courses
- EduBehaviors: Assertion-based Schemas for Auditable Coding of Educational Dialogues — EduBehaviors: Assertion-based Schemas for Auditable Coding of Educational Dialogues
- From Sentiment Classification to Actionable and Responsible Feedback: A Scoping Review and Evidence Map of NLP in Student Evaluation of Teaching, 2015–2026 — From Sentiment Classification to Actionable and Responsible Feedback: A Scoping Review and Evidence Map of NLP in Student Evaluation of Teaching, 2015–2026
- What Fidelity Metrics Miss: A Structural Check on Synthetic Educational Data — What Fidelity Metrics Miss: A Structural Check on Synthetic Educational Data
- Evaluation of pre-trained models for pedagogical assessment of novel AI-assisted educational questions — Evaluation of pre-trained models for pedagogical assessment of novel AI-assisted educational questions