On this page

Educational measurement — the psychometric theory and methods for quantifying and validating learning and its constructs — runs through the knowledge base's Item Response Theory, Knowledge Tracing, and Assessment Validity pages. The Large Language Models (LLMs) era forces measurement to reconcile classical psychometrics with new AI-generated response streams: automated scoring, AI-predicted difficulty, and Multimodal AI traces must be validated against established measurement principles to preserve reliability and validity.

Questions to Consider

  • An LLM predicts that an exam item is easy, but empirical test data says students find it hard. Which do you Trust, and what would convince you to trust the machine's estimate?
  • Educational measurement is about turning observations of learning into defensible quantitative claims. When an AI grades an essay or scores a response, is a 'score' automatically a measurement — or does something have to be validated first? What?
  • Some research suggests assessment instruments may not measure the same thing for humans as they do for LLMs — the latent structure diverges. If the constructs genuinely differ across humans and AI, what does that imply about AI-generated grades or difficulty ratings?
  • AI can score, generate items, and predict difficulty at unprecedented scale. Is 'more measurement' the same as 'better measurement'? What makes a score reliable and valid, and can those standards be preserved when the measurement is done by a generative model?
  • Benchmarks and AI-generated scores are everywhere now. What would you need to see before you'd treat an AI-based assessment as evidence about a learner's actual understanding rather than just a number?

Introduction

Educational measurement is the discipline of turning observations about learning — responses, behaviors, scores — into defensible quantitative claims. It encompasses construct definition, item/test design, scaling, reliability, and validity. In AI in education, measurement questions are everywhere: does a benchmark score measure what we think? Is an AI-generated grade reliable and valid? Do AI-predicted item difficulties agree with empirically estimated ones?

How educational measurement appears in the research

  • AI's decade-long reshaping of the field: Xiong and Li (2026) map AI's impact across three eras (Formative 2015–2018, Expansion 2019–2022, Generative 2023–present) via an Efficiency–Enhancement–Transformation framework, spanning AI on scoring and item generation, psychometric modeling, assessment innovation and process data, and fairness/ethics/equity. They argue for a new paradigm integrating measurement theory with AI methods and for reconceptualizing constructs in the context of human–AI interaction — the same boundary that latent-structure comparison probes empirically.

  • AI-predicted difficulty and calibration: LLM difficulty calibration and item-difficulty prediction use LLMs to estimate item difficulty, which must be validated against psychometric estimates (see Item Response Theory). Razavi and Powers (2026) add a large-scale K-5 test of this premise: across 5,170 math and reading items calibrated under the Rasch IRT model, GPT-4o's zero-shot difficulty ratings correlated moderately-to-strongly with true difficulties (r = 0.83 math, r = 0.81 reading) but were uneven across grades and no better than a grade-mean dummy regressor for grades K and 1. A feature-based approach — LLM-extracted cognitive and linguistic features fed into tree-based models — reached correlations up to r = 0.87, with grade level and word count the top predictors. The study illustrates both the promise and the limits of AI-predicted difficulty as a measurement input, and offers a practical seven-step workflow for testing professionals.

  • Psychometric awareness in AI Assessment: psychometrically aware AI is the standard that AI-based assessment be aligned with measurement theory — calibrated, uncertainty-aware, and validity-preserving (see Confidence Aware AI Assessment).

  • Automated scoring and validity: AI scoring and language bias and multimodal item-parameter estimation examine how automated scoring and multimodal data affect measurement quality.

  • Validity frameworks: Assessment Validity and Educational NLP supply the standards and tools for validating LLM-based measurement.

  • Latent-structure comparison: Strugatski et al. (2026) extend educational measurement to the LLM setting by testing whether assessment instruments show the same factor structure for humans and LLMs. Using EFA, factor congruence, and resampling, they show LLM–human latent structures systematically diverge across chemistry and quantitative-reasoning instruments, implying the constructs measured differ across populations — a necessary check before human validity evidence is assumed to transfer to AI.

  • AI as a rater with measurable error: Cvengros & Kortemeyer treat AI-assigned points as fallible observations with multiple error sources (items, runs, tasks) under generalizability theory and Kane's argument-based validity. Reliability analysis of a 296-student handwritten general-chemistry exam showed high stability of total scores across five AI runs (ICC(A,1) = 0.967, Kendall's W = 0.959, 95% repeatability coefficient 5.33 of 60 points) with lower item-level stability (ICC(A,1) = 0.836) — unsystematic errors partially cancel when summed (a Spearman–Brown aggregation effect that lifts total-score agreement to R² = 0.91), while a small positive intercept with slope < 1 revealed a "timid grader" score-compression bias. This frames run-averaging and aggregation as measurement design decisions, and motivates confidence filters calibrated to IRT risk as a validity safeguard before automated measures are trusted.

  • Calibration that scales with new data rather than the whole history: Jewsbury et al. (2026) confront the operational consequence of AI item generation — banks larger, sparser and more frequently updated than conventional ones, where refitting all accumulated response data at each update is costly and can exceed available memory. Their consensus calibration combines independently calibrated quarterly periods by mapping each period's posterior draws onto a common metric with a per-draw robust Haebara link (so linking uncertainty propagates rather than being treated as a fixed transformation), then aggregating as a product of Gaussian posteriors from which each period's estimated, hierarchical prior is subtracted and a consensus prior reinstated. On operational Duolingo English Test data across four periods the aggregate matched a pooled single-run benchmark on posterior means (r = .998 for difficulty, .991 for log-discrimination) and standard deviations (r = .970 and .920), with residual under-dispersion of 0.91–0.98 in the consensus/pooled SD ratio across exposure tertiles — concentrated in the sparsest item tertile, where prior over-counting bites hardest. The measurement lesson is twofold: an estimated prior cannot be handled by the standard subset-prior devices of parallel Bayesian computation, and posterior dispersion — the uncertainty adaptive selection and scoring consume — is as much a target of the correction as the posterior mean.

  • Identifying a system's dispersion from two published numbers: Restrepo Morales et al. (2026) show that an achievement distribution's standard deviation can be recovered from published summary statistics alone — the mean and the share of students at or above a fixed proficiency cut score — because under a distributional assumption σ = (c − μ) / z(1 − p). Applied to PISA 2025 El Salvador (mathematics mean 346, 12.28% at or above the level-2 cut of 420.07), the implied σ is 63.8 (reading 76.8, science 63.8), well below the international benchmark of 100, as expected for a distribution pressed against the floor of the scale; the three estimates were obtained from independent pairs of figures and lie within about 13 points of one another, a mild internal consistency check. The share at Level 2 is reported for every PISA system and underpins the SDG 4.1.1 indicator, so the identification needs no microdata, and the same paper demonstrates the denominator problem in effect sizes by reporting one mathematics gap three ways — 129 points, 1.29 international standard deviations, or 2.02 standard deviations of the Salvadoran distribution itself — on the argument that standardized effect sizes divided by their own sample's dispersion are not legitimately comparable across studies.

Measurement instruments in the knowledge base

A central function of educational measurement is the development, validation, and use of instruments — the concrete scales, tests, and coding schemes that operationalize constructs. The knowledge base's articles document a wide range of instruments for AI-in-education constructs, which can be categorized by what they measure and by their measurement approach.

AI / GenAI literacy instruments

AI literacy is the construct with the richest instrument coverage in the knowledge base. Two broad families exist: performance-based tests (objective, less susceptible to self-report bias) and self-report scales (subjective, capturing perceived competence). The trade-offs of that second family are the subject of Self-Report Measures: self-report reaches attitudes and perceptions cheaply, but it cannot carry a claim about competence or behavior, and the knowledge base documents a 40% overestimation gap when parallel self-report and performance measures are compared.

Attitudes, acceptance, and motivation instruments

Assessment-quality and validity instruments

Measurement approach matters

The knowledge base's evidence repeatedly shows that how a construct is measured changes the conclusions. Self-reported AI literacy diverges sharply from performance-based measures (How to Assess AI Literacy: Misalignment Between Self-Reported and Objective-Based Measures); LLM annotation of engagement diverges from trained human coding (Measuring Cognitive Engagement in Collaborative Discourse with an Extended ICAP Framework: Comparing Human Annotation, In-Context Learning, and Reflective LLM Agents); and latent structures differ between humans and LLMs (Do Assessment Instruments Measure the Same Thing for Humans and LLMs? A Latent Structure Analysis). Rigorous instrument validation — reliability, structural validity, external/predictive validity — is therefore not a formality but the foundation of trustworthy AI-in-education evidence, connecting to Assessment Validity and Psychometrically Aware AI. A recent material-level study shows how much validation work sits behind a simple acceptance claim: Wang, Chuang and Wu (2026) used a split-half design (EFA on a development sample, CFA on a held-out half) for two age-tiered AI literacy guidebooks, reporting KMO = .820, 67.3% variance explained, acceptable student fit (CFI = .943, RMSEA = .059, SRMR = .059), composite reliability of .84 to .90, and measurement invariance across editions — then documenting where the instrument strains, with HTMT values up to .950 and a constrained playfulness-intention correlation that had to be tested for unity. Their supplementary checks for differential item functioning are a model of the honesty this page argues for: invariance held, but low-variation response patterns among younger students meant one group difference weakened once those respondents were excluded, and the paper reports that alongside the headline result rather than instead of it.

Reliability is partly recoverable from the model itself. Organisciak and Acar (2026) evaluate three cheap upgrades to LLM-based scoring on more than 20,000 responses to the Alternative Uses Test: reading the model's own token-level confidence, taking a probability-weighted mean over its top-n predicted scores instead of a single best guess, and ensembling across models. Each improved agreement with human scores (correlation rising from r = 0.781 to 0.823 and RMSE falling from 0.599 to 0.498 in the best configuration), and their diagnostics locate a specific failure of single-pass scoring: the model's expressed confidence was negatively related to its accuracy (β = −0.602). This is a measurement-approach result rather than a model result — the same model, scored differently, is a more reliable instrument — and it sits alongside the corpus-level agreement evidence above.

A 2026 appraisal of 33 teacher AI literacy instruments shows where instrument development is mature and where it is not. Zainal, Mohd Matore and Maat (2026) graded the instruments against a decision matrix adapted from COSMIN and Terwee et al. (2007) and found internal consistency the strongest domain (28 of 33 at Grade A, 84.8%) and fairness the weakest, with only five instruments (15.2%) reporting measurement invariance or differential item functioning evidence. Structural validity was strong, with 24 instruments (72.7%) at Grade A through CFA, PLS-SEM or IRT modeling, yet content validity rested mostly on qualitative review, with 21 instruments (63.6%) at Grade B for lacking quantitative expert agreement statistics.

Issues and limitations: what measurement can miss or get wrong

Educational measurement is powerful but fallible. Understanding its failure modes is essential to reading AI-in-education evidence critically — and to recognizing where an apparent learning gain or construct claim may be an artifact of measurement rather than a real effect.

  • Reliability limits. Measurement is never perfectly reliable; error variance is always present. When instruments have low internal consistency or test–retest stability, observed differences may be noise. In AI contexts, new failure modes compound this: LLM-generated responses can be scored with high machine agreement yet diverge from human scoring (Measuring Cognitive Engagement in Collaborative Discourse with an Extended ICAP Framework: Comparing Human Annotation, In-Context Learning, and Reflective LLM Agents), and automated scoring can be internally consistent while systematically wrong — precision without validity. The measurement limitations of AIEd research document unreliable instruments as a cross-cutting weakness.
  • Rater agreement sets the practical ceiling. A large-scale validation on Uruguay's Acredita EB found that human rater agreement itself sets the practical ceiling for AI scoring: among ten experts independently scoring 50 texts, no rubric item reached unanimous agreement with consensus, the most divergent rater typically fell below 80% agreement while the best exceeded 90%, and Cohen's Kappa was only slight or fair for several skewed items that almost all responses satisfy. The authors therefore judge automated scores against a reference standard that is itself imperfect — ten raters, a single operational score for most responses, and several items in the conventional 70% acceptability zone — a caution for any measurement claim built on single-rater operational labels.
  • Validity — measuring the wrong thing. Validity asks whether an instrument measures the construct it claims to. Common failures include construct under-representation (an AI-literacy test that samples only technical knowledge, missing ethics) and construct-irrelevant variance (an item that rewards reading fluency rather than the target skill). AI scoring and language bias shows how surface features — language, phrasing, style — can drive automated scores in ways unrelated to the intended construct. Assessment Validity is the guardrail against these threats.
  • The self-report gap. Self-report measures capture perceived competence, not actual competence. The knowledge base repeatedly shows self-reported AI literacy diverging sharply from performance-based measures (How to Assess AI Literacy: Misalignment Between Self-Reported and Objective-Based Measures, ~40% overestimation by teachers) and that self-report fails to predict real AI-assisted performance where performance tests succeed (GLAT). Measures that rely on self-report can systematically overstate constructs and conceal true skill gaps.
  • Constructs that don't transfer across populations. Strugatski et al. show assessment instruments can have a different factor structure for humans and LLMs — meaning the same items may not measure the same latent construct across populations. Even within humans, instruments validated on one group (e.g., Western, resourced Higher Education) may not generalize to others (Global South), a concern for the generalizability of AI-in-education measures.
  • What measurement can miss. Some of the constructs most central to AI-in-education learning are the hardest to measure well — and therefore the most easily missed or distorted by instruments:
    • Process and strategy. Standard outcome measures capture the product of learning, not the process. They can miss how students actually engage with AI — over-reliance, unreflective acceptance, or critical verification. Tracing GenAI Literacy: Student-AI Interaction Patterns in Academic Writing uses interaction traces precisely because self-report and outcome tests miss these dynamics.
    • Longitudinal and durable learning. A single post-test may show inflated performance from AI assistance while missing the erosion of durable, unassisted knowledge (the performance–learning gap). Measurement at one time point can be actively misleading about learning.
    • Equity and access. Instruments that assume uniform device/connectivity access, or that are normed on privileged samples, can miss — or systematically under-measure — the capabilities of under-resourced Learners, misattributing access gaps to ability gaps.
    • Affective and motivational states. Engagement, motivation, and self-efficacy are often measured by self-report, inheriting the self-report gap above; they may not capture the situated, momentary dynamics that drive learning.
  • Automated measurement can be confidently wrong. The combination of high machine confidence and opaque scoring is a distinctive AI-era risk: an LLM grader or annotator can produce highly self-consistent scores that are systematically biased, and the appearance of rigor (large N, high inter-LLM agreement) can mask invalidity. The AI Ed Evaluation and Psychometrically Aware AI frameworks are the antidote — requiring calibration, uncertainty awareness, and validity evidence before automated measures are trusted.

In short, educational measurement can miss what it does not sample (process, durability, access, affect) and can get wrong what it samples poorly (self-perception, surface features, cross-population constructs). Reading AI-in-education findings therefore requires asking not just what was measured but how — and what the instrument may have failed to capture.

  • Auditable coding separates definitional error from model error. EduBehaviors decomposes a construct into separately annotated observable behaviors plus an explicit aggregation rule, so a reviewer can see which evidence produced a label and re-derive labels after a rule change without new model calls; it reached macro-F1 0.673 on Teacher TalkMoves, competitive with direct LLM prompting (Bernado et al., 2026).

Connections

Educational measurement is the foundation for Item Response Theory, Assessment Validity, Knowledge Tracing, and Learner Modeling and Adaptive Instruction. It connects to Learning Analytics (measurement of learning data), Educational NLP (measuring language), and Psychometrically Aware AI (AI aligned with measurement theory). Its validity and reliability concerns underpin AI Ed Evaluation and the measurement limitations of the field. For the constructs it measures, it intersects with AI Literacy, Technology Adoption Models, Self-Efficacy, Motivation, and Student Engagement.

  • Causal modeling of support interventions (2026): a structural causal modeling protocol moves educational assessment beyond associative item-response-theory belief updating toward interventional and counterfactual reasoning (e.g., the effect of hints), with structural equations elicited from experts using purely logical information — illustrated on compulsory-school algorithmic-skills tasks (Causal Modelling of Support Interventions for Student Competency Assessment).

Connected Concepts

Connected Articles

Connected FAQs

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.