On this page

Item response theory (IRT) — a family of psychometric models that estimate latent ability from item responses by modeling the relationship between a learner's ability and the probability of answering each item correctly. IRT models item difficulty and discrimination, enabling measurement precision and adaptive testing. In the AI era, IRT meets LLMs in Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction and Aligning LLM-Simulated and Human Examinees for Psychometric Calibration: A Cognitive Diagnostic Profiling Approach: AI predicts and calibrates item difficulty, potentially improving measurement precision and feeding Adaptive Learning.

Questions to Consider

  • Item response theory treats ability and item difficulty as jointly estimated from response patterns, rather than treating a raw test score as the measure. How might two students with the same number correct actually differ in ability?
  • IRT lets you compare learners on a common scale and estimate precision per person. Why might knowing an item's difficulty and discrimination matter more than just knowing whether a student got it right?
  • One study used IRT person-fit statistics to distinguish human from AI-generated responses on multiple-choice tests — flagging AI responses as 'aberrant.' How could the same measurement machinery that assesses learning also police academic integrity?
  • Researchers use IRT to validate that AI-generated exam questions match expert-written ones in difficulty and discrimination. If an AI writes an item that 'looks' good, why is empirical calibration against fitted IRT parameters still necessary?
  • As AI predicts and calibrates item difficulty, what could go wrong if a model's estimate of difficulty isn't validated against real student response data?
  • IRT connects to adaptive testing and knowledge tracing — using your responses to choose what to ask next. How does estimating your ability from each answer enable a test to become shorter and more precise rather than just longer?

Introduction

IRT treats ability (θ) and item parameters (difficulty, discrimination, sometimes guessing) as jointly estimated from response patterns, rather than treating a raw score as the measure. This makes it possible to compare learners on a common scale, to select items adaptively, and to estimate precision per person rather than globally.

How IRT appears in the research

  • AI-predicted difficulty: LLM item-difficulty prediction uses language models to estimate item difficulty, which must be validated against empirically fitted IRT parameters.

  • Psychometric calibration: LLM psychometric calibration aligns model-based assessment with IRT-based measurement so that AI-generated responses preserve measurement properties.

  • Knowledge tracing and student modeling: IRT is closely related to Knowledge Tracing and Learner Modeling and Adaptive Instruction — models that track learner knowledge over time — sharing the goal of estimating unobservable learner states from observable responses.

  • Bayesian hierarchical field validation: Assessing AI-Generated Exams uses a Bayesian hierarchical 2PL IRT model (with pre-test anchor items to place 1,686 students on a common θ scale) to show that AI-generated questions match expert-written standardized-exam items in difficulty and discrimination — a large-scale demonstration of IRT as the validation backbone for Automated Question Generation.

  • Separating human from GenAI responses with person-fit statistics: Strugatski and Alexandron (2026) apply person-fit statistics (PFS) within IRT to distinguish human from Generative AI responses on multiple-choice assessments. PFS flag GenAI responses as 'aberrant' responders in two authentic contexts (a chemistry test and a national exam), show that different chatbots produce distinct response patterns (a heterogeneous group of 'intelligences'), and reveal that newer GenAI versions become more human-like — positioning IRT as a robust framework for integrity screening in high-stakes testing.

  • LLM difficulty estimation against Rasch IRT parameters: Razavi and Powers (2026) evaluate whether GPT-4o can estimate the difficulty of K-5 math and reading assessment items (N = 5170) calibrated under the Rasch IRT model. A zero-shot direct estimation approach correlated moderately-to-strongly with true Rasch difficulties (r = 0.83 math, r = 0.81 reading) but was uneven across grades and often no better than a grade-mean dummy regressor for grades K and 1, likely due to range restriction in lower-grade item difficulties. A feature-based strategy — LLM-extracted cognitive and linguistic features fed into tree-based models — outperformed direct estimation (correlations up to r = 0.87), with grade level and word count the top predictors. The study underscores that LLM difficulty estimates must be validated against empirically fitted IRT parameters, and that structured feature extraction can sharpen prediction where holistic zero-shot judgment falls short.

  • Item-writing flaws as a pre-deployment screen for IRT parameters: Schmucker and Moore (2026) test whether Item-Writing Flaw (IWF) rubrics — a domain-general, textual evaluation requiring no student data — predict empirically estimated IRT difficulty and discrimination. Across 7,126 multiple-choice questions in STEM (physical science, mathematics, life/earth sciences), they used automated, LLM-assisted coding to show that IWF rubrics carry predictive validity for empirical IRT parameters, offering a scalable pre-deployment screen that complements or partially substitutes resource-intensive pilot testing.

  • IRT-based risk filtering for selective AI grading: Cvengros & Kortemeyer fit a two-parameter logistic IRT model to AI-graded handwritten-chemistry data and define the "risk" of accepting an AI judgment as the absolute deviation between the AI's normalized score and the IRT-expected probability of credit (Risk = |s−p|); accepting only items within a chosen tolerance of this Bayesian expectation flags "surprising" AI scores for human review, turning IRT from a pure score-aggregation tool into an operational acceptance/deferral mechanism for Automated Assessment — one that achieved alignment with human grading similar to simpler partial-credit thresholds but with lower human workload, though its logic is less transparent to non-technical audiences.

  • Divide-and-conquer calibration for continuously evolving banks: Jewsbury et al. (2026) treat IRT recalibration as a scaling problem rather than a fitting problem. When AI-based item generation and feature-based parameter prediction make a bank larger, sparser and continuously updated, refitting the full response history at every update grows steadily costlier; their consensus calibration instead calibrates each time period once and combines a new period with already-computed earlier posteriors. Two features separate it from existing IRT divide-and-conquer work: the periods do not share a latent metric, so each is linked to a reference metric by a robust Haebara criterion solved separately for every posterior draw (carrying linking error into the linked posteriors), and each period is its own hierarchical fit contributing an estimated prior, so the naive product of posteriors must have that prior divided out and a consensus prior reinstated — reducing to the Bayesian committee machine rule when the priors are fixed. Against a pooled benchmark on four quarterly periods of the Duolingo English Test, posterior means agreed at r = .998 (difficulty) and .991 (log-discrimination) with posterior SDs at r = .970 and .920, leaving mild under-dispersion (SD ratio 0.91–0.98) that was largest in the lowest per-period exposure tertile. It is IRT calibration re-engineered for the delivery conditions AI-generated item banks create.

  • How rarely IRT anchors instrument validation: an appraisal of teacher AI literacy instruments quantifies IRT's absence rather than its use. Zainal, Mohd Matore and Maat (2026) graded 33 instruments against a decision matrix adapted from COSMIN and Terwee et al. (2007); structural validity was strong, with 24 (72.7%) at Grade A through CFA, PLS-SEM or IRT modeling, yet none used IRT or Rasch as its primary evidence, and only five instruments (15.2%) reported measurement invariance or differential item functioning evidence. The authors argue for IRT and performance tasks alongside self-assessment to separate validated capability from reported confidence.

Connections

IRT is a foundation of Educational Measurement and Assessment Validity, underpins Adaptive Learning (adaptive item selection) and Learner Modeling and Adaptive Instruction, and connects to Psychometrically Aware AI (AI assessment aligned with measurement theory) and Knowledge Tracing. It features in LLM difficulty calibration for programming assessment.

Connected Concepts

Connected Articles

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.