Research Article
Applying IRT to distinguish between human and generative AI responses to multiple-choice assessments
Synthesis: Strugatski and Alexandron (2026) established a theoretical and empirical foundation for using Item Response Theory to separate generative AI and human responses to multiple-choice Assessments. They investigated whether person-fit statistics (PFS), which evaluate how well an examinee's response pattern fits IRT model expectations, can distinguish between GenAI and human responses. Using data from two authentic assessment contexts — a high-school chemistry test and a national university entrance exam, each with roughly 1,000 human respondents — they demonstrated that PFS reveal significant differences between human responses and those generated by advanced chatbots, which appear as 'aberrant' responders. The study positions IRT as a robust, reproducible framework for characterizing and separating human and GenAI response patterns, supporting Academic Integrity in high-stakes testing.
Core Finding
Person-fit statistics grounded in IRT can reliably separate GenAI from human responses to multiple-choice questions, because chatbot response patterns exhibit systematic "aberrance" relative to IRT models fitted on human data. Yet the method has a clear boundary condition: separation diminishes as the proportion of GenAI responses in the data rises (from significant at 5% contamination to negligible at 25% on the quantitative instrument), since anomaly-based detection weakens as the detected behavior becomes more typical. Separately, different chatbots form a heterogeneous group of "intelligences," and newer model generations become not only more accurate but more human-like in their response patterns.
The Theoretical Rationale
A single MCQ response is just a digit — near-useless for attributing source — but the pattern across a collection of items can be diagnostic. Because GenAI and human cognition rest on fundamentally different architectures, the two are influenced differently by task dimensions, making some items surprisingly easy or hard for chatbots relative to what human-derived assessment models predict. Person-fit statistics quantify how well an individual's response pattern aligns with model-based expectations; they were originally developed to flag aberrancy from guessing or cheating, making them a natural fit for detecting GenAI's systematic misfit.
Methods: Data, Person-Fit Measures, and Designs
- Two authentic instruments: a formative chemistry diagnostic (Moodle, 22 MCQs, 931 high-school students, mean 71.49/100) and the quantitative-reasoning section of the Psychometric national university entrance exam (20 MCQs, over 4,800 examinees, mean 12.45/20). Both were Multimodal AI and dichotomously scored.
- GenAI responses: premium versions of six chatbots across three vendors in two cohorts — Gen1 (ChatGPT-4o, Gemini 1.5 Pro, Claude 3.5 Sonnet) and Gen2 (ChatGPT-5.2 Reasoning, Gemini 3 Pro Thinking, Claude 4.5 Sonnet) — each run 20 times per instrument with temperature held at default.
- Four non-parametric PFS: the Guttman error count (G), its normalized form (G*), and the U3 and ZU3 statistics (Van der Flier), implemented via the PerFit R package.
- Research questions: RQ1 compared human vs. GenAI PFS at 5%, 10%, and 25% contamination (Wilcoxon rank-sum); RQ2 compared the six chatbots' PFS (Kruskal-Wallis plus Dunn's post-hoc); RQ3 compared Gen1 vs. Gen2 convergence to human patterns via effect sizes, distributional tests, and scalogram/Guttman regularity.
Key Findings
- Human–GenAI separation (RQ1): At 5% contamination, GenAI responses were significantly more aberrant than human responses across all four PFS and both instruments (p < .001). At 10%, differences largely remained significant. At 25%, differences for the quantitative instrument became negligible (e.g., U3 p=.3222), while chemistry retained significance — ZU3 was the most contamination-sensitive measure, G the least.
- Between-chatbot heterogeneity (RQ2): Significant PFS divergence among the six chatbots on both instruments. PFS were more sensitive to generation-level shifts than to vendor "family"; older models were statistically indistinguishable from one another, with divergence appearing mainly between older and newer generations. On quantitative reasoning, GPT-5.2 and Gemini 3 formed a distinct cluster vs. older baselines, while Claude 4.5 had a unique profile.
- Generational convergence (RQ3): Gen2 models outperformed Gen1 (e.g., chemistry Gen2 M=80.8% vs. Gen1 71.5% vs. human 71.5%) and simultaneously converged toward human-like patterns — effect sizes for human-divergence shrank sharply (chemistry G: r=0.389→0.060). On the quantitative instrument, Gen2 PFS distributions became statistically non-significant relative to humans; on chemistry they remained significant. Scalograms showed Gen2 shifting toward the Guttman (internally consistent) pattern, meaning their perceived item difficulty more closely mirrors human learners'. The cross-instrument difference may reflect vendors' heavy investment in quantitative reasoning or the Psychometric instrument's open availability to crawlers.
What this means for practice
- Assessment designers. Use person-fit statistics as a probabilistic, wide-brush marker of heavy GenAI use rather than a classifier: at 5% contamination GenAI responses were significantly more aberrant than human responses on all four measures and both instruments (p < .001), but at 25% contamination the quantitative-reasoning differences became negligible (U3 p = .3222).
- Designers. Plan to refit detection as models update, because separation decays as generators improve: the authors' third research question shows newer chatbots becoming more human-like in their response patterns, so a person-fit threshold calibrated on one generation will not hold for the next.
- Researchers. Report which person-fit statistic you rely on — ZU3 was the most contamination-sensitive and the Guttman error count the least — and extend the approach to choice-level rather than dichotomous features, textual responses via NLP, other domains, and adaptive tests.
- Administrators. Apply the flag for prevalence estimation, such as comparing proctored with unproctored settings, and for routing aberrant respondents to in-person verification, remembering that elevated person-fit can also arise from learning disabilities or test anxiety rather than GenAI use.
Limitations
- The approach is anomaly-based, so its sensitivity declines as generative AI use becomes widespread, and it assumes responses span the whole instrument; selective mixing of human and AI answers is unexplored.
- It relies on labeled, clean response data, which is exactly what a real assessment in which some answers are AI-generated cannot guarantee.
- The between-chatbot comparison on the chemistry instrument had low statistical power, which the authors name rather than reading as a null result.
- Elevated person-fit scores can arise from legitimate sources such as learning disabilities or anxiety, and external validity is limited to six tools and two instruments.
Connected Concepts
- Item Response Theory
- Academic Integrity
- Generative AI
- Large Language Models (LLMs)
- Assessment
- Educational Measurement
Connected Articles
- Ensuring Academic Integrity through Automated Online Exam Proctoring: A Decade-Long Systematic Review — Automated online exam proctoring review
- Multimodal Item Parameter Estimation using Simulated Response Probabilities — Multimodal item parameter estimation
- Can AI Evaluate Assessment? A Study of Large Language Model Meta-Assessment Performance — Can AI evaluate assessment
- Do Assessment Instruments Measure the Same Thing for Humans and LLMs? A Latent Structure Analysis — Latent structure of human vs. LLM assessment
Citation
Strugatski, A., & Alexandron, G. (2026). Applying IRT to distinguish between human and generative AI responses to multiple-choice assessments. Computers and Education: Artificial Intelligence, 100668.