On this page

Synthesis: Proposes Cognitive Diagnostic Profiling (CDP), a zero-shot framework that dramatically improves Large Language Models (LLMs)-simulated examinee alignment with human test-takers. With CDP, IRT difficulty Spearman correlations rose from 0.24 to 0.90, and RMSE fell from 6.31 to 0.90. Makes LLM-simulated examinees practical for operational test development.

Relevance to AI in Education: This paper contributes to the understanding of Automated Assessment, Personalized Learning, and Student Experience. The findings have implications for Adaptive Learning systems, Formative Assessment design, and the broader Edtech Platform landscape. Future work should explore how these results generalize across STEM Education and Higher Education contexts.

This research connects to the growing body of work on AI Literacy and Teaching, highlighting both the promise and limitations of AI tools in educational settings.

What this means for practice

  • Researchers. Condition simulated examinees on explicit cognitive profiles rather than asking a model for generic answers: profile conditioning raised 1PL item-difficulty agreement from a Spearman correlation of 0.24 to 0.90 in the strongest configuration.
  • Researchers. Report alignment at all three levels the authors use — ability-distribution overlap (OVL), mastery-profile correlation and item-difficulty recovery — since a model can look good on one and poor on another.
  • Designers. Give simulated examinees a mastery profile and sample it under a realistic population distribution; the uninformative condition raised overlap in seven of eight configurations, and the informative one improved alignment further in seven of eight.
  • Administrators. Use LLM-simulated examinees to triage new items before committing to costly human pretesting, then confirm final parameters on a human sample.

Limitations

  • The evidence comes from a single instrument: the Tatsuoka fraction-subtraction dataset with 15 items, five attributes and 536 examinees, and only one five-attribute decomposition was examined.
  • Even with CDP, the best configuration produced 133 distinct response patterns against 267 from human examinees, so simulated diversity still falls well short of the human benchmark.
  • The items are publicly distributed in the R package CDM and widely analyzed, so exposure in LLM training corpora cannot be ruled out; the authors call for replication on secure, unreleased item pools.
  • The informative condition used an in-sample prior estimated from the same 536 examinees that define the evaluation reference, so it marks an upper bound on prior benefit, and each cell was generated once at default sampling settings, leaving generation variability unquantified.

Citation

Wenjie Zhou, Yunting Liu, Renjiao Tang, Mark Wilson (2026). Aligning LLM-Simulated and Human Examinees for Psychometric Calibration: A Cognitive Diagnostic Profiling Approach. arXiv preprint.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.