Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction

Created: 2026-06-29 | Tags: assessmentllmlearning-analyticshigher-edstudent-modelingbenchmarkk-12

Chenguang Wang, Ming Li, Xinyue Zeng, Zhuochun Li, Hong Jiao (2026) โ€” Computation and Language (cs.CL). ๐Ÿ“„ Full text (arXiv)

This paper introduces Epi2Diff (Episode to Difficulty), a framework that maps LLM reasoning traces into cognitively grounded episode sequences for predicting human item difficulty in educational assessment. The authors argue that difficulty should be viewed not only as a property of item text but also as an observable consequence of problem-solving burden. By analyzing reasoning traces from large reasoning models (LRMs), Epi2Diff extracts compact cognitive episodes that capture reasoning scale, effort allocation, and state transitions โ€” enabling interpretable student modeling without costly human calibration.

The work connects to knowledge tracing and IRT by offering a process-level view of item difficulty that complements traditional outcome-based models. It has implications for adaptive learning systems, where more precise difficulty estimates enable better personalized item selection, and for learning analytics, where reasoning trace analysis can provide instructors with fine-grained diagnostic information about which cognitive steps students find challenging.

This approach represents a novel intersection of LLM benchmarking and assessment design, suggesting that reasoning models can serve as cognitive proxies for human test-takers in K-12 and higher education settings. It connects to AI learning transfer by examining how model reasoning processes mirror human cognitive processes during problem-solving.

Related Pages

Citation

APA: Chenguang Wang, Ming Li, Xinyue Zeng, Zhuochun Li, Hong Jiao (2026). Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction. arXiv:2606.28186. Computation and Language (cs.CL).