Lan Anh Do, Hanling Jiang, Shuchin Aeron, Ayanna K. Thomas โ CogSci 2026 (accepted full paper). ๐ Full text (arXiv)
Synthesis
This study applies an extended 7-point ICAP framework (Interactive, Constructive, Active, Passive) to characterize cognitive engagement in collaborative dialogue, comparing trained human annotators with LLM-based labeling: in-context learning (ICL), zero-shot prompting, and self-reflective agents.
Human interrater reliability was robust across framework refinement stages (kappa = 0.906โ0.998), far exceeding ICL-based annotation (kappa = 0.541โ0.609) โ a large gap between human and LLM labeling of engagement.
The human-refined framework improved human agreement (ฮkappa = 0.10) but gave only modest gains to ICL LLMs (ฮkappa < 0.04); agent-refined frameworks improved cross-model agreement but stayed below human-refined performance.
Findings highlight the promise of reflective-agent approaches for scaling engagement measurement while underscoring that LLM annotation of learning processes still trails trained humans โ relevant for learning analytics pipelines that rely on automated discourse coding.
Related Pages
Citation
APA: Do, L. A., Jiang, H., Aeron, S., & Thomas, A. K. (2026). Measuring cognitive engagement in collaborative discourse with an extended ICAP framework. CogSci 2026. arXiv:2607.28651.