Research Article
Comparing human and LLM ordered coding of qualitative data: How coding differences cascade through temporal analysis
Synthesis: Misiejuk and colleagues (2026) test whether a modern generative LLM can automate ordered coding of learner text, where a single utterance may carry several codes and the assignment order matters — a case classic language models such as BERT cannot handle. Coded by two researchers and by gemini-2.5-flash under a sliding five-message context window, 6,013 Discord messages from three master's-level courses reveal systematic, statistically significant differences between machine and human coding across structural, transitional and code-level metrics (χ2(7) = 1427, p < 0.001): the model over-assigned Reaction (2,107 vs. 823) and Discussion (1,092 vs. 756) while humans produced far more Monitoring (1,185 vs. 283) and CoRegulation (1,352 vs. 770), binary accuracy stayed at 0.693–0.888 with very low recall for Feedback (0.131) and Monitoring (0.100), and the model detected far more multi-step transition patterns than its human counterparts. The paper contributes two evaluation approaches for ordered coding and a consistent-context prompting method.
Key Findings
- Automating qualitative coding of learner text has been a long-standing goal of learning analytics because it is an essential step toward timely, scalable feedback — yet it is especially hard for ordered coding schemes required by temporal analytical methods, where a single utterance can carry more than one code and the assignment order matters.
- The problem goes beyond multi-class and multi-label classification, which means it cannot be easily tackled with classic language models such as BERT; the study instead evaluates modern generative LLMs — here, gemini-2.5-flash, prompted with a sliding context window of up to five preceding messages from the same conversation.
- The empirical base was 6,013 Discord messages from small-group collaboration in three master's-level courses at the University of Eastern Finland, coded by two researchers using a hybrid eight-code scheme covering cognitive and regulatory processes, and in parallel by the Large Language Models (LLMs).
- The paper makes two main contributions: it presents two evaluation approaches for assessing the quality of ordered data coding and the usability of LLMs in automatically coding ordered processes, and it demonstrates an LLM prompting method that leverages a consistent context window.
- Results reveal systematic and statistically significant differences between LLM and human coding across structural, transitional, and code-level metrics, for both binary and ordered tasks (overall code-frequency association χ2(7) = 1427, p < 0.001).
- Frequency discrepancies were large: the LLM over-assigned Reaction (2,107 vs. 823 human codes) and Discussion (1,092 vs. 756), while humans produced far more Monitoring (1,185 vs. 283) and CoRegulation (1,352 vs. 770); only Coordination (2,591 vs. 2,439) and Socializing (1,329 vs. 1,182) aligned within chance.
- Binary agreement was modest: overall accuracy ranged 0.693–0.888, recall was very low for Feedback (0.131) and Monitoring (0.100), and Cohen's kappa ranged 0.090–0.539 with most categories below 0.400.
- The LLM detected multi-step transition patterns far more often than humans (e.g., Coordination→Socializing: 46 vs. 378 occurrences; Discussion→Coordination: 19 vs. 268), and its detection was position-dependent — Discussion and Monitoring found early in messages, socio-emotional codes late — whereas human coding was more evenly distributed across message positions.
- Because classification errors can propagate through automated feedback systems, relying on LLM outputs risks amplifying inaccuracies and producing misleading interpretations of learning processes.
Study Design & Method
The researchers treat coding quality not as a single accuracy number but along multiple dimensions: structural properties of the coded sequence, transitional patterns between consecutive codes, and code-level agreement for binary and ordered assignments. LLM outputs were produced with a prompting strategy that maintains a consistent context window, giving the model access to surrounding textual context needed for accurate interpretation (up to five preceding messages per utterance, with position slots T1–T8 inside messages). Each LLM-coded output was then compared against human coding using the two proposed evaluation approaches, including Transition Network Analysis (TNA) with permutation tests and centrality comparisons — e.g., significant betweenness and in-strength differences for Discussion, Feedback, Socializing, CoRegulation, and Consolidation.
What this means for practice
- Researchers. Treat LLM coding as a human-in-the-loop proposition: validate its output with the proposed evaluation metrics before feeding it into automated feedback, which matters for Automated Assessment, Educational NLP and any Learning Analytics analysis built on LLM-coded transcripts.
- Learning analytics designers. Expect systematic rather than random deviation: the model foregrounded surface-level social exchange (Socializing, Reaction) and under-represented regulatory and collaborative processes (Coordination, CoRegulation, Monitoring), which risks painting a more rigid, socially-driven picture of group learning than human coders would produce.
- Learning analytics designers. Protect the temporal analyses downstream: plausible-looking coding can cascade its deviations into later results, so the size of the disagreement matters more than how fluent the coding looks.
- Researchers. Reconsider what counts as ground truth: human coding is treated as the reference here even though coding is interpretive and, in an ordered scheme, disagreement can come from where the text is segmented as much as from which code applies.
Limitations
- One dataset, one institution. The evaluation used three courses at a single institution with Discord as the collaboration medium, so generalizability to other learning processes, discourse types and educational levels is untested.
- One model, one prompt, one codebook. Results rest on a single LLM (gemini-2.5-flash), a single prompting strategy and a single codebook; even slight prompt changes can alter results, and a simpler codebook might have performed better.
- A fixed context window, untested. The window was set at up to five preceding messages, with no systematic sensitivity analysis of its size.
- Human coding treated as ground truth. Qualitative coding is inherently interpretive, and in ordered coding disagreement can stem not only from whether a code is present but from how the text is segmented into codes.
Citation
Misiejuk, K., López-Pernas, S., Oliveira, E. A., Eagan, B., & Saqr, M. (2026). Comparing human and LLM ordered coding of qualitative data: How coding differences cascade through temporal analysis.