On this page

Synthesis: Borchers, Jansen, and Weidlich (2026) conduct a rapid scoping review of how student interactions with large language models are defined and categorized, finding that research remains conceptually fragmented. Across 46 categorizations drawn from 33 studies, they identify substantial variation in data sources, category-construction approaches, and units of analysis, which prevents comparison across studies and understanding of when LLM use supports learning. The review argues for a convergent taxonomy of learning-oriented Large Language Models (LLMs) interactions. It provides a methodological foundation for the Student Experience literature and for interpreting Learning Analytics and self-regulatory evidence across heterogeneous studies.

Key Findings

  1. Across 46 categorizations extracted from 33 studies, the literature shows no shared meta-characteristic (Nickerson et al.'s term) for what student–LLM interaction categories describe — studies variously classify outputs, intentions, activities, dialogue moves, and strategies, so similar labels often name different phenomena and vice versa.
  2. Four recurring taxonomy types emerge: Type A (output/product), Type B (function/activity), Type C (dialogue move), and Type D (strategy/pattern). Type B dominates the corpus while Type D is least common, indicating most research abstracts away from concrete interaction sequences.
  3. Category construction is heterogeneous: roughly 43.5% of taxonomies were built inductively, 32.6% ad hoc, and 23.9% deductively, split about evenly between self-report and interaction-log data — so categories often reflect the measurement approach as much as the interaction itself.
  4. The authors propose a hierarchical taxonomy of interaction episodes — goal-directed, temporally bounded sequences of observable exchange between student and system — with six broad classes: knowledge acquisition, evaluative feedback, strategic guidance, dialogic inquiry, artifact refinement, and co-AI Regulation in Education.
  5. The episode-level unit occupies an intermediate level between isolated dialogue moves and broad cross-task strategies, connecting observable Student-AI Interaction behavior to established learning processes and enabling future integration with Learning Analytics and adaptive systems.

A Fragmented Field

Generative LLMs are now widely used by students, yet research on their role in learning remains conceptually fragmented. Existing studies describe student interactions using heterogeneous and often incompatible categories, ranging from speculative use cases to fine-grained dialogue acts, making it difficult to compare findings across studies and to understand when and how LLM use supports learning. The problem is fundamentally one of classification: Learning Theories that explain why interaction processes matter, such as the ICAP Framework and Self-Regulated Learning, do not supply a consistent vocabulary for categorizing and comparing learner–LLM interaction. Research on Knowledge Tracing, Scaffolding, and instructional support in LLM-based environments remains underdeveloped, and LLM exchanges lack the predefined event structures (feedback events, error states) that anchor Intelligent Tutoring analytics.

Method and Variation

This rapid scoping review follows the PRISMA-ScR extension, searching Scopus (July 2025) and pragmatically screening the 200 highest-ranked of 326 records, yielding 33 studies and 46 distinct taxonomies or taxonomic levels. It examines how interaction types are defined and constructed, coding each categorization's data source, construction approach, and unit of analysis. The authors find substantial variation reflecting divergent theoretical commitments and research purposes: self-report studies tend to capture perceptions and intentions, whereas log-based studies capture observable conversational behavior, and the two are rarely integrated. Categories are frequently posited ad hoc or built from qualitative coding, and the field is producing categories faster than it is integrating them, impeding cumulative knowledge building and systematic comparison.

Four Taxonomy Types

The manual coding distinguishes four analytical layers of student–LLM use. Type A taxonomies classify the products or deliverables an interaction yields — summaries, outlines, generated code — but say nothing about how interaction unfolds. Type B taxonomies, the dominant group, classify broad functions or learning activities such as brainstorming, verification, and feedback seeking, yet may conflate goals, cognitive processes, and observable actions. Type C taxonomies capture turn-level conversational behavior — follow-up questions, clarification, acknowledgment — but are too granular to connect readily to task and learning goals. Type D taxonomies describe recurring strategies spanning episodes or tasks, such as task decomposition or iterative prompting, but abstract from how strategies are enacted within particular exchanges. Coder disagreements clustered precisely at the boundaries between these types, especially where studies mixed functional purposes with observable conversational behavior.

The Interaction Episode and Six Classes

To integrate these divergent categorizations, the review adopts the interaction episode as the organizing unit: a goal-directed, temporally bounded sequence of observable exchange between a student and an LLM, occupying an intermediate level between isolated dialogue moves and broad strategies. The resulting hierarchical taxonomy groups episodes into six broad classes: (1) knowledge acquisition — seeking information, explanations, worked examples, and conceptual understanding; (2) evaluative feedback — submitting answers or artifacts for checking, critique, and revision guidance; (3) strategic guidance — requesting direction on how to learn, study, or organize work; (4) dialogic inquiry — advancing understanding through reasoning, hypothesis testing, and guided probing; (5) artifact refinement — improving a student-produced artifact through iterative editing and style refinement; and (6) co-regulation — reflecting on goals, progress, and strategies with LLM support, where regulation unfolds through observable exchange. This structure deliberately preserves distinctions among outputs, episodes, dialogue moves, and cross-episode strategies, focusing on enacted behavior rather than system functionality or outcomes.

Toward a Shared Taxonomy

The lack of shared terminology motivates a convergent taxonomy of learning-oriented use. The authors situate these interaction categories within Self-Regulated Learning and Learning Analytics, supporting Research Methods in AIED for synthesizing evidence on AI Feedback Quality and the conditions under which student-LLM engagement produces learning rather than mere completion. They position the taxonomy as complementary to other pathways toward convergence: meta-theoretical frameworks distinguishing learning- versus performance-oriented engagement, application of established theories such as ICAP and Bloom, and new AI-specific theories like Agentivism. Integrating interaction taxonomies with Learning Analytics and adaptive educational systems is identified as a promising direction for correlating instructional dialogue acts with rates of skill acquisition and Learning Gains.

What this means for practice

  • Researchers. Adopt the interaction episode — a goal-directed, temporally bounded sequence of exchange — as a shared unit of analysis, so findings built at different levels of analysis can be compared and cumulated.
  • Researchers. Integrate self-report and interaction-log evidence within the same study rather than letting the measurement approach dictate the categories.
  • Learning analytics designers. Instrument AI-mediated environments to capture episode sequences, then correlate instructional dialog acts with rates of skill acquisition to test taxonomies against real learning outcomes.
  • Learning analytics designers. Distinguish learning-oriented episodes (knowledge acquisition, dialogic inquiry, co-regulation) from completion-oriented ones (artifact refinement) when specifying what a tool should support.
  • Instructors. Reframe questions about whether LLMs support learning into questions about which interaction episodes support which learning processes, and specify the intended episode type in feedback and assessment designs.

Limitations

  • The rapid design relied on a single database (Scopus), pragmatically screened the 200 highest-ranked of 326 records, and supplemented the search with four author-identified studies, yielding 33 studies and 46 categorizations that should be read as selective rather than exhaustive.
  • The reviewed literature was concentrated in higher education and drawn from a rapidly changing period (2021–2025).
  • The proposed taxonomy was synthesized from reported categorizations and grounded in established frameworks rather than validated against an independent corpus of learner–LLM dialogues, so its coverage, episode boundaries, and coding reliability still require empirical examination.

Citation

Borchers, C., Jansen, S., & Weidlich, J. (2026). Toward convergence in student-LLM interactions: A rapid scoping review and taxonomy for learning-oriented use. EdArXiv preprint.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.