On this page

Synthesis: Tsiligkiris (2026) examines how the depth of students' Large Language Models (LLMs) interaction relates to task quality and immediate recall. In a controlled session, 22 postgraduate students completed a pre-test, an LLM-assisted neuroeconomics case task, and an immediate post-test, with fine-grained interaction logs capturing Depth (proportion of "why/how/explain" prompts), Volume, and Pacing. Students showed large immediate learning gains (Cohen's dz = 2.12), and interaction depth was positively associated with independently marked task quality (β = 6.27) — what students ask matters more than how much they ask. However, depth was not associated with immediate recall; gain scores were driven by baseline knowledge. This dissociation between performance quality and short-term recall aligns with the distinction between elaboration-driven comprehension and retrieval-driven consolidation.

Key Findings

  1. Explanation-seeking depth predicts task quality, not volume. Depth (the proportion of explanation-seeking "why/how/explain" prompts) was positively associated with task quality beyond baseline knowledge and overall Volume (β = 6.27, p = .006) — a one-SD increase in explanation-seeking prompts corresponded to ~6 additional marks on a 0–100 scale. The composition of interaction matters more than the amount.
  2. Depth did not predict immediate recall. Depth showed a null association with immediate post-test recall (β = −0.014, p = .728); gain scores were strongly associated with baseline knowledge (β = −0.161), consistent with reduced headroom among higher-baseline students.
  3. A dissociation between performance and retention. The pattern aligns with cognitive psychology: explanation-seeking (elaboration) improves comprehension and applied performance, while retrieval practice — not fluent explanation — consolidates retention. In LLM-supported study without explicit retrieval demands, learners may experience high fluency with limited need to retrieve knowledge unaided.
  4. Depth may function as a productive scaffold for applied outputs. Depth-oriented use appears to scaffold applied task performance (aligned with constructive engagement and the ICAP framework), even when it does not translate into improved recall — a nuance on the cognitive-offloading account.
  5. Methodological contribution. The study demonstrates a replicable, privacy-preserving instrumentation pipeline linking turn-level conversational telemetry (Depth/Volume/Pacing) to learning outcomes — a process–outcome modeling approach for LLM interactions.

Comprehension vs. retention in LLM-supported learning

The central theoretical contribution is separating comprehension from retention in LLM-mediated learning. Depth-oriented dialogue can enhance reasoning, integration, and production quality (benefiting applied tasks) while leaving memory encoding largely unaffected — because the LLM supplies complete, coherent explanations on demand, learners may allocate less effort to internal retrieval and reconstruction, consistent with Cognitive Offloading and the "illusion of competence" literature. The author argues this is not merely a null result but a theoretically informative dissociation: elaboration drives comprehension; retrieval practice drives consolidation. This echoes the performance-vs-learning distinction central to the knowledge base.

Pedagogical and practical implications

To translate comprehension gains into durable retention, the author recommends: (1) embedding retrieval demands after LLM use via closed-tool outputs (short-answer questions, concept maps from memory, teach-back explanations without AI); (2) separating scaffolding from checking — use LLMs for clarification and feedback during learning but include distinct checkpoints where learners demonstrate independent recall without the model; and (3) prompt design that requires learner-generated reasoning (e.g. asking the model to pose questions, generate Misconceptions about AI or counterexamples, or critique the learner's own explanation) rather than producing finished answers. These align with Scaffolding and productive struggle principles.

Limitations

The study has a single-group design (no causal claims), a modest sample (n = 22) leaving moderation/clustering underpowered, immediate-only testing (no delayed retention), and a keyword-based depth proxy that captures the surface form of explanation-seeking rather than its underlying quality.

Connected Concepts

Connected Articles

Citation

Tsiligkiris, V. (2026). What students ask matters: LLM interaction depth, task quality, and immediate recall in higher education. International Journal of Educational Technology in Higher Education, 23(44).

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.