On this page

Student-AI interaction — the patterns, processes, and cognitive work in how learners engage with generative AI systems during learning and problem solving. Research here characterizes what students ask of AI, how prompts and dialogues evolve, and how interaction quality relates to learning outcomes, Cognitive Offloading, and Learner Agency.

Questions to Consider

  • Think about the last few prompts you (or a student) wrote to an AI. Would you describe most of them as asking for the answer, or asking the AI to explain, probe, or evaluate? What do you suspect that pattern does to learning?
  • The research finds that a small subset of question types accounts for most student inquiries, and that the questions change as a task progresses. Why do you think students' questioning narrows, and what does that suggest about how they're using the tool?
  • The page claims shallow, answer-seeking prompts are associated with reduced learning and over-reliance, while reflective, verification-oriented interaction supports understanding. What do you think separates a 'good' prompt from a 'bad' one — and is that the student's responsibility or the tool's design?
  • If interaction quality is shaped by task context and scaffolding rather than being a fixed trait of the student, how might a course or tool be redesigned to invite a wider, more productive range of inquiry?
  • How would you know whether a student's fluent AI dialogue reflects genuine learning or just skilled delegation — and what would you check to find out?

Introduction

Student-AI interaction is the observable surface of learners' engagement with generative AI — the questions they pose, the prompts they write, the way they negotiate and verify AI output, and how those patterns shift across task stages and over time. It sits at the intersection of Student Experience, Prompt Engineering, and Learning Analytics, and is central to debates about whether AI use in education represents genuine learning or over-reliance. Where Human AI Collaboration frames the high-level division of cognitive labor between people and models, student-AI interaction is the concrete, measurable enactment of that relationship — the specific inquiries, prompts, and negotiation moves learners make moment to moment.

What students ask AI

A core strand of research measures the types and quality of student inquiries. Studies apply taxonomies of question types — for example the Graesser et al. 18-type taxonomy — to classify student-AI interactions, often using few-shot classifiers to scale the analysis across hundreds or thousands of interactions. Findings indicate that a small subset of question types accounts for the majority of student inquiries, and that the questions students ask change substantially as a task progresses (e.g., Analysis of Types of Inquiries in Student-AI Interaction: A case study of two CS2 tasks). This task-dependence matters: interaction quality is not a fixed trait of the student but is shaped by problem context, Scaffolding, and the affordances of the AI tool. At the youngest ages, Vahedian Movahed & Martin (2025) found children (ages 6–14) actively tested a chatbot's credibility by posing known-answer questions (e.g., "how big is a t rex") — an expression of epistemic self-agency — while the modal child asked only 1–3 questions and a standout first-grader asked 21, underscoring how developmental and individual variation shapes the questions learners pose.

The taxonomies themselves do not yet agree: across 46 categorizations from 33 studies, similar labels name different phenomena, so the review proposes the interaction episode — a goal-directed, temporally bounded exchange — as a unit spanning knowledge acquisition, evaluative feedback, strategic guidance, dialogic inquiry, artifact refinement, and co-regulation (Borchers, Jansen & Weidlich (2026)).

Complementing these taxonomy studies, Yan et al. (2026) characterize the dialogue form of student inquiries. Among 38 undergraduates completing unsupervised argumentative essays, 76.32% used a single-turn ask–get-answer–stop pattern — typically pasting the assessment title without specifying their needs and resubmitting identical prompts when dissatisfied — and 78.94% touched GenAI only at the start (ideas, background) or end (polishing, length) of a task, keeping it separate from reading and independent writing; only 23.68% sustained iterative back-and-forth dialogue with follow-up questions and their own reasoning. The authors place these patterns on a spectrum from cognitive outsourcing to cognitive reallocation — the GenAI-era analogue of surface versus deep approaches to learning — noting that most students conceived the tool as an upgraded search engine, which constrained them to the outsourcing end.

Reframing these patterns as epistemic work, an analysis of 200 co-programming chat sessions found 78.8% of student–GenAI interactions ran on non-mastery aims and strategies such as outsourcing or verification-seeking, and only 11.1% coupled mastery-oriented aims with epistemic justification (Wu (2026)).

Coding 50 sampled exchanges gives a three-way typology of the same behavior: Zahra (2026) classified 46% as Passive Review, where the model acts as an oracle and output is accepted with little scrutiny, 18% as Direct Question, and 36% as Strategic Dialogue, pairing the distribution with a constraint-first design argument (kappa = .48 between coders). A seven-assignment survey of 211 computing students reaches the same conclusion from the other direction: Nizamani et al. (2026) measured LLM adoption from 89.6% in algorithms to 15.2% in software engineering and attributed the spread to assignment complexity, verifiability and scaffolding rather than to the subfield.

Interaction quality and learning

A complementary strand links the form of interaction to learning. Shallow or habitually narrow prompts (asking AI to produce the answer rather than to explain, probe, or evaluate) are associated with reduced learning and increased over-reliance, whereas reflective, verification-oriented interaction supports Metacognition and durable understanding. This connects student-AI interaction directly to Intelligent Tutoring design: systems can be built to invite a wider, more productive range of inquiry and to scaffold question-asking rather than merely answering. Interaction need not run through answer-seeking prompts at all — when AI critiques students' own work, the exchange becomes a reflective, verification-oriented dialogue: in Oppenheimer, Cash & Connell Pensky (2025), learners' responses to Large Language Models (LLMs) essay Feedback showed reflection in 92.7%, acceptance in 93.6%, and active rebuttal of LLM claims in 87.8% of cases (inter-rater κs = 0.81–0.89), and their response-to-feedback quality improved across iterations as a learnable skill. The reflective side of interaction need not run through direct prompting — in CoBi, students engaged with classroom-level AI visualizations of their own collaborative speech and deliberated over when the AI's classifications seemed off, turning apparent misclassifications into opportunities for calibrating their understanding of the AI's capabilities and limits (Trust Calibration) rather than purely accepting its output. The link between interaction form and learning has a limit, though: Page (2026) separates conversational iteration from cognitive iteration on the ground that a learner can refine an output across many turns while the mental model behind it stays unchanged, so the evidence of learning is a change in the learner's cognitive position rather than the length or fluency of the exchange — turn count is not a proxy for engagement depth.

The interaction, not the tool, fixes the pathway: in a field experiment with nearly 1,000 high-school mathematics students, unrestricted GPT-4 raised practice scores 48% but cut unassisted exam scores 17%, while the same model restricted to teacher-designed hints raised practice 127% and largely erased the deficit (Lee, 2026).

Whether students prompted effectively, not how much, tracked success: AI Query Efficiency and AI-Driven Problem-Solving were the strongest predictors of academic performance across 128 engineering students, and stayed significant after controlling for GPA (Isaza Dominguez et al. (2026)).

Behavioral context is a separate signal from the question itself: in TutorTrace's four deployments (480 learners, ~180,000 IDE events), conditioning help on a learner's recent behavioral state cut intervals between queries with no independent work from 50.0% to 20.7%, and imminent queries were predictable from behavior alone (AUROC = .726) (Barron et al. (2026)).

Density is not the mechanism: alternating students between voice and text nearly doubled dialogue turns per minute (1.34 versus 0.75) without changing weekly mastery, and the authors read the typed channel's median 26-second deliberation before the first keystroke as the encoding act itself (Yang, Van Alstyne and Dellarocas (2026)).

A verification loop can still stay shallow: lag sequential analysis of a four-week LLM-agent deployment in an undergraduate database course found a significant Query–Evaluation–Query cycle, yet strong self-transition loops within lower-order states and only 3.92% of interactions reaching higher-order cognition (Li et al. (2026)).

Bernstein and Sibia (2026) document an iterative-filter pattern in how CS2 students handle GenAI explanations (Flawed but Memorable: Student Critical Reception of Interest-Personalized GenAI Analogies in Computing Education): they cross-reference against lecture notes, demand provenance ("I would be a lot more doubtful... without one"), and probe with follow-up questions for inconsistency rather than issuing a single accept-or-reject judgment. Students also read explanations for whose knowledge and background they assumed — assumed prior knowledge beyond the syllabus, default sport and gaming references ("the more male-dominated side of computing"), and excessive repetition all functioned as signals about the imagined reader, with over-scaffolding read as condescending rather than merely inefficient.

Consulting AI at the right point in a task matters as much as the wording of the individual prompt: in the same study, Yan et al. (2026) found the reallocation-oriented minority alternated independent work with GenAI consultation and reported unchanged total effort but a shifted focus — moving resources from searching to checking argumentative quality and balance, and writing reflection notes after sessions to counter shallow retention — whereas learners with mastery goals but weak AI literacy fell into an "efficiency paradox," offloading "not by intention, but by default."

From interaction to pedagogy

Characterizing student-AI interaction informs Learning Design: instructors can notice when students' questioning patterns are narrow or shallow and design interventions that broaden inquiry; Teaching shifts toward coaching students to interact productively with AI. It also grounds AI Literacy curricula that treat effective prompting and verification as learnable skills rather than innate abilities.

Design levers can shift that default: in a roughly 70-student asynchronous physics course the graded artifact was the dialogue transcript rather than an answer, yet many students still asked long lists of questions without answering the model's Socratic follow-ups — the failure mode the criteria were written to catch (Rodriguez & Wulff (2026)).

Teacher-authored prompts are one such lever, but enacted rigor lags the target: across 1,479 conversations, 38% under-reached the teacher's intended Depth-of-Knowledge level, approaching 50% at DOK 3, while explicit finish lines narrowed that gap by 0.22 levels and a "no direct answers" guardrail cut AI final-answer rates by 8.5 percentage points (Liu et al. (2026)).

Non-use is itself an interaction pattern that Pedagogies and Teaching Strategies must plan for. Zou et al. (2026), studying 85 student teachers in three courses where GenAI use in Assessment was explicitly permitted, found that 62.4% (53 of 85) declined to use it at all, far below the 79–83% adoption seen in comparable UK and Australian surveys, and that adopters' use was shallow and corrective rather than generative (proofreading 43.8%, clarity checks 34.4%, text generation only 18.8%). Their choices tracked assessment design and institutional culture rather than technical difficulty: 41.5% of non-adopters feared being wrongly accused of plagiarism, and nine of eleven interviewees read the permissive policy itself as a possible "trap." The gap between 32 survey-reported users and 28 self-declarations shows that students' reported AI interaction is shaped by graded consequences — a measurement caveat for learning-analytics accounts of student-AI interaction.

Discipline and Cognitive Engagement in Student-AI Chat

  • Discipline-associated cognitive engagement in student-AI chat. Chang and Li (2026) analyze student prompts to AI across 116 courses with a within-person, cross-discipline design, showing that student-AI conversations reflect discipline-associated cognitive engagement rather than fixed individual interaction styles. Roughly 62% of prompts encoded higher-order cognitive demand overall, but Bloom-level profiles differed sharply by discipline: STEM courses elicited Apply-prevalent prompts (20.8%), language courses Understand-prevalent (31.7%), and social science courses Create-prevalent (33.8%). Paired within-person comparisons confirmed the same students produced significantly more higher-order prompts in social science than STEM courses (pooled n = 16, p < .001), and course-level variation exceeded student-level variation — a strong argument that AI teaching assistants should be designed and evaluated with disciplinary context in mind.

Connected Concepts

Connected Articles

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.