On this page

Synthesis: Using Transition Network Analysis (TNA), this study opens the "black box" of learner-chatbot interaction in scaffolded EFL writing, modeling how 119 Japanese junior-high learners moved between writing, requesting feedback, revising, and chatting across 4,651 sessions and 21,061 interactions with "Penny," a GPT-4o-powered writing chatbot. Two dominant behavioral loops emerged — a "Revision Loop" (feedback → successful error correction) and a "Chat Loop" (feedback → sustained dialogue → more feedback) — revealing that AI-scaffolded writing is a non-linear, dialogic process, not a linear submit-and-correct cycle. Critically, English proficiency significantly shaped interaction: high-proficiency learners engaged more in open dialogue and negotiation of meaning, while low-proficiency learners relied more heavily on repetitive corrective-feedback cycles (using the "check my writing" button rather than clarifying or negotiating). This highlights the need for differentiated chatbot design that moves beyond simple error correction to foster deeper cognitive engagement for all learners.

Key Findings

  • The revision-dialogue divergence. Following chatbot feedback, learner behavior split almost evenly: revise_writing (43.8%) and user_chat dialogue (37.3%). This dual pathway shows learners use the chatbot both as a corrective tool and as a partner for negotiating meaning.
  • Effective revision when chosen. When learners chose to revise, the transition to successful_uptake was 69% — the chatbot reliably facilitated error repair. Successful uptake (7.25% of events) was observed significantly more often than unsuccessful (1.66%) or no uptake (1.55%).
  • The Chat Loop. Dialogue often looped back: user_chat → penny_feedback (0.44) or → penny_chat non-corrective talk (0.56), creating sustained feedback-and-response cycles. This aligns with interactionist SLA theories where "pushed output" and dialogue facilitate hypothesis testing.
  • Proficiency shapes strategy. A chi-squared test (χ²=25.4, p<.003) showed interaction frequencies differ by proficiency. High-proficiency learners showed significant over-representation in conversational states (user_chat) and under-use of check_button; low-proficiency learners were the inverse — over-using check_button and under-engaging in dialogue, often triggering further corrective feedback (the chatbot functioned more as a persistent corrector than a partner for them).
  • Network structure. The TNA network (10 nodes, 26 directed edges, density 0.29, reciprocity 0.31) had strong "gravitational" hubs (penny_feedback), showing learners favored specific behavioral sequences rather than transitioning randomly.
  • Large Language Models (LLMs)-based classification was validated. Chatbot responses and learner uptake were auto-coded with gpt-4o-mini and validated against human coders with substantial inter-rater agreement (Fleiss' κ = 0.70 and 0.71).

Why this matters for education

This study is a model demonstration of Transition Network Analysis applied to AI-in-education log data — treating the learner-chatbot interaction as a process to be modeled temporally rather than judged by product (final essay score). For language learning and AI tutoring generally, it shows that learners of different proficiency levels interact with the same chatbot in qualitatively different ways, and that the pedagogical value of AI feedback depends on how learners actually engage with it. The finding that lower-proficiency learners get trapped in a corrective loop while higher-proficiency learners negotiate meaning suggests that AI writing tools may inadvertently widen proficiency gaps unless designed to actively scaffold negotiation and dialogue for less-advanced learners — a directly actionable implication for Scaffolding and adaptive chatbot design.

What this means for practice

  • Instructors. Design chatbot interactions to break the corrective loop: for lower-proficiency learners, prompt them to explain why a correction is needed or to propose their own fix before the answer is revealed, instead of letting them repeat the "check my writing" cycle.
  • Instructors. Differentiate support by proficiency: because high- and low-proficiency learners engaged the same chatbot in qualitatively different ways, adapt scaffolding, feedback framing, and dialogue prompts to each learner's level and metacognitive readiness.
  • Designers. Build negotiation into the chatbot itself — clarification requests, meaning-negotiation prompts, metalinguistic questions — since low-proficiency learners' dialogue disproportionately triggered further corrective feedback rather than genuine negotiation.
  • Researchers. Model the process, not just the product: use transition network analysis (and related sequence or temporal methods) on interaction logs to see whether feedback produces uptake, dialogue, or disengagement, because output metrics and self-reports can diverge from actual revision behavior and error-rate or final-score measures alone cannot evaluate an AI writing tool.

Limitations

  • The study covers 119 Japanese junior-high learners in a single course using one GPT-4o-powered chatbot, so the transition patterns are tied to that cultural, curricular, and model context and should not be read as general chatbot behavior.
  • It is a cross-sectional analysis of logged interactions (4,651 sessions; 21,061 interactions): it identifies behavioral differences by proficiency but cannot establish that AI scaffolding caused them.
  • The user_chat node is coarse — off-task talk, clarification requests, frustration, and social pleasantries are not distinguished — and automated coding of feedback versus chat reached only substantial agreement (Fleiss' κ = 0.70 and 0.71), so the transition probabilities are approximations rather than exact measures.
  • Immediate successful_uptake measures error repair only; it does not establish long-term language acquisition or retention, which would require longitudinal data.

Citation

Woollaston, S., Flanagan, B., Toyokawa, Y., & Ogata, H. (2026). Penny: Transition Network Analysis of Learner-Chatbot Interactions in Scaffolded EFL Writing. LAK '26 Transition Network Analysis Workshop.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.