Research Article
AI-Generated Interactive Fiction for Educational Use: A Pilot Study of Perceived Comprehensibility, Coherence, and Engagement
Synthesis: This pilot study (N = 22 STEM higher-education students) evaluates AI-generated interactive fiction as an educational medium. Narrative clarity and length acceptance rated positively, engagement hovered near neutral, and story-content coherence was the weakest dimension — with quiz integration emerging as the main usability bottleneck. The authors derive concrete design implications for interactive, narrative learning experiences.
Key Findings
- Generation is not enough. Generative AI can produce interactive narrative content at scale, but scenarios that are confusing, narratively inconsistent, or unengaging are unlikely to be useful in practice — quality must be validated with users, not assumed.
- Mixed user ratings. Participants rated narrative clarity (M = 4.11) and length acceptance (M = 4.14) positively, engagement sat near the neutral midpoint (M = 3.08, 95% CI straddling it), and story-content coherence was the weakest dimension by a clear margin (M = 2.92, with three of four items below the mid-point).
- Quiz integration is the bottleneck. Qualitative feedback identified artificial in-fiction motivation for quiz prompts (six of ten responses), abrupt setting/character changes that coincided with the lowest coherence scores (≤ 1.75), and the absence of story-level consequences for wrong answers as the main friction points.
- Self-report is decoupled from performance. Correlations between sub-dimension scores and gameplay metrics were all small-to-moderate (|ρ| ≤ 0.38, none significant), so low coherence ratings are not explained by frustration over quiz performance or by the English-language stimulus (negligible language-barrier correlations).
Background and Motivation
Generative AI is lowering the authoring bottleneck for Game-Based Learning in Higher Education: interactive fiction — short, choice-driven text-based scenarios — embeds educational content into branching narratives that learners explore through decision-making. Prior reviews of Large Language Models (LLMs)-based game and narrative generation find that runnable outputs are within reach while narrative quality and the meaningful integration of game rules and story remain difficult, with coherence flagged as a recurring weakness.
This pilot evaluates perceived quality rather than generation success. It builds on a domain-agnostic pipeline (SINE) that auto-generates interactive-fiction scenarios from structured content via an open-weight LLM with a frozen prompt, a deterministic validator, and a repair stage. The central design principle is intrinsic integration: learning content should be woven into the decisions and causal structure of the story rather than layered on as a separate quiz. The study therefore treats perceived story-content coherence as the pivotal dimension for whether AI-driven narrative instruction is usable at all.
Methods
The pipeline generated a controlled stimulus pool from 20 fixed educational seeds in a single STEM sub-domain (media technology — sampling, quantization, lossless and lossy compression, JPEG image compression), producing 48 scenarios after automated playability and validation filtering. Participants (N = 22 STEM-affiliated adults at the University of Lübeck) each played one randomly assigned scenario in the browser (5–10 minutes), then completed a ten-item positively-worded questionnaire adapted from the Narrative Engagement Scale, Transportation Scale–Short Form, MEEGA+, and the Intrinsic Motivation Inventory — grouped into four subscales: Narrative Clarity, Story–Content Coherence, Engagement, and Length Acceptance — plus an open-ended prompt. Analyses were primarily descriptive, with internal-consistency and exploratory Spearman correlations.
Results
Narrative clarity and length acceptance were rated positively, engagement hovered near the neutral midpoint with its confidence interval straddling it, and story-content coherence was the clear bottleneck. Internal consistency was acceptable-to-good for the multi-item subscales (Cronbach's α = 0.87 for coherence, 0.82 for engagement; Spearman-Brown = 0.63 for the two-item clarity scale). Qualitative feedback from 45% of participants converged on the coherence result: artificial in-fiction motivation for quiz prompts, abrupt location or character changes without narrative bridge, and the possibility of "click-through" without story-level consequences for wrong answers. Notably, the lowest coherence ratings came from participants with mid-to-high first-try-correct rates, and positive remarks appeared even in mixed-score sessions — suggesting the interactive-fiction format itself is accepted even where the current stimulus implementation is criticized.
Design Implications
Two families of design targets emerge. First, participant-driven changes: make the in-fiction motivation for why characters ask the player for technical knowledge explicit (surreal or fictional settings may free the scenario from real-world causal plausibility), suppress abrupt setting/character shifts in the generation or repair stage, attach story-level consequences to wrong answers rather than looping "try again," and mix purely narrative choices with learning choices. Second, pipeline-level changes: replace verbatim quiz-fidelity checks with a semantic-equivalence check (e.g. an LLM judge with human review) and strengthen playability validation from start-to-end reachability to traversal- or objective-coverage, integrated into the repair loop rather than applied post-hoc.
The study is a small convenience-sample pilot in a single STEM sub-domain with single-exposure sessions and deliberately no learning-outcome measure. Findings are specific to the pipeline-model combination (Qwen3 14B), and the single-rater qualitative coding, borderline two-item clarity reliability, and potential pro-technology self-selection are acknowledged limitations.
What this means for practice
- Instructors. Pilot any AI-generated scenario with real learners before classroom use: narrative clarity (M = 4.11) and length acceptance (M = 4.14) were rated positively, but story–content coherence was the weakest dimension (M = 2.92, with three of four items below the midpoint) and engagement sat at the neutral midpoint.
- Designers. Motivate quiz prompts inside the fiction instead of inserting them mechanically — six of ten open-ended responses identified artificial in-fiction motivation for quizzes as the main friction point.
- Designers. Attach story-level consequences to wrong answers rather than looping "try again", and suppress abrupt location or character changes in the generation or repair stage, because those defects coincided with the lowest coherence ratings (≤ 1.75).
- Instructors. Pair structured ratings with qualitative feedback and gameplay telemetry when trialing Pedagogical Agent narrative tools; here self-report was decoupled from gameplay performance (|ρ| ≤ 0.38, none significant), so low coherence ratings were not simply frustration over quiz scores.
- Researchers. Replace verbatim quiz-fidelity checks with a semantic-equivalence check (for example an LLM judge with human review) and extend playability validation from start-to-end reachability to traversal or objective coverage inside the repair loop.
Limitations
- Small convenience-sample pilot: 22 STEM-affiliated adults at a single university (University of Lübeck) each played one randomly assigned scenario, and pro-technology self-selection is acknowledged.
- Single exposure in one sub-domain: stimuli came from 20 fixed seeds in media technology (48 scenarios survived filtering), with 5–10 minute sessions (median play 5.7 minutes) and no repeated use or second domain.
- No learning-outcome measure by design: the study evaluates perceived quality only, so it cannot support claims about learning effectiveness.
- Measurement caveats: the ten-item instrument omitted reverse-scored items, the two-item clarity scale had borderline reliability (Spearman-Brown = 0.63), open-ended responses were coded by a single rater, and with N = 22 all correlations are reported as exploratory.
Citation
Rogosch, F., & Schrader, A. (2026). AI-Generated Interactive Fiction for Educational Use: A Pilot Study of Perceived Comprehensibility, Coherence, and Engagement. EDULEARN26 Proceedings, Article 1075.