Research Article
Semantic Variability of LLM-Generated Replies Across LLMs: Implications for Designing Conversation-Based Assessment
Synthesis: Hao (2026) examines whether LLM-generated replies remain semantically consistent when the underlying model changes, using messages from real collaborative problem-solving conversations to compare generated replies across four LLMs under conditions with and without preceding chat history. Results show that model choice and conversational context both affect response similarity and alignment with human replies: within-model similarity consistently exceeds between-model similarity, and including history meaningfully changes response content. The findings indicate that prompting and conversational context alone may not suffice to preserve response consistency across LLMs, highlighting the need for infrastructure and design strategies that maintain stable, comparable responses amid the rapid evolution of LLMs — a core validity challenge for conversation-based Assessment.
The adaptivity-standardization tension
Large language models and generative AI create a new paradigm for assessing complex skills such as communication and collaboration, which are difficult to measure at scale with conventional item formats because they are manifested through situative, dynamic interactions rather than isolated responses. LLM-based agents can engage learners in natural, adaptive interactions that elicit richer, more authentic evidence of these skills. Yet LLM-enabled interaction introduces a fundamental tension between adaptivity and standardization: standardized assessment assumes comparable learner inputs produce comparable assessment conditions, whereas LLM-based agents generate responses dynamically, adapting to learner input, conversational context, and model behavior. The same learner input may receive semantically different replies depending on the LLM, prompt, deployment setting, or prior chat history — a challenge compounded by the rapid evolution of LLMs as new models are released and older ones are retired.
The measurement stakes
From a measurement perspective, the key issue is not whether AI responses vary but whether they remain sufficiently consistent for measuring the intended constructs. AI-generated replies need not be identical in wording, but they should preserve the intended assessment function, sustain task-relevant interaction patterns, and provide comparable opportunities for learners to demonstrate the target construct. Uncontrolled variability risks introducing construct-irrelevant variance that threatens validity, reliability, fairness, and comparability — the foundational requirements of educational assessment (AERA, APA, & NCME, 2014). This connects directly to the knowledge base's Assessment Validity and Educational Measurement concepts and the broader literature on psychometrically aware AI. The goal is not to eliminate variability but to ensure it stays within acceptable bounds for the intended use.
Neuro-symbolic design and response consistency
Traditional conversational systems, such as those in intelligent tutoring systems, use modular architectures of intent detection, dialogue state tracking, and rule-based response selection. LLM-based systems instead rely primarily on prompts and conversation history to guide generation, enabling more natural interaction but also greater response variability. To optimize this tradeoff, neuro-symbolic approaches are increasingly adopted in conversation-based learning and assessment: LLMs interpret input, detect intent, classify dialogue moves, or generate responses, while rule-based actions are executed by a separate module or by an LLM prompted to follow predefined rules. Such designs depend on two LLM capabilities: accurately classifying utterances and generating semantically consistent responses in similar conversational contexts. Hao et al. have demonstrated the former; this study examines the latter.
Empirical findings
The study used dyadic online collaborative problem-solving science data (99 teams), focusing on 61 late-stage focal messages with highly relevant human replies. Four LLMs were evaluated (GPT-4o mini, GPT-5.4, GPT-5.4 mini, GPT-5.4 nano), each generating 100 responses per focal message under both no-history and with-history conditions, with semantic similarity measured via text-embedding cosine similarity. Key findings:
- Including chat history does not make LLM replies more similar to one another in all cases; its effect depends on the LLM. Replies remain consistently dissimilar from human replies, though history can slightly improve alignment.
- Within-model similarity consistently exceeds between-model similarity: within-LLM similarity ranged from 0.715–0.795, while between-LLM similarity ranged from 0.443–0.604. Models within the GPT-5.4 family were more similar to one another than to GPT-4o mini, indicating model architecture and lineage shape consistency.
- In the no-history condition, semantically similar focal messages elicited more similar replies; adding history introduced context that made replies less dependent on focal-message similarity alone.
- Median cross-history similarity was ~0.40–0.45, meaning that adding chat history meaningfully changed the semantic content of generated replies across all models.
Mixed-effects models confirmed that LLM type, history condition, and their interaction significantly affected both mean pairwise similarity and its variability, with responses more aligned with human replies tending to be more consistent.
Implications for conversation-based assessment
The results show that even for a single conversational turn, the semantic content of LLM-generated replies varies across models and context conditions. This means that maintaining consistent assessment interactions as LLMs evolve is an infrastructure challenge rather than merely a Prompt Engineering challenge. Systems that rely on LLM-generated responses require additional mechanisms to monitor, Benchmark, and control response behavior across model transitions — potentially including symbolic rules, response templates, and validation layers that preserve assessment-relevant functions despite underlying model changes. These considerations are particularly important in high-stakes settings, where variability may affect the comparability of assessment conditions and, ultimately, the validity, reliability, and fairness of AI-based assessment — connecting to Trust Calibration and the design of robust conversational agents for assessment.
Connected Concepts
- LLM
- Conversational AI
- Assessment Validity
- Assessment
- Educational Measurement
- Automated Assessment
- Prompt Engineering
- Generative AI
- Intelligent Tutoring
- Psychometrically Aware AI
- Trust Calibration
- Learning Analytics
Connected Articles
- Assessment Latent Structure Human LLM 2026 — Do assessment instruments measure the same thing for humans and LLMs?
- AI Feedback Enactment Workflow 2026 — Making AI-generated feedback matter: workflows and student enactment
- LLM Formative Feedback Systematic Review 2026 — Systematic review of LLM-based formative feedback
- Studychat Student Dialogues Chatgpt AI Course 2026 — The StudyChat dataset of student–LLM dialogues in an AI course
- Conversational AI Agents Umbrella Review 2026 — Umbrella review of conversational AI agents
- Conversational AI Informal Learning — Conversational AI in informal learning
- Socratic Tests Conversational Assessment — Socratic tests in conversational assessment
- Viberg Efficiency Effectiveness SRL LLM Help Seeking 2026 — LLM-mediated help-seeking in STEM
- Evaluation Age AI Output Evidence 2026 — Evaluation in the age of AI: output as evidence of learning
- LLM Tutoring Feedback Diagnosis Gap — LLM tutoring, feedback, and the diagnosis gap
Citation
Hao, J. (2026). Semantic variability of LLM-generated replies across LLMs: Implications for designing conversation-based assessment. arXiv:2608.24920 / AI in Measurement and Education (AIME 2026).