On this page

Synthesis: Large Language Models (LLMs)-Assisted Sentiment Analysis for Mixed-Methods Education Research demonstrates how LLMs can serve as scalable qualitative research assistants, enabling researchers to investigate multiple demographic variables simultaneously rather than being limited to simple binary comparisons. Using 151 longitudinal written reflections from a study abroad program, the authors show that LLM-assisted sentiment analysis combined with statistical testing can uncover granular patterns: prior experience living abroad was the only personal variable that significantly impacted students' sentiments about their language and communication behaviors. This workflow bridges computational and qualitative methods, suggesting that LLMs can reduce the bottleneck of manual qualitative coding without replacing the interpretive depth of thematic analysis. The approach has implications for Higher Education research methodology, complementing existing Learning Analytics pipelines and extending mixed-methods capabilities beyond what has been possible with Automated Grading and Formative Assessment systems alone. The paper connects to discussions about Educational Development in equipping researchers with AI literacy for methodological innovation, and relates to AI Literacy as both a tool for researchers and a consideration in how computational methods change the practice of qualitative inquiry.

  • LLM-assisted sentiment analysis enables comparison across 7 identity/lived-experience variables simultaneously
  • Only prior experience living abroad significantly impacted students' communication sentiments
  • The workflow preserves qualitative depth while adding statistical power
  • Implications for Learning Analytics and Student Experience research methodology

What this means for practice

  • Researchers. Treat the model as an additional rater rather than ground truth: label a sample by hand as well, compute Cohen's kappa against the LLM at each time point (this study reported κ = 0.52 to 0.67), and reexamine quotations where the two disagree.
  • Researchers. Partition sentiment counts by each identity and lived-experience variable and run the tests per partition (Shapiro-Wilk for normality, then Student's t-test or the Wilcoxon rank-sum test) instead of reporting one pooled trend — only prior experience living abroad produced statistically significant differences across all three reflection time points.
  • Researchers. Keep manual quote extraction in the pipeline and delegate only the well-scoped labeling task; the LLM produced irrelevant quote lists here, so all analyzed quotations were extracted by hand.
  • Researchers. Budget for prompt iteration, output filtering, and the scripting and data-management setup before promising laboratory efficiencies — the study's prompts needed repeated tuning, the model produced extraneous commentary and duplicate labels for single quotes, and the authors question the payoff for small or one-off datasets.
  • Researchers. Use human–LLM divergence as an analytic resource: here the model assigned affect to statements the human coder read as neutral and descriptive, and those disagreements drove inter-rater discussion that sharpened coding criteria.

Limitations

  • The 151 reflections came from 51 of 80 students at a single large, research-intensive institution in the southwestern United States.
  • Measured effects were small to medium (Glass rank biserial coefficients from -.289 to .383 for the significant comparisons), so findings may not extend to other study abroad programs, particularly language immersion programs or longer stays.
  • The program was a month-long condensation of a semester-long course, and the analysis used three of the four reflection time points (two students omitted Reflection 3), with free time and pre-program arrival in Japan likely shaping what students wrote.
  • Model agreement was weakest at the first time point: 20% of Llama3's sentiment labels for Reflection 1 were not identified by the human coder, compared with 8% for Reflection 3 and 7% for Reflection 4.

Citation

Xiomara Gonzalez, Gabriella Coloyan Fleming, Andrew Katz, Maya Denton, Jessica Deters (2026). LLM-assisted sentiment analysis for integrated computational and qualitative mixed methods education research: A case study of students' written reflection assignments. arXiv preprint.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.