Research Article
LLMs in text linguistics teaching: An exploratory study with genAI novices in higher education
Synthesis: Brocca and Garassino (2026) use an action research design to examine how Generative AI novices in higher education design prompts and evaluate Large Language Models (LLMs) outputs in text linguistics teaching. Students' reports show novices refine prompts through trial and error, occasionally use in-context examples, or simplify complex instructions; end-of-term reflections reveal limited prompting competence but growing confidence in applying subject-specific knowledge. Students often attribute unsatisfactory results to the LLM rather than their own prompt formulation, and challenges such as anthropomorphizing the models and overgeneralising limited outcomes emerge. The study finds that students engage in metacognitive reflection with LLMs chiefly when disciplinary knowledge is well consolidated, and concludes that limited prompt-design knowledge remains a major obstacle — signaling the need for explicit Prompt Engineering instruction in disciplinary AI use.
Key Findings
Novice prompting behavior. genAI novices refine prompts through trial and error and occasionally use in-context examples or simplify complex instructions — a natural but limited progression without structured guidance.
Limited prompting competence. End-of-term reflections indicate limited prompting competence despite growing confidence in applying subject-specific knowledge, exposing a gap between perceived and actual AI skill.
Attribution bias. Students often attributed unsatisfactory results to the LLM's limitations rather than to their own prompt formulation, a form of misattribution relevant to Trust Calibration and AI Literacy.
Metacognition and disciplinary grounding. Students reported metacognitive reflection with LLMs particularly when disciplinary knowledge was already well consolidated, suggesting Metacognition in AI use depends on domain foundations.
Pedagogical implication. Limited knowledge of prompt design is a major obstacle; the authors argue for explicit prompt-engineering instruction within disciplinary teaching to help novices harness LLM potential in text analysis.
What this means for practice
- Instructors. Teach prompt design explicitly before the task: with no instruction, the ten novices in this study defaulted to a teacher–student dialogue format, refined prompts by trial and error, and rarely reached for in-context examples or few-shot prompting.
- Instructors. Treat a failed output as the next design problem rather than a verdict on the model: six of the eight students who commented on output evaluation stopped refining prompts after concluding the model could not detect anaphors, a conclusion the authors call overgeneralization from a single prompting experience.
- Instructors. Require students to compare their own manual analysis with the LLM's output on the same text; in this study that comparison produced perceived competence gains, reduced replacement anxiety, and was named by six of eight students as improved output-evaluation skill.
- Instructors. Front-load the disciplinary knowledge the activity depends on and assess it beforehand, since students reflected metacognitively only where subject knowledge was already consolidated and prior knowledge was never checked across the pairs.
- Teacher educators. Make attribution of failure an explicit debrief question in methods courses: here novices overwhelmingly blamed the LLM rather than their own prompt formulation, and only one of the eight students attributed unsatisfactory results to prompt strategy.
Limitations
- Ten students (nine women), aged 19–29 (M = 21.3, SD = 3.2), all native Italian speakers in one Italian-language linguistics seminar at a single Swiss university; the authors state this small, demographically unbalanced sample restricts external validity and prevents statistical comparison across gender or age.
- Evidence about perceptions and learning comes from self-reports and guided written reflections, which the authors concede may introduce response bias, with confirmation bias a risk in the qualitative coding; behavioral prompt logs covered only the first research question.
- Prior knowledge of text linguistics was never assessed before the intervention, so the authors cannot rule out that differences in disciplinary knowledge shaped how the pairs evaluated LLM output.
- The intervention ran with no preliminary prompt-engineering introduction, and the authors note the rapid pace of model development may limit long-term generalizability and reproducibility.
Citation
Brocca, N., & Garassino, D. (2026). LLMs in text linguistics teaching: An exploratory study with genAI novices in higher education. Computers and Education Open, 100414. https://doi.org/10.1016/j.caeo.2026.100414