Distinguishing Artificial from Authentic: Evaluating LLMs for Detecting LLM-Generated Content

Created: 2026-07-24 | Tags: ai-detectionacademic-integrityllmhigher-ed

Leinonen & Denny (2026) โ€” Aalto University; University of Auckland. arXiv preprint (cs.CL). ๐Ÿ“„ Full text (arXiv)

As students increasingly use llms to draft written responses and program code, this study asks whether LLMs can reliably detect their own generated content across educational task types โ€” programming exercises, reflective writing, and short-answer questions. Using authentic student responses alongside multiple LLM-generated variants, the authors evaluate detection under varied prompting strategies and output formats. Detection proves highly task-dependent: it is substantially more reliable for programming tasks and longer reflective responses, but performs poorly for short-answer questions, where LLMs frequently judge their own outputs as more human-like than authentic student work. Prompt framing and response verbosity strongly affect detectability in reflective writing, with minor prompt variations sharply reducing accuracy, while programming detection is comparatively robust. The results highlight both the promise and the limits of LLM self-detection for academic-integrity, cautioning against standalone reliance and complementing dedicated ai-detection and plagiarism-detection work. They connect to identity-detection challenges in socially-fluent-ai-identity-detection and student-side dynamics in student-rationalization-ai-writing.

Related Pages

Citation

APA: Leinonen & Denny (2026). Distinguishing Artificial from Authentic: Evaluating LLMs for Detecting LLM-Generated Content. arXiv:2607.20446. arXiv preprint (cs.CL).