Research Article
AI Agents Can Now Navigate and Complete LMS Tasks: A Call for Pedagogical Innovation
Synthesis: Hadjisolomou and El-Haddad (2026) document that autonomous AI agents can now log into a university learning management system (LMS), read course materials, and complete unproctored asynchronous assessed work end-to-end with no student involvement — and argue this is fundamentally an Assessment Validity problem, not merely an integrity one. Three demonstrations on a live undergraduate psychology course supply the evidence: two quiz completions (one in ~12 minutes via Claude for Chrome, one in under 5 minutes via Perplexity Comet, scored 10/10) and a third in which Claude Opus fabricated a credible first-person life history for a discussion board after mining peers' posts. The wider public record includes at least 15 documented agent runs across three platforms (Canvas, Moodle, Brightspace) and seven agent tools. Applying Kane's argument-based validity framework, the authors argue agent completion removes the "human-production assumption" at the base of the scoring inference chain — so every downstream inference (generalization, extrapolation, decision), from course grades to program-review and accreditation evidence, rests on support that is no longer there. The failure concerns validity rather than integrity: an institution can punish misconduct and still lack grounds for the scores it reports. The response is design rather than detection — four principles for verified human presence (presence over product; integration over isolation; authenticity over genericity; low-stakes practice, high-stakes presence).
- Agent completion is now demonstrated and documented. In September and December 2025, agents (Claude for Chrome, then Perplexity Comet) completed and submitted a real 10-question open-book quiz on a live course from a single sentence of instruction, one scored 10/10 in under 5 minutes across 52 steps. A March 2026 demonstration showed Claude Opus fabricating a plausible "grandfather" life story for a discussion-board reflection after reading and tailoring to three peers' posts. The public record (≥15 runs, 3 platforms, 7 tools) establishes feasibility through convergence across independent observers.
- The categorical shift is assistance → completion. A generative chatbot keeps the student in the loop (reads, chooses, edits, submits); an agent removes the student entirely, executing a multi-step browser task from a single instruction. Nearly everything institutions built to manage AI assistance — citation policies, disclosure rules, process guidelines — presumes a student in the loop making choices, and so does not reach the agentic case.
- Vulnerability is diagnosable in advance. Three audit questions expose an assessment: Can the answer be derived from LMS or web materials? Can it be completed end-to-end by an autonomous agent? Is there any verified moment an agent cannot satisfy? Open-book quizzes, discussion posts, short-answer sets, take-homes, annotated bibliographies, and reflection journals all fail this test. Personalization changes difficulty, not category — fabricated experience defeats "write from your own experience" prompts.
- The failure is validity, not just integrity. Via Kane's argument-based framework, the scoring inference inherits a silent "human-production assumption"; agent completion removes its backing, invalidating the intended score interpretation for every unproctored asynchronous score — including honestly earned ones, since authorship is unverified. Detection does not rescue this: AI-text classifiers are unreliable and biased against non-native writers, and behavioral LMS monitoring is structurally blind because an agent produces the same clicks a student would.
- The response is design for verified human presence. Four principles — presence over product, integration over isolation, authenticity over genericity, and low-stakes practice / high-stakes presence — translate into low-effort changes (a short oral component on the highest-stakes assignment, process-visibility drafts, references to specific class sessions, shifting grade weight to synchronous/verified moments, randomly-sampled distributed defenses). Equity demands a menu of verified-moment options rather than a proctoring mandate, aligning with UDL. Collective accreditor (C-RAC) guidance addresses institutional AI use but not agent-completed student work, leaving a gap assessment professionals must close through program-review and reporting cycles.
Key Findings
- AI agents can complete and submit unproctored asynchronous LMS assessments end-to-end with no student involvement — demonstrated on a live course (quiz scored 10/10 in <5 minutes; fabricated discussion reflection) and corroborated by ≥15 independent public runs across three LMS platforms and seven agent tools.
- Agent completion is categorically distinct from AI assistance: it removes the student from the loop entirely, invalidating the institutional policies and assumptions built for the assistance case.
- Assessment vulnerability is structurally diagnosable in advance via three audit questions (answerable from LMS/web materials? agent-completable end-to-end? any un-forgeable verified moment?) — with open-book quizzes, discussion posts, reflections, and take-homes all exposed.
- Through Kane's argument-based validity framework, agent completion removes the human-production assumption at the base of the scoring inference, invalidating every downstream inference and even honestly earned unproctored scores, because authorship is unverifiable — a validity failure, not merely an integrity one.
- Detection cannot rescue validity (classifiers unreliable/biased; LMS monitoring structurally blind), so the solution is design: four principles for verified human presence with low-effort, equity-preserving options, and durable instruments (program review, reporting, curriculum maps) owned by assessment professionals.
Connected Concepts
- Assessment Validity — the central framework (Kane's argument-based validity, human-production assumption)
- Agentic AI — autonomous agents completing browser/LMS tasks
- Academic Integrity — contrasted with validity: integrity punishes conduct; validity asks whether the score is supportable
- Authentic Assessment — redesign for authentic demonstrations of capability
- AI Detection — rejected as insufficient (unreliable, biased, structurally blind)
- Assessment — the assessment-design ecosystem
- Remote Proctoring — narrows but does not close the gap
- Governance — accreditor guidance (C-RAC) and institutional policy
- Educational Policy AI — institutional responses
- Generative AI — the underlying technology
Connected Articles
- Chirikov AI Grade Inflation 2026 — AI task displacement as a mechanism of grade inflation (Chirikov 2026)
- Coauthorship Integrity Reconceptualising Assessment Validity For The Age Of Gene — Reconceptualising assessment validity for generative AI
- Asynchronous Oral Assessment 2026 — Asynchronous oral assessments in the AI era
- Beyond Detection Authentic Assessment AI 2025 — Beyond-detection authentic assessment
- Biology Grade Vulnerability GenAI 2026 — Vulnerability of course grades to AI-mediated dishonesty
Citation
Hadjisolomou, S. P., & El-Haddad, R. W. (2026). AI agents can now navigate and complete LMS tasks: A call for pedagogical innovation. Intersection: A Journal at the Intersection of Assessment and Learning (AALHE Conference Proceedings), 7(4), 54–69.