On this page

Synthesis: Hadjisolomou and El-Haddad (2026) document that autonomous AI agents can now log into a university learning management system (LMS), read course materials, and complete unproctored asynchronous assessed work end-to-end with no student involvement — and argue this is fundamentally an Assessment Validity problem, not merely an integrity one. Three demonstrations on a live undergraduate psychology course supply the evidence: two quiz completions (one in ~12 minutes via Claude for Chrome, one in under 5 minutes via Perplexity Comet, scored 10/10) and a third in which Claude Opus fabricated a credible first-person life history for a discussion board after mining peers' posts. The wider public record includes at least 15 documented agent runs across three platforms (Canvas, Moodle, Brightspace) and seven agent tools. Applying Kane's argument-based validity framework, the authors argue agent completion removes the "human-production assumption" at the base of the scoring inference chain — so every downstream inference (generalization, extrapolation, decision), from course grades to program-review and accreditation evidence, rests on support that is no longer there. The failure concerns validity rather than integrity: an institution can punish misconduct and still lack grounds for the scores it reports. The response is design rather than detection — four principles for verified human presence (presence over product; integration over isolation; authenticity over genericity; low-stakes practice, high-stakes presence).

  1. Agent completion is now demonstrated and documented. In September and December 2025, agents (Claude for Chrome, then Perplexity Comet) completed and submitted a real 10-question open-book quiz on a live course from a single sentence of instruction, one scored 10/10 in under 5 minutes across 52 steps. A March 2026 demonstration showed Claude Opus fabricating a plausible "grandfather" life story for a discussion-board reflection after reading and tailoring to three peers' posts. The public record (≥15 runs, 3 platforms, 7 tools) establishes feasibility through convergence across independent observers.
  2. The categorical shift is assistance → completion. A generative chatbot keeps the student in the loop (reads, chooses, edits, submits); an agent removes the student entirely, executing a multi-step browser task from a single instruction. Nearly everything institutions built to manage AI assistance — citation policies, disclosure rules, process guidelines — presumes a student in the loop making choices, and so does not reach the agentic case.
  3. Vulnerability is diagnosable in advance. Three audit questions expose an assessment: Can the answer be derived from LMS or web materials? Can it be completed end-to-end by an autonomous agent? Is there any verified moment an agent cannot satisfy? Open-book quizzes, discussion posts, short-answer sets, take-homes, annotated bibliographies, and reflection journals all fail this test. Personalization changes difficulty, not category — fabricated experience defeats "write from your own experience" prompts.
  4. The failure is validity, not just integrity. Via Kane's argument-based framework, the scoring inference inherits a silent "human-production assumption"; agent completion removes its backing, invalidating the intended score interpretation for every unproctored asynchronous score — including honestly earned ones, since authorship is unverified. Detection does not rescue this: AI-text classifiers are unreliable and biased against non-native writers, and behavioral LMS monitoring is structurally blind because an agent produces the same clicks a student would.
  5. The response is design for verified human presence. Four principles — presence over product, integration over isolation, authenticity over genericity, and low-stakes practice / high-stakes presence — translate into low-effort changes (a short oral component on the highest-stakes assignment, process-visibility drafts, references to specific class sessions, shifting grade weight to synchronous/verified moments, randomly-sampled distributed defenses). Equity demands a menu of verified-moment options rather than a proctoring mandate, aligning with UDL. Collective accreditor (C-RAC) guidance addresses institutional AI use but not agent-completed student work, leaving a gap assessment professionals must close through program-review and reporting cycles.

Key Findings

  1. AI agents can complete and submit unproctored asynchronous LMS assessments end-to-end with no student involvement — demonstrated on a live course (quiz scored 10/10 in <5 minutes; fabricated discussion reflection) and corroborated by ≥15 independent public runs across three LMS platforms and seven agent tools.
  2. Agent completion is categorically distinct from AI assistance: it removes the student from the loop entirely, invalidating the institutional policies and assumptions built for the assistance case.
  3. Assessment vulnerability is structurally diagnosable in advance via three audit questions (answerable from LMS/web materials? agent-completable end-to-end? any un-forgeable verified moment?) — with open-book quizzes, discussion posts, reflections, and take-homes all exposed.
  4. Through Kane's argument-based validity framework, agent completion removes the human-production assumption at the base of the scoring inference, invalidating every downstream inference and even honestly earned unproctored scores, because authorship is unverifiable — a validity failure, not merely an integrity one.
  5. Detection cannot rescue validity (classifiers unreliable/biased; LMS monitoring structurally blind), so the solution is design: four principles for verified human presence with low-effort, equity-preserving options, and durable instruments (program review, reporting, curriculum maps) owned by assessment professionals.

Connected Concepts

Connected Articles

Citation

Hadjisolomou, S. P., & El-Haddad, R. W. (2026). AI agents can now navigate and complete LMS tasks: A call for pedagogical innovation. Intersection: A Journal at the Intersection of Assessment and Learning (AALHE Conference Proceedings), 7(4), 54–69.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.