On this page

You have already made the discovery that forces this question. Somewhere in your asynchronous course is an assignment whose first draft is now free — a discussion post, a reflection, a case analysis, a problem set — and the work that arrives is fluent, on-topic and tells you nothing about what the student can do. You are not in the room to watch the thinking happen, and the artifact that used to stand in for it no longer does.

That was the version of the problem in which a student decides, at two in the morning, to let a model write the post. The harder version is already here. In three demonstrations on a live undergraduate psychology course, autonomous agents logged into the learning management system, read the course materials, and completed unproctored assessed work with no student involved: two quizzes finished in about 12 minutes and in under 5 minutes, both scoring 10/10, and a discussion post in which the agent mined its peers' posts and then fabricated a credible first-person life history to answer them. The wider record the authors assemble covers at least 15 documented agent runs across Canvas, Moodle and Brightspace using seven agent tools. Their argument is that this is an Assessment Validity problem rather than only an integrity one: agent completion removes the assumption that the submitted work was produced by the person whose learning is being assessed (Hadjisolomou and El-Haddad, 2026).

The bottom line: an asynchronous course has no presence or process visibility to fall back on, so the design has to manufacture them. Move the evidence of learning away from the finished product, keep one measurement the student must produce unassisted, guardrail whatever AI you supply, and pace the term deliberately — self-regulation is what self-paced formats quietly assume and rarely teach. Put plainly: stop designing around the assumption that producing the artifact demonstrates learning. That is not an argument for banning AI or for making tasks AI-proof, but for designing the course so that having AI produce the artifact is insufficient to accomplish the learning.

The short version

  • Make the thinking the deliverable. Staged drafts, decision logs, and self-explanation are harder to outsource than a final artifact, and they show you the reasoning you otherwise cannot see.
  • Keep one unassisted measurement on the assessments that certify competence. Population data show the AI-access effect on retention disappears under proctoring, which tells you the shortcut is off-platform and conditional on the conditions you set.
  • Use asynchronous oral assessments where the stakes justify them — just-in-time prompts, time-limited unrevised recordings, rubrics embedded, transcripts automatic.
  • Guardrail the AI you provide. Hint-not-answer tutoring eliminated the exam penalty that unguarded access produced in a randomized trial; supplying a model without constraints is the version that harms learning.
  • Embed the support inside the course rather than linking out to a generic chatbot. The Open University's embedded assistant doubled time on task in its trial; the same institution's experience with an external, unattached chatbot went the other way.
  • Pace the term on purpose. Early-warning signals exist up to 7–8 days before module deadlines, and self-regulation and well-being decline across a term in step with assessment clustering — so spacing deadlines is a design decision with measurable consequences, not a scheduling detail.
  • Facilitate discussions sparingly, and sequence them deliberately. LLM facilitators are markedly more eager to intervene than human facilitators, and the summary tools that genuinely help students navigate a large forum do not keep participation from declining on their own. Ask for a committed position, a challenge, and a documented reconsideration rather than a post count.
  • Commit before consulting. Have students produce their own position, prediction, outline or decision first, then bring in the model for critique; the thinking that has to be learned should happen before the tool is opened.
  • Pair a vulnerable task with a twin. Keep the take-home analysis and add a closely scheduled second task that assesses the same outcome in a way the first cannot be delegated for.
  • Give the AI a stated role, and be able to finish the sentence: the AI's job here is X, the student's job is Y. If Y is thin, the activity is not ready.

Decide what the assessment is actually measuring

The reason an asynchronous course is more exposed than a face-to-face one is not that students there are less honest. It is that the course's evidence of learning is the submission, and the submission is now cheap to produce. Two pieces of causal evidence show what that costs.

In a randomized trial with roughly 1,000 high-school mathematics students, unguarded AI assistance raised practice performance by 48% while reducing unassisted, closed-book exam scores by 17% — students who never had access outperformed those who did. The critical detail is the fix: a guardrailed tutor that gave hints rather than answers eliminated the harm. The problem was never the model's presence; it was the absence of a constraint on what it would do.

Population-scale behavioral evidence points the same way. Across 3.2 million ALEKS interactions, study time on AI-susceptible problems fell 26.9% after ChatGPT's public release, with a 25% decline in the odds of answering proctored retention items correctly — an effect that vanished once proctoring was in place. That disappearance is the important part for course design: it locates the shortcut off-platform and shows the harm is a property of the conditions you set, not of the learner's character.

There is also a subtler failure to design against. Metacognitively discordant completion names the state of submitting correct, complete work while privately knowing that understanding never arrived. In a classroom you might catch that in a student's hesitation; in an asynchronous course the submission is your only channel, and a correct submission reads as success. Assessment design is therefore the first lever, not the last resort.

Decide what counts as evidence when there is no room to walk into

If the finished product no longer proves the thinking, the evidence has to come from the process or from a condition the student cannot delegate.

Process-revealing artifacts are the cheapest change and work at scale: staged submissions with a required intermediate decision log, a one-paragraph self-explanation attached to each answer, an annotation of a source, a plan revised after feedback. These capture the reasoning rather than its residue, and they cost the instructor comparatively little to scan.

Asynchronous oral assessment is the strongest option when competence must be certified. Pentland, Lowenthal and Krier (2026) evaluated web-based assessments in which prompts are delivered just-in-time, students record brief, time-limited webcam responses they cannot revisit, and instructors grade against embedded rubrics while transcripts generate automatically. Across two studies — an intermediate accounting pilot and a data analytics course — students scored higher on these assessments than on in-person multiple-choice exams (significant in the second study, a positive but non-significant trend in the first), with moderate cross-format correlations supporting convergent validity. Students reported preparing differently, using more active study strategies, and treating the format as professionally relevant. It answers the asynchronous problem directly: the thinking is performed live, at a time and place of the student's choosing, at administrative cost that does not scale with cohort size.

One unassisted measurement belongs in any course whose grade certifies knowledge. The ALEKS result above is the argument: when proctoring was in place, the retention gap disappeared. Close the loop by telling students why the condition exists — the reasoning belongs in your writing a course AI policy.

Decide the sequence, not just the artifact

Once the finished product is unreliable evidence, the order in which the work happens becomes part of the assessment design.

Brcic and Frljic (2026) make placement the center of the argument: allow-or-ban is a false dichotomy, and the design question that matters is where the tool sits. The causal evidence they collect shows the outcome flipping on placement alone — the same unguarded helper that left high-school students about 17% worse on an unaided exam did no harm once rebuilt to withhold answers, while a well-engineered tutor roughly doubled learning. Their six-move frame for placing the tool (prime, probe, point, attach, strengthen, test) is one usable menu, and their one-line diagnostic is the one to keep: if letting AI in makes the task feel effortless, it is in the wrong place. Poorly placed assistance does not merely fail to help; it leaves an illusion of learning that collapses on the unaided task — the state metacognitively discordant completion names for a submission that is correct and not understood.

For an asynchronous course that translates into a sequence you can write into the assignment itself:

Think → Commit → Use AI → Critique → Revise → Explain.

Students attempt the intellectual work and commit to an initial interpretation, prediction, outline or decision. Only then does the model enter, for feedback, alternatives, counterarguments or critique. Then students evaluate what it produced, revise, and explain what changed and why. The commitment step is what makes the rest assessable: an uncommitted student has nothing to revise, and an artifact written from scratch by a model has no revision history at all.

One caution about the overcorrection: collecting every prompt and forcing documentation of every keystroke produces workload without evidence. Ask for meaningful decision points instead — which evidence was selected, which alternative rejected, what changed after critique — and for the reasoning you would actually read.

Pair a vulnerable task with a twin

The other structural option is to stop choosing between a pedagogically valuable task and a defensible one. Roe, Perkins and Giray (2026) propose assessment twins: pair the AI-vulnerable task — a take-home essay or analysis — with a second, less vulnerable task that assesses the same outcomes, scheduled closely enough for cross-verification, and marked interdependently. Their mapping runs across Messick's six strands of validity evidence, and the design process has three steps: identify the vulnerabilities, align outcomes and choose the twin, then develop marking that connects the two. The take-home analysis keeps its value for extended inquiry; its twin might be a short case variation, an explanation of one key decision, or a brief oral or multimedia defense.

The companion question is ownership rather than authorship. Ebrahimzadeh, Shibani and Buckingham Shum argue that coauthorship with generative AI undermines several forms of validity evidence, and propose coauthorship integrity as a source of validity evidence in its own right — violated when a student submits AI-generated content they do not understand. To check ownership at scale rather than by inspection, they report work on an AI Viva: a conversational agent that runs a hybrid viva voce, asking comprehension questions of controllable type and complexity, validated in depth by expert educators and assessment specialists. That is worth weighing against proctoring everything: a spoken defence of one's own reasoning scales in a way that room monitoring does not, and it produces evidence about understanding rather than about who was in the room.

Decide what the AI inside your course is allowed to do

The comparison that matters is not AI versus no AI. It is embedded, constrained AI versus an unattached chatbot, and the two produce different results.

The Open University's AIDA assistant is the best-documented embedded case: six iterative design-based studies over 18 months with 498 students and 20 staff, at an institution serving 200,000-plus learners across more than 50 countries. In an exploratory randomized trial, students using AIDA spent twice as long and visited more pages in their course than the control group, and 96% wanted the assistant available in their formal studies. Purpose-built and in-environment beat generic and external.

Lock, Arteaga and Johnson's (2025) review — 63 citations across 32 countries — adds a social condition to the design condition: students who used ChatGPT alongside teacher tutoring perceived greater Learning Gains than those who used it alone. The assistant supplements the instructor or it replaces the relationship, and only one of those is a course design.

Two cautions are worth taking seriously. KhanMigo's failure was not technical: learners simply did not engage with the chatbot and evidence of gains was limited, which points at organizational readiness and instructional fit rather than model quality. And the instructor side is not automatically ready: in a comparison of South African teacher preparation, self-reported TPACK for AI-integrated science teaching was 64.0% at a campus-based university against 47.4% at a distance university, with pedagogical knowledge the weakest domain in both (Online Teaching and Learning). Deploying an assistant into an asynchronous course does not train the people who must judge its output.

Decide how the asynchronous discussion is actually facilitated

Discussion forums are where asynchronous courses either build community or quietly become submission boxes. Two findings should shape what you automate there.

First, timing is the hard part, not topic detection. Tsirmpas and colleagues built the PEFK corpus to compare facilitation datasets, then ran the first survey of facilitation timing with expert human facilitators and LLM judges. Humans were more cautious about intervening; LLMs were excessively eager. Both were more certain when judging that facilitation was not needed than when judging that it was. Trained classifier models outperformed the LLM setups, and even then the existing datasets capped performance. The practical reading: do not hand autonomous moderation to an agent, and if you use AI in a discussion at all, configure it to stay quiet by default.

Second, AI helps students navigate a forum without making them participate. With 128 university students across three iterations, Hao and Cukurova (2026) found AI-generated discussion summaries broadened exposure to peers' contributions and strengthened network connectedness — the weak ties social-capital theory calls bridging — by lowering the effort of finding meaningful posts. They did not stop viewing activity declining across the course. Summarization is a navigational scaffold, not a substitute for generating participation.

The structure that mitigates this is a sequence rather than a thread: Position → Challenge → Reconsideration. Students commit to an interpretation grounded in the course material, then meet another perspective, counterexample or critique, then explain whether and how their reasoning moved. What you grade is the movement between ideas, not the number of posts — and a submitted post plus two replies no longer demonstrates it, since the third demonstration above has an agent mining a peer's posts and impersonating a classmate's life history to order.

Because what an agent imitates is surface behavior, the instructor's own contributions become the scarce resource. Replying mechanically to dozens of individual posts is the least valuable form of teaching presence; synthesizing across the discussion is the most — three assumptions keep appearing in your analyses; several of you read this evidence differently; this argument is convincing until we introduce this counterexample. That is orchestration rather than message production, and it is the part no agent in this literature performs.

The framework question underneath both is accountability. Ba, Gašević, Lim and Anderson (2026) argue that generative AI unsettles the Community of Inquiry assumption that presence indicators can be attributed to humans at all: they treat GenAI as an epistemic condition, and presences as sociotechnical accomplishments whose relationship to inquiry quality depends on where human accountability sits. The design implication for an asynchronous course is concrete — decide and state who is accountable for what in each exchange, rather than assuming presence arises because a forum exists.

Decide how students keep pace without a room to walk into

Self-paced formats assume self-regulation and rarely teach it. The evidence on which behaviors actually correlate with staying on task is unusually practical.

Surveying 530 college students with association-rule mining and clustering, Shi et al. (2026) found that self-regulated learning behaviors — goal setting, environment structuring and time management — co-occurred most consistently with low digital distraction, along with learner–instructor and learner–content engagement and technical competence. Notably, reliance on peer help-seeking and learner–learner engagement appeared less often in the low-distraction profiles. Structure the environment and the schedule before you design another group activity.

Self-regulation is also not a fixed trait you can assume or dismiss. Across a full semester with 75 first-year students, Song et al. (2026) found SRL functioning as both a stable aptitude and a fluctuating state: individual baselines held steady while metacognitive knowledge and Well-Being declined systemically across the term, driven by curriculum demands such as major assessment deadlines. Clustering deadlines may be an efficient administrative choice, and it is also a measurable tax on the regulation the format depends on.

The intervention window is knowable. Zhang, Jeffries and Koprinska (2025) showed that interpretable machine learning on content-interaction logs predicts module-level progress and flags dropout outcomes up to 7–8 days before module deadlines in large-scale online programming courses. That is enough notice to send a specific nudge to a specific student, and it beats discovering the failure at grading time. For adult and distance learners, the AI-ALOE design guidelines add the constraint that matters most: mobile access, offline capability, and genuine asynchronous availability, since these learners study in fragments of time between other obligations.

What AI can take off your plate as the designer

Two uses have reasonable evidence behind them, and both target instructor workload rather than learner thinking.

Production cost. MAIC reframes the MOOC's "one video for N students" broadcast as "N agents for one student," using specialized teacher, assistant, classmate and analyzer agents on a shared model foundation. Its authors report collapsing course production from roughly $25,000 and 60 hours per MOOC to under $2 and 30 minutes, piloted at Tsinghua across two courses with 100,000+ learning records from 500+ students, and released as open-source OpenMAIC. Personalized media is the same idea at the asset level: in a large online course, Tomlinson et al. (2026) found students preferred AI-generated personalized videos over non-personalized human-recorded ones, with the personalization effect outweighing the value placed on a human presenter.

Forum navigation. The summary scaffold above reduced the cost of locating good contributions without burdening instructors with the summarizing.

What not to automate. Deciding whether a piece of writing is the student's own thinking, judging whether a discussion needs intervention, and calibrating difficulty are the tasks the evidence says are either unreliable or accountability-bearing. Generated material also needs pedagogical review before it carries credit — cheap production is not the same as sound design.

Give the AI a stated role, then check the student still has one

"Students may use AI" is too broad a specification to design against. Naming the role is what makes an activity designable, and the roles carry consequences for presence: a study of generative AI in marketing education distinguishes tutor, teammate and tool, shows each influencing teaching, social and cognitive presence differently, and lists the familiar ethical exposures — data privacy, plagiarism, dependency and assessment fairness (GenAI in Marketing Education). The same model can be a Socratic questioner before an assignment, a critic after a first draft, a simulated stakeholder during a case analysis, an opponent whose argument must be rebutted, a hint-giving tutor during practice, or an editor brought in only after the substantive reasoning exists. Those are different activities, and conflating them is how "AI is allowed" quietly comes to mean "nothing was designed".

Scaffolding supplies the test: support should help learners do what they cannot yet do alone, and should fade as competence develops. A model that keeps supplying complete solutions is not a scaffold — it is doing the task. So for every AI-enabled activity, the designer should be able to complete this sentence without hesitation:

The AI's job here is ______, while the student's job is ________ .

If the second blank contains little meaningful thinking, the activity needs redesigning rather than a stricter policy.

Teach evaluation, not just operation

AI Literacy is not the ability to operate a chatbot; it includes critical evaluation, metacognitive judgment and knowing when not to use the tool at all — and the knowledge base is consistent that self-reported confidence is a poor proxy for demonstrated competence. That points at a change any asynchronous course can adopt immediately: sometimes give students the AI output yourself. Hand over two competing answers and ask which is stronger and why. Ask for the weakest claim, the missing assumption, the unverifiable source, the plausible explanation that is nonetheless wrong. Ask what evidence would change their mind.

Framed that way, the capability being graded is no longer only can this student produce an answer? but can this student recognize whether an answer deserves to be trusted? — closer to what the discipline actually requires, and far harder to delegate. The verification practices and the AI literacy material are where the technique lives; what matters here is putting evaluation inside the graded task instead of leaving it as advice.

The objections you will hear

"These are adults who chose an asynchronous course. If they let AI write it, that is their decision." The choice argument would hold if the course certified nothing. It does. The ALEKS evidence shows the shortcut produces an appearance of competence that does not survive a proctored check, and the cost lands on the student later — in the next course, the licensure exam, or the job. There is an equity edge too: the learners most likely to offload are often those with the least time, which is precisely the group an asynchronous course exists to serve.

"Detection is the answer." Detection is contested, and the proctoring result above shows why design beats policing: the harm disappeared when the conditions changed, not when policing intensified. Detection also carries false-positive costs and turns instruction into an arms race. Redesign the task and keep one unassisted measurement instead — see How Can I Reduce AI Cheating in My Course? and How Do I Redesign Assessment So That a Grade Still Tells Me Something Defensible About What the Student Knows or Can Do?.

"Teaching presence is impossible at a distance, so async is inherently inferior." Presence in an asynchronous course is designed rather than implied, and the CoI reconceptualization above says its indicators are sociotechnical accomplishments in the GenAI era. The AIDA trial is a useful counter-example: embedded generative support doubled time on task in a course with no synchronous meetings at all.

"Proctoring is surveillance and I will not impose it." A defensible position, and it does not leave you without options. Process-revealing artifacts, asynchronous oral assessment, self-explanation and the AI Viva described above all capture reasoning without monitoring anyone's room. If you do use a proctored condition, say so in the syllabus, explain the reasoning, and keep it to the assessments that certify competence.

"Then make it all synchronous and proctored." Many students choose asynchronous study precisely because they have employment, caregiving responsibilities, disabilities or geographic constraints that make synchronous attendance difficult, and detection carries its own equity costs: tools that flag non-native writers disproportionately produce false positives that penalize honest work (AI Detection). The proportionate alternative is a small number of verification moments inside an otherwise flexible course: a short recorded explanation, a personalized application, a response to an instructor-selected question, an annotated decision trail, a low-stakes individual check. Verification should raise the validity of your evidence without removing the flexibility that made the format worth offering.

What the evidence does not settle

The systematic review of AI in distance education (56 articles, 2020–2025) covers personalization (24 studies), assessment and feedback (19), human–AI interaction (17) and governance and equity, and concludes that the base is short-term and cross-sectional, with little longitudinal work. The review of AI and student engagement in online learning (24 studies) is limited to one database, treats engagement only, and explicitly conflates synchronous with asynchronous contexts — so its conclusions should not be read as asynchronous-specific. Asynchronous oral assessment rests on two studies in two courses. MAIC is a pilot. Facilitation-timing datasets cap performance even with trained classifiers. And no study here follows an asynchronous cohort long enough to show whether redesigned assessment produces durable learning rather than better evidence of it. Treat all of it as strong enough to change your next course and too thin to justify a policy claim.

Do this week

In ten minutes: pick your highest-stakes assessment and add one condition the student completes unassisted. Tell them why it exists.

Before the next assignment goes out: take the assignment AI now completes end-to-end and rewrite it so the process is the artifact — a decision log, a staged draft, a self-explanation, a comparison of two attempts. If the task can be finished without any of that reasoning, the reasoning was never required.

This term: convert one major assessment to an asynchronous oral defense with embedded rubrics; set the AI permission per task in the syllabus and place it where students will actually read it; space the major deadlines instead of clustering them; and open a 7–8-day pre-deadline window in which you contact students whose interaction data has gone quiet.

Run three questions over your weakest assignment. (1) Could an AI system complete this activity without the student understanding the material? (2) What cognitive activity is supposed to produce the learning here? (3) What evidence will show that the student actually performed that activity? If the first answer is yes and the other two are hard to answer, the problem is the learning design rather than the AI policy — and no wording of the policy will fix it.

The goal is not a course that AI cannot participate in. It is a course where AI can participate without displacing the learning the course exists to produce.

Where to go next: the course-level rules belong in How Do I Write a Course AI Policy and Communicate It to Students?; the assessment rebuild is covered in How Do I Redesign Assessment So That a Grade Still Tells Me Something Defensible About What the Student Knows or Can Do?; the reliance problem underneath it in How Do I Keep Students from Over-Relying on AI?; instructor workloads and what AI realistically removes in How Can AI Save Me Time as an Instructor?; and the design-level view of embedding AI into a learning experience in How Should AI Be Designed Into the Learning Experience?.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.