On this page

You have warned them. The chatbot is not a search engine, check the sources, use it responsibly — and the same fabricated citations and confident wrong answers keep arriving in the submissions. The warnings fail because they ask for a disposition. Careful students have it, hurried ones do not, and there is nothing in the sentence a student can act on at 11 p.m. with an output on screen that reads perfectly well.

The research gathered here says that framing is too coarse to teach or grade. Checking is a set of analytically separate actions, and the broad label "responsible use" hides the difference between a student who attempted a check and one who settled whether the output was correct. The bottom line: verification becomes teachable the moment you stop treating it as a trait and start treating it as a sequence — a standard fixed before the check, a check that produces a record, and a decision the student must defend on domain grounds. That is a design problem you can solve this week — and a grading problem you can solve with artifacts the assignment already generates.

This page addresses the mechanics of verification itself. For what students should understand about how these systems work, and how that understanding is defined, sequenced and assessed, see How Should I Incorporate AI Literacy into My Course? and What Is the Evidence on AI Literacy Interventions in Higher Education?. A student can know a great deal about large language models and verify nothing.

What verification actually requires of a student

Wei and Shang (2026) separate seven targets that "critical AI use" usually collapses into one: epistemic evaluation, verification initiation, process quality, verification success, reliance decisions, immediate task performance, and independent learning. Your syllabus rubric says "checked sources," which sits at step two and says nothing about whether the check was competent (process quality) or whether it changed anything (success, then the reliance decision). Initiation is not success: a strong process can end inconclusive, a weak one can land on the right answer, and a student who ran a search and stayed confused scores the same as one who resolved the question.

The second distinction is between Trust and reliance. The same review defines reliance calibration as a judgment about whether a reliance decision was appropriate given the actual quality of the AI output — an output-contingent classification, not a stage on a timeline and not a score on a trust scale. One audited study found that false ChatGPT information shifted participants' reported trust, but never recorded whether they later accepted or rejected a specific recommendation. Another logged accept and reject decisions on ChatGPT Feedback for 78 translation students, but without independent expert evaluation an acceptance cannot be called appropriate, nor a rejection justified. Without ground truth about output quality, reliance cannot be classified at all — which tells you what every verification task needs: a defensible answer key.

That is why Jaidka and Cai (2026) treat miscalibration as a design problem rather than a student deficit. Their typology crosses ability to verify with motivation to verify, and the profiles need different remedies: the high-ability, low-motivation user is prone to complacent overtrust, while the low-ability, low-motivation user is the most exposed — so two students failing the same check may need opposite remedies.

How to teach it

Wei and Shang's synthesis drew on 493 deduplicated records from Web of Science Core searches and mapped 14 priority empirical studies onto verification, reliance and outcome columns. Its constructive output is the sequence to build tasks around. Adjudicate output quality first, capture whether verification was initiated, code process quality, score whether it succeeded, record the accept–revise–reject decision, and evaluate that decision against the adjudicated quality. One warning belongs here too: interventions that reduce inappropriate acceptance must also be checked for the unintended rejection of correct assistance, so leave room for a student to accept good output for a stated reason.

Predict before the output, then compare. Gousopoulos (2026) builds a measurement program for construction tasks around predict before you run: the learner states what should happen, then the artifact's behavior is judged against that prediction on domain grounds rather than by whether it runs. Its audit of 24 studies found the same tool producing opposite outcomes under different task structures: a conventional ChatGPT setup ended significantly lower in achievement, Self-Efficacy and flow, while a condition adding verification requirements and error-reflection modules showed stronger higher-order thinking. Because the prediction precedes the output, the reasoning is on record — and the record is what you grade.

Require source checks against the record, with a named target. Denny et al. (2026) traced 113,588 references from 5,225 computing education papers published since 2021 and manually verified 828 suspicious records, finding 30 containing verifiably fabricated bibliographic information across 14 papers, all from 2025 and 2026. Thirteen were entirely fabricated; the other 17 combined a real title with fabricated or incorrect authorship, venue or year. At the SIGCSE Technical Symposium the count rose from 3 in the 2025 proceedings to 17 in 2026, or 2.3% of 2026 proceedings papers. Author fields fail more often than any other part of a generated reference, so "open the source and read the author list" targets the field most likely to be wrong. Do not ask students to "make sure sources are real"; ask them to confirm authors, venue and year against the record itself. The paper asks that every cited work be verified and that any checker stay human-in-the-loop — the practice a citation-verification requirement rehearses.

Teach the discrimination, not only the caution. Gousopoulos formalizes verification as a signal-detection problem with two independent parameters: sensitivity, the ability to discriminate correct from flawed output, and criterion, where the learner sets the threshold for rejection. Over-reliance splits accordingly — warnings, checklists and hallucination prompts shift the criterion, changing when a student rejects, while domain instruction, worked comparisons and seeded-error practice raise sensitivity, because the learner can then tell. The prerequisite follows: verification instruction is educative only where sensitivity can exceed zero, which requires enough domain knowledge to distinguish correct from incorrect output. In the same audit, novices asked to explain Large Language Models (LLMs)-generated code succeeded on roughly a third of tasks, which is why judging AI feedback or code against one's own reasoning is a real check only where that reasoning has substance. If your students cannot yet do the domain thinking unaided, more warnings buy you nothing; teach the content first and attach the check to it.

Forewarn, immediately before the task. Vu, Cummings and Park (2026) showed a generic inoculation (forewarning) message immediately before two tasks to 100 US-based students, 40 domestic and 60 international EFL. Inoculated students were significantly more likely to verify the academic-source-summary task (M = 0.34 versus 0.18), while self-reported verification intentions did not move. Two lessons: a short warning at the point of use changes enacted behavior, and the effect was task-dependent — appearing for the source-summary task but not uniformly for a mathematics quiz on exponentiation and large-number multiplication.

Make correction cost something. In an error-correction paradigm the review audited, effort during correction mattered for learning; simple answer substitution is unlikely to deliver the same benefit. Requiring a student to reproduce a step, rewrite a passage, or state the domain reason an output is wrong converts a check into work. Venetsanos (2026) sets the bar for what may be checked mechanically: documented criteria, comparison against established knowledge without interpretive judgment, and a single correct answer or pre-specified acceptable alternatives. Anything interpretive fails that bar, so separate the mechanical layer — dates, formulas, citations, calculations — from the judgment layer, where the student's own reading is the instrument of the check.

How to assess it without policing students

Gousopoulos draws the consequence directly: if the AI can produce the artifact, the artifact cannot be the assessment, so evaluation relocates to the specification, the validation reasoning and the interpretation. Its model authorship construct has four facets — specification, conceptual model, verification, interpretation — at four ordered levels from delegated to authored, where the authored level requires a verifiable specification preceding the first prompt, rejection of output on domain grounds with a stated reason, and interpretation beyond the artifact's own report of itself. Because the rubric is scored from materials a construction task already generates, the same instrument serves as formative assessment and as a research measure. That is the answer to the busywork worry: you are not adding an assignment, you are scoring a byproduct.

Three decisions keep this from becoming surveillance. First, grade the check against adjudicated quality, not effort — otherwise you build an incentive to perform checking theater. Second, state provenance, as Venetsanos requires: students must understand the provenance, nature and limitations of the feedback they receive — which parts were machine-verified, which evaluatively judged, and that human judgment has primacy — because students cannot weigh feedback they cannot situate. Third, keep enforcement off detection tools. Bassett et al. (2026) argue that AI detection should not be used in education at all: its estimates are probabilistic and cannot be independently verified because real-world text origin is unknown; its scores do not meet the balance-of-probabilities standard integrity investigations require; and the human-versus-AI dichotomy is meaningless for work created with, rather than by, AI. Their conclusion is that detection "does not safeguard academic integrity; it undermines it" — surveillance regimes foster suspicion and erode student Trust. Keep the two senses of detection apart: detecting AI-generated text has no defensible evidentiary role here, while teaching students to detect errors in AI output is the point of the exercise.

Separation is architectural, and worth telling students about: Li, Zhang and Botelho (2026) check output with a second model rather than embedding the check in the generator — and a student who sees why can see why "the tool said it was right" is not a verification argument.

Why students skip the check

Calibration. Jaidka and Cai's typology puts the high-ability, low-motivation student at risk of complacent overtrust while the low-ability, low-motivation student is the most exposed. The audited literature shows the same split in miniature: a design that supplied advice correct about half the time, where the weight students gave it varied with prior knowledge. Calibration failures are not fixed by exhortation: the student's own confidence signal is what is miscalibrated.

Fluency. Output that reads well is treated as right. Novices asked to explain Large Language Models (LLMs)-generated code succeeded on roughly a third of tasks, and a field study of student–ChatGPT quiz conversations found that following correct guidance still produced a wrong answer. When sensitivity is near zero, skipping the check costs the student nothing they can perceive — the failure is invisible until it is graded.

Social proof and intention. Students report that they intend to verify, and the report is worthless: in Vu, Cummings and Park's study, self-reported verification intentions did not move while enacted behavior did (M = 0.34 versus 0.18). Do not grade the intention. Norm-based appeals are not a reliable fallback either — message-based norm nudges showed no significant effect in a direct tournament Jaidka and Cai report. Time pressure is a documented reason students skip checks, and a syllabus that leaves verification unscheduled says it is optional.

Three objections

"They can just do it for the grade." Partly true — so grade materials the AI cannot produce on a student's behalf. A specification written before the first prompt, a prediction recorded before running the artifact, a source list checked against the record, a stated domain reason for rejection: each of these requires the student to hold a position the generator cannot supply. At the authored level, Gousopoulos requires rejection of output on domain grounds with a stated reason — a fabricated reason is as visible as a fabricated citation. The residual risk is bounded by the sensitivity threshold: where a student has no domain knowledge to check against, verification is theater, and no rubric rescues it. That is an argument for teaching the domain first, not for abandoning the check.

"There is no time in the syllabus." Forewarning is one message immediately before the task; the seeded-error task is twenty items, eight with a domain-level error, and it tells you whether students can discriminate at all; and the rubric rides on materials the task already generates, so it costs marking time on evidence you would otherwise lack rather than a new assignment. What costs time is the alternative: unverified submissions, fabricated references, and the integrity conversations that follow. The constraint is still genuine — time pressure is a documented reason students skip checks, so verification that is not scheduled into the task will not happen.

"The tool is usually right." Then the check is cheap and the exceptions are the point. The record on references is not reassuring: 30 verifiably fabricated items across 14 papers, all from 2025 and 2026, 17 of them pairing a real title with fabricated or incorrect authorship, venue or year, 13 entirely fabricated, and a jump from 3 in the 2025 SIGCSE proceedings to 17 in 2026, or 2.3% of 2026 proceedings papers. Accuracy on the parts you happen to notice says nothing about the parts you do not — author fields fail more often than any other part of a generated reference. And an accepted answer can be wrong even when the guidance was right: the field study of student–ChatGPT quiz conversations recorded exactly that. The reason to teach verification is not that the tool is usually wrong, but that the student cannot tell which case they are in — precisely the skill your course exists to build.

What remains unknown

  • Whether any intervention improves verification success and the reliance decision that follows, judged against adjudicated output quality: no study among Wei and Shang's 14 priority cases measured both, and few followed immediate performance with delayed retention or transfer.
  • Whether the components relate in the order the map implies. The seven targets are an analytic ordering, not a validated causal model.
  • Whether design propositions work. Jaidka and Cai's eight propositions are untested predictions, and message-based norm nudges showed no significant effect in a direct tournament they report.
  • Whether checking built into feedback develops self-verification habits or dependency on external validation, which Venetsanos raises and leaves open.
  • Whether verification instruction works below a domain-knowledge threshold. Gousopoulos predicts it cannot, and notes that some domains furnish an external criterion — physics, chemistry, ecology, epidemiology — while history, literature and Ethics largely do not.

What to do this week

  • Fix the standard before the check. Decide what adjudicated correctness means for the task so a check has something to resolve against.
  • Ask for the prediction first. Have students commit in writing to what should happen, then compare the output on domain grounds, not fluency.
  • Require citation verification with a named target. Author lists are the least reliable part of a generated reference; have students open the record and confirm authors, venue and year.
  • Diagnose sensitivity and criterion separately. A student who accepts flawed output because they cannot tell needs domain practice; one who can tell and accepts anyway needs the threshold moved.
  • Grade the checking, not only the artifact. Score specifications, prediction records, validation logs, source checks and stated reasons for rejection.
  • State provenance. Tell students which feedback was machine-verified and which was judged by a person.
  • Keep enforcement off detectors, and schedule verification. Time pressure is a documented reason students skip checks.
  • Run the seeded-error task. Twenty items, eight with a domain-level error, tells you whether students can discriminate at all.
  • Say the warning at the point of use. Forewarn per task type, not once per syllabus.

For the surrounding work, see how to incorporate AI literacy into a course and what the AI literacy evidence shows — the two pages that build the understanding this one puts to work.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.