On this page

Synthesis: Venetsanos (2026) proposes a tripartite feedback framework for AI-assisted assessment of complex written reports — laboratory, research, design and case reports in STEM and social science — that sorts feedback into three analytically distinct levels by epistemic status rather than by cognitive hierarchy: low-level structural and presentational feedback (formatting, referencing, grammar, sentence-level clarity), which is rule-based and technologically mature; intermediate-level factual and procedural verification (whether a calculation used a specified method correctly, whether a stated fact matches authoritative references), which compares a claim against established knowledge rather than judging reasoning; and high-level critical evaluation and synthesis, which is inherently interpretive and requires disciplinary and pedagogical judgment. The framework's contribution is not a working system but an allocation rule: the question is not what technology can technically perform but which feedback tasks carry the epistemic status that makes automation defensible. Its most contested space is the intermediate level, where AI support is conditional on five non-negotiable boundary principles — assessor-curated retrieval-grounded knowledge bases, human oversight and final authority, strict limitation to bounded verification, transparency and attribution, and security by design — that must hold simultaneously. The paper explicitly presents no empirical validation, and supplies the mixed-methods evaluation design that would be required to test it.

Key Findings

  1. Three levels defined by epistemic status, not by difficulty. The taxonomy separates rule-based checking, factual verification, and interpretive evaluation, and the author argues these are distinct dimensions rather than a ladder: a calculation can be correctly presented (low-level) while arithmetically wrong (intermediate) and methodologically inappropriate (high-level). This is deliberately different from Hattie and Timperley's (2007) framework, which explains where feedback operates in learning (task, process, self-regulation, self) but is silent on whether a task involves rule-checking or judgment — the distinction that actually determines automation suitability. The two frameworks are complementary: the same content can attract intermediate verification ("is the arithmetic right?"), high-level process feedback ("was this the right approach?"), and self-regulation feedback ("how would you check this next time?").
  2. The clarity caveat is the sharpest boundary case. Sentence-level clarity (grammar, syntax, word choice) is genuinely surface-level and safely automatable; argumentative clarity is not. Unclear argumentation frequently signals incomplete conceptual grasp rather than a presentational lapse, so the author specifies that flagged instances of argumentative unclarity should trigger human review rather than automated correction, since correcting the prose would leave the underlying reasoning problem untouched.
  3. The intermediate level is defined by verification through comparison. Unlike low-level feedback it engages substantive content; unlike high-level feedback it verifies bounded claims against established knowledge rather than evaluating the quality of reasoning. A chemistry laboratory report illustrates the boundary: a missing figure caption is low-level, a calculated yield checked against the correct formula is intermediate, and whether the student chose an appropriate method for the reaction conditions — and justified it — is high-level. The same visible task (method selection) can look like fact-checking on the surface while requiring disciplinary reasoning underneath, which is why the framework requires all four bounded-verification criteria rather than surface resemblance.
  4. Principle 1 — assessor-curated knowledge bases. Whatever the AI checks against must be explicitly defined, bounded, and curated exclusively by qualified human assessors (assignment briefs, rubrics, laboratory manuals, worked examples, licensed textbooks, lecture content), and the system must not autonomously expand that base, reach external sources without validation, or rely on pre-trained parametric knowledge. Grounded in the documented hallucination problem, the rationale is that retrieval-augmented generation bounds the domain in which errors can occur; violation converts bounded verification into unconstrained assertion, whose practical consequence is confident, authoritative-sounding feedback that is factually wrong with no mechanism to catch it before it reaches students.
  5. Principle 2 — human oversight and final authority. Assessors review all AI outputs before students see them, hold absolute discretion to override, modify or reject, and remain fully accountable for all feedback students receive; any conflict between AI output and human judgment resolves to the human without exception. The grounds are both empirical (Boud and Dawson's demonstration that effective feedback depends on professional judgment about student development, context and disciplinary norms) and normative — accountability for decisions affecting students' standing is a core element of higher education's ethical purposes, not merely a safeguard. Violation creates an accountability vacuum with no meaningful recourse for a student disadvantaged by incorrect AI feedback.
  6. Principle 3 — bounded verification only, with conservative escalation. AI involvement is limited to tasks satisfying all four criteria simultaneously: unambiguous and explicitly documented assessment criteria; evaluation by direct comparison against established knowledge without interpretive judgment; no assessment of alternative valid approaches; and a single correct answer or a well-defined pre-specified set of acceptable alternatives. Tasks failing any criterion fall outside appropriate AI involvement "regardless of how routine or straightforward they may appear," and ambiguous cases escalate to human assessment by default.
  7. Principle 4 — transparency and attribution. Students must understand the provenance, nature and limitations of every piece of feedback, including what AI was involved, for which tasks, under what conditions, and the primacy of human judgment; criteria must distinguish which dimensions of the work are verified and which are evaluatively judged. The argument draws on feedback literacy — students cannot weigh feedback they cannot situate — and on institutional honesty: undifferentiated feedback of mixed provenance invites misplaced confidence, while opacity about AI involvement undermines the relational trust assessment depends on.
  8. Principle 5 — security and integrity by design, assuming adversarial use. Input sanitisation must detect embedded instructions hidden in white or small text, encoded strings (hexadecimal, base64, Unicode), text inside images, and content in document metadata or comments; systems need suspicious-pattern monitoring, enforced integrity policies addressing AI manipulation, and independent security audits. The stated rationale is empirical: within large cohorts a meaningful minority will treat automated assessment as a technical challenge to exploit, and successful circumvention spreads rapidly through student networks. If a system cannot be made secure against adversarial use, the efficiency and consistency gains are negated by integrity costs — with the burden falling disproportionately on students who engage honestly.
  9. A four-stage implementation pathway with explicit human-in-the-loop requirements. Stage 1 deploys mature commercial tools for structural and presentational feedback; Stage 2 develops AI-supported intermediate checking framed as flagging potential issues for expert review rather than producing feedback, starting with high-frequency human oversight of all outputs and scaling back only as evidence demonstrates acceptable reliability and security; Stage 3 has staff generate high-level feedback themselves, deciding whether time freed earlier is reinvested in depth or elsewhere; Stage 4 integrates all three into student-facing delivery with clear attribution. The operational procedure (Report Assessment & Feedback) defines the fact-check points and rubric first, has the system produce low- and intermediate-level reports, has the assessor modify them, add high-level feedback, then rank all comments into a Final Feedback Report that the assessor approves — with assessment of both the frequency and the type of human action retained as evidence.
  10. Human-in-the-loop frequency is a design decision with a default. High-frequency oversight buys quality control, rapid error detection, accountability and calibration but can erode the efficiency gains that motivated AI use and bottleneck peak periods; low-frequency spot-checking scales and speeds turnaround but risks undetected error propagation across submissions, weaker accountability, and inequity if some students receive more thorough review than others. Rather than prescribing a universal answer, the framework requires the trade-off to be made deliberately against disciplinary norms, assessment stakes, cohort size and institutional resources, and sets a default: begin with high-frequency oversight and reduce it only on substantial evidence of acceptable reliability, security and fairness — the burden of proof resting on demonstrating that reduced oversight is safe.
  11. The pedagogical questions are left open rather than answered. The paper asks whether automated correction of formatting and grammar builds professional communication skill or merely produces correctly formatted documents; whether automated factual checking develops self-verification habits or creates dependency on external validation that undermines autonomy; whether students engage with why an error occurred or simply fix flagged items; and whether multi-source feedback (human, AI verification, automated checking) supports evaluative judgment or multiplies information in a way that reinforces transmission models of feedback. Its equity mandate is similarly a requirement rather than a result: training-data bias, differential treatment of non-standard expression styles, Accessibility for disabled students, and unequal institutional access to AI-assisted feedback are named as live risks to be audited with disaggregated outcome monitoring.
  12. An evaluation design is specified as the price of adoption. Quantitative evidence would cover efficiency (turnaround time, staff time per submission, total workload against control modules), consistency (inter-marker reliability for high-level feedback, low- and intermediate-level consistency, AI error frequency), quality (student satisfaction, independently rated specificity and actionability) and learning (subsequent performance, validated feedback-literacy and self-assessment instruments, learning gains versus control). Qualitative evidence would cover student focus groups on clarity, usefulness and trustworthiness, staff interviews on workload, judgment processes and professional identity, process documentation, and systematic analysis of AI errors, security incidents and equity impacts. Matched control modules are required; randomized assignment gives the strongest inference but may be infeasible, so the paper accepts pre–post comparison across years or mixed-cohort within-subject designs, with a staged approach beginning from low-cost proxies such as turnaround time, flagged error rates and human–AI agreement in spot checks.

Relationship to Existing Frameworks

The paper positions itself against three substantive gaps. Hattie and Timperley (2007) and Boud and Molloy (2013) are, in its reading, powerful on what feedback should accomplish and how to design it for agency, but were written before capable AI and are silent on which tasks technology may appropriately support and which must stay human. Boud and Dawson (2023) and Carless and Boud (2018) establish that feedback provision needs professional judgment and that students must be able to interpret feedback, yet neither offers automation criteria. Existing AI-in-assessment discussions, meanwhile, remain at the level of general principle — AI should support rather than replace judgment, feedback should be transparent — without stating what conditions must be met before AI involvement is justifiable or what safeguards are non-negotiable, which leaves institutions without criteria or accountability standards. The framework's own comparison table scores prior frameworks on six dimensions (analytical purpose, epistemic status of tasks, human–AI allocation criteria, explicit engagement with AI, hybrid accountability, feedback literacy in hybrid contexts) and finds none addressing epistemic status or allocation — the two dimensions on which automation decisions turn. The stated contribution is therefore structural rather than prescriptive: categories, boundary conditions, and validation criteria for context-sensitive institutional decision-making.

Limitations

The paper is deliberately conceptual and reports no empirical validation and no implemented pilot, so which Large Language Models (LLMs) is used, how it is prompted, and how the RAG pipeline is configured remain open — and different models and configurations may satisfy the five boundary principles to varying degrees, particularly knowledge-base grounding and bounded verification. The taxonomy's boundaries may be harder to operationalize than stated: in reports with open-ended design, contested methodological choices or multiply interpretable data, tasks that look like straightforward fact-checking can contain embedded interpretive dimensions, and the conservative escalation default addresses this procedurally without resolving it. The five principles have not been tested for simultaneous feasibility — assessor time for knowledge-base curation, infrastructure for robust security, and the workflow demands of high-frequency oversight could make intermediate-level AI support impractical in most institutions, so the conditions may be so demanding that the framework limits its own operational value. It was developed for report genres in STEM and social science and should not be assumed to transfer to essays, portfolios, dissertations or reflective writing. Its central pedagogical assumption — that differentiating feedback by level and source supports rather than undermines feedback literacy — is untested, and the author concedes it is equally plausible that multi-source feedback confuses students, weakens trust, or multiplies delivered information without building the capacity to seek and use feedback. Security may constrain viability more than capability if the adversarial landscape evolves faster than institutional response. Equity problems are named but cannot be pre-specified, and may be systematically differential in ways invisible without sustained disaggregated monitoring. The paper also acknowledges that the principles may shift staff effort rather than reduce it, leaving net efficiency gains an open question, and that a bounded framework could still ease future expansion of AI's role — which is why the principles are framed as a standing check rather than a one-off safeguard.

Connected Concepts

Connected Articles

Citation

Venetsanos, D. T. (2026). A Tripartite Feedback Framework for AI-Assisted Assessment of Complex Reports in Higher Education. AI in Education, 2(3), 26.

Connected FAQs

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.