On this page

Synthesis: This pilot study examines how higher-education students perceive GenAI-driven adaptive assessment, Critical Thinking, Cognitive Offloading and Academic Integrity. Combining secondary literature with a Google Form survey of 13 students collected on 14 September 2026 through convenience sampling, it tests seven perception hypotheses against the neutral midpoint of a five-point scale using Wilcoxon signed-rank and sign tests. The strongest signal is offloading: 61.5% agreed that excessive reliance on real-time AI causes cognitive offloading (mean 4.00, p = .005), and 61.5% rated offloading in unmonitored work as frequent (mean 3.77, p = .010). Students endorsed interactive hints for breaking down multi-step problems (69.2%, p = .011) but were less sure that Scaffolding encourages verification of AI outputs (p = .106). They were undecided on whether process-based assessment reduces ghostwriting and divided on institutional guidance (p = .958). The authors recommend hints over answers, graded verification of AI outputs, gradual process evidence and clear AI-use rules, while cautioning that a non-random, self-reported sample of 13 makes the findings indicative, not generalizable.

Key Findings

  1. Interactive AI hints aid problem decomposition. Nine of 13 respondents (69.2%) agreed that interactive AI hints improve students' ability to break down multi-step problems independently, supporting H2 (p = .011).
  2. Cognitive offloading is the survey's strongest signal. Eight of 13 (61.5%) agreed that excessive reliance on real-time AI causes cognitive offloading and reduces critical-thinking capacity, with no disagreement (mean 4.00, p = .005).
  3. Unmonitored work is seen as offloading-prone. Eight respondents (61.5%) rated the frequency of offloading in unmonitored assignments at 4 or 5, with a mean of 3.77 and H4 supported (p = .010).
  4. Verification claims are weaker. Seven respondents agreed that AI scaffolding encourages verification of AI output, but four were neutral and two disagreed; H1 was not supported (p = .106).
  5. Process-based assessment remained undecided. Seven respondents were neutral on whether tracking draft history and thought progression would reduce ghostwriting more than take-home exams; H5 was not supported (p = .119).
  6. Institutional guidance was divided. Six respondents agreed their institution provides clear, enforceable guidelines on GenAI in assessment while four disagreed; the median response was neutral and H7 was not supported (p = .958).
  7. Students believed they could tell ethical use from misconduct. Seven respondents (53.8%) agreed that students can distinguish ethical AI assistance from academic misconduct, with none disagreeing (p = .008), though the authors stress this is self-perception.

Context and study design

The paper sits between two literatures. On one side, controlled studies show that design determines whether AI helps: Kestin et al. (2025) found that a purpose-built AI tutor let college physics students learn more in less time than an Active Learning comparison, while Bastani et al. (2025) found that unrestricted GPT-4 access improved practice performance but reduced later examination performance once access was removed—a harm largely avoided by teacher-created hints rather than direct answers. On the other, correlational work links heavy AI use to reduced cognitive effort: Gerlich (2025) associated higher AI-tool use with cognitive offloading and lower critical-thinking scores. The present study is a cross-sectional, descriptive-exploratory pilot survey with a small qualitative component, positioned as the first stage of a broader mixed-methods design. A Google Form questionnaire reached respondents through the researcher's peer network via non-probability convenience sampling, yielding 13 complete responses: 12 undergraduates and 1 postgraduate, with no teaching staff participating. The instrument had 15 items—five profile and AI-use items (Q1–Q5), nine closed items (Q6–Q14) and one open-ended item (Q15)—scored on a five-point Likert scale, and analysis combined frequencies, non-parametric tests against the midpoint of 3 and thematic coding. The authors frame their gap against Intelligent Tutoring experiments, offloading surveys of adults and conceptual integrity-detection scholarship, arguing that student views on scaffolding, offloading, integrity and institutional guidance are rarely studied together in one instrument.

Perceived benefits and the offloading result

Respondents were predominantly undergraduates and regular GenAI users: eight reported daily academic use, and more than half belonged to Law and Policy Studies. Familiarity with adaptive assessment tools was moderate (mean 3.62 on a five-point scale), and immediate explanations of incorrect answers were the most commonly observed feature; only two respondents identified step-by-step guidance that withholds the final answer, the design the authors argue matters most for reducing dependence. The clearest benefit concerned problem structuring—9 of 13 respondents (69.2%) agreed that interactive AI hints improve the ability to break down multi-step problems independently (p = .011)—which supports the study's second hypothesis and aligns with the promise of Personalized Learning through Formative Assessment. The strongest signal overall was risk. Eight of 13 (61.5%) agreed that excessive reliance on real-time AI causes cognitive offloading and reduces critical-thinking capacity (mean 4.00, p = .005, no disagreement), passing the conservative Bonferroni threshold of .0071, and the same proportion rated offloading in unmonitored assignments as frequent (mean 3.77, p = .010). Verification was weaker: 7 respondents agreed that AI scaffolding encourages verification of AI output, but 4 were neutral and 2 disagreed (p = .106), a result the authors read as students being less confident that such tools develop Evaluative Judgment. They note that the survey measures perceived benefit, not actual critical-thinking performance, and that Feedback design—hints versus finished answers—is the deciding factor.

Integrity, institutional readiness and open responses

On integrity, 7 respondents (53.8%) agreed that students can distinguish ethical AI assistance from academic misconduct, with none disagreeing (p = .008)—a result the authors treat cautiously because it reflects self-perception rather than tested knowledge or conduct. The sample was drawn largely from Law and Policy Studies, so Legal Education scenarios framed much of the discussion: a prompt asking a learner to identify legal issues, distinguish facts and apply a rule may support analysis, but the student must still verify the statute, case authority and reasoning. Respondents were undecided about process-based assessment: 7 were neutral on whether tracking draft history and thought progression would reduce ghostwriting more effectively than conventional take-home examinations (p = .119), which the authors suggest may reflect limited experience with Process-Oriented Assessment rather than rejection of it. Institutional guidance was divided—6 agreed their institution provides clear and enforceable guidelines on GenAI in assessment while 4 disagreed, with a neutral median (p = .958)—and the authors warn that unclear rules can lead to inconsistent student conduct and enforcement. Exploratory analysis produced three uncorrected associations below .05: frequent GenAI users were more familiar with adaptive tools, familiar respondents were more likely to agree that scaffolding supports verification, and frequent users were less likely to agree that excessive reliance causes offloading. None survived a correction for multiple testing. Seven respondents gave usable open-ended answers, coded into 8 segments dominated by cognitive erosion and cheating, with smaller themes of skills and access gaps, efficiency, fairness and a preference for traditional methods—evidence that the risks of Student-AI Interaction and Hallucination Risk are recognized alongside the benefits.

What this means for practice

  • Instructors. Design for hints, not answers: adaptive systems should guide students through questions, prompts and partial feedback rather than reveal a finished answer, since only 2 of 13 respondents identified step-by-step guidance that withholds the final answer as the feature they encountered.
  • Assessment designers. Make verification a graded step—require students to check AI-generated claims, identify errors and cite reliable sources—and introduce process evidence such as draft histories and short oral explanations gradually before giving it substantial weight in integrity decisions.
  • Administrators. Publish clear and enforceable guidelines stating what AI uses are permitted, restricted or prohibited and when disclosure is required; institutional guidance was the one item where agreement and disagreement were nearly balanced.
  • Researchers. Treat the three sub-.05 associations as leads, not findings—none survived correction for multiple testing—and follow the authors' proposed quasi-experiment with a validated critical-thinking measure and a delayed test after AI support is withdrawn.

Limitations

  • The sample of 13 convenience respondents is too small and unrepresentative for generalization.
  • Respondents were predominantly undergraduate (12 of 13) and more than half (7) from Law and Policy Studies; no faculty member participated.
  • The survey relies on self-reported perceptions rather than actual critical-thinking performance or observed integrity behavior.
  • Q6 was unusable because of a form-design error, some scale anchors were not exported, and some questions contained more than one proposition.
  • Correlations and subgroup comparisons were exploratory; none survived a correction for multiple testing.
  • Qualitative responses were coded by one researcher only.

Citation

Raj, A., & Kumari, S. (2026). Evaluating the Impact of Generative AI-Driven Adaptive Assessment Frameworks on Critical Thinking and Academic Integrity in Higher Education. EdArXiv preprint.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.