On this page

Synthesis: Keith, Wood and Posey ask a measurement question that much Generative AI research in education assumes away: does how much someone uses AI tell us anything about how they use it? Their answer is an eight-mode AI Engagement Typology — Oracle, Production Assistant, Tutor, Collaborative Problem-Solver, Verification Agent, Creative Expander, Critical Challenger and Problem Setter — organised as theoretical prototypes of increasing user cognitive governance across Passivity, Partnership and Agency tiers on Hammond's intuitive-analytic continuum. Study 1 measures reported typical use with a 29-item AI Interaction Typology in 359 U.S. working adults; Study 2 classifies 1,353 participant turns from 487 student transcripts and pairs them against the same instrument for 171 students. The modes carry different cognitive signatures: Oracle shows the broadest adverse adjusted pattern, Verification the only consistently favourable agency-tier one, while seven of eight modes correlate positively with both perceived productivity and dependency. Yet reported and observed profiles correspond at near zero — no correlation survives false-discovery correction across three coding-protocol generations, and behaviour predicts quiz scores where self-report does not. The takeaway is that exposure-based AI research may be measuring the wrong variable.

Key Findings

  1. Eight modes, not one exposure variable. The AI Interaction Typology (AIT) names Oracle, Production Assistant, Tutor, Collaborative Problem-Solver, Verification Agent, Creative Expander, Critical Challenger and Problem Setter, grouped into Passivity, Partnership and Agency tiers that the authors explicitly treat as theoretical prototypes of cognitive governance rather than a validated bipolar latent scale.
  2. Oracle carries the broadest adverse adjusted pattern. In 319 complete-case working adults, Oracle was positively associated with Cognitive Offloading (β = .39), the cognitive-harm composite (β = .32), atrophy of decision-making behaviour (β = .41) and the dependency composite (β = .33) — all surviving Benjamini-Hochberg false-discovery correction across 64 mode-outcome tests — and negatively with AI calibration (β = −.23).
  3. Verification is the only agency-tier mode with favourable associations surviving correction. Verification showed lower cognitive harm (β = −.19), lower atrophy of decision-making behaviour (β = −.26) and lower dependency (β = −.15), plus higher AI calibration (β = .55) and AI-specific Metacognition (β = .21). No Problem Setter, Creative Expander or Critical Challenger coefficient survived the 64-test correction.
  4. Production Assistant is productive but not protective. It correlated with perceived productivity (β = .28) and AI-specific metacognition (β = .15) but showed no independent association with cognitive offloading or atrophy of decision-making behaviour, so the simple passive-tier prediction of H1a held only partially.
  5. Simultaneous entry suppresses the agency modes. The agency-tier modes are strongly intercorrelated — Problem Setter alone shares 69.7% of its variance with the other seven — and 13 of 56 mode-outcome coefficients reversed sign from unadjusted to adjusted models, every reversal turning a positive association negative. Taken alone, every agency-tier mode was positively associated with perceived productivity and flourishing.
  6. Productivity and dependency travel together for seven of eight modes. Seven modes correlated positively with both perceived productivity (r = .20 to .50) and dependency (r = .08 to .48); Verification was the exception, at r = .04 with productivity and r = −.22 with dependency. Equivalence tests placed its productivity correlation inside ±.15 but not ±.10, and its dependency correlation within a small-negative rather than zero-equivalent band (TOST p = .63 at ±.20).
  7. Mode profiles add variance beyond usage quantity. Adding the eight modes to usage-only models raised in-sample R² by .23 on average (range .05 to .37) and cross-validated R² by .20, with seven of eight outcomes retaining a positive cross-validated increment; flourishing did not (ΔR² = −.01). Verification also held the largest zero-order correlations with AI calibration (r = .50), AI-specific metacognition (r = .42) and Productive Friction (r = .39).
  8. Self-report and observed behaviour do not correspond. Across 171 students with complete paired profiles, no AIT-to-classifier correlation survived FDR correction and none reached the conventional |r| ≥ .30 convergent-validity threshold; Verification came closest at r = .199 (95% CI [.05, .34], q = .072) yet met neither the ±.20 nor the ±.25 equivalence bound that five and seven other modes respectively satisfied. Mean total-variation distance between paired profiles was .701, only .002 closer than randomly re-paired profiles (permutation p = .010).
  9. Behaviour, not self-report, tracks quiz performance. In the learning task, behavioural composition raised R² by .166 (F = 5.82, permutation p = .005) against .031 for self-report (ns); the .135 gap's participant-bootstrap interval included zero (95% CI [−.02, .31]). The best-supported association was Tutor: a higher share of explanation-seeking turns accompanied higher quiz scores (r = .33, p < .001, populated by 134 of 154 profiles), though all models had negative cross-validated R², so the increment is in-sample only.
  10. The mode profile hides who initiated the work. Protocol 4.2 records whether each coded turn was self-originated or AI-proposed: 49% of engagement moves in the learning task were AI-initiated against roughly 6% in the two production tasks, and AI-proposed moves comprised 48.7% of codeable learning-task turns. This co-regulated share did not significantly predict quiz performance (r = .14, p = .08).

How the typology was derived and validated

The eight modes describe the functional role AI plays in a person's work rather than the amount of AI used. Table 1 holds a single referent task constant — drafting an analytical paper on an unfamiliar topic — and shows how the mode changes the prompt and the cognitive demand: Oracle asks what the literature says and recruits automation bias and algorithm appreciation; Production Assistant asks for the paper and recruits Cognitive Offloading and over-reliance; Tutor asks for a step-by-step walkthrough and recruits metacognitive monitoring; Collaborative Problem-Solver proposes a hypothesis and asks what evidence would test it; Verification Agent submits a draft for critique and recruits calibrated trust; Creative Expander asks for divergent lenses; Critical Challenger asks for the strongest objection and recruits critical thinking under AI and the illusion of understanding; Problem Setter questions the format itself and recruits Hammond's account of competent judgment under uncertainty. The tiers predict limited monitoring when AI supplies an answer, reciprocal work when it teaches or collaborates, and user-governed work when it checks, expands, challenges or reframes. The authors stress that a named active mode is not by itself constructive enactment — iterative dialogue can remain AI-led — and that Verification is especially narrow, because the user must form a judgment or artifact before seeing evaluation and then resolve disagreement; reviewing an AI-produced artifact is Production or Collaboration with an evaluation step, not independent Verification.

Measurement proceeded in two layers in parallel. The 29-item AIT supplies person-level self-report, scored as unweighted behavioural frequency indices, with observed alphas from .60 (Verification) to .83 (Challenger) in the adult sample and .52 (Production) to .81 (Challenger) in the student sample; because the items sample distinct constituent behaviours, the authors report internal consistency descriptively rather than as a validity gate. The message-level classifier codes each participant turn into one of the eight modes or mode 0 (no codeable move) under the AIEM Mode Coding Standard, protocol version 4.2, using a structured-prompt call to Anthropic Claude Haiku 4.5 at temperature 0 with the preceding assistant turn visible. A five-model consistency panel reached Fleiss κ = .78 (unanimity 74%, pairwise κ = .71 to .85), while the two-coder human reference produced under an earlier manual gave inter-coder κ = .49; agreement with that reference rose monotonically across protocol generations to κ = .65, and to κ = .82 where four of five models concurred. The authors are explicit that model agreement bounds consistency, not validity, and that no reported figure is an inter-rater reliability of protocol 4.2 itself.

The structural evidence is mixed and the paper reports it that way. A six-factor self-report structure outperformed a three-tier collapse, and the eight-factor model was not statistically distinguishable from a six-factor solution merging Tutor with Collaborative Problem-Solver and Creative Expander with Problem Setter. The classifier can assign eight labels, but sparse assignment does not establish eight behaviourally separable constructs, and human validation confirms label accuracy only for the two corpus-dominant modes.

The eight modes and what distinguishes them

The distinction the typology is built to make is between modes that look identical in a usage log but differ in who does the cognitive work. Large Language Models (LLMs) sessions of the same duration can contain copying an answer, seeking an explanation, co-producing an artifact, or testing a user-authored judgment, and the tier structure orders these by how much cognitive governance stays with the user. Verification is drawn most sharply: the user authors first, inspects AI output against their own judgment, and resolves disagreement, which is why acceptance is explicitly not verification under the coding standard and why the authors argue appropriate reliance requires observing uptake of correct and incorrect advice rather than inferring it from instructions or interface labels.

The Student-AI Interaction literature already supplies engagement levels such as the ICAP Framework, engagement domains, functional roles, and observed person profiles, so the authors claim no component-level novelty. What they claim is integration: eight theoretically grounded modes measured in the same participants by instrument and by transcript, a person-level correspondence test, a behaviour-versus-self-report incremental-validity contrast, turn-level initiation records, and a versioned, agreement-audited coding protocol. Mode 0 exists to keep non-engagement visible: bare acknowledgements, uploaded artifacts and sequence-closing assessments made up 5.2% of the 1,353 coded turns and were excluded from mode denominators while remaining in the turn count.

Cognitive signatures attached to each mode

Each mode carries a predicted signature, and Study 1 tested them simultaneously with usage covariates (task percentage, use frequency, session duration, number of AI tools) and dispositional covariates (critical-thinking dispositions, AI Self-Efficacy, need for cognition) in 359 eligible U.S. working adults recruited through Prolific in late April 2026. Oracle's signature is the least favourable and the widest: elevated offloading, cognitive harm, atrophy of decision-making behaviour and dependency, with depressed AI calibration. Production Assistant's is productivity without measured cognitive cost. Tutor, Collaborative Problem-Solver, Creative Expander, Critical Challenger and Problem Setter show no adjusted coefficient that survives correction, which the authors read cautiously because suppression rather than absence may be responsible.

Verification is the pivot of the paper's cognitive claim. It is the only agency-tier mode whose coefficients do not reverse between specifications and remain stable, and it shows the strongest positive association with AI calibration (β = .55) as well as a negative dependency association — the pattern the authors read as a Verification-related self-concept. They caution that this is not proof of behavioural skill, that the proposed artifact-authorship mechanism was never measured, and that the gap between Verification and the other agency modes is narrower in the unadjusted specification than the adjusted table implies. The Trust Calibration framing matters because elsewhere five therapists in a fixed-order study accepted erroneous AI advice that converted an average of three previously correct answers to wrong ones despite instructions to verify — evidence that verification cannot be inferred from access or stated intention.

The self-report–behavior gap and what it means for survey research

The divergence is the paper's most consequential result for anyone who studies AI in education by questionnaire. Reported typical engagement and task-specific enacted engagement showed near-zero correspondence, stable across three coding-protocol generations, with a different single near-threshold correlation in each generation (Oracle twice, Verification in the reported one), so the authors refuse any mode-specific convergence reading. In the learning task the behaviour layer added R² = .166 while the self-report layer added .031, but the direct bootstrap interval on the difference included zero, so H4b-ii was unsupported and the behavioural increment is in-sample only. Mean profiles were 82% Tutor in the learning task and 85% and 84% Production in the writing and evaluation tasks, which the assignments explicitly afforded — a reminder that task design constrains measured mode variance.

The authors refuse the easy interpretation in both directions. They do not treat the weak correspondence as invalidating either layer: the transcript codes carry criterion validity the self-reports lack through quiz prediction, the self-report scales show differentiated FDR-corrected outcome associations, and what fails is the assumption that the two channels index the same construct. Candidate explanations listed include different reporting windows, limited self-awareness, coder error, and task constraints. Because the same outcomes were never measured alongside both engagement layers, the studies cannot establish a general belief-versus-performance architecture; a decisive test would measure reported typical engagement, enacted task behaviour, productivity, dependency and independent performance in the same participants. The contribution this extends is the discrepancy between logged and self-reported digital use, moved from amount to functional mode.

Implications, design guidance and limitations

For practice, the design implication is deliberately narrow. With the corpus-dominant labels showing strong agreement with the human development reference, mode feedback could show a user a task-specific distribution over validated modes and ask whether it matches their goal, but feedback on rare agency-tier modes awaits construct-revised classification, and the authors argue current evidence supports experimental mode induction and reflective feedback rather than automatic restrictions or prescriptive mode-task alerts. Mode induction has been shown to change interaction patterns, not to make one mode universally superior, and existing interventions are AI-scaffolded without non-AI controls. For Higher Education specifically, the practical warning is that Learning Analytics dashboards built on usage quantity will misdescribe what students are doing, and Self-Regulated Learning supports premised on students accurately reporting their own AI use rest on an assumption this study does not support. The theoretical frame attached to the results is a conditional assistance-removal hypothesis: assistance may fail to transfer when it replaces the cognitive process later tested, and more readily when it elicits explanation, retrieval, comparison, revision or independent production.

The limitations are stated at length and are unusually specific. Both studies are cross-sectional, use U.S. samples, and cannot establish temporal change or causation; Study 1 shares response method between predictors and outcomes, leaving self-concept, reverse direction and omitted variables viable. Study 2's behavioural measure carries four disclosed limitations: the measure moved as coding defects were repaired across protocol generations, so every finding carries its protocol version; no reported figure is an inter-rater reliability of protocol 4.2 itself; turn-level coding of problem setting extends Schön's practitioner-level construct with boundary rules partly derived on other corpora; and tasks constrained mode variance so that the five non-Tutor learning-task modes remain sparse and quiz models do not generalise out of sample. Several short scales had α < .70, the tiers are unvalidated prototypes, and competitor constructs were never administered. A strict narrative sensitivity analysis restricted to the 59 low- or some-concerns reports left every headline literature conclusion unchanged. The authors' closing position is that the evidence supports an integrative measurement framework, not a causal tier theory: Student Engagement with AI is multidimensional, the two measurement channels must be run in parallel rather than substituted, and construct-revised rare-mode classification, competitor measures, randomized mode induction and delayed independent outcomes are the next tests.

Connected Concepts

  • Cognitive Offloading — the outcome most strongly associated with Oracle and the mechanism the Passive tier is theorised around
  • Self-Report Measures — the AIT's person-level channel and the layer that failed to correspond with observed behaviour
  • Educational Measurement — the convergent-validity, equivalence-testing and incremental-variance apparatus used throughout
  • Assessment Validity — the auditability frame the authors apply to their own classifier, including what model agreement bounds
  • Trust Calibration — the theoretical foundation of the Verification Agent mode and the AICal principle scale
  • Metacognition — AI-specific metacognition (AIMeta) as a mode-associated principle and the mechanism behind the Tutor mode
  • Critical Thinking — critical-thinking dispositions as covariates and the foundation of the Critical Challenger mode
  • Student-AI Interaction — the broader literature on how students actually converse with AI systems
  • Human AI Collaboration — the human-first participation designs the paper contrasts with Human Confirmation
  • Learner Agency — cognitive governance, the dimension the Passivity–Partnership–Agency tiers order
  • ICAP Framework — the established engagement-levels framework the typology positions itself against
  • Learning Analytics — the usage-level measures whose explanatory value the mode profiles exceed
  • Self-Regulated Learning — co-regulated and self-originated engagement, recorded per turn through the initiation property
  • Student Engagement — multidimensional engagement with AI as the paper's headline construct claim

Connected Articles

Citation

Keith, M. J., Wood, D. A., & Posey, C. (2026).The eight-mode AI engagement typology: Differential cognitive signatures and a self-report–behavior gap. Preprint (v19, 2026-08-18), submitted to Scientific Reports. Brigham Young University, Marriott School of Business. Unrefereed preprint; no public DOI assigned at the time of writing.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.