Research Article
Generative AI and the Productivity Divide: Human-AI Complementarities in Education
Synthesis: Idan & Anand (2026) conduct an RCT showing that GenAI access significantly increases task performance on average — but the gains are highly uneven, NOT predicted by GPA or prior knowledge, but by AI Interaction Competence (AIC): the ability to elicit, filter, and verify model outputs. High-AIC participants realized outsized gains while low-AIC saw limited or negative returns. A scaffolding intervention (conceptual maps) reduced outcome variance, showing that standardized workflows can mitigate the new "AI productivity divide."
Key Findings
- GenAI access increased mean task performance — approximately a 17% productivity lift — but the gains were concentrated among high-AIC users while low-AIC users saw limited or even negative marginal returns.
- GPA and prior knowledge did NOT predict GenAI-augmented performance; AI Interaction Competence (AIC) was the decisive moderator, supplanting traditional academic credentials as the axis of a new productivity divide.
- A conceptual-map scaffolding intervention reduced outcome variance by nearly 40% without lowering the mean, disproportionately helping low-AIC novices by "leveling the floor."
- Extending mandated study time (three to four hours) produced no significant performance gains, revealing diminishing returns to time-on-task in the absence of effective interaction skills.
- Participants strongly preferred LLMs (69%) over lectures, YouTube, and textbooks, and baseline-condition participants dropped out at more than twice the rate — attrition as a revealed-preference signal.
Background and Motivation
Generative AI is transforming how firms create, process, and apply knowledge, yet the heterogeneity of its productivity effects across users remains poorly understood. Drawing on the management tradition that technology creates value only when paired with complementary human and organizational capabilities, the authors argue that GenAI heightens this complementarity: unlike earlier automation that standardized routine tasks, large language models require users to engage in iterative problem solving — formulating prompts, interpreting probabilistic outputs, and verifying content quality. These interactional skills are tacit, unevenly distributed, and rarely taught. The paper introduces AI Interaction Competence (AIC) as a new dimension of human capital that determines how individuals translate AI access into performance, and frames GenAI adoption as a problem of capability design rather than tool procurement. The result is a new productivity divide driven by interaction skill rather than domain expertise.
Method
The study is a randomized controlled experiment with 179 participants recruited primarily from engineering programs at Texas A&M University, approximating the conditions of early-career knowledge workers learning and applying unfamiliar technical information. After a profiling survey and a 15-item pre-intervention exam, participants were randomized into a Baseline condition (traditional resources only) or an LLM condition (restricted to ChatGPT). Within the LLM condition, novice learners were further randomized into Scaffolding sub-conditions: a baseline-LLM arm, a time-on-task arm (four rather than three hours daily), a conceptual-roadmap scaffolding arm, and a peer-collaboration arm. The outcome was post-intervention exam performance normalized to the unit interval, analyzed with OLS regressions and interaction terms controlling for baseline performance, GPA, and other covariates.
Results
Participants' self-assessments correlated meaningfully with measured performance (strongest for general Machine Learning knowledge, ρ = .71), yet the dimensions they could introspect accurately were not the ones driving inequality. LLM-condition participants scored significantly higher on the post-intervention exam (M = .56 vs. .48) and dropped out at far lower rates, with preferences and attrition converging as revealed-preference signals. Crucially, GPA (p = .59) and prior knowledge (p = .2) showed no significant interaction with treatment, while the Treatment × AIC interaction was positive and significant. A three-way interaction showed that novices with low prior knowledge benefited most — but only when they had high AIC. Scaffolding compressed heterogeneity (SD dropped from .22 to .14), while additional study time yielded no gains.
Discussion and Implications
The findings reframe the educational and organizational stakes of GenAI. Because AIC — an emerging, untested, and unevenly distributed AI literacy — now determines who advances and who falls behind, educational institutions and firms should pair access with short AIC micro-training (60–90 minutes covering prompting logic, verification heuristics, and synthesis structure) and light scaffolds such as standard operating procedures, prompt templates, and review checklists. These process interventions reduce performance variance by about one-third without diminishing the mean, directly addressing the equity challenge of AI-mediated learning. The paper's central message is that the effective management of GenAI hinges less on the technology itself and more on the design of complementary routines that embed consistency, discipline, and Feedback into human–AI interaction — a shift toward organizational capability design.
What this means for practice
- Learners. Invest in AI Interaction Competence rather than in more access: the productivity lift went to participants who could elicit, filter, and verify model outputs, while GPA (p = .59) and prior knowledge (p = .2) showed no significant interaction with treatment.
- Learners. When you are new to a topic, study from a conceptual roadmap that sequences the material, and check topics off as you go; scaffolded novices outscored unguided ones (M = .45 vs. .38) and the benefit was concentrated among low-AIC participants.
- Instructors. Teach 60–90 minutes of prompting logic, verification heuristics, and synthesis structure instead of extending study time; adding a fourth daily hour to the required three produced no significant gains.
- Administrators. Pair any AI access rollout with light process scaffolds — prompt templates, standard procedures, review checklists — which compressed the standard deviation of outcomes from .22 to .14 without lowering the mean.
- Researchers. Measure AI interaction competence as a moderator in AI-in-education experiments. It, not prior attainment, carried the significant treatment interaction, so studies that only control for GPA will miss the effect.
Limitations
- Single institution, students standing in for workers. 179 participants recruited mainly from engineering programs at one university, studied as an approximation of early-career knowledge workers; the population is not the workforce the framing invokes.
- Short intervention, immediate outcome. Three consecutive days of self-study with a post-intervention exam normalized to the unit interval and no delayed measure of retention or transfer.
- Weakly significant scaffolding effects. The scaffolded-versus-unguided difference and the scaffolding × AIC interaction were only weakly significant (p > .05, p < .10), and the authors note that subgroup contrasts had limited statistical power.
- Self-reported AIC and preferences. Interaction competence is measured from self-assessments whose alignment with measured performance varied by domain (ρ = .71 for general machine-learning knowledge), so the decisive moderator may partly reflect self-perception.
Citation
Idan, L., & Anand, B. (2026). Generative AI and the Productivity Divide: Human-AI Complementarities in Education.