Research Article
Does generative AI narrow education-based productivity gaps? Evidence from a randomized experiment
Synthesis: In a randomized controlled trial with 1,174 participants, Cruces et al. find that generative AI substantially narrows education-based productivity gaps, closing approximately three-quarters of the initial performance difference between higher- and lower-education workers. Critically, gains are not purely from delegation — lower-education participants retain part of their improvement after AI is removed, and follow-up performance improves when intensive AI use is combined with sustained effort. This study provides causal evidence that AI tools can serve as productivity equalizers in workplace tasks.
Experimental Design
The study employed a randomized online experiment with 1,174 adults aged 25-45 in Argentina completing workplace-style Problem Solving tasks:
- Treatment group: Access to a generative AI assistant during the main task
- Control group: No AI assistance
- Follow-up module: Both groups completed an unassisted module to measure learning retention
Access was randomized across education-defined groups (high school versus postsecondary), holding the task fixed to identify how AI changes the productivity advantage associated with formal education. Responses were scored by an AI-assisted grading procedure validated against independent human graders. Chat logs were analyzed to understand differential AI usage patterns across education levels.
Key Findings
- AI substantially narrows the education-based productivity gap: Without AI, high-education participants outperform low-education participants by 0.548 SD; with AI, the gap shrinks to 0.139 SD, closing about 75% of the baseline difference. AI raises performance for both groups, with larger gains for lower-education participants (1.242 SD vs. 0.834 SD).
- The gains are not purely delegation: On an unassisted follow-up module, treated participants do not perform worse than controls once AI is removed, and low-education participants retain part of their improvement (0.171 SD), though a sizable education gap (0.200 SD) re-emerges.
- Carry-over depends on engagement: Intensive AI assistance produces strong submitted answers even with low task engagement, but follow-up performance is substantially higher only when intensive AI use is combined with sustained effort.
- The gap narrows but does not disappear because of effective use: Lower-education participants obtain substantial assistance from AI, while higher-education participants use the tool somewhat more effectively across margins such as prompt detail and workflow structure.
The Equalizing Effect on Task Performance
The main outcome was an overall task score standardized relative to the low-education control group. In the no-AI control condition, higher-education participants outperform lower-education participants by 0.548 SD, and access to a GPT-based assistant closes about three-quarters of that gap, leaving a residual 0.139 SD difference that is large but marginally insignificant. The treatment effect is 0.408 SD larger for low-education participants, with gains operating through both the content and writing components of the score, so the result is not driven by writing improvements alone. Because only about 13% of treated low-education participants reach the top score category, the convergence is not an artifact of ceiling effects — the task remains cognitively demanding, and participants must still evaluate, select, and integrate AI-generated content, so human judgment continues to shape answers.
Why Gains Carry Over After AI Is Removed
The authors paired the assisted task with an immediate non-AI-assisted follow-up module to test whether gains reflect delegation to the tool rather than productive use. The results do not support a pure delegation interpretation: treated participants do not perform worse than controls once AI is removed, and low-education participants show a modest 0.171 SD gain in follow-up performance, consistent with some internalization of the task rather than mere exposure to a better answer. However, low-education treated participants still lag their high-education counterparts by 0.200 SD, indicating that underlying human capital continues to shape unassisted performance. A four-way split on AI assistance and task engagement sharpens the picture: intensive AI use predicts strong task scores even when engagement is low, but follow-up performance is substantially higher only when intensive assistance is combined with sustained effort — exposure to a correct AI-generated answer is not by itself sufficient.
Patterns of AI Use: Why the Gap Persists
Chat-log analysis of 471 treated participants who used the assistant explains why the gap narrows substantially but does not disappear. There are no meaningful differences between education groups in the intensity of interaction: both send a similar number of messages and request help on a comparable share of task components (around 61% in both groups). Differences instead emerge along qualitative dimensions of use. Higher-education participants provide more detailed instructions guiding the assistant's reasoning, are more likely to begin with a highly specific prompt and a structured workflow, and are more likely to use AI output as an input into their own writing rather than fully copying it. Lower-education participants are 10 percentage points more likely to copy-paste AI-generated text into their final answer. This suggests that effective use of AI remains partly shaped by underlying human capital, even when access and onboarding are equal.
Robustness Checks
The main results are stable across the ten iterations of the LLM-based grading used to score responses, and correlate highly (above 0.9) with manual grading by independent human graders and with alternative grading procedures, including an Elo-based approach. Estimates also hold under alternative definitions of the high-education group and after controlling for observable characteristics such as age, gender, employment status, and work experience.
What this means for practice
- Policymakers. Fund productive adoption, not just tool access: with AI the 0.548 SD baseline gap between higher- and lower-education participants shrank to 0.139 SD, but a 0.200 SD gap reappeared in the unassisted follow-up, so equalizing task-level capability is not the same as equalizing outcomes.
- Instructors. Teach use strategies explicitly, because they are not picked up from access alone: higher-education participants gave more detailed instructions and used AI output as an input to their own writing, while lower-education participants were 10 percentage points more likely to copy-paste generated text.
- Administrators. Require sustained effort alongside AI assistance when the goal is learning rather than completed work: intensive AI use produced strong submitted answers even at low engagement, but follow-up performance improved substantially only when intensive assistance was combined with sustained effort.
- Researchers. Measure assisted and unassisted performance in the same design: the immediate non-AI follow-up module is what separates internalization from delegation, including the 0.171 SD retained gain among lower-education participants.
Limitations
- 1,174 participants aged 25 to 45 in Argentina, recruited from three commercial panel companies out of 74,449 invited, with the sample split 520 low-education and 654 high-education; the personal-computer requirement excluded more low-education applicants, so it is 44% low- and 56% high-education.
- One self-contained workplace-style task that participants were told would take 20 minutes (they averaged 21 minutes), followed by an immediate unassisted module — there is no delayed retest or workplace follow-up.
- The design deliberately abstracts from firms, wages, and organizational task allocation, so the estimates are task-level capability effects and the authors state they should not be read as predictions about wage or aggregate inequality.
- Outcomes rest on an LLM-based grading procedure validated against human graders on a random 10% subsample of 117 responses (correlation above 0.9); the equalizing pattern is also specific to one generation of AI capability and may attenuate or reverse as models change.
Citation
Cruces, G., Fernandez Meijide, D., Galiani, S., Galvez, R., & Lombardi, M. (2026). Does generative AI narrow education-based productivity gaps? Evidence from a randomized experiment. v1.