Research Article
Generative Artificial Intelligence and Creativity in K–12 Education: A Systematic Scoping Review
Synthesis: Rahimi, Babaee, Esmaeiligoujar and Dede ask how Generative AI has been used to support or assess Creativity in K-12 education, and answer with a PRISMA-guided scoping review of 45 studies published between 2017 and 2025. The corpus began as 2,321 records drawn from Web of Science, ERIC, Scopus, ScienceDirect, ProQuest and Google Scholar plus eight creativity and AI-in-education journals; after 399 duplicates were removed and 129 full texts were screened against the inclusion criteria by two researchers, a hand-search of 11 key researchers' recent work and the reference lists of included papers added 13 studies. Four categories organize the field: creativity assessment (4 studies), Human AI Collaboration or co-creativity (3), stakeholders' perceptions (14) and creativity enhancement (34), with several studies counted in more than one category. Language-based storytelling and writing dominate (17 studies); only one study addresses music and only two explicitly address equity or Accessibility. Large language models can score creativity in close agreement with human raters and well-designed scaffolds raise creative output, but the authors find thin theoretical grounding, weak Assessment and genuine risks of homogenization and over-reliance. The takeaway for K–12 practice: design GenAI to augment student authorship, not to automate it.
Key Findings
- The corpus is 45 studies, and it grew out of 2,321 records. Six databases (Web of Science, ERIC via ProQuest, Scopus, ScienceDirect, ProQuest, Google Scholar) and eight targeted journals produced 2,321 records; 399 duplicates were removed, 129 full texts were screened by two researchers, and a hand-search of 11 prominent creativity-and-AI researchers plus reference-list checking added 13 more studies. Three researchers coded inductively and reconciled differences; a fourth expert in AI in education reviewed the themes.
- Four categories structure the field. Creativity assessment (4 studies), human–AI co-creativity (3), stakeholders' perceptions (14, with four sub-themes) and creativity enhancement (34, across creativity support tools, specific effects and broad effects). Because a study can sit in more than one category, the category counts exceed the 45 included articles.
- GenAI can score creativity close to human raters. Acar's MOTES showed high internal consistency (H = .89) and AI–human agreement of rs = .79–.91; Goecke's multilingual XLM-RoBERTa correlated r = .80 with human originality ratings in German science writing; Rahimi's GPT-4o scoring of student-built Physics Playground levels correlated r = .81 with human ratings when a rubric covered elaboration, aesthetics, originality and surprise.
- Language tasks crowd out other modalities. Storytelling and writing were the most common creative tasks (n = 17 of the reviewed creativity-enhancement studies); creative coding and programming were prominent (n = 7). Exactly one tool study targeted music (Citizen DJ, early childhood) and none addressed embodied or spatial creativity, which the authors treat as a narrowing of how creativity is conceived in the field.
- Custom-built tools outperformed off-the-shelf platforms. Among 11 creativity support tools, 8 were purpose-built and 4 adapted general platforms such as ChatGPT, DALL·E, Canva and Leonardo.AI. ChatScratch raised the creativity support index (M = 84.0 vs. 75.4, p < .05) and MindScratch beat Scratch on code quality, creativity and Computational Thinking; the authors attribute this to age-appropriate scaffolds such as voice input, storyboards and visual mind maps, which general platforms lack.
- Stakeholders consistently frame GenAI as scaffold rather than author. Across all 14 perception studies, preserving student agency and creative ownership was the one universal sub-theme; 11 emphasized creativity as a social and cultural practice, 10 the need for transparency and trust, and 9 emotional resonance and personal meaning. In Marrone's survey of 80 Australian secondary students, many rejected the idea that AI could match human creativity at all.
- The authors' own concern is homogenisation. They cite findings that LLMs shrink linguistic diversity and amplify dominant stylistic traits, that detailed prompting did not stop AI-assisted college admissions essays from converging, and that increasingly AI-contaminated training data compounds the problem. They add metacognitive laziness among co-writers and note that over-reliance appeared as a risk in perception, co-creativity and curriculum studies alike.
- Assessment rigour is the weakest link. Most included studies never defined creativity or used an established framework (the 4C model of Creativity and Csikszentmihalyi's systems theory are both noted as rare), and the broad-effects studies inferred creativity from engagement or novelty without validated measures. Where GenAI scoring was tested, threats remained: dataset bias, representational limits, hallucination, black-box scoring, privacy and rubric overfitting.
- Theory and equity are the stated gaps. Only two studies explicitly addressed equity and accessibility; most assumed fluent readers with unrestricted device access and looked at English-language, largely Western or East Asian classroom literacy contexts. The five-cycle ethical framework (access, representation, algorithmic bias, interpretation, citizenship) is put forward as the lens the field mostly omitted.
How the review was conducted and what it covers
The authors chose a scoping review because the intersection of Generative AI, creativity and schooling is young, methodologically mixed and not yet ready for an effect-size synthesis; they followed PRISMA reporting and the Arksey and O'Malley / Peters guidance for scoping work. The search covered 2017–2025, with 2017 taken as the start because that was the year "Attention Is All You Need" introduced the transformer architecture behind the Large Language Models (LLMs) models and text-to-image systems the review tracks.
Inclusion required: published in 2017–2025; written in English; focused on K–12 education; creativity as a primary or secondary outcome; use of GenAI tools such as ChatGPT or Midjourney; and full-text availability. Exclusions mirrored these — pre-2017, non-English, adult or Higher Education or informal settings, no attention to creativity, traditional or rule-based AI rather than generative models, and no full text. All manuscript types were admitted, including empirical studies, reviews and technical reports, which is why the evidence base ranges from controlled experiments to conceptual papers.
How GenAI is used for creativity in K–12
The review's organizing distinction is between assessing creativity with GenAI and enhancing it. In the four assessment studies, Large Language Models (LLMs) models scored divergent-thinking tasks, creative products, drawings and game levels, with the shared aim of replacing labor-intensive subjective scoring with scalable, consistent judgment. The three co-creativity studies put the model inside open-ended tasks instead: children co-told stories with voice agents, co-wrote sustainability-themed texts with ChatGPT, and took turns writing with a GPT-2-based storytelling partner that preserved their authorship while assigning the agent limited social standing.
The larger picture comes from the 34 creativity-enhancement studies. Custom-built tools clustered into creative coding environments (ChatScratch, MindScratch, CoRemix, App Planner), narrative and storytelling systems (AIStory, collaborative storytelling, creative mathematical writing) and one Multimodal AI music system for preschoolers. Off-the-shelf use, by contrast, gave students fast access to novelty through prompting but offered no mechanism to reflect, revise or explore alternatives. Several interventions embedded Generative AI in whole curricula — a nationwide Korean creativity program, an AI curriculum for 9–12 year olds in Taiwan, an LLM-and-robot dual-teacher classroom — which is where Project-Based Learning and Inquiry-Based Learning structures met GenAI directly. Creative coding and Game-Based Learning design recur throughout, and possibility thinking, with its "what if" and "as if" moves, is the pedagogical vocabulary the authors borrow from Beghetto's Human × AI bots.
Positive findings, and the concerns the authors add
The positive case rests on consistency between machine and human scoring, on engagement gains, and on evidence that scaffolded tools lift measured creative output: CoRemix increased remixing activity and rated enjoyment, expressiveness and immersion; Magic Camera stories were rated significantly more creative and fluent than stories made without it; the KAIT Design Thinking model reported 75% of students thinking more creatively and 55.6% feeling empowered; and Fermi problem-based learning with AI outperformed traditional instruction on 21st-century skills (M = 18.25 vs. 13.5). Custom scaffolds that preserved authorship — voice input, visual planning, divergent "many options" prompts — recurred in the strongest examples.
The concerns, however, come from the authors as much as from the corpus. They state plainly their own doubt that GenAI can be an excellent co-writing partner, because of structural properties of Large Language Models (LLMs) models: lowered linguistic diversity, phrasing that seldom "colours outside the lines", homogenization of style and the contamination of future training data by AI-generated text. They flag metacognitive laziness in co-writing, over-reliance that can undercut creative confidence and self-awareness, and a corpus where concern about homogenization is largely absent from the co-creation literature itself. Stakeholder worries were more cultural and emotional: students rejected AI stories that ignored their norms, valued outputs that lacked "feeling" or "soul", and doubted machine creativity, echoing Runco's argument that AI produces "artificial creativity" rather than the human kind. The mitigation strategies the authors report are technical rather than pedagogical — retrieval-augmented generation, context engineering that injects the learner's own voice and knowledge, fine-tuning small specialized models, and locally hosted small language models to contain cost and access barriers.
How creativity is assessed in these studies: the measurement problem
Assessment is the review's sharpest diagnosis. Four studies carried the construct seriously, each aligning machine scores with expert human ratings and grounding tasks in creativity theory, whether divergent thinking (Alternate Uses, Instances, Sentence Completion) or a rubric spanning elaboration, aesthetics, originality and surprise. Together they showed convergence across verbal, scientific and visual-spatial domains, and Acar's historical synthesis traced four phases of AI scoring — semantic statistics, latent semantic analysis, word embeddings and now large language modeling — toward greater reliability and generality, while warning that models must be transparent and bias-aware rather than enforcing narrow norms.
Elsewhere the construct thinned. Most studies measured products or engagement and never said how they defined creativity; only two treated it as multidimensional, spanning fluency, originality, elaboration and emotional expression; and assessment of the creative process — preparation, incubation, illumination, verification — was essentially absent. The authors' prescription is measurement-first: creativity has to be assessed well before it can be enhanced reliably, so they point to evidence-centered design and stealth assessment, which score unobtrusively inside digital tasks and so fit the open-ended, Game-Based Learning environments where GenAI creativity actually happens, and to creativity's social-emotional side, above all creative Self-Efficacy, which most studies only implied. They also flag validity threats they say cannot be ignored: dataset bias, hallucination, rubric overfitting, data privacy and the cost and opacity of the large models doing the scoring.
Design implications and the frameworks the authors propose
The chapter's constructive contribution is a pair of frameworks it applies rather than invents. The five-cycle ethical framework from Dieterle and colleagues names five interconnected divides — access, representation, algorithmic bias, interpretation and citizenship — and the review uses it to show what the corpus missed: hardware and connectivity gaps, Western-centric training data that omits African, Latin American and Middle Eastern languages and worldviews, stereotyping in generated stories, opaque outputs, and the wider societal stakes of who gets to create. The second is creativity augmentation or intelligence augmentation, in which GenAI supplies ideation breadth and humans supply ethical reasoning, context and ownership; possibility thinking supplies the classroom tactics.
Five practical recommendations close that argument: professional development on GenAI ethics and equity, so teachers can address bias, surveillance and unequal access and move beyond efficiency toward genuinely new instruction; student AI Literacy programs built on functional, critical and rhetorical literacies, in which learners critique and revise AI suggestions rather than accept them; clear classroom policies that permit flexible use while requiring disclosure of AI-assisted work and preserving the student's own voice; explicit instruction in critical engagement, comparing AI and human outputs and reflecting on what makes a voice authentic; and continuous monitoring of whether students are becoming more autonomous and reflective or more dependent. The design message is consistent: Scaffolding should be developmentally appropriate, culturally responsive, transparent, and structured to extend the creative process rather than substitute for it.
What this means for practice
- Teachers. Choose purpose-built tools with age-appropriate scaffolding - voice input, storyboards, visual mind maps - over off-the-shelf platforms: ChatScratch, which supplies these, lifted the creativity support index to M = 84.0 against 75.4 for the control (p < .05).
- Educators. Keep GenAI in an augmentation role in which the student supplies ethical reasoning, context and ownership while the model supplies ideation breadth, and treat it as a scaffold rather than an author.
- Curriculum designers. Write classroom policy that permits flexible use but requires disclosure of AI-assisted work and preserves the student's own voice, and teach students to compare AI and human outputs critically.
- Curriculum designers. Adopt the review's five recommendations: professional development on GenAI ethics and equity, functional-critical-rhetorical AI literacy for students, disclosure-based policy, explicit critical engagement, and continuous monitoring of whether students become more autonomous or more dependent.
- Researchers. Define Creativity with an established framework and measure it with evidence-centered design or stealth assessment before claiming enhancement; most reviewed studies never stated how they defined the construct.
Limitations
The review also circulates under a second title — Generative AI for Creative Learning in K-12 Education: Insights from a Systematic Scoping Review — with the same authors and the same corpus of 45 studies, so this page covers both versions.
- Genre. A scoping review maps breadth rather than estimating effects, and the authors report no effect sizes, risk-of-bias appraisal or quality weighting for the 45 studies, so nothing in it establishes that GenAI causes creativity gains; the closest the corpus comes is small-scale experimental comparison — ChatScratch against Scratch, MindScratch against Scratch, an AI-supported problem-based curriculum against conventional teaching — mostly with short exposures.
- Thin evidence where it matters most. The review's own question rests on four assessment studies and three co-creativity studies, plus a broad-effects subcategory where creativity "emerged" from open-ended activity without formal measurement.
- Coverage. The corpus is English-language only by criterion, heavily skewed to language and writing tasks, and contains a single music study and no embodied or spatial work; only two explicit equity studies means the review can say little about who benefits. Much of the creativity-tool work appears in human–computer interaction and interaction-design venues such as CHI and IDC rather than in the assessment or creativity-measurement literature that supplies the psychometric standards the authors invoke, and several "studies" are frameworks, prototypes or conceptual papers, so design insight and demonstrated learning gain are not the same evidence.
- Currency. The Generative AI capabilities under study move faster than publication cycles: the review's 2025 tools will date even if the problems of theory, measurement and equity that it diagnoses do not, and the authors' five-cycle and measurement arguments will need re-testing on the next generation of models.
Connected Concepts
- Creativity — the construct at the center of the review, which most included studies left undefined
- Generative AI — the technology the corpus applies to K–12 creative learning, in assessment, co-creation and enhancement roles
- Human AI Collaboration — co-creativity as alternating agency between student and model in storytelling and writing tasks
- Storytelling in Education — the dominant creative task type, present in 17 of the reviewed enhancement studies
- Multimodal AI — GenAI's capacity for visual, text and audio expression, and the research gap outside language modalities
- Automated Assessment — LLM scoring of divergent thinking, drawings and game levels against human ratings
- Psychometrically Aware AI — the construct validity, fairness and reliability standards the authors demand of creativity scoring
- Equity — access, representation and citizenship divides that only two studies addressed directly
- Learner Agency — creative ownership, the one universal stakeholder sub-theme and the review's design criterion
- AI Literacy — functional, critical and rhetorical literacies as the precondition for reflective use
- Early Childhood Education — the age band where custom scaffolds and child-friendly design mattered most
- Bias Mitigation — algorithmic bias, cultural stereotyping and how students were taught to notice and counter them
- Meta-Analysis and Systematic Review — the PRISMA-guided scoping method used to map 45 studies from 2,321 records
- Limitations in AIEd Research — thin theory, weak measurement and non-comparable designs as the field's standing problems
Connected Articles
- TriKoNet: The Trivalence Model of Potential Co-Creativity in Socio-Technical Networks — Co-creativity modeled as a socio-technical network, complementing this review's co-creativity category
- In the AI era: A project-based digital storytelling framework for art and design education — Project-based digital storytelling as a structured frame for AI-supported creative work
- Tailoring AI Agents for Early Learning: The Creative Project Approach — Developmentally tailored AI agents and the Project Approach for early-childhood creativity
- Generative AI in Design Thinking Pedagogy: Enhancing Creativity, Critical Thinking, and Ethical Reasoning in Higher Education — GenAI across design thinking's stages, with the same authorship and bias concerns
- Evaluating the Impact of AI-Supported Inquiry-Based Learning on Students' Creative Mathematical Performance, Critical Problem-Solving Skills, and Attitudes Toward Mathematics — Inquiry-based learning with AI measured on creative mathematical performance
- The cognitive impact of ChatGPT in higher education: A systematic review of critical and creative thinking outcomes — A systematic review of ChatGPT's effects on critical and creative thinking
- LLMs to Support K-12 Teachers in Culturally Relevant Pedagogy: An AI Literacy Example — LLMs used to build culturally relevant K–12 pedagogy, addressing the representation divide
- Generative Artificial Intelligence (GAI) in Teaching and Learning Processes at the K-12 Level: A Systematic Review — A systematic review of GenAI in K–12 teaching and learning more broadly
- The Evidence Base on AI in K-12: A 2026 Review — The causal-evidence picture for AI in K–12 classrooms
- Cultivating Design Creativity of Vocational Students: A Model of Project-Based Learning in AI-Enabled Immersive Virtual Environments — Design creativity measured in an AI-enabled immersive project-based learning environment
Citation
Rahimi, S., Babaee, M., Esmaeiligoujar, S., & Dede, C. (2026). Generative artificial intelligence and creativity in K–12 education: A systematic scoping review. PsyArXiv Preprints. In press in R. A. Beghetto (Ed.), The Oxford handbook of AI and creativity in education (Oxford University Press).