An analysis of 310 AI-generated lesson plans (2,230 individual activities) produced by ChatGPT (GPT-4o), Gemini (1.5 Flash), and Copilot (GPT-4 based) for all 53 Massachusetts eighth-grade civics standards. Each standard received two prompts: a basic "write a lesson plan" and a "highly interactive" variant.
Large-scale empirical evaluation of AI-generated civics lesson plans reveals that without teacher revision, AI tools overwhelmingly produce lower-order thinking activities and monocultural content โ fundamentally at odds with the goals of civic education.
The Study
Trust et al. (2025) analyzed 310 AI-generated lesson plans (2,230 individual activities) produced by ChatGPT (GPT-4o), Gemini (1.5 Flash), and Copilot (GPT-4 based) for all 53 Massachusetts eighth-grade civics standards. Each standard received two prompts: a basic "write a lesson plan" and a "highly interactive" variant.
Key Findings
Lower-Order Thinking Dominates
Using Bloom's Revised Taxonomy:
| Level | Share |
|---|
| Remember | 45% |
| Understand | 21% |
| Apply | 24% |
| Subtotal (lower-order) | 90% |
| Analyze | 4% |
| Evaluate | 2% |
| Create | 4% |
90% of activities demanded only recall, comprehension, or simple application. Activities like "write definitions," "list three facts," and "answer comprehension questions" were pervasive. Even prompting for "highly interactive" lessons made minimal difference.
Near-Total Absence of Multicultural Content
Using Banks' Four Levels of Integration of Multicultural Content:
94% of activities contained no discernible multicultural content (2,086 of 2,230).Of the 144 activities that did, 137 were at the lowest "Additive" level (mentioning diverse figures without restructuring curriculum).Only 1 activity reached "Transformation" (restructuring the curriculum to include diverse perspectives).Zero activities reached "Social Action" (empowering students to address social issues).This is especially damning for civic education, where multicultural perspectives and critical engagement with power structures are essential learning goals.
Formulaic Outputs Across All Chatbots
All three chatbots produced structurally identical lesson plans: Introduction โ Activities 1-4 โ Conclusion โ Assessment โ Extension โ Homework. This factory-line format was applied regardless of whether the standard addressed constitutional principles, civil rights, or local government โ homogenization that strips away the disciplinary texture of civic education.
Implications for AI in Education
The "Trust But Verify" Mandate
This study provides concrete evidence for why AI literacy for teachers is not optional โ it's a prerequisite. AI tools reliably produce plausible-looking but pedagogically impoverished lesson plans. Teachers must:
1. Recognize the pattern of lower-order thinking bias.
2. Inject higher-order activities (analysis, evaluation, creation).
3. Add multicultural perspectives the AI omits.
Connection to Broader AI Alignment Problems
This finding parallels Hardy & Kim's Educational LLM Alignment โ AI tools may appear competent (producing well-formatted lesson plans) while failing at the intended impact (fostering critical civic thinking). The homogenized output reflects shared pretraining patterns that embed narrow pedagogical assumptions.
The Teacher's Role Is Enhanced, Not Replaced
Far from making teachers obsolete, these results reinforce the critical oversight role of educators. AI can generate drafts, but human judgment is essential for:
Elevating cognitive demand beyond recall/application.Integrating multicultural and critical perspectives.Adapting plans to specific classroom contexts and student needs.This aligns with evidence that teacher prompting instruction can improve AI output quality โ but only when teachers understand what to look for.
The Civic Education Context Matters
Civic education is a uniquely high-stakes domain for AI application because:
It explicitly aims to develop critical thinking about power, justice, and democracy โ skills AI tools systematically suppress in their default outputs.Multicultural content is not a "nice to have" but a core learning objective.Formulaic lesson structures undermine the domain's inherent demand for perspective-taking and deliberation.Open Questions
Would fine-tuned educational LLMs (e.g., EduQwen) produce more cognitively demanding and multiculturally-aware lesson plans?How do these findings generalize to other subjects (math, science, language arts)?Can better prompt engineering (e.g., explicitly requesting higher-order thinking and multicultural integration) close the gap?What does the teacher revision process look like in practice โ do teachers have the time and training to meaningfully redesign AI outputs?Connected Concepts
AI LiteracyAutomated Question GenerationFormative AssessmentRegulationHuman In The Loop AIK 12 AI EducationLLM Cultural Relevance K12Pedagogical LLM TrainingTeacher AI CompetencyK 12Teacher RoleConnected Articles
Educational LLM Alignment โ Educational LLM AlignmentNsmq Riddles Science Math Benchmark โ NSMQ Riddles: A Benchmark of Scientific and Mathematical Riddles for Quizzing Large Language ModelsAaai2026 Prompting Literacy K12 โ Learning to Use AI for Learning: Teaching Responsible Use of AI Chatbot to K-12 Students Through an AI Literacy ModuleAccess Not Enough AI Tutoring 2026 โ Access is Not Enough: Human Support Improves Engagement with AI TutoringAdapt Adaptive Lesson Plan Transformer โ AdaPT: Adaptive Lesson Plan Transformer for Cross-Regional and Differentiated InstructionAdaptive Pretesting Retention โ Do Gains from Generative AI-Enabled Adaptive Pretesting Persist? Evidence from a Retention StudyAgency Gap AI Writing โ The agency gap in AI-supported writing: how reactive and proactive agent designs shape multimodal reasoningAgent Voice Accents K12 Group Learning โ Exploring How Agent Voice Accents Shape Human-AI Collaboration in K-12 Group LearningAgentic AI Education Scoping Review โ Agentic AI in Education: A Scoping Review of Research Landscape, Capabilities, and the Frontier Agent ParadigmAgentic AI Pedagogical Best Practice 2026 โ Agentic AI and Pedagogical Best Practice: The Tension Between Automation and LearningAgentic Literacy Debt โ Agentic Literacy Debt: A Structural Problem the AI Literacy Field Has Not Yet NamedAgreement Not Quality LLM Coding Verification โ Agreement Is Not Quality: Blind Expert Verification of Human and LLM Qualitative Coding When Human Consensus Is Not G...AI Adoption Training Public Sector โ The Main Barrier to AI Adoption in the Public Sector is Lack of TrainingAI Assessment Human Tutors โ AI-Driven Assessment of Human Tutors: Linking Training Performance to Real-Life PracticeAI Assessment Scale Reform โ A bit of chaos and madness": The AI Assessment Scale and the work of assessment reformAI Assistance Discretionary Feedback โ AI Assistance for Discretionary Work: Increasing Feedback Provision in Higher EducationAI Assisted Learning Modes Eeg โ An exploratory behavioral and electroencephalographic study of artificial intelligence-assisted learning modes in hig...AI Changing Teaching Workflows โ How AI Is Changing Teaching WorkflowsAI Education Global Capacity โ What AI in Education Needs Next: Lessons from Youth Leaders Across Five CountriesAI Engineering Education Balancing Act โ Using AI in engineering education: a balancing act, driven by clear purposeAI Ethics Education Public Discourse โ A Longitudinal Analysis of Public Discourse on AI Ethics in Education Using Twitter DataAI Fatigue Academic Contexts โ Defining AI Fatigue in Academic Contexts: Dimensions, Indicators, and a Stage-Based Model Using Grounded TheoryAI Generated Feedback Higher Ed โ Artificial intelligence and feedback in university education: effectiveness and student perceptionsAI Generated Slides Student Perception โ AI-Generated Slides: Are They Good? Can Students Tell?AI Higher Ed Bridge Gap โ Higher Education Must Bridge the AI GapCitation
(2025), A.T.T.M.R.X.C.P.K., 25(3), J.C.I.I.T.A.T.E., Name].", I.A.H.I.L.F., |, B.L.T.A.O., levels, O.A.A.A.R.U.O.A., & |, B.L.T. (2026). AI-Generated Lesson Plans in Civic Education