Research Article
Pedagogical Promise and Peril of AI: A Text Mining Analysis of ChatGPT Research Discussions in Programming Education
Synthesis: > Across 69 Scopus-indexed studies on ChatGPT in programming education, a text mining analysis of term frequencies, phrase patterns, and LDA topic models reveals a persistent dual framing: ChatGPT is cast simultaneously as a learning aid that supports explanation, feedback, and efficiency and as a pedagogical risk linked to overreliance, unreliable outputs, and academic integrity concerns. The literature concentrates overwhelmingly on classroom practice and learner engagement (nearly half the corpus), while Assessment design, prompting, and institutional governance remain comparatively underexplored. The authors conclude that ChatGPT's benefits—motivation, self-efficacy, computational thinking, faster grading—materialize mainly under structured pedagogy and teacher facilitation, and that responsible integration demands clearer policies, authentic assessment practices, and equitable access.
Key Findings
- Text mining of 229 retrieved documents (69 after deduplication and screening) yields four dominant themes: pedagogical implementation, student-centered learning and engagement, AI infrastructure and human-AI collaboration, and assessment, prompting, and model evaluation.
- ChatGPT is consistently framed as both a productive learning aid (explanation, feedback, code generation, efficiency) and a pedagogical risk (overreliance, cognitive dependency, unreliable or hallucinated outputs, integrity concerns).
- The corpus skews toward classroom practice and learner experience, with comparatively limited attention to systematic assessment design, prompt design, and institutional governance.
- Controlled, structured applications report gains in motivation, engagement, self-perceived competence, computational thinking, and grading efficiency (e.g. ChatGPT-4 grading at 0.91 correlation with instructors; ~75% grading-time reduction), whereas unmoderated use is associated with reduced persistence, weaker independent debugging, and limited genuine learning-gain.
- The paper recommends course-level AI policies, verification procedures (code walkthroughs, oral assessment), faculty AI literacy and prompt competence, and equitable infrastructure as preconditions for responsible adoption.
The Dual Character of ChatGPT in Programming Education
The chapter opens by situating ChatGPT against a long history of AI support for computer science education. Early intelligent tutoring systems and automated assessment platforms delivered adaptive feedback and improved performance on topics such as loops, recursion, and data structures, but were limited in scale and struggled with complex, open-ended programming tasks. Large language models like ChatGPT redirected this pursuit: they produce explanations, examples, and code corrections through natural conversation, functioning as tutor, debugging assistant, and feedback tool. The chapter's central claim is that this technology carries a dual character—simultaneously a scaffold for learning and a source of academic and ethical challenges—and that this duality is visible across the research literature itself.
Method: Text Mining the Corpus
The empirical core is a computational analysis of published discourse. The dataset began as 229 documents from Scopus (open-access journal articles and conference papers matching a ChatGPT-and-programming-education query), reduced through deduplication and relevance screening to 69 documents. Each text was preprocessed (lowercasing, removal of punctuation and digits, tokenization, stopword removal, Porter stemming) and then analyzed with three complementary procedures:
- Term frequency analysis to identify the dominant research concepts (student, ChatGPT, AI, education, and program rank highest);
- Phrase pattern analysis of bigrams and trigrams to surface conceptual relationships ("AI tools," "programming education," "Problem Solving skills," "AI-generated content"); and
- LDA topic modeling, configured with four topics after coherence iterations, to reveal underlying themes.
Each document was also reviewed manually to link computational patterns to the reported opportunities, challenges, and limitations.
Four Dominant Themes
The topic modeling surfaces four themes that organize scholarly discussion. Pedagogical Use and Classroom Implementation (19% of the corpus) emphasizes teacher facilitation, structured integration, and ethical awareness. Student-Centered Learning and Engagement (49%) portrays ChatGPT as an interactive tutor supporting motivation, engagement, and coding performance—the largest and most learner-focused cluster. AI Infrastructure and Human-AI Collaboration (23%) turns to institutional and technical dimensions such as transparency, readiness, and accountability. Assessment, Prompting, and Model Evaluation (9%) is the smallest cluster, addressing prompt design, feedback generation, and model accuracy. The distribution itself is a finding: the literature privileges the classroom and the learner while devoting the least attention to assessment and governance.
Benefits and Opportunities
Where ChatGPT is integrated through structured instructional frameworks, the reported benefits are substantial. The R5E model improved student performance and critical thinking; PyChatAI delivered real-time bilingual feedback that aided debugging; a mobile learning system enhanced motivation, Self-Efficacy, and coding accuracy; and a GPT-based code review system reduced academic dishonesty while improving feedback precision. On the instructor side, ChatGPT-4 graded submissions at a 0.91 correlation with human raters, the GreAIter system cut grading time by over 75% without losing accuracy, and ChatGPT-3.5 generated coherent exam questions that reduced preparation time. Personalized and adaptive applications—automated grading, fuzzy-memory feedback, Scaffolding in simulated environments—are credited with supporting accessibility and individualized learning.
Risks and Limitations
The risks cluster into four areas. Academic integrity and misuse: teachers regard ChatGPT as a major contributor to exam dishonesty; students who received assistance often copied inaccurate outputs; AI-content detectors performed poorly at distinguishing AI-generated from human code. Negative cognitive impacts: frequent reliance is associated with memorizing incorrect explanations, reduced motivation to solve problems independently, weaker software-testing, and limited reasoning when explaining AI-generated code—consistent with cognitive dependency. Technical unreliability: students encounter incomplete or incorrect code, one study found only 30% of ChatGPT outputs fully usable, and repeated errors and hallucinated responses are documented—tying directly to Hallucination Risk. Ethical and social concerns: plagiarism, Privacy, unequal access, reduced Creativity, and uncertainty about authorship all recur, alongside equity questions (including about support for learners with disabilities).
Solutions and Future Directions
The authors' recommendations center on course-level policies that distinguish guided learning from dishonesty, verification procedures (code walkthroughs, oral assessment), pedagogy that requires students to compare, critique, and justify AI-generated code, faculty training in AI literacy and prompt construction, and equitable infrastructure such as campus-wide licenses and accessibility features. Future research should move beyond short-term classroom experiments to longitudinal studies of how consistent exposure affects problem-solving, code quality, persistence, and higher-order skills such as abstraction and design thinking, and should compare instructional frameworks and cross-institutional/cross-cultural readiness.
What this means for practice
- Learners. Treat generated code as a draft to verify, not an answer to submit: in the reviewed studies students frequently encountered incomplete or incorrect code responses, and one study found only 30% of ChatGPT outputs fully usable.
- Learners. Attempt the problem yourself before prompting, and practice explaining what generated code does: the corpus links frequent reliance to reduced motivation to solve problems independently, fewer and less effective software tests, and limited reasoning when students account for AI-generated code (cognitive offloading).
- Instructors. Place ChatGPT inside a structured instructional framework instead of leaving it open: the controlled applications reviewed reported gains in performance and critical thinking, whereas unmoderated use was associated with weaker persistence and debugging.
- Instructors. Make students justify what they submit — code walkthroughs, oral assessments, and critique-and-compare tasks — because AI-content detectors performed poorly at distinguishing AI-generated from human-written code and cannot carry integrity checks on their own.
- Administrators. Write course-level rules and secure equitable access now rather than waiting for consensus: institutional governance and assessment are the thinnest clusters in the corpus, and campus-wide licenses plus accessibility features are the remedy the authors recommend against a digital divide.
Limitations
- The corpus is indexed and open-access, not the whole field: 229 Scopus records were reduced to 81 for initial review and then to 69 documents, and only open-access journal articles and conference papers were eligible.
- The themes rest on one analytic configuration — LDA was fixed at four topics after iterations for coherence, and each document was assigned to a single primary topic by its highest probability score — so a paper spanning assessment and classroom practice is counted only once.
- The reported cluster sizes are proportions of documents rather than of evidence: pedagogical implementation 0.19 (13 documents), student-centered learning and engagement 0.49 (34), AI infrastructure and human-AI collaboration 0.23 (16), and assessment, prompting, and model evaluation 0.09 (6).
- The analysis measures published discourse through term frequencies, bigram and trigram patterns, and topic models, so it cannot show whether structured ChatGPT use improved student outcomes; the performance figures it repeats (a 0.91 grading correlation with instructors, a grading-time reduction of over 75 percent) come from the primary studies it reviews, not from this analysis.
Citation
Grume et al. (2026). Pedagogical Promise and Peril of AI: A Text Mining Analysis of ChatGPT Research Discussions in Programming Education. Pedagogical Innovations in CS Education (IGI Global).