Research Article
Constructing Epistemic AI Literacy: Detecting Epistemic Aims and Processes in Student-AI Co-Programming
Synthesis: Epistemic thinking — understanding how knowledge is constructed and justified — plays a central role in AI Literacy, particularly when students co-program with generative AI. This paper introduces a framework for detecting epistemic aims and processes in Student Experience during programming activities. The analysis reveals that students engage in question construction, AI output evaluation, and solution integration as distinct epistemic processes. These findings inform Scaffolding design for programming education and connect to broader discussions of Agentic Education with AI Coding Assistants where students maintain agency while leveraging AI assistance.
Key Findings
- The paper introduces the conceptual framework of Epistemic AI Literacy (EAIL), reframing AI literacy as a process-oriented epistemic phenomenon that emerges through dynamic human-AI interactions, drawing on the AIR framework of epistemic aims, ideals, and reliable epistemic processes.
- Using a large dialogue dataset of human-AI co-programming, the study identifies observable dimensions of epistemic aims (mastery-oriented aims) and epistemic processes (outsourcing, explanation seeking, verification seeking, prompt monitoring, and epistemic justification).
- A subset of interactions was manually annotated to ground the constructs, which then informed scalable automatic labeling using complementary approaches — few-shot prompting and regex-based scripts — applied interactively.
- Results reveal a prevalent lack of EAIL: 78.8% of student-GenAI interactions relied on non-mastery-oriented aims and less reliable epistemic strategies such as outsourcing and verification-seeking.
- Only 11.1% of interactions showed high epistemic engagement, where mastery-oriented aims were coupled with advanced strategies like epistemic justification in a more reliable epistemic process.
- While GenAI facilitates task success, robust epistemic performance and genuine learning rarely emerge without deliberate instructional and design support.
Study Design & Method
The study operationalizes epistemic constructs that are normally hard to observe. Epistemic aims and processes were detected in student-AI co-programming interaction data, with manual annotation of a subset grounding the constructs. Complementary automated approaches — few-shot prompting with large language models and regex-based scripts — were then used interactively to label the full dataset at scale, providing a path from small-scale qualitative insight to large-scale measurement. The design responds to a limitation identified in a 2022 UNESCO report: AI education has typically taken a technology-oriented approach, ignoring the human and in-depth ethical questions of how AI is actually used in practice.
What this means for practice
- Instructors. Design for epistemic aims rather than task completion: only 11.1% of interactions coupled mastery-oriented aims with advanced strategies such as epistemic justification, so access to GenAI did not by itself produce learning-oriented use.
- Instructors. Prompt the epistemic moves directly — construct a question, evaluate the AI output, justify why the answer is integrated — because 78.8% of interactions relied on outsourcing and verification-seeking.
- Designers. Instrument the five observable strategies (outsourcing, explanation seeking, verification seeking, prompt monitoring, epistemic justification) so that assessment can surface which epistemic process a student is actually running.
- Faculty developers. Move AI Literacy curricula from tool mechanics to epistemic practice: teach students how to decide what to trust and why, not only how to operate a model.
- Researchers. Use the framework's detection pipeline as a measurement instrument for Student Experience at scale rather than relying on small hand-coded samples.
Limitations
- The analysis is a secondary analysis of the StudyChat corpus: 200 complete co-programming chat sessions (about 2,000 prompts) randomly sampled from one undergraduate AI course at a single large U.S. research university, so the 78.8% and 11.1% figures describe that course and context.
- Only 499 turns were manually annotated to establish the gold standard, and the scalable labeling that produced the full-dataset percentages was validated against that small set; the authors call for comparing additional labeling approaches and multiple LLMs within the same pipeline to test robustness.
- The authors state the dataset's size and scope should be expanded as resources permit to enable stronger generalization and re-validation of the epistemic patterns.
- Epistemic aims are inferred from discourse markers and linguistic indicators in logged dialogue, so the constructs are operationalizations of observable talk rather than direct evidence of a learner's intent.
Citation
Mengqian Wu (2026). Constructing Epistemic AI Literacy: Detecting Epistemic Aims and Processes in Student-AI Co-Programming.