On this page

Synthesis: Artificial intelligence assistants deployed in online learning environments create new opportunities to collect large volumes of learner interaction data and generate insights to improve student outcomes. Architecture for AI-Augmented Learning (A4L) is a modular data architecture that enables the collection, integration, and analysis of learner interaction data from educational AI systems, supporting the generation of instructional insights that facilitate personalized learning and reinforce the bidirectional feedback loop between instructors and learners. This study examines the modular design of the A4L Data Analytics Pipeline, an extensible data infrastructure that enables the ingestion, processing, and analysis of heterogeneous datasets generated by educational AI assistants. We describe the design principles and development process used to extend the pipeline's analytical capabilities while preserving flexibility across domains. We evaluate the pipeline through case studies spanning three research domains corresponding to three educational AI assistants deployed in online learning environments at Georgia Tech.

  • Reusable analytics infrastructure: Bai et al. present the A4L Data Analytics Pipeline as modular, domain-agnostic infrastructure for analyzing learner interaction data from educational AI assistants. The pipeline is designed to ingest heterogeneous datasets across different courses and tutoring contexts without rebuilding analytics from scratch.
  • Cross-domain validation: The pipeline was evaluated through three case studies at Georgia Tech, each involving a different educational AI assistant deployed in real online learning environments. Results showed that a common set of statistical methods could be consistently applied across datasets with varying structures and instructional contexts.
  • Extensibility demonstrated: Analytical capabilities initially developed for one domain were successfully extended to support richer analyses in another domain, proving the pipeline's extensibility. This positions the A4L pipeline as reusable infrastructure for future Learning Analytics systems.
  • Bidirectional feedback loop: The architecture supports a feedback loop between instructors and learners, enabling Personalized Learning insights derived from AI-augmented learning data. This connects to broader conversations in Edtech Platform design about how analytics infrastructure should scale across domains.
  • EDULEARN26 publication suggests growing academic interest in systematizing analytics for educational AI, complementing work on Smaller, Younger, and More Impactful: How AI-Assisted Writing Transforms Research Teams and LLM-assisted sentiment analysis for integrated computational and qualitative mixed methods education research: A case study of students' written reflection assignments which explore different facets of AI-augmented education research.

What this means for practice

  • Instructors. Read the user-versus-non-user comparison for a course before judging whether an AI assistant worked: the replicated SAMI analysis found no bias in adoption and a higher sense of belonging among users than non-users, and the VERA analysis found adopters had significantly higher need-for-cognition scores.
  • Instructors. Check who the assistant reaches, not only whether use tracks with outcomes — the Jill Watson analysis of Fall 2023 and Spring 2024 data was framed around whether demographic factors influenced adoption and whether adoption affected course performance.
  • Researchers. Re-run one statistical option across new datasets by changing configuration values rather than code: a Welch's t-test was configured separately for JW, VERA, and SAMI, and a power calculation first applied to VERA was extended to SAMI through a new analysis payload.
  • Software developers. Add capability through analysis options and configuration payloads, and rely on the daily scheduled job that re-runs only the payloads whose datasets changed.
  • Administrators. Fund shared analytics infrastructure instead of per-course analysis scripts: one modular platform reproduced published findings from three different learning analytics deployments, each collected in a single graduate-level computer science course.

Limitations

  • The extension work was performed by research team members already familiar with the system's architecture; the authors name this as the study's one limitation and propose observing an outside researcher extend the pipeline as the evidence they still lack.
  • All three case studies replicate analyses of a single institution's existing data — graduate-level computer science courses at Georgia Tech, spanning Fall 2023 to Fall 2024 for JW and SAMI and a single Summer 2023 offering for VERA.
  • No new assistant's data was ingested: the analysis configuration is written using "XYZ" as a placeholder for a future assistant, so the generalizability claim rests on reanalysis of datasets that already existed in the A4L environment.
  • It is a system-development case study with no comparison condition and no student outcome measure; the demonstrated result is that three analyses were reproduced and one capability extended, not that the pipeline improves learning.

Citation

Yallen Bai, Ploy Thajchayapong, & Ashok Goel (2026). Generalizing a Highly Configurable Analytics Pipeline to Replicate and Support Educational Research Across Multiple Domains. EDULEARN26.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.