Research Article
A didactical-driven teacher assistant for a dimensional modeling course
Synthesis: Brisson, Segarra and Smits present a didactically-driven Large Language Models (LLMs) teacher assistant for a university dimensional modeling (data warehousing) course. Unlike most educational chatbots that delegate pedagogical decisions implicitly to the LLM, their system makes content selection and didactic structuring explicit and traceable: tutoring strategy is encoded in an external didactic layer that the LLM executes, so tutoring behavior can be evaluated and reproduced. The design responds directly to the opacity critique raised in Rethinking Scaffolding in LLM Tutors: The Interactional Mismatch Between Benchmarks and Real-World Deployments and complements retrieval-grounded designs such as Retrieval-Augmented Tutoring for Algorithm Tracing and Problem-Solving in AI Education and safety-layered tutors like EduGuard: A Safe RAG-Based LLM Tutor for Programming Education. As an instructor-facing pedagogical agent it sits alongside TeachingCoach: A Fine-Tuned Scaffolding Chatbot for Instructional Guidance to Instructors, and its explicit didactic structuring exemplifies principled Learning Design applied to LLM tutoring in CS Education.
What this means for practice
- Instructors. Write down the intent-to-didactic-approach mapping you already apply when answering students, then encode it as the system's routing layer — the paper's architecture keeps content selection and didactic structuring outside the LLM so tutoring behavior can be traced and reproduced.
- Do not rely on raw semantic retrieval to surface definitions and comparisons: on the 195 authentic questions, EXPLAIN queries retrieved at roughly 50% while COMPARE fell to 7.5%, so resolve intent and concept before retrieving content.
- Treat abstention as a recovery step rather than a system failure, since a wrong intent propagates silently into a pedagogically incorrect answer while a refusal triggers a reformulation request — the pattern detector correctly refused 37 out-of-scope or social queries.
- Designers. Evaluate intent detection, concept resolution, and retrieval as separate modules, which is what makes each failure mode diagnosable instead of hidden behind a single end-to-end quality score.
- Expect a lexicon-based detector to cover only a fraction of authentic questions: the pattern baseline answered 32% of queries, and adding an LLM fallback cut LLM calls by about a third while lowering joint detection accuracy below the LLM alone.
Limitations
- Evaluation is technical only: 195 authentic student questions gathered from 24 students in a single French-language dimensional modeling course, with no measure of response quality, perceived usefulness, or learning outcomes.
- The pattern detector reached only 32% coverage on the real corpus, and the combined pattern+Gemini configuration scored a pair F1 of 63.50, significantly below Gemini-2.0-flash alone at 67.81 (p = 0.005) — the pattern introduced errors on queries the LLM handled correctly.
- The retrieval corpus is mono-authored, drawn entirely from one instructor's teaching notes, so system coverage is bounded by that single source.
- The system is reactive and atomistic, with no learner model, no awareness of the course timeline, and no capacity to initiate interaction, and its transfer to courses built on theorem derivation, algorithmic processes, or code production is untested.
Citation
Laurent Brisson, Maria Segarra, Grégory Smits (2026). A didactical-driven teacher assistant for a dimensional modeling course.