On this page

Synthesis: SlidesQAQA is a Flask-based system that extracts text and rendered images from PDF lecture slides and processes them through a four-stage Large Language Models (LLMs) pipeline: window planning (segment extraction), deck synthesis (cross-slide reasoning), slide annotation (per-slide question generation), and reconciliation (deck-level revision to reduce redundancy and improve coverage). The key innovation is joint reasoning about slide modality and pedagogical role, with a bounded question budget that forces prioritization of important content.

How It Works

Unlike earlier Automated Question Generation systems that generate questions slide-by-slide in isolation, SlidesQAQA reasons across the entire presentation. This enables deck-level scaffolding — questions build on each other across the slide sequence, matching the intended instructional flow. The reconciliation stage filters non-instructional slides and revises draft annotations to eliminate redundancy, producing structured JSON output with deck-level goals, section structure, slide summaries, question sets, and evaluation scores.

This approach contrasts with Generate-Then-Validate: Question Generation for Education frameworks by front-loading pedagogical reasoning rather than post-hoc validation. Where AI-Generated Slides: Are They Good? Can Students Tell? research has shown that AI-generated slide content can be perceived as lower quality, SlidesQAQA focuses on question quality rather than slide generation itself. It also differs from AISSA: AI-based Student Slides Analysis Tool for Academic Presentations systems that analyze slides for Accessibility rather than pedagogical question extraction.

Pedagogical Design

The bounded question budget per slide forces the system to make pedagogical decisions about what content merits a question — an implicit form of Scaffolding that prioritizes key concepts. Initial experiments on two technical lecture decks demonstrated successful filtering of non-instructional slides and generation of pedagogically coherent questions for visually complex content. This has implications for Formative Assessment automation at scale.

What this means for practice

  • Instructors. Run an existing deck through the pipeline before rebuilding your question bank: the system zero-budgeted administrative and transition slides in the "Self-Attention and Transformers" deck, so non-instructional slides are filtered rather than quizzed.
  • Instructors. Review the generated items against your own learning goals before deploying them — the three metrics (Coverage, Fidelity, Scaffolding) are scored on a 1-5 scale by the pipeline itself, and scaffolding scores ranged from 3 to 5 rather than uniformly high.
  • Software developers. Build deck-level reasoning into question generation rather than generating slide by slide: the reconciliation stage is what reduces redundancy and balances coverage across the presentation.
  • Instructors. Keep the multimodal path — text plus rendered images — for mechanism-heavy slides, since the system grounded questions on visual evidence such as "horizontal arrows pointing left and right between the boxes" that text-only pipelines miss.
  • Learners. Expect generated questions to work best as foundational comprehension checks on complex material; the evaluated items functioned as prerequisites that progress logically through a deck rather than as exam-level assessment.

Limitations

  • The evaluation rests on only two technical lecture decks — "Self-Attention and Transformers" and "Neural Constituency Parsing" — drawn from NLP and deep learning courses, so generalization to other disciplines is untested.
  • Coverage, Fidelity, and Scaffolding were scored automatically by the system on a 1-5 scale from its own pipeline logs; the authors name comparison against human-authored question sets as future work, meaning no human rating or benchmark currently validates the scores.
  • The pipeline relies on a specific proprietary LLM (Gemini), which the authors flag as a reproducibility and long-term stability risk if the underlying API model is updated or deprecated.
  • The multi-pass architecture incurs substantial LLM inference latency and API costs when processing large decks.

Citation

Salsman, J. (2026). Slide Deck Q&A Quality Assurance App: A Multi-Stage Pipeline for Pedagogical Question Generation.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.