Research Article
Evidence-Grounded Multimodal Knowledge Graph Construction for Multi-Lecture Educational Reasoning
Synthesis: This paper introduces an evidence-grounded Multimodal AI pipeline that constructs provenance-rich knowledge graphs from lecture videos by integrating speech transcripts, slide OCR, and vision-language model analysis. Processing three neural-network lectures, the pipeline extracted 172 canonical concepts and 282 typed relationships with 90.38% endpoint coverage, achieving perfect retrieval accuracy on a preliminary test. The approach addresses a key challenge in educational AI: converting rich multimodal lecture content into structured, queryable knowledge representations without losing the evidential provenance that makes them trustworthy.
Key Findings
- The pipeline fuses speech transcription, slide and diagram OCR, and vision-language analysis into a single evidence-grounded workflow, so every extracted item is tied to lecture evidence rather than inferred.
- Across three neural-network lectures it processed 3,118 frames, 756 transcript segments, and 559 semantic anchors, retaining 1,022 concept mentions and 312 relationship mentions.
- Canonicalization collapsed those mentions into 172 canonical concepts and 282 typed relationships, achieving 90.38% relationship endpoint coverage.
- A preliminary three-question retrieval test reached 100% top-1 and top-3 accuracy and 100% mean top-5 recall, though the authors frame it as a sanity check rather than a benchmark.
- Validation and confidence thresholds (0.55 minimum) plus evidence-pool checking filter out unsupported claims, making the resulting graph inspectable and auditable rather than a black-box extraction.
Motivation
Lecture videos distribute knowledge across speech, slide text, diagrams, equations, and presentation order, which transcript-only retrieval does not fully preserve. Students integrate these signals to connect definitions, examples, and prerequisites. Automated lecture question answering therefore needs a representation that retains concepts, relationships, evidence, and temporal context rather than only flat transcript chunks. Conventional retrieval-augmented generation grounds language models in external context but is weakest on explicit dependencies, concept evolution, and visually communicated information. A knowledge graph instead represents concepts as nodes and typed relations as edges while retaining the lecture, timestamp, frame, and evidence quotation that justify each extraction. This auditability matters in education because unsupported claims can mislead learners, and instructors need to inspect or correct extracted knowledge.
Pipeline Architecture
The multimodal pipeline processes lecture videos through several stages:
- Transcription: Speech-to-text conversion of lecture audio, with each segment aligned to a primary frame via its midpoint
- Semantic Anchor Selection: Identification of key concept-bearing segments, targeting about 18% of frames with temporal spacing to suppress near-duplicates
- OCR Extraction: Text extraction from slide content and diagrams, capturing labels, annotations, symbols, and equations that may not be spoken
- Vision-Language Analysis: Concept and relationship extraction with evidential grounding under a constrained JSON prompt
- Validation and Canonicalization: Cross-referencing mentions against multiple evidence sources and merging duplicates
- Knowledge Graph Construction: Typed relationships with provenance tracking in a NetworkX MultiDiGraph
Grounded Extraction and Validation
For every semantic anchor, the vision-language model receives the frame, a transcript window spanning 15 seconds before through 22 seconds after the anchor timestamp, OCR text, and locally derived candidate terms. It returns strict JSON arrays for concepts and relationships, where empty arrays are permitted. A concept records a name, definition, evidence quotation, source modality, and confidence; a relationship records source and target concepts, one of eight permitted types (prerequisite_of, component_of, uses, optimizes, computed_by, example_of, contrasts_with, or related_to), evidence quotation, source modality, and confidence.
Validation removes missing or low-information fields, unsupported evidence, invalid relation types, and claims below a 0.55 confidence threshold. Transcript- and OCR-sourced claims must occur in the evidence pool, while visual-only concepts require stronger confidence. This validation improves auditability but, as the authors note, does not by itself guarantee factual correctness — reducing the risk of a structurally plausible but unsupported graph.
Canonicalization and Graph Construction
Validated mentions are merged using aliases, normalized and fuzzy string matching, token overlap, and embedding similarity. Each canonical node retains its identifier, display name, aliases, definitions, mention list, lecture coverage, evidence count, and average confidence. Relationship endpoints are mapped only after concept deduplication, which prevents otherwise-valid edges from being lost when raw endpoint names differ from canonical names. The result is a graph whose edges carry relation type, lecture, timestamp, anchor, evidence quotation, and confidence, with multiple edge instances allowed between the same nodes when supported by different evidence.
Retrieval and Question Answering
Canonical concept text combines name, definition, aliases, and evidence snippets and is embedded using BGE-large English. The retrieval score combines semantic similarity, exact name and alias matches, fuzzy similarity, and an evidence-count prior. The top six concepts are expanded to a one-hop subgraph, and the answer prompt receives only formatted definitions, relationships, and evidence, requesting lecture identifiers and timestamps when available. This graph-grounded retrieval contrasts with video question answering systems that reason over a fixed clip without an explicit intermediate representation, which cannot easily aggregate evidence for a single concept across multiple, separately recorded lectures.
Results
Across the three lectures, the model returned 1,155 raw concept and 400 raw relationship mentions; evidence and confidence checks retained 1,022 concepts (88.48%) and 312 relationships (78.00%). Canonicalization produced 172 concepts, and endpoint mapping retained 282 edges (90.38% of validated relationships). The reduction from 1,022 mentions to 172 nodes is expected because course concepts recur across anchors and lectures. Central training concepts dominate the evidence distribution — weights, loss function, neuron, bias, and gradient descent carry the most evidence — though remaining singular and plural variants reveal conservative but incomplete entity resolution. Definition, relation, prerequisite, and cross-lecture trial questions ranked their target concepts first and within the top three, though generated answers sometimes added correct background knowledge not explicitly supported by retrieved evidence.
What this means for practice
- Designers. Require an evidence quotation, frame, or OCR string for every extracted concept and relation, and expose it: the pipeline retained 1,022 of 1,155 raw concept mentions and 312 of 400 relationship mentions through evidence and confidence checks, and that provenance is what lets an instructor inspect or correct a wrong extraction.
- Designers. Set an explicit confidence floor and hold visual-only claims to a stricter bar than transcript- or OCR-backed ones: the 0.55 minimum plus evidence-pool checking is credited with reducing the risk of a structurally plausible graph that no source supports.
- Designers. Track endpoint coverage as a build diagnostic: canonicalization collapsed 1,022 mentions into 172 concepts, and the 90.38% endpoint retention after deduplication is what distinguishes relation losses caused by merging from losses caused by extraction.
- Designers. Use the graph as the structural layer under adaptive and student-modeling features: typed prerequisite_of edges are the dependency structure a sequencing or diagnosis component reads, and the evidence attached to each edge keeps that layer inspectable.
- Instructors. Use the typed relations for cross-lecture review and sequencing: edges carry relation type, lecture, and timestamp, so prerequisite_of and component_of edges give learners an evidence-backed path rather than a flat transcript search.
Limitations
- The dataset is three lectures from a single neural-network series (3,118 frames, 756 transcript segments, 559 semantic anchors), with no cross-course or cross-domain evidence.
- Extraction has no manually annotated concept-and-relation gold standard, so the three-question retrieval result (100% top-1 and top-3) is a sanity check the authors say cannot support statistical claims.
- Canonicalization stays incomplete — singular and plural variants such as weight/weights and bias/biases survive — and isolated or noisy nodes may encode generic terms, numeric labels, or visual artifacts.
- OCR and vision-language accuracy depend on frame resolution, handwriting, slide transitions, and diagram complexity, and answer generation sometimes adds correct background knowledge that the retrieved evidence does not support.
Citation
Al Farib, S., Meem, M. A., Islam, S. R., & Raihan, M. T. (2026). Evidence-Grounded Multimodal Knowledge Graph Construction for Multi-Lecture Educational Reasoning. v1.