Research Article
EduGage: Methods and Dataset for Sensor-Based Momentary Assessment of Engagement in Self-Guided Video Learning
Synthesis: EduGage (Leng et al., 2026) shows that fine-grained, segment-level engagement estimation in self-guided video learning is feasible but inherently noisy. Using wearable and camera-based sensors (PPG, ECG, EDA, EEG, IMU, heart rate, temperature, eye tracking) across a 16-participant user study, the authors' Multimodal AI model achieved an MAE of 0.81 and 83.75% within-1 accuracy, outperforming sensor-free, statistical, deep-temporal, foundation-model, and Large Language Models (LLMs)-based baselines. The authors argue that practical systems should prioritize lightweight combinations of behavioral and physiological signals over full multimodal instrumentation, and they release the EduGage dataset — synchronized multimodal sensor streams, probe-aligned momentary engagement labels, video metadata, quizzes, and study materials — to support reproducible research.
Key Findings
- Fine-grained engagement estimation is feasible but noisy. Segment-level engagement in self-guided video learning can be estimated from wearable sensing, but inherently noisy, so systems should treat predictions as approximate.
- Strong accuracy vs. baselines. Across participant-based cross-validation, the EduGage model achieves an MAE of 0.81, 83.75% within-1 accuracy, 73.93% binary accuracy, and 68.45% binary Macro-F1, beating sensor-free, statistical, deep temporal, foundation-model, and LLM-based baselines.
- Lightweight sensing wins for deployment. Practical systems should prioritize lightweight combinations of behavioral and physiological signals over full multimodal instrumentation, balancing predictive value with real-world feasibility.
- A reusable dataset. EduGage contains 16 participants, 64 video-viewing sessions, 715 probe-aligned windows, and ~12 hours of synchronized multimodal recordings, enabling reproducible research on sensor-based engagement modeling.
The Engagement Problem in Video Learning
EduGage (Leng et al., 2026) addresses a core challenge in online and video-based learning: learners must self-regulate their engagement with instructional materials without continuous instructor feedback. Engagement is widely understood as a multidimensional construct spanning behavioral, emotional, and cognitive components, and prior work consistently emphasizes the importance of cognitive engagement for meaningful learning because it reflects the mental effort learners invest in understanding, monitoring, and mastering material.
This is especially relevant in higher education and professional training, where prerecorded instructional videos are now common and learners must regulate their own attention, motivation, and understanding. Because engagement is a dynamic, partially latent process, it can shift substantially within a single lesson: learners may appear behaviorally attentive while mind-wandering, or show little outward response while still actively processing difficult material.
Dimensions of Engagement
| Dimension | Measurement | Relevance to Learning |
|---|---|---|
| Attentional | Eye tracking, gaze patterns | Sustained focus on content |
| Emotional | Facial expression, sentiment | Positive affect supports persistence |
| Cognitive | Physiological signals, task performance | Deep processing vs. superficial viewing |
Sensor-Based Momentary Assessment
Traditional engagement measures are limited:
- Post-hoc surveys: Retrospective bias, low temporal resolution
- Self-reports: Introspection difficulty, social desirability bias
- Behavioral proxies: Viewing time and interaction logs capture activity at scale but not momentary internal experience
EduGage contributes real-time sensor fusion for momentary assessment during video learning, combining repeated in-situ self-reports with fine-grained sensing. Following experience sampling and ecological momentary assessment methodology, the study operationalizes momentary cognitive engagement as the perceived difficulty of sustaining attention during the preceding segment, collected via a brief single-item probe on a 5-point Likert scale.
Technical Approach
- Sensors: Webcam (eye tracking), interaction logs, and wearable devices capturing PPG, ECG, EDA, EEG, IMU, heart rate, and temperature
- Assessment: Momentary (in-the-moment) vs. retrospective, aligned to ~44-second probe-aligned windows
- Feedback loop: Real-time reflection prompts based on engagement state
Study Design and the EduGage Dataset
The study recruited 16 college students who watched instructional videos from the MIT Open Learning Library across four domains (X-Rays, Aerospace Engineering, Environmental Science, Business), with A/B variant pairs and counterbalanced ordering. Participants wore multiple sensors concurrently — a PPG ring, Microsoft Band 2 (EDA/heart rate), an earable IMU, a Polar H10 chest strap (ECG), a Muse S Athena headband (EEG/PPG/IMU), and a webcam eye tracker — while periodically reporting their momentary attention difficulty at natural stopping points. Pre- and post-video quizzes provided a secondary learning-related measure to check whether higher engagement related to learning gains.
The resulting EduGage dataset provides temporally aligned wearable streams, repeated probe-level self-reports, video-segment metadata, and study materials at a fine temporal resolution — a resource suited for modeling how engagement fluctuates across specific moments of instruction.
The Multimodal Prediction Framework
Each probe-aligned window is a prediction sample. The framework processes each modality with its own temporal encoder, then uses a context-informed gated fusion mechanism: separate gating networks learn a soft contribution weight for each modality, conditioned on the modality embedding and a shared context vector (relative video progress). The weighted modality embeddings are fused via a normalized gated average and passed to a regression head that outputs the predicted engagement level. This design lets the model adaptively emphasize informative streams while down-weighting noisier or less reliable ones, reflecting the heterogeneous signal quality across wearable sensing channels.
Results and Sensing Tradeoffs
Across participant-based cross-validation, the model achieved an MAE of 0.81, 83.75% within-1 accuracy, 73.93% binary accuracy, and 68.45% binary Macro-F1, outperforming sensor-free, statistical, deep temporal, foundation-model, and LLM-based baselines. The key practical conclusion is that fine-grained estimation is feasible but noisy, and that full multimodal instrumentation may not be worth its deployment burden — lightweight combinations of behavioral and physiological signals offer a better balance of predictive utility and real-world feasibility.
Connection to Adaptive Learning
This enables adaptive interventions in video learning:
- Detect disengagement (gaze diversion, prolonged pauses, attention difficulty)
- Trigger scaffolds (reflection prompt, content re-summarization, revisiting a segment)
- Close loop: Learner reflects → re-engages → improved outcomes
This aligns with Adaptive Learning principles: real-time learner modeling → personalized intervention. Momentary sensing can support self-directed learning, student modeling, and post-hoc content refinement without replacing the teacher role.
What this means for practice
- Software developers. Prioritize lightweight signal combinations — ring PPG, EEG, IMU, and EDA carried most of the usable signal here — over a full multimodal rig, since the full setup buys accuracy at an instrumentation and Privacy cost that rarely survives deployment. This is the same trade-off seen in other multimodal work, such as Multimodal Dialogue in STEM Education and Affective Tutoring.
- Instructors. Use segment-level attention-difficulty estimates to find where within a video attention drops, and trigger a just-in-time response (reflection prompt, content re-summarization, or replay of that segment) instead of acting on coarse session-level metrics.
- Designers. Present engagement estimates as approximations and pair low-confidence readings with a self-regulation prompt, because even the best model reached an MAE of 0.81 on a 5-point scale and the authors describe fine-grained estimation as inherently noisy.
- Researchers. Build on the released EduGage dataset — 16 participants, 64 video-viewing sessions, 715 probe-aligned windows, and approximately 12 hours of synchronized recordings — and its participant-based splits as a reproducible Benchmark for sensor-based engagement modeling.
Limitations
- The study recruited 16 college students from one university through mailing lists and online postings, so the cohort is small, single-site, and not representative of other learner populations.
- Data were collected in a controlled laboratory session using four instructional videos from the MIT Open Learning Library; the authors state that extending the approach to larger, longer-term, and more naturalistic environments such as homes and libraries remains future work.
- The prediction target is a single self-report item ("How difficult was it to pay attention during the last part of the lecture?") on a 5-point Likert scale that captures only the attentional dimension of engagement, and the authors note that exact self-reported scores remain difficult to predict.
- All results come from participant-based 4-fold cross-validation over 715 probe-aligned windows of 44 seconds each rather than a deployed system; an LLM few-shot baseline reaches only 46.86% binary accuracy on the same task, showing how little signal a text-only approach extracts.
Citation
Leng, Z., Eyal, E., Shi, Y., He, J., Liu, Y., & Plötz, T. (2026). EduGage: Methods and Dataset for Sensor-Based Momentary Assessment of Engagement in Self-Guided Video Learning