On this page

Synthesis: Tabletop exercises (TTXs) let learner teams rehearse high-stakes workplace tasks such as cybersecurity incident response, but their open-ended, collaborative nature makes Formative Assessment difficult: teams often receive delayed or incomplete feedback. This full research-to-practice paper compares assessment methods that exploit the action and communication logs captured by TTX platforms to evaluate how well teams meet learning objectives.

The work situates team Problem Solving assessment within CS Education and broader STEM Education contexts, showing how logged interaction data can drive faster, richer Feedback Loops than manual grading. By operationalizing teamwork behaviors as measurable signals, it connects to Learning Analytics and the Student Experience of collaborative crisis-response training, with implications for Higher Education computing courses where TTXs are increasingly used but rubric reliability remains a barrier.

What this means for practice

  • Learners. Treat your platform logs as part of the assessment record: the study scored teams from logged milestone IDs and timestamps, so whether and when the team reaches each milestone in the exercise is evidence in its own right.
  • Learners. Ask for cluster-level feedback, which lets instructors give one round of assessment-based advice to several comparable teams; the analytics ran in under 2 minutes on a standard laptop.
  • Learners. Check how your written incident-response communication was scored against the shared rubric, which human assessors applied as a three-dimensional score vector (2 = satisfied, 1 = partially satisfied, 0 = not satisfied).
  • Learners. Expect process to be judged, not one right answer: both exercises (EXF and PHI) are open-ended collaborative tasks, and the assessment methods target how a team deviates from effective practice.
  • Learners. Ask for human review of any LLM-based assessment of your work: the authors recommend instructor oversight of automated decisions for accountability and fairness, and the local clustering keeps your data from being shared with an external service.

Limitations

  • The sample is small and domain-specific: 36 cybersecurity students at one Czech university in the EXF exercise and 11 teams in the Estonian PHI exercise, 76 participants across 24 teams after one team declined consent, and the authors say findings generalize reliably only to learners who already have basic cybersecurity knowledge.
  • Automated assessment was restricted to specific LLMs (GPT-4o in the pilot, GPT-5.2 in the validation study), and because the models' internals are inaccessible the authors cannot determine why LLM scores sometimes differed from human scores - a limitation of all black-box methods.
  • LLM-based assessment depended on querying an external service, so an unavailable model or a changed version would alter the assessment's validity and reliability, unlike the clustering, which ran locally.
  • Clustering features came only from activity logs (milestone IDs, timestamps, action sequences and tool use) and exclude communication and external factors, so the method indicates similarity of team process rather than quality of outcome.

Citation

Valdemar Švábenský, Jan Vykopal, Sukrit Leelaluk, Pavel Čeleda, et al. (2026). Assessment in Team Problem-Solving Exercises in Computing Education. .

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.