On this page

Synthesis: Deploying Large Language Models (LLMs) tutors in K-12 raises concerns around privacy, cost, and reliance on proprietary models, motivating small language models (SLMs) as an alternative. The authors introduce CSTutorBench, a Benchmark evaluating language models as CS tutors in VEX VR, a block-based robotics environment. It comprises 17 scenario-based questions scored against a pedagogical rubric grounded in tutoring and feedback research, using a human-in-the-loop LLM-as-judge pipeline. Across 11 models (4B–120B parameters), models handled surface-level criteria (vocabulary, tone) well but struggled with deeper pedagogical behaviors — especially avoiding answer leakage and engaging with student debugging histories. Model family and instruction-tuning predicted tutoring quality better than parameter count; a targeted prompt revision improved scores for 10 of 11 models.

What this means for practice

  • Designers. Benchmark candidate models on the target tutoring domain before deployment: across 11 models (4B–120B parameters), models met surface criteria such as vocabulary and tone but struggled to avoid answer leakage and to engage with student debugging histories.
  • Designers. Do not select by parameter count: an 8B model reached 77% while qwen3-coder at 30B scored 52%.
  • Designers. Spend a Prompt Engineering iteration before discarding a model: the rubric-grounded prompt revision improved 10 of 11 models, by 6.6 to 16.2 percentage points (mean 11.2).
  • Instructors. Ask for criterion-level results rather than aggregate scores, because the revision mainly lifted the four type-specific criteria while accuracy and actionability moved little.

Limitations

  • The benchmark contains only 17 questions, and builds_on_success is scored on a single question, which limits the precision of the per-criterion comparisons.
  • The revised prompt was written in response to the first prompt's weaknesses and tested on the same 17 questions with no held-out subset, so the reported gains may conflate improvement with overfitting.
  • Every item is single-turn, so the benchmark cannot capture the multi-turn dialogue of real tutoring, and no students were evaluated — higher rubric scores may not predict learning outcomes.
  • The automated judge (Claude Sonnet 4) showed instancing inconsistency, varying how it weighted or combined criteria across model-trial combinations.

Citation

Lane, H. C., & Kageler, B. (2026). CSTutorBench: Benchmarking Small Language Models as Tutors for Block-Based Programming.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.