On this page

Synthesis: Boateng et al. (2026) introduce NSMQ Riddles, a benchmark of 1.8K scientific and mathematical riddles drawn from 11 years of Ghana's National Science and Maths Quiz — a live TV competition for senior secondary school students. This is one of the first AI benchmarks originating from the Global South for educational evaluation.

Core Contribution

Boateng et al. (2026) introduce NSMQ Riddles, a benchmark of 1.8K scientific and mathematical riddles drawn from 11 years of Ghana's National Science and Maths Quiz — a live TV competition for senior secondary school students. This is one of the first AI benchmarks originating from the Global South for educational evaluation.

Why It's Distinctive

Unlike standard benchmark datasets (MMLU, GSM8K), NSMQ Riddles:

  • Features progressive clue revelation — early clues are vague (worth more points), testing incremental reasoning
  • Covers biology, chemistry, physics, and math at the high school level
  • Evaluates models against human student performance in a competitive format
  • Represents African educational content, addressing geographic bias in The Evidence Base on AI in K-12: A 2026 Review

The benchmark found that even state-of-the-art models (GPT-5.4, Gemini 3.1 Pro, Claude Opus 4.6) underperform the best student contestants, highlighting gaps in Large Language Models (LLMs) scientific reasoning.

Connections to Knowledge Base

This benchmark connects to TeachBench - Evaluating LLM Teaching Ability as another syllabus-grounded evaluation framework, but from a Global South perspective. It complements the The Aftermath of DrawEduMath: Vision Language Models Underperform with Struggling Students and Misdiagnose Errors work on DrawEduMath by providing a text-based STEM reasoning benchmark. The focus on competitive quizzing connects to Automated Question Generation research and Civic education in the age of AI: Should we trust AI-generated lesson plans concerns about AI-generated educational content quality.

The finding that LLMs lag behind top human students on these riddles reinforces The Evidence Base on AI in K-12: A 2026 Review concerns — general LLMs may not match specialized educational needs, especially in non-Western contexts.

Open Questions

  • How well do Training Pedagogical LLMs for Tutoring approaches like EduQwen perform on NSMQ compared to general LLMs?
  • Can the benchmark be extended to other African and Global South educational systems?
  • What does the clue-progression format reveal about LLM reasoning vs. retrieval?

What this means for practice

  • Developers. Do not read benchmark-topping scores as instructional readiness: GPT-5.4 with high reasoning effort led the offline benchmark at 86.54% EM accuracy, yet in the real-time proxy the best LLM result of 75.64% EM accuracy and 366 points still trailed the best student teams at 78.21% EM accuracy and 377 points.
  • Developers. Build practice items with progressive clue revelation, where earlier clues are vaguer and worth more (5 points on the first, 4 on the second, 3 thereafter), so that exercises reward incremental reasoning rather than a single retrieval step.
  • Developers. Evaluate K-12 STEM reasoning against a human baseline and on Global South content: NSMQ Riddles draws 1.8K riddles from 11 years of Ghana's National Science and Maths Quiz, a coverage that general benchmarks such as MMLU and GSM8K do not provide.
  • Developers. Instrument how many clues a model needs as well as whether it answers correctly, since accuracy alone hides the difference between recognizing an answer early and arriving at it only after most of the riddle has been revealed.

Limitations

  • The real-time proxy evaluation used only one year of the NSMQ (2019) because annotations of the required metadata and compute resources were limited; the authors list more years as future work.
  • That evaluation is a proxy rather than a true competition simulation: annotated audio of the contests was unavailable, and actual points depend on whether the student or the model answers first in a live round.
  • The riddles appear publicly on YouTube, so contamination of model training data is possible; the authors state they did not assess it and list de-contamination analysis as future work.
  • The student comparison uses retrospective real-world team performance rather than matched conditions, and only 156 riddles carried the metadata needed for the points analysis.

Citation

Boateng, G., Ibrahim, N. D., John, S., et al. (2026). NSMQ Riddles: A Benchmark of Scientific and Mathematical Riddles for Quizzing Large Language Models.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.