🧠 AI Ed Wiki

Boateng et al. (2026) introduce NSMQ Riddles, a benchmark of 1.8K scientific and mathematical riddles drawn from 11 years of Ghana's National Science and Maths Quiz — a live TV competition for senior secondary school students. This is one of the first AI benchmarks originating from the Global South for educational evaluation.

NSMQ Riddles: Educational Benchmark from Ghana

Core Contribution

Boateng et al. (2026) introduce NSMQ Riddles, a benchmark of 1.8K scientific and mathematical riddles drawn from 11 years of Ghana's National Science and Maths Quiz — a live TV competition for senior secondary school students. This is one of the first AI benchmarks originating from the Global South for educational evaluation.

Why It's Distinctive

Unlike standard benchmark datasets (MMLU, GSM8K), NSMQ Riddles:

  • Features progressive clue revelation — early clues are vague (worth more points), testing incremental reasoning
  • Covers biology, chemistry, physics, and math at the high school level
  • Evaluates models against human student performance in a competitive format
  • Represents African educational content, addressing geographic bias in Stanford Evidence Base AI K12 2026
  • The benchmark found that even state-of-the-art models (GPT-5.4, Gemini 3.1 Pro, Claude Opus 4.6) underperform the best student contestants, highlighting gaps in LLM scientific reasoning.

    Connections to Wiki

    This benchmark connects to Teachbench LLM Teaching Evaluation as another syllabus-grounded evaluation framework, but from a Global South perspective. It complements the Educational Vlm Evaluation work on DrawEduMath by providing a text-based STEM reasoning benchmark. The focus on competitive quizzing connects to Automated Question Generation research and Civic Education AI Lesson Plans concerns about AI-generated educational content quality.

    The finding that LLMs lag behind top human students on these riddles reinforces Tutoring Specific Vs General AI concerns — general LLMs may not match specialized educational needs, especially in non-Western contexts.

    Open Questions

  • How well do Pedagogical LLM Training approaches like EduQwen perform on NSMQ compared to general LLMs?
  • Can the benchmark be extended to other African and Global South educational systems?
  • What does the clue-progression format reveal about LLM reasoning vs. retrieval?
  • Connected Concepts

  • Automated Question Generation
  • Pedagogical LLM Training
  • Connected Articles

  • Stanford Evidence Base AI K12 2026
  • Teachbench LLM Teaching Evaluation
  • Educational Vlm Evaluation
  • Civic Education AI Lesson Plans
  • Tutoring Specific Vs General AI
  • Citation

    al, A.G.B.N.I.S.J.E., and, N.R.A.B.O.S., Large, M.R.F.Q., Models, L., Yeboah3,4, P.A.J.A.M.K.T., and, W.E.A.K.M.N.S.Y., Kumbol2,3, V., & Zurich, E. (2026). NSMQ Riddles: A Benchmark of Scientific and Mathematical Riddles for Quizzing Large Language Models