📄 Research Article
NSMQ Riddles: A Benchmark of Scientific and Mathematical Riddles for Quizzing Large Language Models
Boateng et al. (2026) introduce NSMQ Riddles, a benchmark of 1.8K scientific and mathematical riddles drawn from 11 years of Ghana's National Science and Maths Quiz — a live TV competition for senior secondary school students. This is one of the first AI benchmarks originating from the Global South for educational evaluation.
NSMQ Riddles: Educational Benchmark from Ghana
Core Contribution
Boateng et al. (2026) introduce NSMQ Riddles, a benchmark of 1.8K scientific and mathematical riddles drawn from 11 years of Ghana's National Science and Maths Quiz — a live TV competition for senior secondary school students. This is one of the first AI benchmarks originating from the Global South for educational evaluation.
Why It's Distinctive
Unlike standard benchmark datasets (MMLU, GSM8K), NSMQ Riddles:
The benchmark found that even state-of-the-art models (GPT-5.4, Gemini 3.1 Pro, Claude Opus 4.6) underperform the best student contestants, highlighting gaps in LLM scientific reasoning.
Connections to Wiki
This benchmark connects to Teachbench LLM Teaching Evaluation as another syllabus-grounded evaluation framework, but from a Global South perspective. It complements the Educational Vlm Evaluation work on DrawEduMath by providing a text-based STEM reasoning benchmark. The focus on competitive quizzing connects to Automated Question Generation research and Civic Education AI Lesson Plans concerns about AI-generated educational content quality.
The finding that LLMs lag behind top human students on these riddles reinforces Tutoring Specific Vs General AI concerns — general LLMs may not match specialized educational needs, especially in non-Western contexts.
Open Questions
Connected Concepts
Connected Articles
Citation
al, A.G.B.N.I.S.J.E., and, N.R.A.B.O.S., Large, M.R.F.Q., Models, L., Yeboah3,4, P.A.J.A.M.K.T., and, W.E.A.K.M.N.S.Y., Kumbol2,3, V., & Zurich, E. (2026). NSMQ Riddles: A Benchmark of Scientific and Mathematical Riddles for Quizzing Large Language Models