Core Contribution
Boateng et al. (2026) introduce NSMQ Riddles, a benchmark of 1.8K scientific and mathematical riddles drawn from 11 years of Ghana's National Science and Maths Quiz โ a live TV competition for senior secondary school students. This is one of the first AI benchmarks originating from the Global South for educational evaluation.
Why It's Distinctive
Unlike standard benchmark datasets (MMLU, GSM8K), NSMQ Riddles:
- Features progressive clue revelation โ early clues are vague (worth more points), testing incremental reasoning
- Covers biology, chemistry, physics, and math at the high school level
- Evaluates models against human student performance in a competitive format
- Represents African educational content, addressing geographic bias in ai-k12-evidence-base
The benchmark found that even state-of-the-art models (GPT-5.4, Gemini 3.1 Pro, Claude Opus 4.6) underperform the best student contestants, highlighting gaps in LLM scientific reasoning.
Connections to Wiki
This benchmark connects to teachbench-llm-teaching-evaluation as another syllabus-grounded evaluation framework, but from a Global South perspective. It complements the educational-vlm-evaluation work on DrawEduMath by providing a text-based STEM reasoning benchmark. The focus on competitive quizzing connects to automated-question-generation research and civic-education-ai-lesson-plans concerns about AI-generated educational content quality.
The finding that LLMs lag behind top human students on these riddles reinforces tutoring-specific-vs-general-ai concerns โ general LLMs may not match specialized educational needs, especially in non-Western contexts.
Open Questions
- How well do pedagogical-llm-training approaches like EduQwen perform on NSMQ compared to general LLMs?
- Can the benchmark be extended to other African and Global South educational systems?
- What does the clue-progression format reveal about LLM reasoning vs. retrieval?