NSMQ Riddles: A Benchmark of Scientific and Mathematical Riddles for Quizzing Large Language Models

Created: 2026-05-08 | Tags: benchmarkstem-educationk-12llmefficacy-study

Core Contribution

Boateng et al. (2026) introduce NSMQ Riddles, a benchmark of 1.8K scientific and mathematical riddles drawn from 11 years of Ghana's National Science and Maths Quiz โ€” a live TV competition for senior secondary school students. This is one of the first AI benchmarks originating from the Global South for educational evaluation.

Why It's Distinctive

Unlike standard benchmark datasets (MMLU, GSM8K), NSMQ Riddles:

The benchmark found that even state-of-the-art models (GPT-5.4, Gemini 3.1 Pro, Claude Opus 4.6) underperform the best student contestants, highlighting gaps in LLM scientific reasoning.

Connections to Wiki

This benchmark connects to teachbench-llm-teaching-evaluation as another syllabus-grounded evaluation framework, but from a Global South perspective. It complements the educational-vlm-evaluation work on DrawEduMath by providing a text-based STEM reasoning benchmark. The focus on competitive quizzing connects to automated-question-generation research and civic-education-ai-lesson-plans concerns about AI-generated educational content quality.

The finding that LLMs lag behind top human students on these riddles reinforces tutoring-specific-vs-general-ai concerns โ€” general LLMs may not match specialized educational needs, especially in non-Western contexts.

Open Questions

Related Pages