IKS-Instruct: A 24,000-Example Multilingual Dataset for Teaching Language Models Indian Knowledge Systems

Created: 2026-07-31 | Tags: llmpersonalized-learningeducational-theorylanguage-learningopen-source

Shwetha Singaravelu, Gayathri Muruganantham, Lakshmi Rajendran, Santhosh Sivasubramani (2026) โ€” arXiv:2607.23322 (cs.CL, cs.CY)

๐Ÿ“„ Full text (arXiv)

Summary

Presents a 24,795-example multilingual instruction dataset for teaching LLMs to deliver educational content grounded in Indian Knowledge Systems. Spans seven languages and bridges a gap in non-Western pedagogical content for instruction tuning. Demonstrates that domain-specific educational datasets improve LLM performance on culturally grounded knowledge tasks.

The work connects to broader discussions in AI and education around llm, personalized-learning, multilingual-learning, contributing to our understanding of how llm shapes educational practice.

Key Contributions

Related Pages

Citation

APA: Shwetha Singaravelu, Gayathri Muruganantham, Lakshmi Rajendran, Santhosh Sivasubramani (2026). IKS-Instruct: A 24,000-Example Multilingual Dataset for Teaching Language Models Indian Knowledge Systems. arXiv:2607.23322. cs.CL, cs.CY.