📄 Research Article
Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages
Synthesis: Roy & Roy (2026) argue that the infrastructure of AI — training corpora, tokenization, benchmarks, deployment architectures — systematically disadvantages speakers of underrepresented languages before a model is trained, reframing dataset scarcity as a structural barrier rather than an isolated technical limitation. Using Bengali as a case in AI-assisted education, they document four interlocking failures: a web-presence gap (<0.5% of global content for ~4% of the population), a 67:1 English↔Bengali training-token deficit, a tokenization penalty from the alphasyllabary script, and connectivity exclusion (36.5% rural vs 71.4% urban internet penetration). They position offline-first design as an equity-oriented infrastructure strategy. The work connects to Equity, Language Learning, and Digital Divide debates in educational AI.
Four Interlocking Infrastructure Failures
The paper identifies four structural barriers that compound to exclude underrepresented languages from AI-assisted education:
These failures reflect longstanding resource-allocation decisions, institutional priorities, and design defaults that did not center underrepresented languages in mainstream AI development.
Reframing Scarcity as Structure
The authors argue dataset scarcity should be understood as a structural barrier rather than an isolated technical limitation. They recommend treating offline-first design as an equity-oriented infrastructure strategy for AI-assisted education in low-connectivity environments, and outline directions for linguistics and AI research aimed at reducing these structural inequalities.
Connected Concepts
Connected Articles
Citation
Roy, A., & Roy, P. (2026). Structural silence: When AI infrastructure fails speakers of underrepresented languages. arXiv:2608.12278.