Publications and Research
Document Type
Article
Publication Date
Fall 8-10-2026
Abstract
Large language models are increasingly deployed as general-purpose educational and technical assistance systems, but their basic infrastructure does not treat languages equally. One underexamined source of disparity is tokenization: semantically equivalent content can require substantially different token counts across languages, affecting API cost, latency, and usable context length before a model is even invoked. We introduce the Tokenization Equity Audit (TEA), a reproducible benchmark for measuring tokenization premiums in technical tutoring content. TEA evaluates three widely used tokenizers, GPT-4o’s o200k base, Qwen2.5-7B, and Mistral-7B, on a 120-item Python debugging corpus translated from English into Bengali, Hindi, Arabic, Tamil, and Yoruba. Bengali and Hindi are the primary validated cases; the remaining languages provide exploratory cross-script and cross-family comparisons. Across this corpus, Bengali requires 1.56× as many GPT-4o tokens as English, reducing a nominal 128k-token context window to an effective 82k-token equivalent for the same semantic content. On Qwen2.5 and Mistral tokenizers, Bengali reaches 4.5× the English token count. Yoruba, despite using Latin script, shows the highest GPT-4o penalty at 2.37×, indicating that tokenization inequity is not reducible to script family alone. We demonstrate that tokenization creates measurable economic and functional barriers that must be addressed as an equity-relevant infrastructure layer for underserved language communities, especially where educational systems depend on low-cost or offline-capable AI tools.
Included in
Artificial Intelligence and Robotics Commons, Computational Linguistics Commons, Science and Technology Studies Commons

Comments
Presented orally and as a poster at the Language Models for Underserved Communities (LM4UC) Workshop, IJCAI-ECAI 2026, Bremen, Germany, August 16, 2026.
Preprint: arXiv:2608.09046
Benchmark data and code: github.com/HeyAvijitRoy/tea-benchmark
ORCID: Avijit Roy (0009-0007-8036-0952); Proma Roy (0009-0004-1060-9116); Hrishitva Patel (0000-0001-7887-6641)