Speaker
Description
Modern Large Language Models (LLMs) inherently exhibit a profound architectural bias toward English and other Latin-script languages, inadvertently erecting a severe “script barrier” for the vast majority of the world’s linguistic
diversity. This barrier stems primarily from the ine{cient subword tokenization of non-Roman scripts, such as Devanagari, where standard algorithms aggressively fragment text into high-fertility sequences. This fragmentation not only drastically shrinks the ezective context window but also quadratically ampli}es the computational cost of self-attention. To circumvent this tokenization bottleneck, this research investigates romanization—the transliteration of native scripts into the Latin alphabet—as a highly efficient computational bridge. By aligning the input representation with the pre-existing orthographic strengths of English centric models, romanization serves as a pragmatic interface layer rather than a linguistic replacement, fundamentally mitigating the computational penalties imposed by standard tokenizers.
Our comprehensive empirical analysis, integrating a primary case study with }ndings from the ROMANSETU framework, demonstrates that romanizing Hindi text yields a consistent 2.5x to 4x reduction in token count. This efficiency directly translates to competitive or superior performance across a wide array of Natural Language Understanding (NLU) and Natural Language Generation (NLG) tasks, particularly in generative and knowledge-retrieval domains.
Furthermore, we formalize this computational overhead by deriving an attention
ampli}cation factor, revealing that native Devanagari processing requires over an order of magnitude more attention computation per unit of semantic content compared to its Romanized equivalent. We also systematically characterize the limitations of this pipeline, notably the risks of transliteration error propagation and the nuanced performance degradation on complex morpho-syntactic
reasoning tasks. Ultimately, while romanization provides a powerful and immediately deployable strategy for enhancing multilingual AI e{ciency, its necessity highlights the pressing, long-term requirement for fundamentally script-agnostic
tokenization and multilingual model architectures.
Keywords: Large Language Models, Romanization, Subword Tokenization, Devanagari, Token Fertility, Attention Ampli}cation, Multilingual NLP, Indic NLP
Any other info we should know?
Former Speaker OOSC 3.0 (IIT Kanpur)
Session author's bio
Yash Mishra ( National Institute of Technology Karnataka Surathkal ) ,
| In Person Attendance | In-person |
|---|---|
| Please confirm that there are included headshots of all speakers in their profiles | Yes |
| Social Media | @YashMis70584151 (Twittter ) |
| Level of Difficulty | Intermediate |
| Agree to Privacy Policy and Notice | I agree |