Loading film...
0:00 / 6:16
Thought Leadership
AI Has a Huge Problem (India Might’ve Fixed It)
6:16
The same sentence takes seven tokens in English and twenty-four in Hindi. That is a design decision, not a bug. Here’s what one Bangalore startup did about it.
The company is Sarvam AI, and the fix starts at the tokenizer. The film explains why an English-weighted vocabulary fragments Indian scripts, what fertility measures, and how theirs lands near English efficiency where most multilingual models need four to eight tokens a word. It then covers the mixture of experts model built on top, the GPUs allocated under the India AI mission, and the open question of whether a technical edge survives a rival with a hundred times the distribution.