📣 Send us your press release
Site updates every 15 minutes
Technology

India must counter AI censorship and exclusion with fairer sovereign data

India faces a growing threat from AI censorship and data exclusion, necessitating the development of fairer, sovereign datasets. Limited linguistic resources and foreign models pose significant challenges.

3 September 2026
India must counter AI censorship and exclusion with fairer sovereign data
Image is an AI-generated illustration

India must address the increasing threat of artificial intelligence (AI) censorship and data exclusion by developing its own, fairer, and sovereign data resources. The rapid advancement of large language models (LLMs) risks perpetuating distorted or censored information, particularly in low-resource languages and concerning Indian subjects.

Many current LLM models are predominantly trained on Western or Chinese datasets, leading to the dominance of these perspectives. This can result in inadequate or inaccurate representations of Indian philosophies, such as Advaita Vedanta. Furthermore, text-to-image models have repeatedly depicted Indians using stereotypes, exacerbating exclusion.

State-sponsored censorship within datasets and the LLMs built upon them poses a significant risk to India's information ecosystem and sovereignty. As these models become more sophisticated and are deployed in critical sectors like defense and law enforcement, censored data could shape their operations and decision-making.

China has already invested heavily in building national, validated datasets, though these are strictly state-censored. While Western datasets also moderate content, they draw from a broader range of open sources. India needs to establish its own diverse and uncensored data repositories to ensure AI's equitable and accurate utilization.

Original source: medianama.com