What Is BharatGen? Inside India’s Own AI Foundation Model


0

What Is BharatGen? Inside India’s Own AI Foundation Model

SevenFeeds Explained – we make complicated technology understandable

Meenakshi – SevenFeeds
Published: August 12, 2026
Updated: August 12, 2026

Quick answer: BharatGen is India’s government-backed national AI initiative – a consortium led by IIT Bombay building foundation models for text, speech, and documents, trained from scratch on Indian languages instead of adapted from English-first models like GPT or Gemini. Its models are released as open-weight technology, meaning Indian startups and government bodies can use them without paying for API access. It has received roughly Rs 1,223.6 crore in government backing – including Rs 988.6 crore as the top beneficiary of the IndiaAI Mission 2025 – and in March 2026 signed a deal with L&T to build its own AI chips and data centers in India.

Table of Contents

  1. The Problem BharatGen Was Built to Solve
  2. What Exactly Is BharatGen?
  3. Meet BharatGen’s Model Family
  4. Bharat Data Sagar: The Data Behind the Models
  5. Real Applications Already Live
  6. Who’s Behind BharatGen
  7. The BharatGen-L&T Deal: Building India’s Own AI Chips
  8. BharatGen vs Sarvam AI vs Krutrim vs Bhashini
  9. How to Access BharatGen’s Models
  10. Strengths and Honest Limitations
  11. What Comes Next
  12. People Also Ask FAQs
  13. Final Verdict

The Problem BharatGen Was Built to Solve {#problem}

Type a question in Hindi into ChatGPT, Gemini, or Claude and you’ll get a competent answer. But the model behind it was trained overwhelmingly on English text, then adapted to handle other languages afterward. That approach has a real cost: global models generate several times more tokens for the same sentence in Hindi than in English, making every Hindi conversation slower and more expensive to run. They also miss the code-mixing – switching between Hindi and English mid-sentence

  • that defines how hundreds of millions of Indians actually type and speak.

India has more than 20 official languages and over 100 dialects. A country where the large majority of the population is more comfortable outside English cannot build its digital future entirely on top of AI models designed around English first.

That is the gap BharatGen exists to close – not by fine-tuning a Western model to sound more Indian, but by training foundation models from the ground up on Indian languages, data, and cultural context.

What Exactly Is BharatGen? {#what-is}

BharatGen is India’s sovereign AI stack – a national consortium led by IIT Bombay, built from scratch on Indian data, with models that work across text, speech, and vision, released so startups, governments, and companies can build on top of them.

The Ministry of Science & Technology officially launched BharatGen in 2024 as a generative AI initiative designed to enhance public service delivery, built around four defining features: multilingual and multimodal models, Bhartiya dataset-based training, an open-source platform, and a generative AI research ecosystem for India. It is, by design, the world’s first government-funded Multimodal Large Language Model project built specifically for Indian languages.

The initiative began with an initial Rs 235 crore allocated by India’s Department of Science and Technology in 2024 under the National Mission on Interdisciplinary Cyber-Physical Systems (NM-ICPS). It gained serious momentum on September 18, 2025, when the Ministry of Electronics and Information Technology awarded a further Rs 988.6 crore under the IndiaAI Mission – announced by Union Minister Ashwini Vaishnaw at an event in New Delhi. That single allocation made BharatGen the foremost beneficiary of the entire Rs 1,500 crore IndiaAI Mission 2025 round, bringing total government backing to approximately Rs 1,223.6 crore.

That funding isn’t just for maintaining existing models – it is earmarked to build BharatGen’s next generation of Large Language and Multimodal Models with up to one trillion parameters, alongside smaller specialised models, trained on advanced supercomputing clusters.

On November 7, 2025, the project took a structural step most academic research initiatives never take: it incorporated. The BharatGen Technology Foundation was registered as a Section 8 not-for-profit company with the Registrar of Companies in Mumbai, headquartered on IIT Bombay’s Powai campus – making IIT Bombay the first Indian academic institution to directly own and operate an AI-focused company.

BharatGen’s profile has risen internationally too – it was referenced by Prime Minister Narendra Modi at the AI Summit in Paris as an example of India’s push into sovereign generative AI.

Meet BharatGen’s Model Family {#models}

<!– IMAGE 2 –> <!– File: bharatgen-model-family-2026.jpg | 1200x600px –> <!– Alt: BharatGen model family 2026 – Param Patram Shrutam Sooktam compared by parameters and use case –> <!– Canva: clean infographic, 5 model cards in a row, each showing model name, parameter count badge, and one-line use case, saffron and navy palette –>

Unlike a single chatbot product, BharatGen is building a full stack of specialised models under its official product line – Patram, DocBodh, Bharat Data Sagar, the BharatGen Text Model, Sooktam, Shrutam, Krishi Sathi, and e-VikrAI.

Model

Parameters

What It Does

Param-1

2.9 billion

Bilingual Hindi-English text generation, pretrained on 5 trillion tokens, built to handle code-mixing and Indian cultural references

Param2-17B-A2.4B-Thinking

17 billion total, 2.4 billion active (Mixture-of-Experts)

BharatGen’s larger, reasoning-focused successor – high capacity while keeping inference costs closer to a much smaller model

Patram

7 billion

India’s first document-vision model – reads and understands Indian forms, IDs, certificates, and handwritten documents

Shrutam

Compact ASR model

Automatic speech recognition trained on real-world noisy Hindi audio – call centres, busy streets

Sooktam

Compact TTS model

Text-to-speech across multiple Indic languages including Marathi, Bengali, Tamil, Telugu, Punjabi, Gujarati, and Malayalam

Why Param-1’s size is the point, not a limitation. At 2.9 billion parameters, Param-1 is deliberately smaller than frontier models like GPT-5 or Gemini 3.1 Pro. It was pretrained on 5 trillion tokens across English and Hindi, built to be genuinely fluent in how India speaks, cheap enough to run at population scale, and small enough that Indian startups can actually afford to deploy it.

Param-2 is where BharatGen gets ambitious. Officially released as Param2-17B-A2.4B-Thinking and unveiled at the AI Impact Summit 2026, this model uses a Mixture-of-Experts architecture – meaning it has 17 billion total parameters but activates only around 2.4 billion of them for any given query. The “Thinking” in its name signals a focus on reasoning ability, not just fluency – and with government funding now targeting models up to a trillion parameters, Param-2 looks like a stepping stone rather than BharatGen’s ceiling.

Patram solves a problem no global model was built for. Indian administrative documents – ration cards, land records, handwritten government forms – don’t look like the clean digital documents most document-AI models are trained on. Patram was trained specifically on these formats, which is why it powers DocBodh, BharatGen’s own document question-answering product for Indian citizens.

Bharat Data Sagar: The Data Behind the Models {#data}

A foundation model is only as good as what it’s trained on – and this is where BharatGen’s approach differs most from adapting a Western model. Bharat Data Sagar is BharatGen’s own initiative to build the world’s largest dataset focused specifically on underrepresented Indian data: text, speech, and images tied to India’s languages, culture, history, and philosophy.

The scale is genuinely large – the project targeted over 15,000 hours of annotated voice data across 22 Indian languages by the end of 2025, covering everything from rural dialects to urban conversational speech. This is the raw material that makes models like Shrutam usable in noisy, real-world Indian environments rather than only in clean lab conditions.

Alongside the dataset work, BharatGen runs an upskilling arm – funding MTech and PhD researchers, running AI courses, and hosting hackathons aimed at building India’s own AI talent pipeline rather than only its models.

Real Applications Already Live {#applications}

BharatGen isn’t only a research project sitting in papers and GitHub repos – it already powers working applications, several of which have been demonstrated directly to government ministers:

Krishi Sathi – India’s first farm bot, a voice-enabled WhatsApp advisory tool that lets farmers ask agricultural questions in their own language and get instant spoken answers, no typing required.

e-VikrAI – an AI assistant that automatically generates product descriptions from a single photo, helping small Indian sellers list products online without writing marketing copy themselves.

DocBodh – a document question-answering platform powered by Patram that lets ordinary citizens ask questions about complex official documents in plain language and get an understandable answer back.

MahaGPT – launched with MITRA and IIT Bombay at the AI Impact Summit 2026, MahaGPT enables AI-powered search across roughly 150,000 government resolutions for the Maharashtra state government, aimed at strengthening governance and citizen participation.

These examples share a pattern: none of them are chatbots for tech-savvy urban professionals. They are infrastructure for the “last mile” – farmers, small sellers, and citizens navigating paperwork and government records – exactly the audience global AI products consistently underserve.

Who’s Behind BharatGen {#who}

BharatGen is anchored by IIT Bombay but structured as a genuine national consortium. Academic partners include IIT Madras, IIT Kanpur, IIT Hyderabad, IIT Mandi, IIT Kharagpur, IIT Delhi, IIM Indore, and IIIT Hyderabad, alongside close collaboration with Bhashini, the government’s language-data initiative.

Leadership sits with two people: Professor Ganesh Ramakrishnan of IIT Bombay, the Principal Investigator, who has publicly framed the mission as building models that “sound Indian, think Indian, and work reliably in Indian environments,” and Rishi Bal, Executive Vice President, who bridges the academic research side with enterprise-scale execution.

On the industry side, BharatGen has built partnerships that extend well beyond government funding:

  • IBM – a formal collaboration announced in September 2025 to combine IBM’s model-training and governance technology with BharatGen’s Indic data and models, including deployment templates on IBM Watsonx and Red Hat OpenShift AI
  • L&T Semiconductor Technologies and L&T-Vyoma – a March 2026 MoU to build sovereign AI chips and data center infrastructure (more below)
  • Zoho and NASSCOM – India-based enterprise software and industry partners
  • State and central government bodies – including Maharashtra (MahaGPT) and the Ministry of Water and Sanitation, for citizen-facing applications

The BharatGen-L&T Deal: Building India’s Own AI Chips {#compute}

<!– IMAGE 3 –> <!– File: bharatgen-lt-sovereign-compute-2026.jpg | 1000x500px –> <!– Alt: BharatGen L&T MoU 2026 – India sovereign AI chips and data center infrastructure explained –> <!– Canva: three-pillar infographic – chip icon, data center icon, brain/model icon – labelled Silicon, Infrastructure, Models, navy and saffron palette –>

In March 2026, BharatGen made a move that goes well beyond software. It signed a landmark Memorandum of Understanding with L&T Semiconductor Technologies and L&T-Vyoma (Larsen & Toubro’s data center and cloud services arm) to jointly design, build, and deploy a complete sovereign AI compute platform for India.

The partnership rests on three pillars:

Indian AI silicon. L&T Semiconductor Technologies will design and develop custom AI ASIC and xPU chips optimised specifically for BharatGen’s workloads – language models and multimodal systems built to run efficiently on India-designed hardware rather than exclusively on imported chips.

Sovereign AI infrastructure. L&T-Vyoma will provide AI-ready data center infrastructure, including an upcoming 30 MW data center facility in Kanchipuram, to support large-scale compute needs.

Foundational AI models. BharatGen will define and co-optimise the actual AI workloads – LLMs, smaller language models, and multimodal systems – tailored to run on this new infrastructure.

The signing ceremony was attended by Ajay Kumar Sood, Principal Scientific Adviser to the Government of India, underlining that this is being treated as a matter of national technological significance, not just a corporate partnership. Taken together with its models and data, this MoU pushes BharatGen toward a genuinely complete sovereign AI stack – silicon, infrastructure, and software, designed and built in India.

BharatGen vs Sarvam AI vs Krutrim vs Bhashini {#comparison}

<!– IMAGE 4 –> <!– File: india-sovereign-ai-comparison-2026.jpg | 1200x600px –> <!– Alt: BharatGen vs Sarvam AI vs Krutrim vs Bhashini comparison India sovereign AI models 2026 –> <!– Canva: 4-column comparison table graphic, each column headed with model/company name, navy and saffron palette –>

BharatGen isn’t building India’s sovereign AI stack alone. It sits alongside three other major players, each with a different focus:

 

BharatGen

Sarvam AI

Krutrim

Bhashini

Built by

IIT Bombay consortium (govt-backed)

Ex-AI4Bharat researchers, backed by Peak XV & Lightspeed

Ola’s Bhavish Aggarwal

Government of India

Focus

Open public infrastructure, now including compute

Enterprise voice + API infra

Consumer Hindi/Hinglish chat

Free language translation

Cost to use

Open-sourced subset, free

Paid API

Paid API / app

Free

Best for

Startups & govt building their own tools

B2B developers needing production voice AI

Consumer-facing Hindi apps

Budget-conscious builders, NGOs, students

Notable model

Param-1, Param2-17B-A2.4B-Thinking

Sarvam-105B (open-sourced Feb 2026)

Krutrim-2 (open-sourced)

22-language translation engine

2026 status

Rs 1,223.6cr govt-backed; building own chips with L&T

Raising $300-350M at $1.5B valuation

India’s first AI unicorn (early 2024)

Government infrastructure, continuously expanded

The honest performance picture: on Indian-language benchmarks, BharatGen and Sarvam AI are genuine achievements – Sarvam’s largest model reportedly beats GPT-4, Claude, and Gemini in roughly 90% of head-to-head Indian-language comparisons. But on general-purpose global reasoning benchmarks, neither comes close to frontier models. That’s not a failure – it’s a different design goal entirely. These models exist to solve India-specific problems cheaply and at scale, not to compete with OpenAI or Anthropic on raw intelligence.

Where BharatGen specifically stands apart: it is the only one of the four simultaneously government-owned, releasing open-source models, and now building its own chip and data center supply chain through the L&T partnership. Sarvam and Krutrim are venture-backed companies that need to eventually charge for access and rely on imported compute like everyone else. BharatGen’s bet is that foundational AI infrastructure for a country of 1.4 billion people should be sovereign at every layer – not just the model weights.

How to Access BharatGen’s Models {#access}

If you’re a developer, researcher, or startup founder wanting to actually use BharatGen’s models rather than just read about them, here’s how:

Step 1 – Visit the Hugging Face repository. BharatGen publishes its open-sourced models at huggingface.co/bharatgenai, including Param-1 and Param2-17B-A2.4B-Thinking.

Step 2 – Pick the right model for your use case. Param-1 or Param-2 for bilingual Hindi-English text generation. Patram for processing Indian forms and documents. Shrutam for Hindi speech recognition. Sooktam for text-to-speech output.

Step 3 – Fine-tune on your own data. Because the open-sourced releases are open-weight, you can fine-tune them for a specific domain – a healthtech startup, for instance, could fine-tune Param-1 on medical consultation transcripts in Hindi.

Step 4 – Deploy through IBM Watsonx or your own servers. Thanks to the IBM collaboration, BharatGen models can be deployed via IBM Watsonx and Red Hat OpenShift AI, or self-hosted since no licensing fee applies to the open-sourced models.

Strengths and Honest Limitations {#limitations}

What BharatGen genuinely gets right:

  • Open-sourced models at zero licensing cost – a real advantage for cash-strapped Indian startups and government departments
  • Trained on Indian data from the ground up, backed by its own Bharat Data Sagar dataset initiative
  • A complete multimodal stack – text, speech, and document vision – that only a handful of countries have built end to end
  • Real deployed applications (Krishi Sathi, MahaGPT, DocBodh), not just benchmark papers
  • Now actively building domestic AI chips and data centers through the L&T partnership, directly addressing India’s import dependency

Where it honestly still falls short:

  • Param-1 at 2.9 billion parameters is not close to frontier-model reasoning ability, though Param-2 and the trillion-parameter roadmap aim to close that gap
  • The L&T sovereign compute platform is a March 2026 MoU, not yet operational infrastructure – it will take time before India-designed chips are actually training BharatGen’s models at scale
  • BharatGen is transitioning toward enterprise licensing revenue alongside its government funding – a transition that will decide how much this ecosystem can grow independent of further public money
  • Compared to Sarvam AI’s polished commercial API products, BharatGen’s open-weight models require more in-house engineering effort to actually deploy

None of this makes BharatGen unimportant. It makes it exactly what it claims to be: essential, unglamorous infrastructure – closer to a national highway system than a flashy new car.

What Comes Next {#next}

Three things to watch over the next 12-18 months:

Whether the trillion-parameter models materialise. The Rs 988.6 crore IndiaAI Mission allocation was explicitly earmarked to fund models up to one trillion parameters. Whether BharatGen can actually train and ship something at that scale – not just announce it – is the single biggest thing to watch.

The L&T sovereign compute platform going live. Custom AI chips and a 30 MW data center in Kanchipuram don’t appear overnight. 2026 and 2027 will show whether this MoU turns into working India-designed silicon actually training Indian AI models.

Increasing convergence with Sarvam, Krutrim, and vertical models. India’s AI ecosystem is layering – foundation models like BharatGen and Sarvam at the base, applied vertical models like MahaGPT and Tech Mahindra’s education-focused Project Indus building on top. Expect more shared infrastructure between these players rather than pure competition, since they’re all ultimately trying to solve the same national problem: reducing India’s dependency on foreign frontier AI for its digital future.

People Also Ask FAQs {#faqs}

What is BharatGen?

BharatGen is India’s national sovereign AI initiative – a consortium led by IIT Bombay building foundation models for text, speech, and documents trained specifically on Indian languages and data, operating through the not-for-profit BharatGen Technology Foundation with backing from India’s Department of Science and Technology and the IndiaAI Mission.

Who built BharatGen and who leads it?

BharatGen is led by Professor Ganesh Ramakrishnan of IIT Bombay as Principal Investigator, alongside Rishi Bal as Executive Vice President. Consortium partners include IIT Madras, IIT Kanpur, IIT Hyderabad, IIT Mandi, IIT Kharagpur, IIT Delhi, IIM Indore, and IIIT Hyderabad.

How much funding has BharatGen received?

Rs 235 crore from the DST in 2024, plus Rs 988.6 crore from MeitY under the IndiaAI Mission in September 2025 – making BharatGen the top beneficiary of that Rs 1,500 crore funding round. Total backing is approximately Rs 1,223.6 crore, partly earmarked to build models up to one trillion parameters.

What models has BharatGen released?

Param-1 (2.9B parameters, trained on 5 trillion Hindi-English tokens), Param2-17B-A2.4B-Thinking (17B total, 2.4B active, Mixture-of-Experts), Patram (7B parameters, document and vision understanding), Shrutam (speech recognition), and Sooktam (text-to-speech across Indic languages).

Is BharatGen building its own AI chips?

Yes. In March 2026, BharatGen signed an MoU with L&T Semiconductor Technologies and L&T-Vyoma to build custom AI chips and a sovereign data center platform, including a 30 MW facility in Kanchipuram – moving toward a complete India-built AI stack from silicon to software.

Is BharatGen better than Sarvam AI or Krutrim?

They serve different purposes. Sarvam AI focuses on production-grade voice and API infrastructure for businesses. Krutrim focuses on consumer Hindi and Hinglish chat. BharatGen focuses on open, government-backed public infrastructure – and is now the only one of the three also building its own domestic chip and data center supply chain.

Can developers use BharatGen’s models for free?

Yes, for the open-sourced subset. Param-1 and Param2-17B-A2.4B-Thinking are available on Hugging Face to download, fine-tune, and deploy, with deployment support through BharatGen’s collaboration with IBM Watsonx and Red Hat OpenShift AI.

Final Verdict {#verdict}

BharatGen matters not because it will out-perform GPT-5 or Gemini on a leaderboard – it was never trying to. It matters because it answers a question most global AI coverage never asks: what happens to the other 80% of the internet’s languages when the entire AI industry optimises for English first?

For India specifically, BharatGen is a bet that sovereign, open, publicly funded AI infrastructure can do for language technology what public highways did for commerce – unglamorous, essential, and built to be used by everyone, not owned by anyone. The March 2026 deal with L&T to build domestic chips and data centers shows this bet is expanding beyond software into the physical infrastructure underneath it.

Whether it succeeds depends less on model architecture and more on something far less technical: whether India can turn over Rs 1,200 crore of public investment – and now a domestic chip supply chain – into a self-sustaining ecosystem before the grant money runs out. 2026 and 2027 are the years that answer starts to show.

<!– RELATED ARTICLES NOTE: This is Article #1 on the rebuilt site, so there are no existing SevenFeeds articles to link to yet. Add this section once these two follow-up articles are published (both were suggested as strong Article #2/#3 candidates): -> How Many AI Startups Does India Have in 2026? -> What Is the IndiaAI Mission GPU Pool? Explained Once live, add a “Related on SevenFeeds” block linking to both, per the internal linking rule (1 pillar + related cluster articles). –>

Sources:

  • BharatGen official site (bharatgen.com) – products, funding announcement, L&T MoU
  • Drishti IAS / PIB – original 2024 launch announcement
  • IIT Bombay – How IIT Bombay Is Powering India’s AI Revolution
  • The Tribune – Jitendra Singh Hails BharatGen as India’s First Sovereign LLM
  • IBM Newsroom – IBM and BharatGen Collaboration Announcement
  • arXiv – PARAM-1: BharatGen 2.9B Model (2507.13390)
  • Rest of World / explainx.ai – India’s sovereign AI ecosystem, 2026

Tags: what is bharatgen, bharatgen ai india, bharatgen foundation model, iit bombay bharatgen, param-1 bharatgen, bharatgen vs sarvam ai, bharatgen vs chatgpt, india sovereign ai model, indiaai mission, bharatgen funding, bharatgen l&t

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ SECTION 4 — IMAGE PROMPTS ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

IMAGE 1 – HERO (1200x630px) File: what-is-bharatgen-hero.jpg Alt: What is BharatGen – India’s sovereign AI foundation model explained Canva: dark navy background, glowing India map outline made of small circuit-node dots, bold white text “What Is BharatGen?”, subtext “India’s Own AI Foundation Model, Explained”

IMAGE 2 – MODEL FAMILY (1200x600px) File: bharatgen-model-family-2026.jpg Alt: BharatGen model family 2026 – Param Patram Shrutam Sooktam compared by parameters and use case Canva: 5 model cards in a row, each with name, parameter badge, one-line use case, saffron and navy palette

IMAGE 3 – L&T SOVEREIGN COMPUTE (1000x500px) File: bharatgen-lt-sovereign-compute-2026.jpg Alt: BharatGen L&T MoU 2026 – India sovereign AI chips and data center infrastructure explained Canva: three-pillar infographic – chip icon, data center icon, brain/model icon – labelled Silicon, Infrastructure, Models

IMAGE 4 – COMPARISON TABLE (1200x600px) File: india-sovereign-ai-comparison-2026.jpg Alt: BharatGen vs Sarvam AI vs Krutrim vs Bhashini comparison – India sovereign AI models 2026 Canva: 4-column comparison graphic, navy and saffron palette

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ SECTION 5 — WORDPRESS PUBLISH CHECKLIST ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

CONTENT [ ] Focus keyword “what is bharatgen” in H1, first 100 words, 2+ H2s [ ] Meta title (54 chars) + meta description (159 chars) set [ ] Slug: what-is-bharatgen [ ] Category set to AI, tagged “SevenFeeds Explained” [ ] Word count: 3,700+ confirmed [ ] Comparison table + model table included [ ] FAQ section (7 Qs) matches FAQ schema exactly [ ] Funding figures consistent throughout: Rs 988.6cr (MeitY) + Rs 235cr (DST) = Rs 1,223.6cr total

SCHEMA [ ] Article schema pasted [ ] FAQ schema pasted [ ] HowTo schema pasted

IMAGES [ ] All 4 images created in Canva and uploaded [ ] Alt text matches exactly as specified [ ] Hero set as Featured Image + Open Graph image

TECHNICAL [ ] Canonical URL set [ ] Author: Pranay Sharma, linked to author page [ ] Published + Updated dates visible

POST-PUBLISH [ ] Submit to Google Search Console -> Request Indexing [ ] Share on LinkedIn – the “India building its own AI chips” (L&T deal) is the strongest, freshest hook for a founder/tech audience [ ] Come back and add the Related Articles block once Article #2 and #3 are live (see note in Section 3)




Like it? Share with your friends!

0

What's Your Reaction?

hate hate
0
hate
confused confused
0
confused
fail fail
0
fail
fun fun
0
fun
geeky geeky
0
geeky
love love
0
love
lol lol
0
lol
omg omg
0
omg
win win
0
win
sevenfeeds.7

0 Comments

Your email address will not be published. Required fields are marked *