Foundation models and generative AI in India, explained
By Abha Lohia · Startup Decoded
A foundation model is a large AI system trained on huge amounts of data that can be adapted to many tasks, such as writing, translating or answering questions. Generative AI is the use of such models to create text, speech, images or code. In India the focus is on local languages, voice and low cost.
What is a foundation model?
A foundation model is trained once on very large amounts of text, speech or images, and then reused for many jobs. Instead of building a separate program to translate, summarise and answer questions, a company trains one large model and adapts it. Training means showing the model enormous amounts of examples until it learns patterns, and it needs thousands of powerful chips running for weeks or months.
A large language model, or LLM, is the best known type. It works on words. Other models work on speech, images or video. Generative AI is the name for the use of these models to produce new content. When a chatbot writes a reply, a model is generating it one small piece at a time.
Training from scratch is very costly. Many startups instead take an open model that already exists and fine-tune it, which means training it a little more on their own data for a narrower job. This is far cheaper and is how many Indian teams start.
Why does India build its own models?
Global models are strongest in English. India has many languages, many dialects, and people who mix languages in one sentence. A model built with Indian data can serve these users better and more cheaply, especially by voice, where many first-time internet users prefer speaking to typing.
There is also a policy reason. The government wants Indian organisations to have models they can run inside the country, on their own terms, for sensitive work in public services and defence. This idea is often called sovereign AI. Sarvam AI, which builds models for Indian languages, is one of the companies named under the IndiaAI Mission for this work, and Soket AI and Gnani.ai are among others picked for foundation model projects.
A third reason is cost. Serving millions of users in a country with low spending per user needs models that are cheap to run. This pushes Indian teams toward smaller, efficient models tuned to specific tasks.
How do these startups earn money?
The common routes are model access through an interface (an API), where customers pay per amount of text or speech processed; enterprise licences for companies that want a model set up for their needs, sometimes inside their own systems; and government or institutional projects funded by programmes such as the IndiaAI Mission.
Voice is a strong early product. Gnani.ai sells enterprise voice AI, used for tasks like automated customer calls. ElevenLabs, a global voice AI company, shows how large this market can become. Some firms also build applications on top of their own models, for example assistants or search tools, to show what the model can do and to earn directly.
Few Indian model labs earn large steady revenue yet. Most depend on investor money, grants and a small number of big customers while products mature.
What does it cost and who funds it?
The largest cost is computing power. A GPU is the chip that does the training maths, and a serious training run needs a great many of them. Indian startups rent GPUs from providers such as E2E Networks, Neysa or Krutrim, and can apply for subsidised access through the IndiaAI compute portal. A March 2026 government statement said more than 38,000 GPUs had been onboarded to that portal.
The second cost is people. Researchers who can train large models are rare and paid near global levels. The third is data: collecting, cleaning and licensing good Indian language data takes time and money.
Funders in the last 12 months in the StopDown data include Peak XV Partners, Accel, Lightspeed, Khosla Ventures and Nvidia, with Prosus and Bessemer Venture Partners also active in the wider sector. Investors in this area usually back the team and its approach, because revenue comes later.
What rules apply and what are the risks?
Training data often includes personal information, so the Digital Personal Data Protection Act, 2023 matters. Its Rules were notified in November 2025, and most company duties on notice, consent and security are due by about May 2027. Copyright questions about training on published text are not settled in India as of October 2026, so companies are cautious about what data they use. There is no dedicated AI law.
The business risks are clear. Global labs release strong models quickly, often free or cheap, and an Indian model must beat them on language, price or privacy. Training failures can waste months of computing. Customers may also switch models easily, since many products can use more than one model.
This is general information, not legal advice. A startup handling personal data or training on third-party content should check with a lawyer.
How do you judge a model company?
Because revenue is early, readers need other signals. Ask whether the model is open or closed, and whether others can build on it. Ask which languages and tasks it handles well, and whether it is tested on real Indian speech and text rather than only on clean benchmarks. A benchmark is a standard test that lets models be compared, but it can be gamed or may not match real use.
Look at who actually uses the model. Paying enterprise customers, government deployments and developers building products on it say more than a launch announcement. Look at cost to serve, because a model that is accurate but too expensive for Indian price levels will struggle to grow.
Finally, look at the team and the compute plan. A lab with strong researchers but no clear access to GPUs, or a plan that depends on one subsidy, carries extra risk. Recent news in the StopDown data shows Sarvam AI naming an ex-Google Cloud head as president, a sign that labs are adding commercial leadership as they move from research to sales.
What is changing next?
Three shifts stand out. Models are getting smaller and cheaper to run, which helps Indian use cases. Voice is moving from a feature to the main way people use AI. And agents, which act on a user's behalf, are changing what customers expect from a model.
The breakdown
Business models
| Model | How it makes money | Who uses it |
|---|---|---|
| API access | Pay per amount of text or speech processed | Model labs and voice AI firms |
| Enterprise licence | Annual fee to set up and run a model for a company | Labs serving banks and telecom firms |
| Government project | Funded build or deployment of a sovereign model | Labs selected under public programmes |
| Own applications | Subscriptions or usage fees from assistants and tools | Labs that also sell products |
The numbers that matter
- Training cost depends mostly on GPU hours, so access to cheap or subsidised computing changes what a startup can attempt.
- Serving cost per user decides margins, so smaller efficient models are attractive in India.
- Talent is a large fixed cost because senior researchers are scarce.
- Revenue is often lumpy, coming from a few large enterprise or government contracts.
Rules and regulators
| Regulator or law | What it means |
|---|---|
| DPDP Act, 2023 and Rules, 2025 | Personal data used in training or in products needs lawful handling; most duties are due by about May 2027. |
| IndiaAI Mission | Offers subsidised GPU access and backs selected foundation model projects. |
| Copyright Act, 1957 | How it applies to training on published content is not settled; companies take care with data sources. |
| IT Act, 2000 | Applies to harmful or unlawful content produced or shared through AI products. |
Risks
- Global models improving faster than local ones.
- High and uncertain training costs.
- Thin revenue while products mature.
- Unsettled rules on training data and copyright.
- Low switching costs for customers.
Foundation models & generative AI: latest on StopDown
- Inner Sky Labs launches physical AI foundation models 8 October 2026
- ElevenLabs reaches six hundred million dollars in annual recurring revenue 24 September 2026
- Gnani AI raises $14.19 million in Series B 11 September 2026
- Sarvam AI raises $300-$350 million and turns unicorn 9 September 2026
- Bajaj Finance buys 5% stake in TrueFan AI 9 September 2026
- Sarvam AI co-founder named among TIME's top AI leaders 30 August 2026
- Gnani.ai unveils sovereign AI stack Artha for Indian enterprises 28 August 2026
- Sarvam AI raises $309 million to build India-specific AI models and infrastructure 24 August 2026
Every Foundation models & generative AI story →
Most active investors here
- HCLTech (6 rounds)
- Nvidia (5 rounds)
- Bessemer Venture Partners (4 rounds)
- Activate (2 rounds)
- IAN Alpha Fund (2 rounds)
- Khosla Ventures (2 rounds)
- Lightspeed (2 rounds)
- Peak XV Partners (2 rounds)
Rounds StopDown covered in the last 12 months. Activity is not a measure of quality.
Questions people ask
What is the difference between a foundation model and generative AI?
A foundation model is the large trained system. Generative AI is what you do with it: create text, speech, images or code. The same model can power many generative products.
Why do Indian language models matter?
Most global models work best in English. Models trained on Indian languages and speech serve more people, especially by voice, and can run at lower cost for local use.
Can a startup afford to train its own model?
Training from scratch needs a lot of GPUs, so few can. Many start by fine-tuning an open model, and some use subsidised GPU access through the IndiaAI compute portal.
Which Indian companies work on foundation models?
Sarvam AI, Soket AI and Gnani.ai are among those working on Indian foundation models, and Krutrim builds AI infrastructure. These are examples, not a ranking.
Which Foundation models & generative AI startups in India raised money recently?
Gnani.ai ($14.19 million, Series B); Sarvam AI ($300-$350 million); Sarvam AI ($309 million, Series B); Sarvam AI ($300 million, Series B); Ema ($80 million, growth).
Who invests in Foundation models & generative AI startups in India?
Among the most active backers in StopDown's coverage over the last year: HCLTech, Nvidia, Bessemer Venture Partners, Activate, IAN Alpha Fund.
Which Foundation models & generative AI companies are in the news?
Recent stories on StopDown cover Inner Sky Labs, ElevenLabs, Gnani.ai, Sarvam AI, TrueFan AI, Murf AI, Ola Krutrim, Ema.
More in AI & Deep Tech
Startup Decoded · Glossary · Sectors explained · Investor directory · FAQs