LLM API pricing in India — GPT, Gemini, Claude, Llama & more

30 models compared — 16 open-source, 14 proprietary. Per-million-token costs with INR conversion, 18% GST/RCM notes, open-weight flags, and data-residency context. Sorted by effective cost after GST.

Manually verified data — last checked 2026-07-27. The daily GPU pipeline does not refresh this page. Always verify current rates at provider links. Prices in USD per 1M tokens unless marked INR. FX rate: ₹96.28/$1.
How this table works: Sorted by effective INR cost per 1M output tokens. For USD-billed providers, we add 18% GST/RCM to show the real cost for Indian businesses. Sarvam AI and other domestic providers marked 🇮🇳 India bill in INR with GST included — no RCM burden. open-source models can be self-hosted on India GPU clouds to eliminate data-residency concerns. Data residency notes are editorial context, not legal advice.
Provider / ModelInput / 1M tokensOutput / 1M tokensAfter 18% GSTPrice barContextIndia regionGST/RCMSource
Sarvam AI🇮🇳 IndiaSarvam-2B-v0.5Natively priced in INR. No RCM/GST complication — domestic vendor.
₹1₹4₹4 incl.
4KIndiaNone
Mistral AIopen-sourceMistral NemoOpen-weight Apache 2.0. GDPR-compliant EU servers. Free tier available. RCM GST for Indian businesses.
$0.02$0.04₹4.54 +18%
128KRCM +18%
Meta / Groqopen-sourceLlama 3.1 8B (Groq)Llama 3.1 8B hosted on Groq LPU — 10x faster than GPU inference. Open-weight: self-hostable on India GPUs.
$0.05$0.08₹9.09 +18%
128KRCM +18%
Meta / HuggingFaceopen-sourceLlama 3.1 8B (HF Providers)Free tier available (rate-limited). HF routes to Groq/Cerebras for fastest inference. OpenAI-compatible. Good for prototyping — upgrade to provider directly for production.
$0.05$0.08₹9.09 +18%
128KRCM +18%
Sarvam AI🇮🇳 IndiaSarvam-30BNatively priced in INR. Supports 10+ Indian languages.
₹2.5₹10₹10 incl.
32KIndiaNone
Sarvam AI🇮🇳 IndiaSarvam-105BAt ₹16/1M output tokens, dramatically cheaper than equivalent frontier models after RCM GST.
₹4₹16₹16 incl.
128KIndiaNone
Alibaba / Together AIopen-sourceQwen3.5 9B (Together AI)Open-weight. Cheaper alternative to 70B for most tasks. RCM GST applies.
$0.17$0.25₹28.4 +18%
128KRCM +18%
Meta / HuggingFaceopen-sourceLlama 3.3 70B (HF Providers)HuggingFace passes provider pricing through with zero markup. PRO plan ($9/mo) gives 2M credits/month free. Good entry point to test before committing to a provider directly.
$0.26$0.26₹29.54 +18%
128KRCM +18%
DeepSeekopen-sourceDeepSeek V4 Flash18% GST RCM for Indian businesses. Open-weight: can self-host on India GPU cloud to avoid data-residency concern.
$0.14$0.28₹31.81 +18%
1000KRCM +18%
Mistral AIopen-sourceMistral Small 3Open-weight. ~25x cheaper input vs GPT-5.4. GDPR compliance useful for EU data. RCM GST.
$0.10$0.30₹34.08 +18%
32KRCM +18%
Meta / DeepInfraopen-sourceLlama 3.3 70B (DeepInfra)Cheapest hosted 70B option. RCM GST applies. Open-weight: self-hostable.
$0.23$0.40₹45.44 +18%
128KRCM +18%
Alibaba / Qwenopen-sourceQwen3.5 FlashOpen-weight — widely downloaded. Data residency concern: servers in China. RCM GST applies.
$0.10$0.40₹45.44 +18%
1000KRCM +18%
OpenAIGPT-4o mini18% GST RCM for Indian businesses. India data residency requires Azure OpenAI (separate pricing).
$0.15$0.60₹68.17 +18%
128KRCM +18%
Meta / Groqopen-sourceLlama 3.3 70B (Groq)250+ tok/s on Groq. Open-weight: can self-host on India H100s (~₹191/hr from NeevCloud) for data residency.
$0.59$0.79₹89.75 +18%
128KRCM +18%
Meta / Together AIopen-sourceLlama 3.3 70B (Together AI)Flat input/output rate. OpenAI-compatible — swap with base URL change. Wide catalog for experimenting. RCM GST applies.
$0.88$0.88₹99.98 +18%
128KRCM +18%
Google🇮🇳 IndiaGemini 3.1 Flash-LiteAvailable in India region via Vertex AI. 18% GST RCM for non-GST-registered Indian businesses.
$0.25$1.50₹170.42 +18%
1000KGoogle Cloud MumbaiRCM +18%
DeepSeekopen-sourceDeepSeek R1Cheapest reasoning model API. Open-weight — widely hosted by Groq, Together AI, etc. also at lower latency.
$0.55$2.19₹248.81 +18%
64KRCM +18%
Alibaba / Qwenopen-sourceQwen3.5 PlusOpen-weight. 1M context at competitive pricing. Strongest open-source option for long-context tasks.
$0.40$2.40₹272.66 +18%
1000KRCM +18%
Google🇮🇳 IndiaGemini 3.0 Flash
$0.50$3.00₹340.83 +18%
1000KGoogle Cloud MumbaiRCM +18%
DeepSeekopen-sourceDeepSeek V4 ProPeriodic 75% promotional discount drops to ~$0.44/$0.87. RCM GST applies. Open-weight: self-hostable.
$1.74$3.48₹395.36 +18%
1000KRCM +18%
Alibaba / Qwenopen-sourceQwen3.5 397BOpen-weight MoE (only 22B active params). Strong math/coding benchmarks. RCM GST + China data residency concern.
$0.60$3.60₹409 +18%
128KRCM +18%
Google🇮🇳 IndiaGemini 3.5 Flash
$0.75$4.50₹511.25 +18%
1000KGoogle Cloud MumbaiRCM +18%
AnthropicClaude Haiku 4.518% GST RCM applies for Indian businesses.
$1.00$5.00₹568.05 +18%
200KRCM +18%
Mistral AIopen-sourceMistral Large 2Open-weight. 60% cheaper output than GPT-5.4. GDPR-compliant EU hosting. RCM GST applies.
$2.00$6.00₹681.66 +18%
128KRCM +18%
OpenAIGPT-4o
$2.50$10.00₹1,136.1 +18%
128KRCM +18%
Google🇮🇳 IndiaGemini 3.1 Pro≤200K context price shown. Longer context incurs higher rates.
$2.00$12.00₹1,363.32 +18%
2000KGoogle Cloud MumbaiRCM +18%
OpenAIGPT-5.4
$2.50$15.00₹1,704.16 +18%
128KRCM +18%
AnthropicClaude Sonnet 5
$3.00$15.00₹1,704.16 +18%
200KRCM +18%
AnthropicClaude Opus 4.8
$5.00$25.00₹2,840.26 +18%
200KRCM +18%
AnthropicClaude Fable 5
$10.00$50.00₹5,680.52 +18%
200KRCM +18%

"After 18% GST" column shows the effective INR cost of 1M output tokens for an Indian business on RCM. Sarvam AI is GST-inclusive (no RCM). FX ₹96.28/$1 as of 2026-07-27.

Self-hosting open-source models on India GPU clouds

Models marked open-source have publicly available weights under open licences. Indian teams handling personal data under the DPDP Act can self-host these on India-based GPU clouds (E2E Networks, NeevCloud, Neysa, Jarvislabs) to keep data in-country — removing the China/US data-residency concern that comes with using DeepSeek or Qwen APIs directly. At sustained workloads (>50M tokens/day), self-hosting a 70B model on an H100 often becomes cheaper than hosted APIs; at lower volumes, managed APIs win on cost and ops overhead.

Frequently asked questions

Which LLM API is cheapest for Indian developers?

Sarvam AI's models are dramatically cheaper when you factor in 18% GST/RCM. Sarvam-30B costs ₹2.5/1M input tokens in INR with no RCM tax. For open-source models, DeepSeek V4 Flash at $0.14/1M input (+18% RCM = ~₹16) is the cheapest frontier-class hosted API. Llama 3.1 8B on Groq is just $0.05/1M input for lighter tasks.

What are the best open-source LLM APIs for Indian developers?

The best open-source hosted LLM APIs in 2026 are: Llama 3.3 70B on DeepInfra ($0.23/$0.40 per 1M tokens — cheapest 70B), DeepSeek V4 Flash ($0.14/$0.28 — cheapest frontier-class), Mistral Small 3 ($0.10/$0.30 — GDPR-compliant EU hosting), and Qwen3.5 Flash ($0.10/$0.40 — 1M context). All have open weights, so Indian teams can self-host them on India-based GPU clouds (E2E Networks, NeevCloud, Neysa) to eliminate data residency concerns under the DPDP Act — at the cost of managing infrastructure.

Do Indian businesses pay GST on OpenAI / Anthropic / Google APIs?

Yes. Foreign LLM API providers are classified as online information and database access services (OIDAR) under Indian GST rules. Indian businesses must self-assess and pay 18% GST under the Reverse Charge Mechanism (RCM) when importing these digital services. This increases the effective cost by 18%. Domestic providers like Sarvam AI bill in INR and collect GST directly — no RCM burden on the buyer.

Which LLM API providers have data centres in India?

Sarvam AI operates servers in India. Google offers India-region access via Vertex AI (Mumbai, ap-south1). Azure OpenAI Service has India regions, but requires an Azure enterprise contract rather than direct OpenAI API access. DeepSeek and Qwen run servers in China — data sent to their API leaves India, which is a concern under the DPDP Act. Anthropic does not currently offer an India-region endpoint. Open-source models (Llama, Mistral, DeepSeek, Qwen) can be self-hosted on India-based GPU clouds to keep data in-country.

How much does GPT-4o cost per month for a typical Indian startup?

A typical startup making 100 million tokens of API calls per month (roughly 50M input + 50M output) would pay: GPT-4o at $2.50 input + $10.00 output = $625, plus 18% RCM GST = ~₹71,200/month effective. By switching to Llama 3.3 70B on DeepInfra ($0.23 + $0.40) the same workload costs $31.50 + 18% RCM = ~₹3,590/month. Self-hosting Llama on India GPU (e.g. H100 at NeevCloud ₹191/hr) removes data residency risk and can reduce cost further for sustained workloads.

Understanding API costs: a plain-English primer

What are tokens?

Tokens are chunks of text — roughly 0.75 words in English, more in Indic languages. A 1,000-word essay is about 1,333 tokens. All LLM API pricing is per token.

Why output costs more

Output tokens (model responses) require more GPU compute than input tokens (your prompt). Typical ratio: output costs 4–10× more than input. For chatbots and RAG, budget for 3× more output than input.

GST / RCM on imports

Importing digital services (APIs) from foreign providers attracts 18% GST under Reverse Charge Mechanism. Indian businesses must self-assess and pay this directly to the government. Adds 18% to any USD-billed API cost.

GST invoice from provider

For input tax credit, you need a GST-valid invoice. Indian providers (Sarvam, E2E) issue one automatically. Google Workspace and Azure provide GST invoices. Direct OpenAI API invoices are USD only — you pay RCM yourself.

DPDP and data residency

Under India's DPDP Act, personal data handling obligations apply regardless of where data is stored. For regulated sectors, many teams prefer self-hosting open-source models on India DCs or using providers with India endpoints.

Open-source = self-hostable

Open-source models (Llama, Mistral, DeepSeek, Qwen) can be downloaded and run on your own GPU infrastructure. This removes per-token API costs and keeps data on-premises — at the cost of managing GPU infrastructure yourself.

Compare GPU cloud costs for self-hosting

Running open-source LLMs in India? These GPU providers have India data centres and offer hourly billing — no commitment required for experimentation.