| Provider / Model | Input / 1M tokens | Output / 1M tokens | After 18% GST | Price bar | Context | India region | GST/RCM | Source |
|---|---|---|---|---|---|---|---|---|
Sarvam AI🇮🇳 IndiaSarvam-2B-v0.5Natively priced in INR. No RCM/GST complication — domestic vendor. | ₹1 | ₹4 | ₹4 incl. | 4K | ✓ India | None | ↗ verify 2026-07-21 | |
Mistral AIopen-sourceMistral NemoOpen-weight Apache 2.0. GDPR-compliant EU servers. Free tier available. RCM GST for Indian businesses. | $0.02 | $0.04 | ₹4.54 +18% | 128K | ✗ | RCM +18% | ↗ verify 2026-07-21 | |
Meta / Groqopen-sourceLlama 3.1 8B (Groq)Llama 3.1 8B hosted on Groq LPU — 10x faster than GPU inference. Open-weight: self-hostable on India GPUs. | $0.05 | $0.08 | ₹9.09 +18% | 128K | ✗ | RCM +18% | ↗ verify 2026-07-21 | |
Meta / HuggingFaceopen-sourceLlama 3.1 8B (HF Providers)Free tier available (rate-limited). HF routes to Groq/Cerebras for fastest inference. OpenAI-compatible. Good for prototyping — upgrade to provider directly for production. | $0.05 | $0.08 | ₹9.09 +18% | 128K | ✗ | RCM +18% | ↗ verify 2026-07-27 | |
Sarvam AI🇮🇳 IndiaSarvam-30BNatively priced in INR. Supports 10+ Indian languages. | ₹2.5 | ₹10 | ₹10 incl. | 32K | ✓ India | None | ↗ verify 2026-07-21 | |
Sarvam AI🇮🇳 IndiaSarvam-105BAt ₹16/1M output tokens, dramatically cheaper than equivalent frontier models after RCM GST. | ₹4 | ₹16 | ₹16 incl. | 128K | ✓ India | None | ↗ verify 2026-07-21 | |
Alibaba / Together AIopen-sourceQwen3.5 9B (Together AI)Open-weight. Cheaper alternative to 70B for most tasks. RCM GST applies. | $0.17 | $0.25 | ₹28.4 +18% | 128K | ✗ | RCM +18% | ↗ verify 2026-07-27 | |
Meta / HuggingFaceopen-sourceLlama 3.3 70B (HF Providers)HuggingFace passes provider pricing through with zero markup. PRO plan ($9/mo) gives 2M credits/month free. Good entry point to test before committing to a provider directly. | $0.26 | $0.26 | ₹29.54 +18% | 128K | ✗ | RCM +18% | ↗ verify 2026-07-27 | |
DeepSeekopen-sourceDeepSeek V4 Flash18% GST RCM for Indian businesses. Open-weight: can self-host on India GPU cloud to avoid data-residency concern. | $0.14 | $0.28 | ₹31.81 +18% | 1000K | ✗ | RCM +18% | ↗ verify 2026-07-21 | |
Mistral AIopen-sourceMistral Small 3Open-weight. ~25x cheaper input vs GPT-5.4. GDPR compliance useful for EU data. RCM GST. | $0.10 | $0.30 | ₹34.08 +18% | 32K | ✗ | RCM +18% | ↗ verify 2026-07-21 | |
Meta / DeepInfraopen-sourceLlama 3.3 70B (DeepInfra)Cheapest hosted 70B option. RCM GST applies. Open-weight: self-hostable. | $0.23 | $0.40 | ₹45.44 +18% | 128K | ✗ | RCM +18% | ↗ verify 2026-07-21 | |
Alibaba / Qwenopen-sourceQwen3.5 FlashOpen-weight — widely downloaded. Data residency concern: servers in China. RCM GST applies. | $0.10 | $0.40 | ₹45.44 +18% | 1000K | ✗ | RCM +18% | ↗ verify 2026-07-21 | |
OpenAIGPT-4o mini18% GST RCM for Indian businesses. India data residency requires Azure OpenAI (separate pricing). | $0.15 | $0.60 | ₹68.17 +18% | 128K | ✗ | RCM +18% | ↗ verify 2026-07-21 | |
Meta / Groqopen-sourceLlama 3.3 70B (Groq)250+ tok/s on Groq. Open-weight: can self-host on India H100s (~₹191/hr from NeevCloud) for data residency. | $0.59 | $0.79 | ₹89.75 +18% | 128K | ✗ | RCM +18% | ↗ verify 2026-07-21 | |
Meta / Together AIopen-sourceLlama 3.3 70B (Together AI)Flat input/output rate. OpenAI-compatible — swap with base URL change. Wide catalog for experimenting. RCM GST applies. | $0.88 | $0.88 | ₹99.98 +18% | 128K | ✗ | RCM +18% | ↗ verify 2026-07-27 | |
Google🇮🇳 IndiaGemini 3.1 Flash-LiteAvailable in India region via Vertex AI. 18% GST RCM for non-GST-registered Indian businesses. | $0.25 | $1.50 | ₹170.42 +18% | 1000K | ✓ Google Cloud Mumbai | RCM +18% | ↗ verify 2026-07-21 | |
DeepSeekopen-sourceDeepSeek R1Cheapest reasoning model API. Open-weight — widely hosted by Groq, Together AI, etc. also at lower latency. | $0.55 | $2.19 | ₹248.81 +18% | 64K | ✗ | RCM +18% | ↗ verify 2026-07-21 | |
Alibaba / Qwenopen-sourceQwen3.5 PlusOpen-weight. 1M context at competitive pricing. Strongest open-source option for long-context tasks. | $0.40 | $2.40 | ₹272.66 +18% | 1000K | ✗ | RCM +18% | ↗ verify 2026-07-21 | |
Google🇮🇳 IndiaGemini 3.0 Flash | $0.50 | $3.00 | ₹340.83 +18% | 1000K | ✓ Google Cloud Mumbai | RCM +18% | ↗ verify 2026-07-21 | |
DeepSeekopen-sourceDeepSeek V4 ProPeriodic 75% promotional discount drops to ~$0.44/$0.87. RCM GST applies. Open-weight: self-hostable. | $1.74 | $3.48 | ₹395.36 +18% | 1000K | ✗ | RCM +18% | ↗ verify 2026-07-21 | |
Alibaba / Qwenopen-sourceQwen3.5 397BOpen-weight MoE (only 22B active params). Strong math/coding benchmarks. RCM GST + China data residency concern. | $0.60 | $3.60 | ₹409 +18% | 128K | ✗ | RCM +18% | ↗ verify 2026-07-21 | |
Google🇮🇳 IndiaGemini 3.5 Flash | $0.75 | $4.50 | ₹511.25 +18% | 1000K | ✓ Google Cloud Mumbai | RCM +18% | ↗ verify 2026-07-21 | |
AnthropicClaude Haiku 4.518% GST RCM applies for Indian businesses. | $1.00 | $5.00 | ₹568.05 +18% | 200K | ✗ | RCM +18% | ↗ verify 2026-07-21 | |
Mistral AIopen-sourceMistral Large 2Open-weight. 60% cheaper output than GPT-5.4. GDPR-compliant EU hosting. RCM GST applies. | $2.00 | $6.00 | ₹681.66 +18% | 128K | ✗ | RCM +18% | ↗ verify 2026-07-21 | |
OpenAIGPT-4o | $2.50 | $10.00 | ₹1,136.1 +18% | 128K | ✗ | RCM +18% | ↗ verify 2026-07-21 | |
Google🇮🇳 IndiaGemini 3.1 Pro≤200K context price shown. Longer context incurs higher rates. | $2.00 | $12.00 | ₹1,363.32 +18% | 2000K | ✓ Google Cloud Mumbai | RCM +18% | ↗ verify 2026-07-21 | |
OpenAIGPT-5.4 | $2.50 | $15.00 | ₹1,704.16 +18% | 128K | ✗ | RCM +18% | ↗ verify 2026-07-21 | |
AnthropicClaude Sonnet 5 | $3.00 | $15.00 | ₹1,704.16 +18% | 200K | ✗ | RCM +18% | ↗ verify 2026-07-21 | |
AnthropicClaude Opus 4.8 | $5.00 | $25.00 | ₹2,840.26 +18% | 200K | ✗ | RCM +18% | ↗ verify 2026-07-21 | |
AnthropicClaude Fable 5 | $10.00 | $50.00 | ₹5,680.52 +18% | 200K | ✗ | RCM +18% | ↗ verify 2026-07-21 |
"After 18% GST" column shows the effective INR cost of 1M output tokens for an Indian business on RCM. Sarvam AI is GST-inclusive (no RCM). FX ₹96.28/$1 as of 2026-07-27.
Self-hosting open-source models on India GPU clouds
Models marked open-source have publicly available weights under open licences. Indian teams handling personal data under the DPDP Act can self-host these on India-based GPU clouds (E2E Networks, NeevCloud, Neysa, Jarvislabs) to keep data in-country — removing the China/US data-residency concern that comes with using DeepSeek or Qwen APIs directly. At sustained workloads (>50M tokens/day), self-hosting a 70B model on an H100 often becomes cheaper than hosted APIs; at lower volumes, managed APIs win on cost and ops overhead.
Frequently asked questions
Which LLM API is cheapest for Indian developers?
Sarvam AI's models are dramatically cheaper when you factor in 18% GST/RCM. Sarvam-30B costs ₹2.5/1M input tokens in INR with no RCM tax. For open-source models, DeepSeek V4 Flash at $0.14/1M input (+18% RCM = ~₹16) is the cheapest frontier-class hosted API. Llama 3.1 8B on Groq is just $0.05/1M input for lighter tasks.
What are the best open-source LLM APIs for Indian developers?
The best open-source hosted LLM APIs in 2026 are: Llama 3.3 70B on DeepInfra ($0.23/$0.40 per 1M tokens — cheapest 70B), DeepSeek V4 Flash ($0.14/$0.28 — cheapest frontier-class), Mistral Small 3 ($0.10/$0.30 — GDPR-compliant EU hosting), and Qwen3.5 Flash ($0.10/$0.40 — 1M context). All have open weights, so Indian teams can self-host them on India-based GPU clouds (E2E Networks, NeevCloud, Neysa) to eliminate data residency concerns under the DPDP Act — at the cost of managing infrastructure.
Do Indian businesses pay GST on OpenAI / Anthropic / Google APIs?
Yes. Foreign LLM API providers are classified as online information and database access services (OIDAR) under Indian GST rules. Indian businesses must self-assess and pay 18% GST under the Reverse Charge Mechanism (RCM) when importing these digital services. This increases the effective cost by 18%. Domestic providers like Sarvam AI bill in INR and collect GST directly — no RCM burden on the buyer.
Which LLM API providers have data centres in India?
Sarvam AI operates servers in India. Google offers India-region access via Vertex AI (Mumbai, ap-south1). Azure OpenAI Service has India regions, but requires an Azure enterprise contract rather than direct OpenAI API access. DeepSeek and Qwen run servers in China — data sent to their API leaves India, which is a concern under the DPDP Act. Anthropic does not currently offer an India-region endpoint. Open-source models (Llama, Mistral, DeepSeek, Qwen) can be self-hosted on India-based GPU clouds to keep data in-country.
How much does GPT-4o cost per month for a typical Indian startup?
A typical startup making 100 million tokens of API calls per month (roughly 50M input + 50M output) would pay: GPT-4o at $2.50 input + $10.00 output = $625, plus 18% RCM GST = ~₹71,200/month effective. By switching to Llama 3.3 70B on DeepInfra ($0.23 + $0.40) the same workload costs $31.50 + 18% RCM = ~₹3,590/month. Self-hosting Llama on India GPU (e.g. H100 at NeevCloud ₹191/hr) removes data residency risk and can reduce cost further for sustained workloads.
Understanding API costs: a plain-English primer
Tokens are chunks of text — roughly 0.75 words in English, more in Indic languages. A 1,000-word essay is about 1,333 tokens. All LLM API pricing is per token.
Output tokens (model responses) require more GPU compute than input tokens (your prompt). Typical ratio: output costs 4–10× more than input. For chatbots and RAG, budget for 3× more output than input.
Importing digital services (APIs) from foreign providers attracts 18% GST under Reverse Charge Mechanism. Indian businesses must self-assess and pay this directly to the government. Adds 18% to any USD-billed API cost.
For input tax credit, you need a GST-valid invoice. Indian providers (Sarvam, E2E) issue one automatically. Google Workspace and Azure provide GST invoices. Direct OpenAI API invoices are USD only — you pay RCM yourself.
Under India's DPDP Act, personal data handling obligations apply regardless of where data is stored. For regulated sectors, many teams prefer self-hosting open-source models on India DCs or using providers with India endpoints.
Open-source models (Llama, Mistral, DeepSeek, Qwen) can be downloaded and run on your own GPU infrastructure. This removes per-token API costs and keeps data on-premises — at the cost of managing GPU infrastructure yourself.
Compare GPU cloud costs for self-hosting
Running open-source LLMs in India? These GPU providers have India data centres and offer hourly billing — no commitment required for experimentation.