Chatbot Development Company: 2026 Costs to Ship a Production Support Bot in India

Verified 2026 per-resolution and per-token prices, and the arbitrage between them.

Read time
13 min
Word count
2.2K
Sections
8
FAQs
8
Share
AI support chatbot cost comparison between outcome pricing and raw model tokens in 2026
Published 2026 per-resolution and per-token rates for production support bots.
On this page · 8 sections
  1. What the outcome vendors actually charge
  2. The token arithmetic, done properly
  3. Read the resolution rate's denominator
  4. India-specific considerations
  5. How we scope a support bot
  6. FAQ
  7. How eCorpIT can help
  8. References

Summary. A support bot handling 50,000 conversations a month, at a 50% automated-resolution rate, costs $24,750 on Intercom Fin at its published $0.99 per outcome. The same 50,000 conversations, at 8 turns each with 3,000 input and 300 output tokens per turn, cost about $384 a month in raw model tokens on GPT-5.6 Luna at $0.20 and $1.20 per million, or about $1,800 on Claude Haiku 4.5 at $1 and $5. That is a 64x gap at the cheap end and roughly 7x even against Claude Sonnet 5 at $3,600. The gap is not fraud; it is the price of a product rather than an API. But it is the whole reason to know your own numbers before signing, because at Indian support volumes the fee is denominated in dollars and the value of the platform has to clear a very high bar. Two other things distort the decision: published resolution rates are measured against denominators the vendors choose, and Indian-language traffic costs more per unit of meaning than English does.

What the outcome vendors actually charge

Vendor Billing unit Published rate What triggers the charge
Intercom Fin Outcome $0.99 per outcome; $9.99 for lead qualification Resolution, procedure handoff or disqualification; one outcome per conversation maximum
Zendesk, committed Automated resolution £1.50 on the India-facing page A request resolved by the AI agent with no escalation to a human
Zendesk, pay as you go Automated resolution £2.00 on the India-facing page Same definition, no commitment
Freshworks Freddy AI Agent Session $49 per 100 sessions, first 500 included A 72-hour interaction window, resolved or not
Freshdesk seats Agent $19 Growth, $55 Pro, $89 Enterprise per agent per month, billed annually Per human agent

Three details in that table decide more than the headline rate.

Freshworks bills sessions, not resolutions. A session is defined as a unique interaction, and for email agents as a 72-hour window from the customer's first message, with every AI reply inside that window counting once. You pay whether or not the bot resolved anything. Run the same 50,000 conversations through it and, after the 500 free sessions, you buy 495 packs at $49, which is $24,255. That lands within 2% of the Fin bill in total, but the two behave in opposite directions: halve the resolution rate and Fin's bill halves while the Freshworks bill does not move at all.

Intercom's resolution is generous to Intercom in one specific way. Its own documentation says an outcome counts when the customer confirms the answer was satisfactory or exits the conversation without requesting further assistance, and that disengaging for 24 hours is treated as an assumed resolution. A customer who gives up and emails you instead is, on that definition, a resolution.

The currency is the third detail, and it is the one Indian buyers feel. None of the three vendors publishes a rupee price. Zendesk's India-facing page is stranger still: it prices seats in US dollars and automated resolutions in pounds, because the resolution rows are served from UK content fragments. An Indian buyer carries full exchange exposure on every one of these products.

The token arithmetic, done properly

Assume 50,000 conversations a month, 8 turns each, 3,000 input tokens and 300 output tokens per turn. That is a realistic shape for a retrieval-augmented support prompt: a long fixed system prompt plus retrieved documents on input, a short answer on output. It works out to 400,000 turns, 1,200 million input tokens and 120 million output tokens a month.

Model Published input and output per million tokens Monthly cost at this volume
GPT-5.6 Luna $0.20 and $1.20 $384
Gemini 3.1 Flash-Lite, global $0.25 and $1.50 $480
Gemini 3.7 Flash, introductory rate to 31 December 2026 $0.75 and $3.75 $1,350
Claude Haiku 4.5 $1 and $5 $1,800
Claude Sonnet 5 $2 and $10 $3,600
Claude Opus 5 $5 and $25 $9,000

Set that against $24,750 for Fin. Even Sonnet 5, a model far stronger than a support bot needs for most tickets, costs less than a sixth of the outcome fee. Inference is not the expensive part of an AI support bot in 2026 and has not been for some time.

Two dated caveats belong on those numbers. Google's introductory pricing on Gemini 3.7 and 3.6 Flash runs to 31 December 2026, and standard pricing of $1.50 and $7.50 applies from 1 January 2027, which doubles that line. And Anthropic has confirmed that Claude Sonnet 5's $2 and $10 rates, originally announced as introductory through 31 August 2026, are now the standard price; the scheduled rise to $3 and $15 will not happen.

Prompt caching moves the floor further down. Anthropic and OpenAI both price a cache read at roughly a tenth of base input. A support prompt is mostly a fixed system prompt, so if 70% of input tokens are cache hits, the Haiku 4.5 line falls to about $1,044 and Luna to about $233, excluding cache-write charges. Batch processing halves both vendors' rates again for anything that does not need to be synchronous.

None of this is the total cost of a bot. It excludes the vector database, embeddings, retrieval infrastructure, observability, an evaluation harness and the engineering time to build and maintain all of it. That is precisely the work an outcome fee is buying you out of, and for a small team it can be worth every cent. The question is whether it is still worth it at 50,000 conversations a month, and past a certain volume the answer changes.

Read the resolution rate's denominator

Intercom's marketing states an average resolution rate of 76% across its customer base. Its own help documentation defines the terms: automation rate equals involvement rate multiplied by resolution rate, and resolution rate is conversations Fin resolved divided by conversations Fin was involved in. The worked example in that document has 1,000 conversations, Fin involved in 500, and 250 resolved, producing a resolution rate of 50% and an automation rate of 25%.

Those are the same 250 conversations described two ways. The number that matters to your cost model is the automation rate, because it tells you what share of total contacts never reached a human. The number in the marketing is the other one.

Zendesk states that its AI resolves up to 80% of even complex service issues, with no footnote, sample size or window on the page. Freshworks states up to 80% in marketing, and 65% average deflection in a press release describing an early-access cohort with no sample size given. Zendesk's own explainer on the metric contains no percentage at all, and warns that different tools classify "resolved" in different ways, which can inflate performance or hide gaps. That warning is the most useful sentence any of the three vendors has published on the subject.

The best-specified vendor figure we found is Salesforce reporting that Agentforce has handled 4.3 million inquiries on its own help property and resolved 70% of them, though a separate Salesforce page puts the figure at 76% over 1.7 million conversations, and its customer examples run at 50% and 60%. The one rigorously A/B-tested study, published by Nubank at KDD 2026 across a base of about 100 million users, deliberately reports only a 29 percentage point improvement in self-service rate rather than an absolute level, and its authors note the trade-off directly: routing hard cases and frustrated users straight to humans raises both the AI satisfaction score and the apparent resolution rate.

Plan on 50% to 70% with a named denominator. Treat 80% as a best case that has been promoted into a capability claim.

India-specific considerations

Indian-language traffic costs more per unit of meaning, and no vendor will tell you how much. Anthropic's glossary says only that a token approximately represents 3.5 English characters and that the exact number varies by language. Google says a token can be characters, words or phrases. There is no published Indic multiplier from any model vendor.

The measured picture, from independent work rather than vendors, is that the penalty has fallen sharply and is still real. Peer-reviewed analysis of tokenisation inequality found differences up to 15 times across languages and argued this produces unfair costs for some language communities. More recent measurement against current tokenizers puts the Indic penalty far lower, in the region of 1.3 times for Hindi and 2.5 to 2.9 times for Tamil, Telugu, Kannada and Malayalam, though that work is a single-author preprint and should be treated as indicative rather than settled. The practical instruction is narrow: if you are still costing an Indic deployment using a multiplier taken from 2023-era tokenizers, you are overstating it substantially, and if you are costing it at parity with English you are understating it.

Quality reporting is thinner than pricing. Google enumerates 15 Indian languages that Gemini models can understand and respond in. Anthropic publishes relative-to-English scores for exactly two, Hindi at 96.7% and Bengali at 95.4% for Sonnet 4.5, and 92.4% and 90.4% for Haiku 4.5. OpenAI's last published per-language Indic figures were MMLU Language scores of roughly 0.90 for Hindi and 0.89 for Bengali; its two most recent system cards contain no multilingual section at all. No vendor publishes an accuracy score for Tamil, Telugu or Marathi on a current frontier text model. Benchmark your own traffic, in your own languages, including the code-mixed Hinglish and Tanglish that AI4Bharat points out is the primary mode of communication for millions and the mode that models trained on clean text handle worst.

There is an Indian alternative with a rupee rate card. Sarvam AI publishes Sarvam 105B at Rs 29.28 per million input tokens, Rs 10.98 cached input and Rs 73.20 output, covering the ten most-spoken Indian languages plus English in native script, romanized and code-mixed input. Worth benchmarking against, not least because it removes the currency exposure.

The compliance shape is set by the Digital Personal Data Protection Act, 2023. A chat transcript is data about an identifiable individual in digital form, so you are the Data Fiduciary and your model vendor is a Data Processor, engaged under section 8(2) only through a valid contract. Section 8(1) makes you liable for processing carried out on your behalf, which means a breach at your model vendor is your notification obligation. Section 8(6) requires intimation of a breach to both the Board and each affected data principal.

Section 11(1)(b) is the clause that should change your architecture. A data principal can require you to disclose the identities of every other data fiduciary and processor with whom their personal data has been shared, along with a description of what was shared. If a conversation flows to a model API, an embedding provider, a vector database and an observability vendor, all four are disclosable, per customer, on request. Design the transcript pipeline so that list is short and machine-generatable. Section 8(7) then requires erasure on consent withdrawal or once the purpose is served, and requires you to cause your processor to erase too, which is a direct problem for transcript archives kept as a permanent retrieval corpus.

How we scope a support bot

  1. Measure your real conversation shape before pricing anything. Turns per conversation and tokens per turn are the two inputs the whole model rests on, and neither can be assumed.
  1. Compute the crossover. There is a monthly volume at which an outcome fee stops being cheaper than owning the stack, and it is a specific number for your traffic, not a rule of thumb.
  1. Define resolution yourself, and instrument it against total contacts rather than bot-involved contacts. Otherwise you cannot compare a vendor's number with your own.
  1. Benchmark your own languages. Run a held-out set of real tickets, including code-mixed ones, across two or three model tiers before choosing.
  1. Draw the data map for section 11. Every hop a transcript takes is a disclosure obligation and a deletion obligation.
  1. Build the evaluation harness before the bot. Without it, a model or prompt change is a guess, and model versions move faster than release cycles.

Our engagement model follows that order. A fixed-scope discovery produces the measured conversation shape, a crossover analysis against at least two commercial options, a language benchmark on your own tickets and the DPDP data map. Several of these end with a recommendation to buy a platform and spend the budget on knowledge quality instead, which is usually the right answer below the crossover. Above it, the build runs on milestones with the evaluation harness delivered first.

For related reading, our AI chatbots customer service cost reduction analysis covers the deflection economics in more depth, enterprise AI agents production use cases covers what agents are actually doing in production, and AI agent security and prompt injection guardrails covers the failure mode that turns a transcript store into a breach.

FAQ

How eCorpIT can help

eCorpIT is a Gurugram-based engineering organisation, founded in 2021, that builds retrieval-augmented support bots and the evaluation harnesses that keep them honest. We start by measuring your real conversation shape, computing the crossover against the commercial options, and benchmarking model tiers on your own tickets in your own languages, so the buy-or-build decision rests on your numbers. We design systems aligned with Digital Personal Data Protection Act, 2023 requirements, including the section 11 processor disclosure and section 8(7) erasure paths, and work as senior-led, multi-disciplinary teams under CMMI Level 5 assessed processes. Talk to our team before you sign an outcome-priced contract.

References

  1. Intercom pricing and Fin outcome pricing, Intercom
  1. Fin AI agent outcomes, Intercom Help Center
  1. Update to Fin performance metrics, Intercom Help Center
  1. Zendesk pricing, Zendesk India
  1. Freshdesk pricing, Freshworks
  1. Claude model pricing, Anthropic developer platform
  1. Multilingual support, Anthropic developer platform
  1. OpenAI API pricing, OpenAI developers
  1. Generative AI on Vertex AI pricing, Google Cloud
  1. Gemini model language support, Google Cloud documentation
  1. Sarvam AI pricing, Sarvam documentation
  1. Neural machine translation and Indic language evaluation, AI4Bharat, IIT Madras
  1. Language model tokenizers introduce unfairness between languages, Petrov et al.
  1. Improving self-service rate and satisfaction in an AI customer support agent, Nubank, KDD 2026
  1. Agentforce customer support results, Salesforce
  1. The Digital Personal Data Protection Act, 2023, India Code

Last updated: 16 August 2026.

Frequently asked

Quick answers.

01 How much does an AI support chatbot cost per month?
It depends on what you buy. At 50,000 conversations a month with a 50% resolution rate, Intercom Fin bills 25,000 outcomes at $0.99, so $24,750. The underlying model tokens for the same traffic cost about $384 on GPT-5.6 Luna or $1,800 on Claude Haiku 4.5.
02 Why is per-resolution pricing so much more than the token cost?
Because you are buying a product rather than an API. The fee covers retrieval infrastructure, the vector store, observability, evaluation tooling and the engineering behind them. That is genuinely valuable below a certain volume. Above it, the same capability costs less to own than to rent.
03 What is the difference between a resolution and a session?
Intercom and Zendesk bill resolutions, meaning a request the AI closed without escalation. Freshworks bills sessions, defined for email agents as a 72-hour window from the customer's first message. Sessions are charged whether the bot resolved anything or not, so a poor resolution rate does not reduce the bill.
04 Are the 80% deflection rates vendors quote realistic?
Treat them cautiously. Zendesk and Freshworks both say "up to 80%" with no sample size or methodology published. Intercom's 76% is resolutions divided by conversations the bot was involved in, not total conversations; its own worked example turns a 25% automation rate into a 50% resolution rate.
05 Do Indian languages cost more to run through an LLM?
Yes, though less than they used to and no vendor publishes a figure. Independent measurement against current tokenizers suggests roughly 1.3 times for Hindi and 2.5 to 2.9 times for Tamil, Telugu, Kannada and Malayalam. Older estimates based on 2023-era tokenizers substantially overstate the penalty now.
06 Which model should a support bot use?
Benchmark rather than assume. The cheap tiers are strong enough for most support traffic, and the published spread between GPT-5.6 Luna at $0.20 per million input tokens and Claude Opus 5 at $5 is more than 20x. Run held-out real tickets in your own languages across two or three tiers.
07 What does the DPDP Act require of chat transcripts?
Transcripts are personal data, so you are the Data Fiduciary and your model vendor is a processor engaged under a valid contract. Section 8(6) requires breach intimation to the Board and affected individuals, section 11(1)(b) requires disclosing every processor the data was shared with, and section 8(7) requires erasure.
08 Should we buy a platform or build our own bot?
Buy below your crossover volume and build above it. The crossover is a real number specific to your conversation shape, resolution rate and language mix. Compute it from measured traffic rather than from a vendor calculator, because the calculator's assumptions favour the vendor.

About the author

Manu Shukla

Founder & Director

Founder of eCorpIT. Hands-on engineer leading senior-only delivery for AI apps, custom software, and cloud systems for global clients.

Subscribe

One engineering note a week. No fluff, no spam.

Senior-architect playbooks on AI agents, mobile apps, cloud, security, data, and marketing — delivered every Wednesday.

Past the reading

Read enough. Let's build something.

A senior architect responds in 24 working hours with scope, indicative cost, and a timeline. NDA before any technical conversation.