What does an AI chatbot really cost? Our own numbers
The raw compute of one AI answer costs us about a hundredth of a cent. An AI chatbot gets expensive not through tokens but through the pricing model: price per resolution, seats, quotas. Here are our numbers and the maths behind them.
What an answer costs technically
Language models count in tokens — pieces of words. You pay per million tokens, split into input (question, conversation history, knowledge-base excerpts) and output (the answer).
SilentChat's default model is Mistral Small 24B in the IONOS AI Model Hub in Germany; other AI providers, such as US ones, are blocked. According to the IONOS price list, the model costs 0.11 US dollars per million input tokens and 0.33 US dollars per million output tokens. The vectors for knowledge-base search are computed by bge-m3 at 0.02 US dollars per million tokens.
Our measurement
On 14 September 2026, our cost overview showed 737 AI calls on Mistral Small with 417,000 tokens in total — just under 570 tokens per call on average. That comes to about 7 US cents, roughly a hundredth of a cent per call. The IONOS console showed 0.13 euros for the month up to then.
These are measurements from young operations with few tenants. Longer conversations and larger knowledge bases mean more input tokens per answer and therefore higher costs per conversation.
A rounding error that charged a hundred times too much
Our own cost display first showed 7.37 US dollars for the same 737 calls. The cause: every call was rounded up to a full cent. With amounts of a hundredth of a cent, that is a factor of a hundred — and because our cost caps added these amounts up, customers would have hit the AI pause far too early.
Since then we calculate in millionths of a cent and carry remainders per tenant. For the caps we book with a safety margin of at most twenty times — better too cautious than an open bill.
Why AI chatbots still don't cost next to nothing
- Price per resolution: Intercom charges 0.99 US dollars per “outcome” for its AI agent Fin — for example when a customer confirms their issue is resolved, or asks for no further help after Fin's answer, or Fin completes a workflow (“Procedure”), including handoffs. It is charged at most once per conversation. That is a price for the result, not for the compute.
- Operations and protection: hosting, monitoring, support and abuse protection cost more than the tokens. An AI that answers in a loop or is addressed by bots needs limits.
- Larger models: top-tier models cost many times more per token.
How we price it
Our plans include AI conversations as a quota: 50 a month on Starter (€19), 300 on Growth (€49), 1,000 on Professional (€129) and 3,000 on Enterprise (€349) — monthly prices excl. VAT, about 20 % less with annual billing (as of October 2026). Because a single conversation can run arbitrarily long, monthly AI usage is also capped (fair use). Once the quota or that cap is reached, the bot pauses until the end of the month; your team takes over, and live chat and all other features keep running. On top come a cost cap per tenant and a kill switch for the whole AI. The free plan includes no AI.
What to watch for when comparing
- What is counted — answers, conversations or resolutions? And what counts as “resolved”?
- Is there a cap, or does the bill grow with every visitor?
- Which model runs, and where? Is your data used for training?
- Can a human take over at any time?
Our plans in detail: silentchat.de/pricing. How the search vectors are created and what they cost is covered in RAG embeddings: the invisible costs.
Sources
- IONOS AI Model Hub: pricing (retrieved 5 October 2026)
- IONOS: Data Handling (retrieved 5 October 2026)
- Intercom: Pricing (retrieved 5 October 2026)