Token limits & budget
Last updated: September 16, 2026
Each AI request consumes tokens — the unit of measurement for AI usage. Tokens are approximately 0.75 words.
Token Consumption per Request
A typical AI request consumes:
| Component | Tokens |
|---|---|
| System Prompt | 100–500 |
| Chat History (Context) | 200–1000 |
| KB Chunks (RAG) | 300–1500 |
| Bot Response (Output) | 100–500 |
| Total per Request | ~700–3500 |
Limits
Monthly Limit
The monthly token limit is determined by your plan:
- Growth: Base limit defined in the plan
- Professional: Higher limit
- Enterprise: Custom or unlimited
The limit can be overridden per tenant via Monthly Token Limit in the AI configuration (if the plan allows).
Daily Limit
Optionally, you can set a daily limit to avoid peaks:
Settings → AI Chatbot → Daily Token Limit
- 0: No daily limit (only monthly limit applies)
- > 0: Maximum tokens per day
Budget Display
Under AI Chatbot → Usage you can see:
- Consumed Tokens (current month)
- Remaining Budget
- Daily Consumption (graph)
- Consumption per Conversation (top list)
- Consumption per Widget (if multiple widgets)
What Happens When the Budget is Exhausted?
When the token budget is used up:
- The bot sends the fallback message
- New visitor messages are forwarded directly to agents
- The bot is no longer called until the budget renews (next month/day)
Estimate Costs
Costs depend on the model (all at IONOS):
| Model | Input (per 1M Tokens) | Output (per 1M Tokens) |
|---|---|---|
| Qwen 3.5 9B (IONOS) | $0.11 | $0.17 |
| Mistral Small 24B (IONOS) | $0.11 | $0.33 |
| GPT-OSS 120B (IONOS) | $0.17 | $0.71 |
| Llama 3.3 70B (IONOS) | $0.71 | $0.71 |
Tip: Start with an inexpensive model (Qwen 3.5 9B or Mistral Small 24B) and switch if needed.