What Does Each Chatbot Conversation Cost?
By CodexierPublished 5 min read
When you run a chatbot on a language model, the model provider bills you for usage, not per user or per month. That makes the running cost feel unpredictable. It is not, once you understand three things: tokens, model tiers and how much text is sent with each message. This guide explains them without maths anxiety and shows how to estimate from your own conversation volume.
The short answer
Tokens explained without maths anxiety
A token is a piece of text the model reads or writes. In English a token is on average a little under a word; Swedish, with its long compound words, usually needs more tokens for the same content. Providers publish a price per million input tokens and a higher price per million output tokens. You do not need the exact figures to reason about cost; you need to know which parts of a conversation become tokens.
- The system prompt: your instructions, tone and rules. Sent with every single message.
- Retrieved context: extracts from your documents or website that the bot uses to answer. Also sent every time.
- The conversation history: everything the visitor and the bot have said so far.
- The answer itself: output tokens, which cost more per token than input.
Why model choice changes cost
Providers sell model tiers: small, fast models and large, more capable ones. The price gap between tiers is large, often an order of magnitude. For a customer-service bot answering questions about opening hours, delivery times and returns, a small model with good retrieval is usually enough. Complex reasoning, long documents or tasks where mistakes are costly may justify a larger model. A common pattern is routing: a cheap model handles most messages and only hard cases go to the expensive one.
| Choice | Effect on cost | When it fits |
|---|---|---|
| Small model for everything | Lowest | FAQ-style support, short answers, clear sources |
| Large model for everything | Highest | Advisory tasks where quality outweighs cost |
| Routing between tiers | Low to moderate | Mixed traffic with a minority of hard questions |
| Prompt caching | Lower input cost on repeated text | Long, stable system prompts sent every time |
Context length and retrieval costs
Context is where cost quietly grows. If your system prompt is long and each message pulls in several pages from your documents, you may send thousands of tokens to produce a two-sentence answer. And because history is resent, the tenth message in a chat costs far more than the first. Practical fixes: keep the system prompt tight, retrieve fewer but better extracts, summarise long histories and cap conversation length before handing over to a human. Good retrieval also makes the bot more accurate, as we explain in how to stop a chatbot making things up.
Estimating from your own volume
- Count tokens in your system prompt and a typical retrieved context. Most provider dashboards or tokenizer tools show this.
- Pick a typical conversation: say four visitor messages and four answers of average length.
- For each turn, add system prompt plus context plus history so far as input, and the answer as output.
- Sum the turns to get input and output tokens per conversation, then multiply by the provider's current prices.
- Multiply by expected conversations per month, and add a buffer for longer chats and growth.
Do the calculation with current list prices from the provider's pricing page, since they change often and have generally fallen over time. For most small Swedish businesses the model usage ends up small next to the cost of building and maintaining the bot. Our overview of what an AI chatbot costs covers the whole picture, and the chatbot setup page shows what we include.
Keeping costs predictable
- Set a monthly spending limit and alerts in the provider's console.
- Limit answer length in the instructions; long answers cost more and are read less.
- Rate-limit per visitor to stop abuse and runaway loops.
- Log tokens per conversation so you can see which question types are expensive.
- Review monthly for the first quarter, then quarterly.
When you do not need a custom bot at all: if your visitors ask a handful of the same questions, a well-written FAQ page answers them for free. If you want help running the numbers for your own traffic, book a free call.
Frequently asked questions
Is it cheaper to pay per conversation or per token?
Many chatbot platforms bundle usage into a monthly fee with a conversation cap. Direct API use is billed per token. Bundles are easier to budget; direct use is usually cheaper at volume but needs spending limits and monitoring.
Why is Swedish more expensive than English?
Tokenizers are trained mostly on English text, so Swedish words, especially compounds, are split into more tokens. The same answer in Swedish therefore uses somewhat more tokens. It is a real but modest effect.
Does a longer system prompt make the bot better?
Not necessarily. Clear, specific instructions beat long ones. Every extra line is paid for on every message, so trim anything the bot does not actually need.
What happens if the bill suddenly spikes?
Usually a loop, abuse from a bot or a change that pulls in far more context. Spending limits stop the damage; token logging per conversation shows the cause.
Want an estimate for your volume?
Bring rough numbers on visitors and common questions. In fifteen minutes we can sketch the running cost and the model tier that fits.
Book a free 15-minute call