What AI replies actually cost per conversation
Model pricing is published per million tokens, which is a unit nobody experiences. The number that matters is what it costs to answer one customer, and that is built from parts the price list does not show.
What you pay for on every reply
- The instructions: your tone rules, catalogue, policies and examples, sent on every single message.
- The conversation so far, which grows as the thread does.
- The reply itself, which is usually the smallest part.
That ordering surprises people. A two-line answer to "how much per square metre" can cost more in instructions than in the answer, because the model has to be told who it is before it can respond.
Caching changes the arithmetic
Because the instructions are identical every time, they can be cached and re-read at a fraction of the price. That is often the difference between an expensive setup and a cheap one, and it is why two models with the same headline rate can differ several times over in practice.
Long conversations cost more than short ones
Every message added to a thread is re-sent with the next one. A forty-message conversation is not forty times a one-message conversation, but it is not free either. Conversations that reach a human quickly are cheaper as well as better.
Measure, do not estimate
The only reliable number is your own: real token usage from real replies, priced at the rate you are actually paying. Once you have that per-reply figure, a model change stops being a guess and becomes arithmetic — and it is usually a smaller number than people expect, well under the cost of the sale it protects.
whacl records tokens and cost for every reply, broken down per customer and per model, so switching models is a decision you can justify rather than defend.