Retour au blog
August 4, 2026 · 6 min de lecture

How to choose an AI model for customer replies

How to choose an AI model for customer replies

Every model launch comes with a chart showing it beating the last one. None of those charts were produced by asking your customers questions in your customers' language about your products. The gap between "scores well" and "answers this business well" is where most disappointment lives.

Test on your own conversations, not on prompts you invent

A made-up test message tells you how a model handles a made-up test message. Real threads carry the things that actually break models: half-typed words, dialect, a question that only makes sense because of something said four messages earlier, a customer who changes their mind mid-sentence.

Take twenty or thirty conversations you already have. Prefer the awkward ones. Replay them through each model with the exact instructions your live assistant uses, and read the answers side by side.

Score the things that cost money

  • Did it hand over when it should have? A model that writes beautifully but never escalates leaves hot leads sitting.
  • Did it stay in the customer's language, including dialect and script?
  • Did it invent a price, a stock level or a delivery date?
  • Did it respect opening hours instead of promising an immediate callback at 2am?
  • What did the reply actually cost, in money rather than tokens?

Tokens are not comparable, cost is

Different providers count tokens differently, so "uses fewer tokens" means nothing across vendors. Compare the price of a finished reply. And remember that caching changes the picture: a model that caches your instructions gets cheaper on repeat traffic in a way a headline rate does not show.

Latency matters less than people think

A reply in nine seconds still lands long before a human would have picked up the phone. Treat a two-times difference as signal and a twenty-percent difference as noise, and do not trade answer quality for a second.

Decide with a backup in place

Whatever you pick will occasionally time out or return nothing. That is a normal property of third-party APIs, not a reason to avoid the good one. Nominate a second model to catch failures and the choice stops being risky.

whacl has this built in: replay real conversations against Claude, GPT and Gemini side by side, rate the answers, see the real cost per reply, and switch from a single setting once you have decided.

WT
The whacl team
WhatsApp Business, day to day

Written and reviewed by the team that builds whacl and runs assistants on real business numbers every day.

See it answer your own customers.
Connect your WhatsApp number, load your catalog, watch it reply. Live the same day.
Commencer gratuitement →