Measuring whether your AI chat is working
The headline number on most chat dashboards was invented for a support department. Here is what a shop should watch instead, and how much time to spend watching it.
By The Starly team · · 12 min read
Most chat tools put deflection rate at the top of the dashboard, or something wearing a different name that measures the same thing. That number was invented for a support department whose budget is justified by tickets that never reached a person. You run a shop. You are measured on orders. The headline number is answering somebody else's question, and answering it well. This page replaces it with four that answer yours, then puts a clock on the routine.
One thing to set straight before the list. Starly gives you full conversation transcripts and a per-message trace of what the assistant read before it answered. It does not give you a metrics dashboard. Everything below is counted by hand or dropped into a spreadsheet, which for most shops is about twenty minutes a month. If you were expecting a chart, this is the trade: you read the conversations, and you learn considerably more than a chart would have told you.
Deflection rate flatters everyone
Deflection rate counts conversations that did not reach a human, as a share of all conversations. Read that twice. It counts the customer who got a clear answer about shipping to Germany and went on to check out. It also counts the customer who asked the same thing three ways, got nothing usable, closed the window and bought elsewhere. Both improve the number. An assistant that quietly annoys people into leaving scores beautifully.
It also gets worse the harder you push on it. Burying the handoff raises it. Asking three qualifying questions before answering raises it. Replying with a link to a policy page instead of the fact somebody asked for raises it. Any metric you can improve in an afternoon by degrading the experience is not one to steer by.
It is not useless, though. It is a capacity number, and as a capacity number it is good. Say 400 conversations arrive in a week and 300 of them end without a person. The other 100 are your team's actual workload from this channel. Multiply by the minutes a reply takes you and you know whether one person covers it, whether the queue survives a holiday week, and at what volume you would have to hire. That is a real planning question and deflection answers it properly. It just says nothing about whether those 300 people were helped.
Four numbers that answer a shop's question
The replacements are not exotic. Three are counts you take straight from transcripts. The fourth is a list rather than a rate, and it is the one that changes what you do next week.
- Conversations that reached a cart or an order. Not attributed revenue, just the count and the share.
- Second questions: conversations where the customer asked again after the first answer.
- Handoffs that arrived with context, as a share of all handoffs.
- Questions the assistant could not answer, grouped by subject rather than counted.
Cart share is the closest thing to your question, and it is easier to count than it sounds, because reaching a cart leaves a mark in the transcript. What that mark looks like differs by channel and the mechanics of it belong elsewhere; for counting purposes all you need is that there is a visible moment where the conversation turned into a cart. You are therefore counting something that happened rather than inferring intent from wording, which is the difference between this number and most of the ones you will be offered. Count it within the same session only, and treat the result as a floor.
Second questions is the one people misread. A conversation with one question and no follow-up is usually not a success. It is often someone who tried, got something they could not use, and left without saying so. Asking again means the first answer was worth building on, and two or three exchanges is what a working conversation looks like. If nearly all of yours are one turn long, stop counting and read the answers. Count only new questions here. Somebody rephrasing the same question because the answer missed is the opposite signal, and it should have fired a handoff rather than shown up in this column; if your assistant is configured the way the handoff page argues for, those conversations are already sitting in somebody's inbox and are not yours to count twice.
The handoff number measures the quality of the transfer, not how often it happens. A handoff that reaches your team carrying the product, the question and what has already been said takes a few minutes to resolve. One that arrives as "customer would like to speak to someone" costs what it would have cost with no chat at all, and the customer repeats themselves, which is the thing they most resent. Track the share that arrive complete. If that share is low, the fix is in your content and your handoff setup, not in the number.
Unanswered questions are a backlog, not a score
The fourth number is the one most shops throw away, and it is the most valuable thing the channel produces. Every question the assistant could not answer is a customer telling you, unprompted and for free, what your shop does not explain. A failure rate is a verdict you can do nothing with. A grouped list is a work queue with the priorities already filled in.
Group by subject, not by wording. Say you end the month with eleven questions about whether a product is dishwasher safe, seven about delivery to one country, four about what happens when a size is wrong. Written as a percentage that is a small failure rate you would shrug at. Written as a list it is next month's writing, already sorted, and you start with the group of eleven. The sizes are the ordering, which is why the grouping is worth the five minutes it costs.
The list surfaces things beyond chat, too. If fourteen people asked for the same measurement, the product page is missing a measurement, and fixing the page removes the question from every channel at once rather than teaching the assistant to answer it. Treat the biggest groups as a site to-do list first and a chat to-do list second.
Do not chase the tail. There will always be questions asked exactly once, and answering them one by one is how you lose a weekend for nothing. Groups of one are noise. Groups of five and up are content. Turning one of those groups into a card that actually gets retrieved is a separate craft, dealt with where writing is; this page stops at finding the backlog and putting it in order.
Segment by channel before you conclude anything
Website chat, WhatsApp and helpdesk tickets behave nothing alike, and a number computed across all three blends populations with nothing in common. Website visitors are mostly pre-purchase, in a hurry and anonymous. WhatsApp arrives from people who already bought, often days after an order and often in another language. Helpdesk tickets are the hard ones by definition, because the easy version of that question got answered somewhere else.
Averaged together they hide the one that is broken. The usual version is a shop whose overall picture looks fine because a healthy website widget is carrying a WhatsApp assistant nobody has read since the day it was connected. Compute all four numbers per channel, every time. If that feels like too much work, it means you have more channels open than you are willing to maintain, and closing one is a legitimate answer.
One bookkeeping note: exclude your shareable test links from every count. Those conversations are you, your colleague and whoever you sent the link to, and they are wildly better behaved than real traffic. A dozen test conversations in a slow month will move your cart share by several points and tell you nothing.
The four-week rule
The first week measures your setup, not your assistant. Content is still wrong in ways you have not found yet, unexpected questions arrive, and you will change things mid-week, which invalidates the week you were measuring. Shops cancel in week one over one embarrassing transcript and renew over one good one. Both are decisions made on noise.
There is an awkward version of this at Starly specifically. The first week is free, and the free week is week one, which is the week that tells you least. If you want the trial to answer anything, spend it fixing content rather than judging output, and put the real decision in week four.
Judge it in week four against the same four weeks before you had it. That comparison is imperfect. Seasonality leaks into it, so does a campaign, so does a supplier delay, and you should say so out loud when you look at the two columns. It still beats week four against week one, which mostly measures how much content you fixed in between. Put the review date in the calendar on day one, while you still have no opinion to defend.
Twenty transcripts a week, in a calendar slot
No metric replaces reading. Twenty transcripts a week, whole conversations rather than opening messages, tells you more about what your assistant is doing to your customers than any dashboard would. Pick them without curating. If you can pick them out, take five that ended in a handoff and five that ran past six messages, since long conversations and handoffs are where things go wrong. Otherwise take twenty in a row and do not skip the boring ones.
Make it a scheduled fifteen minutes on the same day each week rather than an aspiration. Fifteen minutes is enough for twenty conversations, because most are short and obviously fine and you are scanning for the two that make you wince. Real problems appear in transcripts weeks before they appear in any number, usually as one confidently wrong sentence about a policy nobody ever wrote down. When you find that sentence, the per-message trace shows you which knowledge card the assistant was reading, so the fix takes about as long as the finding did.
What to refuse to measure
Refuse precise revenue attribution for a single conversation. Someone asks about a fabric on Tuesday, returns through an email link on Friday and buys on a different device. Any tool that hands you a figure for that conversation's contribution has guessed and then rounded the guess to look authoritative. Count the conversations that reached a cart in the same session, call it a floor, say plainly that the real number is higher by an amount you cannot know, and decide on the floor. A floor you trust beats a total you do not.
Refuse satisfaction scores built from a handful of responses. Chat surveys are answered by the delighted and the furious and nearly nobody else. Four responses is an anecdote with a decimal point attached, and it will swing a whole point for reasons unrelated to your assistant. Below about thirty responses in a month, read the comments and ignore the average.
Refuse two more numbers that get offered to you mainly because they are easy to produce. Average response time is close to meaningless for an assistant, which answers everything in seconds including the questions it should have refused, so it measures your infrastructure and tells you nothing about your shop. Messages per conversation is worse, because it points in two directions at once: a long conversation is either somebody being walked through a real decision or somebody failing to get an answer and trying again. You already have an instrument for the second case, and it is the second-question count read together with what the second question actually said.
Some of this is unmeasurable at the volume a normal shop has, and the honest substitute is reading. That is unwelcome advice, because a dashboard promises that judgment can be handed over to arithmetic and reading promises the opposite. Anyone selling you a number for something that cannot be counted at your scale is selling the feeling of knowing.
When there is nothing here worth counting
Under roughly two hundred conversations a month, every rate on this page is noise. Two extra carts move your cart share by a point that means nothing, and a bad week is three people in a bad mood. At that volume, skip the rates entirely and read every transcript, because you can. Keep the unanswered-question list, which works at any volume because it is a list rather than a rate.
If your chat is overwhelmingly post-purchase, these four numbers are measuring the wrong shop. Cart share cannot rise on conversations that were never going to end in a cart, and the unanswered list fills with order lookups no assistant here can perform at any standard of content. That is a boundary rather than a result, and mistaking it for a result is how a shop spends three months writing cards against a number that was never going to move. So decide which shop you are before you interpret anything: high-consideration products generate pre-purchase chat, high repeat volume on thin margins generates support chat, and only the first is the case these measurements were built for.
Finally, do not produce a backlog you have no intention of working through. If this month has no writing time in it, skip the grouping and keep the fifteen minutes of reading. An unactioned list of gaps generates guilt and nothing else, and by next month it is stale enough that you distrust it and start again.
The cost side, in one paragraph
Cost per conversation is a pricing question and it is worked through properly there, including what happens when the vendor rather than you defines a conversation. What this page contributes is the denominator. Divide whatever you pay by the conversations you actually had this month, counted the way you counted everything else above rather than the way an invoice counted them, and set the result beside your average order value multiplied by your cart share. Two of those three numbers came out of reading rather than out of a bill, which is the only reason the comparison is worth anything.
The monthly routine, on a clock
Here is the whole thing with a time budget. About ninety minutes a month, mostly in small weekly pieces.
- Weekly, 15 minutes: read twenty transcripts, same slot every week.
- Weekly, 5 minutes: add whatever the assistant could not answer to a running list, grouped by subject as you go.
- Monthly, 20 minutes: count the four numbers, per channel, and write them next to last month's in the same place.
- Monthly, 10 minutes: sort the unanswered list by group size and pick the top three to write.
- Quarterly, 15 minutes: cost per conversation against what a conversation is worth, and a decision about continuing.
Ninety minutes is the budget, not the floor. If it starts costing three hours you have built a reporting habit rather than a shop habit, and the first thing to cut is the monthly counting, not the weekly reading. The numbers tell you whether to keep going. The transcripts tell you what to do on Monday.
Common questions
What is a good deflection rate for an ecommerce chatbot?
There is no benchmark worth quoting, and any figure you are handed deserves suspicion, because deflection rate counts every conversation that did not reach a human, including everyone who gave up and left. A frustrating assistant scores well on it. Use deflection for capacity planning, so you know how much human time the channel needs and at what volume you would have to hire, and judge quality on four other things: conversations that reached a cart, whether customers asked a second question, whether handoffs arrive with context, and what the assistant could not answer.
How long should I run an AI chatbot before deciding whether it works?
Four weeks. The first week measures your setup rather than the assistant: content is still wrong in ways you have not found, unexpected questions arrive, and you will change things mid-week. Judge it in week four against the same four-week period before you had it, not against week one, since that comparison mostly measures how much content you fixed. Put the review date in the calendar on day one so the decision is deliberate rather than a reaction to one good or bad transcript.
How do I know if my chatbot is actually driving sales?
Count the conversations that reached a cart or an order in the same session and treat that as a floor, not a total. Precise attribution is not achievable for a customer who chats on Tuesday, returns through an email link on Friday and buys on another device, and any tool reporting an exact revenue figure for that conversation has guessed. The floor is enough to decide with: compare it against what the tool costs you and against the same period before you had it.
Should I measure chatbot performance separately for each channel?
Yes, always, before drawing any conclusion. Website chat, WhatsApp and helpdesk tickets serve different populations: website visitors are mostly pre-purchase and anonymous, WhatsApp tends to be existing customers days after an order, and helpdesk tickets are the hard cases by definition. Averaged together they hide the broken one, and the usual pattern is a healthy website widget masking a WhatsApp assistant nobody has reviewed since it was connected. Leave your test-link conversations out of the counts entirely.
How many chat transcripts should I read, and how often?
Twenty a week, in a scheduled fifteen-minute slot on the same day. Read whole conversations rather than opening messages, and pick them without curating: if you can, take five that ended in a handoff and five that ran past six messages, and otherwise take twenty in a row. Fifteen minutes covers twenty conversations because most are short and obviously fine, and you are scanning for the two that are not. Real problems show up in transcripts weeks before they show up in any metric.