AI chat solutions for online stores, compared honestly

Four different products get sold under the same phrase. Working out which one you are looking at tells you more than any feature table, and one question about how a conversation should end settles most of the decision.

By · · 10 min read

Almost every tool sold as AI chat for an online store is one of four things wearing the same words: a help-center search box, a general assistant with a copy of your site pasted into it, an AI feature bolted onto a helpdesk, or an assistant wired into your product catalog. They fail in four different places. Which one a vendor is showing you matters more than anything on a comparison grid, and it is rarely stated outright. Below are the four categories, three questions that identify which one a demo is, and one question about how a conversation should end that settles most of the rest.

1. The help-center search box

It matches the words a visitor typed against articles you already wrote and hands back the closest one. The pro: nothing it says was written by a machine, so it cannot invent a fact or a price. Every sentence was approved by a person before it shipped. If you keep a help center current and your product line changes a few times a year, this deflects a real share of repeat questions with close to no risk of an answer you have to apologize for. Whether it matches keywords or uses a language model to rank the same articles is a mechanism question, not a category question.

The specific thing it fails at is the question nobody wrote an article about, phrased the way customers phrase things. "Will the medium fit a 42-inch chest" when your sizing information is a table. "Can I use the 15 percent code on the sale bundle." There is nothing to match, so it hands back the three nearest articles and lets the customer sort it out. The ceiling is your article library, and the library ages the way your product pages age. Every new question you want covered is another article somebody has to sit down and write.

2. The general assistant with your site pasted in

A language model handed a copy of your pages as context. The pro is that it handles real phrasing on day one, in whatever language the customer typed, and it answers questions you never anticipated as long as the answer sits somewhere in the text it was given. Setup is an afternoon: point it at a URL, wait for the crawl, paste in an embed snippet. That is why so many of them exist.

The specific thing it fails at is anything that changed after the copy was taken. It will quote the price you cut last Tuesday, describe a color you discontinued in spring, and sound equally certain about both, because a snapshot does not know it is a snapshot. Ask a demo how often the content refreshes and what triggers a refresh. If the answer is monthly, every price change carries up to 30 days of wrong answers behind it. The second failure is quieter: a crawl reads what your pages say, and your pages do not say the return window is 14 days if that line only exists in an email your team keeps re-sending.

3. The helpdesk add-on

AI attached to the ticketing system your team already works in. The pro is organizational rather than technical. Replies land in the queue your agents already watch, escalation to a person is native instead of bolted on, and the macros written over the last two years become answer content on day one. Nothing else on this list starts with that much material already written, and nothing else keeps one customer's history in one place.

The specific thing it fails at is the pre-purchase question. Its job is to close a ticket, so when a question is ambiguous its instinct is to resolve the conversation rather than extend it. Ask it "which of these two should I get for a six-month-old" and watch what happens. It points at a buying guide and marks the ticket handled. That is the right outcome for a support tool and a lost sale for a store. It also has nowhere for a conversation to end except a closed thread; a queue is built to finish conversations, not to sell anything.

4. The catalog-aware assistant

It reads your product data as data, next to policies you wrote, and it can do something at the end of the conversation: put one specific variant in the cart, quote a discount code that exists in the store, capture a lead and route it to a person. The pro is that it can finish a sale rather than describe one, and prices and variants come from product records rather than from a copy of your pages.

The specific thing it fails at is anything your catalog does not say. If your descriptions never mention material, neither will the assistant. If two variants are both named "Standard," it picks the wrong one about as often as a confused customer would. Messy product data shows up in the first ten answers, not in month three. This category also needs a written policy layer more than the others do, because a product record knows the price and knows nothing about your return window, your lead time on custom orders, or which two items you refuse to ship in the same box.

Three questions that identify the category in a demo

All four categories use the same words on a homepage. Three questions asked of a live demo separate them in about five minutes. Ask them of the product, not of the salesperson, and insist on seeing the answer rather than hearing about it.

The first question matters longest. A tool that cannot tell you where an answer came from cannot be debugged, and you spend the first month guessing why one reply in twenty was wrong. If it can show you which article, card or product record sat behind a reply, you fix one line of content and move on. Ask to see that view yourself during the demo rather than being told it exists.

The third question is the one vendors are least prepared for, so ask it with a real example: something a customer asks weekly that is nowhere in your written material. Three outcomes are possible and they are not equally bad. Saying it does not know and offering a person is fine. Reasoning from adjacent facts and labeling the reasoning is fine. Producing a confident specific number that appears nowhere in your data is the disqualifying one, and it is invisible in a scripted demo, because a scripted demo only asks questions with written answers. You have to bring the question yourself.

Match by shop shape, not by feature count

Five customer questions a week points at the help-center search box, or at a well-written FAQ page and nothing else. At that volume you spend more hours writing answer content than you currently spend answering people, and the tool never sees enough traffic to show you what it gets wrong. Three thousand SKUs is the opposite case and points straight at the catalog-aware category. Nobody is going to write articles for three thousand products, and a page snapshot of a catalog that size is out of date within a week of being taken.

A B2B site with ten pages of capability copy and no catalog to speak of should look at the general assistant first, and judge it on what happens after a conversation rather than during one: whether a lead reaches a named person, and where one submitted at midnight actually lands. How many fields to ask for and when to ask is a decision about your visitors rather than about the category, and it is made on the lead capture page. A shop whose support already runs inside a helpdesk queue, with tagging and assignment and a second agent, should look at the add-on for that helpdesk first. Not because it answers better, but because the alternative is two inboxes and one customer's history split across both.

One arrangement is worth naming, because four tidy categories make it sound impossible: running two of them on purpose. A shop whose pre-purchase and post-purchase traffic look nothing alike can put a catalog-aware assistant on the storefront, where the job is to finish a sale, and keep the helpdesk add-on on email and tickets, where the job is to close a thread. The cost of that is real and it is not the second subscription. It is that one customer's history now lives in two systems, and somebody has to decide which of them owns a conversation that begins as a sizing question and becomes a complaint. If you cannot answer that in one sentence today, run one tool imperfectly rather than two tools ambiguously.

The switching cost nobody asks about

Feature comparisons assume you buy once. Assume instead that you will look at this decision again in a year, and ask what you would carry out of the building. The answer content is the asset. If you write 40 short entries covering shipping, sizing, returns and the two products customers always confuse, those 40 entries are worth more next year than any feature you compared this week. Ask, before you pay, what format they are stored in and whether you can read them back as plain text.

Three things to check specifically. First, whether the knowledge you write is plain sentences you own or structure locked inside a proprietary editor. Second, whether conversation transcripts are visible and retrievable, and how far back, since transcripts are the only record of what customers actually asked and they are what you would use to set up the next tool. Third, how much of the work transfers in principle: content written as prose moves anywhere, while decision trees, tagging schemes and macro hierarchies move nowhere. A tool that ingests your site and keeps everything derived from it in its own shape leaves you nothing to take with you.

That applies to us as well. Starly's knowledge is cards you write in plain text and can read back at any time, and transcripts plus per-message traces are visible to you in the dashboard. If you set the cards up by connecting Claude, ChatGPT or Cursor over MCP, the assistant writes those same plain-text cards, so the output is still text you can read and reuse. Hold every vendor you talk to, including this one, to the same two questions: can I read my content, and can I read my transcripts.

Where Starly is the wrong choice

Four cases, stated as boundaries rather than argued again. Customers cannot look up their own order status here, and nothing sends an automated back-in-stock alert; whether the absence of those two disqualifies you is decided on the readiness page and not by picking a different category off this list. If your support runs on multi-agent ticket workflows with SLAs, assignment rules and escalation tiers, buy the AI add-on for the helpdesk you already run and keep everything in one queue rather than splitting it.

If you need phone, SMS or Instagram DMs, Starly has none of them. What exists is a website widget, your own WhatsApp number, Gorgias tickets and shareable test links. And if the real problem is that nobody is answering the email inbox, no assistant in any of these four categories fixes it. It answers the easy half and hands the hard half to the same unattended inbox, which leaves the customer waiting twice instead of once.

The one question that decides it

What do you want the conversation to end in? A closed ticket points at the helpdesk add-on. An order points at the catalog-aware assistant. A deflected repeat question with no risk attached points at the help-center search box. A qualified lead on a site with no catalog points at the general assistant. Feature tables cannot answer this for you, because all four categories list most of the same features.

If the honest answer is "a person," buy a better inbox instead of a bot. A shared inbox with assignment and a two-hour response target beats every category on this page for a shop whose actual failure is that messages sit unread until Monday. Fix the answering first. An assistant is worth buying once the answering already happens reliably and you want it to also happen while your team is asleep, in the customer's language, and to end in a cart.

Common questions

How do I tell which kind of AI chat tool a vendor is selling me?

Ask three questions during a live demo. Where did that specific answer come from, an article, a page snapshot, a saved macro, or the product record? What would it have told a customer this morning if you changed a price yesterday, and how long does a change take to reach the answer? And what does it do with a question that is true about the product but absent from every page on your site? The answers place a tool in one of four categories: help-center search box, general assistant with your site pasted in, helpdesk add-on, or catalog-aware assistant.

Which category fits a shop that already runs support inside a helpdesk?

The AI add-on for that helpdesk, first. Replies land in the queue your agents already watch, escalation is native, and your existing macros become answer content immediately. The trade is pre-purchase questions: a ticket tool is built to close threads, so an ambiguous "which of these should I buy" tends to end in a buying-guide link and a handled ticket. If a real share of your chat volume is people deciding whether to order, look at the catalog-aware category for the storefront and keep the helpdesk for post-purchase.

Which AI chat tool fits a store with thousands of products?

A catalog-aware assistant, meaning one that reads product data rather than a copy of your pages. At that catalog size nobody will write help articles per product, and a site snapshot goes out of date within about a week, so prices and variants in the answers drift away from the store. The trade is that answers can only be as good as the product data behind them: if descriptions omit the material or two variants share a name, the assistant will be wrong in exactly those places.

Can I move my chatbot answer content to another tool later?

It depends on the format, and this is worth asking before you pay rather than after. Answer content written as plain sentences you can read back moves to another tool in an afternoon. Decision trees, tag hierarchies and macro structures built inside one vendor's editor do not transfer at all. Ask two questions of any vendor: can I read my written knowledge back as text, and can I see and retrieve conversation transcripts, since those transcripts are the only record of what customers actually asked.

When is a shared inbox a better buy than an AI chatbot?

When the real problem is that nobody is answering. If customer emails and messages currently sit unread until Monday, adding an assistant handles the easy half and passes the hard half to the same unattended inbox, so the customer waits twice. A shared inbox with assignment and a response target fixes the underlying failure. An assistant becomes worth buying once answering reliably happens and you want it to also happen outside your hours, in another language, or to end in a cart rather than a reply.