How to stop an AI assistant making things up

Invention is two different failures wearing the same face. Most shops fix the wrong one and rewrite a policy that was already correct.

Ask a general language model when your returns window closes and it will answer. Not because it knows, but because producing the most plausible next sentence is the whole mechanism. It has read an enormous number of returns policies, and what comes out is a blend of them, phrased to sound like it came from you. Invention is two problems, though, not one. Sometimes the fact was never written down anywhere the assistant can reach. Sometimes the fact is sitting in your own knowledge and the assistant never pulled it up. The instinct is to assume the first and start rewriting content. When the real problem is the second, the rewrite changes nothing and the same wrong answer comes back a week later.

Invention is the design, not a defect the vendor will patch

It helps to stop calling this a malfunction. A model that refused to produce anything it could not verify would be useless for most of what chat actually is: phrasing, sizing, comparison, translation, tone. The same machinery that answers “is this warm enough for Berlin in November” in the customer's own language is the machinery that answers “what is your returns window” with a number nobody at your shop has ever said. Fluency and invention are one capability pointed at two different questions.

So the useful question is never which model invents least. All of them invent, and as models get more fluent the invented answers get harder to spot, not rarer. The question is how short the leash is: what the assistant is allowed to draw on, what it is required to do when it finds nothing there, and whether you can read afterwards what was actually in front of it. A shorter leash and a readable trace are the two levers you have. A better model is not one of them, and neither is a disclaimer.

Failure A: nobody ever wrote it down

The first failure is a content gap. The policy exists in your head, or in the way your shop manager handles it at the counter, or across three different Instagram replies from last spring. It is not in a document anywhere. The assistant looks, finds nothing, and fills the hole from the general average, because filling holes is the thing it does. The reply comes out fluent, specific, and entirely invented.

The fix is boring and it is the one you expect: write the fact down. What matters on this page is that you correctly identify that this is what happened. If you assume a content gap when the real problem is retrieval, you will write a second copy of a policy you already had and be no better off than before.

Failure B: it was written down and the assistant never saw it

The second failure looks identical from the customer's side and has nothing in common with the first. The fact is there. Somebody wrote a clear card six months ago saying sale items are exchange-only within fourteen days. The assistant composed its reply without that card in front of it, so it improvised, and it improvised confidently, because nothing tells a model that something is missing.

This is the easier of the two to miss. The content is correct, so a content review turns up nothing wrong, and the conclusion at the end of that review is usually that the assistant is just unreliable. The content was never the problem. Retrieval failed, and editing the card will not touch it.

Telling them apart in one specific bad reply

Do not reason about this in the abstract. Take one reply you know is wrong and ask exactly one question: can you find the correct fact in your own knowledge? Go and look. Search your cards for it the way you would search for anything, using the words the customer used rather than the words you filed it under.

If it is not there, you have a content gap and you write it. If it is there, in a card you can point at, then you have a retrieval failure and no amount of rewriting that card will help, because the card was never what broke. Same symptom, opposite fix, about a minute to tell them apart. Do this per reply. A general sense that the assistant is inaccurate cannot be acted on; one bad reply with a known cause can.

Retrieval turns mostly on how a thing is titled

The reason a correct card goes unretrieved is usually not that it was badly written. It is that it was badly labeled. Knowledge is pulled into a reply largely by how it is named and framed, not by how good the prose inside it is, so a card called “Terms and Conditions” that happens to contain your sale-item rule is invisible to a customer asking whether they can return a discounted jacket. The paragraph inside is perfect. Nothing ever asks for it.

That is why the diagnostic order matters: suspect retrieval before you suspect wording. When a reply is wrong and the fact exists, look at the title of the card that should have answered and ask whether a customer's own phrasing would ever land on it. Two symptoms show up here again and again: a title in your internal vocabulary rather than the customer's, and a title so broad that it covers half the shop. Naming and splitting cards properly is a craft with its own rules and its own page. The job here is narrower: stop misdiagnosing the failure.

Make “I do not know” a success condition

Left to itself, a model treats not answering as the worse outcome. Invert that. When the assistant cannot find the fact in what you wrote, saying so and fetching a person is the successful ending rather than the degraded one, and it has to be written down as such or the model will keep quietly ranking it last. How that fetch should be worded and where it should land belong to the handoff page. What belongs here is the standard it is measured against. A customer who reaches a human in ten seconds has been served. A customer given an invented delivery date has been given a promise your shop now has to honour or explain, made in writing, by you, to somebody who has no reason to doubt it.

Configure this explicitly, and configure it as three separate prohibitions, because an unknown has three tempting exits. No guess. No fallback to general knowledge about what shops like yours usually do. No hedged half-answer. The hedge is the worst of the three: “most stores in our category allow around thirty days” reads as caution to you and as authority to the customer, who will quote it back at you. Where that handoff goes, who gets alerted and how they are reached is a separate setup question with its own page.

Per-message traces are the only honest evidence

You cannot debug retrieval from a transcript. The transcript tells you what the assistant said. What you need is what it had: which knowledge cards and which catalog rows were actually in front of it at the moment it composed that specific reply. With that, failure A and failure B separate immediately, because you can see whether the right card was present and ignored or simply absent.

Reading one is a short job. Open the conversation, open the message that was wrong, and look at the knowledge that was supplied for that turn. Either the card that answers the question is in that list or it is not. If it is not, you are looking at retrieval and the title is the first suspect. If it is in the list and the reply still contradicts it, look for a second card that disagrees with the first, because the assistant was given two answers and picked one. That case is content again, and the fix is deleting a card rather than adding one.

A third thing shows up in traces and it is the one owners find hardest to accept. Sometimes the right source was retrieved, quoted faithfully, and the customer still got something untrue, because the source itself was wrong. The product row says the shelf holds 40kg because somebody typed 40 into a spreadsheet two years ago. The card says next-day dispatch because it did on the morning it was written. Nothing in the chain malfunctioned: the assistant read your shop out loud and your shop was inaccurate. Keep that case separate from the other two, because it is the only one where no change to the assistant fixes anything, and because it is the one that comes back every few weeks if you file it under invention and move on.

Starly keeps full transcripts and per-message traces, and you can read both. Whatever tool you use, that is the capability to check before you buy, and check it by asking to see a real trace rather than by reading a feature list. If a vendor cannot show you what the assistant was looking at when it produced a given sentence, then every conversation you ever have about accuracy ends in an assurance instead of an answer, and you will be rewriting content on a guess.

A standing monthly test of the money questions

Once a month, ask the same short list. Not “what do you sell”, which everything passes. Ask the ones where being wrong costs money:

These test the policy answer, not a lookup. Starly does not check the status of an individual order, so the late-parcel question is asking what your shop does about a late parcel and who the customer should talk to. If the assistant answers that one with a specific date, you have found an invention before a customer did.

Keep the wording identical month to month so you are comparing like with like, and run them somewhere customers cannot see: a test link is enough. Paste the five replies into a document with the date on it. The value of a fixed set is that it catches drift, and drift is invisible without last month's answer next to this month's. Someone edited a card, a product changed, a policy moved, and nothing announced any of it. Five questions, five minutes.

The failure nobody tests for: the slightly wrong paraphrase

A wholly invented answer is easy to spot. The dangerous one is the answer that is mostly right. Your policy says thirty days from delivery, unworn, receipt required. The reply says you have about a month to send it back. Nothing in that sentence sounds alarming, and two things in it are wrong: about a month is not thirty days from delivery, and the receipt has quietly disappeared.

Drift runs in both directions and both directions cost you. Slightly stricter loses a sale and earns a complaint. Slightly more generous is a promise the customer will hold you to, in writing, with a screenshot. So check the reply against the policy on numbers rather than on gist: the day count, the free-shipping threshold, the currency, the conditions attached. If your policy contains four numbers, count four numbers in the reply. This is the check most owners skip, because the answer read fine.

Why a disclaimer under the chat window fixes nothing

“Answers may be inaccurate” is not a control. Nobody reads it, it does not change what the customer was told about your returns window, and it will not help you in the argument that follows. The only thing it protects is the feeling that something was done. Put the effort into what the assistant is allowed to say instead.

What a short leash costs you

A short leash makes the assistant duller, and that is a real cost rather than a polite concession. It will hand off things it could probably have handled, and it will say it does not know to questions a general model would have answered acceptably. If you sell something where being wrong is an annoyance rather than a refund, you are paying in conversion for a safety margin you may not need.

Some shops should not install a customer-facing assistant at all yet, and the readiness page argues that case at length. Read in the terms of this one, it comes to this: a shop with nothing written down has only failure A, in every conversation, and no leash is short enough to help. The same is true when nobody can take the handoff, because an assistant configured to fetch a person is worse than no chat at all on the days no person arrives. And where a wrong answer in your category is a regulatory problem rather than a commercial one, dosage, medical suitability, anything legally binding, keep it out of a chat window at any leash length. The two-failure diagnosis on this page assumes a wrong answer is something you can afford to find afterwards.

Two limits on Starly specifically. Its sources are the imported catalog and the cards you typed, and there is nothing behind them, so if what you want is an assistant that discusses the category in general or answers about products you do not stock, this is the wrong tool and there is no setting that loosens it. And it will not tell you which of the two failures you are looking at. It gives you the transcript and the trace. Reading them and deciding is your job.

What good looks like after a month

The target is not zero invented answers. The realistic target is that when one happens you can tell within a couple of minutes whether the fact existed, read what the assistant actually had in front of it, and then make the correct one of two fixes. Get there and the same policy stops coming back every few weeks, because each fix lands on the thing that was broken. Miss it and you will keep editing content that was already right.

Common questions

Why does my AI chatbot make up answers about returns and shipping?

Because producing a plausible answer is what a language model is built to do. If your returns rule is not written down anywhere it can reach, it will produce a blend of the many returns policies it has read. There is a second cause that looks identical: the rule is written down, but the assistant did not retrieve it before replying, so it improvised anyway. Check whether the fact exists in your own knowledge before deciding which one you have.

Will a newer or more expensive model stop the invented answers?

No. Every language model invents, and a more fluent one produces inventions that are harder to catch, which is worse for you rather than better. What reduces invention is restricting the assistant to your own catalog and your own written knowledge, requiring it to hand off when it finds nothing there, and being able to read a per-message trace afterwards to see what it was working from.

How do I tell whether the assistant invented a fact or just failed to find it?

Take the bad reply and search your own knowledge for the correct fact. If it is not there, it is a content gap and you write the fact down. If it is there in a card you can point at, it is a retrieval failure, and rewriting that card will not help. Look instead at how the card is titled, since retrieval turns largely on titles and framing rather than on the quality of the text inside.

Is a disclaimer under the chat window enough to protect me?

No. A line saying answers may be inaccurate is not read by customers, does not change what they were told, and does not undo a delivery date or a refund promise the assistant invented. The control has to sit in what the assistant is able to say: a restricted source of facts, and a handoff to a person when the fact is not there.

What should the assistant do when it genuinely does not know?

Say so in one plain sentence and pass the conversation to a person in the same turn, without asking permission and without a hedged partial answer first. Ten seconds to a human is a good customer experience. A confident guess about a refund, a delivery date or an allergy is the outcome that costs you money, and a hedged half-answer is worse than both because it reads as authoritative.

The reply was close to my policy but not exact. Does that count?

Yes, and it is the version most owners let through. Compare the numbers rather than the gist: the day count, the threshold, the currency, the conditions attached. A policy of thirty days from delivery with a receipt is not the same promise as about a month, and the customer with the screenshot will be holding you to whichever version was more generous.