Why the demo impresses and the product disappoints

A demo is built around the cases that work. The product meets all of them. What to check before you trust an AI demo.

You watched a demo of an AI tool, everything ran smoothly, and you decided to buy. After a week of real work it no longer looks the same. That does not necessarily mean you were misled. A demo is one run on a case chosen in advance, and that is exactly what it does best.

What is different in real work?

First, the input. In a demo someone types a clean sentence. In real work you get requests in other languages, with typos, a one-word request like "returns", or a long conversation where the important thing was said twenty messages ago. OpenAI lists exactly these kinds of input among the edge cases to test.

Second, variability. OpenAI writes that generative AI is variable: the same input can produce different output. A run that worked in a demo does not promise the next twenty will.

Third, selection. Whoever prepares a demo picks the cases that show the tool at its best. That is not cheating. It is what a demonstration is.

What do the people who build the tools say?

OpenAI's guide to evals warns against a few mistakes: building a test set that does not reflect real traffic, relying on a feeling like "it seems like it's working", and waiting until you ship before you test. Anthropic's engineering team suggests starting with a few dozen tasks taken from real failures.

What do you ask for before you trust a demo?

This is my own list, not a standard.

Run the tool on twenty cases of your own, including a few messy ones. The guide on 20 cases before launch shows how to build a set like that.

Run each case several times and see whether the answer is stable.

Send short requests, requests with typos, and requests in the language your customers use.

And ask the vendor directly: where does it fail, and what does a failure look like? A vendor with no answer to that has given you an answer.

So what do you do with this?

A demo is an invitation to start testing, not a reason to stop. If the tool holds up on your twenty cases, you have far better reasons to trust it than any demonstration gives you.

Sources and method