Buyer Guides · 5 MIN

How to evaluate an AI demo before you buy

A buyer checklist for judging an AI demo: ask for held-out test results, failure cases, live inputs, cost per item, and what happens when the model is wrong.

By NactorePublished 12 Sep 2026All articles

Judge an AI demo by what it shows about failure, not success. A polished demo proves a system can work on chosen inputs. What you need to know is how often it works on yours, what it does when it is wrong, and what it costs to run. Ask for measured results on held-out data, live input from you, and an honest account of the failures.

Key takeaways
  • A demo shows the best case on curated inputs, so ask to see the failures.
  • Insist on measured output quality (evals) on data the builder did not tune against.
  • Bring your own inputs and run them live, including awkward ones.
  • Ask for cost and latency per item, not only accuracy.
  • Probe what happens when the model is wrong, uncertain, or attacked.
  • Nactore ships to production with evals, so we expect buyers to ask for the numbers behind any demo.

Why are AI demos so easy to get wrong?

Language models are fluent by default, so a demo almost always looks good. The same system that handles ten hand-picked examples can fail on a meaningful share of real traffic, and a demo will not reveal that unless you ask.

Three things make demos misleading. The inputs are chosen by the presenter. The system may have been tuned on exactly those examples. And a single impressive answer says nothing about consistency across thousands of runs. Your job as a buyer is to replace anecdote with measurement.

What evidence should you ask for?

Request these items in the order below. A serious team will have most of them ready.

  1. A scored test set. Ask how many examples, who labeled them, and what the scoring rule is.
  2. Held-out results. Confirm the test items were kept separate from anything used to tune prompts or models.
  3. Failure categories. Ask for the top ways it fails, with real examples.
  4. Cost and latency per item. Accuracy at an unaffordable or slow price is not a result.
  5. A change history. How scores moved as the system improved, which shows whether the team measures at all.

For the engineering side of this, see how to evaluate LLM output quality.

How do you test a demo yourself?

Bring a small set of your own inputs and run them live. Include some that are messy, ambiguous, or deliberately hard.

TestWhat you learn
A normal, typical inputBaseline quality on your real work
A messy or incomplete inputWhether it copes with production data
An input outside its scopeWhether it says "I cannot do this" or invents an answer
The same input run three timesConsistency
An adversarial or odd instructionBasic resistance to misuse
A large batchSpeed and cost at volume

Pay attention to confidence. A good system flags uncertainty and routes hard cases to a person. A poor one answers everything with the same assurance.

What questions expose a weak demo?

The following questions tend to separate real systems from stage-managed ones.

  • What was this tuned on? If the answer includes your sample inputs, the result is inflated.
  • What happens when it is wrong? Look for review queues, fallbacks, and logging, not "it rarely is."
  • Where does our data go? Ask which model providers see it and under what terms. Have your counsel review the written terms.
  • Can we see the prompts and eval set? You should be able to rerun the measurement later.
  • What does it cost at our volume? Ask for a per-item estimate based on your sample, not a general figure.
  • What would you not use this for? A confident team can name the limits.
Pro tip

If a vendor cannot show you a single failure case, they have either not tested enough or will not tell you. Both are reasons to slow down.

What does a demo not tell you?

Even a strong demo leaves gaps. It rarely shows integration with your systems, behavior under sustained load, monitoring after launch, or how the system degrades when a model provider changes something. Those are production questions, and they are why we favor a short pilot over buying on a demo alone. The checklist for that is in how to scope an AI pilot.

How should you score what you see?

Use a simple scorecard so the decision does not rest on the most charismatic presenter. Rate each area from one to five, with notes.

  1. Quality on your inputs. Measured, not claimed.
  2. Handling of failure. Uncertainty flags, fallbacks, human review.
  3. Cost and speed. Per item, at your volume.
  4. Data handling. Where data flows, retained or not, and under what terms.
  5. Path to production. Monitoring, ownership, and the plan after the demo.

A system that scores well on quality but poorly on failure handling should not go live unattended.

Frequently asked questions

Is a demo ever enough to make a buying decision?

Rarely. A demo is a reason to run a pilot, not a substitute for one. It helps you shortlist, and a measured pilot on your data should decide.

What is a held-out test set?

It is a set of examples kept separate from anything used to build or tune the system. Scores on it tell you how the system handles inputs it has not seen, which is closer to production.

How many test inputs do we need?

Enough that you can see patterns in the failures. For an early evaluation, a hundred or so representative items, labeled by your experts, is a common starting point. Your risk tolerance should set the final number.

Should we worry if the demo uses a well-known model?

The model matters less than the system around it. Prompts, retrieval, guardrails, and evals determine most of the quality you experience.

Ask for the numbers

The most useful single request you can make is "show me the scores and the failures." Teams that build with evals answer easily. Teams that do not will change the subject.

Want this built for your team? Book a free 30-minute call.

Want to apply this to your business?

Book a free 30-minute call. We will tell you what we would do first.