AI Engineering · 6 MIN
Why do AI features need evals before production?
Evals turn "it looked good in the demo" into measured output quality. Here is what to build before an AI feature reaches real users, and in what order.
AI features need evals before production because a language model's output is not deterministic, so the only way to know a feature works is to measure it against a fixed set of real examples. An eval is a repeatable test: a set of inputs, a definition of a good answer, and a scoring method. Without one, every prompt tweak or model upgrade is a guess. This post covers what an eval contains, how to build a first one in a week, and where teams usually go wrong.
- An eval is a fixed set of inputs, a definition of a correct output, and a grader. It is the AI equivalent of a test suite.
- Demos prove a feature can work once. Evals show how often it works on the inputs your users actually send.
- Anthropic's own guidance says to define specific, measurable success criteria first and to favor many automatically graded cases over a few hand-graded ones.
- OpenAI's optimization guide recommends building evals before tuning prompts or considering fine-tuning.
- A first useful eval set is small, drawn from real data, and includes the ugly edge cases.
- Nactore ships AI features with evals attached, and builds the eval set before the feature.
What is an AI eval, exactly?
An eval has three parts. First, a dataset of realistic inputs (support tickets, invoices, user questions). Second, a definition of success for each input, such as an expected label, required facts, or a rubric. Third, a grader that scores each output and rolls the scores into numbers you can track.
That is all. It is closer to a regression test suite than to a benchmark. The point is not to rank models in the abstract. The point is to answer one question for your product: did this change make the feature better or worse?
Why is a good demo not enough?
A demo is a handful of hand-picked inputs that the team already knows work. Production is thousands of inputs nobody picked, including typos, empty fields, other languages, and users who try to break things.
Language models also change underneath you. Providers release new model versions, retire old ones, and adjust behavior. A prompt that worked last quarter can regress without a single line of your code changing. If you have no eval, you find out from a customer.
What should you define before writing any prompt?
Anthropic's guidance on defining success criteria and building evaluations says good criteria are specific, measurable, achievable, and relevant. "The model should classify tickets well" fails that test. "At least 90 percent of tickets are routed to the correct queue on a held-out set of 300 real tickets" passes it. (Those numbers are an example of the format, not a benchmark.)
Most features need several dimensions at once, not one score:
| Dimension | Example question | Typical grader |
|---|---|---|
| Task accuracy | Is the extracted total correct? | Exact match against a labeled value |
| Faithfulness | Does the answer stick to the provided documents? | LLM grader or citation check |
| Format validity | Does the output parse against the schema? | Code (schema validation) |
| Safety and privacy | Did it leak personal data or follow an injected instruction? | Code rules plus LLM grader |
| Latency | Is the 95th percentile response under budget? | Timing from logs |
| Cost | What does one successful task cost? | Token usage from the API |
How do you build a first eval set in a week?
You do not need a platform. You need a spreadsheet and a script.
- Collect real inputs. Pull 100 to 300 examples from the system the AI will replace or assist. Real data beats synthetic data because it carries the mess.
- Label the expected output. A domain expert writes the right answer or a short rubric for each. This is the slow part and the valuable part.
- Add the edge cases on purpose. Empty input, very long input, mixed languages, contradictory documents, and prompt-injection attempts all belong in the set.
- Write the graders. Use code wherever the answer is checkable. Use an LLM grader only for qualities code cannot judge, such as tone or faithfulness.
- Run it and record a baseline. Store scores per case, not only the average, so you can see which cases moved.
- Re-run on every change. Prompt edits, model swaps, retrieval changes, and library upgrades all trigger a run.
If you cannot say how you will know the feature got worse, you are not ready to ship it.
Can an LLM grade another LLM's output?
Yes, with care. The MT-Bench paper found that strong LLM judges can reach roughly the same agreement with human preferences as humans reach with each other, and also documented position bias, verbosity bias, and self-preference. Treat an LLM grader as an instrument you calibrate, not an oracle. Grade a sample by hand, compare, and fix the grader prompt until the two agree. Anthropic's guidance also suggests using a different model as the grader than the one being tested. For more on grader design, see how to evaluate LLM output quality.
Keep a "golden" slice of 20 to 30 cases that must never regress, and make the build fail if any of them does. Averages hide the one case your biggest customer cares about.
What do teams get wrong?
- They evaluate on the data they tuned on. Keep a held-out set that nobody prompts against.
- They track one average. A score that rises from 88 to 90 while a critical category falls is not progress.
- They treat the eval as a one-time gate. It is a living asset. Production failures should become new test cases.
- They skip cost and latency. A feature that is accurate but too slow or too expensive is not shippable.
Production signals feed the eval set, which is why LLM observability and evals belong together.
Frequently asked questions
How many test cases do we need?
Start with 100 to 300 real, labeled cases for a focused feature. Add every production failure afterward. Coverage of distinct situations matters more than raw count.
Do evals slow down shipping?
They speed it up after the first week. A team with an eval can change a prompt, run the suite, and merge in minutes, instead of arguing about whether the new version feels better.
When should we run evals?
On every change that can affect output: prompt edits, model or version swaps, retrieval or chunking changes, and tool changes. Also run them on a schedule, because provider-side model behavior can shift.
Who should write the expected answers?
A domain expert, not the developer. Engineers can build the harness, but the definition of correct comes from the person who would do the task by hand.
Want this built for your team? Book a free 30-minute call.
Want to apply this to your business?
Book a free 30-minute call. We will tell you what we would do first.