Field Notes · 7 MIN

What is the checklist for taking an AI prototype to production?

A production checklist for AI features built from failures we hit, covering evals, logging, cost, infrastructure, review gates, and how to verify before you report.

By NactorePublished 21 Aug 2026All articles

A prototype becomes production when it has measured quality, visible failures, a known cost, safe infrastructure, a review gate, and a habit of checking facts before reporting them. This checklist comes from failures we actually hit, not from a template. Use it before launch, and treat every unchecked line as a decision you are making on purpose.

Key takeaways
  • A demo shows the best case. Production needs an eval that shows the distribution, including wrong answers and abstentions.
  • Logging, timeouts, and cold-start behavior fail quietly, so set them up before launch, not after the first incident.
  • Measure the model's share of cost before optimizing it, and check infrastructure quotas that can switch a whole stack off.
  • Gate interface changes behind a reviewer looking at real screenshots, and verify any fact before you report it.
  • Nactore is an AI-native software engineering partner, and we ship AI features to production with evals instead of demos.

What separates a prototype from production?

A prototype answers "can this work?" Production answers "what happens when it does not?" The gap is rarely the model. It is everything that makes failure visible, bounded, and recoverable. The checklist below is ordered by what hurt us most in practice.

AreaPrototype behaviorProduction requirement
QualityA few good examplesAn eval with right, wrong, and no-answer rates
FailuresSilent or logged to a consoleExplicit logging and clear user messages
CostUnknownModel share of each transaction measured
InfrastructureFree tier, default settingsKnown quotas, secrets in the host, tested restarts
InterfaceLooks fine in codeReviewed on real screenshots
ReportingConfident statementsFacts re-checked in isolation

1. How do you prove quality before launch?

Build an evaluation set before launch and report three numbers separately. Right, wrong, and no answer. A wrong answer that sounds confident is a different failure from an abstention, and they cost the user differently. We describe the full approach in an attribution engine with evals and the general method in AI evals before production.

  • Say what the set covers. A small set gives a number about that set. Never quote it as a general rate.
  • Allow abstention. If the system cannot say "I do not know," it will guess.
  • Re-run on every change. Diff per-case results, not only totals.

2. Where do failures go?

They go nowhere unless you build the path. In one production build, Django logged no tracebacks with DEBUG=False because, per the Django logging documentation, the default configuration only displays log records when DEBUG=True. Every other bug became harder to find.

  1. Configure logging first. Send application and framework logs to stdout at an error-inclusive level before the first deploy.
  2. Log third-party failure bodies. A bare status code often says nothing. A mail API returned a bare 403 until we logged the response body.
  3. Set timeouts and retries on model calls. Decide what the user sees when a call fails or is slow.
  4. Record model inputs and outputs in a way you can search. See LLM observability: what to log.

3. What does the model call actually cost?

Measure it as a share of what one successful transaction earns, before you optimize. In one paid product the model cost was a tiny fraction of each sale, and the real constraint was acquiring customers. If your ratio looks like that, spend effort elsewhere. If your ratio is large, start with LLM cost control and prompt caching.

Also check the settings that quietly change cost. Reasoning modes can be on by default and bill as output tokens, as covered in LLM settings for temperature and cost. Decide per call whether the task needs them.

4. What infrastructure traps should you check?

These are the ones that took things down for us.

  • Quota-capped free tiers. A serverless database with an automatic usage quota refused to start after the quota ran out, and the whole stack went down with it. Production should not rest on a dependency that can be switched off by usage.
  • Idle pauses and cold starts. Free tiers can pause or sleep when idle. Test the first request after a quiet period.
  • Address families. A database host reachable from a laptop can be unreachable from the app host if it is IPv6-only. Use the pooler or an IPv4 endpoint.
  • Vendor SDKs. One payment SDK failed to import on a modern runtime. When the protocol is small, plain HTTP and the standard library can be less code. See what breaks when you ship an AI product with payments.
  • Env parsing. Our env reader treated the whole line as the value, so an inline comment broke a flag silently.
  • Secrets. Live keys only in the production host's environment, never in the repository, and public values only in client bundles.
  • Firewall and crawlers. A firewall rule returned errors to AI crawlers while serving Googlebot normally. Test with each crawler's user agent, as covered in Cloudflare for AI products.

5. How do you gate interface changes?

Do not judge an interface from its code. We had a first version of an owner console rejected by the founder, and a review of real screenshots at desktop and phone widths found problems that code review had missed, including a stray typeface from marketing styles and global CSS collisions that only appear on screen.

Since then, no interface ships without a PASS from a design reviewer who screenshots the real build at 1440 and 390 pixel widths and scores it against a rubric. The loop is fix, re-review, repeat until pass. In our case, one console went through several rounds before it passed. Scoping the new CSS under its own container, with explicit resets, stopped styles from leaking in from other parts of the site.

6. How do you stop reporting wrong facts?

This one is about people, and it cost us credibility twice. Both times, one command and one glance became a claim.

  • A shell pipeline bound tighter than intended and reported a credential that did not exist.
  • A tool returned a cached listing, and acting on it sent someone to fix something that was not broken.

The rule is to re-run the narrowest possible check on its own, with no pipe and no compound operator, before stating that something exists or does not. When a listing tool disagrees with what you expect, try the real operation before asking anyone to change configuration. And when a report turns out wrong, correct it in one plain sentence and keep going.

The same discipline applies to measurement. Query the real system you claim to report on, from a cold session, and save the evidence, as in measuring AI answers with Playwright.

Pro tip

Make "I did not find it" and "it does not exist" two different sentences. Most wrong reports we have seen, ours included, came from treating the first as the second.

What does a launch review look like?

We run it as a short, written pass before any production release.

  1. Eval results with right, wrong, and no-answer counts and the sample size.
  2. Logging confirmed by forcing an error and reading it.
  3. Cost share of the model call in one transaction.
  4. Infrastructure quotas listed, with the ones that can halt service marked.
  5. Interface review result with screenshots.
  6. Facts rechecked for any claim that will be reported to a customer.

For scoping the work before it gets this far, read how to scope an AI pilot.

Frequently asked questions

How long does it take to harden a prototype?

It depends on what the prototype skipped. Evals, logging, and infrastructure checks are usually the bulk of it, and the checklist above shows where to look first.

Do I need evals if the demo works?

Yes. A demo shows the best case. An eval shows how often the system is right, wrong, or silent across realistic inputs.

What is the most common production failure you see?

Quiet failures around the model, such as missing logs, a dead integration, or a quota that shuts a dependency off, rather than bad model output.

Can I skip the interface review for an internal tool?

We would not. Problems that only show on screen affect internal users too, and a review is cheap compared with a rejected release.

Want this built for your team? Book a free 30-minute call.

Want to apply this to your business?

Book a free 30-minute call. We will tell you what we would do first.