Buyer Guides · 7 MIN

How to scope an AI pilot so it ends in a decision

A practical checklist for scoping a 4-week AI pilot: one workflow, a measurable bar, real data, a named owner, and a written decision at the end.

By NactorePublished 31 Aug 2026All articles

A good AI pilot is one workflow, one measurable quality bar, real data, a named owner on your side, and a fixed end date that forces a go or no-go decision. If any of those five is missing, the pilot will produce a demo and a debate instead of an answer. This guide gives you the scoping checklist we would want a buyer to bring to us.

Key takeaways
  • Scope a pilot around one workflow with a clear start, end, and owner, not around a technology.
  • Write the pass bar as a measurable output-quality threshold before any code exists.
  • Real data access on day one is the most common schedule risk, so confirm it before you sign.
  • A pilot should end with a written decision: ship, extend with changes, or stop.
  • Four weeks is enough to answer "does this work on our data" and not enough to build a full platform.
  • Nactore builds AI work with evals, scoped to each team, and expects a clear go or no-go decision before you ship.

What should an AI pilot actually prove?

A pilot exists to answer one question that a slide deck cannot: does this approach work well enough on our data, inside our constraints, to justify building it properly? It is not a smaller version of the final product. It is an experiment with a result.

That framing changes what you scope. You are not buying features. You are buying evidence that lets you decide with confidence, so everything in scope should serve that decision.

Which workflow should you pick for the pilot?

Pick a workflow that is frequent, painful, and bounded. Frequent means enough volume to measure. Painful means someone will notice if it improves. Bounded means a clear input, a clear output, and a person who can say whether the output is right.

A short test for candidate workflows:

QuestionGood signWarning sign
Is there a clear input and output?A ticket in, a category and draft reply out"Make the whole team smarter"
Can a human judge correctness?An expert can grade 50 samples in an afternoonQuality is a matter of taste
Is the data reachable?Exports or an API exist todayData lives in someone's inbox
Is the blast radius small?Errors are caught before a customer sees themA wrong answer triggers a payment or a legal notice
Does one person own it?A named operations or product leadA committee with no tiebreaker

If you are unsure which of your workflows scores best, read which workflows to automate first.

How do you write a pass bar before building anything?

Write the success criteria as a measurable threshold, agreed in writing, before the first line of code. Otherwise the bar moves to fit whatever the prototype can do.

Good pass bars are specific about the measure and the sample. For example, "at least this share of a 100-item expert-labeled test set is handled correctly without human edits, and the remaining items are routed to a person." You pick the share with your own domain experts, because only you know what an acceptable error costs.

This measured output quality is what we call evals. The mechanics are covered in AI evals before production. For scoping, you only need three things agreed up front.

  1. The test set. A sample of real, representative inputs with the correct answers labeled by your team.
  2. The metric. What counts as correct, partially correct, and wrong.
  3. The threshold. The score at which you would ship, and the score at which you would stop.

What data and access does the pilot need on day one?

Most pilots slip because of access, not because of models. Before the clock starts, confirm the following.

  • Sample data. A representative export of real inputs, with sensitive fields handled according to your policy.
  • System access. Read access, or a sandbox, to the systems the workflow touches.
  • A subject-matter expert. One person with a few hours per week to label samples and review outputs.
  • Security sign-off. Your approval for which model providers and data paths are acceptable.

If any of these takes more than a week to arrange, say so at the start and move the start date instead of compressing the work.

What does a 4-week pilot look like week by week?

A fixed four-week window keeps the work honest. This is a reasonable shape, and you should expect your partner to propose something similar.

  1. Week 1: Baseline. Lock the test set, metric, and threshold. Connect to data. Measure how the workflow performs today, by humans or by existing rules.
  2. Week 2: First working version. Build the simplest approach that could pass. Run it against the test set and publish the first score.
  3. Week 3: Improve against the evals. Fix the failure categories the scores reveal. Add guardrails and the human-review path.
  4. Week 4: Decision. Final scores, cost per item, failure analysis, and a written recommendation with a production plan if the answer is yes.
Pro tip

Ask for a score on the test set at the end of week 2, not week 4. A partner who cannot show you any number halfway through is not running a pilot.

What should be explicitly out of scope?

Write down what the pilot will not do. This protects both sides. Typical exclusions are a polished interface, full integration into every system, multi-language support, high-volume load testing, and edge cases below a stated frequency.

Out-of-scope items are not forgotten. They go on a list for the production phase, where you can price and plan them with real results in hand. For more on how fixed scope compares to open-ended billing, see fixed-scope AI pilot vs time and materials.

What should you have in hand when the pilot ends?

You should leave with more than a prototype. Ask for these deliverables in the scope document.

  • The scorecard. Results against the agreed threshold, with failure categories.
  • The test set and eval harness. So you can rerun the measurement yourself later.
  • A cost and latency profile. What each processed item costs and how long it takes.
  • A risk list. Data, security, and quality risks found, with proposed mitigations.
  • A recommendation. Ship, extend with named changes, or stop, with reasoning.

Frequently asked questions

How long should an AI pilot take?

Four weeks is enough to test one bounded workflow on real data and reach a decision. Longer pilots tend to drift into building the product without a decision gate, which defeats the purpose.

What if the pilot fails the bar?

That is a valid and useful result. You learned in weeks, at a fixed cost, that this approach does not work for this workflow, or that the data needs work first. A written stop recommendation is a successful pilot outcome.

Do we need clean data before starting?

No, but you need representative data. Messy real inputs are better than tidy samples, because production will be messy. Plan for some cleaning inside the pilot.

Who on our side needs to be involved?

At minimum, one workflow owner who decides, one subject-matter expert who labels samples, and one technical contact who can grant access.

Scope for a decision, not a demo

A pilot that ends in a written go or no-go saves you from the most expensive outcome, a half-built system nobody trusts. Scope one workflow, set the bar first, secure the data, and hold the date.

Want this built for your team? Book a free 30-minute call.

Want to apply this to your business?

Book a free 30-minute call. We will tell you what we would do first.