AI Engineering · 6 MIN

How do you choose an LLM for production?

Choose a production LLM by testing candidates on your own eval set across quality, latency, cost, and risk, not by leaderboard rank. A step-by-step method.

By NactorePublished 18 Sep 2026All articles

You choose an LLM for production by running two or three candidates on your own eval set and comparing quality, latency, and cost per successful task, then picking the cheapest model that clears your quality bar. Public leaderboards are a way to build a shortlist and not a way to make the decision, because they do not use your data, your prompts, or your edge cases. This post gives a selection method, the criteria that matter, and how to keep the choice reversible.

Key takeaways
  • Build the eval set first. Anthropic's selection guide calls a good evaluation set the most important step.
  • Shortlist two or three models from different tiers, then test them on your real prompts and data.
  • Compare cost per successful task, including retries and escalations, not price per token.
  • Settings like reasoning effort can move cost and latency within one model, and may be a better lever than switching models.
  • Mixed setups, where a small model handles bulk work and a stronger model handles hard cases, are common and often cheaper.
  • Nactore builds model-agnostic pipelines with an eval harness, so swapping a model is a measured, same-day change.

What criteria should drive the choice?

Anthropic's guide to choosing a model frames it as balancing capability, speed, and cost, and adds effort settings as a fourth lever. In practice, add two more that matter in a company setting.

CriterionQuestion to answerHow to measure
QualityDoes it clear the bar on our eval set?Pass rate per category
LatencyIs time to first token and total time acceptable for the user?Median and 95th percentile
CostWhat does one successful task cost?Tokens times price, including retries
CapabilitiesTool use, structured output, vision, long context needed?Feature checklist plus tests
ReliabilityRate limits, uptime, regional availabilityProvider docs and load test
Data termsRetention, training use, region, certificationsContract and data policy
Lock-in riskHow hard is it to switch?Abstraction layer and eval harness

The last two are not in a benchmark. For US, UK, and EU customers, data residency and retention terms can rule a model in or out before quality is even compared.

How do you run the comparison?

  1. Write the eval set. Real, labeled cases, including edge cases. See why AI features need evals before production.
  2. Pick candidates. One strong model, one mid-tier model, and one small model is a good spread. Include a model from a second provider if lock-in matters.
  3. Hold everything else constant. Same prompts, same retrieval, same schema. Then tune each prompt separately, because models respond differently to wording.
  4. Run and record. Quality per category, latency, token usage, and cost.
  5. Look at failures, not just averages. Two models with the same score can fail on different cases. One failure mode may be tolerable and another not.
  6. Decide on cost per successful task. A model that needs two retries to match a pricier model's accuracy is not cheaper.

Should you start with the best model or the cheapest one?

Anthropic's guide describes both. An efficiency-first start begins with a faster, lower-cost model, tests it thoroughly, and upgrades only if there is a capability gap. A capability-first start begins with the strongest model for the task and then optimizes downward through lower effort or smaller models once the workflow is stable.

Use efficiency-first for high-volume, straightforward work such as classification, routing, and extraction. Use capability-first for complex reasoning or long-running agent work, where you want to learn what is possible before optimizing. Either way, the decision is made by your evals, not by the label on the model.

Is switching models better than tuning settings?

Not always. The same guide notes that some models expose an effort parameter that trades intelligence for latency and cost within a single model, and that tuning it is often a better lever than switching. Check your prompt first, then the settings, then the model. Also consider prompt caching, which can cut input cost for long stable prompts. See LLM cost control and prompt caching.

When does a multi-model setup make sense?

When task difficulty varies a lot. A router or a first-pass model handles the easy majority, and hard cases escalate to a stronger model. Anthropic's guide also describes patterns where a lower-cost model does bulk work under a more capable orchestrator or advisor. Build this only after you have a single-model baseline, because routing adds complexity and its own failure modes. Your eval set should include the cases a router might misjudge.

Pro tip

Never choose a model from a demo prompt. Pick the five ugliest real inputs you have, run every candidate on them, and read the outputs yourself before looking at any score.

How do you keep the choice reversible?

Models are retired, repriced, and replaced often. Design for change.

  • Put an abstraction layer between your code and the provider, limited to the features you actually use.
  • Keep prompts, schemas, and graders in version control with the eval set.
  • Pin model versions where the provider allows it, and schedule upgrades deliberately.
  • Re-run the eval suite on every upgrade, and on a schedule, since provider-side behavior can shift. Catching regressions is the job of LLM observability plus evals.
  • Keep a tested fallback model for outages and rate limits.

What about open-weight models?

They can fit when data cannot leave your environment, when volume makes per-token pricing painful, or when you need to customize behavior. They also shift work to you: hosting, scaling, monitoring, and security updates. Compare the total cost of ownership, not only the token price, and test them on the same eval set as hosted models. If the main driver is specialized behavior, read RAG or fine-tuning first.

Frequently asked questions

How many models should we test?

Two to four is typical: a strong model, a mid-tier model, a small model, and optionally one from a second provider. More than that rarely changes the decision and slows the work.

Can we trust public benchmarks?

Use them to build a shortlist. They do not reflect your data, prompts, or edge cases, so the final call should come from your own eval set.

How often should we revisit the model choice?

At each major provider release or price change, and on a schedule such as quarterly. With an automated eval suite, a re-test costs hours and not weeks.

What if the best model is too slow?

Try lower reasoning effort, a smaller model with a stronger prompt, streaming for perceived speed, or caching. Measure the quality impact of each against the eval set.

Want this built for your team? Book a free 30-minute call.

Want to apply this to your business?

Book a free 30-minute call. We will tell you what we would do first.