Buyer Guides · 6 MIN
Red flags when hiring AI developers or an AI engineering team
Ten warning signs when hiring AI developers or a partner, from demo-only portfolios to no evals, plus the questions that expose them. A buyer checklist.
The biggest red flags when hiring AI developers are a portfolio of demos with no production systems, no measurement of output quality, vague answers about failure, and a pitch built around a model name instead of your problem. Any one is worth a follow-up question. Several together should end the conversation. This checklist gives you the signs and the questions that surface them.
- Demos are cheap. Ask for evidence of systems running in production and how they are monitored.
- Anyone serious about AI can describe how they measure output quality.
- Honest engineers can name what a system should not be used for and how it fails.
- Be cautious of promises about certainty, speed, or results before seeing your data.
- Vague answers on data handling and ownership are a risk, not a detail.
- Nactore expects buyers to ask these questions and answers them with evals, not adjectives.
Why is hiring in AI harder to judge than in other engineering?
Language models make it easy to produce something impressive in an afternoon. That compresses the visible gap between a prototype builder and an engineer who can run a system reliably for years. Resumes and demos look similar. The difference lies in what happens with messy data, rare failures, cost, and monitoring, none of which show up in a polished walkthrough.
So the test is not "can they build it" but "can they prove it works and keep it working." The signs below are about that difference.
What are the ten red flags?
| Red flag | Why it matters | Question that exposes it |
|---|---|---|
| Portfolio is only demos | No evidence of production reliability | "Which of these run in production today, and how do you know they work?" |
| No measurement of quality | Quality is a guess | "How do you score output, and on what data?" |
| Cannot describe failures | Untested or not candid | "Show me a case where it failed and what you changed." |
| Leads with a model name | Tool before problem | "Why this approach for our workflow?" |
| Promises certainty | AI output is probabilistic | "What error rate should we plan for?" |
| Guarantees results before seeing data | Cannot know yet | "What would you need to see to estimate this?" |
| Vague on data and IP | Exposure risk | "Where does our data go, and who owns the outputs?" |
| No plan for monitoring | System decays silently | "What do you log, and who is alerted?" |
| No cost awareness | Surprise bills at scale | "What does each processed item cost?" |
| Bench swap after signing | Different people do the work | "Who exactly will work on this?" |
How can you tell if someone really ships to production?
Ask for specifics that are hard to fake. A person who has run production systems can talk about the unglamorous parts. They mention retries, timeouts, rate limits, monitoring, alerting, cost caps, and rollbacks without prompting.
Ask them to walk you through the last time a system misbehaved after launch. Listen for how they found out, how they diagnosed it, and what they changed. Detail and humility are good signs. Blame on the model alone, or no story at all, are not.
Do they measure output quality?
This is the single most useful filter. A team that ships AI without a way to score output is guessing, and so are you. Ask to see an example eval set, the metric, and how scores changed as the system improved. Strong teams treat evals as a core part of the work, covered in AI evals before production.
A simple screen: ask "how will we know when it is good enough to ship?" If the answer is a feeling, a demo, or "when you are happy," keep looking.
What about data, security, and ownership?
Red flags here are evasive or casual answers. A capable team can explain which providers touch your data, what is retained, how access is controlled, and who owns the code, prompts, and eval sets. If they cannot, that is a delivery risk as well as a legal one. Use the checklist in IP and data terms for AI projects, and have your counsel review the actual terms.
What do good signs look like?
Contrast is useful. These are signs that a person or team is worth continuing with.
- They ask about your data first. They want to see real inputs before proposing anything.
- They define success in numbers. A threshold and a test set, agreed before building.
- They name limits. They tell you what the system should not do and where humans stay in the loop.
- They propose a small, bounded first step. A short pilot with a decision at the end.
- They are comfortable saying no. They will tell you if the use case is poorly suited to AI.
For how to frame that first step, see how to scope an AI pilot.
How should you run the vetting conversation?
- Start with your problem, not their pitch. Describe a real workflow and watch how they respond.
- Ask for a production example. Probe for failure stories and monitoring.
- Request the measurement approach. Ask how they would test on your data.
- Check who does the work. Meet the engineers who will actually be on the project.
- Run a bounded trial. A fixed pilot reveals more than any interview.
Frequently asked questions
Is a strong portfolio enough to trust a team?
A portfolio shows what was built, not how reliably it runs. Ask about production performance, monitoring, and failure handling, and speak to the engineers who would work on your project.
Should we be worried if a team will not promise accuracy numbers up front?
No. Honest teams cannot promise results before seeing your data. They should, however, propose how to measure and a threshold to agree on.
How do we vet an individual developer versus a team?
Use the same questions. For individuals, also ask about backup, continuity, and who covers when they are unavailable.
What is the fastest way to reduce hiring risk?
A short, fixed-scope pilot with a written pass bar. It shows how the team works with your data and gives you evidence before a larger commitment.
Hire for proof, not polish
Look for evidence of production systems, measured quality, candor about failure, and clear answers on data. Those four separate strong teams from impressive demos.
Want this built for your team? Book a free 30-minute call.
Want to apply this to your business?
Book a free 30-minute call. We will tell you what we would do first.