AI Engineering · 6 MIN
How do you evaluate LLM output quality?
Choose the grader that fits the output: exact match, code checks, LLM judges, or human review. A practical method for measuring LLM output quality.
You evaluate LLM output quality by matching the grading method to the kind of output: exact match or code checks for anything with a verifiable answer, an LLM grader with a written rubric for qualities like tone or faithfulness, and sampled human review to keep the other two honest. No single metric covers a whole feature. This post lays out the grading methods, when each fits, and how to calibrate an LLM judge so its scores mean something.
- Pick the cheapest grader that can actually judge the property in question. Code beats LLM judges where the answer is checkable.
- LLM judges are useful and biased. Published research documents position, verbosity, and self-preference bias.
- A rubric with concrete score levels beats a vague "rate this 1 to 10" instruction.
- Grade the parts of a pipeline separately. A bad answer can come from bad retrieval, a bad prompt, or a bad model.
- Calibrate every automated grader against a hand-labeled sample before you trust its numbers.
- Nactore builds these graders into every AI feature it ships, so quality is a number on a dashboard and not an opinion.
What are the main ways to grade LLM output?
There are four, and a real feature usually uses three of them.
| Method | Best for | Weakness |
|---|---|---|
| Exact or fuzzy match | Classification, extraction of IDs, dates, amounts | Fails on valid paraphrases |
| Code checks | Schema validity, required fields, banned strings, arithmetic, link validity | Cannot judge meaning |
| LLM grader with rubric | Tone, faithfulness to sources, completeness, helpfulness | Biased and needs calibration |
| Human review | Calibration, ambiguous cases, high-stakes outputs | Slow and expensive |
Anthropic's guide to building evaluations lists these same families (exact match, similarity measures, and LLM-based Likert, binary, and ordinal scales) and recommends automating where possible and favoring volume over a few hand-graded cases. Start there, then add human review only where the others cannot reach.
How do you grade extraction and classification outputs?
This is the easy case, and teams still overcomplicate it. If the task is to pull an invoice total, a due date, or a category label, you have a ground truth. Compare the output to the label with code.
Normalize before comparing. Strip whitespace, standardize date formats, and round currency. Then report precision and recall per field, not one blended score. A model that gets totals right and due dates wrong needs a different fix than one that is uniformly mediocre. For the extraction use case specifically, see document extraction with LLMs.
How do you grade open-ended answers?
Open-ended output has no single correct string, so you grade properties instead. Write each property as a yes or no question, or as a short scale with described levels.
- Faithfulness. "Does every claim in the answer appear in the provided documents? List any that do not."
- Completeness. "Does the answer address all parts of the question?"
- Policy compliance. "Does the answer avoid promising refunds, legal advice, or delivery dates?"
- Tone. "Is the reply polite, direct, and free of filler?"
Binary questions are more stable than 1 to 10 scales. When you need a scale, define each level. "3 means the answer is correct but omits one required step" is gradable. "3 means average" is not.
Ask the grader to give its reasoning first and its verdict last, and keep the verdict in a strict format so code can read it. Anthropic's hallucination guidance describes a related pattern for the generator itself: have the model find a supporting quote for each claim and retract any claim it cannot support. The same pattern works as a grader.
Can you trust an LLM judge?
Trust it after you test it. The MT-Bench and Chatbot Arena paper reports that strong LLM judges can agree with human preferences at roughly the level humans agree with each other, and it also names position bias (favoring the first or last answer), verbosity bias (favoring longer answers), and self-enhancement bias (favoring answers from similar models).
Practical countermeasures:
- Swap the order in pairwise comparisons and keep only results that survive the swap.
- Penalize length explicitly in the rubric, or cap the allowed length.
- Use a different model as the grader than the one that produced the answer.
- Calibrate. Hand-label 50 outputs, run the grader, and look at disagreements. Edit the rubric until agreement is acceptable, then freeze it.
- Re-check periodically. If you change the grader model, repeat the calibration.
Version your graders like code. When a score moves, you need to know whether the product changed or the judge did.
How do you evaluate a pipeline instead of a single prompt?
Most production features are pipelines: retrieve, build a prompt, call a model, validate, act. Grade each stage on its own so you can locate a failure.
For retrieval-augmented features, measure whether the right passage was retrieved at all (a retrieval hit rate against labeled questions) before you measure whether the answer was good. Anthropic's contextual retrieval write-up measures its improvements as the share of questions where the correct chunk did not appear in the top 20 results, which is a good example of grading retrieval separately from generation. For agents, grade the tool calls (right tool, right arguments) separately from the final message.
What does a lightweight scorecard look like?
A one-page scorecard per feature is enough to start:
- Task accuracy on the held-out set, per field or per category.
- Faithfulness pass rate from the calibrated grader.
- Schema-valid output rate.
- Refusal and fallback rate.
- Median and 95th percentile latency, and cost per successful task.
Track these on every change, and add a failing production case to the set each time one appears. The wider case for doing this before launch is in why AI features need evals before production.
Frequently asked questions
Should we use the same model to generate and to grade?
Avoid it where you can. Judges tend to favor outputs that resemble their own, and a shared blind spot means errors pass unnoticed. Use a different model for grading, and calibrate it against human labels.
Are similarity metrics like ROUGE or embedding similarity good enough?
They are useful for consistency checks and rough summarization quality. They miss factual errors that keep the wording similar, so pair them with a faithfulness check.
How many human labels do we need to calibrate a judge?
A sample of around 50 outputs, spread across easy and hard cases, is a workable start. Add more when the grader disagrees with humans in a pattern you cannot explain.
What if two good graders disagree?
Treat it as a signal that the rubric is ambiguous. Read the disagreements, tighten the criterion in writing, and re-run.
Want this built for your team? Book a free 30-minute call.
Want to apply this to your business?
Book a free 30-minute call. We will tell you what we would do first.