AI Automation · 6 MIN

How do you extract data from documents with LLMs reliably?

LLMs can turn invoices, contracts and forms into structured data. Learn the pipeline, validation checks and evals that make extraction reliable.

By NactorePublished 27 Aug 2026All articles

You extract data from documents reliably by treating the model as one step in a pipeline: parse the file, ask for a schema-bound answer, validate every field with code, and send anything doubtful to a person. The model reads messy layouts well. Your code proves the result is usable. This post walks through each stage and the evals that tell you when it is ready.

Key takeaways
  • Modern models read PDFs directly, including tables and scanned layouts, so many projects no longer need a separate OCR stage.
  • A JSON schema guarantees the shape of the output, not the truth of the values. Validation code supplies the truth check.
  • Measure extraction per field against a labeled set of real documents, not by eyeballing a few samples.
  • Route low-confidence or failed-validation documents to a review queue rather than forcing a guess.
  • Nactore builds extraction pipelines with field-level evals scoped to each team.

What is the pipeline from file to structured data?

Most reliable systems share the same five stages.

  1. Ingest. Accept PDFs, images and emails, and normalize them into a consistent form.
  2. Read. Send the document to a model that can handle text and page images together.
  3. Extract. Ask for output that matches a strict JSON schema.
  4. Validate. Check types, ranges, totals and cross-field rules in code.
  5. Route. Pass clean records to your system and send exceptions to a review queue.

Anthropic documents that Claude can work with PDFs including text, charts and tables, and lists document extraction into structured formats as a core use case. The documentation also publishes request size and page limits, so check them against your longest documents before you design the pipeline.

Do we still need OCR?

Often not. Earlier systems split the job into OCR first, then rules or a text model. That chain loses layout, and a table flattened into text loses its meaning. Models that see the page image alongside the text keep the structure. For clean digital PDFs the extracted text is usually enough. For scans, photos and skewed pages, sending page images is typically more reliable.

A dedicated OCR or layout service can still earn its place for very high volume, strict latency or cost limits, or on-premises requirements. Decide with a measurement, not a preference. Run both approaches on the same labeled set and compare field accuracy and cost.

How do we get output we can trust?

Use a schema, then distrust it. OpenAI's structured output documentation explains that the feature makes responses adhere to a supplied JSON Schema, and that refusals or truncated responses can still occur. This removes parsing failures. It does not tell you whether the invoice total was read correctly.

So add checks that only code can do:

CheckExampleWhat it catches
Type and formatDate parses, currency has two decimalsMalformed values
ArithmeticLine items sum to the subtotal, tax applied correctlyMisread digits
Cross-fieldDue date is after issue dateSwapped fields
Reference dataVendor exists in your master listHallucinated names
EvidenceEach value includes the quoted source text or pageInvented values

Asking the model to return the source snippet for each field is one of the cheapest reliability gains available. If the quoted text does not appear in the document, reject the field.

Pro tip

Make "not found" a legal answer. If the schema forces a value for every field, the model will invent one. Allow null, and treat null as a signal, not a failure.

How do we measure extraction quality?

Build a labeled set of real documents, ideally a few hundred across your vendors, layouts and languages. Label the correct value for every field you plan to extract. Then score each field separately.

  • Exact match rate per field. Invoice number and total need near-perfect accuracy, while a free-text description can tolerate fuzziness.
  • Null accuracy. Count how often the model correctly says a field is absent.
  • End-to-end pass rate. The share of documents where every required field is correct and validation passes.
  • Review rate. The share routed to humans, and how many of those humans actually had to correct.

Weight fields by business risk. A wrong payment amount is serious, and a wrong secondary reference is minor. Our general method for this is in how to evaluate LLM output quality, and the schema side is covered in structured output and JSON reliability.

What should the human review step look like?

Review should be fast, or people will approve without reading. Show the original page beside the extracted fields, highlight values that failed validation or had low confidence, and let the reviewer correct with one click. Store every correction.

Those corrections feed back into your test set and your prompt examples. Over a few weeks the review rate usually falls as you fix the document types that cause the most errors. See human-in-the-loop automation for how to design the queue.

What goes wrong in production?

  • New layouts. A vendor changes their template and accuracy drops. Monitor per-vendor accuracy so you notice.
  • Long documents. Fields buried on page 40 get missed. Split by section or extract in passes.
  • Sensitive data. Check your provider's data retention terms and your own regulatory duties before sending documents. This matters especially for finance and health records.
  • Prompt injection in documents. A document can contain text that tries to instruct the model. Treat all document content as untrusted and never let extracted text trigger actions directly.
  • Cost surprises. Page images cost more than text. Estimate cost on your real document mix during the pilot.

Frequently asked questions

Which model should we use?

Test two or three on your own labeled set. Documents differ enough that public benchmarks rarely predict your results. Pick the cheapest model that clears your per-field accuracy bar.

Can this handle handwriting or poor scans?

Sometimes, with lower accuracy. Measure it on real samples and route low-confidence cases to review. Do not assume performance from a clean demo file.

How long does a first version take?

A focused pilot on one document type fits in four weeks, including the labeled set, the pipeline, the validation rules and the review queue.

Do we have to replace our existing OCR vendor?

Not necessarily. You can compare approaches on the same test set and keep whichever clears the bar at the lowest cost.

Want this built for your team? Book a free 30-minute call.

Want to apply this to your business?

Book a free 30-minute call. We will tell you what we would do first.