AI Engineering · 6 MIN

How do you get reliable JSON from an LLM?

Use schema-constrained structured outputs, validate anyway, and handle refusals and truncation. A reliability checklist for JSON from LLMs in production.

By NactorePublished 4 Jul 2026All articles

You get reliable JSON from an LLM by using the provider's schema-constrained structured output feature instead of asking nicely in the prompt, then validating the result in your own code and handling the cases where the output cannot match the schema. Constrained generation makes invalid JSON and missing required fields effectively a solved problem. It does not make the values correct, and it has documented edge cases such as refusals and truncation. This post covers the options, the gaps, and a pattern that holds up in production.

Key takeaways
  • There are three tiers of reliability: prompting for JSON, JSON mode (valid JSON only), and schema-constrained structured outputs.
  • Both OpenAI and Anthropic document structured outputs that constrain generation to a supplied JSON Schema.
  • Valid structure does not mean correct content. A well-formed object can still hold a wrong total.
  • Refusals and token-limit truncation can produce output that does not match the schema, so check stop reasons.
  • Schema design affects accuracy: clear field names, enums for closed sets, and an explicit way to say "unknown."
  • Nactore validates structure in code and measures content accuracy with evals, so downstream systems never receive unchecked model output.

What are the ways to get JSON from a model?

ApproachWhat it guaranteesWhere it breaks
Prompt only ("respond in JSON")NothingExtra prose, code fences, missing fields, wrong types
JSON modeOutput parses as JSONFields and types can still differ from what you expect
Structured outputs (schema-constrained)Output conforms to your JSON SchemaUnsupported schema features, refusals, truncation, content errors
Strict tool useTool arguments conform to the tool's schemaSame limits, plus the model may choose the wrong tool

OpenAI's structured outputs guide draws exactly this line: structured outputs ensure the response follows your schema, while JSON mode only ensures valid JSON. Anthropic's structured outputs documentation describes constrained sampling against a compiled grammar so the response is valid and type-safe.

Use the highest tier your model supports. Prompt-only JSON is acceptable for a prototype and a liability in production.

When does valid JSON still fail?

Even with constrained decoding, three cases need code.

  1. Refusals. If a model declines a request for safety reasons, the response carries a refusal instead of schema-conforming content. OpenAI exposes a refusal field. Anthropic returns a refusal stop reason and notes the output may not match your schema.
  2. Truncation. If generation hits the token limit or a filter, the output is cut off. OpenAI marks the response as incomplete, and Anthropic documents a max_tokens stop reason where output may not match the schema. Check the stop reason before parsing.
  3. Schema limits. Providers support a subset of JSON Schema. Anthropic's docs list unsupported features such as recursive schemas and numeric or string length constraints, and describe complexity limits. First use of a new schema can add latency while it is processed.

Anthropic's docs also note a subtle one: enum values may not keep your exact capitalization, so compare enums case-insensitively. Read the current limits for your provider and model before you design the schema, since they change.

How should you design the schema?

The schema is part of your prompt. The model sees field names and descriptions, so write them like instructions.

  • Name fields clearly. invoice_total_incl_tax beats total.
  • Describe each field. One sentence on format and meaning, including units and currency.
  • Use enums for closed sets. Categories, statuses, and priority levels should be enums, not free text.
  • Allow "unknown" explicitly. Make a field nullable or add an unknown enum value. Forcing the model to fill a required field it cannot answer pushes it to guess. The same logic appears in Anthropic's guidance on hallucinations, which recommends giving the model permission to say it does not know.
  • Keep it flat where possible. Deep nesting and many optional fields raise complexity and error rates.
  • Put reasoning before the answer. If you want the model to think, add a short reasoning string ahead of the decision field, because fields are generated in order.
  • Version the schema. Log the schema version with each trace. See LLM observability: what to log.

What does a safe handling pattern look like?

  1. Call with the schema and a max_tokens high enough for the largest realistic output.
  2. Check the stop reason. If it is a refusal or truncation, branch before parsing.
  3. Parse and validate with your own validator (Pydantic, Zod, or similar), even though the provider promised conformity. It protects against future provider changes.
  4. Apply business rules. Totals add up, dates are plausible, IDs exist, enums are known.
  5. Retry once on validation failure, feeding the specific error back, then fall back to a safe default or a human queue.
  6. Log the outcome so failures become eval cases.
Pro tip

Do not treat a schema-valid response as a trusted one. If the object triggers an action such as a refund or an email, the business-rule checks in step 4 are what stand between a model error and a customer.

How do you measure content accuracy, not just structure?

Structure is the easy 5 percent. The hard part is whether the extracted or classified values are right. Build a labeled set, compare field by field, and report accuracy per field, as described in how to evaluate LLM output quality. Track the schema-valid rate separately, since a drop there points at truncation, refusals, or a provider change. For a concrete application, see document extraction with LLMs.

Frequently asked questions

Is JSON mode good enough?

It guarantees parseable JSON and nothing more. Fields can be missing or the wrong type. Use schema-constrained structured outputs when your provider and model support them.

Do we still need a validator if the API enforces the schema?

Yes. It guards against refusals, truncation, provider changes, and fallback to a different model, and it is where business rules live.

Why does the model fill fields with plausible but wrong values?

A required field with no allowed way to say "unknown" pushes the model to guess. Make uncertain fields nullable and measure accuracy with a labeled set.

Can we use structured outputs with streaming?

Check your provider's current documentation. When you stream, you generally cannot validate until the object is complete, so design the UI to render partial results safely.

Want this built for your team? Book a free 30-minute call.

Want to apply this to your business?

Book a free 30-minute call. We will tell you what we would do first.