LLM Structured Outputs: Reliable JSON From AI in Production

LLM Structured Outputs: Reliable JSON From AI in Production — cover image

The first version of most AI features works beautifully in a demo. The model reads an email and returns a tidy summary. Then it ships, and one morning the integration breaks because the model wrapped its answer in friendly commentary, used a different field name, or returned a date in a format nobody expected. The business logic downstream was written for data, and the model sent prose.

Structured outputs are the discipline of making an LLM return data your software can rely on. This article covers why free-form responses fail, how schemas and validation fit together, and a step-by-step approach for building extraction, classification and routing features that behave predictably in production.

Why free-form text is a bad contract

When a program asks a language model for an answer, it implicitly expects a contract: certain fields, certain types, certain allowed values. Plain text offers none of that. A prompt that says "reply in JSON" works most of the time, which is the most dangerous failure rate in software, because it passes testing and fails in front of customers.

Typical failure modes include trailing explanations after the JSON, missing required fields, numbers returned as strings, invented enum values ("urgent-ish" instead of "high"), and truncated output when a response hits a length limit. Each one produces an exception or, worse, silent bad data stored in your database.

The three layers of reliability

Dependable structured output comes from stacking three layers, each catching a different class of error.

Layer 1: Constrain the format

Modern model APIs let you pass a JSON schema and have the provider enforce it during generation, or let you define tools with typed parameters. This removes most syntax problems. Define required fields, types, enums for closed sets of values, and length limits. Write clear field descriptions, because the model reads them as instructions.

Layer 2: Validate the data

Even enforced schemas do not guarantee sensible values. Parse the response with a validation library in your own language and add rules that reflect business reality: a refund cannot exceed the order total, a date cannot be in the distant past, a phone number must match your supported regions, and a referenced ID must exist in your database.

Layer 3: Handle failure gracefully

Plan what happens when validation fails. Retry with the error message, fall back to another approach, or send the item to a person. A pipeline that has no failure path will eventually stall or corrupt data. Observability matters here too, and our guide to AI observability and tracing for LLM apps explains how to see which prompts and inputs cause the failures.

A real-world example scenario

Imagine a logistics SME that receives shipment requests by email in dozens of formats. They want a system that reads each message and creates a booking record with pickup address, delivery address, weight, preferred date and service level.

A first attempt asks the model to "extract the details as JSON." In testing it works. In production, one customer writes weight as "about 2 tonnes," another writes dates as "next Tuesday," and a third attaches a PDF with no body text. The system either crashes or stores strange values.

A stronger design defines a schema where weight is a number in kilograms with a separate field for the original text, the date is an ISO string plus a confidence field, and service level is an enum with an explicit unknown value. Validation checks that the weight is plausible and the date is in the future. Anything that fails twice, or comes back with unknown in a critical field, lands in a review queue where a coordinator confirms it in seconds. The model does the heavy reading, while the system protects the data store. This pattern also underlies the workflows in our look at AI document search and RAG for enterprises, where extraction quality decides everything downstream.

Step-by-step: building a dependable extraction pipeline

  1. Define the output you actually need. Start from the downstream consumer. List the fields your database or API requires and the allowed values for each. Resist the urge to ask for everything the model could produce.
  2. Write the schema with descriptions. Mark required fields, use enums for closed sets, and describe each field in one clear sentence, including format and units. Add an explicit "unknown" or null option so the model is not forced to guess.
  3. Enable schema enforcement. Use your provider's structured output or tool calling feature rather than relying on prompt wording alone.
  4. Parse and validate in code. Use a typed validation library, then apply business rules and cross-checks against your own data.
  5. Return errors to the model. On failure, send the original input, the invalid output and the validation message back, and ask for a corrected response. Cap retries at one or two.
  6. Add a fallback path. Options include a default value, a different model, a rules-based extractor or a human review queue. Decide per field how risky a wrong value is.
  7. Capture evidence. Ask for a short quote or source reference for important fields, so reviewers can verify quickly and so you can detect hallucinated values.
  8. Build a test set. Collect real examples, including ugly ones, and run them whenever the prompt, schema or model changes. Our guide to LLM evals and regression suites shows how to set this up.
  9. Monitor in production. Track validation failure rate, retry rate and review queue size. A rising failure rate is an early warning that inputs or the model changed.
  10. Version everything. Treat schemas like API contracts. Add fields in a backward-compatible way and record which schema version produced each record.

Schema design tips that save pain later

Function calling versus schema-constrained output

These two features overlap, and teams often mix them up. Tool or function calling lets the model decide which action to take and supply typed arguments, which suits agents that choose between searching, creating a ticket or sending a message. Schema-constrained output simply forces the response into a shape and suits extraction, classification and report generation. When you build multi-step agents, structured arguments become the glue between steps, a topic covered in our article on AI agent orchestration patterns.

Security and safety considerations

Structured output does not remove prompt injection risk. If the model reads untrusted text such as an email or web page, an attacker can embed instructions that influence field values. Treat model output as untrusted input. Validate it, escape it before inserting into queries or HTML, and never let a model-generated value directly decide a privileged action without a check. Allow-lists for actions and values are your friend.

Key benefits of structured outputs

Treat the model as a talented but unreliable colleague: let it do the reading, and let your code decide what gets accepted.

Common mistakes to avoid

A practical checklist before you ship

Before releasing any structured output feature, walk through a short list. Does every field have a clear description and, where relevant, a unit or format? Is there an explicit way for the model to say it does not know? Are business rules enforced in code rather than trusted to the prompt? Is there a capped retry and a fallback path? Do you have at least a few dozen real examples in a test set, including messy ones? Can you see failure rates on a dashboard? If any answer is no, fix that first. These checks take a day or two of work and prevent weeks of cleaning up bad records later.

It also helps to involve the people who will use the output. A support lead or finance analyst can quickly tell you which fields matter and which mistakes are tolerable, and that knowledge shapes your schema more than any technical guideline.

Conclusion

Reliable AI features are mostly engineering around the model, not clever prompts. Define a clear schema, enforce it at generation time, validate against real business rules, retry once with feedback, and send hard cases to people. That layered approach turns a promising demo into a component you can run in front of customers every day.

If you are planning an AI feature that has to plug into real systems, our AI development team can help design the schemas, validation and fallbacks that make it dependable.

Frequently Asked Questions

What are structured outputs in LLM applications?
Structured outputs are model responses that follow a defined format, usually a JSON schema, instead of free-form text. They let your code read fields like category, amount or priority directly, without fragile text parsing.
Is a schema enough to guarantee correct data?
No. A schema guarantees shape, not truth. A model can return perfectly valid JSON with a wrong amount or an invented customer name. You still need business-rule validation, source grounding where possible, and review steps for high-risk fields.
Should I use function calling or JSON mode?
Use whichever your provider supports with schema enforcement. Function or tool calling is natural when the model chooses between actions. Schema-constrained JSON output is a better fit for extraction and classification. Both need validation on your side.
How many times should I retry when the output fails validation?
One or two retries is typical. Feed the validation error back to the model so it can correct itself, and cap retries to control cost and latency. If it still fails, route to a fallback such as a simpler model, a default value or human review.
Do structured outputs make prompts shorter or longer?
Often shorter, because the schema carries the formatting instructions. Good field descriptions inside the schema also act as lightweight prompting, so you can cut repeated formatting rules from the prompt text.