Structured Output (JSON) from LLMs
Structured output is a response from a large language model that follows a predefined machine-readable format, most commonly a JSON object that matches a schema. It allows application code to parse, validate and store model output directly, instead of interpreting unstructured prose.
- JSON: JavaScript Object Notation, a text format of keys and values that nearly every programming language can parse.
- Schema: A specification of the required fields, their data types and their allowed values, often written as JSON Schema.
- Prompt-based formatting: The prompt describes the fields and requests JSON only, a basic prompt engineering technique that works with any model but offers no guarantee.
- Constrained decoding: Some model APIs restrict generation so that the output always matches a supplied schema.
- Validation: Application code checks the LLM JSON output against the schema before using it, whichever method produced it.
For example, a triage service sends a plain-text bug report to a model and receives a JSON object with a title, a severity, a component and reproduction steps for the issue tracker.
Key Characteristics of Structured Output
- Deterministic shape: Field names and data types remain constant across requests, even though the extracted values differ each time, and a low temperature further reduces variation.
- Machine-readable: The output can be passed directly to a database, an internal API or a user interface without manual editing.
- Schema-driven: The schema functions as a contract between the prompt, the model and the consuming application code.
- Enumerations: Fields such as severity can be restricted to a fixed list, for example low, medium, high and critical.
- Fallibility: Prompt-based JSON can contain missing fields, extra fields, wrong types or invalid values.
- Foundation for tools: Tool calling depends on the same mechanism, since the model emits function arguments as structured JSON.
How Structured Output Works
- Define the schema: List each field, its data type, whether it is mandatory and any permitted enumeration values.
- Write the extraction prompt: Describe the fields, supply the input text and request a single JSON object without commentary.
- Generate the response: The model produces JSON, either guided by the prompt alone or restricted by constrained decoding.
- Parse the JSON: Locate the object within the response text and convert it into native data structures with a standard JSON parser.
- Validate the data: Verify the required fields, the data types, the permitted values and the absence of unexpected keys.
- Retry or accept: Return the validation errors to the model for one corrected attempt, or store the valid object.
A typical extraction prompt for this task looks like the following.
Extract the bug report below into one JSON object with these fields:
title (string), severity (one of: low, medium, high, critical),
component (string), steps (list of strings).
Return only the JSON object, with no other text.
Bug report: Clicking Pay with an empty cart shows a 500 error page.Example: Validating JSON Output in Python
The program below parses a hard-coded sample model response and validates it against a simple schema using only the standard library.
import json
# Expected fields for a bug report, with their Python types
SCHEMA = {"title": str, "severity": str, "component": str, "steps": list}
SEVERITIES = {"low", "medium", "high", "critical"}
# Sample model response, hard-coded for illustration (no model is called)
sample_response = """Here is the extracted bug report:
{"title": "Checkout returns 500 for an empty cart",
"severity": "urgent", "component": "backend",
"steps": ["Empty the cart", "Click Pay"], "assignee": "Priya"}"""
def extract_json(text):
start, end = text.find("{"), text.rfind("}")
return json.loads(text[start:end + 1])
def validate(data):
errors = []
for field, expected in SCHEMA.items():
if field not in data:
errors.append(f"missing field: {field}")
elif not isinstance(data[field], expected):
errors.append(f"{field} must be {expected.__name__}")
if data.get("severity") not in SEVERITIES:
errors.append(f"severity '{data.get('severity')}' not in {sorted(SEVERITIES)}")
for field in sorted(set(data) - set(SCHEMA)):
errors.append(f"unexpected field: {field}")
return errors
data = extract_json(sample_response)
errors = validate(data)
print("Parsed fields:", list(data))
print("Valid:", not errors)
for e in errors:
print(" -", e)
if errors:
print("Retry prompt: Fix these problems and return only JSON:", "; ".join(errors))Parsed fields: ['title', 'severity', 'component', 'steps', 'assignee']
Valid: False
- severity 'urgent' not in ['critical', 'high', 'low', 'medium']
- unexpected field: assignee
Retry prompt: Fix these problems and return only JSON: severity 'urgent' not in ['critical', 'high', 'low', 'medium']; unexpected field: assignee- Parsing is not validation: The response is valid JSON, yet it still fails two checks, an invalid severity and an invented assignee field.
- Tolerance: Locating the first and last braces is a simple way to extract JSON from LLM responses that open with an introductory sentence despite instructions.
- Actionable retry: Specific error messages give the model precise corrections, which is more effective than repeating the original prompt.
Applications of Structured Output
- Issue triage: Converting unstructured bug reports and support emails into consistent issue tracker fields.
- Log analysis: Extracting the service name, the error code and the timestamp from unstructured application log messages.
- Code review automation: Returning review comments as objects containing a file path, a line number and a severity classification.
- Classification pipelines: Producing a category label and a confidence level for every document in a large batch.
- Agent orchestration: Passing typed arguments between successive steps in AI agents and multi-step workflows.
Advantages
- Reliable integration: Downstream services consume predictable, typed fields instead of parsing descriptive prose.
- Automatic quality checks: Validation identifies malformed or incomplete responses before they reach production systems or users.
- Efficiency: A compact JSON object usually requires fewer output tokens than an equivalent descriptive paragraph.
- Testability: Expected objects can be compared field by field in automated evaluation suites.
Limitations
- Incorrect values: Valid JSON can still contain wrong or invented values, a form of LLM hallucination.
- Truncation: Prompt-based JSON can be truncated or wrapped in extra text, especially when the output limit is reached.
- Trade-off: Forcing an immediate JSON answer can reduce accuracy on tasks that benefit from chain of thought prompting.
- Complexity: Deeply nested or highly detailed schemas increase both the error rate and the prompt length.
- Portability: Constrained decoding features and the supported schema keywords vary considerably between model providers.
Quick Quiz
Pick an answer to check yourself. Nothing is saved.
Question 1 / 3
1. What does a schema define in structured output?
Frequently Asked Questions
How do I get JSON output from an LLM?
Describe the required fields in the prompt, ask for a single JSON object with no other text, and validate the response in code. Where the model API supports schema-constrained output, supply the schema as well.
What is the difference between JSON mode and LLM structured output?
JSON mode generally guarantees syntactically valid JSON but not a particular set of fields. Schema-based structured output also enforces the field names, types and allowed values defined in a schema.
Why does an LLM return invalid JSON?
Without constraints the model predicts text token by token, so it can add commentary, omit a closing brace or stop at the output limit. Validation and a retry step handle these cases.
Is validation still needed when the API enforces a schema?
Yes. A schema guarantees the shape of the data, not the correctness of the values. Business rules, such as whether a component name exists, still need checks in application code.
Related Articles
- What is Prompt EngineeringLearn what prompt engineering is, how to design a prompt step by step, its key characteristics, uses and limits, with a Python code review prompt example.
- Tool Calling (Function Calling) in LLMLearn how tool calling works in LLMs: JSON Schema tool definitions, model tool calls, argument validation and tool results, with a Python DevOps example.
- Chain of Thought PromptingLearn chain of thought prompting: how step-by-step reasoning improves LLM accuracy and how to parse the final answer, with a Python log triage example.
- What is Context EngineeringLearn what context engineering is: selecting, ranking and fitting information into an LLM context window, with a Python token budget example and limits.
- LLM Hallucination: Causes and FixesLearn what LLM hallucination is, why models invent facts and code APIs, and how to reduce it with RAG and validation, with a runnable Python checker.
- System Prompt vs User PromptSystem prompt vs user prompt: who writes each, what belongs in each, how models prioritise them, with a comparison table and a Python code review example.