AI EngineeringGuides & Tutorials

Structured Output from LLMs Explained: JSON Mode, Schemas, and Why Your Agent Breaks Without It

Stop regex-ing prose into data and retrying failed parses. Learn the full ladder of LLM structured output, from prompting and JSON mode to schema-enforced structured outputs and tool calling, with vendor-by-vendor limits, latency gotchas, and when not to force a schema.

Toolbit AI - Team
15 min read
Structured Output from LLMs Explained: JSON Mode, Schemas, and Why Your Agent Breaks Without It

It is 2 a.m. and your phone is buzzing. Your agent, the one that was supposed to quietly summarize support tickets all night, has crashed. The fix for this whole class of failure is LLM structured output: give the model a JSON Schema, and constrained decoding makes malformed output structurally impossible instead of merely discouraged. You open the logs and find the model's reply:

Sure! Here's the JSON you asked for: {"priority": "high", ...

json.loads() threw, the pipeline died, and nobody told the retry loop to stop crying. Or the reply got cut off mid-object. Or the field was customer_name on Monday and name on Tuesday. One bad token, one dead pipeline.

If you have ever regex-ed prose into data, built retry loops, or raised a shrine of try/except blocks around your parser, you know this pain. Anthropic's own docs list the failure modes builders hit even with careful prompting: parsing errors from invalid JSON syntax, missing required fields, inconsistent data types, and schema violations that force error handling and retries.

Here is the good news, and it is genuinely good: there is an escalating ladder of fixes. Prompt harder. Then JSON mode. Then schema-constrained structured outputs. Then tool calling. Every rung up buys you a stronger guarantee, and you get to pick the lowest rung that satisfies your app. Most agents need the schema rung. Let's climb.

In short

  • Prompting for JSON guarantees nothing. In OpenAI's own vendor-reported testing, an older model prompted for a schema scored under 40% adherence on their evals.
  • JSON mode guarantees valid JSON, but not YOUR JSON. No schema adherence.
  • Schema-constrained structured outputs make malformed output structurally impossible via constrained decoding: the model can only emit schema-valid tokens.
  • Tool calling is a second structured channel with the same strict guarantees, used when the model should act, not just answer.
  • Even with a schema: enum casing can drift, first requests can pay grammar-compile latency, and schema-valid is not semantically correct. Validate values.

The night your agent broke

Quick autopsy: each failure shape kills a different line of your code.

  1. Prose-wrapped JSON. The model opens with "Sure! Here's the JSON:" or wraps everything in markdown fences. Your parser wanted a bare object and got a conversation.
  2. Missing or renamed fields. You asked for ticket_id, the model decided id felt friendlier. Your code reads a key that isn't there and returns undefined all the way up the stack.
  3. Type drift. The model emits "3" today and 3 tomorrow. One is a string, one is a number, and your arithmetic is about to meet NaN.

The nastiest part is that the failure is non-deterministic. Most calls parse fine, so it ships, and the odd failure wakes you at 2 a.m. Retries are a tax rather than a cure: you pay latency, you pay tokens, and you still handle the failure again next time. Treating a retry loop as the fix is one of the most common automation mistakes builders make with AI pipelines, and it deserves the same skepticism here.

This is not a fringe complaint. When OpenAI introduced Structured Outputs, they described years of developers "working around the limitations of LLMs in this area via open source tooling, prompting, and retrying requests repeatedly" just to make outputs interoperate with their systems. The whole industry was duct-taping this.

So here is the ladder of real fixes. Each rung converts an error you handle at runtime into a constraint you enforce at generation time:

  1. Level 0: Prompt for JSON. Ask nicely. No guarantee.
  2. Level 1: JSON mode. Valid JSON guaranteed. Your keys, types, and enums: not guaranteed.
  3. Level 2: Schema-constrained structured outputs. Malformed output becomes structurally impossible.
  4. Level 3: Tool calling. The same strict guarantees, applied to the arguments the model uses to act.

Ask nicely: prompting for JSON

The four rungs from prompting to typed tool channels

Prompting for JSON is the only rung with no guarantee: the LLM can still wrap output in prose, rename your fields, or drift types on any given call.

Level 0 is what most people try first, and honestly, it is a reasonable instinct. You write "Respond ONLY with valid JSON matching this schema" in the prompt, paste the schema below it, and wrap the call in a validator, maybe Pydantic in Python or Zod in JavaScript, plus a retry or three.

What does it guarantee? Nothing. The model treats your schema as a suggestion. In OpenAI's vendor-reported evals, a gpt-4-0613-class model with prompting alone scored under 40% on schema adherence, while gpt-4o-2024-08-06 with strict Structured Outputs scored 100% on the same evals. Treat those numbers as the vendor grading its own homework, but the gap matches what every retry loop has been screaming: asking nicely is not enforcement.

To be fair, Level 0 is not useless. If you are on a legacy model with no JSON mode, or writing a one-off script where a retry loop is cheaper than an API migration, prompting plus validation will carry you.

The mistake is stopping here and calling it solved. Everything that broke at 2 a.m. can still break on this rung.

A contract the model cannot break

Constrained decoding lets only schema-valid token paths survive

The contract comes in two strengths: JSON mode promises valid JSON, and schema-constrained structured outputs promise your JSON.

Level 1: JSON mode, the first guarantee

JSON mode guarantees the output is valid, parseable JSON, but not that it matches your schema: keys, types, and enums stay unconstrained.

JSON mode is the first time the API itself makes you a promise. On OpenAI, you enable it with a response format of type: "json_object", and from that moment the output is valid JSON. Syntactically parseable, always. That is the entire promise.

OpenAI's own comparison table is refreshingly blunt about it: JSON mode, valid JSON yes, adheres to schema no. You get an object; you have no idea which object. It might be missing every field you care about, and that's still "valid JSON."

Even this rung has a famous gotcha. OpenAI's docs warn that if the word "JSON" doesn't appear in your prompt context, the model "may generate an unending stream of whitespace" until it hits the token limit, so the API errors out if "JSON" is absent. A mode that can fail by silently emitting infinite whitespace is a mode with a sense of humor.

OpenAI calls Structured Outputs "the evolution of JSON mode" and recommends it over JSON mode whenever possible, reserving JSON mode for older models.

Level 2: schema-constrained structured outputs

Here is the rung most agents need. You supply a JSON Schema, and the vendor guarantees the output conforms to it. Not "usually conforms." Not "conforms if you prompt well." Conforms, structurally. This is schema enforcement at the generation level: the contract is enforced while tokens are sampled, not checked after the fact.

You write a small schema, say:

{
  "type": "object",
  "properties": {
    "priority": { "type": "string", "enum": ["low", "medium", "high"] },
    "ticket_id": { "type": "string" }
  },
  "required": ["priority", "ticket_id"],
  "additionalProperties": false
}

And the API makes it impossible for the model to reply with prose, with Priority instead of priority, with ticket_id: 3 as a number, or with a surprise extra field. The surfaces differ by vendor: OpenAI uses a json_schema response format with strict: true, Anthropic uses output_config.format for JSON outputs plus strict: true on tools, and Gemini takes a schema alongside response_format with mime_type: application/json.

How it actually works: constrained decoding

The mechanism is called constrained decoding, and it is beautiful. As OpenAI explains in their launch post, they convert your JSON Schema into a context-free grammar (CFG): a set of rules that defines what counts as valid in the language of your schema. At every step of generation, the model samples only from tokens that are valid under that grammar. Everything off-schema is simply not on the menu. Malformed output doesn't get punished; it gets prevented.

Why a context-free grammar rather than a simpler finite state machine (FSM)? OpenAI's answer: "FSMs cannot generally express recursive types," so FSM-based approaches "may struggle to match parentheses in deeply nested JSON." A CFG handles recursion, which is also why OpenAI can support recursive schemas via $ref at all.

Anthropic names the same family of technique on their structured outputs page: responses are "guarantee[d] schema-compliant" "through constrained decoding," with schemas "compiled into a grammar that constrains Claude's output," a process they call grammar-constrained sampling. Different vendors, same load-bearing idea.

What still breaks at this rung

A hard guarantee of structure is not a guarantee of content. Three escape hatches remain:

  • Refusals bypass the schema. If the model refuses for safety reasons, the refusal escapes your format entirely. OpenAI exposes a refusal field; Anthropic returns stop_reason: "refusal". Both are detectable in code, which is the point.
  • Truncation. If you hit max_tokens, you get incomplete JSON (stop_reason: "max_tokens"). Retry with a higher cap.
  • Schema-valid but semantically wrong. Gemini's docs put it plainly: "While output is syntactically correct JSON, always validate values in your application," because schema-compliant-but-incorrect outputs happen. A confident priority: "low" on a house fire is valid JSON and a bad answer.

The guarantee is structural, not cognitive. The model physically cannot emit malformed output, but nothing in the grammar forces the content to be true.

Tool calling as a structured channel

Here is the insight that ties the whole ladder together: function calling, the thing you use to let a model act, is structured output through another door.

When you define a tool, you give it a schema for its arguments, input_schema on Anthropic or a function definition with strict: true on OpenAI. When the model decides to call the tool, the arguments it emits get the exact same grammar-constrained guarantees. All three vendors publish the same decision rule for choosing between them. Structured outputs are for shaping the model's answer. Function calling is for the model taking an action mid-conversation: asking you to run a function before it continues. OpenAI says it directly: if you're connecting the model to tools or systems in your app, use function calling; if you want to structure the model's response to the user, use a structured text format. Gemini's docs draw the same line in their comparison table.

Anthropic frames the combo as the agentic power move: strict tool use plus JSON outputs gives you "reliable tool calls AND structured final outputs" in one workflow. The tool mechanics, the tool_result round-trips, server tools, are their own topic. Here you only need the pattern: actions go through the tool channel, answers through the schema channel, both enforced.

Two things still break. Tool definitions add per-request token overhead: in Anthropic's example, Opus 5 goes from 286 tokens with no tools to 406 with tools. And a forced tool call on unrelated input can hallucinate arguments: OpenAI advises instructing the model to return empty params instead.

What each vendor actually ships

The nicest part of this story is how much the vendors converged. All three independently landed on JSON Schema as the constraint language and on the same two surfaces: a response-format schema for answers, and strict tool parameters for actions. Here is what each actually ships today.

OpenAIAnthropicGemini
API surfacejson_schema response format with strict: true; responses.parse() with Pydantic/Zodoutput_config.format with type: "json_schema"; strict: true on toolsresponse_format with mime_type: application/json + schema; Pydantic/Zod via GenAI SDKs
Recursive schemasYes, via $refNo, 400 errorNot documented
Optional fieldsAll fields required; emulate with union of type and nullUp to 24 optional paramsNull via type arrays like ["string", "null"]
Complexity limits5,000 properties, 10 nesting levels, 120k chars, 1,000 enums20 strict tools, 24 optional params, 16 unions, 180s compile timeout"Not all JSON Schema features supported"; large or deeply nested schemas rejected

OpenAI Structured Outputs: strict: true with a json_schema response format, available from gpt-4o onward. Supported in the Responses API, Chat Completions, Assistants, Batch, and Fine-tuning, with streaming included. The strict subset has firm rules: every field required (optional means a union with null), additionalProperties: false everywhere, no top-level anyOf, and hard limits of 5,000 properties, 10 nesting levels, 120,000 characters across property names and enum values, and 1,000 enum values. Full details are in their structured outputs guide.

Anthropic: a dedicated, generally available structured outputs page (docs.anthropic.com now redirects to platform.claude.com) describes two independent features: JSON outputs via output_config.format and strict tool use via strict: true, each usable alone or together. Notable limits: no recursive schemas (400 error), no min/max string lengths, a regex pattern subset only, at most 20 strict tools, 24 optional parameters, and 16 union-type parameters per request. It works with streaming, token counting, and the Batch API at a 50% discount, but not with citations or message prefilling. Note the SDK migration: output_format moved to output_config.format, and Python SDK v1.0+ raises a TypeError on the old parameter.

Gemini: response_format with mime_type: application/json plus a schema, with Pydantic and Zod supported in the Google GenAI SDKs. The subset covers core types, enums, format values like date-time, and anyOf. Their limitations section is candid: "Not all JSON Schema features are supported," and "very large or deeply nested schemas may be rejected." Their mechanism docs are sparser, with no documented grammar-compile cache. Best practices in the Gemini structured output docs are worth reading twice.

The convergence is the evergreen takeaway. Three teams, building separately, arrived at the same answer: JSON Schema in, grammar-constrained tokens out. The reliable-agent stack sits on constrained decoding.

The gotchas that still bite

Structured outputs kill the 2 a.m. parse crash. They do not kill everything. Here is the field guide of what still bites, all vendor-doc confirmed.

  • Schema design friction. OpenAI's strict mode forces every field to be required, so you emulate optional fields with a union like "type": ["string", "null"], and additionalProperties: false is mandatory on objects. Keep your JSON Schema and your Pydantic or Zod types in sync or they drift: use the SDK-native support or generate one from the other in CI.
  • First-request latency. On Anthropic, the first request with a new schema pays grammar-compilation time, then compiled grammars are cached for 24 hours. The cache invalidates when the schema structure or tool set changes, but name- and description-only edits don't invalidate it. Optional parameters and unions carry what Anthropic calls "exponential compilation cost," enforced by a 180-second compilation timeout and 400 "Schema is too complex for compilation" errors. Their tips: fewer optional params, flatter nesting, only critical tools strict, split across requests or sub-agents. Details in the Anthropic structured outputs page.
  • Enum casing drift. Anthropic documents that capitalization is not guaranteed: your schema says "Conversation topic 3", the model emits "Conversation Topic 3", the call completes normally, no error, no special stop reason. Compare enums case-insensitively, and never define two enums that differ only by casing.
  • Recursive schemas split the vendors. OpenAI supports them via $ref (this is exactly why a CFG beats an FSM). Anthropic rejects them with a 400. Port your schema between vendors before you port your code.
  • Cost shape. Enforcement itself is free: the same per-token price. But Anthropic auto-injects an explanatory system prompt, billed like any system prompt; tool definitions add token overhead per request; and changing the format parameter invalidates prompt caching for the thread.
  • When NOT to force a schema. On open-ended or reasoning-heavy tasks, constraining the token space can degrade quality; OpenAI's docs admit outputs "can still contain mistakes," and their own workaround pattern is to let the model think inside a steps array before a constrained final_answer. If you need interleaved citations, strict JSON is incompatible on Anthropic (400 error). If the model should refuse, remember that refusals escape the schema by design. And if the goal is action rather than a formatted answer, use function calling. One more judgment call: over-constraining an interactive flow so it can never ask the user anything is its own failure mode, the over-asking counterpart of an agent that interrogates users into rage-quitting.

From fragile parsing to reliable agents

The whole ladder in one line: every rung converts an error you handle at runtime into a constraint you enforce at generation time.

That matters most for agents. OpenAI calls structured data extraction "one of the core use cases," including multi-step agentic workflows. Anthropic's complexity tips literally recommend splitting across sub-agents. Gemini lists "generate structured inputs for tools or APIs" as a primary use case. Agents are machines for moving structured data between steps, and every step you harden removes a whole class of failure from the machine.

But an agent that reliably emits structured JSON has only solved half its job. That data still needs somewhere to go: tools to call, context to carry. That's where the MCP protocol and context engineering pick up the story, because a schema guarantees shape, and shape still has to be routed and remembered.

Close on the thing worth remembering: the models will keep changing, the pricing will keep changing, but the mechanism outlives them all. Constrained decoding, a grammar compiled from your schema, tokens that cannot go off the rails: that is the load-bearing wall of the reliable-agent stack. Build on it, and 2 a.m. gets a lot quieter.

Structured output FAQs

Why is my first structured-output request to Claude slow?

The first request with a new schema pays grammar-compilation latency: Anthropic compiles your JSON Schema into a grammar before generation starts. The compiled grammar is cached for 24 hours, and the cache invalidates only when the schema structure or tool set changes, not for name or description edits.

Why did my enum come back with different capitalization than the schema?

Anthropic documents that enum and const capitalization is not guaranteed, so "Conversation topic 3" can arrive as "Conversation Topic 3". The call completes normally with no error and no distinct stop reason. Compare enums case-insensitively and never define two enums that differ only by casing.

Are recursive JSON schemas supported?

It depends on the vendor, not on JSON Schema itself. OpenAI supports recursion via $ref, which works precisely because schemas are compiled into a context-free grammar rather than a finite state machine. Anthropic rejects recursive schemas with a 400 error, so check your schema before porting.

When should I not force a structured output schema?

Skip it on open-ended or reasoning-heavy tasks, where constraining the token space can degrade quality (let the model think inside a steps array and constrain only the final answer). Skip it when you need interleaved citations, which are incompatible with strict JSON on Anthropic. And skip it when the goal is action rather than a formatted answer: that is what tool calling is for.

Share this article

Related articles

Continue exploring similar guides and insights