Every AI pipeline that ships to production eventually learns the same lesson the hard way. The demo works, the JSON looks perfect, and then one Tuesday the model starts wrapping its output in a markdown fence, or invents a priority level called "kinda urgent," or quietly stops returning the due_date field that your billing system depends on. Nothing crashed. Nothing announced itself. The data just went wrong, and nobody noticed until a customer did.
If you are building anything where an LLM feeds structured data into code, an API, a database, or an agent, you have two engineering problems, not one. The first is making the output valid right now: JSON that parses, fields that exist, enums that stay inside the list you defined. The second is keeping it valid next month: when the provider quietly updates a model, when a teammate rewrites your prompt "just a little," and when you swap in a cheaper model that behaves nothing like the old one. This drift problem is a cousin of the one behind why AI hallucinates: the model gives you a confident, well-formed answer that is quietly wrong.
The good news as of September 2026 is that both problems have real, boring, well-understood solutions. Constrained decoding handles the first. Regression testing handles the second. This is reliability engineering, and the tools have finally caught up to the discipline.
Where unvalidated LLM output actually breaks
Before picking tools, it helps to know your enemy. Production failures from LLM output fall into a few distinct patterns, and each one has a different fix.
Malformed JSON. The classic. The model adds a friendly preamble ("Sure! Here is your JSON:"), wraps the object in triple backticks, truncates mid-object because it ran out of tokens, or emits a stray brace. If you asked for JSON in the prompt and hoped for the best, this is the failure you got. Google's own docs are blunt about it: setting the response MIME type to application/json without a schema is "only a strong hint" that "carries a small risk of generating malformed JSON."
Valid JSON, wrong shape. Sneakier. The output parses beautifully, and then you discover the model omitted three required fields, returned "priority": "highish" instead of an enum value, or made a number into a string. OpenAI's older JSON mode guaranteed syntax but explicitly not schema conformance, which is why the company published a strict mode at all.
Schema-valid, semantically wrong. The nastiest one, because no decoder can fix it. Every field is present, every type is correct, and the answer is still wrong: the model classified a furious billing complaint as "positive sentiment" or hallucinated an evidence quote that isn't in the source document. Constrained decoding guarantees shape, not substance. Your tests have to cover substance separately.
Silent drift. The slow killer. Your pipeline is fine for months, then a provider ships a model update, or someone edits a system prompt, and extraction quality drops 8%. There is no diff to review, no error log, no crash. Just quietly worse data flowing downstream. This is the failure mode that turns prompt regression testing from a nice-to-have into table stakes.
Each fix layer below targets a specific one of these. Use them stacked.

Layer one: constrained decoding, so invalid JSON becomes impossible
The strongest guarantee in this whole space is not a prompt trick or a post-hoc parser. It is changing how the model samples tokens. Constrained decoding compiles your JSON Schema into a grammar, and at every generation step the model can only pick tokens that keep the output valid. Malformed JSON stops being unlikely and becomes mechanically impossible.
OpenAI Structured Outputs is the reference implementation. Launched in August 2024 and now the default way to get JSON out of the Responses API, it lets you pass a JSON Schema with strict: true and get back a response that matches it. OpenAI's own published evaluations put strict mode at 100% schema adherence, versus roughly 86% for function calling and lower still for plain prompting (the platform docs have the full schema subset and examples). Treat the 100% as vendor-reported, but the mechanism is real and the strict subset it enforces tells you a lot: every object must set additionalProperties: false, every property must be listed in required (optional fields are done as unions with null, not omitted keys), and oneOf is not supported. There are hard limits too, around 100 total properties and five levels of nesting. Schemas that feel fine in a notebook can bounce at the API gateway, so validate them in CI.
Two production details that bite people. First, strict mode adds a separate refusal field: if the model declines a request on safety grounds, you get a populated refusal and no parsed object. Code that never checks this ships silent NULL bugs into downstream tables. Second, the first request against a brand-new schema pays a compilation penalty while the grammar is built, then it is cached. Keep schemas stable across requests instead of regenerating them per call.
Anthropic shipped the same capability for Claude, and by 2026 it has moved out of beta: the old output_format parameter migrated to output_config.format, the beta header is gone, and the SDKs offer Pydantic and Zod helpers (Anthropic's structured outputs docs walk through the current API shape). Anthropic splits the feature in two: JSON outputs for the response body, and strict: true on tool definitions so agent tool calls get grammar-validated arguments. If you are still forcing a tool and reading its arguments to fake structured output, you can stop.
Google Gemini takes the same idea and expresses it through responseSchema in the generation config. One gotcha worth knowing: Gemini's schema is a subset of the OpenAPI 3.0 schema object, not JSON Schema, and it has a field nobody else has, propertyOrdering, which controls the order in which properties get generated. Gemini's docs also recommend pairing the schema with response_mime_type: application/json for guaranteed validity, and the SDKs will parse and validate the response against your Pydantic or Zod type in one step.
For local and self-hosted models, the tool that built this category is Outlines, maintained by the .txt (dottxt) team and still actively released (version 1.3.3 shipped in August 2026). Outlines compiles your JSON Schema, regex, or context-free grammar into a finite-state machine and masks invalid tokens during generation. It is baked into the serving stacks a lot of teams already run: vLLM, SGLang, TGI, and llama.cpp all carry some form of it. The same pattern exists across the open ecosystem, and dottxt also offers a hosted, OpenAI-compatible API endpoint that enforces your full JSON Schema if you do not want to run your own GPUs.
A quick archaeology note, because older blog posts will send you down dead ends. Jsonformer, the clever 2023 library that filled in fixed JSON tokens and only generated the content, was influential but is effectively dormant today, with a tiny commit history and no real activity in years. Guidance, from Microsoft Research, is still around for token-level control and template-style programs, but it remains pre-1.0 and is more of a decoding-control layer than a JSON reliability tool. For new work, native provider features or Outlines are the mainstream paths.
Layer two: validation libraries, for everything the schema cannot say
Constrained decoding guarantees structure. It cannot guarantee that start_date is before end_date, that email is a real email, that confidence is between 0 and 1 in spirit rather than letter, or that the summary actually reflects the document. That is what validation libraries are for.
Instructor is the de facto standard here. It is a thin patch over provider SDKs: you define a Pydantic model, pass it as response_model, and get back a typed, validated instance instead of raw text. When validation fails, Instructor retries automatically, feeding the validation error back to the model so it can correct itself, three attempts by default. The current entry point is instructor.from_provider("openai/gpt-4o") (or anthropic, gemini, ollama, and a dozen others), which makes the same call work across providers. It streams partial objects, handles nested models, and is very actively maintained.
Here is the honest framing for 2026: native structured outputs absorbed the baseline JSON-validity problem, so Instructor's job shifted. It is now the cross-provider layer and the business-rules layer. Use it when you need one code path across OpenAI, Anthropic, and Gemini; when you have cross-field invariants no JSON Schema can express; or when you want Pydantic validators as your last line of defense. The Instructor team themselves draw the boundary clearly: Instructor for extraction, and their README now points agent use cases to PydanticAI, the Pydantic team's agent runtime.
If you do not want a library, you can get 80% of this layer yourself: validate every response against your schema with Ajv (JavaScript) or Pydantic (Python) at the boundary, and on failure, retry with the error message appended. That retry-with-feedback loop is genuinely effective, because models are good at fixing errors when you show them exactly what went wrong. But it is a retry tax, in latency and cost, which is exactly what layer one exists to eliminate.
The architecture most teams land on: strict mode at the provider for structure, Pydantic or Zod validation at the boundary for defense in depth and business rules, and no retry loop at all for syntax, because the decoder already made bad JSON impossible.
Layer three: prompt regression testing, so it stays working
Layers one and two fix today. The third layer answers the question that actually keeps production teams up at night: did my change break something? That is what prompt regression testing is, and it works like software testing, because it is software testing.
The premise is the one from this article's opening: run the same suite of prompts against your change (a new prompt, a new model, a new provider, a temperature tweak) and compare against a baseline, tracking the failures that matter: invalid JSON, field omissions, semantic regressions.
The tool that owns this space is Promptfoo, and it just had the most eventful year of any tool in this article. In March 2026, OpenAI announced it was acquiring the company, with the plan to fold Promptfoo's evaluation and red-teaming technology into Frontier, OpenAI's enterprise agent platform (both OpenAI and Promptfoo published the announcement). Two facts matter for anyone choosing it right now. First, Promptfoo has committed that the open-source project stays open source and keeps supporting every provider, not just OpenAI's. Second, the scale is real: the company reports more than 350,000 developers have used it, 130,000 monthly active, and teams at over a quarter of the Fortune 500.
What makes Promptfoo fit this reliability story is its assertion system, which covers all three failure classes from earlier:
- Structural: the
is-jsonassertion checks that the entire output is valid JSON, and optionally validates it against a full JSON Schema (loaded from a file, so your schema lives in one place and your eval references it). There is a softercontains-jsonfor when prose around the JSON is acceptable, but if your API consumes the whole response, useis-json: it fails on markdown fences exactly the way your strict parser will. - Field-level: deterministic
javascriptorpythonassertions run real logic against the parsed output, so you can assert thatevidenceis empty when the input contains no evidence, or thatqueueis a value your routing system actually accepts. - Semantic: model-graded assertions like
llm-rubriccompare outputs against a rubric, which is how you catch the regression where every field is present and the answer is still wrong. - Operational: latency and cost assertions keep your cheap replacement model honest.
And the part that makes it a reliability tool rather than a nice dashboard: the same prompt suite runs across providers, so "GPT vs Claude vs Gemini vs local model" becomes a matrix you regenerate whenever you want, with results in the web viewer or as JUnit XML in CI. When a provider ships a model update, you re-run the suite and read the diff. (The assertion docs cover every check type, including the JSON schema ones.) The same cross-model thinking applies when you are picking the models in the first place, and if you work across vendors you already know that prompting ChatGPT, Claude, and Gemini identically gives identically mediocre results.
Promptfoo is not the only option, just the center of gravity. DeepEval and Braintrust cover similar ground with different ergonomics, and a plain pytest harness with Pydantic validation is a legitimate v1 if your needs are narrow. But for the specific workflow in this article, one suite across many models, Promptfoo's provider matrix is the thing that made it popular.
A starter setup you can build in an afternoon
Here is a concrete setup for a support ticket classifier, the shape that generalizes to most extraction pipelines.
First, write the contract once. Create schemas/ticket.json, a JSON Schema with exactly four top-level fields: queue (enum: billing, identity, technical, general), priority (enum: low, medium, high), summary (bounded string), and evidence (array, capped). This single file is now your API contract: the provider schema, the validation schema, and the test schema all reference it. One definition, no drift.
Second, enforce it at the provider. Turn on strict structured outputs (OpenAI strict: true, Anthropic output_config.format, Gemini responseSchema), or wrap the call in Instructor with the Pydantic twin of your schema. Keep a validator at the boundary anyway; with constrained decoding it should never fire, and the day it does you will want to know.
Third, write the eval config. A promptfoo YAML with your prompt file, three providers (say an OpenAI model, a Claude model, and a Gemini model), and a dozen test cases: normal tickets, a blank ticket, an ambiguous ticket, a furious ticket, a ticket with no evidence. Attach three assertions per test: is-json with your schema file (structure), a JavaScript check that evidence.length matches expectations (field behavior), and an llm-rubric for the tricky cases (semantics).
Fourth, wire it into CI. Run the suite on every pull request that touches the prompt, the schema, or the model list, and export the results. Green means your change is safe across all three models. Red with a specific assertion failure tells you exactly which of the three failure classes you just introduced.
That is the whole stack: one schema file as the contract, constrained decoding to make structure automatic, validation for the rules structure cannot express, and a regression suite that re-proves the whole thing whenever anything changes. None of it is exotic. All of it is the same discipline you would apply to any API that mattered, and the same discipline that separates the agent deployments that hold up in production from the real risks of AI agents nobody talks about until something breaks.

Two questions come up constantly when teams set this up:
If strict mode guarantees the schema, do I still need tests? Yes, and the reason is the third failure class. Strict mode guarantees that priority is one of low, medium, or high. It cannot guarantee the model picked the right one, or that your summary matches the ticket. Structure is solved; meaning is not. Semantic assertions are the only layer that covers meaning.
Is Promptfoo still safe to adopt now that OpenAI owns it? The open-source commitment is public and explicit, and the acquisition is recent enough that the practical answer is unchanged: the CLI, the config format, and the multi-provider support are all still there. If vendor ownership is a hard constraint for your org, pin your version, keep your eval configs in your own repo (they are just YAML), and remember that the entire setup above degrades gracefully: your schema, your assertions, and your test cases are portable even if the runner changes.
Pricing and plan details for the tools above are as published by the vendors around September 2026 and can change; confirm on the official sites before committing.




