Enforcing Structured Outputs from Language Models
Practical frameworks for validating and enforcing structured LLM outputs — schemas, retry loops, constrained decoding, and production-grade validation pipelines.

Language models output text. Production systems need structured data. The gap between "usually returns valid JSON" and "always returns valid JSON" is where most AI features break in production. Here's how to close that gap completely.
The structured output problem
Ask an LLM for JSON and you'll get valid JSON 92-97% of the time, depending on the model and prompt complexity. That means 3-8% of production requests return malformed data that crashes your downstream pipeline.
Three approaches exist, each with different trade-offs:
- Schema-constrained generation — force the model to only produce valid tokens
- Post-generation validation with retry — validate output, retry on failure
- Hybrid approach — constrain structure, validate semantics
Approach 1: Schema-constrained generation
Modern APIs support native structured output modes that guarantee schema compliance at the token level.
from pydantic import BaseModel, Field, validator
from typing import Literal
import anthropic
class ExtractedEntity(BaseModel):
name: str = Field(description="Entity name as it appears in text")
entity_type: Literal["person", "organization", "location", "product"]
confidence: float = Field(ge=0.0, le=1.0, description="Extraction confidence")
context: str = Field(description="Surrounding sentence for verification")
@validator("name")
def name_not_empty(cls, v):
if not v.strip():
raise ValueError("Entity name cannot be empty")
return v.strip()
class ExtractionResult(BaseModel):
entities: list[ExtractedEntity] = Field(default_factory=list)
summary: str = Field(max_length=500)
processing_notes: str | None = None
async def extract_entities(text: str) -> ExtractionResult:
client = anthropic.AsyncAnthropic()
response = await client.messages.create(
model="claude-sonnet-4-20250514",
max_tokens=2048,
messages=[{"role": "user", "content": f"Extract entities from: {text}"}],
tools=[{
"name": "submit_extraction",
"description": "Submit extracted entities",
"input_schema": ExtractionResult.model_json_schema()
}],
tool_choice={"type": "tool", "name": "submit_extraction"}
)
# Tool use guarantees schema compliance
tool_input = response.content[0].input
return ExtractionResult.model_validate(tool_input)
This approach gives you structural guarantees — the output will always match your schema. But it doesn't guarantee semantic correctness. A valid JSON object with hallucinated entities is still wrong.
Approach 2: Validation with intelligent retry
When schema-constrained generation isn't available or sufficient, validate and retry with the error fed back to the model.
import { z } from "zod";
import Anthropic from "@anthropic-ai/sdk";
const SentimentSchema = z.object({
sentiment: z.enum(["positive", "negative", "neutral", "mixed"]),
score: z.number().min(-1).max(1),
reasoning: z.string().min(10).max(500),
key_phrases: z.array(z.string()).min(1).max(10),
language: z.string().length(2),
});
type SentimentResult = z.infer<typeof SentimentSchema>;
async function analyzeSentiment(
text: string,
maxRetries: number = 3,
): Promise<SentimentResult> {
const client = new Anthropic();
let lastError: string | null = null;
for (let attempt = 0; attempt < maxRetries; attempt++) {
const systemPrompt = `Analyze sentiment. Return valid JSON matching this schema:
${JSON.stringify(SentimentSchema.shape, null, 2)}
${lastError ? `\nPrevious attempt failed validation: ${lastError}\nFix the issue.` : ""}`;
const response = await client.messages.create({
model: "claude-sonnet-4-20250514",
max_tokens: 1024,
system: systemPrompt,
messages: [{ role: "user", content: text }],
});
const content = response.content[0].type === "text" ? response.content[0].text : "";
// Extract JSON from response
const jsonMatch = content.match(/\{[\s\S]*\}/);
if (!jsonMatch) {
lastError = "No JSON object found in response";
continue;
}
try {
const parsed = JSON.parse(jsonMatch[0]);
const validated = SentimentSchema.parse(parsed);
return validated;
} catch (e) {
lastError = e instanceof z.ZodError
? e.errors.map((err) => `${err.path.join(".")}: ${err.message}`).join("; ")
: `JSON parse error: ${e.message}`;
}
}
throw new Error(`Validation failed after ${maxRetries} attempts: ${lastError}`);
}
The key insight: feed validation errors back into the prompt. Models are excellent at self-correcting when told exactly what went wrong. First-retry success rate is typically 95%+.
Approach 3: Multi-layer validation
For high-stakes outputs, validate at multiple levels:
from dataclasses import dataclass
from enum import Enum
from typing import Any, Callable
class ValidationLevel(Enum):
STRUCTURAL = "structural" # Valid JSON, matches schema
SEMANTIC = "semantic" # Values make sense in context
FACTUAL = "factual" # Claims can be verified
@dataclass
class ValidationResult:
level: ValidationLevel
passed: bool
errors: list[str]
confidence: float
class MultiLayerValidator:
def __init__(self):
self.validators: list[tuple[ValidationLevel, Callable]] = []
def add_validator(self, level: ValidationLevel, fn: Callable):
self.validators.append((level, fn))
async def validate(self, output: Any, context: dict) -> list[ValidationResult]:
results = []
for level, validator_fn in self.validators:
result = await validator_fn(output, context)
results.append(ValidationResult(
level=level,
passed=result["passed"],
errors=result.get("errors", []),
confidence=result.get("confidence", 1.0)
))
# Stop at first failing level — no point checking semantics if structure is wrong
if not result["passed"]:
break
return results
# Usage
validator = MultiLayerValidator()
# Level 1: Schema validation
validator.add_validator(ValidationLevel.STRUCTURAL, validate_schema)
# Level 2: Semantic checks (e.g., date ranges make sense, amounts are reasonable)
validator.add_validator(ValidationLevel.SEMANTIC, validate_semantics)
# Level 3: Factual verification (e.g., company names exist, URLs resolve)
validator.add_validator(ValidationLevel.FACTUAL, verify_facts)
Production patterns
Pattern: Graceful degradation
When validation fails after all retries, don't crash. Degrade gracefully:
| Failure Mode | Fallback Strategy |
|---|---|
| Schema invalid after 3 retries | Return partial result with confidence=0 |
| Semantic validation fails | Flag for human review, serve cached response |
| Factual check fails | Serve with disclaimer, log for correction |
| Timeout on validation | Return unvalidated with warning flag |
Pattern: Validation caching
If you validate the same output shape repeatedly, cache validation results by schema hash + output hash. Repeated identical outputs (common with cached/batched responses) skip re-validation.
Pattern: Output type negotiation
Different consumers need different output formats. Let the API negotiate:
from enum import Enum
class OutputFormat(Enum):
JSON = "json"
MARKDOWN = "markdown"
PLAIN_TEXT = "plain_text"
async def generate_with_format(
prompt: str,
output_format: OutputFormat,
schema: dict | None = None,
) -> dict:
"""Generate with format-specific validation."""
match output_format:
case OutputFormat.JSON:
# Use tool_use for guaranteed schema compliance
return await generate_with_tool_use(prompt, schema)
case OutputFormat.MARKDOWN:
# Validate markdown structure (headers, code blocks)
return await generate_and_validate_markdown(prompt)
case OutputFormat.PLAIN_TEXT:
# Minimal validation — length, language, no PII
return await generate_and_validate_text(prompt)
Metrics to track
| Metric | Healthy | Investigate |
|---|---|---|
| First-attempt validation pass rate | >95% | <90% |
| Retry success rate | >99% | <95% |
| Average retries needed | <1.1 | >1.5 |
| Validation latency overhead | <50ms | >200ms |
| Hard failures (all retries exhausted) | <0.1% | >1% |
Key takeaways
- Use schema-constrained generation (tool_use, function calling) as your first choice — it eliminates structural failures entirely
- When retry is needed, feed validation errors back into the prompt — first-retry success rate exceeds 95%
- Validate at multiple levels: structure, semantics, and facts
- Design graceful degradation for every validation failure mode
- Cache validation results for repeated outputs
- Monitor validation pass rates — degradation in pass rate signals prompt or model changes
The goal isn't 100% first-attempt success. The goal is 100% eventual correctness with bounded latency. Build the validation pipeline once, and every LLM integration in your system benefits.
Recommended reading

The State of Agentic AI in 2026: Capabilities, Limitations, and Production Readiness
Comprehensive analysis of agentic AI in 2026 covering production capabilities, current limitations, and enterprise readiness benchmarks with real deployment data.

Observability for AI Agents: Tracing Multi-Step Reasoning Chains in Production
How to implement production observability for AI agents including distributed tracing, reasoning chain analysis, and debugging multi-step failures.

Measuring and Reducing AI Workload Carbon Emissions: A Practical Engineering Guide
Building a carbon-aware scheduling system for ML training and inference workloads that reduced our AI infrastructure emissions by 42% while maintaining SLA commitments.

Comments
No comments yet. Be the first to share your thoughts.