Enforcing Structured Outputs from Language Models

Practical frameworks for validating and enforcing structured LLM outputs — schemas, retry loops, constrained decoding, and production-grade validation pipelines.

#llm#validation#structured-output#ai
Cover image for the article: Enforcing Structured Outputs from Language Models

Language models output text. Production systems need structured data. The gap between "usually returns valid JSON" and "always returns valid JSON" is where most AI features break in production. Here's how to close that gap completely.

The structured output problem

Ask an LLM for JSON and you'll get valid JSON 92-97% of the time, depending on the model and prompt complexity. That means 3-8% of production requests return malformed data that crashes your downstream pipeline.

Three approaches exist, each with different trade-offs:

  1. Schema-constrained generation — force the model to only produce valid tokens
  2. Post-generation validation with retry — validate output, retry on failure
  3. Hybrid approach — constrain structure, validate semantics

Approach 1: Schema-constrained generation

Modern APIs support native structured output modes that guarantee schema compliance at the token level.

LLM Output Validation Pipeline

from pydantic import BaseModel, Field, validator
from typing import Literal
import anthropic

class ExtractedEntity(BaseModel):
    name: str = Field(description="Entity name as it appears in text")
    entity_type: Literal["person", "organization", "location", "product"]
    confidence: float = Field(ge=0.0, le=1.0, description="Extraction confidence")
    context: str = Field(description="Surrounding sentence for verification")

    @validator("name")
    def name_not_empty(cls, v):
        if not v.strip():
            raise ValueError("Entity name cannot be empty")
        return v.strip()

class ExtractionResult(BaseModel):
    entities: list[ExtractedEntity] = Field(default_factory=list)
    summary: str = Field(max_length=500)
    processing_notes: str | None = None

async def extract_entities(text: str) -> ExtractionResult:
    client = anthropic.AsyncAnthropic()

    response = await client.messages.create(
        model="claude-sonnet-4-20250514",
        max_tokens=2048,
        messages=[{"role": "user", "content": f"Extract entities from: {text}"}],
        tools=[{
            "name": "submit_extraction",
            "description": "Submit extracted entities",
            "input_schema": ExtractionResult.model_json_schema()
        }],
        tool_choice={"type": "tool", "name": "submit_extraction"}
    )

    # Tool use guarantees schema compliance
    tool_input = response.content[0].input
    return ExtractionResult.model_validate(tool_input)

This approach gives you structural guarantees — the output will always match your schema. But it doesn't guarantee semantic correctness. A valid JSON object with hallucinated entities is still wrong.

Approach 2: Validation with intelligent retry

When schema-constrained generation isn't available or sufficient, validate and retry with the error fed back to the model.

import { z } from "zod";
import Anthropic from "@anthropic-ai/sdk";

const SentimentSchema = z.object({
  sentiment: z.enum(["positive", "negative", "neutral", "mixed"]),
  score: z.number().min(-1).max(1),
  reasoning: z.string().min(10).max(500),
  key_phrases: z.array(z.string()).min(1).max(10),
  language: z.string().length(2),
});

type SentimentResult = z.infer<typeof SentimentSchema>;

async function analyzeSentiment(
  text: string,
  maxRetries: number = 3,
): Promise<SentimentResult> {
  const client = new Anthropic();
  let lastError: string | null = null;

  for (let attempt = 0; attempt < maxRetries; attempt++) {
    const systemPrompt = `Analyze sentiment. Return valid JSON matching this schema:
${JSON.stringify(SentimentSchema.shape, null, 2)}
${lastError ? `\nPrevious attempt failed validation: ${lastError}\nFix the issue.` : ""}`;

    const response = await client.messages.create({
      model: "claude-sonnet-4-20250514",
      max_tokens: 1024,
      system: systemPrompt,
      messages: [{ role: "user", content: text }],
    });

    const content = response.content[0].type === "text" ? response.content[0].text : "";
    
    // Extract JSON from response
    const jsonMatch = content.match(/\{[\s\S]*\}/);
    if (!jsonMatch) {
      lastError = "No JSON object found in response";
      continue;
    }

    try {
      const parsed = JSON.parse(jsonMatch[0]);
      const validated = SentimentSchema.parse(parsed);
      return validated;
    } catch (e) {
      lastError = e instanceof z.ZodError
        ? e.errors.map((err) => `${err.path.join(".")}: ${err.message}`).join("; ")
        : `JSON parse error: ${e.message}`;
    }
  }

  throw new Error(`Validation failed after ${maxRetries} attempts: ${lastError}`);
}

The key insight: feed validation errors back into the prompt. Models are excellent at self-correcting when told exactly what went wrong. First-retry success rate is typically 95%+.

Approach 3: Multi-layer validation

For high-stakes outputs, validate at multiple levels:

from dataclasses import dataclass
from enum import Enum
from typing import Any, Callable

class ValidationLevel(Enum):
    STRUCTURAL = "structural"   # Valid JSON, matches schema
    SEMANTIC = "semantic"       # Values make sense in context
    FACTUAL = "factual"        # Claims can be verified

@dataclass
class ValidationResult:
    level: ValidationLevel
    passed: bool
    errors: list[str]
    confidence: float

class MultiLayerValidator:
    def __init__(self):
        self.validators: list[tuple[ValidationLevel, Callable]] = []

    def add_validator(self, level: ValidationLevel, fn: Callable):
        self.validators.append((level, fn))

    async def validate(self, output: Any, context: dict) -> list[ValidationResult]:
        results = []
        for level, validator_fn in self.validators:
            result = await validator_fn(output, context)
            results.append(ValidationResult(
                level=level,
                passed=result["passed"],
                errors=result.get("errors", []),
                confidence=result.get("confidence", 1.0)
            ))
            # Stop at first failing level — no point checking semantics if structure is wrong
            if not result["passed"]:
                break
        return results

# Usage
validator = MultiLayerValidator()

# Level 1: Schema validation
validator.add_validator(ValidationLevel.STRUCTURAL, validate_schema)

# Level 2: Semantic checks (e.g., date ranges make sense, amounts are reasonable)
validator.add_validator(ValidationLevel.SEMANTIC, validate_semantics)

# Level 3: Factual verification (e.g., company names exist, URLs resolve)
validator.add_validator(ValidationLevel.FACTUAL, verify_facts)

Production patterns

Pattern: Graceful degradation

When validation fails after all retries, don't crash. Degrade gracefully:

Failure ModeFallback Strategy
Schema invalid after 3 retriesReturn partial result with confidence=0
Semantic validation failsFlag for human review, serve cached response
Factual check failsServe with disclaimer, log for correction
Timeout on validationReturn unvalidated with warning flag

Pattern: Validation caching

If you validate the same output shape repeatedly, cache validation results by schema hash + output hash. Repeated identical outputs (common with cached/batched responses) skip re-validation.

Pattern: Output type negotiation

Different consumers need different output formats. Let the API negotiate:

from enum import Enum

class OutputFormat(Enum):
    JSON = "json"
    MARKDOWN = "markdown"
    PLAIN_TEXT = "plain_text"

async def generate_with_format(
    prompt: str,
    output_format: OutputFormat,
    schema: dict | None = None,
) -> dict:
    """Generate with format-specific validation."""
    match output_format:
        case OutputFormat.JSON:
            # Use tool_use for guaranteed schema compliance
            return await generate_with_tool_use(prompt, schema)
        case OutputFormat.MARKDOWN:
            # Validate markdown structure (headers, code blocks)
            return await generate_and_validate_markdown(prompt)
        case OutputFormat.PLAIN_TEXT:
            # Minimal validation — length, language, no PII
            return await generate_and_validate_text(prompt)

Metrics to track

MetricHealthyInvestigate
First-attempt validation pass rate>95%<90%
Retry success rate>99%<95%
Average retries needed<1.1>1.5
Validation latency overhead<50ms>200ms
Hard failures (all retries exhausted)<0.1%>1%

Key takeaways

  • Use schema-constrained generation (tool_use, function calling) as your first choice — it eliminates structural failures entirely
  • When retry is needed, feed validation errors back into the prompt — first-retry success rate exceeds 95%
  • Validate at multiple levels: structure, semantics, and facts
  • Design graceful degradation for every validation failure mode
  • Cache validation results for repeated outputs
  • Monitor validation pass rates — degradation in pass rate signals prompt or model changes

The goal isn't 100% first-attempt success. The goal is 100% eventual correctness with bounded latency. Build the validation pipeline once, and every LLM integration in your system benefits.

Comments

    No comments yet. Be the first to share your thoughts.