Making AI Agents Reliable Enough for Production

Battle-tested patterns for building AI agents that handle failures gracefully — retry strategies, state machines, human-in-the-loop fallbacks, and observability.

#ai-agents#reliability#error-handling#production
Cover image for the article: Making AI Agents Reliable Enough for Production

AI agents that work 95% of the time in development will fail spectacularly in production. The gap between a demo agent and a production agent is entirely about reliability engineering. Here are the patterns that close that gap.

Why agents fail differently

Traditional software fails predictably — a null pointer, a timeout, a missing field. Agents fail in novel ways:

  • Reasoning loops: The agent gets stuck retrying the same failed approach
  • Tool misuse: Correct tool, wrong parameters, plausible-looking output
  • Context drift: Accumulated errors in long conversations that compound
  • Confidence without correctness: The agent produces wrong answers with no signal it's uncertain

These failure modes demand different solutions than traditional error handling.

Pattern 1: State machine with guardrails

Every production agent needs an explicit state machine. Without one, agents wander through unbounded execution paths.

Agent State Machine Architecture

from enum import Enum
from dataclasses import dataclass
from typing import Any, Callable

class AgentState(Enum):
    PLANNING = "planning"
    EXECUTING = "executing"
    VALIDATING = "validating"
    RECOVERING = "recovering"
    ESCALATING = "escalating"
    COMPLETE = "complete"
    FAILED = "failed"

@dataclass
class AgentContext:
    state: AgentState = AgentState.PLANNING
    attempts: int = 0
    max_attempts: int = 3
    history: list[dict] = None
    errors: list[str] = None

    def __post_init__(self):
        self.history = self.history or []
        self.errors = self.errors or []

class ReliableAgent:
    def __init__(self, tools: list[Callable], max_steps: int = 20):
        self.tools = tools
        self.max_steps = max_steps
        self.step_count = 0

    async def run(self, task: str) -> dict[str, Any]:
        ctx = AgentContext()

        while ctx.state not in (AgentState.COMPLETE, AgentState.FAILED):
            self.step_count += 1
            if self.step_count > self.max_steps:
                ctx.state = AgentState.ESCALATING
                break

            match ctx.state:
                case AgentState.PLANNING:
                    plan = await self._plan(task, ctx)
                    if plan.is_valid:
                        ctx.state = AgentState.EXECUTING
                    else:
                        ctx.state = AgentState.RECOVERING

                case AgentState.EXECUTING:
                    result = await self._execute(plan, ctx)
                    ctx.state = AgentState.VALIDATING

                case AgentState.VALIDATING:
                    if await self._validate(result, task):
                        ctx.state = AgentState.COMPLETE
                    else:
                        ctx.attempts += 1
                        if ctx.attempts >= ctx.max_attempts:
                            ctx.state = AgentState.ESCALATING
                        else:
                            ctx.state = AgentState.RECOVERING

                case AgentState.RECOVERING:
                    recovery = await self._recover(ctx)
                    if recovery.should_retry:
                        ctx.state = AgentState.PLANNING
                    else:
                        ctx.state = AgentState.ESCALATING

                case AgentState.ESCALATING:
                    await self._escalate_to_human(task, ctx)
                    ctx.state = AgentState.FAILED

        return {"state": ctx.state.value, "result": result, "steps": self.step_count}

The key insight: never let an agent execute indefinitely. Bound execution with step limits, attempt counters, and mandatory validation checkpoints.

Pattern 2: Tool call validation

Agents hallucinate tool parameters. Validate every tool call before execution.

import { z } from "zod";

interface ToolDefinition {
  name: string;
  schema: z.ZodType;
  execute: (params: unknown) => Promise<unknown>;
  dangerLevel: "safe" | "moderate" | "dangerous";
}

class ValidatedToolExecutor {
  private tools: Map<string, ToolDefinition>;
  private executionLog: Array<{
    tool: string;
    params: unknown;
    result: unknown;
    timestamp: number;
  }> = [];

  async execute(
    toolName: string,
    params: unknown,
  ): Promise<{ success: boolean; result?: unknown; error?: string }> {
    const tool = this.tools.get(toolName);
    if (!tool) {
      return { success: false, error: `Unknown tool: ${toolName}` };
    }

    // Validate parameters against schema
    const validation = tool.schema.safeParse(params);
    if (!validation.success) {
      return {
        success: false,
        error: `Invalid params: ${validation.error.message}`,
      };
    }

    // Check for dangerous operations
    if (tool.dangerLevel === "dangerous") {
      const approved = await this.requestHumanApproval(toolName, params);
      if (!approved) {
        return { success: false, error: "Human rejected dangerous operation" };
      }
    }

    // Execute with timeout
    try {
      const result = await Promise.race([
        tool.execute(validation.data),
        this.timeout(30_000),
      ]);

      this.executionLog.push({
        tool: toolName,
        params: validation.data,
        result,
        timestamp: Date.now(),
      });

      return { success: true, result };
    } catch (e) {
      return { success: false, error: e.message };
    }
  }

  private timeout(ms: number): Promise<never> {
    return new Promise((_, reject) =>
      setTimeout(() => reject(new Error("Tool execution timeout")), ms),
    );
  }
}

Pattern 3: Progressive retry with backoff

Not all failures are equal. Transient errors (rate limits, timeouts) need retries. Logic errors (wrong approach) need replanning.

from enum import Enum
import asyncio
import random

class FailureType(Enum):
    TRANSIENT = "transient"    # Retry same approach
    LOGIC = "logic"            # Replan entirely
    PERMANENT = "permanent"    # Escalate immediately

def classify_failure(error: Exception, context: dict) -> FailureType:
    """Classify failure to determine recovery strategy."""
    if isinstance(error, (TimeoutError, ConnectionError)):
        return FailureType.TRANSIENT
    if "rate_limit" in str(error).lower():
        return FailureType.TRANSIENT
    if isinstance(error, ValidationError):
        return FailureType.LOGIC
    if context.get("attempts", 0) >= 2:
        # Same error twice = not transient
        return FailureType.PERMANENT
    return FailureType.LOGIC

async def retry_with_strategy(
    func,
    failure_type: FailureType,
    max_retries: int = 3
) -> Any:
    for attempt in range(max_retries):
        try:
            return await func()
        except Exception as e:
            ft = classify_failure(e, {"attempts": attempt})

            if ft == FailureType.PERMANENT:
                raise

            if ft == FailureType.TRANSIENT:
                # Exponential backoff with jitter
                delay = (2 ** attempt) + random.uniform(0, 1)
                await asyncio.sleep(delay)

            elif ft == FailureType.LOGIC:
                # Don't retry the same thing - signal replanning needed
                raise ReplanRequired(f"Logic error after {attempt + 1} attempts: {e}")

Pattern 4: Conversation checkpointing

Long-running agents accumulate context that can drift. Checkpoint state at critical decision points so you can rollback.

The implementation stores snapshots of agent memory, tool state, and conversation history at each major transition. When validation fails, you rollback to the last valid checkpoint instead of starting over or continuing with corrupted state.

Production metrics that matter

Track these for every agent in production:

MetricTargetAlert Threshold
Task completion rate>95%<90%
Avg steps to completion<8>15
Escalation rate<5%>10%
Tool call validation failures<10%>20%
Mean time to completion<30s>60s
Loop detection triggers<2%>5%

The human-in-the-loop escape valve

Every production agent needs a path to human intervention. Not as a failure state, but as a designed capability.

Design escalation triggers:

  • Confidence below threshold on high-stakes decisions
  • Multiple failed recovery attempts
  • Novel situations not covered by the agent's training
  • Actions that cross a danger threshold (deleting data, spending money, external communications)

The best agents escalate early and gracefully rather than grinding through uncertain territory.

Key takeaways

  • Bound all agent execution with step limits and attempt counters — infinite loops are the #1 production incident
  • Validate every tool call against a schema before execution
  • Classify failures by type and match recovery strategy accordingly
  • Checkpoint state so you can rollback, not just retry
  • Design escalation as a feature, not a failure mode
  • Monitor agent behavior with the same rigor you'd apply to any distributed system

Reliability isn't a feature you add later. It's the architecture you start with. Build the state machine and validation layer first, then make the agent smarter within those constraints.

Comments

    No comments yet. Be the first to share your thoughts.