Making AI Agents Reliable Enough for Production
Battle-tested patterns for building AI agents that handle failures gracefully — retry strategies, state machines, human-in-the-loop fallbacks, and observability.

AI agents that work 95% of the time in development will fail spectacularly in production. The gap between a demo agent and a production agent is entirely about reliability engineering. Here are the patterns that close that gap.
Why agents fail differently
Traditional software fails predictably — a null pointer, a timeout, a missing field. Agents fail in novel ways:
- Reasoning loops: The agent gets stuck retrying the same failed approach
- Tool misuse: Correct tool, wrong parameters, plausible-looking output
- Context drift: Accumulated errors in long conversations that compound
- Confidence without correctness: The agent produces wrong answers with no signal it's uncertain
These failure modes demand different solutions than traditional error handling.
Pattern 1: State machine with guardrails
Every production agent needs an explicit state machine. Without one, agents wander through unbounded execution paths.
from enum import Enum
from dataclasses import dataclass
from typing import Any, Callable
class AgentState(Enum):
PLANNING = "planning"
EXECUTING = "executing"
VALIDATING = "validating"
RECOVERING = "recovering"
ESCALATING = "escalating"
COMPLETE = "complete"
FAILED = "failed"
@dataclass
class AgentContext:
state: AgentState = AgentState.PLANNING
attempts: int = 0
max_attempts: int = 3
history: list[dict] = None
errors: list[str] = None
def __post_init__(self):
self.history = self.history or []
self.errors = self.errors or []
class ReliableAgent:
def __init__(self, tools: list[Callable], max_steps: int = 20):
self.tools = tools
self.max_steps = max_steps
self.step_count = 0
async def run(self, task: str) -> dict[str, Any]:
ctx = AgentContext()
while ctx.state not in (AgentState.COMPLETE, AgentState.FAILED):
self.step_count += 1
if self.step_count > self.max_steps:
ctx.state = AgentState.ESCALATING
break
match ctx.state:
case AgentState.PLANNING:
plan = await self._plan(task, ctx)
if plan.is_valid:
ctx.state = AgentState.EXECUTING
else:
ctx.state = AgentState.RECOVERING
case AgentState.EXECUTING:
result = await self._execute(plan, ctx)
ctx.state = AgentState.VALIDATING
case AgentState.VALIDATING:
if await self._validate(result, task):
ctx.state = AgentState.COMPLETE
else:
ctx.attempts += 1
if ctx.attempts >= ctx.max_attempts:
ctx.state = AgentState.ESCALATING
else:
ctx.state = AgentState.RECOVERING
case AgentState.RECOVERING:
recovery = await self._recover(ctx)
if recovery.should_retry:
ctx.state = AgentState.PLANNING
else:
ctx.state = AgentState.ESCALATING
case AgentState.ESCALATING:
await self._escalate_to_human(task, ctx)
ctx.state = AgentState.FAILED
return {"state": ctx.state.value, "result": result, "steps": self.step_count}
The key insight: never let an agent execute indefinitely. Bound execution with step limits, attempt counters, and mandatory validation checkpoints.
Pattern 2: Tool call validation
Agents hallucinate tool parameters. Validate every tool call before execution.
import { z } from "zod";
interface ToolDefinition {
name: string;
schema: z.ZodType;
execute: (params: unknown) => Promise<unknown>;
dangerLevel: "safe" | "moderate" | "dangerous";
}
class ValidatedToolExecutor {
private tools: Map<string, ToolDefinition>;
private executionLog: Array<{
tool: string;
params: unknown;
result: unknown;
timestamp: number;
}> = [];
async execute(
toolName: string,
params: unknown,
): Promise<{ success: boolean; result?: unknown; error?: string }> {
const tool = this.tools.get(toolName);
if (!tool) {
return { success: false, error: `Unknown tool: ${toolName}` };
}
// Validate parameters against schema
const validation = tool.schema.safeParse(params);
if (!validation.success) {
return {
success: false,
error: `Invalid params: ${validation.error.message}`,
};
}
// Check for dangerous operations
if (tool.dangerLevel === "dangerous") {
const approved = await this.requestHumanApproval(toolName, params);
if (!approved) {
return { success: false, error: "Human rejected dangerous operation" };
}
}
// Execute with timeout
try {
const result = await Promise.race([
tool.execute(validation.data),
this.timeout(30_000),
]);
this.executionLog.push({
tool: toolName,
params: validation.data,
result,
timestamp: Date.now(),
});
return { success: true, result };
} catch (e) {
return { success: false, error: e.message };
}
}
private timeout(ms: number): Promise<never> {
return new Promise((_, reject) =>
setTimeout(() => reject(new Error("Tool execution timeout")), ms),
);
}
}
Pattern 3: Progressive retry with backoff
Not all failures are equal. Transient errors (rate limits, timeouts) need retries. Logic errors (wrong approach) need replanning.
from enum import Enum
import asyncio
import random
class FailureType(Enum):
TRANSIENT = "transient" # Retry same approach
LOGIC = "logic" # Replan entirely
PERMANENT = "permanent" # Escalate immediately
def classify_failure(error: Exception, context: dict) -> FailureType:
"""Classify failure to determine recovery strategy."""
if isinstance(error, (TimeoutError, ConnectionError)):
return FailureType.TRANSIENT
if "rate_limit" in str(error).lower():
return FailureType.TRANSIENT
if isinstance(error, ValidationError):
return FailureType.LOGIC
if context.get("attempts", 0) >= 2:
# Same error twice = not transient
return FailureType.PERMANENT
return FailureType.LOGIC
async def retry_with_strategy(
func,
failure_type: FailureType,
max_retries: int = 3
) -> Any:
for attempt in range(max_retries):
try:
return await func()
except Exception as e:
ft = classify_failure(e, {"attempts": attempt})
if ft == FailureType.PERMANENT:
raise
if ft == FailureType.TRANSIENT:
# Exponential backoff with jitter
delay = (2 ** attempt) + random.uniform(0, 1)
await asyncio.sleep(delay)
elif ft == FailureType.LOGIC:
# Don't retry the same thing - signal replanning needed
raise ReplanRequired(f"Logic error after {attempt + 1} attempts: {e}")
Pattern 4: Conversation checkpointing
Long-running agents accumulate context that can drift. Checkpoint state at critical decision points so you can rollback.
The implementation stores snapshots of agent memory, tool state, and conversation history at each major transition. When validation fails, you rollback to the last valid checkpoint instead of starting over or continuing with corrupted state.
Production metrics that matter
Track these for every agent in production:
| Metric | Target | Alert Threshold |
|---|---|---|
| Task completion rate | >95% | <90% |
| Avg steps to completion | <8 | >15 |
| Escalation rate | <5% | >10% |
| Tool call validation failures | <10% | >20% |
| Mean time to completion | <30s | >60s |
| Loop detection triggers | <2% | >5% |
The human-in-the-loop escape valve
Every production agent needs a path to human intervention. Not as a failure state, but as a designed capability.
Design escalation triggers:
- Confidence below threshold on high-stakes decisions
- Multiple failed recovery attempts
- Novel situations not covered by the agent's training
- Actions that cross a danger threshold (deleting data, spending money, external communications)
The best agents escalate early and gracefully rather than grinding through uncertain territory.
Key takeaways
- Bound all agent execution with step limits and attempt counters — infinite loops are the #1 production incident
- Validate every tool call against a schema before execution
- Classify failures by type and match recovery strategy accordingly
- Checkpoint state so you can rollback, not just retry
- Design escalation as a feature, not a failure mode
- Monitor agent behavior with the same rigor you'd apply to any distributed system
Reliability isn't a feature you add later. It's the architecture you start with. Build the state machine and validation layer first, then make the agent smarter within those constraints.
Recommended reading

The State of Agentic AI in 2026: Capabilities, Limitations, and Production Readiness
Comprehensive analysis of agentic AI in 2026 covering production capabilities, current limitations, and enterprise readiness benchmarks with real deployment data.

Observability for AI Agents: Tracing Multi-Step Reasoning Chains in Production
How to implement production observability for AI agents including distributed tracing, reasoning chain analysis, and debugging multi-step failures.

Measuring and Reducing AI Workload Carbon Emissions: A Practical Engineering Guide
Building a carbon-aware scheduling system for ML training and inference workloads that reduced our AI infrastructure emissions by 42% while maintaining SLA commitments.

Comments
No comments yet. Be the first to share your thoughts.