Observability for AI Agents: Tracing Multi-Step Reasoning Chains in Production
How to implement production observability for AI agents including distributed tracing, reasoning chain analysis, and debugging multi-step failures.

Why Traditional Observability Breaks Down for AI Agents
Traditional application observability assumes deterministic execution paths. You instrument endpoints, trace requests through services, and alert on latency or error thresholds. AI agents break every one of these assumptions.
An agent processing the same input twice may take entirely different execution paths, call different tools in different orders, reason for variable durations, and produce different (but equally valid) outputs. Your existing Datadog dashboards are not equipped for this.
After instrumenting over 40 production agent deployments, I have converged on a set of patterns that actually work. The key insight: you are not monitoring a service — you are monitoring a decision-making process. Teams already running LLMs in production will recognize many of these challenges from their initial inference deployments. This builds on the broader operational discipline required when running LLMs in production.
The Three Layers of Agent Observability
Layer 1: Execution Tracing
Execution tracing captures what the agent did — every tool call, every model invocation, every state transition. This is the foundation.
// Minimal agent trace structure
interface AgentTrace {
traceId: string;
sessionId: string;
taskDescription: string;
startTime: number;
endTime: number;
steps: AgentStep[];
outcome: 'success' | 'failure' | 'escalated' | 'timeout';
totalTokens: { input: number; output: number };
totalCost: number;
metadata: Record<string, unknown>;
}
interface AgentStep {
stepId: string;
parentStepId?: string;
type: 'reasoning' | 'tool_call' | 'observation' | 'decision';
model?: string;
input: string;
output: string;
duration: number;
tokens: { input: number; output: number };
toolName?: string;
toolArgs?: Record<string, unknown>;
confidence?: number;
alternatives?: string[]; // Other options the agent considered
}
Layer 2: Reasoning Analysis
Execution tracing tells you what happened. Reasoning analysis tells you why. This layer extracts and evaluates the quality of the agent's decision-making at each step.
Key signals to extract from reasoning:
| Signal | How to Measure | Failure Indicator |
|---|---|---|
| Goal coherence | Semantic similarity between current action and original task | Score drops below 0.6 |
| Plan stability | Number of plan revisions per task | >3 revisions for simple tasks |
| Confidence calibration | Agent's stated confidence vs. actual success rate | >20% miscalibration |
| Information utilization | % of tool outputs referenced in subsequent reasoning | <40% utilization |
| Reasoning loops | Repeated similar reasoning steps | >2 near-duplicate steps |
Layer 3: Outcome Evaluation
The final layer connects agent actions to business outcomes. Did the agent's work actually achieve the intended result?
// Outcome evaluation pipeline
interface OutcomeEvaluation {
traceId: string;
taskType: string;
// Immediate outcomes
taskCompleted: boolean;
correctness: number; // 0-1, measured by automated checks
humanOverrideRequired: boolean;
// Downstream outcomes (measured async)
codeDeployed: boolean;
testsPassedInCI: boolean;
incidentCreated: boolean;
customerSatisfied: boolean;
// Cost efficiency
costVsHumanBaseline: number; // ratio
timeVsHumanBaseline: number; // ratio
}
Implementing Distributed Tracing for Multi-Agent Systems
When agents orchestrate other agents, you need distributed tracing that preserves the hierarchical relationship between orchestrator decisions and worker executions.
The Span Hierarchy
[Orchestrator Trace]
├── [Planning Span] - Task decomposition
│ ├── reasoning: "This task requires code changes and infrastructure updates"
│ └── decision: split into code_agent + infra_agent
├── [Delegation Span] - code_agent
│ ├── [Worker Trace - code_agent]
│ │ ├── [Tool Span] - file_read("src/api/routes.ts")
│ │ ├── [Reasoning Span] - "Need to add middleware before route handler"
│ │ ├── [Tool Span] - file_write("src/api/middleware/rateLimit.ts")
│ │ └── [Tool Span] - test_run("src/api/__tests__/rateLimit.test.ts")
│ └── result: success, 3 files modified
├── [Delegation Span] - infra_agent
│ ├── [Worker Trace - infra_agent]
│ │ ├── [Tool Span] - terraform_plan("modules/api-gateway")
│ │ ├── [Reasoning Span] - "Rate limit config needs WAF rule update"
│ │ └── [Tool Span] - terraform_apply("modules/api-gateway")
│ └── result: success, infrastructure updated
└── [Verification Span] - Results aggregation
├── reasoning: "Both sub-tasks completed, verifying integration"
└── outcome: success
Implementing with OpenTelemetry
OpenTelemetry provides the foundation, but you need custom semantic conventions for agent-specific attributes:
import { trace, SpanKind, context } from '@opentelemetry/api';
const tracer = trace.getTracer('agent-system', '1.0.0');
// Custom semantic conventions for agent traces
const AGENT_ATTRIBUTES = {
'agent.name': 'agent.name',
'agent.model': 'agent.model',
'agent.step_type': 'agent.step_type',
'agent.tool_name': 'agent.tool_name',
'agent.tokens_input': 'agent.tokens.input',
'agent.tokens_output': 'agent.tokens.output',
'agent.confidence': 'agent.confidence',
'agent.reasoning_summary': 'agent.reasoning_summary',
'agent.plan_revision_count': 'agent.plan_revision_count',
} as const;
async function traceAgentStep(
stepType: string,
fn: () => Promise<AgentStepResult>
): Promise<AgentStepResult> {
return tracer.startActiveSpan(
`agent.${stepType}`,
{ kind: SpanKind.INTERNAL },
async (span) => {
try {
const result = await fn();
span.setAttributes({
[AGENT_ATTRIBUTES['agent.step_type']]: stepType,
[AGENT_ATTRIBUTES['agent.tokens_input']]: result.tokensUsed.input,
[AGENT_ATTRIBUTES['agent.tokens_output']]: result.tokensUsed.output,
[AGENT_ATTRIBUTES['agent.confidence']]: result.confidence,
});
return result;
} catch (error) {
span.recordException(error as Error);
span.setStatus({ code: 2, message: (error as Error).message });
throw error;
} finally {
span.end();
}
}
);
}
Critical Alerts for Production Agent Systems
Not all agent failures look like errors. Many manifest as subtle degradation patterns that traditional alerting misses entirely.
Alert Configuration
| Alert | Condition | Severity | Action |
|---|---|---|---|
| Reasoning Loop | >3 semantically similar steps in sequence | P2 | Kill task, escalate |
| Cost Explosion | Task cost >10x median for task type | P3 | Throttle, notify |
| Confidence Collapse | Agent confidence drops >50% between steps | P2 | Pause, request review |
| Tool Hallucination | Tool called with params matching no valid schema | P1 | Abort immediately |
Detecting these hallucinated tool calls is a specific case of LLM hallucination detection applied to structured outputs rather than free text. | Goal Drift | Task-action similarity <0.4 for 3+ steps | P2 | Reset to last checkpoint | | Cascading Failure | >3 agents failing within 60s window | P1 | Circuit break all agents |
Implementing Goal Drift Detection
Goal drift — where an agent gradually deviates from its original task — is the most insidious production failure. It does not trigger errors; it just produces wrong results confidently. Detecting drift early is closely related to the challenge of LLM hallucination detection and mitigation.
import { cosineSimilarity } from './embeddings';
interface GoalDriftDetector {
originalTaskEmbedding: number[];
threshold: number;
windowSize: number;
scores: number[];
}
function checkGoalDrift(
detector: GoalDriftDetector,
currentActionEmbedding: number[]
): { drifting: boolean; score: number; trend: 'stable' | 'declining' | 'recovering' } {
const score = cosineSimilarity(detector.originalTaskEmbedding, currentActionEmbedding);
detector.scores.push(score);
const recentScores = detector.scores.slice(-detector.windowSize);
const avgScore = recentScores.reduce((a, b) => a + b, 0) / recentScores.length;
// Detect trend
let trend: 'stable' | 'declining' | 'recovering' = 'stable';
if (recentScores.length >= 3) {
const firstHalf = recentScores.slice(0, Math.floor(recentScores.length / 2));
const secondHalf = recentScores.slice(Math.floor(recentScores.length / 2));
const firstAvg = firstHalf.reduce((a, b) => a + b, 0) / firstHalf.length;
const secondAvg = secondHalf.reduce((a, b) => a + b, 0) / secondHalf.length;
if (secondAvg < firstAvg - 0.1) trend = 'declining';
else if (secondAvg > firstAvg + 0.1) trend = 'recovering';
}
return {
drifting: avgScore < detector.threshold,
score: avgScore,
trend,
};
}
Debugging Multi-Step Failures: A Systematic Approach
When an agent fails in production, you need to answer: At which step did the reasoning go wrong, and why?
The Five-Phase Debug Protocol
-
Identify the divergence point: Compare the failed trace against successful traces for the same task type. Find where they diverge.
-
Classify the failure mode:
- Planning failure (wrong decomposition)
- Tool-use failure (wrong tool or wrong parameters)
- Observation failure (misinterpreted tool output)
- Reasoning failure (correct information, wrong conclusion)
- External failure (API timeout, rate limit, data corruption)
-
Check the context window: Was critical information present in the context when the decision was made? Context window overflow is responsible for 23% of agent failures in production. For a production-tested approach to this problem, see how Kiro handles complex DevOps workflows with structured context management.
-
Evaluate alternatives: Did the agent consider the correct action? If it appeared in the reasoning but was rejected, the issue is evaluation. If it never appeared, the issue is generation.
-
Test the fix in simulation: Replay the trace with modified prompts, tools, or guardrails. Measure whether the fix resolves the failure without regressing other cases.
Production Debugging Dashboard Metrics
| Metric | Description | Healthy Range |
|---|---|---|
| Steps per Success | Average steps for successful task completion | 3-7 |
| Failure Step Distribution | Which step number most failures occur at | Even distribution |
| Tool Error Rate | % of tool calls returning errors | <5% |
| Context Utilization | % of context window used at failure | <80% |
| Reasoning Token Ratio | Reasoning tokens / total tokens | 30-50% |
| Recovery Success Rate | % of self-corrections that resolve the issue | >60% |
Building Your Observability Stack
Recommended Architecture
# Agent observability stack
collection:
- opentelemetry-sdk # Trace collection
- custom-agent-exporter # Agent-specific span enrichment
storage:
- tempo/jaeger # Trace storage (short-term)
- clickhouse # Analytics (long-term)
- s3 # Full trace archives (compliance)
analysis:
- grafana # Dashboards and alerting
- langsmith # Agent-specific replay and evaluation
- custom-eval-pipeline # Automated quality scoring
alerting:
- pagerduty # P1/P2 alerts
- slack # P3/P4 notifications
- automated-rollback # Circuit breaker triggers
Cost of Observability
Agent observability is not free. The overhead matters:
- Token overhead for extracting reasoning metadata: 5-8% additional tokens
- Latency overhead for trace export: <50ms per step (async)
- Storage costs: ~2KB per step compressed, ~$0.15/million steps in cloud storage
- Evaluation costs (running LLM-as-judge on traces): $0.02-0.05 per trace
For a system processing 100,000 agent tasks per month, expect $2,000-5,000/month in observability infrastructure costs. This is 3-7% of total agent operational cost — well within acceptable overhead.
Key Takeaways
- Traditional observability tools are insufficient for AI agents — you need reasoning-aware tracing
- Implement three layers: execution tracing, reasoning analysis, and outcome evaluation
- Goal drift detection is critical — agents fail silently by deviating from their original task
- Budget 3-7% of agent operational costs for observability infrastructure
- The debugging protocol starts with divergence detection, not log searching
- Alert on behavioral patterns (loops, drift, confidence collapse) not just errors
Frequently Asked Questions
Can I use existing APM tools for agent observability?
Partially. Tools like Datadog and New Relic can capture the execution layer (API calls, latencies, errors), but they lack the reasoning analysis capabilities needed for agent-specific failure modes. You will need to supplement with agent-specific tooling like LangSmith, Arize, or custom evaluation pipelines.
How much trace data should I retain?
Retain full traces (including reasoning content) for 7-14 days for debugging. Archive compressed trace metadata (without full content) for 90 days for trend analysis. For compliance-regulated industries, archive full traces for the required retention period (often 7 years) in cold storage.
What is the minimum viable agent observability setup?
At minimum: structured logging of every tool call (input, output, duration), task-level success/failure tracking, and cost per task. This gives you enough to debug most issues. Add reasoning analysis and goal drift detection as you scale past 1,000 tasks per day.
How do I handle observability for multi-agent systems where agents call other agents?
Use distributed tracing with context propagation — exactly like microservices. The orchestrator's trace ID propagates to all child agents. Each child creates child spans under the parent. This preserves the full execution tree for debugging.
Recommended reading

The State of Agentic AI in 2026: Capabilities, Limitations, and Production Readiness
Comprehensive analysis of agentic AI in 2026 covering production capabilities, current limitations, and enterprise readiness benchmarks with real deployment data.

Measuring and Reducing AI Workload Carbon Emissions: A Practical Engineering Guide
Building a carbon-aware scheduling system for ML training and inference workloads that reduced our AI infrastructure emissions by 42% while maintaining SLA commitments.

AI-Assisted Capacity Planning: Surviving Black Friday Without Over-Provisioning
How we used ML forecasting models to predict Black Friday traffic patterns, pre-provision infrastructure with surgical precision, and handle 47x normal load without wasting $180K on idle capacity.

Comments
No comments yet. Be the first to share your thoughts.