Claude API Production Integration Patterns

Battle-tested patterns for integrating Claude into production systems with proper error handling, retries, and observability.

#claude#anthropic#ai#api-design
Cover image for the article: Claude API Production Integration Patterns

Integrating Claude into a production system is not the same as calling an API in a Jupyter notebook. After running Claude-powered features serving millions of requests per month, I've distilled the patterns that keep things reliable when the pressure is real.

The problem with naive integration

Most teams start with something like this: a direct API call wrapped in a try-catch. It works in development. Then production happens — rate limits hit at 2 AM, responses take 45 seconds during peak load, and malformed JSON crashes your parsing layer.

The root cause is treating an LLM API like a deterministic REST endpoint. It isn't. Claude's responses are variable-length, variable-latency, and occasionally surprising in structure. Your integration layer needs to account for all of this.

Architecture: the resilient Claude client

Claude API Integration Architecture

The architecture I've landed on after multiple iterations has four layers:

  1. Request formation — prompt assembly, token budgeting, parameter selection
  2. Transport — retries, circuit breaking, timeout management
  3. Response processing — validation, parsing, fallback handling
  4. Observability — cost tracking, latency percentiles, quality scoring

Layer 1: Request formation with token budgeting

Before sending any request, calculate your token budget explicitly. Don't rely on the API to truncate — you'll get unpredictable behavior.

import anthropic
from dataclasses import dataclass
from typing import Optional

@dataclass
class TokenBudget:
    max_input: int = 180_000
    max_output: int = 4_096
    reserved_for_system: int = 2_000

    @property
    def available_for_context(self) -> int:
        return self.max_input - self.reserved_for_system

class ClaudeRequestBuilder:
    def __init__(self, client: anthropic.Anthropic, budget: TokenBudget):
        self.client = client
        self.budget = budget
        self.token_counter = client.count_tokens

    def build_messages(
        self,
        system_prompt: str,
        user_content: str,
        context_documents: list[str],
    ) -> dict:
        system_tokens = self._count(system_prompt)
        user_tokens = self._count(user_content)
        remaining = self.budget.available_for_context - system_tokens - user_tokens

        # Greedily pack context documents within budget
        included_docs = []
        for doc in context_documents:
            doc_tokens = self._count(doc)
            if doc_tokens <= remaining:
                included_docs.append(doc)
                remaining -= doc_tokens

        context_block = "\n---\n".join(included_docs)
        full_user_content = f"{context_block}\n\n{user_content}" if included_docs else user_content

        return {
            "model": "claude-sonnet-4-20250514",
            "max_tokens": self.budget.max_output,
            "system": system_prompt,
            "messages": [{"role": "user", "content": full_user_content}],
        }

    def _count(self, text: str) -> int:
        return self.token_counter(model="claude-sonnet-4-20250514", messages=[
            {"role": "user", "content": text}
        ]).input_tokens

Layer 2: Transport with exponential backoff and circuit breaking

The transport layer handles the reality of distributed systems. Rate limits, network blips, and overloaded endpoints are not edge cases — they're Tuesday.

import Anthropic from "@anthropic-ai/sdk";

interface RetryConfig {
  maxRetries: number;
  baseDelay: number;
  maxDelay: number;
  retryableStatuses: number[];
}

const DEFAULT_RETRY: RetryConfig = {
  maxRetries: 3,
  baseDelay: 1000,
  maxDelay: 30000,
  retryableStatuses: [429, 500, 502, 503, 529],
};

class ResilientClaudeClient {
  private client: Anthropic;
  private circuitOpen = false;
  private failureCount = 0;
  private lastFailure = 0;
  private readonly failureThreshold = 5;
  private readonly recoveryWindow = 60_000;

  constructor(apiKey: string) {
    this.client = new Anthropic({ apiKey });
  }

  async createMessage(
    params: Anthropic.MessageCreateParams
  ): Promise<Anthropic.Message> {
    if (this.isCircuitOpen()) {
      throw new Error("Circuit breaker open — Claude API unavailable");
    }

    for (let attempt = 0; attempt <= DEFAULT_RETRY.maxRetries; attempt++) {
      try {
        const response = await this.client.messages.create(params);
        this.recordSuccess();
        return response;
      } catch (error: any) {
        if (!this.isRetryable(error) || attempt === DEFAULT_RETRY.maxRetries) {
          this.recordFailure();
          throw error;
        }
        const delay = Math.min(
          DEFAULT_RETRY.baseDelay * Math.pow(2, attempt) + Math.random() * 1000,
          DEFAULT_RETRY.maxDelay
        );
        await this.sleep(delay);
      }
    }
    throw new Error("Exhausted retries");
  }

  private isCircuitOpen(): boolean {
    if (!this.circuitOpen) return false;
    if (Date.now() - this.lastFailure > this.recoveryWindow) {
      this.circuitOpen = false;
      this.failureCount = 0;
      return false;
    }
    return true;
  }

  private recordSuccess(): void {
    this.failureCount = 0;
  }

  private recordFailure(): void {
    this.failureCount++;
    this.lastFailure = Date.now();
    if (this.failureCount >= this.failureThreshold) {
      this.circuitOpen = true;
    }
  }

  private isRetryable(error: any): boolean {
    return DEFAULT_RETRY.retryableStatuses.includes(error?.status);
  }

  private sleep(ms: number): Promise<void> {
    return new Promise((resolve) => setTimeout(resolve, ms));
  }
}

Layer 3: Response validation

Never trust the structure of an LLM response. Even with explicit JSON instructions, Claude occasionally wraps output in markdown code fences or includes preamble text.

My approach: define a Zod schema for every expected response shape, attempt parsing, and retry with the validation error injected into the prompt if it fails. Two retries maximum — after that, fall back to a degraded experience.

Layer 4: Observability that matters

Track these metrics from day one:

MetricWhy it matters
P50/P95/P99 latencyCapacity planning and SLA compliance
Token cost per requestBudget forecasting — costs creep silently
Retry rateEarly warning for API degradation
Circuit breaker tripsTells you about systemic issues
Response validation failure ratePrompt drift or model behavior changes

Benchmarks from production

After implementing this pattern across three services:

  • Error rate dropped from 4.2% to 0.3% — mostly by handling transient failures gracefully
  • P99 latency improved 40% — circuit breaking prevents cascading timeouts
  • Cost per request dropped 18% — token budgeting eliminated wasted context

Key takeaways

  1. Treat Claude like an unreliable network peer, not a function call. Design for failure at every layer.
  2. Budget tokens explicitly. Don't rely on API truncation — you lose control over what gets included.
  3. Circuit breakers are not optional. One degraded upstream service should not take down your entire system.
  4. Validate every response structurally. Schema validation with retry is cheaper than debugging malformed data downstream.
  5. Measure cost as a first-class metric. Token usage is your cloud bill for AI — treat it with the same rigor.

These patterns have held up across queue-based batch workloads, real-time user-facing features, and internal tooling. The investment in the resilience layer pays for itself within the first week of production traffic.

Comments

    No comments yet. Be the first to share your thoughts.