AWS Bedrock Model Routing Strategies: Optimizing Cost and Quality at Scale

How we built an intelligent model routing layer that reduced our LLM inference costs by 62% while maintaining output quality above our SLA thresholds.

#aws#bedrock#ai#llm
Cover image for the article: AWS Bedrock Model Routing Strategies: Optimizing Cost and Quality at Scale

When we first integrated LLMs into our product, every request went to Claude 3.5 Sonnet. The output quality was excellent. The invoice was not. At 40,000 API calls per day, we were spending $34,000/month on a single model endpoint. Most of those calls were simple classification tasks, summarizations under 200 tokens, or structured data extraction that a smaller model handles perfectly well.

We built a routing layer that dispatches requests to the optimal model based on task complexity, latency requirements, and cost constraints. Monthly LLM costs dropped from $34,000 to $12,900 while quality scores on our evaluation suite remained above 94%.

The Problem: One Model Does Not Fit All Tasks

Our platform uses LLMs for seven distinct task categories:

  • Document classification — Categorize incoming documents (simple, low-token)
  • Summarization — Condense long documents into briefs (medium complexity)
  • Entity extraction — Pull structured fields from unstructured text (structured output)
  • Content generation — Draft customer-facing responses (high quality required)
  • Code analysis — Review code snippets for issues (specialized reasoning)
  • Translation — Multi-language support (medium complexity)
  • Complex reasoning — Multi-step analysis and decision support (highest quality)

Sending all seven categories to the same frontier model is like using a Ferrari for grocery runs. It works, but the cost-per-mile is absurd.

Model Routing Architecture

Architecture: The Router Pattern

Our routing layer sits between the application and Bedrock endpoints. It makes routing decisions based on three inputs:

  1. Task metadata — Declared by the calling service (task type, quality tier, max latency)
  2. Input characteristics — Token count, language, structural complexity
  3. Historical performance — Cached quality scores from our evaluation pipeline

The router selects from a tiered model pool:

TierModelCost (per 1M input tokens)Use Case
EconomyHaiku 3.5$0.80Classification, simple extraction
StandardSonnet 3.5$3.00Summarization, translation, moderate reasoning
PremiumOpus 4$15.00Complex reasoning, high-stakes content
SpecializedMistral Large$4.00Code analysis, structured output

Implementation: The Routing Decision Engine

import {
  BedrockRuntimeClient,
  InvokeModelCommand,
} from '@aws-sdk/client-bedrock-runtime';

interface RoutingRequest {
  taskType: TaskType;
  qualityTier: 'economy' | 'standard' | 'premium';
  maxLatencyMs: number;
  inputTokenEstimate: number;
  input: string;
  systemPrompt: string;
}

interface ModelConfig {
  modelId: string;
  maxTokensPerSecond: number;
  costPerInputToken: number;
  costPerOutputToken: number;
  qualityScore: Record<TaskType, number>;
}

const MODEL_REGISTRY: Record<string, ModelConfig> = {
  'haiku-3.5': {
    modelId: 'anthropic.claude-3-5-haiku-20241022-v1:0',
    maxTokensPerSecond: 4000,
    costPerInputToken: 0.0000008,
    costPerOutputToken: 0.000004,
    qualityScore: {
      classification: 0.96,
      summarization: 0.82,
      extraction: 0.91,
      generation: 0.74,
      code_analysis: 0.78,
      translation: 0.85,
      reasoning: 0.68,
    },
  },
  'sonnet-3.5': {
    modelId: 'anthropic.claude-3-5-sonnet-20241022-v2:0',
    maxTokensPerSecond: 2000,
    costPerInputToken: 0.000003,
    costPerOutputToken: 0.000015,
    qualityScore: {
      classification: 0.98,
      summarization: 0.94,
      extraction: 0.96,
      generation: 0.93,
      code_analysis: 0.91,
      translation: 0.94,
      reasoning: 0.89,
    },
  },
  'opus-4': {
    modelId: 'anthropic.claude-opus-4-20250514-v1:0',
    maxTokensPerSecond: 1000,
    costPerInputToken: 0.000015,
    costPerOutputToken: 0.000075,
    qualityScore: {
      classification: 0.99,
      summarization: 0.97,
      extraction: 0.98,
      generation: 0.98,
      code_analysis: 0.96,
      translation: 0.97,
      reasoning: 0.97,
    },
  },
};

function selectModel(request: RoutingRequest): ModelConfig {
  const candidates = Object.values(MODEL_REGISTRY);

  // Filter by latency constraint
  const latencyFiltered = candidates.filter((model) => {
    const estimatedLatency =
      (request.inputTokenEstimate / model.maxTokensPerSecond) * 1000;
    return estimatedLatency < request.maxLatencyMs * 0.7; // 30% buffer
  });

  // Filter by quality threshold based on tier
  const qualityThreshold =
    request.qualityTier === 'premium'
      ? 0.95
      : request.qualityTier === 'standard'
        ? 0.88
        : 0.78;

  const qualityFiltered = latencyFiltered.filter(
    (model) => model.qualityScore[request.taskType] >= qualityThreshold
  );

  // Select cheapest model that meets all constraints
  return qualityFiltered.sort(
    (a, b) => a.costPerInputToken - b.costPerInputToken
  )[0];
}

The critical design decision: we select the cheapest model that meets both latency and quality constraints rather than the highest-quality model within budget. This naturally pushes simple tasks to cheaper models.

Fallback and Quality Monitoring

The router includes automatic fallback when a model returns low-confidence results:

async function invokeWithFallback(
  request: RoutingRequest,
  client: BedrockRuntimeClient
): Promise<ModelResponse> {
  const primaryModel = selectModel(request);
  const response = await invokeModel(client, primaryModel, request);

  // Check response quality signals
  if (response.stopReason === 'max_tokens' && request.qualityTier !== 'economy') {
    // Response was truncated — retry with a higher-tier model
    const upgradedRequest = { ...request, qualityTier: 'premium' as const };
    const fallbackModel = selectModel(upgradedRequest);
    return invokeModel(client, fallbackModel, request);
  }

  // Log routing decision for offline evaluation
  await logRoutingDecision({
    requestId: response.requestId,
    taskType: request.taskType,
    selectedModel: primaryModel.modelId,
    inputTokens: response.inputTokens,
    outputTokens: response.outputTokens,
    latencyMs: response.latencyMs,
    cost: calculateCost(primaryModel, response),
  });

  return response;
}

Cost Impact: Before and After

After 60 days of production routing with 1.2 million daily requests:

Task TypeBefore (model)After (model)Cost Reduction
ClassificationSonnet ($4,200/mo)Haiku ($560/mo)86.7%
SummarizationSonnet ($8,100/mo)Sonnet ($8,100/mo)0%
ExtractionSonnet ($5,400/mo)Haiku ($720/mo)86.7%
GenerationSonnet ($7,200/mo)Sonnet ($7,200/mo)0%
Code AnalysisSonnet ($3,600/mo)Mistral ($1,920/mo)46.7%
TranslationSonnet ($2,700/mo)Haiku ($360/mo)86.7%
ReasoningSonnet ($2,800/mo)Opus ($4,040/mo)-44.3%
Total$34,000/mo$12,900/mo62.1%

Notice reasoning tasks actually cost more now. We upgraded those from Sonnet to Opus because our quality evaluation showed Sonnet was producing errors in 11% of complex reasoning tasks. Spending more on the hard tasks while spending less on the easy ones was the right tradeoff.

Model Routing Cost Comparison

Quality Evaluation Pipeline

Cost savings mean nothing if quality degrades. We run a nightly evaluation pipeline:

  1. Sample 500 requests per task type from production logs
  2. Send the same inputs to all candidate models
  3. Score outputs using a combination of automated metrics (ROUGE, exact match for extraction) and a judge model (Opus evaluating Haiku/Sonnet outputs)
  4. Update the quality score matrix in the model registry
  5. Alert if any model's score drops below threshold for any task type

This caught a regression in Haiku's extraction quality after a model update. The router automatically shifted extraction tasks to Sonnet until the quality scores recovered.

Lessons Learned

Declare task type at call site, not in the router. Early versions tried to auto-classify task complexity by analyzing the prompt. This added 200ms latency and was wrong 15% of the time. Having the calling service declare its task type is simpler and more accurate.

Cache routing decisions, not model responses. Model responses for the same prompt can legitimately vary. But routing decisions for a given (task_type, quality_tier, token_estimate) tuple are deterministic. Caching these saves the routing computation.

Start with two tiers, not four. We began with just Haiku and Sonnet. Only added Opus and Mistral after 30 days of data showed where Sonnet was either overkill or insufficient. Let production data drive tier expansion.

Monitor cost per successful output, not cost per request. If a cheap model fails 20% of the time and requires retry on a more expensive model, the effective cost is higher than routing to the better model initially.

Conclusion

Intelligent model routing is the single highest-leverage optimization for production LLM applications. Our 62% cost reduction came not from degrading quality but from matching model capability to task requirement. The investment was roughly two engineer-weeks to build the router and evaluation pipeline. At $21,100/month in savings, payback was under a week.

Comments

    No comments yet. Be the first to share your thoughts.