Systematic Model Evaluation Frameworks for Claude in Production

Building comprehensive evaluation and benchmarking systems for Claude deployments with custom metrics, regression detection, and continuous monitoring.

#claude#evaluation#benchmarking#ai
Cover image for the article: Systematic Model Evaluation Frameworks for Claude in Production

You can't improve what you don't measure. And you definitely can't catch regressions in what you don't continuously evaluate. After running Claude in production across multiple features, I've built an evaluation framework that answers the three questions every engineering team needs answered: Is the model performing well? Is it getting worse? Should we switch to a different model or version?

The evaluation gap in AI engineering

Most teams ship an LLM feature, eyeball a few responses, call it good, and move on. Then a model update lands, quality subtly degrades, and nobody notices for weeks until users start complaining. Or worse — quality degrades on a subset of inputs that the team never tests manually.

The evaluation gap exists because:

  1. LLM outputs are non-deterministic — the same input produces different outputs
  2. "Quality" is multi-dimensional and subjective
  3. Traditional software testing (exact match assertions) doesn't work
  4. Evaluation is expensive — it often requires another LLM call

The framework I'll describe addresses all four of these challenges.

Evaluation Framework Architecture

Architecture: the evaluation pipeline

Component 1: Evaluation dataset management

Your eval dataset is a living asset. It grows over time, captures real production edge cases, and is versioned alongside your code.

from dataclasses import dataclass, field
from enum import Enum
from typing import Optional
import json
from pathlib import Path
from datetime import datetime

class Difficulty(Enum):
    TRIVIAL = "trivial"
    STANDARD = "standard"
    COMPLEX = "complex"
    ADVERSARIAL = "adversarial"

class EvalCategory(Enum):
    ACCURACY = "accuracy"
    FORMATTING = "formatting"
    SAFETY = "safety"
    EDGE_CASE = "edge_case"
    REGRESSION = "regression"  # Added after a bug was found

@dataclass
class EvalExample:
    id: str
    input_data: dict
    expected_output: Optional[dict] = None  # For exact-match checks
    evaluation_criteria: list[str] = field(default_factory=list)
    difficulty: Difficulty = Difficulty.STANDARD
    category: EvalCategory = EvalCategory.ACCURACY
    added_date: str = field(default_factory=lambda: datetime.now().isoformat())
    source: str = "manual"  # manual, production_failure, user_report
    metadata: dict = field(default_factory=dict)

class EvalDatasetManager:
    """Manage versioned evaluation datasets."""

    def __init__(self, dataset_path: Path):
        self.dataset_path = dataset_path
        self.examples: list[EvalExample] = self._load()

    def add_from_production_failure(
        self,
        input_data: dict,
        failed_output: str,
        correct_behavior: str,
        category: EvalCategory = EvalCategory.REGRESSION,
    ) -> EvalExample:
        """Add a new eval example from a production failure."""
        example = EvalExample(
            id=f"prod_fail_{len(self.examples):04d}",
            input_data=input_data,
            evaluation_criteria=[
                f"Must NOT produce: {failed_output[:200]}",
                f"Should satisfy: {correct_behavior}",
            ],
            difficulty=Difficulty.COMPLEX,
            category=category,
            source="production_failure",
            metadata={"failed_output_sample": failed_output[:500]},
        )
        self.examples.append(example)
        self._save()
        return example

    def get_by_category(self, category: EvalCategory) -> list[EvalExample]:
        return [e for e in self.examples if e.category == category]

    def get_by_difficulty(self, difficulty: Difficulty) -> list[EvalExample]:
        return [e for e in self.examples if e.difficulty == difficulty]

    def stats(self) -> dict:
        return {
            "total": len(self.examples),
            "by_category": {
                cat.value: len(self.get_by_category(cat))
                for cat in EvalCategory
            },
            "by_difficulty": {
                diff.value: len(self.get_by_difficulty(diff))
                for diff in Difficulty
            },
        }

    def _load(self) -> list[EvalExample]:
        if not self.dataset_path.exists():
            return []
        with open(self.dataset_path) as f:
            return [EvalExample(**item) for item in json.load(f)]

    def _save(self) -> None:
        with open(self.dataset_path, "w") as f:
            json.dump(
                [vars(e) for e in self.examples],
                f,
                indent=2,
                default=str,
            )

Component 2: Multi-dimensional scoring

Quality isn't one number. Evaluate across multiple dimensions, each with its own scoring method:

import Anthropic from "@anthropic-ai/sdk";

interface DimensionScore {
  dimension: string;
  score: number; // 0.0 to 1.0
  method: "deterministic" | "llm_judge" | "hybrid";
  details: string;
}

interface EvalResult {
  exampleId: string;
  modelVersion: string;
  scores: DimensionScore[];
  overallScore: number;
  latencyMs: number;
  tokenUsage: { input: number; output: number };
  timestamp: string;
}

class MultiDimensionalEvaluator {
  private client: Anthropic;
  private judgeModel = "claude-sonnet-4-20250514";

  constructor(client: Anthropic) {
    this.client = client;
  }

  async evaluate(
    modelOutput: string,
    example: EvalExample,
    dimensions: EvalDimension[]
  ): Promise<DimensionScore[]> {
    const scores: DimensionScore[] = [];

    for (const dimension of dimensions) {
      let score: DimensionScore;

      switch (dimension.method) {
        case "deterministic":
          score = this.evalDeterministic(modelOutput, dimension);
          break;
        case "llm_judge":
          score = await this.evalWithJudge(modelOutput, example, dimension);
          break;
        case "hybrid":
          const det = this.evalDeterministic(modelOutput, dimension);
          const judge = await this.evalWithJudge(
            modelOutput,
            example,
            dimension
          );
          score = {
            dimension: dimension.name,
            score: det.score * 0.4 + judge.score * 0.6,
            method: "hybrid",
            details: `Deterministic: ${det.score}, Judge: ${judge.score}`,
          };
          break;
      }

      scores.push(score);
    }

    return scores;
  }

  private evalDeterministic(
    output: string,
    dimension: EvalDimension
  ): DimensionScore {
    let score = 1.0;
    const checks: string[] = [];

    // Format validation
    if (dimension.expectedFormat === "json") {
      try {
        JSON.parse(output);
        checks.push("Valid JSON: PASS");
      } catch {
        score = 0;
        checks.push("Valid JSON: FAIL");
      }
    }

    // Length constraints
    if (dimension.maxLength && output.length > dimension.maxLength) {
      score *= 0.5;
      checks.push(`Length ${output.length} > max ${dimension.maxLength}`);
    }

    // Required content
    if (dimension.mustContain) {
      for (const term of dimension.mustContain) {
        if (!output.toLowerCase().includes(term.toLowerCase())) {
          score *= 0.8;
          checks.push(`Missing required: "${term}"`);
        }
      }
    }

    return {
      dimension: dimension.name,
      score,
      method: "deterministic",
      details: checks.join("; "),
    };
  }

  private async evalWithJudge(
    output: string,
    example: EvalExample,
    dimension: EvalDimension
  ): Promise<DimensionScore> {
    const judgePrompt = `You are evaluating an AI model's output on the dimension: ${dimension.name}

Evaluation criteria: ${dimension.criteria}

User input that was given to the model:
${JSON.stringify(example.inputData)}

Model's output:
${output}

Rate the output on a scale of 0.0 to 1.0 for the dimension "${dimension.name}".
Respond with ONLY a JSON object: {"score": <number>, "reasoning": "<brief explanation>"}`;

    const response = await this.client.messages.create({
      model: this.judgeModel,
      max_tokens: 256,
      messages: [{ role: "user", content: judgePrompt }],
    });

    const text =
      response.content[0].type === "text" ? response.content[0].text : "";
    const parsed = JSON.parse(text);

    return {
      dimension: dimension.name,
      score: parsed.score,
      method: "llm_judge",
      details: parsed.reasoning,
    };
  }
}

Component 3: Regression detection

The most valuable part of the framework: automatically detecting when quality degrades.

SignalDetection methodAlert threshold
Overall score dropCompare running avg to baseline>5% drop sustained 24h
Dimension-specific regressionPer-dimension tracking>10% drop on any dimension
New failure patternsCluster analysis on failures>3 similar failures in 24h
Latency regressionP95 tracking>50% increase from baseline
Cost regressionPer-request cost tracking>20% increase

Run evaluations on two triggers:

  1. On model/prompt changes — full eval suite before deploying any change
  2. Continuous sampling — evaluate 5% of production traffic nightly to catch drift

Model comparison framework

When Anthropic releases a new model version or you consider switching between Sonnet and Haiku, run a structured comparison:

@dataclass
class ModelComparison:
    model_a: str
    model_b: str
    dataset_size: int
    results: dict

    def summary(self) -> str:
        wins_a = sum(
            1 for r in self.results["comparisons"]
            if r["winner"] == self.model_a
        )
        wins_b = sum(
            1 for r in self.results["comparisons"]
            if r["winner"] == self.model_b
        )
        ties = self.dataset_size - wins_a - wins_b

        return (
            f"Model comparison ({self.dataset_size} examples):\n"
            f"  {self.model_a}: {wins_a} wins ({wins_a/self.dataset_size:.1%})\n"
            f"  {self.model_b}: {wins_b} wins ({wins_b/self.dataset_size:.1%})\n"
            f"  Ties: {ties} ({ties/self.dataset_size:.1%})\n"
            f"  Cost ratio: {self.results['cost_ratio']:.2f}x\n"
            f"  Latency ratio: {self.results['latency_ratio']:.2f}x"
        )

Production evaluation results

Our evaluation framework across three production features (200 eval examples each):

FeatureBaseline scoreCurrent scoreRegressions caughtFalse alerts
Customer support0.890.9241
Document extraction0.910.9320
Code review assistant0.860.8862

Every regression caught was a real quality degradation — either from a prompt change or a model update. Without the framework, these would have been discovered weeks later via user complaints.

Key takeaways

  1. Evaluation is infrastructure, not a one-time task. Build it once, run it continuously, grow it from production failures.
  2. Multi-dimensional scoring reveals what single scores hide. A feature can score 0.9 overall while being 0.6 on a critical dimension.
  3. LLM-as-judge plus deterministic checks is the winning combination. Neither alone is sufficient for production evaluation.
  4. Every production failure becomes a test case. Your eval dataset should grow faster from production than from manual curation.
  5. Run evals on two cadences. Every change (pre-deploy) and continuous sampling (post-deploy). Both are required.

The teams that ship AI features confidently are the ones that measure quality systematically. Everything else is hope-driven development.

Comments

    No comments yet. Be the first to share your thoughts.