Systematic Model Evaluation Frameworks for Claude in Production
Building comprehensive evaluation and benchmarking systems for Claude deployments with custom metrics, regression detection, and continuous monitoring.

You can't improve what you don't measure. And you definitely can't catch regressions in what you don't continuously evaluate. After running Claude in production across multiple features, I've built an evaluation framework that answers the three questions every engineering team needs answered: Is the model performing well? Is it getting worse? Should we switch to a different model or version?
The evaluation gap in AI engineering
Most teams ship an LLM feature, eyeball a few responses, call it good, and move on. Then a model update lands, quality subtly degrades, and nobody notices for weeks until users start complaining. Or worse — quality degrades on a subset of inputs that the team never tests manually.
The evaluation gap exists because:
- LLM outputs are non-deterministic — the same input produces different outputs
- "Quality" is multi-dimensional and subjective
- Traditional software testing (exact match assertions) doesn't work
- Evaluation is expensive — it often requires another LLM call
The framework I'll describe addresses all four of these challenges.
Architecture: the evaluation pipeline
Component 1: Evaluation dataset management
Your eval dataset is a living asset. It grows over time, captures real production edge cases, and is versioned alongside your code.
from dataclasses import dataclass, field
from enum import Enum
from typing import Optional
import json
from pathlib import Path
from datetime import datetime
class Difficulty(Enum):
TRIVIAL = "trivial"
STANDARD = "standard"
COMPLEX = "complex"
ADVERSARIAL = "adversarial"
class EvalCategory(Enum):
ACCURACY = "accuracy"
FORMATTING = "formatting"
SAFETY = "safety"
EDGE_CASE = "edge_case"
REGRESSION = "regression" # Added after a bug was found
@dataclass
class EvalExample:
id: str
input_data: dict
expected_output: Optional[dict] = None # For exact-match checks
evaluation_criteria: list[str] = field(default_factory=list)
difficulty: Difficulty = Difficulty.STANDARD
category: EvalCategory = EvalCategory.ACCURACY
added_date: str = field(default_factory=lambda: datetime.now().isoformat())
source: str = "manual" # manual, production_failure, user_report
metadata: dict = field(default_factory=dict)
class EvalDatasetManager:
"""Manage versioned evaluation datasets."""
def __init__(self, dataset_path: Path):
self.dataset_path = dataset_path
self.examples: list[EvalExample] = self._load()
def add_from_production_failure(
self,
input_data: dict,
failed_output: str,
correct_behavior: str,
category: EvalCategory = EvalCategory.REGRESSION,
) -> EvalExample:
"""Add a new eval example from a production failure."""
example = EvalExample(
id=f"prod_fail_{len(self.examples):04d}",
input_data=input_data,
evaluation_criteria=[
f"Must NOT produce: {failed_output[:200]}",
f"Should satisfy: {correct_behavior}",
],
difficulty=Difficulty.COMPLEX,
category=category,
source="production_failure",
metadata={"failed_output_sample": failed_output[:500]},
)
self.examples.append(example)
self._save()
return example
def get_by_category(self, category: EvalCategory) -> list[EvalExample]:
return [e for e in self.examples if e.category == category]
def get_by_difficulty(self, difficulty: Difficulty) -> list[EvalExample]:
return [e for e in self.examples if e.difficulty == difficulty]
def stats(self) -> dict:
return {
"total": len(self.examples),
"by_category": {
cat.value: len(self.get_by_category(cat))
for cat in EvalCategory
},
"by_difficulty": {
diff.value: len(self.get_by_difficulty(diff))
for diff in Difficulty
},
}
def _load(self) -> list[EvalExample]:
if not self.dataset_path.exists():
return []
with open(self.dataset_path) as f:
return [EvalExample(**item) for item in json.load(f)]
def _save(self) -> None:
with open(self.dataset_path, "w") as f:
json.dump(
[vars(e) for e in self.examples],
f,
indent=2,
default=str,
)
Component 2: Multi-dimensional scoring
Quality isn't one number. Evaluate across multiple dimensions, each with its own scoring method:
import Anthropic from "@anthropic-ai/sdk";
interface DimensionScore {
dimension: string;
score: number; // 0.0 to 1.0
method: "deterministic" | "llm_judge" | "hybrid";
details: string;
}
interface EvalResult {
exampleId: string;
modelVersion: string;
scores: DimensionScore[];
overallScore: number;
latencyMs: number;
tokenUsage: { input: number; output: number };
timestamp: string;
}
class MultiDimensionalEvaluator {
private client: Anthropic;
private judgeModel = "claude-sonnet-4-20250514";
constructor(client: Anthropic) {
this.client = client;
}
async evaluate(
modelOutput: string,
example: EvalExample,
dimensions: EvalDimension[]
): Promise<DimensionScore[]> {
const scores: DimensionScore[] = [];
for (const dimension of dimensions) {
let score: DimensionScore;
switch (dimension.method) {
case "deterministic":
score = this.evalDeterministic(modelOutput, dimension);
break;
case "llm_judge":
score = await this.evalWithJudge(modelOutput, example, dimension);
break;
case "hybrid":
const det = this.evalDeterministic(modelOutput, dimension);
const judge = await this.evalWithJudge(
modelOutput,
example,
dimension
);
score = {
dimension: dimension.name,
score: det.score * 0.4 + judge.score * 0.6,
method: "hybrid",
details: `Deterministic: ${det.score}, Judge: ${judge.score}`,
};
break;
}
scores.push(score);
}
return scores;
}
private evalDeterministic(
output: string,
dimension: EvalDimension
): DimensionScore {
let score = 1.0;
const checks: string[] = [];
// Format validation
if (dimension.expectedFormat === "json") {
try {
JSON.parse(output);
checks.push("Valid JSON: PASS");
} catch {
score = 0;
checks.push("Valid JSON: FAIL");
}
}
// Length constraints
if (dimension.maxLength && output.length > dimension.maxLength) {
score *= 0.5;
checks.push(`Length ${output.length} > max ${dimension.maxLength}`);
}
// Required content
if (dimension.mustContain) {
for (const term of dimension.mustContain) {
if (!output.toLowerCase().includes(term.toLowerCase())) {
score *= 0.8;
checks.push(`Missing required: "${term}"`);
}
}
}
return {
dimension: dimension.name,
score,
method: "deterministic",
details: checks.join("; "),
};
}
private async evalWithJudge(
output: string,
example: EvalExample,
dimension: EvalDimension
): Promise<DimensionScore> {
const judgePrompt = `You are evaluating an AI model's output on the dimension: ${dimension.name}
Evaluation criteria: ${dimension.criteria}
User input that was given to the model:
${JSON.stringify(example.inputData)}
Model's output:
${output}
Rate the output on a scale of 0.0 to 1.0 for the dimension "${dimension.name}".
Respond with ONLY a JSON object: {"score": <number>, "reasoning": "<brief explanation>"}`;
const response = await this.client.messages.create({
model: this.judgeModel,
max_tokens: 256,
messages: [{ role: "user", content: judgePrompt }],
});
const text =
response.content[0].type === "text" ? response.content[0].text : "";
const parsed = JSON.parse(text);
return {
dimension: dimension.name,
score: parsed.score,
method: "llm_judge",
details: parsed.reasoning,
};
}
}
Component 3: Regression detection
The most valuable part of the framework: automatically detecting when quality degrades.
| Signal | Detection method | Alert threshold |
|---|---|---|
| Overall score drop | Compare running avg to baseline | >5% drop sustained 24h |
| Dimension-specific regression | Per-dimension tracking | >10% drop on any dimension |
| New failure patterns | Cluster analysis on failures | >3 similar failures in 24h |
| Latency regression | P95 tracking | >50% increase from baseline |
| Cost regression | Per-request cost tracking | >20% increase |
Run evaluations on two triggers:
- On model/prompt changes — full eval suite before deploying any change
- Continuous sampling — evaluate 5% of production traffic nightly to catch drift
Model comparison framework
When Anthropic releases a new model version or you consider switching between Sonnet and Haiku, run a structured comparison:
@dataclass
class ModelComparison:
model_a: str
model_b: str
dataset_size: int
results: dict
def summary(self) -> str:
wins_a = sum(
1 for r in self.results["comparisons"]
if r["winner"] == self.model_a
)
wins_b = sum(
1 for r in self.results["comparisons"]
if r["winner"] == self.model_b
)
ties = self.dataset_size - wins_a - wins_b
return (
f"Model comparison ({self.dataset_size} examples):\n"
f" {self.model_a}: {wins_a} wins ({wins_a/self.dataset_size:.1%})\n"
f" {self.model_b}: {wins_b} wins ({wins_b/self.dataset_size:.1%})\n"
f" Ties: {ties} ({ties/self.dataset_size:.1%})\n"
f" Cost ratio: {self.results['cost_ratio']:.2f}x\n"
f" Latency ratio: {self.results['latency_ratio']:.2f}x"
)
Production evaluation results
Our evaluation framework across three production features (200 eval examples each):
| Feature | Baseline score | Current score | Regressions caught | False alerts |
|---|---|---|---|---|
| Customer support | 0.89 | 0.92 | 4 | 1 |
| Document extraction | 0.91 | 0.93 | 2 | 0 |
| Code review assistant | 0.86 | 0.88 | 6 | 2 |
Every regression caught was a real quality degradation — either from a prompt change or a model update. Without the framework, these would have been discovered weeks later via user complaints.
Key takeaways
- Evaluation is infrastructure, not a one-time task. Build it once, run it continuously, grow it from production failures.
- Multi-dimensional scoring reveals what single scores hide. A feature can score 0.9 overall while being 0.6 on a critical dimension.
- LLM-as-judge plus deterministic checks is the winning combination. Neither alone is sufficient for production evaluation.
- Every production failure becomes a test case. Your eval dataset should grow faster from production than from manual curation.
- Run evals on two cadences. Every change (pre-deploy) and continuous sampling (post-deploy). Both are required.
The teams that ship AI features confidently are the ones that measure quality systematically. Everything else is hope-driven development.
Recommended reading

The State of Agentic AI in 2026: Capabilities, Limitations, and Production Readiness
Comprehensive analysis of agentic AI in 2026 covering production capabilities, current limitations, and enterprise readiness benchmarks with real deployment data.

Observability for AI Agents: Tracing Multi-Step Reasoning Chains in Production
How to implement production observability for AI agents including distributed tracing, reasoning chain analysis, and debugging multi-step failures.

Measuring and Reducing AI Workload Carbon Emissions: A Practical Engineering Guide
Building a carbon-aware scheduling system for ML training and inference workloads that reduced our AI infrastructure emissions by 42% while maintaining SLA commitments.

Comments
No comments yet. Be the first to share your thoughts.