Systematic Prompt Engineering for Claude in Production

A data-driven methodology for prompt engineering with measurable quality metrics, regression testing, and continuous improvement loops.

#claude#prompt-engineering#ai#methodology
Cover image for the article: Systematic Prompt Engineering for Claude in Production

Prompt engineering is not art — it's software engineering with a different compilation target. After shipping dozens of Claude-powered features, I've developed a systematic approach that treats prompts as testable, measurable, and improvable artifacts. No vibes. No "it feels better." Numbers.

The problem with ad-hoc prompting

Most teams iterate on prompts by running a few examples manually, eyeballing the output, and shipping. This works until it doesn't — which is usually three weeks after launch when a user hits an edge case that makes the model produce garbage.

The failure mode is always the same: someone edits a prompt to fix Case A, unknowingly breaks Cases B through F, and nobody notices until users complain. This is the exact same problem we solved in traditional software with automated testing decades ago.

The systematic prompt development lifecycle

Prompt Engineering Lifecycle

Phase 1: Define success criteria before writing a single word

Every prompt needs a specification. What does "good output" look like? Define it in terms that a machine can evaluate:

  • Structural correctness — Does the response match the expected format (JSON schema, specific sections, length constraints)?
  • Factual accuracy — Does the response contain only information from the provided context?
  • Behavioral compliance — Does it follow the persona/tone/restriction rules?
  • Task completion — Does it actually solve the user's problem?
from dataclasses import dataclass, field
from enum import Enum

class EvalDimension(Enum):
    STRUCTURE = "structure"
    ACCURACY = "accuracy"
    COMPLIANCE = "compliance"
    TASK_COMPLETION = "task_completion"

@dataclass
class PromptSpec:
    name: str
    version: str
    description: str
    dimensions: list[EvalDimension]
    golden_examples: list[dict] = field(default_factory=list)
    failure_cases: list[dict] = field(default_factory=list)
    minimum_scores: dict[str, float] = field(default_factory=lambda: {
        "structure": 0.95,
        "accuracy": 0.90,
        "compliance": 0.98,
        "task_completion": 0.85,
    })

# Example specification for a document summarizer
summarizer_spec = PromptSpec(
    name="document_summarizer",
    version="1.0.0",
    description="Summarizes technical documents into structured briefs",
    dimensions=[
        EvalDimension.STRUCTURE,
        EvalDimension.ACCURACY,
        EvalDimension.TASK_COMPLETION,
    ],
    minimum_scores={
        "structure": 0.95,
        "accuracy": 0.92,
        "task_completion": 0.88,
    },
)

Phase 2: Build a golden evaluation dataset

Your eval dataset is your most valuable asset. It needs:

  • 50-200 examples minimum — fewer gives you noisy signal
  • Representative distribution — include easy cases, hard cases, edge cases, adversarial cases
  • Human-annotated expected outputs — not generated by another LLM
  • Versioned and immutable — never edit existing examples, only add new ones

I store these as JSONL files in the repository:

{"id": "sum_001", "input": {"document": "...", "audience": "executive"}, "expected": {"format": "bullet_points", "max_items": 5, "must_include": ["revenue_impact", "timeline"]}, "difficulty": "standard"}
{"id": "sum_002", "input": {"document": "...", "audience": "technical"}, "expected": {"format": "structured_sections", "must_include": ["architecture", "tradeoffs"]}, "difficulty": "complex"}

Phase 3: Implement automated evaluation

Evaluation is a mix of deterministic checks and LLM-as-judge scoring. Use both:

import Anthropic from "@anthropic-ai/sdk";

interface EvalResult {
  exampleId: string;
  scores: Record<string, number>;
  passed: boolean;
  failures: string[];
}

async function evaluatePrompt(
  client: Anthropic,
  prompt: string,
  dataset: EvalExample[]
): Promise<EvalResult[]> {
  const results: EvalResult[] = [];

  for (const example of dataset) {
    const response = await client.messages.create({
      model: "claude-sonnet-4-20250514",
      max_tokens: 4096,
      system: prompt,
      messages: [{ role: "user", content: example.input }],
    });

    const output = response.content[0].type === "text"
      ? response.content[0].text
      : "";

    // Deterministic checks
    const structureScore = checkStructure(output, example.expectedFormat);

    // LLM-as-judge for semantic quality
    const qualityScore = await judgeQuality(client, {
      input: example.input,
      output,
      criteria: example.evaluationCriteria,
    });

    const scores = { structure: structureScore, quality: qualityScore };
    const passed = Object.entries(scores).every(
      ([dim, score]) => score >= MINIMUM_SCORES[dim]
    );

    results.push({
      exampleId: example.id,
      scores,
      passed,
      failures: passed ? [] : identifyFailures(scores, MINIMUM_SCORES),
    });
  }

  return results;
}

Phase 4: Iterate with version control

Every prompt change gets:

  1. A git branch
  2. A full eval run against the golden dataset
  3. A comparison against the previous version's scores
  4. A review where the team looks at the delta

This is CI/CD for prompts. We run this in GitHub Actions:

# .github/workflows/prompt-eval.yml
name: Prompt Evaluation
on:
  pull_request:
    paths: ['prompts/**']
jobs:
  evaluate:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - run: npm ci
      - run: npm run eval:prompts
      - run: npm run eval:compare -- --baseline=main

Measurable improvements from this approach

After adopting this methodology across four production prompts:

MetricBefore (ad-hoc)After (systematic)
Prompt-related incidents/month4-60-1
Time to iterate on prompt2-3 days4-6 hours
Regression rate on changes~30%<5%
Eval dataset size (avg)10 examples150+ examples
Confidence in shipping changesLowHigh

Advanced techniques that compound

Prompt decomposition

Complex prompts are like complex functions — break them down. Instead of one 2000-token system prompt trying to do everything, chain smaller focused prompts:

  1. Classification prompt — determine what type of request this is
  2. Extraction prompt — pull relevant information from context
  3. Generation prompt — produce the final output
  4. Validation prompt — verify the output meets constraints

Each piece is independently testable and improvable.

A/B testing prompts in production

Run two prompt versions simultaneously with traffic splitting. Compare on real user signals: thumbs up/down, task completion, support tickets. Lab evals tell you about quality; production A/B tells you about user satisfaction.

Prompt regression alerts

Set up monitoring that runs your eval suite nightly against the production prompt with the latest model version. Model updates from Anthropic can subtly change behavior — you want to know before your users do.

Key takeaways

  1. Define success numerically before you start. If you can't measure it, you can't improve it.
  2. Your eval dataset is your most important asset. Invest in it like you invest in your test suite.
  3. Every prompt change goes through CI. No exceptions, no "quick fixes" in production.
  4. Decompose complex prompts into testable, composable units.
  5. Monitor continuously. Model updates and data drift will erode quality silently.

Systematic prompt engineering takes more upfront investment than ad-hoc iteration. But the compound returns — fewer incidents, faster iteration, higher confidence — make it the only sane approach for production systems.

Comments

    No comments yet. Be the first to share your thoughts.