Measuring AI-Generated Code Quality in Production

A metrics framework for evaluating AI code generation — functional correctness, maintainability, security, and long-term impact on engineering velocity.

#ai#code-generation#quality#metrics
Cover image for the article: Measuring AI-Generated Code Quality in Production

Every engineering team is adopting AI code generation. Few are measuring whether it's actually improving outcomes. "Lines of code generated" is a vanity metric. Here's a framework for measuring what matters: correctness, maintainability, security, and velocity impact.

The measurement problem

AI code generation creates a paradox: it increases code output velocity while potentially decreasing code quality. Without measurement, you can't tell whether your team is shipping faster or accumulating technical debt faster.

The metrics that matter aren't about the AI tool — they're about the code it produces and how that code performs in production.

The quality framework

AI Code Quality Metrics Framework

Four dimensions of AI-generated code quality:

  1. Functional correctness — Does it work as intended?
  2. Maintainability — Can humans understand and modify it?
  3. Security — Does it introduce vulnerabilities?
  4. Velocity impact — Net effect on team productivity?

Dimension 1: Functional correctness

Automated metrics

from dataclasses import dataclass
from typing import Literal
import subprocess
import json

@dataclass
class CorrectnessMetrics:
    test_pass_rate: float          # % of generated code that passes tests
    first_attempt_pass: float      # % passing without human edits
    bug_introduction_rate: float   # Bugs found within 7 days per 1K lines
    type_error_rate: float         # Type check failures per generation
    runtime_error_rate: float      # Exceptions in first 24h of deployment

class CorrectnessEvaluator:
    def __init__(self, repo_path: str):
        self.repo_path = repo_path
        self.results: list[dict] = []

    async def evaluate_generation(
        self,
        generated_code: str,
        test_file: str,
        language: Literal["python", "typescript"],
    ) -> CorrectnessMetrics:
        """Evaluate a code generation against its test suite."""

        # Step 1: Static analysis
        type_errors = await self._run_type_check(generated_code, language)

        # Step 2: Run test suite
        test_result = await self._run_tests(test_file)

        # Step 3: Check for common AI generation errors
        ai_errors = self._detect_ai_antipatterns(generated_code)

        return CorrectnessMetrics(
            test_pass_rate=test_result["pass_rate"],
            first_attempt_pass=1.0 if test_result["pass_rate"] == 1.0 else 0.0,
            bug_introduction_rate=len(ai_errors) / max(1, self._count_lines(generated_code)) * 1000,
            type_error_rate=len(type_errors),
            runtime_error_rate=0.0,  # Measured post-deployment
        )

    def _detect_ai_antipatterns(self, code: str) -> list[str]:
        """Detect common AI code generation mistakes."""
        issues = []

        # Hallucinated imports (modules that don't exist)
        # Incomplete error handling
        # Placeholder/TODO code passed as complete
        # Hardcoded secrets or credentials
        # Missing null checks on optional values

        if "TODO" in code or "FIXME" in code:
            issues.append("Contains placeholder code")
        if "api_key = " in code and "os.environ" not in code:
            issues.append("Possible hardcoded credential")
        if code.count("except:") > 0 or code.count("except Exception:") > 2:
            issues.append("Overly broad exception handling")

        return issues

    async def _run_type_check(self, code: str, language: str) -> list[str]:
        if language == "python":
            result = subprocess.run(
                ["mypy", "--strict", "--no-error-summary", "-"],
                input=code, capture_output=True, text=True
            )
            return result.stdout.strip().split("\n") if result.returncode != 0 else []
        elif language == "typescript":
            result = subprocess.run(
                ["npx", "tsc", "--noEmit", "--strict"],
                capture_output=True, text=True
            )
            return result.stdout.strip().split("\n") if result.returncode != 0 else []
        return []

Benchmark results from our team

After 6 months of tracking AI-generated code (3 engineers, ~2000 generations):

MetricAI-GeneratedHuman-WrittenDelta
First-attempt test pass rate68%N/A—
Post-edit test pass rate94%97%-3%
Bugs found within 7 days (per 1K lines)3.22.1+52%
Type errors per generation0.80.3+167%
Runtime errors (first 24h)1.1%0.4%+175%

AI code passes tests at similar rates after human editing, but introduces more subtle bugs that tests miss initially.

Dimension 2: Maintainability

interface MaintainabilityMetrics {
  cyclomaticComplexity: number;
  cognitiveComplexity: number;
  duplicateCodeRatio: number;
  avgFunctionLength: number;
  commentDensity: number;
  namingConsistency: number; // 0-1 score
  modificationFrequency: number; // Times modified within 30 days
}

async function evaluateMaintainability(
  filePath: string,
  isAiGenerated: boolean,
): Promise<MaintainabilityMetrics> {
  // Run static analysis tools
  const eslintResult = await runEslint(filePath);
  const complexityResult = await runComplexityAnalysis(filePath);

  // Git-based metrics
  const gitLog = await getFileHistory(filePath, 30); // 30-day window

  return {
    cyclomaticComplexity: complexityResult.cyclomatic,
    cognitiveComplexity: complexityResult.cognitive,
    duplicateCodeRatio: await detectDuplication(filePath),
    avgFunctionLength: complexityResult.avgFunctionLines,
    commentDensity: calculateCommentRatio(filePath),
    namingConsistency: evaluateNaming(filePath),
    modificationFrequency: gitLog.commitCount,
  };
}

// Track over time: AI-generated files vs human-written
async function compareMaintainability(
  repoPath: string,
): Promise<{ aiGenerated: MaintainabilityMetrics; humanWritten: MaintainabilityMetrics }> {
  const aiFiles = await getFilesTaggedAsAiGenerated(repoPath);
  const humanFiles = await getFilesNotTaggedAsAiGenerated(repoPath);

  const aiMetrics = await Promise.all(aiFiles.map((f) => evaluateMaintainability(f, true)));
  const humanMetrics = await Promise.all(humanFiles.map((f) => evaluateMaintainability(f, false)));

  return {
    aiGenerated: averageMetrics(aiMetrics),
    humanWritten: averageMetrics(humanMetrics),
  };
}

Maintainability findings

MetricAI-GeneratedHuman-Written
Avg cyclomatic complexity8.25.6
Avg function length (lines)3422
Duplicate code ratio12%4%
Modification frequency (30d)3.1 edits1.8 edits
Comment density15%8%

AI generates more comments (often verbose or obvious ones) but more complex functions with more duplication. The higher modification frequency confirms: AI code needs more rework.

Dimension 3: Security

Track security-specific signals in AI-generated code:

Vulnerability TypeAI Rate (per 1K lines)Human RateRisk Level
SQL injection potential0.80.1Critical
Missing input validation2.10.6High
Hardcoded secrets0.30.05Critical
Insecure deserialization0.40.1High
Missing auth checks1.20.4Critical
Verbose error messages1.80.8Medium

AI code is 3-5x more likely to contain security vulnerabilities. The most common: missing input validation and overly permissive error handling that leaks internal state.

Dimension 4: Velocity impact

The ultimate question: is AI code generation making your team faster, net of the quality costs?

@dataclass
class VelocityMetrics:
    """Net velocity impact of AI code generation."""
    generation_time_saved_hours: float  # Time saved by AI generating initial code
    review_time_added_hours: float      # Additional review time for AI code
    bug_fix_time_hours: float           # Time spent fixing AI-introduced bugs
    rework_time_hours: float            # Time rewriting AI code that wasn't good enough
    net_velocity_hours: float           # Net time saved (positive = faster)

    @property
    def roi_multiplier(self) -> float:
        """How many hours saved per hour spent on AI code management."""
        cost = self.review_time_added_hours + self.bug_fix_time_hours + self.rework_time_hours
        return self.generation_time_saved_hours / max(cost, 0.1)

Our team's velocity data (6-month average, per engineer per week)

ActivityHours/Week
Time saved by AI generation8.2
Additional review time1.8
Bug fixes (AI-introduced)1.2
Rework (rewrites)0.9
Net time saved4.3
ROI multiplier2.1x

For every hour spent managing AI-generated code quality, we save 2.1 hours. The net is positive but not as dramatic as "10x productivity" claims suggest.

The tracking system

Tag AI-generated code at the commit level so you can measure longitudinally:

  • Use commit message conventions: [ai-assisted] or [ai-generated]
  • Track in your git metadata which files were AI-generated
  • Build dashboards that compare metrics by generation source
  • Review monthly: is quality trending up or down?

Recommendations

Based on 6 months of data:

  1. Always review AI code for security — the vulnerability introduction rate is unacceptable without review
  2. Use AI for boilerplate, not business logic — correctness rates are highest on repetitive patterns
  3. Set complexity thresholds — reject AI generations with cyclomatic complexity >10
  4. Track net velocity, not gross output — "more code" isn't valuable if it requires proportionally more maintenance
  5. Invest in test coverage first — AI code without tests is a liability; with tests, bugs surface quickly

Key takeaways

  • AI code generation delivers ~2x ROI on engineering time when properly managed
  • Without measurement, you're likely accumulating technical debt faster than you realize
  • Security is the highest-risk dimension — AI code is 3-5x more likely to contain vulnerabilities
  • Maintainability degrades measurably: more complexity, more duplication, more rework
  • The measurement framework itself improves outcomes — teams that track quality generate better AI code over time

Don't ban AI code generation. Measure it, set quality bars, and let the data guide your adoption strategy.

Comments

    No comments yet. Be the first to share your thoughts.