AI Code Review in CI/CD: Catching Architectural Violations Before They Ship

How we integrated Claude into our CI/CD pipeline to catch architectural violations, security issues, and performance anti-patterns, blocking 340 problematic PRs in 6 months.

#claude#code-review#cicd#automation
Cover image for the article: AI Code Review in CI/CD: Catching Architectural Violations Before They Ship

Static analysis catches syntax errors. Linters catch style violations. But neither catches "this service now depends on a module it shouldn't" or "this query will cause N+1 in production" or "this pattern violates our event sourcing invariants." We integrated Claude into our CI/CD pipeline as an architectural reviewer — a system that understands our architecture decisions and enforces them at PR time.

In 6 months, it blocked 340 PRs that would have introduced architectural violations, caught 47 security issues that passed our SAST tools, and reduced architecture review bottlenecks by 62%.

The Problem

With 40 engineers across 6 teams, architectural consistency was eroding:

  • Senior architects couldn't review every PR (they reviewed ~30% of merged code)
  • New engineers unknowingly violated boundaries (3-4 boundary violations per week)
  • Performance anti-patterns slipped through (discovered weeks later in production)
  • Security patterns were inconsistently applied (especially in newer services)

We needed automated enforcement that understood intent, not just syntax.

Architecture

The system integrates as a GitHub Actions step that runs on every PR targeting main branches.

CI/CD Code Review Architecture

Pipeline Integration

import Anthropic from '@anthropic-ai/sdk';
import { Octokit } from '@octokit/rest';

interface ReviewResult {
  status: 'approved' | 'changes_requested' | 'comment';
  violations: Violation[];
  suggestions: Suggestion[];
  securityFindings: SecurityFinding[];
  summary: string;
}

interface Violation {
  severity: 'critical' | 'high' | 'medium' | 'low';
  category: 'architecture' | 'security' | 'performance' | 'reliability';
  file: string;
  line: number;
  description: string;
  rule: string;
  suggestion: string;
}

class CICDReviewPipeline {
  private anthropic: Anthropic;
  private octokit: Octokit;
  private architectureRules: string;

  constructor() {
    this.anthropic = new Anthropic();
    this.octokit = new Octokit({ auth: process.env.GITHUB_TOKEN });
    this.architectureRules = this.loadArchitectureRules();
  }

  async reviewPullRequest(prNumber: number, repo: string): Promise<ReviewResult> {
    // Fetch PR diff and context
    const diff = await this.fetchPRDiff(prNumber, repo);
    const changedFiles = await this.fetchChangedFiles(prNumber, repo);
    const prDescription = await this.fetchPRDescription(prNumber, repo);

    // Fetch architectural context for affected modules
    const moduleContext = await this.getModuleContext(changedFiles);

    // Run the review
    const response = await this.anthropic.messages.create({
      model: 'claude-sonnet-4-20250514',
      max_tokens: 8192,
      system: `You are a senior software architect reviewing code changes.
Your job is to enforce architectural decisions, catch security issues,
and identify performance problems that static analysis tools miss.

## Architecture Rules
${this.architectureRules}

## Module Boundaries
${moduleContext.boundaries}

## Known Anti-Patterns
${moduleContext.antiPatterns}

Review the PR diff and report violations. Be specific about line numbers
and provide actionable fix suggestions. Do NOT flag style issues —
only architectural, security, and performance concerns.`,
      messages: [{
        role: 'user',
        content: `## PR Description
${prDescription}

## Changed Files
${changedFiles.map(f => f.filename).join('\n')}

## Diff
${diff}

Analyze this PR for:
1. Architecture boundary violations
2. Security vulnerabilities (beyond what SAST catches)
3. Performance anti-patterns
4. Reliability concerns (missing error handling, retry logic, etc.)

Return JSON:
{
  "status": "approved|changes_requested|comment",
  "violations": [...],
  "suggestions": [...],
  "securityFindings": [...],
  "summary": "brief overview"
}`
      }]
    });

    return JSON.parse(response.content[0].text);
  }

  private loadArchitectureRules(): string {
    // Load from repository's ARCHITECTURE.md or .architecture/ directory
    return `
1. Services in /services/payments/ must NOT import from /services/users/ directly — use events
2. All database queries must go through repository classes — no direct ORM usage in handlers
3. External API calls must use the circuit breaker wrapper from @internal/resilience
4. PII fields must use the @Encrypted decorator — never store plaintext
5. Event handlers must be idempotent — check for duplicate processing
6. GraphQL resolvers must not make more than 3 database calls (use DataLoader)
7. Background jobs must have explicit timeout configuration
8. All new endpoints require rate limiting middleware`;
  }
}

Contextual Review with Dependency Analysis

The reviewer doesn't just look at the diff — it understands the broader context of what modules are affected and what rules apply.

import anthropic
from pathlib import Path
import subprocess

class ArchitecturalContextBuilder:
    def __init__(self, repo_root: str):
        self.repo_root = Path(repo_root)
        self.client = anthropic.Anthropic()

    def build_context(self, changed_files: list[str]) -> dict:
        """Build architectural context for the review."""
        
        # Identify affected modules
        modules = set()
        for file_path in changed_files:
            module = self._identify_module(file_path)
            modules.add(module)

        # Load module-specific architecture docs
        module_docs = {}
        for module in modules:
            doc_path = self.repo_root / module / "ARCHITECTURE.md"
            if doc_path.exists():
                module_docs[module] = doc_path.read_text()

        # Analyze dependency changes
        dep_changes = self._analyze_dependency_changes(changed_files)

        # Check for new cross-module imports
        boundary_violations = self._check_boundaries(changed_files, modules)

        return {
            "modules": list(modules),
            "module_docs": module_docs,
            "dependency_changes": dep_changes,
            "potential_violations": boundary_violations,
            "affected_apis": self._find_affected_apis(changed_files)
        }

    def _analyze_dependency_changes(self, changed_files: list[str]) -> list[dict]:
        """Detect new imports that cross module boundaries."""
        violations = []
        
        for file_path in changed_files:
            full_path = self.repo_root / file_path
            if not full_path.exists():
                continue

            # Get the diff for this file specifically
            diff = subprocess.run(
                ["git", "diff", "main", "--", file_path],
                capture_output=True, text=True, cwd=self.repo_root
            ).stdout

            # Extract added import lines
            added_imports = [
                line[1:].strip()
                for line in diff.split('\n')
                if line.startswith('+') and ('import' in line or 'require' in line)
                and not line.startswith('+++')
            ]

            for imp in added_imports:
                source_module = self._identify_module(file_path)
                target_module = self._resolve_import_module(imp)
                if target_module and source_module != target_module:
                    violations.append({
                        "file": file_path,
                        "import": imp,
                        "from_module": source_module,
                        "to_module": target_module
                    })

        return violations

GitHub Integration and Feedback

Reviews are posted as PR comments with inline annotations on specific lines.

class GitHubReviewPoster {
  private octokit: Octokit;

  constructor(token: string) {
    this.octokit = new Octokit({ auth: token });
  }

  async postReview(
    owner: string,
    repo: string,
    prNumber: number,
    result: ReviewResult
  ): Promise<void> {
    // Create inline comments for each violation
    const comments = result.violations.map(v => ({
      path: v.file,
      line: v.line,
      body: `**${v.severity.toUpperCase()}** — ${v.category}\n\n${v.description}\n\n**Rule:** ${v.rule}\n\n**Suggested fix:**\n${v.suggestion}`
    }));

    // Submit the review
    await this.octokit.pulls.createReview({
      owner,
      repo,
      pull_number: prNumber,
      event: result.status === 'approved' ? 'APPROVE' :
             result.status === 'changes_requested' ? 'REQUEST_CHANGES' : 'COMMENT',
      body: this.formatSummary(result),
      comments
    });
  }

  private formatSummary(result: ReviewResult): string {
    const criticalCount = result.violations.filter(v => v.severity === 'critical').length;
    const highCount = result.violations.filter(v => v.severity === 'high').length;

    return `## AI Architecture Review

${result.summary}

| Severity | Count |
|----------|-------|
| Critical | ${criticalCount} |
| High | ${highCount} |
| Security | ${result.securityFindings.length} |

${criticalCount > 0 ? '⛔ **Blocking:** Critical violations must be resolved before merge.' : ''}
${highCount > 0 ? '⚠️ **Warning:** High-severity issues should be addressed.' : ''}
${criticalCount === 0 && highCount === 0 ? '✅ No blocking issues found.' : ''}`;
  }
}

What It Catches That Other Tools Miss

Real examples from our first 6 months:

  1. N+1 query patterns — A GraphQL resolver fetching user profiles in a loop instead of batching (SAST can't detect this)
  2. Event sourcing violations — Direct database mutations in a module that should only emit events
  3. Secret exposure — API keys constructed from environment variables in a way that logs them on error
  4. Retry storms — A new retry wrapper that doesn't implement exponential backoff or jitter
  5. Circular dependencies — Subtle import cycles that TypeScript allows but cause runtime issues

Benchmarks

MetricValue
PRs reviewed (6 months)4,218
Violations caught340 (8.1% of PRs)
False positive rate12%
Avg. review latency47 seconds
Security issues caught (missed by SAST)47
Architecture review bottleneck reduction62%
Monthly API cost$890

The 12% false positive rate is acceptable because violations are suggestions, not hard blocks (except critical severity). Engineers can dismiss with a comment explaining their reasoning.

Tuning False Positives

We maintain a .ai-review-overrides.yml that captures approved exceptions:

overrides:
  - rule: "no-cross-module-import"
    file: "services/billing/legacy-adapter.ts"
    reason: "Legacy adapter intentionally bridges billing and users during migration"
    approved_by: "@architect-team"
    expires: "2026-09-01"

The review system consults this file before flagging, reducing false positives over time.

Conclusion

AI code review in CI/CD isn't about replacing human reviewers — it's about catching the architectural and security violations that humans consistently miss under time pressure. The system pays for itself by preventing production incidents that would cost orders of magnitude more to fix post-deployment. Start by documenting your architecture rules explicitly, then let Claude enforce them consistently on every PR.

Comments

    No comments yet. Be the first to share your thoughts.