AI Code Review in CI/CD: Catching Architectural Violations Before They Ship
How we integrated Claude into our CI/CD pipeline to catch architectural violations, security issues, and performance anti-patterns, blocking 340 problematic PRs in 6 months.

Static analysis catches syntax errors. Linters catch style violations. But neither catches "this service now depends on a module it shouldn't" or "this query will cause N+1 in production" or "this pattern violates our event sourcing invariants." We integrated Claude into our CI/CD pipeline as an architectural reviewer — a system that understands our architecture decisions and enforces them at PR time.
In 6 months, it blocked 340 PRs that would have introduced architectural violations, caught 47 security issues that passed our SAST tools, and reduced architecture review bottlenecks by 62%.
The Problem
With 40 engineers across 6 teams, architectural consistency was eroding:
- Senior architects couldn't review every PR (they reviewed ~30% of merged code)
- New engineers unknowingly violated boundaries (3-4 boundary violations per week)
- Performance anti-patterns slipped through (discovered weeks later in production)
- Security patterns were inconsistently applied (especially in newer services)
We needed automated enforcement that understood intent, not just syntax.
Architecture
The system integrates as a GitHub Actions step that runs on every PR targeting main branches.
Pipeline Integration
import Anthropic from '@anthropic-ai/sdk';
import { Octokit } from '@octokit/rest';
interface ReviewResult {
status: 'approved' | 'changes_requested' | 'comment';
violations: Violation[];
suggestions: Suggestion[];
securityFindings: SecurityFinding[];
summary: string;
}
interface Violation {
severity: 'critical' | 'high' | 'medium' | 'low';
category: 'architecture' | 'security' | 'performance' | 'reliability';
file: string;
line: number;
description: string;
rule: string;
suggestion: string;
}
class CICDReviewPipeline {
private anthropic: Anthropic;
private octokit: Octokit;
private architectureRules: string;
constructor() {
this.anthropic = new Anthropic();
this.octokit = new Octokit({ auth: process.env.GITHUB_TOKEN });
this.architectureRules = this.loadArchitectureRules();
}
async reviewPullRequest(prNumber: number, repo: string): Promise<ReviewResult> {
// Fetch PR diff and context
const diff = await this.fetchPRDiff(prNumber, repo);
const changedFiles = await this.fetchChangedFiles(prNumber, repo);
const prDescription = await this.fetchPRDescription(prNumber, repo);
// Fetch architectural context for affected modules
const moduleContext = await this.getModuleContext(changedFiles);
// Run the review
const response = await this.anthropic.messages.create({
model: 'claude-sonnet-4-20250514',
max_tokens: 8192,
system: `You are a senior software architect reviewing code changes.
Your job is to enforce architectural decisions, catch security issues,
and identify performance problems that static analysis tools miss.
## Architecture Rules
${this.architectureRules}
## Module Boundaries
${moduleContext.boundaries}
## Known Anti-Patterns
${moduleContext.antiPatterns}
Review the PR diff and report violations. Be specific about line numbers
and provide actionable fix suggestions. Do NOT flag style issues —
only architectural, security, and performance concerns.`,
messages: [{
role: 'user',
content: `## PR Description
${prDescription}
## Changed Files
${changedFiles.map(f => f.filename).join('\n')}
## Diff
${diff}
Analyze this PR for:
1. Architecture boundary violations
2. Security vulnerabilities (beyond what SAST catches)
3. Performance anti-patterns
4. Reliability concerns (missing error handling, retry logic, etc.)
Return JSON:
{
"status": "approved|changes_requested|comment",
"violations": [...],
"suggestions": [...],
"securityFindings": [...],
"summary": "brief overview"
}`
}]
});
return JSON.parse(response.content[0].text);
}
private loadArchitectureRules(): string {
// Load from repository's ARCHITECTURE.md or .architecture/ directory
return `
1. Services in /services/payments/ must NOT import from /services/users/ directly — use events
2. All database queries must go through repository classes — no direct ORM usage in handlers
3. External API calls must use the circuit breaker wrapper from @internal/resilience
4. PII fields must use the @Encrypted decorator — never store plaintext
5. Event handlers must be idempotent — check for duplicate processing
6. GraphQL resolvers must not make more than 3 database calls (use DataLoader)
7. Background jobs must have explicit timeout configuration
8. All new endpoints require rate limiting middleware`;
}
}
Contextual Review with Dependency Analysis
The reviewer doesn't just look at the diff — it understands the broader context of what modules are affected and what rules apply.
import anthropic
from pathlib import Path
import subprocess
class ArchitecturalContextBuilder:
def __init__(self, repo_root: str):
self.repo_root = Path(repo_root)
self.client = anthropic.Anthropic()
def build_context(self, changed_files: list[str]) -> dict:
"""Build architectural context for the review."""
# Identify affected modules
modules = set()
for file_path in changed_files:
module = self._identify_module(file_path)
modules.add(module)
# Load module-specific architecture docs
module_docs = {}
for module in modules:
doc_path = self.repo_root / module / "ARCHITECTURE.md"
if doc_path.exists():
module_docs[module] = doc_path.read_text()
# Analyze dependency changes
dep_changes = self._analyze_dependency_changes(changed_files)
# Check for new cross-module imports
boundary_violations = self._check_boundaries(changed_files, modules)
return {
"modules": list(modules),
"module_docs": module_docs,
"dependency_changes": dep_changes,
"potential_violations": boundary_violations,
"affected_apis": self._find_affected_apis(changed_files)
}
def _analyze_dependency_changes(self, changed_files: list[str]) -> list[dict]:
"""Detect new imports that cross module boundaries."""
violations = []
for file_path in changed_files:
full_path = self.repo_root / file_path
if not full_path.exists():
continue
# Get the diff for this file specifically
diff = subprocess.run(
["git", "diff", "main", "--", file_path],
capture_output=True, text=True, cwd=self.repo_root
).stdout
# Extract added import lines
added_imports = [
line[1:].strip()
for line in diff.split('\n')
if line.startswith('+') and ('import' in line or 'require' in line)
and not line.startswith('+++')
]
for imp in added_imports:
source_module = self._identify_module(file_path)
target_module = self._resolve_import_module(imp)
if target_module and source_module != target_module:
violations.append({
"file": file_path,
"import": imp,
"from_module": source_module,
"to_module": target_module
})
return violations
GitHub Integration and Feedback
Reviews are posted as PR comments with inline annotations on specific lines.
class GitHubReviewPoster {
private octokit: Octokit;
constructor(token: string) {
this.octokit = new Octokit({ auth: token });
}
async postReview(
owner: string,
repo: string,
prNumber: number,
result: ReviewResult
): Promise<void> {
// Create inline comments for each violation
const comments = result.violations.map(v => ({
path: v.file,
line: v.line,
body: `**${v.severity.toUpperCase()}** — ${v.category}\n\n${v.description}\n\n**Rule:** ${v.rule}\n\n**Suggested fix:**\n${v.suggestion}`
}));
// Submit the review
await this.octokit.pulls.createReview({
owner,
repo,
pull_number: prNumber,
event: result.status === 'approved' ? 'APPROVE' :
result.status === 'changes_requested' ? 'REQUEST_CHANGES' : 'COMMENT',
body: this.formatSummary(result),
comments
});
}
private formatSummary(result: ReviewResult): string {
const criticalCount = result.violations.filter(v => v.severity === 'critical').length;
const highCount = result.violations.filter(v => v.severity === 'high').length;
return `## AI Architecture Review
${result.summary}
| Severity | Count |
|----------|-------|
| Critical | ${criticalCount} |
| High | ${highCount} |
| Security | ${result.securityFindings.length} |
${criticalCount > 0 ? '⛔ **Blocking:** Critical violations must be resolved before merge.' : ''}
${highCount > 0 ? '⚠️ **Warning:** High-severity issues should be addressed.' : ''}
${criticalCount === 0 && highCount === 0 ? '✅ No blocking issues found.' : ''}`;
}
}
What It Catches That Other Tools Miss
Real examples from our first 6 months:
- N+1 query patterns — A GraphQL resolver fetching user profiles in a loop instead of batching (SAST can't detect this)
- Event sourcing violations — Direct database mutations in a module that should only emit events
- Secret exposure — API keys constructed from environment variables in a way that logs them on error
- Retry storms — A new retry wrapper that doesn't implement exponential backoff or jitter
- Circular dependencies — Subtle import cycles that TypeScript allows but cause runtime issues
Benchmarks
| Metric | Value |
|---|---|
| PRs reviewed (6 months) | 4,218 |
| Violations caught | 340 (8.1% of PRs) |
| False positive rate | 12% |
| Avg. review latency | 47 seconds |
| Security issues caught (missed by SAST) | 47 |
| Architecture review bottleneck reduction | 62% |
| Monthly API cost | $890 |
The 12% false positive rate is acceptable because violations are suggestions, not hard blocks (except critical severity). Engineers can dismiss with a comment explaining their reasoning.
Tuning False Positives
We maintain a .ai-review-overrides.yml that captures approved exceptions:
overrides:
- rule: "no-cross-module-import"
file: "services/billing/legacy-adapter.ts"
reason: "Legacy adapter intentionally bridges billing and users during migration"
approved_by: "@architect-team"
expires: "2026-09-01"
The review system consults this file before flagging, reducing false positives over time.
Conclusion
AI code review in CI/CD isn't about replacing human reviewers — it's about catching the architectural and security violations that humans consistently miss under time pressure. The system pays for itself by preventing production incidents that would cost orders of magnitude more to fix post-deployment. Start by documenting your architecture rules explicitly, then let Claude enforce them consistently on every PR.
Recommended reading

The State of Agentic AI in 2026: Capabilities, Limitations, and Production Readiness
Comprehensive analysis of agentic AI in 2026 covering production capabilities, current limitations, and enterprise readiness benchmarks with real deployment data.

Observability for AI Agents: Tracing Multi-Step Reasoning Chains in Production
How to implement production observability for AI agents including distributed tracing, reasoning chain analysis, and debugging multi-step failures.

Measuring and Reducing AI Workload Carbon Emissions: A Practical Engineering Guide
Building a carbon-aware scheduling system for ML training and inference workloads that reduced our AI infrastructure emissions by 42% while maintaining SLA commitments.

Comments
No comments yet. Be the first to share your thoughts.