Achieving 94% Test Coverage with AI-Generated Tests
How we used Claude to generate meaningful test suites that increased coverage from 43% to 94% while catching real bugs that manual tests missed.

Our codebase had 43% test coverage. Not because engineers were lazy — they were shipping features faster than they could write tests. The coverage gap was concentrated in the oldest, most complex modules — exactly the code most likely to break. We built a Claude-powered test generation system that analyzed our source code, generated comprehensive test suites, and increased coverage to 94% in 8 weeks. More importantly, the generated tests caught 23 real bugs that had been lurking undetected.
This isn't about generating trivial assertion tests. It's about producing meaningful tests that verify behavior, edge cases, and failure modes — tests a senior engineer would write if they had unlimited time.
The Problem
Our coverage breakdown by module:
- Core business logic: 67% (reasonable, engineers prioritized this)
- API handlers: 52% (tested happy paths, missed edge cases)
- Data access layer: 38% (mocking complexity discouraged testing)
- Utility functions: 71% (easy to test, well-covered)
- Integration points: 19% (hardest to test, most critical)
- Error handling paths: 12% (almost entirely untested)
The 12% on error handling was terrifying. We were shipping code with almost no verification of how it behaved when things went wrong.
Architecture
The test generation pipeline operates in four phases: analysis, generation, validation, and integration.
Source Code Analysis
Before generating tests, we build a comprehensive understanding of what each module does, its dependencies, and its failure modes.
import Anthropic from '@anthropic-ai/sdk';
import * as ts from 'typescript';
interface FunctionAnalysis {
name: string;
filePath: string;
signature: string;
returnType: string;
parameters: { name: string; type: string; optional: boolean }[];
dependencies: string[];
sideEffects: string[];
errorPaths: string[];
branchComplexity: number;
existingTests: string[];
uncoveredBranches: string[];
}
class TestAnalyzer {
private anthropic: Anthropic;
constructor() {
this.anthropic = new Anthropic();
}
async analyzeForTesting(filePath: string, sourceCode: string): Promise<FunctionAnalysis[]> {
// Static analysis for structure
const staticAnalysis = this.performStaticAnalysis(filePath, sourceCode);
// Claude analysis for behavior understanding
const response = await this.anthropic.messages.create({
model: 'claude-sonnet-4-20250514',
max_tokens: 4096,
messages: [{
role: 'user',
content: `Analyze this code for test generation. For each function, identify:
1. All possible execution paths (happy + error)
2. Edge cases based on parameter types and constraints
3. Side effects (DB writes, API calls, file operations)
4. Dependencies that need mocking
5. Invariants that should be verified
6. Boundary conditions
Source code:
\`\`\`typescript
${sourceCode}
\`\`\`
Existing test coverage (what's already tested):
${staticAnalysis.existingTestDescriptions.join('\n')}
Return JSON array of analysis objects per function.`
}]
});
const claudeAnalysis = JSON.parse(response.content[0].text);
return this.mergeAnalyses(staticAnalysis, claudeAnalysis);
}
private performStaticAnalysis(filePath: string, source: string) {
const sourceFile = ts.createSourceFile(
filePath, source, ts.ScriptTarget.Latest, true
);
const functions: any[] = [];
const visitor = (node: ts.Node) => {
if (ts.isFunctionDeclaration(node) || ts.isMethodDeclaration(node)) {
functions.push({
name: node.name?.getText(sourceFile) || 'anonymous',
branchCount: this.countBranches(node),
hasAsyncOps: this.detectAsyncOperations(node),
throwStatements: this.findThrowStatements(node)
});
}
ts.forEachChild(node, visitor);
};
ts.forEachChild(sourceFile, visitor);
return { functions, existingTestDescriptions: [] };
}
}
Intelligent Test Generation
The generator creates tests that are meaningful — not just coverage-padding assertions.
import anthropic
import json
from pathlib import Path
class TestGenerator:
def __init__(self):
self.client = anthropic.Anthropic()
def generate_test_suite(
self,
source_code: str,
analysis: dict,
project_context: dict
) -> str:
"""Generate a comprehensive test suite for a module."""
response = self.client.messages.create(
model="claude-sonnet-4-20250514",
max_tokens=8192,
system="""You are a senior test engineer writing comprehensive test suites.
Write tests that:
- Verify behavior, not implementation details
- Cover all execution paths including error cases
- Use descriptive test names that document behavior
- Follow the Arrange-Act-Assert pattern
- Mock external dependencies appropriately
- Include edge cases and boundary conditions
- Test concurrent/async behavior where relevant
Use the project's existing test patterns and frameworks.""",
messages=[{
"role": "user",
"content": f"""Generate a complete test file for this module.
## Source Code
```typescript
{source_code}
Analysis (paths to cover)
{json.dumps(analysis, indent=2)}
Project Test Conventions
- Framework: Jest with TypeScript
- Mocking: jest.mock() for modules, jest.fn() for functions
- Assertions: expect().toBe/toEqual/toThrow patterns
- File naming: *.test.ts alongside source files
- Describe blocks organized by function name
Existing Patterns in Project
{project_context['test_example']}
Requirements
- Cover ALL uncovered branches identified in analysis
- Test error handling explicitly (what happens when deps fail?)
- Include at least one concurrent/race condition test if applicable
- Test boundary values for numeric parameters
- Verify side effects are triggered correctly
- Test idempotency where relevant
Generate the complete test file.""" }] )
return response.content[0].text
def generate_integration_test(
self,
endpoint_code: str,
schema: dict,
related_tests: list[str]
) -> str:
"""Generate integration tests for API endpoints."""
response = self.client.messages.create(
model="claude-sonnet-4-20250514",
max_tokens=4096,
messages=[{
"role": "user",
"content": f"""Generate integration tests for this API endpoint.
Handler Code
{endpoint_code}
Request/Response Schema
{json.dumps(schema, indent=2)}
Test Scenarios Required
- Happy path with valid data
- Authentication failures (missing token, expired token, wrong role)
- Validation failures (each required field missing, invalid types)
- Business logic errors (duplicate, not found, conflict)
- Downstream service failures (timeout, 500, invalid response)
- Rate limiting behavior
- Concurrent request handling
Use supertest for HTTP testing. Mock downstream services. Each test should be independent and not rely on execution order.""" }] )
return response.content[0].text
### Test Validation Pipeline
Generated tests must themselves be valid — they need to compile, run, and verify real behavior.
```typescript
interface TestValidationResult {
compiles: boolean;
runs: boolean;
passes: boolean;
coverageIncrease: number;
mutationScore: number;
issues: string[];
}
class TestValidator {
async validate(testCode: string, targetFile: string): Promise<TestValidationResult> {
const testPath = this.writeTestFile(testCode, targetFile);
try {
// Step 1: Type checking
const compiles = await this.typeCheck(testPath);
if (!compiles.success) {
return this.fixAndRetry(testCode, compiles.errors, targetFile);
}
// Step 2: Run the tests
const runResult = await this.runTests(testPath);
// Step 3: Measure coverage increase
const coverageBefore = await this.measureCoverage(targetFile, null);
const coverageAfter = await this.measureCoverage(targetFile, testPath);
// Step 4: Mutation testing (do tests catch real bugs?)
const mutationScore = await this.runMutationTesting(targetFile, testPath);
return {
compiles: true,
runs: runResult.success,
passes: runResult.allPassed,
coverageIncrease: coverageAfter - coverageBefore,
mutationScore,
issues: runResult.failures || []
};
} finally {
this.cleanupTestFile(testPath);
}
}
private async fixAndRetry(
originalTest: string,
errors: string[],
targetFile: string
): Promise<TestValidationResult> {
// Ask Claude to fix compilation errors
const anthropic = new Anthropic();
const response = await anthropic.messages.create({
model: 'claude-sonnet-4-20250514',
max_tokens: 8192,
messages: [{
role: 'user',
content: `Fix these TypeScript compilation errors in the test file:
Errors:
${errors.join('\n')}
Original test code:
\`\`\`typescript
${originalTest}
\`\`\`
Return the fixed complete test file.`
}]
});
const fixedCode = response.content[0].text;
return this.validate(fixedCode, targetFile);
}
}
Results
After 8 weeks of running the pipeline across 312 source files:
| Metric | Before | After | Change |
|---|---|---|---|
| Line coverage | 43% | 94% | +51 points |
| Branch coverage | 31% | 87% | +56 points |
| Tests generated | 0 | 2,847 | - |
| Tests passing on first generation | - | 78% | - |
| Tests passing after auto-fix | - | 96% | - |
| Real bugs found by new tests | - | 23 | - |
| Mutation score | 34% | 72% | +38 points |
The 23 real bugs were the most compelling result. These included null pointer exceptions in error paths, race conditions in async operations, and off-by-one errors in pagination logic.
Bug Categories Found
The generated tests uncovered bugs in these categories:
- Unhandled null/undefined (8 bugs) — Functions that assumed inputs were always present
- Race conditions (5 bugs) — Async operations that could interleave incorrectly
- Off-by-one errors (4 bugs) — Pagination and array boundary issues
- Error propagation (3 bugs) — Errors that were swallowed silently
- Type coercion (3 bugs) — String/number comparison issues
Quality vs. Quantity
Not all generated tests are equally valuable. We measure test quality using mutation testing — if you introduce a bug, does the test catch it? Our mutation score went from 34% to 72%, meaning the generated tests actually verify behavior rather than just executing code paths.
Cost and Performance
- Generation cost: ~$0.12 per source file (one-time)
- Total generation cost for 312 files: $37.44
- Auto-fix iterations: avg. 1.3 per file
- Generation time: avg. 18 seconds per file
- Human review time saved: estimated 400+ hours vs. writing tests manually
Limitations
The system struggles with:
- Visual/UI tests — Can't generate meaningful Playwright tests without seeing the UI
- Performance tests — Doesn't know what "acceptable latency" means for your context
- Data-dependent tests — Complex database state setup requires domain knowledge
- Flaky test prevention — Sometimes generates tests with timing dependencies
Conclusion
AI-generated tests aren't a replacement for thoughtful test design — but they're an excellent replacement for "no tests at all," which is what most undercovered codebases have. The pipeline pays for itself immediately: $37 in API costs to find 23 bugs that would have cost thousands in production incidents. Run it against your least-covered modules first, validate with mutation testing, and use the generated tests as a starting point that engineers can refine.
Recommended reading

The State of Agentic AI in 2026: Capabilities, Limitations, and Production Readiness
Comprehensive analysis of agentic AI in 2026 covering production capabilities, current limitations, and enterprise readiness benchmarks with real deployment data.

Observability for AI Agents: Tracing Multi-Step Reasoning Chains in Production
How to implement production observability for AI agents including distributed tracing, reasoning chain analysis, and debugging multi-step failures.

Measuring and Reducing AI Workload Carbon Emissions: A Practical Engineering Guide
Building a carbon-aware scheduling system for ML training and inference workloads that reduced our AI infrastructure emissions by 42% while maintaining SLA commitments.

Comments
No comments yet. Be the first to share your thoughts.