Achieving 94% Test Coverage with AI-Generated Tests

How we used Claude to generate meaningful test suites that increased coverage from 43% to 94% while catching real bugs that manual tests missed.

#claude#testing#test-generation#quality
Cover image for the article: Achieving 94% Test Coverage with AI-Generated Tests

Our codebase had 43% test coverage. Not because engineers were lazy — they were shipping features faster than they could write tests. The coverage gap was concentrated in the oldest, most complex modules — exactly the code most likely to break. We built a Claude-powered test generation system that analyzed our source code, generated comprehensive test suites, and increased coverage to 94% in 8 weeks. More importantly, the generated tests caught 23 real bugs that had been lurking undetected.

This isn't about generating trivial assertion tests. It's about producing meaningful tests that verify behavior, edge cases, and failure modes — tests a senior engineer would write if they had unlimited time.

The Problem

Our coverage breakdown by module:

  • Core business logic: 67% (reasonable, engineers prioritized this)
  • API handlers: 52% (tested happy paths, missed edge cases)
  • Data access layer: 38% (mocking complexity discouraged testing)
  • Utility functions: 71% (easy to test, well-covered)
  • Integration points: 19% (hardest to test, most critical)
  • Error handling paths: 12% (almost entirely untested)

The 12% on error handling was terrifying. We were shipping code with almost no verification of how it behaved when things went wrong.

Architecture

The test generation pipeline operates in four phases: analysis, generation, validation, and integration.

Test Generation Pipeline Architecture

Source Code Analysis

Before generating tests, we build a comprehensive understanding of what each module does, its dependencies, and its failure modes.

import Anthropic from '@anthropic-ai/sdk';
import * as ts from 'typescript';

interface FunctionAnalysis {
  name: string;
  filePath: string;
  signature: string;
  returnType: string;
  parameters: { name: string; type: string; optional: boolean }[];
  dependencies: string[];
  sideEffects: string[];
  errorPaths: string[];
  branchComplexity: number;
  existingTests: string[];
  uncoveredBranches: string[];
}

class TestAnalyzer {
  private anthropic: Anthropic;

  constructor() {
    this.anthropic = new Anthropic();
  }

  async analyzeForTesting(filePath: string, sourceCode: string): Promise<FunctionAnalysis[]> {
    // Static analysis for structure
    const staticAnalysis = this.performStaticAnalysis(filePath, sourceCode);

    // Claude analysis for behavior understanding
    const response = await this.anthropic.messages.create({
      model: 'claude-sonnet-4-20250514',
      max_tokens: 4096,
      messages: [{
        role: 'user',
        content: `Analyze this code for test generation. For each function, identify:
1. All possible execution paths (happy + error)
2. Edge cases based on parameter types and constraints
3. Side effects (DB writes, API calls, file operations)
4. Dependencies that need mocking
5. Invariants that should be verified
6. Boundary conditions

Source code:
\`\`\`typescript
${sourceCode}
\`\`\`

Existing test coverage (what's already tested):
${staticAnalysis.existingTestDescriptions.join('\n')}

Return JSON array of analysis objects per function.`
      }]
    });

    const claudeAnalysis = JSON.parse(response.content[0].text);
    return this.mergeAnalyses(staticAnalysis, claudeAnalysis);
  }

  private performStaticAnalysis(filePath: string, source: string) {
    const sourceFile = ts.createSourceFile(
      filePath, source, ts.ScriptTarget.Latest, true
    );

    const functions: any[] = [];
    const visitor = (node: ts.Node) => {
      if (ts.isFunctionDeclaration(node) || ts.isMethodDeclaration(node)) {
        functions.push({
          name: node.name?.getText(sourceFile) || 'anonymous',
          branchCount: this.countBranches(node),
          hasAsyncOps: this.detectAsyncOperations(node),
          throwStatements: this.findThrowStatements(node)
        });
      }
      ts.forEachChild(node, visitor);
    };
    ts.forEachChild(sourceFile, visitor);

    return { functions, existingTestDescriptions: [] };
  }
}

Intelligent Test Generation

The generator creates tests that are meaningful — not just coverage-padding assertions.

import anthropic
import json
from pathlib import Path

class TestGenerator:
    def __init__(self):
        self.client = anthropic.Anthropic()

    def generate_test_suite(
        self,
        source_code: str,
        analysis: dict,
        project_context: dict
    ) -> str:
        """Generate a comprehensive test suite for a module."""
        
        response = self.client.messages.create(
            model="claude-sonnet-4-20250514",
            max_tokens=8192,
            system="""You are a senior test engineer writing comprehensive test suites.
Write tests that:
- Verify behavior, not implementation details
- Cover all execution paths including error cases
- Use descriptive test names that document behavior
- Follow the Arrange-Act-Assert pattern
- Mock external dependencies appropriately
- Include edge cases and boundary conditions
- Test concurrent/async behavior where relevant

Use the project's existing test patterns and frameworks.""",
            messages=[{
                "role": "user",
                "content": f"""Generate a complete test file for this module.

## Source Code
```typescript
{source_code}

Analysis (paths to cover)

{json.dumps(analysis, indent=2)}

Project Test Conventions

  • Framework: Jest with TypeScript
  • Mocking: jest.mock() for modules, jest.fn() for functions
  • Assertions: expect().toBe/toEqual/toThrow patterns
  • File naming: *.test.ts alongside source files
  • Describe blocks organized by function name

Existing Patterns in Project

{project_context['test_example']}

Requirements

  1. Cover ALL uncovered branches identified in analysis
  2. Test error handling explicitly (what happens when deps fail?)
  3. Include at least one concurrent/race condition test if applicable
  4. Test boundary values for numeric parameters
  5. Verify side effects are triggered correctly
  6. Test idempotency where relevant

Generate the complete test file.""" }] )

    return response.content[0].text

def generate_integration_test(
    self,
    endpoint_code: str,
    schema: dict,
    related_tests: list[str]
) -> str:
    """Generate integration tests for API endpoints."""
    
    response = self.client.messages.create(
        model="claude-sonnet-4-20250514",
        max_tokens=4096,
        messages=[{
            "role": "user",
            "content": f"""Generate integration tests for this API endpoint.

Handler Code

{endpoint_code}

Request/Response Schema

{json.dumps(schema, indent=2)}

Test Scenarios Required

  1. Happy path with valid data
  2. Authentication failures (missing token, expired token, wrong role)
  3. Validation failures (each required field missing, invalid types)
  4. Business logic errors (duplicate, not found, conflict)
  5. Downstream service failures (timeout, 500, invalid response)
  6. Rate limiting behavior
  7. Concurrent request handling

Use supertest for HTTP testing. Mock downstream services. Each test should be independent and not rely on execution order.""" }] )

    return response.content[0].text

### Test Validation Pipeline

Generated tests must themselves be valid — they need to compile, run, and verify real behavior.

```typescript
interface TestValidationResult {
  compiles: boolean;
  runs: boolean;
  passes: boolean;
  coverageIncrease: number;
  mutationScore: number;
  issues: string[];
}

class TestValidator {
  async validate(testCode: string, targetFile: string): Promise<TestValidationResult> {
    const testPath = this.writeTestFile(testCode, targetFile);

    try {
      // Step 1: Type checking
      const compiles = await this.typeCheck(testPath);
      if (!compiles.success) {
        return this.fixAndRetry(testCode, compiles.errors, targetFile);
      }

      // Step 2: Run the tests
      const runResult = await this.runTests(testPath);

      // Step 3: Measure coverage increase
      const coverageBefore = await this.measureCoverage(targetFile, null);
      const coverageAfter = await this.measureCoverage(targetFile, testPath);

      // Step 4: Mutation testing (do tests catch real bugs?)
      const mutationScore = await this.runMutationTesting(targetFile, testPath);

      return {
        compiles: true,
        runs: runResult.success,
        passes: runResult.allPassed,
        coverageIncrease: coverageAfter - coverageBefore,
        mutationScore,
        issues: runResult.failures || []
      };
    } finally {
      this.cleanupTestFile(testPath);
    }
  }

  private async fixAndRetry(
    originalTest: string,
    errors: string[],
    targetFile: string
  ): Promise<TestValidationResult> {
    // Ask Claude to fix compilation errors
    const anthropic = new Anthropic();
    const response = await anthropic.messages.create({
      model: 'claude-sonnet-4-20250514',
      max_tokens: 8192,
      messages: [{
        role: 'user',
        content: `Fix these TypeScript compilation errors in the test file:

Errors:
${errors.join('\n')}

Original test code:
\`\`\`typescript
${originalTest}
\`\`\`

Return the fixed complete test file.`
      }]
    });

    const fixedCode = response.content[0].text;
    return this.validate(fixedCode, targetFile);
  }
}

Results

After 8 weeks of running the pipeline across 312 source files:

MetricBeforeAfterChange
Line coverage43%94%+51 points
Branch coverage31%87%+56 points
Tests generated02,847-
Tests passing on first generation-78%-
Tests passing after auto-fix-96%-
Real bugs found by new tests-23-
Mutation score34%72%+38 points

The 23 real bugs were the most compelling result. These included null pointer exceptions in error paths, race conditions in async operations, and off-by-one errors in pagination logic.

Bug Categories Found

The generated tests uncovered bugs in these categories:

  1. Unhandled null/undefined (8 bugs) — Functions that assumed inputs were always present
  2. Race conditions (5 bugs) — Async operations that could interleave incorrectly
  3. Off-by-one errors (4 bugs) — Pagination and array boundary issues
  4. Error propagation (3 bugs) — Errors that were swallowed silently
  5. Type coercion (3 bugs) — String/number comparison issues

Quality vs. Quantity

Not all generated tests are equally valuable. We measure test quality using mutation testing — if you introduce a bug, does the test catch it? Our mutation score went from 34% to 72%, meaning the generated tests actually verify behavior rather than just executing code paths.

Cost and Performance

  • Generation cost: ~$0.12 per source file (one-time)
  • Total generation cost for 312 files: $37.44
  • Auto-fix iterations: avg. 1.3 per file
  • Generation time: avg. 18 seconds per file
  • Human review time saved: estimated 400+ hours vs. writing tests manually

Limitations

The system struggles with:

  • Visual/UI tests — Can't generate meaningful Playwright tests without seeing the UI
  • Performance tests — Doesn't know what "acceptable latency" means for your context
  • Data-dependent tests — Complex database state setup requires domain knowledge
  • Flaky test prevention — Sometimes generates tests with timing dependencies

Conclusion

AI-generated tests aren't a replacement for thoughtful test design — but they're an excellent replacement for "no tests at all," which is what most undercovered codebases have. The pipeline pays for itself immediately: $37 in API costs to find 23 bugs that would have cost thousands in production incidents. Run it against your least-covered modules first, validate with mutation testing, and use the generated tests as a starting point that engineers can refine.

Comments

    No comments yet. Be the first to share your thoughts.