Auto-Generating API Documentation from Codebases with Claude

How we built a documentation pipeline that generates and maintains API docs directly from source code, reducing doc drift to near-zero and saving 15 hours per sprint.

#claude#documentation#api#automation
Cover image for the article: Auto-Generating API Documentation from Codebases with Claude

API documentation is always wrong. Not intentionally — it starts accurate, then drifts as code changes faster than anyone updates the docs. We built a Claude-powered documentation pipeline that generates comprehensive API documentation directly from source code, type definitions, and test files. It runs on every merge to main and produces OpenAPI specs, usage guides, and migration notes automatically.

The result: documentation that's always current, developer satisfaction with docs up from 34% to 89%, and 15 hours per sprint reclaimed from manual documentation work.

The Problem

Our platform exposes 147 REST endpoints and 23 GraphQL operations. Documentation quality was inconsistent:

  • 31% of endpoints had no documentation at all
  • 44% of documented endpoints had at least one inaccuracy
  • Breaking changes were communicated via Slack (and missed by half the consumers)
  • New team members spent an average of 6 hours figuring out undocumented API behavior

The traditional fix — "developers should write docs" — had failed for 3 years. We needed documentation that writes itself.

Architecture

The pipeline has three stages: code analysis, documentation generation, and publication.

API Documentation Pipeline Architecture

Code Analysis Layer

Before Claude generates any documentation, we extract structured information from the codebase using static analysis.

import * as ts from 'typescript';
import Anthropic from '@anthropic-ai/sdk';

interface EndpointMetadata {
  path: string;
  method: string;
  handler: string;
  requestSchema: object | null;
  responseSchema: object | null;
  middleware: string[];
  authentication: string;
  rateLimit: string | null;
  validationRules: object[];
  testCoverage: TestInfo[];
  recentChanges: GitChange[];
}

class CodebaseAnalyzer {
  private program: ts.Program;

  constructor(tsconfigPath: string) {
    const config = ts.readConfigFile(tsconfigPath, ts.sys.readFile);
    const parsedConfig = ts.parseJsonConfigFileContent(
      config.config, ts.sys, '.'
    );
    this.program = ts.createProgram(
      parsedConfig.fileNames,
      parsedConfig.options
    );
  }

  analyzeEndpoint(routeFile: string, handlerName: string): EndpointMetadata {
    const sourceFile = this.program.getSourceFile(routeFile);
    const checker = this.program.getTypeChecker();

    // Extract request/response types from handler signature
    const handler = this.findFunction(sourceFile, handlerName);
    const requestType = this.extractRequestType(handler, checker);
    const responseType = this.extractResponseType(handler, checker);

    // Extract middleware chain
    const middleware = this.extractMiddleware(routeFile, handlerName);

    // Find associated test files
    const tests = this.findRelatedTests(routeFile, handlerName);

    // Get git history for recent changes
    const changes = this.getRecentChanges(routeFile, 30);

    return {
      path: this.extractRoutePath(routeFile, handlerName),
      method: this.extractHttpMethod(routeFile, handlerName),
      handler: handlerName,
      requestSchema: requestType ? this.typeToJsonSchema(requestType, checker) : null,
      responseSchema: responseType ? this.typeToJsonSchema(responseType, checker) : null,
      middleware: middleware,
      authentication: this.detectAuthRequirement(middleware),
      rateLimit: this.extractRateLimit(middleware),
      validationRules: this.extractValidation(handler, checker),
      testCoverage: tests,
      recentChanges: changes
    };
  }

  private typeToJsonSchema(type: ts.Type, checker: ts.TypeChecker): object {
    // Convert TypeScript type to JSON Schema for documentation
    const properties: Record<string, any> = {};
    const required: string[] = [];

    for (const prop of type.getProperties()) {
      const propType = checker.getTypeOfSymbolAtLocation(
        prop, prop.valueDeclaration!
      );
      const isOptional = (prop.flags & ts.SymbolFlags.Optional) !== 0;

      properties[prop.name] = {
        type: this.tsTypeToJsonType(propType, checker),
        description: ts.displayPartsToString(
          prop.getDocumentationComment(checker)
        )
      };

      if (!isOptional) required.push(prop.name);
    }

    return { type: 'object', properties, required };
  }
}

Documentation Generation with Claude

With structured metadata extracted, Claude generates human-readable documentation that goes beyond what automated tools can produce.

import anthropic
import json
from pathlib import Path

class DocumentationGenerator:
    def __init__(self):
        self.client = anthropic.Anthropic()

    def generate_endpoint_doc(self, metadata: dict, existing_doc: str | None = None) -> str:
        """Generate comprehensive documentation for a single endpoint."""
        
        prompt = f"""Generate API documentation for this endpoint.

## Endpoint Metadata
- Path: {metadata['path']}
- Method: {metadata['method']}
- Authentication: {metadata['authentication']}
- Rate Limit: {metadata['rateLimit'] or 'None specified'}

## Request Schema
```json
{json.dumps(metadata['requestSchema'], indent=2)}

Response Schema

{json.dumps(metadata['responseSchema'], indent=2)}

Validation Rules

{json.dumps(metadata['validationRules'], indent=2)}

Test Cases (showing real usage patterns)

{self._format_tests(metadata['testCoverage'])}

Recent Changes

{self._format_changes(metadata['recentChanges'])}

{f"## Previous Documentation (update, don't rewrite from scratch):\n{existing_doc}" if existing_doc else ""}

Generate documentation including:

  1. One-line description

  2. Authentication requirements

  3. Request parameters (path, query, body) with types and constraints

  4. Response format with all possible status codes

  5. 2-3 realistic curl/fetch examples

  6. Common error scenarios and how to handle them

  7. Rate limiting details

  8. Any breaking changes in the last 30 days"""

     response = self.client.messages.create(
         model="claude-sonnet-4-20250514",
         max_tokens=4096,
         system="""You are a technical writer creating API documentation.
    

Write for developers who will integrate with this API. Be precise about types, constraints, and edge cases. Include practical examples that demonstrate real usage patterns. Flag any undocumented behavior you infer from test cases.""", messages=[{"role": "user", "content": prompt}] )

    return response.content[0].text

def generate_changelog(self, old_metadata: dict, new_metadata: dict) -> str:
    """Generate migration notes when an endpoint changes."""
    
    response = self.client.messages.create(
        model="claude-sonnet-4-20250514",
        max_tokens=2048,
        messages=[{
            "role": "user",
            "content": f"""Compare these two versions of an API endpoint and generate a changelog.

Previous Version

{json.dumps(old_metadata, indent=2)}

Current Version

{json.dumps(new_metadata, indent=2)}

Generate:

  1. Summary of changes (breaking vs. non-breaking)

  2. Migration steps if breaking

  3. Deprecation notices if applicable

  4. Before/after code examples""" }] )

     return response.content[0].text
    

    def generate_openapi_spec(self, all_endpoints: list[dict]) -> dict: """Generate a complete OpenAPI 3.1 specification."""

     response = self.client.messages.create(
         model="claude-sonnet-4-20250514",
         max_tokens=16384,
         messages=[{
             "role": "user",
             "content": f"""Generate a valid OpenAPI 3.1 specification from these endpoints.
    

Endpoints: {json.dumps(all_endpoints, indent=2)}

Requirements:

  • Valid OpenAPI 3.1 JSON

  • Include all request/response schemas

  • Add realistic examples for each endpoint

  • Group endpoints by tag (derive from path structure)

  • Include security schemes based on authentication types

  • Add rate limit headers to responses""" }] )

      return json.loads(response.content[0].text)
    

### CI/CD Integration and Diff Detection

The pipeline runs on every merge to main, but only regenerates documentation for changed endpoints.

```typescript
class DocumentationCIPipeline {
  private analyzer: CodebaseAnalyzer;
  private generator: DocumentationGenerator;

  async run(changedFiles: string[]): Promise<PipelineResult> {
    // Step 1: Identify affected endpoints
    const affectedEndpoints = this.findAffectedEndpoints(changedFiles);

    if (affectedEndpoints.length === 0) {
      return { status: 'no-changes', updatedDocs: [] };
    }

    // Step 2: Analyze current state
    const currentMetadata = affectedEndpoints.map(ep =>
      this.analyzer.analyzeEndpoint(ep.file, ep.handler)
    );

    // Step 3: Load previous metadata for diff detection
    const previousMetadata = await this.loadPreviousMetadata(affectedEndpoints);

    // Step 4: Generate/update documentation
    const updates: DocUpdate[] = [];
    for (let i = 0; i < affectedEndpoints.length; i++) {
      const current = currentMetadata[i];
      const previous = previousMetadata[i];

      const doc = await this.generator.generateEndpointDoc(
        current,
        await this.loadExistingDoc(current.path)
      );

      // Generate changelog if breaking change detected
      let changelog: string | null = null;
      if (previous && this.isBreakingChange(previous, current)) {
        changelog = await this.generator.generateChangelog(previous, current);
      }

      updates.push({ endpoint: current.path, documentation: doc, changelog });
    }

    // Step 5: Publish
    await this.publishToDocSite(updates);
    await this.updateOpenAPISpec(currentMetadata);
    await this.notifyConsumers(updates.filter(u => u.changelog));

    return { status: 'updated', updatedDocs: updates };
  }

  private isBreakingChange(previous: EndpointMetadata, current: EndpointMetadata): boolean {
    // Detect breaking changes: removed fields, type changes, new required params
    const prevRequired = new Set(previous.requestSchema?.required || []);
    const currRequired = new Set(current.requestSchema?.required || []);

    // New required fields = breaking
    for (const field of currRequired) {
      if (!prevRequired.has(field)) return true;
    }

    // Removed response fields = breaking
    const prevResponseFields = Object.keys(previous.responseSchema?.properties || {});
    const currResponseFields = new Set(Object.keys(current.responseSchema?.properties || {}));
    for (const field of prevResponseFields) {
      if (!currResponseFields.has(field)) return true;
    }

    return false;
  }
}

Results

After 4 months in production:

MetricBeforeAfterChange
Documented endpoints69%100%Complete coverage
Documentation accuracy56%97%+41 points
Developer doc satisfaction34%89%+55 points
Time spent writing docs15 hrs/sprint1.5 hrs/sprint90% reduction
Breaking change communicationManual/SlackAutomated alerts100% coverage
New integration time3 days avg.4 hours avg.82% faster

Quality Controls

  1. Schema validation — Generated OpenAPI specs are validated against the OpenAPI 3.1 standard
  2. Example testing — Code examples in docs are extracted and run against a test environment
  3. Human review — Major version changes still get technical writer review
  4. Freshness scoring — Docs are flagged if their source code has changed without doc regeneration

Cost Analysis

Monthly costs for documenting 147 endpoints with daily updates:

  • Claude API calls: $340/month (mostly incremental updates)
  • CI/CD compute for analysis: $45/month
  • Doc hosting: $20/month

Total: $405/month — less than 3 hours of a technical writer's time.

Conclusion

Documentation that generates itself from source code eliminates the drift problem permanently. The key insight is that source code, type definitions, tests, and git history contain everything needed to produce excellent documentation — Claude just needs to synthesize it into a developer-friendly format. Stop asking engineers to write docs manually. Build a pipeline that makes accurate documentation an automatic consequence of shipping code.

Comments

    No comments yet. Be the first to share your thoughts.