Document Understanding Pipelines with Claude Vision

Building production document processing systems with Claude multi-modal capabilities for OCR, extraction, and intelligent document understanding.

#claude#vision#document-processing#ai
Cover image for the article: Document Understanding Pipelines with Claude Vision

Traditional OCR extracts text. Claude's vision capabilities understand documents — layout, tables, handwriting, diagrams, and the relationships between elements. This distinction transforms document processing from a character recognition problem into a comprehension problem. Here's how we built a production document understanding pipeline that processes 50,000 documents per day.

Beyond OCR: document understanding

The limitation of traditional OCR pipelines is that they lose structure. A scanned invoice becomes a flat string of text where the relationship between "Total" and "$4,287.50" disappears. Claude vision preserves these relationships because it sees the document the way a human does — spatially.

Use cases where vision-based processing dramatically outperforms text extraction:

  • Invoices and receipts — structured data extraction from varied layouts
  • Technical drawings — understanding diagrams, flowcharts, architecture documents
  • Handwritten notes — meeting notes, whiteboard photos, forms
  • Multi-language documents — mixed-script documents without language detection overhead
  • Tables and charts — preserving structure that OCR flattens

Document Processing Pipeline Architecture

Architecture: the document processing pipeline

Stage 1: Document ingestion and preprocessing

Before sending documents to Claude, optimize them for vision processing:

import anthropic
import base64
from dataclasses import dataclass
from pathlib import Path
from typing import Optional
from PIL import Image
import io

@dataclass
class ProcessedDocument:
    pages: list[bytes]  # Base64-encoded page images
    page_count: int
    original_format: str
    dpi: int

class DocumentPreprocessor:
    """Prepare documents for Claude vision processing."""

    SUPPORTED_FORMATS = {"pdf", "png", "jpg", "jpeg", "tiff", "webp"}
    TARGET_DPI = 150  # Balance between quality and token cost
    MAX_DIMENSION = 1568  # Claude's optimal resolution

    def process(self, file_path: Path) -> ProcessedDocument:
        suffix = file_path.suffix.lower().lstrip(".")
        if suffix not in self.SUPPORTED_FORMATS:
            raise ValueError(f"Unsupported format: {suffix}")

        if suffix == "pdf":
            pages = self._process_pdf(file_path)
        else:
            pages = [self._process_image(file_path)]

        return ProcessedDocument(
            pages=pages,
            page_count=len(pages),
            original_format=suffix,
            dpi=self.TARGET_DPI,
        )

    def _process_image(self, path: Path) -> bytes:
        """Resize and optimize image for Claude vision."""
        with Image.open(path) as img:
            # Convert to RGB if necessary
            if img.mode not in ("RGB", "L"):
                img = img.convert("RGB")

            # Resize if larger than max dimension
            if max(img.size) > self.MAX_DIMENSION:
                ratio = self.MAX_DIMENSION / max(img.size)
                new_size = (int(img.width * ratio), int(img.height * ratio))
                img = img.resize(new_size, Image.LANCZOS)

            # Encode as JPEG for optimal token usage
            buffer = io.BytesIO()
            img.save(buffer, format="JPEG", quality=85, optimize=True)
            return base64.b64encode(buffer.getvalue()).decode()

    def _process_pdf(self, path: Path) -> list[bytes]:
        """Convert PDF pages to optimized images."""
        # Using pdf2image or similar library
        # Each page becomes a separate image for processing
        import pdf2image
        images = pdf2image.convert_from_path(
            path, dpi=self.TARGET_DPI
        )
        pages = []
        for img in images:
            buffer = io.BytesIO()
            if max(img.size) > self.MAX_DIMENSION:
                ratio = self.MAX_DIMENSION / max(img.size)
                new_size = (int(img.width * ratio), int(img.height * ratio))
                img = img.resize(new_size, Image.LANCZOS)
            img.save(buffer, format="JPEG", quality=85)
            pages.append(base64.b64encode(buffer.getvalue()).decode())
        return pages

Stage 2: Intelligent extraction with structured output

The power of Claude vision is in extraction — pulling structured data from unstructured visual inputs:

import Anthropic from "@anthropic-ai/sdk";

interface ExtractionSchema {
  name: string;
  fields: FieldDefinition[];
}

interface FieldDefinition {
  name: string;
  type: "string" | "number" | "date" | "boolean" | "array" | "object";
  description: string;
  required: boolean;
}

interface ExtractionResult {
  success: boolean;
  data: Record<string, unknown>;
  confidence: number;
  processingTime: number;
}

class DocumentExtractor {
  private client: Anthropic;

  constructor(client: Anthropic) {
    this.client = client;
  }

  async extract(
    pageImages: string[],
    schema: ExtractionSchema
  ): Promise<ExtractionResult> {
    const startTime = Date.now();

    const schemaDescription = this.buildSchemaPrompt(schema);

    const content: Anthropic.ContentBlockParam[] = [];

    // Add each page as an image
    for (const pageImage of pageImages) {
      content.push({
        type: "image",
        source: {
          type: "base64",
          media_type: "image/jpeg",
          data: pageImage,
        },
      });
    }

    // Add extraction instructions
    content.push({
      type: "text",
      text: `Extract the following information from this document.
Return ONLY valid JSON matching the schema below. No explanation or commentary.

Schema:
${schemaDescription}

If a field cannot be determined from the document, use null.
For dates, use ISO 8601 format (YYYY-MM-DD).
For currency amounts, return as numbers without currency symbols.`,
    });

    const response = await this.client.messages.create({
      model: "claude-sonnet-4-20250514",
      max_tokens: 4096,
      messages: [{ role: "user", content }],
    });

    const text =
      response.content[0].type === "text" ? response.content[0].text : "";

    try {
      const data = JSON.parse(text);
      const confidence = this.calculateConfidence(data, schema);

      return {
        success: true,
        data,
        confidence,
        processingTime: Date.now() - startTime,
      };
    } catch {
      return {
        success: false,
        data: {},
        confidence: 0,
        processingTime: Date.now() - startTime,
      };
    }
  }

  private buildSchemaPrompt(schema: ExtractionSchema): string {
    const fields = schema.fields.map(
      (f) =>
        `  "${f.name}": ${f.type}${f.required ? " (required)" : " (optional)"} - ${f.description}`
    );
    return `{\n${fields.join(",\n")}\n}`;
  }

  private calculateConfidence(
    data: Record<string, unknown>,
    schema: ExtractionSchema
  ): number {
    const requiredFields = schema.fields.filter((f) => f.required);
    const filledRequired = requiredFields.filter(
      (f) => data[f.name] !== null && data[f.name] !== undefined
    );
    return filledRequired.length / requiredFields.length;
  }
}

Stage 3: Multi-page document understanding

For documents longer than a single page, you need a strategy for maintaining context across pages:

  1. Process pages independently for extraction tasks where each page is self-contained (invoices, forms)
  2. Process pages together when content flows across pages (contracts, reports)
  3. Summarize then synthesize for very long documents — summarize each page, then reason over summaries

Our production system uses a router that classifies documents first, then applies the appropriate multi-page strategy.

Production performance benchmarks

Processing 50,000 documents per day across four document types:

Document typeAccuracyAvg processing timeCost per doc
Invoices96.2%2.1s$0.008
Receipts94.8%1.4s$0.005
Contracts (multi-page)91.3%8.7s$0.032
Handwritten forms89.1%2.8s$0.010

Compared to our previous OCR + rules-based pipeline:

MetricOCR + RulesClaude VisionImprovement
Overall accuracy78%93%+19%
Setup time per new doc type2-4 weeks2-4 hours~95% faster
Maintenance burdenHigh (brittle rules)Low (schema-driven)—
Handling layout variationsPoorExcellent—

Cost management strategies

Vision processing is more expensive than text-only Claude calls. Optimize with:

  1. Resolution optimization — 150 DPI is sufficient for most documents. Don't send 300 DPI scans.
  2. Page routing — skip blank pages, cover pages, and pages without extractable content
  3. Batch processing — use the Batch API for non-urgent document processing (50% cost reduction)
  4. Cache extracted data — don't reprocess the same document twice

Key takeaways

  1. Vision is comprehension, not OCR. Claude understands layout, tables, and relationships between elements. Leverage that.
  2. Preprocessing matters enormously. Resolution, format, and image quality directly impact extraction accuracy and cost.
  3. Schema-driven extraction makes it trivial to add new document types without code changes — just define the schema.
  4. Multi-page strategy must match document type. Don't force a one-size-fits-all approach.
  5. Measure accuracy per field, not just per document. Some fields extract perfectly while others struggle — targeted improvements are more effective.

Document processing with Claude vision eliminates the most painful part of traditional OCR: building and maintaining extraction rules for every layout variation. Define what you want, show the document, get structured data back.

Comments

    No comments yet. Be the first to share your thoughts.