Structured Data Extraction from Unstructured Documents with Claude

Building production data extraction pipelines that convert messy PDFs, emails, and scanned documents into structured data with 97.3% accuracy.

#claude#data-extraction#nlp#automation
Cover image for the article: Structured Data Extraction from Unstructured Documents with Claude

Every company has a pile of unstructured documents that contain critical business data: invoices, contracts, medical records, compliance filings, research papers. Traditional approaches — regex, rule-based parsers, legacy OCR — break constantly as document formats change. We built a Claude-powered extraction pipeline that handles 14 document types across 3 languages with 97.3% field-level accuracy.

This post covers the pipeline architecture, our schema-driven approach, handling edge cases at scale, and the validation framework that makes this production-ready.

The Problem

Our fintech client processes 45,000 documents per month from 200+ counterparties. Each counterparty has slightly different formats for the same document types. The existing rule-based system required constant maintenance:

  • 23 custom parsers for different invoice formats
  • Average 2 parser failures per day requiring engineering intervention
  • 89% accuracy (meaning ~5,000 fields per month needed manual correction)
  • New counterparty onboarding took 2-3 weeks to write custom parsing rules

They needed a system that could handle format variation without per-format engineering.

Architecture

The pipeline processes documents through five stages: ingestion, preprocessing, extraction, validation, and output.

Data Extraction Pipeline Architecture

Document Preprocessing

Before Claude sees any content, we normalize documents into a consistent format.

import anthropic
from pdf2image import convert_from_path
import pytesseract
from pathlib import Path
from dataclasses import dataclass, field

@dataclass
class ProcessedDocument:
    doc_id: str
    doc_type: str
    pages: list[dict]
    raw_text: str
    tables: list[dict]
    metadata: dict = field(default_factory=dict)
    confidence_scores: dict = field(default_factory=dict)

class DocumentPreprocessor:
    def process(self, file_path: Path) -> ProcessedDocument:
        suffix = file_path.suffix.lower()

        if suffix == '.pdf':
            return self._process_pdf(file_path)
        elif suffix in ('.png', '.jpg', '.jpeg', '.tiff'):
            return self._process_image(file_path)
        elif suffix == '.eml':
            return self._process_email(file_path)
        else:
            raise ValueError(f"Unsupported format: {suffix}")

    def _process_pdf(self, file_path: Path) -> ProcessedDocument:
        # Try text extraction first (faster, more accurate for digital PDFs)
        text = self._extract_text_layer(file_path)

        if self._is_text_quality_sufficient(text):
            pages = self._split_into_pages(text)
            tables = self._extract_tables(file_path)
        else:
            # Fall back to OCR for scanned documents
            images = convert_from_path(file_path)
            pages = [
                {"page_num": i + 1, "text": pytesseract.image_to_string(img)}
                for i, img in enumerate(images)
            ]
            tables = self._extract_tables_from_images(images)

        return ProcessedDocument(
            doc_id=self._generate_id(file_path),
            doc_type="unknown",  # classified in next stage
            pages=pages,
            raw_text="\n".join(p["text"] for p in pages),
            tables=tables
        )

    def _is_text_quality_sufficient(self, text: str) -> bool:
        """Check if extracted text is real content vs. garbage from image PDFs"""
        if len(text.strip()) < 50:
            return False
        # Check ratio of alphanumeric to total characters
        alnum_ratio = sum(c.isalnum() for c in text) / max(len(text), 1)
        return alnum_ratio > 0.6

Schema-Driven Extraction

The key innovation is defining extraction schemas declaratively and letting Claude handle the mapping from unstructured content to structured fields.

import Anthropic from '@anthropic-ai/sdk';

interface ExtractionSchema {
  documentType: string;
  fields: FieldDefinition[];
  validationRules: ValidationRule[];
}

interface FieldDefinition {
  name: string;
  type: 'string' | 'number' | 'date' | 'currency' | 'array' | 'object';
  required: boolean;
  description: string;
  examples?: string[];
  constraints?: Record<string, any>;
}

interface ExtractionResult {
  fields: Record<string, any>;
  confidence: Record<string, number>;
  warnings: string[];
  rawEvidence: Record<string, string>;
}

class SchemaBasedExtractor {
  private anthropic: Anthropic;

  constructor() {
    this.anthropic = new Anthropic();
  }

  async extract(
    document: ProcessedDocument,
    schema: ExtractionSchema
  ): Promise<ExtractionResult> {
    const schemaPrompt = this.buildSchemaPrompt(schema);

    const response = await this.anthropic.messages.create({
      model: 'claude-sonnet-4-20250514',
      max_tokens: 4096,
      messages: [{
        role: 'user',
        content: `Extract structured data from this document according to the schema.

## Extraction Schema
${schemaPrompt}

## Document Content
${document.raw_text}

## Tables Found
${JSON.stringify(document.tables, null, 2)}

## Instructions
1. Extract each field defined in the schema
2. For each field, provide your confidence (0-1) 
3. Quote the exact text evidence for each extraction
4. Flag any ambiguities or conflicts in the source
5. If a required field is not found, set confidence to 0 and explain why

Respond as JSON:
{
  "fields": { "field_name": "extracted_value", ... },
  "confidence": { "field_name": 0.95, ... },
  "warnings": ["any concerns about extraction quality"],
  "rawEvidence": { "field_name": "exact quote from document", ... }
}`
      }]
    });

    const result = JSON.parse(response.content[0].text);
    return this.postProcess(result, schema);
  }

  private buildSchemaPrompt(schema: ExtractionSchema): string {
    return schema.fields.map(field => {
      let desc = `- **${field.name}** (${field.type}, ${field.required ? 'required' : 'optional'}): ${field.description}`;
      if (field.examples?.length) {
        desc += `\n  Examples: ${field.examples.join(', ')}`;
      }
      if (field.constraints) {
        desc += `\n  Constraints: ${JSON.stringify(field.constraints)}`;
      }
      return desc;
    }).join('\n');
  }

  private postProcess(result: ExtractionResult, schema: ExtractionSchema): ExtractionResult {
    // Apply type coercion based on schema
    for (const field of schema.fields) {
      if (result.fields[field.name] !== undefined) {
        result.fields[field.name] = this.coerceType(
          result.fields[field.name],
          field.type
        );
      }
    }
    return result;
  }

  private coerceType(value: any, type: string): any {
    switch (type) {
      case 'number': return parseFloat(String(value).replace(/[,$]/g, ''));
      case 'currency': return this.parseCurrency(value);
      case 'date': return this.parseDate(value);
      default: return value;
    }
  }
}

Multi-Pass Validation

Single-pass extraction isn't reliable enough for production. We run a validation pass that cross-references extracted fields.

class ExtractionValidator:
    def __init__(self):
        self.client = anthropic.Anthropic()

    def validate(
        self,
        extraction: dict,
        document: ProcessedDocument,
        schema: dict
    ) -> dict:
        """Cross-validate extracted fields against each other and source."""
        
        # Rule-based validation
        rule_results = self._apply_validation_rules(extraction, schema)
        
        # Claude-based consistency check
        consistency = self.client.messages.create(
            model="claude-sonnet-4-20250514",
            max_tokens=2048,
            messages=[{
                "role": "user",
                "content": f"""Review this data extraction for consistency errors.

Extracted data:
{json.dumps(extraction['fields'], indent=2)}

Source document (abbreviated):
{document.raw_text[:3000]}

Check for:
1. Mathematical consistency (do line items sum to totals?)
2. Date consistency (are dates in logical order?)
3. Reference consistency (do IDs match between sections?)
4. Format consistency (are all currencies in same denomination?)

Return JSON:
{{
    "consistent": true/false,
    "issues": ["list of inconsistencies found"],
    "corrections": {{"field_name": "corrected_value"}}
}}"""
            }]
        )

        consistency_result = json.loads(consistency.content[0].text)

        # Merge corrections if confidence is high
        if not consistency_result["consistent"]:
            extraction = self._apply_corrections(
                extraction,
                consistency_result["corrections"]
            )

        return {
            "extraction": extraction,
            "validation": {
                "rule_results": rule_results,
                "consistency": consistency_result,
                "final_confidence": self._calculate_overall_confidence(
                    extraction, rule_results, consistency_result
                )
            }
        }

Handling Edge Cases

Multi-Language Documents

Documents containing mixed languages (common in international trade) are handled by instructing Claude to extract regardless of language and normalize to English field names:

system_prompt = """Extract data regardless of the document's language.
Field names should always be in English per the schema.
Values should be preserved in their original form (e.g., keep addresses in local script).
Currency amounts should include the currency code (EUR, USD, AED, etc.)."""

Handwritten Annotations

For documents with handwritten notes (common on signed contracts), we use Claude's vision capability on page images alongside the OCR text, giving it both the structured text and the visual context to catch annotations.

Production Benchmarks

Tested across 10,000 documents from 14 different document types:

Document TypeAccuracyAvg. LatencyCost/Doc
Invoices98.1%3.2s$0.04
Contracts96.4%8.7s$0.12
Bank statements97.8%4.1s$0.05
Insurance claims95.9%6.3s$0.08
Medical records94.2%7.8s$0.11
Weighted average97.3%4.8s$0.06

Compared to the previous rule-based system (89% accuracy, $0.02/doc), the cost increase is offset by eliminating ~5,000 manual corrections per month.

Scaling Considerations

At 45,000 documents/month, we process roughly 1,500 per day. Key scaling decisions:

  • Batch API for non-urgent documents (overnight processing at 50% cost reduction)
  • Priority queue for time-sensitive extractions (real-time processing)
  • Caching for repeat document structures (reduces API calls by ~20%)
  • Schema versioning to handle extraction rule changes without reprocessing

Conclusion

Claude-powered data extraction replaces the fragile rules-engine approach with a system that adapts to format variations automatically. The schema-driven design means adding new document types takes hours instead of weeks. The multi-pass validation ensures production-grade accuracy. For any team processing high volumes of unstructured documents, this approach eliminates the constant parser maintenance tax and delivers accuracy that rule-based systems simply can't match.

Comments

    No comments yet. Be the first to share your thoughts.