Structured Data Extraction from Unstructured Documents with Claude
Building production data extraction pipelines that convert messy PDFs, emails, and scanned documents into structured data with 97.3% accuracy.

Every company has a pile of unstructured documents that contain critical business data: invoices, contracts, medical records, compliance filings, research papers. Traditional approaches — regex, rule-based parsers, legacy OCR — break constantly as document formats change. We built a Claude-powered extraction pipeline that handles 14 document types across 3 languages with 97.3% field-level accuracy.
This post covers the pipeline architecture, our schema-driven approach, handling edge cases at scale, and the validation framework that makes this production-ready.
The Problem
Our fintech client processes 45,000 documents per month from 200+ counterparties. Each counterparty has slightly different formats for the same document types. The existing rule-based system required constant maintenance:
- 23 custom parsers for different invoice formats
- Average 2 parser failures per day requiring engineering intervention
- 89% accuracy (meaning ~5,000 fields per month needed manual correction)
- New counterparty onboarding took 2-3 weeks to write custom parsing rules
They needed a system that could handle format variation without per-format engineering.
Architecture
The pipeline processes documents through five stages: ingestion, preprocessing, extraction, validation, and output.
Document Preprocessing
Before Claude sees any content, we normalize documents into a consistent format.
import anthropic
from pdf2image import convert_from_path
import pytesseract
from pathlib import Path
from dataclasses import dataclass, field
@dataclass
class ProcessedDocument:
doc_id: str
doc_type: str
pages: list[dict]
raw_text: str
tables: list[dict]
metadata: dict = field(default_factory=dict)
confidence_scores: dict = field(default_factory=dict)
class DocumentPreprocessor:
def process(self, file_path: Path) -> ProcessedDocument:
suffix = file_path.suffix.lower()
if suffix == '.pdf':
return self._process_pdf(file_path)
elif suffix in ('.png', '.jpg', '.jpeg', '.tiff'):
return self._process_image(file_path)
elif suffix == '.eml':
return self._process_email(file_path)
else:
raise ValueError(f"Unsupported format: {suffix}")
def _process_pdf(self, file_path: Path) -> ProcessedDocument:
# Try text extraction first (faster, more accurate for digital PDFs)
text = self._extract_text_layer(file_path)
if self._is_text_quality_sufficient(text):
pages = self._split_into_pages(text)
tables = self._extract_tables(file_path)
else:
# Fall back to OCR for scanned documents
images = convert_from_path(file_path)
pages = [
{"page_num": i + 1, "text": pytesseract.image_to_string(img)}
for i, img in enumerate(images)
]
tables = self._extract_tables_from_images(images)
return ProcessedDocument(
doc_id=self._generate_id(file_path),
doc_type="unknown", # classified in next stage
pages=pages,
raw_text="\n".join(p["text"] for p in pages),
tables=tables
)
def _is_text_quality_sufficient(self, text: str) -> bool:
"""Check if extracted text is real content vs. garbage from image PDFs"""
if len(text.strip()) < 50:
return False
# Check ratio of alphanumeric to total characters
alnum_ratio = sum(c.isalnum() for c in text) / max(len(text), 1)
return alnum_ratio > 0.6
Schema-Driven Extraction
The key innovation is defining extraction schemas declaratively and letting Claude handle the mapping from unstructured content to structured fields.
import Anthropic from '@anthropic-ai/sdk';
interface ExtractionSchema {
documentType: string;
fields: FieldDefinition[];
validationRules: ValidationRule[];
}
interface FieldDefinition {
name: string;
type: 'string' | 'number' | 'date' | 'currency' | 'array' | 'object';
required: boolean;
description: string;
examples?: string[];
constraints?: Record<string, any>;
}
interface ExtractionResult {
fields: Record<string, any>;
confidence: Record<string, number>;
warnings: string[];
rawEvidence: Record<string, string>;
}
class SchemaBasedExtractor {
private anthropic: Anthropic;
constructor() {
this.anthropic = new Anthropic();
}
async extract(
document: ProcessedDocument,
schema: ExtractionSchema
): Promise<ExtractionResult> {
const schemaPrompt = this.buildSchemaPrompt(schema);
const response = await this.anthropic.messages.create({
model: 'claude-sonnet-4-20250514',
max_tokens: 4096,
messages: [{
role: 'user',
content: `Extract structured data from this document according to the schema.
## Extraction Schema
${schemaPrompt}
## Document Content
${document.raw_text}
## Tables Found
${JSON.stringify(document.tables, null, 2)}
## Instructions
1. Extract each field defined in the schema
2. For each field, provide your confidence (0-1)
3. Quote the exact text evidence for each extraction
4. Flag any ambiguities or conflicts in the source
5. If a required field is not found, set confidence to 0 and explain why
Respond as JSON:
{
"fields": { "field_name": "extracted_value", ... },
"confidence": { "field_name": 0.95, ... },
"warnings": ["any concerns about extraction quality"],
"rawEvidence": { "field_name": "exact quote from document", ... }
}`
}]
});
const result = JSON.parse(response.content[0].text);
return this.postProcess(result, schema);
}
private buildSchemaPrompt(schema: ExtractionSchema): string {
return schema.fields.map(field => {
let desc = `- **${field.name}** (${field.type}, ${field.required ? 'required' : 'optional'}): ${field.description}`;
if (field.examples?.length) {
desc += `\n Examples: ${field.examples.join(', ')}`;
}
if (field.constraints) {
desc += `\n Constraints: ${JSON.stringify(field.constraints)}`;
}
return desc;
}).join('\n');
}
private postProcess(result: ExtractionResult, schema: ExtractionSchema): ExtractionResult {
// Apply type coercion based on schema
for (const field of schema.fields) {
if (result.fields[field.name] !== undefined) {
result.fields[field.name] = this.coerceType(
result.fields[field.name],
field.type
);
}
}
return result;
}
private coerceType(value: any, type: string): any {
switch (type) {
case 'number': return parseFloat(String(value).replace(/[,$]/g, ''));
case 'currency': return this.parseCurrency(value);
case 'date': return this.parseDate(value);
default: return value;
}
}
}
Multi-Pass Validation
Single-pass extraction isn't reliable enough for production. We run a validation pass that cross-references extracted fields.
class ExtractionValidator:
def __init__(self):
self.client = anthropic.Anthropic()
def validate(
self,
extraction: dict,
document: ProcessedDocument,
schema: dict
) -> dict:
"""Cross-validate extracted fields against each other and source."""
# Rule-based validation
rule_results = self._apply_validation_rules(extraction, schema)
# Claude-based consistency check
consistency = self.client.messages.create(
model="claude-sonnet-4-20250514",
max_tokens=2048,
messages=[{
"role": "user",
"content": f"""Review this data extraction for consistency errors.
Extracted data:
{json.dumps(extraction['fields'], indent=2)}
Source document (abbreviated):
{document.raw_text[:3000]}
Check for:
1. Mathematical consistency (do line items sum to totals?)
2. Date consistency (are dates in logical order?)
3. Reference consistency (do IDs match between sections?)
4. Format consistency (are all currencies in same denomination?)
Return JSON:
{{
"consistent": true/false,
"issues": ["list of inconsistencies found"],
"corrections": {{"field_name": "corrected_value"}}
}}"""
}]
)
consistency_result = json.loads(consistency.content[0].text)
# Merge corrections if confidence is high
if not consistency_result["consistent"]:
extraction = self._apply_corrections(
extraction,
consistency_result["corrections"]
)
return {
"extraction": extraction,
"validation": {
"rule_results": rule_results,
"consistency": consistency_result,
"final_confidence": self._calculate_overall_confidence(
extraction, rule_results, consistency_result
)
}
}
Handling Edge Cases
Multi-Language Documents
Documents containing mixed languages (common in international trade) are handled by instructing Claude to extract regardless of language and normalize to English field names:
system_prompt = """Extract data regardless of the document's language.
Field names should always be in English per the schema.
Values should be preserved in their original form (e.g., keep addresses in local script).
Currency amounts should include the currency code (EUR, USD, AED, etc.)."""
Handwritten Annotations
For documents with handwritten notes (common on signed contracts), we use Claude's vision capability on page images alongside the OCR text, giving it both the structured text and the visual context to catch annotations.
Production Benchmarks
Tested across 10,000 documents from 14 different document types:
| Document Type | Accuracy | Avg. Latency | Cost/Doc |
|---|---|---|---|
| Invoices | 98.1% | 3.2s | $0.04 |
| Contracts | 96.4% | 8.7s | $0.12 |
| Bank statements | 97.8% | 4.1s | $0.05 |
| Insurance claims | 95.9% | 6.3s | $0.08 |
| Medical records | 94.2% | 7.8s | $0.11 |
| Weighted average | 97.3% | 4.8s | $0.06 |
Compared to the previous rule-based system (89% accuracy, $0.02/doc), the cost increase is offset by eliminating ~5,000 manual corrections per month.
Scaling Considerations
At 45,000 documents/month, we process roughly 1,500 per day. Key scaling decisions:
- Batch API for non-urgent documents (overnight processing at 50% cost reduction)
- Priority queue for time-sensitive extractions (real-time processing)
- Caching for repeat document structures (reduces API calls by ~20%)
- Schema versioning to handle extraction rule changes without reprocessing
Conclusion
Claude-powered data extraction replaces the fragile rules-engine approach with a system that adapts to format variations automatically. The schema-driven design means adding new document types takes hours instead of weeks. The multi-pass validation ensures production-grade accuracy. For any team processing high volumes of unstructured documents, this approach eliminates the constant parser maintenance tax and delivers accuracy that rule-based systems simply can't match.
Recommended reading

The State of Agentic AI in 2026: Capabilities, Limitations, and Production Readiness
Comprehensive analysis of agentic AI in 2026 covering production capabilities, current limitations, and enterprise readiness benchmarks with real deployment data.

Observability for AI Agents: Tracing Multi-Step Reasoning Chains in Production
How to implement production observability for AI agents including distributed tracing, reasoning chain analysis, and debugging multi-step failures.

Measuring and Reducing AI Workload Carbon Emissions: A Practical Engineering Guide
Building a carbon-aware scheduling system for ML training and inference workloads that reduced our AI infrastructure emissions by 42% while maintaining SLA commitments.

Comments
No comments yet. Be the first to share your thoughts.