Processing 50K Documents Per Day with Multimodal AI
Production architecture for high-throughput document understanding using multimodal AI models, achieving 50K documents/day with 96% extraction accuracy.

Document understanding at scale is one of the highest-ROI applications of multimodal AI. Invoices, contracts, medical records, insurance claims — every enterprise has thousands of documents that need structured data extraction. After building a pipeline that processes 50K+ documents daily with 96% field-level accuracy, here's the architecture that makes it work.
The Problem: Documents Are Messy
Traditional OCR + rule-based extraction breaks down because real-world documents are chaotic:
- Layout variance: The same document type from different vendors has completely different layouts
- Poor scan quality: Skewed pages, coffee stains, low resolution faxes
- Mixed content: Tables, handwritten annotations, stamps, logos, and text interleaved
- Multi-page context: Key information spans multiple pages with cross-references
- Language mixing: Headers in English, body in Arabic, amounts in both
Rule-based systems require per-template maintenance. With 200+ document types and continuous layout changes, the maintenance burden consumed 3 FTEs. Multimodal AI eliminates template maintenance entirely.
Architecture: Pipeline for Scale
The document processing pipeline separates ingestion, preprocessing, understanding, and validation into independent stages that scale horizontally.
Document Preprocessing and Chunking
Before multimodal models see a document, preprocessing ensures consistent quality and splits multi-page documents into processable units.
import asyncio
from dataclasses import dataclass, field
from enum import Enum
from typing import Optional
from pathlib import Path
import io
class DocumentType(Enum):
INVOICE = "invoice"
CONTRACT = "contract"
RECEIPT = "receipt"
MEDICAL_RECORD = "medical_record"
INSURANCE_CLAIM = "insurance_claim"
UNKNOWN = "unknown"
@dataclass
class PageImage:
page_number: int
image_bytes: bytes
width: int
height: int
dpi: int
quality_score: float
@dataclass
class ProcessedDocument:
document_id: str
source_path: str
document_type: DocumentType
pages: list[PageImage]
total_pages: int
metadata: dict = field(default_factory=dict)
class DocumentPreprocessor:
TARGET_DPI = 300
MAX_DIMENSION = 4096
MIN_QUALITY_SCORE = 0.6
def __init__(self, classifier_model, quality_assessor):
self.classifier = classifier_model
self.quality = quality_assessor
async def preprocess(self, document_path: str, document_id: str) -> ProcessedDocument:
raw_pages = await self._extract_pages(document_path)
processed_pages = []
for page in raw_pages:
enhanced = await self._enhance_page(page)
quality_score = self.quality.assess(enhanced.image_bytes)
if quality_score < self.MIN_QUALITY_SCORE:
enhanced = await self._aggressive_enhance(enhanced)
quality_score = self.quality.assess(enhanced.image_bytes)
enhanced.quality_score = quality_score
processed_pages.append(enhanced)
# Classify document type from first page
doc_type = await self._classify_document(processed_pages[0])
return ProcessedDocument(
document_id=document_id,
source_path=document_path,
document_type=doc_type,
pages=processed_pages,
total_pages=len(processed_pages),
)
async def _extract_pages(self, path: str) -> list[PageImage]:
# Convert PDF/TIFF/image to individual page images at target DPI
pages = []
# Implementation uses pdf2image or similar
return pages
async def _enhance_page(self, page: PageImage) -> PageImage:
# Deskew, denoise, contrast normalization, resolution upscaling
return page
async def _aggressive_enhance(self, page: PageImage) -> PageImage:
# Super-resolution + heavy denoising for very poor quality
return page
async def _classify_document(self, first_page: PageImage) -> DocumentType:
prediction = self.classifier.predict(first_page.image_bytes)
return DocumentType(prediction) if prediction in DocumentType._value2member_map_ else DocumentType.UNKNOWN
@dataclass
class ExtractionField:
name: str
value: str
confidence: float
bounding_box: Optional[tuple[int, int, int, int]]
page_number: int
normalized_value: Optional[str] = None
@dataclass
class ExtractionResult:
document_id: str
document_type: DocumentType
fields: list[ExtractionField]
tables: list[dict]
raw_text: str
processing_time_ms: float
model_version: str
class MultimodalExtractor:
def __init__(self, model_client, extraction_schemas: dict[DocumentType, dict]):
self.model = model_client
self.schemas = extraction_schemas
async def extract(self, document: ProcessedDocument) -> ExtractionResult:
import time
start = time.time()
schema = self.schemas.get(document.document_type, self.schemas[DocumentType.UNKNOWN])
# Build multimodal prompt with page images and extraction schema
prompt = self._build_extraction_prompt(schema, document)
# Process pages in groups for multi-page context
page_groups = self._group_pages(document.pages, max_group_size=3)
all_fields = []
all_tables = []
for group in page_groups:
images = [p.image_bytes for p in group]
response = await self.model.generate(
prompt=prompt,
images=images,
max_tokens=4096,
temperature=0.1,
)
fields, tables = self._parse_extraction_response(response, group)
all_fields.extend(fields)
all_tables.extend(tables)
# Deduplicate and merge fields from overlapping page groups
merged_fields = self._merge_fields(all_fields)
return ExtractionResult(
document_id=document.document_id,
document_type=document.document_type,
fields=merged_fields,
tables=all_tables,
raw_text="",
processing_time_ms=(time.time() - start) * 1000,
model_version=self.model.version,
)
def _build_extraction_prompt(self, schema: dict, document: ProcessedDocument) -> str:
field_descriptions = "\n".join(
f"- {name}: {desc}" for name, desc in schema.items()
)
return f"""Extract the following fields from this {document.document_type.value} document.
Return a JSON object with field names as keys. Include confidence scores.
Fields to extract:
{field_descriptions}
For tables, return as arrays of objects with column headers as keys.
If a field is not found, set it to null with confidence 0.
"""
def _group_pages(self, pages: list[PageImage], max_group_size: int) -> list[list[PageImage]]:
groups = []
for i in range(0, len(pages), max_group_size - 1):
groups.append(pages[i:i + max_group_size])
return groups
def _parse_extraction_response(self, response: str, pages: list[PageImage]) -> tuple:
return [], []
def _merge_fields(self, fields: list[ExtractionField]) -> list[ExtractionField]:
seen = {}
for field in fields:
if field.name not in seen or field.confidence > seen[field.name].confidence:
seen[field.name] = field
return list(seen.values())
Validation and Human-in-the-Loop
Extraction results below confidence thresholds route to human review. This creates a feedback loop that continuously improves model performance.
interface ValidationRule {
fieldName: string;
type: 'regex' | 'range' | 'cross_field' | 'lookup';
rule: string | { min: number; max: number } | { field: string; relation: string };
severity: 'error' | 'warning';
}
interface ValidationResult {
documentId: string;
passed: boolean;
fieldResults: Array<{
fieldName: string;
value: string;
confidence: number;
valid: boolean;
validationErrors: string[];
}>;
needsHumanReview: boolean;
reviewReason?: string;
}
class ExtractionValidator {
private rules: Map<string, ValidationRule[]>;
private confidenceThreshold: number;
constructor(rules: ValidationRule[], confidenceThreshold = 0.85) {
this.rules = new Map();
this.confidenceThreshold = confidenceThreshold;
for (const rule of rules) {
const existing = this.rules.get(rule.fieldName) || [];
existing.push(rule);
this.rules.set(rule.fieldName, existing);
}
}
validate(extraction: { documentId: string; fields: Array<{ name: string; value: string; confidence: number }> }): ValidationResult {
const fieldResults = extraction.fields.map(field => {
const rules = this.rules.get(field.name) || [];
const errors: string[] = [];
for (const rule of rules) {
if (!this.checkRule(field.value, rule)) {
errors.push(`Failed ${rule.type} validation: ${JSON.stringify(rule.rule)}`);
}
}
return {
fieldName: field.name,
value: field.value,
confidence: field.confidence,
valid: errors.length === 0,
validationErrors: errors,
};
});
const lowConfidenceFields = fieldResults.filter(f => f.confidence < this.confidenceThreshold);
const invalidFields = fieldResults.filter(f => !f.valid);
const needsReview = lowConfidenceFields.length > 0 || invalidFields.length > 0;
let reviewReason: string | undefined;
if (invalidFields.length > 0) {
reviewReason = `Validation failures: ${invalidFields.map(f => f.fieldName).join(', ')}`;
} else if (lowConfidenceFields.length > 0) {
reviewReason = `Low confidence: ${lowConfidenceFields.map(f => `${f.fieldName}(${f.confidence.toFixed(2)})`).join(', ')}`;
}
return {
documentId: extraction.documentId,
passed: !needsReview,
fieldResults,
needsHumanReview: needsReview,
reviewReason,
};
}
private checkRule(value: string, rule: ValidationRule): boolean {
switch (rule.type) {
case 'regex':
return new RegExp(rule.rule as string).test(value);
case 'range': {
const num = parseFloat(value);
const range = rule.rule as { min: number; max: number };
return !isNaN(num) && num >= range.min && num <= range.max;
}
default:
return true;
}
}
}
Scaling to 50K Documents/Day
Throughput Architecture
- Ingestion: S3 event triggers → SQS FIFO queue (ordered per document batch)
- Processing: ECS Fargate tasks with GPU, auto-scaling 5-50 tasks based on queue depth
- Model serving: Dedicated GPU inference cluster with batched requests
- Storage: Extracted data → DynamoDB, original documents → S3 Glacier after 30 days
Parallelism Strategy
Each document is independent, so horizontal scaling is straightforward. The bottleneck is GPU inference — one A100 processes approximately 3 pages/second with the full extraction pipeline. With 8 GPU instances, we achieve ~2,000 pages/minute sustained.
Benchmarks: Production Performance
| Metric | Value |
|---|---|
| Daily throughput | 52K documents |
| Pages processed/day | 184K |
| Field extraction accuracy | 96.2% |
| Table extraction accuracy | 91.8% |
| Processing latency (median) | 4.2 sec/document |
| Processing latency (p99) | 18 sec/document |
| Human review rate | 7.3% |
| Cost per document | $0.023 |
Accuracy by Document Type
| Document Type | Field Accuracy | Table Accuracy |
|---|---|---|
| Invoices | 97.8% | 94.2% |
| Contracts | 93.1% | 88.5% |
| Receipts | 96.5% | N/A |
| Medical records | 94.7% | 90.1% |
| Insurance claims | 96.9% | 92.3% |
Lessons Learned
Image quality dominates accuracy. Investing in preprocessing (deskew, denoise, resolution enhancement) improved extraction accuracy by 8% across all document types. The model is surprisingly robust to layout changes but sensitive to image artifacts.
Few-shot examples beat long prompts. Including 2-3 example extractions for each document type in the prompt outperforms detailed extraction instructions. The model generalizes better from examples.
Confidence calibration requires effort. Raw model confidence scores are poorly calibrated — a reported 0.9 confidence might only be correct 75% of the time. Calibrate confidence against actual accuracy using a held-out validation set.
Batch similar documents together. Processing invoices together allows the model to leverage shared context and vocabulary. Mixed batches show 3-5% lower accuracy.
Conclusion
Multimodal AI transforms document understanding from a template-maintenance nightmare into a scalable, accurate pipeline. The combination of quality preprocessing, schema-driven extraction prompts, confidence-based human review routing, and horizontal scaling achieves the throughput and accuracy needed for enterprise workloads. Start with your highest-volume document type, validate against manual extraction, and expand document types as confidence in the system grows.
Recommended reading

The State of Agentic AI in 2026: Capabilities, Limitations, and Production Readiness
Comprehensive analysis of agentic AI in 2026 covering production capabilities, current limitations, and enterprise readiness benchmarks with real deployment data.

Observability for AI Agents: Tracing Multi-Step Reasoning Chains in Production
How to implement production observability for AI agents including distributed tracing, reasoning chain analysis, and debugging multi-step failures.

Measuring and Reducing AI Workload Carbon Emissions: A Practical Engineering Guide
Building a carbon-aware scheduling system for ML training and inference workloads that reduced our AI infrastructure emissions by 42% while maintaining SLA commitments.

Comments
No comments yet. Be the first to share your thoughts.