Orchestrating Multi-Agent Workflows with Claude
How we built a multi-agent system where specialized Claude agents collaborate on complex tasks, achieving 3.2x throughput improvement over single-agent approaches.

Single-agent systems hit a ceiling. When a task requires research, analysis, code generation, review, and documentation, a single Claude instance trying to do everything produces mediocre results across the board. We built a multi-agent orchestration framework where specialized Claude agents collaborate on complex tasks — each optimized for their specific role with tailored system prompts, tools, and context windows.
The result: 3.2x throughput improvement, 41% higher quality scores, and the ability to handle tasks that were previously too complex for any single LLM call.
Why Multi-Agent?
The fundamental insight is that LLMs perform better with focused roles. A Claude instance configured as a "security auditor" with security-specific context will catch more vulnerabilities than a general-purpose instance asked to "review this code for everything." Specialization improves both accuracy and efficiency.
Our use case: automated due diligence for M&A transactions. Each deal requires analyzing financial documents, legal contracts, technical architecture, market position, and compliance status — then synthesizing findings into a coherent report. No single prompt can hold all that context effectively.
Architecture
The system uses a directed acyclic graph (DAG) where each node is a specialized agent and edges represent data dependencies.
Agent Definition Framework
Each agent is defined declaratively with its role, tools, context budget, and output schema.
import Anthropic from '@anthropic-ai/sdk';
interface AgentDefinition {
id: string;
role: string;
systemPrompt: string;
model: 'claude-sonnet-4-20250514' | 'claude-haiku-4-20250514';
maxTokens: number;
tools: Anthropic.Tool[];
outputSchema: Record<string, any>;
temperature: number;
}
interface WorkflowDAG {
agents: AgentDefinition[];
edges: { from: string; to: string; dataMapping: Record<string, string> }[];
entryPoints: string[];
exitPoints: string[];
}
const dueDiligenceWorkflow: WorkflowDAG = {
agents: [
{
id: 'financial-analyst',
role: 'Financial document analysis and metric extraction',
systemPrompt: `You are a senior financial analyst conducting M&A due diligence.
Extract key financial metrics, identify trends, flag anomalies.
Focus on: revenue growth, margins, burn rate, unit economics,
customer concentration, and working capital.`,
model: 'claude-sonnet-4-20250514',
maxTokens: 8192,
tools: [/* spreadsheet tools, calculation tools */],
outputSchema: { /* financial metrics schema */ },
temperature: 0.1
},
{
id: 'legal-reviewer',
role: 'Contract and legal document analysis',
systemPrompt: `You are a corporate attorney reviewing legal documents for M&A.
Identify material risks, unusual clauses, change-of-control provisions,
IP assignment gaps, and litigation exposure.`,
model: 'claude-sonnet-4-20250514',
maxTokens: 8192,
tools: [/* document search, clause extraction */],
outputSchema: { /* legal findings schema */ },
temperature: 0.1
},
{
id: 'tech-assessor',
role: 'Technical architecture and code quality assessment',
systemPrompt: `You are a principal engineer assessing technical assets.
Evaluate architecture quality, technical debt, scalability concerns,
security posture, and team capability indicators.`,
model: 'claude-sonnet-4-20250514',
maxTokens: 8192,
tools: [/* code analysis, architecture diagramming */],
outputSchema: { /* tech assessment schema */ },
temperature: 0.2
},
{
id: 'synthesizer',
role: 'Cross-domain synthesis and recommendation',
systemPrompt: `You are a senior M&A advisor synthesizing findings from
financial, legal, and technical analyses. Produce a unified risk
assessment and recommendation with clear go/no-go signals.`,
model: 'claude-sonnet-4-20250514',
maxTokens: 16384,
tools: [],
outputSchema: { /* final report schema */ },
temperature: 0.3
}
],
edges: [
{ from: 'financial-analyst', to: 'synthesizer', dataMapping: { 'financialFindings': 'financials' } },
{ from: 'legal-reviewer', to: 'synthesizer', dataMapping: { 'legalFindings': 'legal' } },
{ from: 'tech-assessor', to: 'synthesizer', dataMapping: { 'techFindings': 'technical' } }
],
entryPoints: ['financial-analyst', 'legal-reviewer', 'tech-assessor'],
exitPoints: ['synthesizer']
};
Orchestration Engine
The orchestrator manages agent execution, data flow, error handling, and resource allocation.
import anthropic
import asyncio
from dataclasses import dataclass
from typing import Any
import networkx as nx
@dataclass
class AgentExecution:
agent_id: str
status: str # "pending", "running", "completed", "failed", "retrying"
input_data: dict
output_data: dict | None
attempts: int
latency_ms: float | None
token_usage: dict | None
class WorkflowOrchestrator:
def __init__(self, workflow: dict):
self.workflow = workflow
self.client = anthropic.Anthropic()
self.dag = self._build_dag(workflow)
self.executions: dict[str, AgentExecution] = {}
async def execute(self, initial_data: dict) -> dict:
"""Execute the full workflow DAG."""
# Identify agents with no dependencies (entry points)
entry_agents = [n for n in self.dag.nodes if self.dag.in_degree(n) == 0]
# Execute in topological order with parallel execution where possible
for generation in nx.topological_generations(self.dag):
tasks = []
for agent_id in generation:
input_data = self._gather_inputs(agent_id, initial_data)
tasks.append(self._execute_agent(agent_id, input_data))
# Run all agents in this generation in parallel
results = await asyncio.gather(*tasks, return_exceptions=True)
for agent_id, result in zip(generation, results):
if isinstance(result, Exception):
await self._handle_failure(agent_id, result)
else:
self.executions[agent_id].output_data = result
self.executions[agent_id].status = "completed"
# Collect outputs from exit points
return self._collect_final_output()
async def _execute_agent(self, agent_id: str, input_data: dict) -> dict:
"""Execute a single agent with retry logic."""
agent_def = self._get_agent_definition(agent_id)
self.executions[agent_id] = AgentExecution(
agent_id=agent_id, status="running",
input_data=input_data, output_data=None,
attempts=0, latency_ms=None, token_usage=None
)
for attempt in range(3):
try:
self.executions[agent_id].attempts = attempt + 1
start = asyncio.get_event_loop().time()
response = self.client.messages.create(
model=agent_def["model"],
max_tokens=agent_def["maxTokens"],
system=agent_def["systemPrompt"],
tools=agent_def.get("tools", []),
temperature=agent_def.get("temperature", 0.2),
messages=[{
"role": "user",
"content": self._format_agent_input(agent_def, input_data)
}]
)
elapsed = (asyncio.get_event_loop().time() - start) * 1000
self.executions[agent_id].latency_ms = elapsed
self.executions[agent_id].token_usage = {
"input": response.usage.input_tokens,
"output": response.usage.output_tokens
}
return self._parse_agent_output(response, agent_def["outputSchema"])
except anthropic.RateLimitError:
await asyncio.sleep(2 ** attempt)
except Exception as e:
if attempt == 2:
raise
def _gather_inputs(self, agent_id: str, initial_data: dict) -> dict:
"""Gather inputs from upstream agents and initial data."""
inputs = {"initial": initial_data}
for pred in self.dag.predecessors(agent_id):
if self.executions.get(pred) and self.executions[pred].output_data:
edge_data = self.dag.edges[pred, agent_id]
for source_key, target_key in edge_data["dataMapping"].items():
inputs[target_key] = self.executions[pred].output_data.get(source_key)
return inputs
Inter-Agent Communication
Agents don't communicate directly — all data flows through the orchestrator. This provides:
- Observability — Every message between agents is logged and traceable
- Schema validation — Outputs are validated against schemas before being passed downstream
- Circuit breaking — If an upstream agent fails, downstream agents receive partial data with failure context
- Cost control — The orchestrator can short-circuit workflows when costs exceed budgets
Benchmarks
Comparing single-agent vs. multi-agent on the same due diligence task (50 test cases):
| Metric | Single Agent | Multi-Agent | Improvement |
|---|---|---|---|
| Quality score (human eval) | 6.2/10 | 8.7/10 | +41% |
| Throughput (docs/hour) | 12 | 38 | 3.2x |
| Avg. completion time | 14 min | 4.3 min | 69% faster |
| Missed findings rate | 23% | 7% | 70% reduction |
| Cost per analysis | $4.80 | $6.20 | +29% |
The multi-agent approach costs 29% more per analysis but produces significantly higher quality output. For due diligence where missed risks have massive consequences, the quality improvement justifies the cost.
Failure Handling Patterns
Graceful Degradation
When one agent fails, the synthesizer receives partial results with explicit gaps noted:
partial_input = {
"financials": financial_results, # success
"legal": {"status": "failed", "reason": "timeout", "partial": partial_legal},
"technical": tech_results # success
}
The synthesizer is instructed to produce the best report possible with available data and clearly mark sections where analysis is incomplete.
Agent Disagreement
When agents produce conflicting findings (e.g., financial data suggesting healthy growth but technical analysis revealing unsustainable scaling), the synthesizer is specifically instructed to surface contradictions rather than resolve them arbitrarily.
Cost Governance
Multi-agent systems can spiral in cost if not governed. Our controls:
- Per-workflow budget limits with early termination
- Agent-level token budgets (hard caps on input/output)
- Model routing: Haiku for simple extraction tasks, Sonnet for analysis
- Caching: repeated document chunks are cached to avoid re-processing
Conclusion
Multi-agent workflows with Claude transform complex analytical tasks from "one giant prompt that does everything poorly" to "specialized experts collaborating efficiently." The key architectural decisions are: define agents with focused roles, use a DAG for data flow, validate outputs at every boundary, and build robust failure handling. The 29% cost increase pays for itself many times over in quality and reliability for high-stakes applications.
Recommended reading

The State of Agentic AI in 2026: Capabilities, Limitations, and Production Readiness
Comprehensive analysis of agentic AI in 2026 covering production capabilities, current limitations, and enterprise readiness benchmarks with real deployment data.

Observability for AI Agents: Tracing Multi-Step Reasoning Chains in Production
How to implement production observability for AI agents including distributed tracing, reasoning chain analysis, and debugging multi-step failures.

Measuring and Reducing AI Workload Carbon Emissions: A Practical Engineering Guide
Building a carbon-aware scheduling system for ML training and inference workloads that reduced our AI infrastructure emissions by 42% while maintaining SLA commitments.

Comments
No comments yet. Be the first to share your thoughts.