Service Discovery at Scale: HashiCorp Consul vs AWS Cloud Map in Production
A head-to-head comparison of Consul and AWS Cloud Map for service discovery at 2,000+ services, covering performance, operational overhead, and multi-cloud considerations.

Service discovery is the nervous system of a microservices architecture. At 2,000+ services with 15,000+ instances across multiple environments, the choice between a self-managed solution like HashiCorp Consul and a cloud-native managed service like AWS Cloud Map has significant implications for operational overhead, performance, and architectural flexibility.
We have operated both in production at scale. This article is not a feature comparison matrix — it is an operational comparison grounded in real production data, failure modes we have experienced, and the hidden costs that only surface at scale.
The Context
Our platform runs 2,147 distinct services across production, staging, and development environments. In production alone, we manage 8,400+ service instances that register, deregister, and health-check continuously. Service-to-service calls generate approximately 2.1 million DNS lookups per minute.
We ran Consul for three years before evaluating Cloud Map as a potential simplification. After running both in parallel for six months, we made a nuanced decision: neither won outright. The right choice depends on your specific constraints.
Architecture Comparison
Consul Architecture
Consul runs as a distributed system with server nodes maintaining the service catalog through Raft consensus and client agents on every compute node handling local registration and health checks.
# Consul server configuration for production cluster
# 5 server nodes across 3 AZs for quorum safety
datacenter = "us-east-1"
data_dir = "/opt/consul/data"
log_level = "WARN"
server = true
bootstrap_expect = 5
# Performance tuning for 2000+ services
performance {
raft_multiplier = 1
}
# Enable gRPC for xDS (Envoy integration)
ports {
grpc = 8502
grpc_tls = 8503
}
# Connect (service mesh) configuration
connect {
enabled = true
ca_provider = "vault"
ca_config {
address = "https://vault.internal:8200"
root_pki_path = "pki-root"
intermediate_pki_path = "pki-intermediate"
}
}
# DNS configuration for service discovery
dns_config {
allow_stale = true
max_stale = "30s"
node_ttl = "10s"
service_ttl {
"*" = "5s"
}
# Critical: enable caching for performance at scale
use_cache = true
cache_max_age = "5s"
}
# Telemetry for monitoring
telemetry {
prometheus_retention_time = "60s"
disable_hostname = true
}
# ACL configuration
acl {
enabled = true
default_policy = "deny"
enable_token_persistence = true
}
AWS Cloud Map Architecture
Cloud Map is a fully managed service — no servers to provision, no quorum to maintain, no client agents to deploy. Services register via API calls, and discovery happens through DNS (Route 53 auto-configured) or the Cloud Map API directly.
// Cloud Map service registration with health checking
import {
ServiceDiscoveryClient,
CreateServiceCommand,
RegisterInstanceCommand,
HealthStatusFilter
} from '@aws-sdk/client-servicediscovery';
const client = new ServiceDiscoveryClient({ region: 'us-east-1' });
// Create a service definition with health checking
async function createService(namespace: string, serviceName: string) {
const command = new CreateServiceCommand({
Name: serviceName,
NamespaceId: namespace,
DnsConfig: {
RoutingPolicy: 'MULTIVALUE',
DnsRecords: [
{ Type: 'A', TTL: 10 },
{ Type: 'SRV', TTL: 10 }
]
},
HealthCheckCustomConfig: {
FailureThreshold: 2
},
Tags: [
{ Key: 'environment', Value: 'production' },
{ Key: 'team', Value: 'payments' }
]
});
return client.send(command);
}
// Register a service instance
async function registerInstance(
serviceId: string,
instanceId: string,
host: string,
port: number,
metadata: Record<string, string>
) {
const command = new RegisterInstanceCommand({
ServiceId: serviceId,
InstanceId: instanceId,
Attributes: {
AWS_INSTANCE_IPV4: host,
AWS_INSTANCE_PORT: port.toString(),
// Custom attributes for routing metadata
version: metadata.version,
region: metadata.region,
canary: metadata.canary || 'false',
...metadata
}
});
return client.send(command);
}
Performance Benchmarks
We ran both systems in parallel for six months, measuring discovery latency, registration speed, and behavior under failure conditions.
Discovery Latency
| Metric | Consul (DNS) | Consul (HTTP API) | Cloud Map (DNS) | Cloud Map (API) |
|---|---|---|---|---|
| Lookup p50 | 0.8ms | 2.1ms | 1.2ms | 4.8ms |
| Lookup p95 | 2.4ms | 5.6ms | 3.1ms | 12.4ms |
| Lookup p99 | 8.2ms | 14.3ms | 6.8ms | 28.6ms |
| Cache hit rate | 94% | N/A | 89% | N/A |
| Lookups/sec capacity | 180K | 45K | 120K (Route 53 limit) | 10K (API throttle) |
Consul's DNS interface wins on raw latency because the client agent maintains a local cache. Cloud Map DNS goes through Route 53, which adds a hop but eliminates client-side infrastructure.
Registration and Deregistration Speed
| Operation | Consul | Cloud Map |
|---|---|---|
| Instance registration | 12ms (local agent) | 180ms (API call) |
| Health check propagation | 2-5s (gossip) | 30-60s (polling) |
| Deregistration on failure | 15-30s (configurable) | 60-120s (custom health) |
| Catalog convergence (2000 services) | 8s | 45s |
This is where the operational difference becomes critical. Consul's gossip protocol propagates changes in seconds. Cloud Map's polling-based health checks can leave stale entries for up to 2 minutes, meaning consumers may route to dead instances.
Operational Overhead
Consul Operational Costs
| Operational Task | Frequency | Team Hours/Month |
|---|---|---|
| Server node upgrades | Monthly | 8 hours |
| ACL token rotation | Quarterly | 4 hours |
| Capacity planning | Monthly | 2 hours |
| Incident response (Consul-specific) | ~2/month | 6 hours |
| Certificate rotation | Monthly | 2 hours |
| Gossip encryption key rotation | Quarterly | 2 hours |
| Total | ~24 hours/month |
Cloud Map Operational Costs
| Operational Task | Frequency | Team Hours/Month |
|---|---|---|
| Namespace management | As needed | 1 hour |
| Quota increase requests | Quarterly | 0.5 hours |
| Stale instance cleanup | Weekly | 2 hours |
| Integration testing | Per deployment | 3 hours |
| Total | ~6.5 hours/month |
Cloud Map requires roughly 75% less operational overhead. But that number hides a critical caveat: when Cloud Map has issues, you have no control plane access to debug. With Consul, you can inspect gossip state, query the Raft log, and understand exactly what went wrong.
Failure Modes We Have Experienced
Consul Failures
-
Raft leader election storm — A network partition isolated 2 of 5 server nodes, causing repeated leader elections. Service discovery continued (stale reads enabled) but write operations (registration) failed for 4 minutes.
-
Gossip pool fragmentation — A misconfigured firewall rule prevented gossip between two AZs. Services in the isolated AZ appeared healthy locally but were invisible to the rest of the cluster. Detection took 12 minutes because health checks were passing locally.
-
Agent memory leak — Consul client agent accumulated connection state under high churn (150+ registrations/min per node). Required rolling restart of all agents.
Cloud Map Failures
-
Route 53 propagation delay — DNS records took 90+ seconds to propagate during a high-registration event (deployment of 400 instances). Consumers hit NXDOMAIN for new services.
-
API throttling during scale event — Auto-scaling event triggered 200 simultaneous registrations, hitting the Cloud Map API rate limit. Registration retries with backoff resolved after 3 minutes, but services were discoverable with delay.
-
Health check false positive — Cloud Map's custom health check API experienced an internal error that marked all instances as unhealthy for 45 seconds. This was an AWS service event — we had zero remediation capability.
Decision Framework
After six months of parallel operation, here is our decision matrix:
Choose Consul when:
- You need sub-second health check propagation
- You operate across multiple clouds or on-premises
- You need service mesh (Connect) integrated with discovery
- You require fine-grained ACL policies on service visibility
- Your team has the operational capacity for distributed systems management
Choose Cloud Map when:
- You are AWS-only and plan to stay that way
- Operational simplicity is more valuable than control
- Your services can tolerate 30-60 second health check propagation
- You have fewer than 500 services
- You want deep integration with ECS, EKS, and Lambda
Our decision: We kept Consul for production workloads requiring sub-second propagation and multi-cloud support, and adopted Cloud Map for development/staging environments and AWS-native services where the operational simplification outweighed the propagation delay.
Conclusion
There is no universal answer to the Consul vs Cloud Map debate. The right choice is a function of your scale, your cloud strategy, your team's operational maturity, and your tolerance for propagation delay. At 2,000+ services, we found that a hybrid approach — Consul for the critical path, Cloud Map for everything else — gives us the best balance of performance and operational efficiency.
Whatever you choose, invest heavily in monitoring your service discovery layer. It is the one piece of infrastructure where silent failure cascades into every service simultaneously.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.