Service Discovery at Scale: HashiCorp Consul vs AWS Cloud Map in Production

A head-to-head comparison of Consul and AWS Cloud Map for service discovery at 2,000+ services, covering performance, operational overhead, and multi-cloud considerations.

#service-discovery#consul#cloud-map#microservices
Cover image for the article: Service Discovery at Scale: HashiCorp Consul vs AWS Cloud Map in Production

Service discovery is the nervous system of a microservices architecture. At 2,000+ services with 15,000+ instances across multiple environments, the choice between a self-managed solution like HashiCorp Consul and a cloud-native managed service like AWS Cloud Map has significant implications for operational overhead, performance, and architectural flexibility.

We have operated both in production at scale. This article is not a feature comparison matrix — it is an operational comparison grounded in real production data, failure modes we have experienced, and the hidden costs that only surface at scale.

The Context

Our platform runs 2,147 distinct services across production, staging, and development environments. In production alone, we manage 8,400+ service instances that register, deregister, and health-check continuously. Service-to-service calls generate approximately 2.1 million DNS lookups per minute.

We ran Consul for three years before evaluating Cloud Map as a potential simplification. After running both in parallel for six months, we made a nuanced decision: neither won outright. The right choice depends on your specific constraints.

Architecture Comparison

Consul Architecture

Consul runs as a distributed system with server nodes maintaining the service catalog through Raft consensus and client agents on every compute node handling local registration and health checks.

Consul Architecture

# Consul server configuration for production cluster
# 5 server nodes across 3 AZs for quorum safety

datacenter = "us-east-1"
data_dir   = "/opt/consul/data"
log_level  = "WARN"

server           = true
bootstrap_expect = 5

# Performance tuning for 2000+ services
performance {
  raft_multiplier = 1
}

# Enable gRPC for xDS (Envoy integration)
ports {
  grpc     = 8502
  grpc_tls = 8503
}

# Connect (service mesh) configuration
connect {
  enabled = true
  ca_provider = "vault"
  ca_config {
    address = "https://vault.internal:8200"
    root_pki_path = "pki-root"
    intermediate_pki_path = "pki-intermediate"
  }
}

# DNS configuration for service discovery
dns_config {
  allow_stale  = true
  max_stale    = "30s"
  node_ttl     = "10s"
  service_ttl {
    "*" = "5s"
  }
  # Critical: enable caching for performance at scale
  use_cache   = true
  cache_max_age = "5s"
}

# Telemetry for monitoring
telemetry {
  prometheus_retention_time = "60s"
  disable_hostname = true
}

# ACL configuration
acl {
  enabled        = true
  default_policy = "deny"
  enable_token_persistence = true
}

AWS Cloud Map Architecture

Cloud Map is a fully managed service — no servers to provision, no quorum to maintain, no client agents to deploy. Services register via API calls, and discovery happens through DNS (Route 53 auto-configured) or the Cloud Map API directly.

// Cloud Map service registration with health checking
import { 
  ServiceDiscoveryClient, 
  CreateServiceCommand,
  RegisterInstanceCommand,
  HealthStatusFilter 
} from '@aws-sdk/client-servicediscovery';

const client = new ServiceDiscoveryClient({ region: 'us-east-1' });

// Create a service definition with health checking
async function createService(namespace: string, serviceName: string) {
  const command = new CreateServiceCommand({
    Name: serviceName,
    NamespaceId: namespace,
    DnsConfig: {
      RoutingPolicy: 'MULTIVALUE',
      DnsRecords: [
        { Type: 'A', TTL: 10 },
        { Type: 'SRV', TTL: 10 }
      ]
    },
    HealthCheckCustomConfig: {
      FailureThreshold: 2
    },
    Tags: [
      { Key: 'environment', Value: 'production' },
      { Key: 'team', Value: 'payments' }
    ]
  });
  
  return client.send(command);
}

// Register a service instance
async function registerInstance(
  serviceId: string, 
  instanceId: string, 
  host: string, 
  port: number,
  metadata: Record<string, string>
) {
  const command = new RegisterInstanceCommand({
    ServiceId: serviceId,
    InstanceId: instanceId,
    Attributes: {
      AWS_INSTANCE_IPV4: host,
      AWS_INSTANCE_PORT: port.toString(),
      // Custom attributes for routing metadata
      version: metadata.version,
      region: metadata.region,
      canary: metadata.canary || 'false',
      ...metadata
    }
  });

  return client.send(command);
}

Performance Benchmarks

We ran both systems in parallel for six months, measuring discovery latency, registration speed, and behavior under failure conditions.

Discovery Latency

MetricConsul (DNS)Consul (HTTP API)Cloud Map (DNS)Cloud Map (API)
Lookup p500.8ms2.1ms1.2ms4.8ms
Lookup p952.4ms5.6ms3.1ms12.4ms
Lookup p998.2ms14.3ms6.8ms28.6ms
Cache hit rate94%N/A89%N/A
Lookups/sec capacity180K45K120K (Route 53 limit)10K (API throttle)

Consul's DNS interface wins on raw latency because the client agent maintains a local cache. Cloud Map DNS goes through Route 53, which adds a hop but eliminates client-side infrastructure.

Registration and Deregistration Speed

OperationConsulCloud Map
Instance registration12ms (local agent)180ms (API call)
Health check propagation2-5s (gossip)30-60s (polling)
Deregistration on failure15-30s (configurable)60-120s (custom health)
Catalog convergence (2000 services)8s45s

This is where the operational difference becomes critical. Consul's gossip protocol propagates changes in seconds. Cloud Map's polling-based health checks can leave stale entries for up to 2 minutes, meaning consumers may route to dead instances.

Operational Overhead

Consul Operational Costs

Operational TaskFrequencyTeam Hours/Month
Server node upgradesMonthly8 hours
ACL token rotationQuarterly4 hours
Capacity planningMonthly2 hours
Incident response (Consul-specific)~2/month6 hours
Certificate rotationMonthly2 hours
Gossip encryption key rotationQuarterly2 hours
Total~24 hours/month

Cloud Map Operational Costs

Operational TaskFrequencyTeam Hours/Month
Namespace managementAs needed1 hour
Quota increase requestsQuarterly0.5 hours
Stale instance cleanupWeekly2 hours
Integration testingPer deployment3 hours
Total~6.5 hours/month

Cloud Map requires roughly 75% less operational overhead. But that number hides a critical caveat: when Cloud Map has issues, you have no control plane access to debug. With Consul, you can inspect gossip state, query the Raft log, and understand exactly what went wrong.

Failure Modes We Have Experienced

Consul Failures

  1. Raft leader election storm — A network partition isolated 2 of 5 server nodes, causing repeated leader elections. Service discovery continued (stale reads enabled) but write operations (registration) failed for 4 minutes.

  2. Gossip pool fragmentation — A misconfigured firewall rule prevented gossip between two AZs. Services in the isolated AZ appeared healthy locally but were invisible to the rest of the cluster. Detection took 12 minutes because health checks were passing locally.

  3. Agent memory leak — Consul client agent accumulated connection state under high churn (150+ registrations/min per node). Required rolling restart of all agents.

Cloud Map Failures

  1. Route 53 propagation delay — DNS records took 90+ seconds to propagate during a high-registration event (deployment of 400 instances). Consumers hit NXDOMAIN for new services.

  2. API throttling during scale event — Auto-scaling event triggered 200 simultaneous registrations, hitting the Cloud Map API rate limit. Registration retries with backoff resolved after 3 minutes, but services were discoverable with delay.

  3. Health check false positive — Cloud Map's custom health check API experienced an internal error that marked all instances as unhealthy for 45 seconds. This was an AWS service event — we had zero remediation capability.

Decision Framework

After six months of parallel operation, here is our decision matrix:

Choose Consul when:

  • You need sub-second health check propagation
  • You operate across multiple clouds or on-premises
  • You need service mesh (Connect) integrated with discovery
  • You require fine-grained ACL policies on service visibility
  • Your team has the operational capacity for distributed systems management

Choose Cloud Map when:

  • You are AWS-only and plan to stay that way
  • Operational simplicity is more valuable than control
  • Your services can tolerate 30-60 second health check propagation
  • You have fewer than 500 services
  • You want deep integration with ECS, EKS, and Lambda

Our decision: We kept Consul for production workloads requiring sub-second propagation and multi-cloud support, and adopted Cloud Map for development/staging environments and AWS-native services where the operational simplification outweighed the propagation delay.

Conclusion

There is no universal answer to the Consul vs Cloud Map debate. The right choice is a function of your scale, your cloud strategy, your team's operational maturity, and your tolerance for propagation delay. At 2,000+ services, we found that a hybrid approach — Consul for the critical path, Cloud Map for everything else — gives us the best balance of performance and operational efficiency.

Whatever you choose, invest heavily in monitoring your service discovery layer. It is the one piece of infrastructure where silent failure cascades into every service simultaneously.

Comments

    No comments yet. Be the first to share your thoughts.