Distributed Tracing Across 47 Microservices with OpenTelemetry

How we implemented end-to-end distributed tracing with OpenTelemetry, reducing debugging time from hours to minutes across our service mesh

#observability#tracing#opentelemetry#microservices
Cover image for the article: Distributed Tracing Across 47 Microservices with OpenTelemetry

When a customer reports a slow checkout, and that request touches 47 microservices across three AWS regions, finding the bottleneck without distributed tracing is guesswork. We migrated from a patchwork of vendor-specific APM tools to a unified OpenTelemetry implementation that provides complete request lineage from edge to database. Here is how we instrumented 47 services in 12 weeks.

The Problem: Observability Fragmentation

Our observability stack had grown organically. The payments team used Datadog, the order team used New Relic, the inventory team used custom Prometheus exporters, and the platform team ran Jaeger for infrastructure traces. When a cross-team incident occurred, correlating signals required manual timestamp matching across four different UIs.

Pain points before OpenTelemetry:

  • Average incident debugging time: 2.4 hours for cross-service issues
  • 4 different tracing vendors with no context propagation between them
  • 23% of traces were incomplete due to missing propagation headers
  • No way to correlate a single customer request end-to-end

Architecture Design

We chose the OpenTelemetry Collector as our central pipeline, deployed as both a sidecar (for high-throughput services) and a gateway (for aggregation and export).

OpenTelemetry Architecture

The architecture uses a two-tier collector deployment: sidecars handle local batching and initial processing, while gateway collectors perform tail-based sampling, service graph computation, and multi-backend export.

Instrumentation Strategy

We wrote a shared library that wraps OpenTelemetry initialization for all our Node.js and Python services. This ensured consistent configuration across teams.

import { NodeSDK } from '@opentelemetry/sdk-node';
import { OTLPTraceExporter } from '@opentelemetry/exporter-trace-otlp-grpc';
import { OTLPMetricExporter } from '@opentelemetry/exporter-metrics-otlp-grpc';
import { Resource } from '@opentelemetry/resources';
import { SemanticResourceAttributes } from '@opentelemetry/semantic-conventions';
import { BatchSpanProcessor } from '@opentelemetry/sdk-trace-base';
import { HttpInstrumentation } from '@opentelemetry/instrumentation-http';
import { ExpressInstrumentation } from '@opentelemetry/instrumentation-express';
import { PgInstrumentation } from '@opentelemetry/instrumentation-pg';
import { RedisInstrumentation } from '@opentelemetry/instrumentation-redis-4';

export function initTracing(serviceName: string) {
  const traceExporter = new OTLPTraceExporter({
    url: process.env.OTEL_COLLECTOR_ENDPOINT || 'grpc://otel-sidecar:4317',
  });

  const sdk = new NodeSDK({
    resource: new Resource({
      [SemanticResourceAttributes.SERVICE_NAME]: serviceName,
      [SemanticResourceAttributes.SERVICE_VERSION]: process.env.APP_VERSION || 'unknown',
      [SemanticResourceAttributes.DEPLOYMENT_ENVIRONMENT]: process.env.NODE_ENV || 'development',
      'service.team': process.env.TEAM_NAME || 'unknown',
      'service.tier': process.env.SERVICE_TIER || 'standard',
    }),
    spanProcessor: new BatchSpanProcessor(traceExporter, {
      maxQueueSize: 2048,
      maxExportBatchSize: 512,
      scheduledDelayMillis: 5000,
    }),
    instrumentations: [
      new HttpInstrumentation({
        ignoreIncomingPaths: ['/health', '/ready', '/metrics'],
      }),
      new ExpressInstrumentation(),
      new PgInstrumentation({ enhancedDatabaseReporting: true }),
      new RedisInstrumentation(),
    ],
  });

  sdk.start();
  return sdk;
}

This initialization runs before any other imports, ensuring all HTTP calls, database queries, and Redis operations are automatically captured.

OpenTelemetry Collector Configuration

The gateway collector configuration handles tail-based sampling, which was critical for our cost management. We only keep 100% of error traces and slow traces, while sampling 5% of successful fast traces.

receivers:
  otlp:
    protocols:
      grpc:
        endpoint: 0.0.0.0:4317
      http:
        endpoint: 0.0.0.0:4318

processors:
  tail_sampling:
    decision_wait: 10s
    num_traces: 100000
    expected_new_traces_per_sec: 5000
    policies:
      - name: errors-always
        type: status_code
        status_code:
          status_codes: [ERROR]
      - name: slow-traces
        type: latency
        latency:
          threshold_ms: 2000
      - name: probabilistic-sampling
        type: probabilistic
        probabilistic:
          sampling_percentage: 5

  batch:
    timeout: 5s
    send_batch_size: 1024
    send_batch_max_size: 2048

  resource:
    attributes:
      - key: collector.version
        value: "1.4.0"
        action: upsert

exporters:
  otlp/tempo:
    endpoint: tempo-distributor.observability:4317
    tls:
      insecure: false
      cert_file: /etc/ssl/certs/collector.pem
  
  prometheusremotewrite:
    endpoint: "http://mimir.observability:9009/api/v1/push"
    resource_to_telemetry_conversion:
      enabled: true

service:
  pipelines:
    traces:
      receivers: [otlp]
      processors: [tail_sampling, batch, resource]
      exporters: [otlp/tempo]
    metrics:
      receivers: [otlp]
      processors: [batch]
      exporters: [prometheusremotewrite]

Tail-based sampling at the gateway level reduced our trace storage costs by 72% while retaining 100% of actionable traces (errors and latency outliers).

Context Propagation Across Boundaries

The hardest part was not instrumentation but propagation. Our services communicate via HTTP, gRPC, SQS, and Kafka. Each required different propagation strategies.

For asynchronous messaging via SQS, we inject trace context into message attributes:

from opentelemetry import trace, context
from opentelemetry.propagate import inject
import json

def publish_to_sqs(queue_url: str, message_body: dict, tracer: trace.Tracer):
    with tracer.start_as_current_span(
        "sqs.publish",
        kind=trace.SpanKind.PRODUCER,
        attributes={
            "messaging.system": "aws_sqs",
            "messaging.destination": queue_url.split("/")[-1],
            "messaging.operation": "publish",
        }
    ) as span:
        # Inject trace context into message attributes
        carrier = {}
        inject(carrier)
        
        message_attributes = {
            key: {"DataType": "String", "StringValue": value}
            for key, value in carrier.items()
        }
        
        sqs_client.send_message(
            QueueUrl=queue_url,
            MessageBody=json.dumps(message_body),
            MessageAttributes=message_attributes,
        )
        
        span.set_attribute("messaging.message_id", response["MessageId"])

On the consumer side, we extract the context from message attributes and link it to the processing span. This gives us end-to-end traces even through asynchronous event-driven flows.

Service Dependency Graph

One unexpected benefit of full tracing was automatic service dependency mapping. We built a Grafana dashboard that renders real-time service graphs from trace data:

Service Dependency Graph

The graph updates every 30 seconds and highlights unhealthy edges (high error rates or elevated latency). During incidents, this became our primary navigation tool for identifying the blast radius.

Results After 12 Weeks

MetricBeforeAfterImprovement
Cross-service debugging time2.4 hours8 minutes-94%
Trace completeness77%99.7%+29%
Observability vendor cost$18.4K/mo$7.2K/mo-61%
Services with tracing31/4747/47100% coverage
Context propagation gaps23% of traces0.3% of traces-99%
Mean time to identify root cause47 minutes3.2 minutes-93%

The cost reduction came from consolidating four vendors into a self-hosted Grafana Tempo backend with tail-based sampling. We kept Grafana Cloud for dashboards and alerting.

Key Takeaways

Shared libraries beat documentation. Teams do not read instrumentation guides. A single initTracing(serviceName) call that auto-instruments everything achieved 100% adoption in four weeks.

Tail-based sampling is non-negotiable at scale. Head-based sampling at 5% would miss most error traces. Tail-based sampling at the collector level gave us 100% error capture with 95% cost reduction on normal traffic.

Async propagation is the real challenge. HTTP propagation works out of the box. SQS, Kafka, and SNS require custom propagation code. Budget twice the time for async boundaries.

Service graphs replace architecture diagrams. Our auto-generated dependency graph from trace data is always accurate, unlike manually maintained architecture documents that drift within weeks.

Conclusion

OpenTelemetry gave us what no single vendor could: unified, vendor-agnostic observability across 47 services, three regions, and four communication protocols. The 94% reduction in debugging time paid for the migration in the first month. Start with the shared instrumentation library, add tail-based sampling immediately, and tackle async propagation boundaries early. The hardest traces to complete are the most valuable ones to have.

Comments

    No comments yet. Be the first to share your thoughts.