Complete Monitoring Stack for 200-Pod Kubernetes Clusters

How to build a production-grade Prometheus and Grafana monitoring stack for large Kubernetes clusters with custom metrics, intelligent alerting, and capacity planning.

#prometheus#grafana#kubernetes#monitoring
Cover image for the article: Complete Monitoring Stack for 200-Pod Kubernetes Clusters

The Problem: Observability at Scale

When your Kubernetes cluster grows beyond 50 pods, the default metrics server and kubectl top become insufficient. At 200+ pods running 34 microservices, you need a monitoring system that answers three questions in under 30 seconds: What's broken? Why is it broken? Is it getting worse?

Our platform serves 8 million requests per hour across three availability zones. Before implementing our current stack, our mean time to detect (MTTD) issues was 12 minutes—often after users reported problems. After building a comprehensive Prometheus and Grafana stack, MTTD dropped to 47 seconds, and 73% of issues are now detected before any user impact.

Architecture Overview

Monitoring Architecture Diagram

Our monitoring stack consists of:

  • Prometheus Operator managing multiple Prometheus instances (sharded by namespace)
  • Thanos for long-term storage and cross-cluster querying
  • Grafana with provisioned dashboards and alert rules
  • Alertmanager with intelligent routing and deduplication
  • Custom exporters for business-level metrics

Deploying the Stack

Prometheus Operator with Sharding

At 200 pods, a single Prometheus instance ingests approximately 180,000 time series. We shard across two instances to keep memory usage predictable:

# prometheus-sharded.yaml
apiVersion: monitoring.coreos.com/v1
kind: Prometheus
metadata:
  name: k8s-prometheus
  namespace: monitoring
spec:
  replicas: 2
  shards: 2
  retention: 7d
  retentionSize: 40GB
  resources:
    requests:
      memory: 8Gi
      cpu: "2"
    limits:
      memory: 12Gi
      cpu: "4"
  storage:
    volumeClaimTemplate:
      spec:
        storageClassName: gp3-encrypted
        resources:
          requests:
            storage: 100Gi
  serviceMonitorSelector:
    matchLabels:
      monitoring: enabled
  ruleSelector:
    matchLabels:
      role: alert-rules
  thanos:
    image: quay.io/thanos/thanos:v0.35.0
    objectStorageConfig:
      key: thanos.yaml
      name: thanos-objstore-config

ServiceMonitor for Application Metrics

Every microservice exposes a /metrics endpoint. ServiceMonitors auto-discover pods via labels:

# servicemonitor-payments.yaml
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
  name: payments-service
  labels:
    monitoring: enabled
spec:
  selector:
    matchLabels:
      app: payments-service
  endpoints:
    - port: metrics
      interval: 15s
      path: /metrics
      relabelings:
        - sourceLabels: [__meta_kubernetes_pod_label_version]
          targetLabel: app_version
        - sourceLabels: [__meta_kubernetes_pod_node_name]
          targetLabel: node
  namespaceSelector:
    matchNames:
      - payments

Custom Recording Rules for Performance

Raw metrics create expensive queries. Recording rules pre-compute frequently used aggregations:

# recording-rules.yaml
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
  name: sli-recording-rules
  labels:
    role: alert-rules
spec:
  groups:
    - name: sli.rules
      interval: 30s
      rules:
        - record: service:request_rate:5m
          expr: |
            sum by (service, namespace) (
              rate(http_requests_total[5m])
            )
        - record: service:error_rate:5m
          expr: |
            sum by (service, namespace) (
              rate(http_requests_total{status_code=~"5.."}[5m])
            ) /
            sum by (service, namespace) (
              rate(http_requests_total[5m])
            )
        - record: service:latency_p99:5m
          expr: |
            histogram_quantile(0.99,
              sum by (service, le) (
                rate(http_request_duration_seconds_bucket[5m])
              )
            )
        - record: node:cpu_saturation:5m
          expr: |
            1 - avg by (node) (
              rate(node_cpu_seconds_total{mode="idle"}[5m])
            )

Grafana Dashboard Design

The Four Golden Signals Dashboard

We provision dashboards as code using Grafana's JSON model. Our primary service dashboard displays the four golden signals: latency, traffic, errors, and saturation.

{
  "title": "Service Overview - Four Golden Signals",
  "panels": [
    {
      "title": "Request Rate",
      "type": "timeseries",
      "targets": [{
        "expr": "service:request_rate:5m{service=\"$service\"}",
        "legendFormat": "{{namespace}}"
      }],
      "fieldConfig": {
        "defaults": {
          "unit": "reqps",
          "thresholds": {
            "steps": [
              {"color": "green", "value": null},
              {"color": "yellow", "value": 5000},
              {"color": "red", "value": 8000}
            ]
          }
        }
      }
    },
    {
      "title": "Error Rate",
      "type": "stat",
      "targets": [{
        "expr": "service:error_rate:5m{service=\"$service\"} * 100",
        "legendFormat": "Error %"
      }],
      "fieldConfig": {
        "defaults": {
          "unit": "percent",
          "thresholds": {
            "steps": [
              {"color": "green", "value": null},
              {"color": "yellow", "value": 0.1},
              {"color": "red", "value": 1.0}
            ]
          }
        }
      }
    }
  ]
}

Grafana Four Golden Signals Dashboard

Intelligent Alerting

Multi-Window, Multi-Burn-Rate Alerts

Simple threshold alerts create noise. We use Google's multi-window burn rate approach for SLO-based alerting:

# alert-rules.yaml
groups:
  - name: slo-burn-rate-alerts
    rules:
      # Fast burn - pages immediately
      - alert: HighErrorBurnRate_Critical
        expr: |
          (
            service:error_rate:5m{service=~".+"} > (14.4 * 0.001)
            and
            service:error_rate:1h{service=~".+"} > (14.4 * 0.001)
          )
        for: 2m
        labels:
          severity: critical
          team: "{{$labels.namespace}}"
        annotations:
          summary: "{{ $labels.service }} burning error budget 14x faster than allowed"
          runbook: "https://runbooks.internal/slo-burn-rate"

      # Slow burn - tickets
      - alert: HighErrorBurnRate_Warning
        expr: |
          (
            service:error_rate:30m{service=~".+"} > (3 * 0.001)
            and
            service:error_rate:6h{service=~".+"} > (3 * 0.001)
          )
        for: 15m
        labels:
          severity: warning
        annotations:
          summary: "{{ $labels.service }} burning error budget 3x faster than allowed"

Alert Routing

Alertmanager routes alerts to the right team through label matching:

# alertmanager-config.yaml
route:
  receiver: default-slack
  group_by: ['service', 'severity']
  group_wait: 30s
  group_interval: 5m
  repeat_interval: 4h
  routes:
    - match:
        severity: critical
      receiver: pagerduty-oncall
      continue: true
    - match:
        severity: critical
      receiver: slack-incidents
    - match_re:
        team: "payments|billing"
      receiver: slack-payments-team

Capacity Planning with Predictive Queries

Beyond alerting, we use Prometheus for capacity planning. Linear prediction queries identify resources that will exhaust within 7 days:

# Predict disk exhaustion within 7 days
predict_linear(
  kubelet_volume_stats_available_bytes{
    namespace="production"
  }[7d], 7 * 24 * 3600
) < 0

# Predict memory pressure within 3 days
predict_linear(
  container_memory_working_set_bytes{
    namespace="production"
  }[3d], 3 * 24 * 3600
) > on(pod) kube_pod_container_resource_limits{resource="memory"}

Results and Metrics

After 6 months of running this stack:

MetricBeforeAfter
Mean Time to Detect12 min47 sec
False positive alerts/week343
Dashboard load time (p95)8.2s1.4s
Prometheus memory usage24 GB (single)9 GB per shard
Metric cardinalityUnbounded180k per shard
Capacity issues caught proactively12%89%

Key Takeaways

  1. Shard Prometheus early. Don't wait until a single instance OOMs at 3 AM. Two shards with 100k series each are more predictable than one with 200k.

  2. Recording rules are essential at scale. Pre-computing aggregations reduces dashboard query time from seconds to milliseconds and prevents Prometheus overload.

  3. Use burn-rate alerts, not thresholds. Multi-window burn rates eliminate 90% of false positives while catching slow-burn degradations that threshold alerts miss.

  4. Provision dashboards as code. Grafana dashboard JSON in version control ensures consistency across environments and enables review processes.

  5. Capacity planning is monitoring's highest ROI. Predictive queries that prevent outages deliver more value than reactive alerting alone.

The goal isn't more data—it's faster answers. Every metric, dashboard, and alert should map directly to a decision someone needs to make.

Comments

    No comments yet. Be the first to share your thoughts.