Complete Monitoring Stack for 200-Pod Kubernetes Clusters
How to build a production-grade Prometheus and Grafana monitoring stack for large Kubernetes clusters with custom metrics, intelligent alerting, and capacity planning.

The Problem: Observability at Scale
When your Kubernetes cluster grows beyond 50 pods, the default metrics server and kubectl top become insufficient. At 200+ pods running 34 microservices, you need a monitoring system that answers three questions in under 30 seconds: What's broken? Why is it broken? Is it getting worse?
Our platform serves 8 million requests per hour across three availability zones. Before implementing our current stack, our mean time to detect (MTTD) issues was 12 minutes—often after users reported problems. After building a comprehensive Prometheus and Grafana stack, MTTD dropped to 47 seconds, and 73% of issues are now detected before any user impact.
Architecture Overview
Our monitoring stack consists of:
- Prometheus Operator managing multiple Prometheus instances (sharded by namespace)
- Thanos for long-term storage and cross-cluster querying
- Grafana with provisioned dashboards and alert rules
- Alertmanager with intelligent routing and deduplication
- Custom exporters for business-level metrics
Deploying the Stack
Prometheus Operator with Sharding
At 200 pods, a single Prometheus instance ingests approximately 180,000 time series. We shard across two instances to keep memory usage predictable:
# prometheus-sharded.yaml
apiVersion: monitoring.coreos.com/v1
kind: Prometheus
metadata:
name: k8s-prometheus
namespace: monitoring
spec:
replicas: 2
shards: 2
retention: 7d
retentionSize: 40GB
resources:
requests:
memory: 8Gi
cpu: "2"
limits:
memory: 12Gi
cpu: "4"
storage:
volumeClaimTemplate:
spec:
storageClassName: gp3-encrypted
resources:
requests:
storage: 100Gi
serviceMonitorSelector:
matchLabels:
monitoring: enabled
ruleSelector:
matchLabels:
role: alert-rules
thanos:
image: quay.io/thanos/thanos:v0.35.0
objectStorageConfig:
key: thanos.yaml
name: thanos-objstore-config
ServiceMonitor for Application Metrics
Every microservice exposes a /metrics endpoint. ServiceMonitors auto-discover pods via labels:
# servicemonitor-payments.yaml
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: payments-service
labels:
monitoring: enabled
spec:
selector:
matchLabels:
app: payments-service
endpoints:
- port: metrics
interval: 15s
path: /metrics
relabelings:
- sourceLabels: [__meta_kubernetes_pod_label_version]
targetLabel: app_version
- sourceLabels: [__meta_kubernetes_pod_node_name]
targetLabel: node
namespaceSelector:
matchNames:
- payments
Custom Recording Rules for Performance
Raw metrics create expensive queries. Recording rules pre-compute frequently used aggregations:
# recording-rules.yaml
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: sli-recording-rules
labels:
role: alert-rules
spec:
groups:
- name: sli.rules
interval: 30s
rules:
- record: service:request_rate:5m
expr: |
sum by (service, namespace) (
rate(http_requests_total[5m])
)
- record: service:error_rate:5m
expr: |
sum by (service, namespace) (
rate(http_requests_total{status_code=~"5.."}[5m])
) /
sum by (service, namespace) (
rate(http_requests_total[5m])
)
- record: service:latency_p99:5m
expr: |
histogram_quantile(0.99,
sum by (service, le) (
rate(http_request_duration_seconds_bucket[5m])
)
)
- record: node:cpu_saturation:5m
expr: |
1 - avg by (node) (
rate(node_cpu_seconds_total{mode="idle"}[5m])
)
Grafana Dashboard Design
The Four Golden Signals Dashboard
We provision dashboards as code using Grafana's JSON model. Our primary service dashboard displays the four golden signals: latency, traffic, errors, and saturation.
{
"title": "Service Overview - Four Golden Signals",
"panels": [
{
"title": "Request Rate",
"type": "timeseries",
"targets": [{
"expr": "service:request_rate:5m{service=\"$service\"}",
"legendFormat": "{{namespace}}"
}],
"fieldConfig": {
"defaults": {
"unit": "reqps",
"thresholds": {
"steps": [
{"color": "green", "value": null},
{"color": "yellow", "value": 5000},
{"color": "red", "value": 8000}
]
}
}
}
},
{
"title": "Error Rate",
"type": "stat",
"targets": [{
"expr": "service:error_rate:5m{service=\"$service\"} * 100",
"legendFormat": "Error %"
}],
"fieldConfig": {
"defaults": {
"unit": "percent",
"thresholds": {
"steps": [
{"color": "green", "value": null},
{"color": "yellow", "value": 0.1},
{"color": "red", "value": 1.0}
]
}
}
}
}
]
}
Intelligent Alerting
Multi-Window, Multi-Burn-Rate Alerts
Simple threshold alerts create noise. We use Google's multi-window burn rate approach for SLO-based alerting:
# alert-rules.yaml
groups:
- name: slo-burn-rate-alerts
rules:
# Fast burn - pages immediately
- alert: HighErrorBurnRate_Critical
expr: |
(
service:error_rate:5m{service=~".+"} > (14.4 * 0.001)
and
service:error_rate:1h{service=~".+"} > (14.4 * 0.001)
)
for: 2m
labels:
severity: critical
team: "{{$labels.namespace}}"
annotations:
summary: "{{ $labels.service }} burning error budget 14x faster than allowed"
runbook: "https://runbooks.internal/slo-burn-rate"
# Slow burn - tickets
- alert: HighErrorBurnRate_Warning
expr: |
(
service:error_rate:30m{service=~".+"} > (3 * 0.001)
and
service:error_rate:6h{service=~".+"} > (3 * 0.001)
)
for: 15m
labels:
severity: warning
annotations:
summary: "{{ $labels.service }} burning error budget 3x faster than allowed"
Alert Routing
Alertmanager routes alerts to the right team through label matching:
# alertmanager-config.yaml
route:
receiver: default-slack
group_by: ['service', 'severity']
group_wait: 30s
group_interval: 5m
repeat_interval: 4h
routes:
- match:
severity: critical
receiver: pagerduty-oncall
continue: true
- match:
severity: critical
receiver: slack-incidents
- match_re:
team: "payments|billing"
receiver: slack-payments-team
Capacity Planning with Predictive Queries
Beyond alerting, we use Prometheus for capacity planning. Linear prediction queries identify resources that will exhaust within 7 days:
# Predict disk exhaustion within 7 days
predict_linear(
kubelet_volume_stats_available_bytes{
namespace="production"
}[7d], 7 * 24 * 3600
) < 0
# Predict memory pressure within 3 days
predict_linear(
container_memory_working_set_bytes{
namespace="production"
}[3d], 3 * 24 * 3600
) > on(pod) kube_pod_container_resource_limits{resource="memory"}
Results and Metrics
After 6 months of running this stack:
| Metric | Before | After |
|---|---|---|
| Mean Time to Detect | 12 min | 47 sec |
| False positive alerts/week | 34 | 3 |
| Dashboard load time (p95) | 8.2s | 1.4s |
| Prometheus memory usage | 24 GB (single) | 9 GB per shard |
| Metric cardinality | Unbounded | 180k per shard |
| Capacity issues caught proactively | 12% | 89% |
Key Takeaways
-
Shard Prometheus early. Don't wait until a single instance OOMs at 3 AM. Two shards with 100k series each are more predictable than one with 200k.
-
Recording rules are essential at scale. Pre-computing aggregations reduces dashboard query time from seconds to milliseconds and prevents Prometheus overload.
-
Use burn-rate alerts, not thresholds. Multi-window burn rates eliminate 90% of false positives while catching slow-burn degradations that threshold alerts miss.
-
Provision dashboards as code. Grafana dashboard JSON in version control ensures consistency across environments and enables review processes.
-
Capacity planning is monitoring's highest ROI. Predictive queries that prevent outages deliver more value than reactive alerting alone.
The goal isn't more data—it's faster answers. Every metric, dashboard, and alert should map directly to a decision someone needs to make.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.