Generating CloudWatch Dashboards from Service Descriptions with Kiro

How Kiro creates comprehensive observability dashboards from natural language service descriptions, ensuring every service ships with proper monitoring from day one.

#kiro#observability#dashboards#automation
Cover image for the article: Generating CloudWatch Dashboards from Service Descriptions with Kiro

Generating CloudWatch Dashboards from Service Descriptions with Kiro

Observability is one of those practices where the gap between "what we know we should do" and "what we actually do" is enormous. Every team agrees that services should ship with dashboards, alerts, and SLO tracking. In reality, monitoring is perpetually the thing that gets added "after launch" — which means it gets added after the first production incident, hastily, and incompletely.

The core problem is that creating good dashboards requires deep knowledge of both the service's architecture and CloudWatch's query language. It is tedious, repetitive work that does not ship features. Kiro eliminates this friction by generating complete CloudWatch dashboard definitions from natural language service descriptions, ensuring every service launches with production-grade observability.

The Observability Gap

We ran a B2B SaaS platform with 28 services. An internal audit revealed the state of our monitoring:

  • 7 services had comprehensive dashboards (the ones that had production incidents)
  • 12 services had basic dashboards (CPU, memory, request count only)
  • 9 services had no dashboards at all

The correlation was clear: services got proper monitoring only after they caused pain. This is backwards. The whole point of observability is to detect problems before they become incidents, not to explain them after the fact.

Creating a proper dashboard for a single service took 4-6 hours of engineering time: identifying the right metrics, writing CloudWatch Metrics Insights queries, laying out widgets logically, setting appropriate time ranges, and configuring alarms. Multiply that by 28 services and you understand why it never got prioritized.

How Kiro Generates Dashboards

We describe services to Kiro in natural language, including their architecture, key business operations, and SLOs. Kiro generates complete CloudWatch dashboard JSON that we deploy via Terraform.

Here is a service description we provided:

## Service: Payment Processing Service

Architecture: ECS Fargate (3 tasks), behind ALB
Database: RDS PostgreSQL (db.r6g.large)
Cache: ElastiCache Redis (cache.r6g.large)
External dependencies: Stripe API, fraud detection service
Queue: SQS for async payment events

Key operations:
- Process payment (p99 < 2s, error rate < 0.1%)
- Refund payment (p99 < 5s, error rate < 0.5%)
- Webhook processing (p99 < 500ms)

Business SLOs:
- 99.95% availability
- Payment success rate > 99.5%
- Webhook processing latency p99 < 500ms

From this description, Kiro generated a complete CloudWatch dashboard. Here is the critical section for the service-level metrics:

{
  "widgets": [
    {
      "type": "metric",
      "x": 0,
      "y": 0,
      "width": 12,
      "height": 6,
      "properties": {
        "title": "Payment Processing - Request Rate & Errors",
        "metrics": [
          [
            "AWS/ApplicationELB",
            "RequestCount",
            "TargetGroup", "payment-service-tg",
            "LoadBalancer", "app/payment-alb/abc123",
            { "stat": "Sum", "period": 60, "label": "Total Requests" }
          ],
          [
            "AWS/ApplicationELB",
            "HTTPCode_Target_5XX_Count",
            "TargetGroup", "payment-service-tg",
            "LoadBalancer", "app/payment-alb/abc123",
            { "stat": "Sum", "period": 60, "label": "5XX Errors", "color": "#d62728" }
          ],
          [
            "AWS/ApplicationELB",
            "HTTPCode_Target_4XX_Count",
            "TargetGroup", "payment-service-tg",
            "LoadBalancer", "app/payment-alb/abc123",
            { "stat": "Sum", "period": 60, "label": "4XX Errors", "color": "#ff7f0e" }
          ]
        ],
        "view": "timeSeries",
        "period": 60,
        "annotations": {
          "horizontal": [
            {
              "label": "Error Rate SLO (0.1%)",
              "value": 0.001,
              "fill": "above"
            }
          ]
        }
      }
    },
    {
      "type": "metric",
      "x": 12,
      "y": 0,
      "width": 12,
      "height": 6,
      "properties": {
        "title": "Payment Processing - Latency (p50, p95, p99)",
        "metrics": [
          [
            "AWS/ApplicationELB",
            "TargetResponseTime",
            "TargetGroup", "payment-service-tg",
            "LoadBalancer", "app/payment-alb/abc123",
            { "stat": "p50", "period": 60, "label": "p50" }
          ],
          [
            "AWS/ApplicationELB",
            "TargetResponseTime",
            "TargetGroup", "payment-service-tg",
            "LoadBalancer", "app/payment-alb/abc123",
            { "stat": "p95", "period": 60, "label": "p95" }
          ],
          [
            "AWS/ApplicationELB",
            "TargetResponseTime",
            "TargetGroup", "payment-service-tg",
            "LoadBalancer", "app/payment-alb/abc123",
            { "stat": "p99", "period": 60, "label": "p99", "color": "#d62728" }
          ]
        ],
        "annotations": {
          "horizontal": [
            {
              "label": "SLO: p99 < 2s",
              "value": 2,
              "fill": "above"
            }
          ]
        }
      }
    }
  ]
}

Kiro also generated widgets for the RDS connection pool, Redis cache hit ratios, SQS queue depth, and external dependency latency — all derived from the service description without any manual specification of metric namespaces or dimensions.

Terraform Integration for Dashboard-as-Code

The generated dashboard JSON integrates directly with our Terraform infrastructure:

resource "aws_cloudwatch_dashboard" "payment_service" {
  dashboard_name = "payment-service-overview"
  dashboard_body = file("${path.module}/dashboards/payment-service.json")
}

# Kiro also generates companion alarms
resource "aws_cloudwatch_metric_alarm" "payment_error_rate" {
  alarm_name          = "payment-service-error-rate-high"
  comparison_operator = "GreaterThanThreshold"
  evaluation_periods  = 3
  threshold           = 0.001
  alarm_description   = "Payment service error rate exceeds 0.1% SLO"
  
  metric_query {
    id          = "error_rate"
    expression  = "errors/requests"
    label       = "Error Rate"
    return_data = true
  }

  metric_query {
    id = "errors"
    metric {
      metric_name = "HTTPCode_Target_5XX_Count"
      namespace   = "AWS/ApplicationELB"
      period      = 300
      stat        = "Sum"
      dimensions = {
        TargetGroup  = "payment-service-tg"
        LoadBalancer = "app/payment-alb/abc123"
      }
    }
  }

  metric_query {
    id = "requests"
    metric {
      metric_name = "RequestCount"
      namespace   = "AWS/ApplicationELB"
      period      = 300
      stat        = "Sum"
      dimensions = {
        TargetGroup  = "payment-service-tg"
        LoadBalancer = "app/payment-alb/abc123"
      }
    }
  }

  alarm_actions = [aws_sns_topic.ops_alerts.arn]
  ok_actions    = [aws_sns_topic.ops_alerts.arn]
  
  tags = {
    service = "payment-service"
    slo     = "error-rate"
  }
}

The SLO Budget Widget

One of Kiro's most valuable generated widgets is the SLO error budget tracker — something most teams never build manually because the math is non-trivial:

{
  "type": "metric",
  "properties": {
    "title": "SLO Error Budget Remaining (30-day rolling)",
    "metrics": [
      [
        {
          "expression": "((1 - (errors_30d / requests_30d)) - 0.9995) / (1 - 0.9995) * 100",
          "label": "Error Budget Remaining (%)",
          "id": "budget"
        }
      ]
    ],
    "annotations": {
      "horizontal": [
        { "label": "Budget Exhausted", "value": 0, "color": "#d62728" },
        { "label": "Warning (25% remaining)", "value": 25, "color": "#ff7f0e" }
      ]
    }
  }
}

This widget shows the team exactly how much error budget remains before they violate their SLO — a metric that drives better operational decisions than raw error counts ever could.

Before and After: Observability Coverage

MetricBefore KiroAfter KiroImprovement
Services with comprehensive dashboards25% (7/28)100% (28/28)4x coverage
Time to create dashboard per service4-6 hours20 minutes92% reduction
MTTD (mean time to detect issues)18 minutes4 minutes78% reduction
Incidents caught by monitors vs. users40% by monitors85% by monitors2.1x improvement
SLO tracking coverage3 services28 servicesComplete coverage

The MTTD improvement alone justified the effort. Detecting issues 14 minutes faster means fewer affected customers, shorter incidents, and less weekend paging.

Kiro Observability Dashboard Architecture

Standardization Across Teams

A secondary benefit: all dashboards now follow the same layout conventions. Widget placement, color coding, time range defaults, and annotation styles are consistent across every service. New team members can look at any service's dashboard and immediately understand the layout because Kiro generates from consistent templates governed by our steering files.

Conclusion

Observability should not be optional or deferred. Every service should ship with monitoring from day one, and the reason most do not is purely the effort required to create proper dashboards. Kiro removes that barrier entirely. Describe your service, get a complete dashboard, deploy it alongside your infrastructure.

For platform engineering leaders, the implication is a shift from "chase teams to add monitoring" to "monitoring is automatic." Kiro-generated dashboards become part of the service provisioning template. New services are born observable. The days of discovering blind spots during production incidents should be behind us.

Comments

    No comments yet. Be the first to share your thoughts.