AI-Powered Postmortem Analysis: Finding Systemic Patterns with Kiro

How Kiro analyzes incident postmortems to identify recurring systemic patterns, predict future failure modes, and generate actionable prevention strategies.

#kiro#incident-management#postmortem#sre
Cover image for the article: AI-Powered Postmortem Analysis: Finding Systemic Patterns with Kiro

AI-Powered Postmortem Analysis: Finding Systemic Patterns with Kiro

We write postmortems religiously. After every SEV-1 and SEV-2 incident, the on-call team produces a detailed timeline, root cause analysis, and action items. We have 147 postmortems in our archive spanning three years. We are excellent at learning from individual incidents. What we were terrible at was learning across incidents — identifying the systemic patterns that cause the same category of failure to recur in different services, with different symptoms, quarter after quarter.

Kiro changed our approach to incident analysis by treating our postmortem corpus as a dataset to be analyzed, not just a collection of individual documents. The patterns it surfaced have driven infrastructure improvements that reduced our incident rate by 43% in two quarters.

The Pattern Blindness Problem

Here is what pattern blindness looks like in practice. Over 18 months, we had:

  • A payment service outage caused by connection pool exhaustion during a traffic spike
  • A notification service degradation caused by thread pool exhaustion during a batch send
  • An authentication service timeout cascade caused by a downstream service being slow

Three different services, three different teams, three separate postmortems. Each postmortem correctly identified the proximate cause and generated appropriate action items. But none of them identified the systemic issue: we had no standard approach to resource pool sizing, backpressure, or graceful degradation under load. The same class of failure kept recurring because we were treating symptoms individually rather than addressing the underlying architectural pattern.

A human reading all 147 postmortems could theoretically spot these patterns. In practice, no one has time to read 147 postmortems looking for cross-cutting themes. That is exactly the kind of analysis that AI excels at.

Feeding Postmortems to Kiro

We structured our postmortem archive and asked Kiro to analyze it for systemic patterns. Each postmortem followed a consistent format:

## Incident: [SERVICE]-[DATE]-[SEQ]
**Severity**: SEV-1 / SEV-2
**Duration**: X hours Y minutes
**Customer Impact**: Description
**Services Affected**: list

### Timeline
- HH:MM - Event description
- ...

### Root Cause
Detailed technical root cause

### Contributing Factors
- Factor 1
- Factor 2

### Action Items
- [ ] Action (Owner, Due Date)

### Tags
category: deployment | scaling | dependency | data | security
trigger: traffic-spike | bad-deploy | upstream-failure | data-corruption

Kiro's analysis produced a systemic pattern report:

## Kiro Systemic Pattern Analysis
### Corpus: 147 postmortems (Jan 2023 - Jun 2026)

---

### Pattern 1: Resource Exhaustion Under Load (23 incidents, 15.6%)

**Recurring theme**: Services fail when a bounded resource (connections, 
threads, file descriptors, memory) is exhausted during traffic spikes or 
batch operations.

**Affected services**: payments (4x), notifications (3x), auth (3x), 
search (3x), orders (2x), + 8 others

**Common contributing factors**:
- No circuit breaker on downstream calls (18/23 incidents)
- Connection pool sized for average load, not peak (15/23)
- No backpressure mechanism (20/23)
- Retry storms amplifying initial resource pressure (12/23)

**Systemic root cause**: No organizational standard for resource pool 
configuration, no load testing requirement before production deployment, 
no circuit breaker library in shared infrastructure.

**Recommended systemic fixes**:
1. Adopt organization-wide circuit breaker library (resilience4j or similar)
2. Mandate load test to 3x average traffic before production deployment
3. Standard resource pool sizing formula: peak_concurrent * 1.5 + buffer
4. Add resource utilization alerts at 70% threshold (early warning)

**Projected impact**: Eliminating this pattern would have prevented 
~15 of these 23 incidents (65%), saving approximately 89 hours of 
incident response time.

---

### Pattern 2: Deployment-Induced Failures (31 incidents, 21.1%)

**Recurring theme**: Production incidents triggered within 30 minutes 
of a deployment, typically due to configuration mismatches, missing 
environment variables, or incompatible schema changes.

**Sub-patterns identified**:
- Missing env vars in new environment (9 incidents)
- Database migration timing vs. code deployment (8 incidents)
- Feature flag not set in production (7 incidents)
- Incompatible API version between services (7 incidents)

**Systemic root cause**: No pre-deployment validation of configuration 
completeness, no automated compatibility checking between service versions.

**Recommended systemic fixes**:
1. Pre-deployment config validation (compare required env vars against 
   deployed secrets)
2. Schema migration must complete and validate before code deployment begins
3. Feature flag audit in deployment pipeline (warn on flags referenced 
   in code but not defined in LaunchDarkly)
4. Contract testing between services (Pact or similar)

**Projected impact**: Would have prevented ~22 of 31 incidents (71%).

From Patterns to Prevention: Generated Policies

Kiro did not just identify patterns — it generated the specific implementation artifacts to prevent them. For Pattern 1 (Resource Exhaustion), it produced a Kiro steering file that prevents future services from launching without proper resource management:

# Steering: Service Resilience Requirements

## Circuit Breakers (Required for all external calls)
Every HTTP client calling an external service must:
- Configure a circuit breaker with failure threshold of 50% over 10 requests
- Define a fallback behavior (cached response, degraded response, or fast-fail)
- Emit circuit breaker state change metrics

## Resource Pool Configuration (Required)
All connection pools must specify:
- Maximum size based on: expected_peak_concurrent_requests * 1.5
- Connection timeout: no longer than 5 seconds
- Idle timeout: aligned with downstream keepalive settings
- Queue/wait timeout: fail fast rather than queue indefinitely

## Load Testing (Required before production)
Services must pass load test demonstrating:
- Stable behavior at 3x average traffic for 15 minutes
- Graceful degradation (not crash) at 5x average traffic
- Recovery to normal within 60 seconds after load subsides

For Pattern 2 (Deployment Failures), Kiro generated a pre-deployment validation GitHub Action:

# .github/workflows/pre-deploy-validation.yml
name: Pre-Deployment Validation

on:
  workflow_call:
    inputs:
      service_name:
        required: true
        type: string
      target_environment:
        required: true
        type: string

jobs:
  validate-config:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      
      - name: Extract required environment variables
        id: required-vars
        run: |
          # Find all env var references in the service code
          grep -roh 'process\.env\.\w\+\|os\.environ\[.\w\+.\]' \
            services/${{ inputs.service_name }}/src/ | \
            sort -u > required_vars.txt
          echo "count=$(wc -l < required_vars.txt)" >> $GITHUB_OUTPUT
          
      - name: Validate against deployed secrets
        run: |
          # Check each required var exists in target environment
          MISSING=""
          while read var; do
            clean_var=$(echo "$var" | grep -oP '\w+$')
            if ! aws ssm get-parameter \
              --name "/${{ inputs.target_environment }}/${{ inputs.service_name }}/${clean_var}" \
              --query 'Parameter.Value' 2>/dev/null; then
              MISSING="${MISSING}\n- ${clean_var}"
            fi
          done < required_vars.txt
          
          if [ -n "$MISSING" ]; then
            echo "::error::Missing environment variables in ${{ inputs.target_environment }}:${MISSING}"
            exit 1
          fi

      - name: Validate database migration status
        run: |
          # Ensure all pending migrations have been applied
          PENDING=$(npm run migration:status -- --env=${{ inputs.target_environment }} 2>&1 | grep "pending" | wc -l)
          if [ "$PENDING" -gt 0 ]; then
            echo "::error::${PENDING} pending database migrations. Run migrations before deploying."
            exit 1
          fi

Before and After: Incident Reduction

MetricBefore Pattern AnalysisAfter Systemic FixesImprovement
Total incidents/quarter18.3 average10.4 average43% reduction
Resource exhaustion incidents5.7/quarter1.2/quarter79% reduction
Deployment-related incidents7.8/quarter2.1/quarter73% reduction
Mean time between incidents5 days8.6 days72% improvement
Repeat pattern incidents67% were repeat patterns31%54% reduction

The most significant insight: 67% of our incidents were instances of patterns we had already seen and documented. We just had not connected the dots across postmortems. After implementing systemic fixes guided by Kiro's analysis, fewer than a third of incidents were repeats.

Continuous Pattern Detection

We now run Kiro's analysis quarterly against our growing postmortem corpus. Each quarter, it identifies emerging patterns before they become entrenched. In Q2 2026, it flagged an emerging pattern of "slow S3 eventual consistency causing stale reads" that had appeared in three incidents across different services — a pattern that was not yet obvious to any individual team.

Kiro Incident Pattern Analysis

Conclusion

Individual postmortems are necessary but insufficient. They prevent the exact same incident from recurring, but they do not prevent the same category of failure from appearing in different services. Kiro's cross-corpus analysis surfaces these systemic patterns and generates the specific policies, automation, and infrastructure changes needed to address them at the root.

For SRE leaders and CTOs, this represents a shift from reactive incident management to proactive failure prevention. Instead of waiting for patterns to emerge through painful repetition, you can identify them early and eliminate entire categories of failure. The postmortem archive stops being a historical record and becomes a strategic asset that actively improves system reliability.

Comments

    No comments yet. Be the first to share your thoughts.