Synthetic Monitors From 12 Regions Catching Issues Before Users

How to implement global synthetic monitoring that detects availability and performance degradation from every major region—catching issues minutes before real users are impacted.

#synthetic-monitoring#availability#sre#testing
Cover image for the article: Synthetic Monitors From 12 Regions Catching Issues Before Users

The Problem: Blind to Regional Failures

Our application serves users across 6 continents. We had server-side metrics and real user monitoring (RUM), but both share the same blind spot: they only detect problems after real users are affected.

When our CDN provider experienced a partial outage in Southeast Asia, we didn't know for 23 minutes—until a customer in Singapore filed a support ticket. Our server-side metrics showed zero errors because traffic from that region never reached our origin. RUM data was sparse because affected users couldn't load the monitoring script.

After deploying synthetic monitors from 12 global regions, we detect regional issues in under 90 seconds, before any real user reports them. Our mean time to detect dropped from 11 minutes (average across incidents) to 47 seconds for issues with regional scope.

What Synthetic Monitoring Catches That Other Monitoring Misses

Failure TypeServer MetricsRUMSynthetic
Origin server downYesDelayedYes
CDN regional outageNoPartialYes
DNS resolution failureNoNoYes
SSL certificate issuesNoPartialYes
Third-party script failureNoYesYes
Regional network partitionNoSparseYes
Slow TTFB from specific regionsNoDelayedYes

Synthetic Monitoring Coverage Map

Architecture: Multi-Region Check Infrastructure

We run synthetic checks from 12 locations using a combination of managed services and self-hosted probes:

# synthetic-monitoring/config.yaml
global:
  check_interval: 60s
  timeout: 30s
  alert_after_consecutive_failures: 2
  
regions:
  - name: us-east-1
    provider: self-hosted  # ECS Fargate
    latitude: 39.0438
    longitude: -77.4874
  - name: eu-west-1
    provider: self-hosted
    latitude: 53.3498
    longitude: -6.2603
  - name: ap-southeast-1
    provider: self-hosted
    latitude: 1.3521
    longitude: 103.8198
  - name: ap-northeast-1
    provider: self-hosted
    latitude: 35.6762
    longitude: 139.6503
  - name: sa-east-1
    provider: self-hosted
    latitude: -23.5505
    longitude: -46.6333
  - name: af-south-1
    provider: self-hosted
    latitude: -33.9249
    longitude: 18.4241
  - name: us-west-2
    provider: self-hosted
    latitude: 45.5155
    longitude: -122.6789
  - name: eu-central-1
    provider: self-hosted
    latitude: 50.1109
    longitude: 8.6821
  - name: ap-south-1
    provider: self-hosted
    latitude: 19.0760
    longitude: 72.8777
  - name: me-south-1
    provider: self-hosted
    latitude: 26.0667
    longitude: 50.5577
  - name: ap-southeast-2
    provider: self-hosted
    latitude: -33.8688
    longitude: 151.2093
  - name: ca-central-1
    provider: self-hosted
    latitude: 45.5017
    longitude: -73.5673

checks:
  - name: homepage-availability
    type: http
    url: https://www.example.com
    assertions:
      - type: status_code
        value: 200
      - type: response_time
        max_ms: 3000
      - type: body_contains
        value: "Welcome"
      - type: ssl_days_remaining
        min: 14
        
  - name: api-health
    type: http
    url: https://api.example.com/health
    method: GET
    headers:
      Accept: application/json
    assertions:
      - type: status_code
        value: 200
      - type: json_path
        path: "$.status"
        value: "healthy"
      - type: response_time
        max_ms: 1000

  - name: checkout-flow
    type: browser
    script: checkout-critical-path
    assertions:
      - type: step_completes
        step: "payment_form_loaded"
        max_ms: 5000

The Synthetic Check Runner

Each region runs a lightweight check runner that executes HTTP and browser-based checks:

// synthetic-runner/src/check-executor.ts
interface CheckResult {
  checkName: string;
  region: string;
  timestamp: Date;
  status: 'pass' | 'fail' | 'degraded';
  responseTimeMs: number;
  assertions: AssertionResult[];
  dnsTimeMs: number;
  tcpTimeMs: number;
  tlsTimeMs: number;
  ttfbMs: number;
  downloadTimeMs: number;
}

interface AssertionResult {
  type: string;
  expected: string;
  actual: string;
  passed: boolean;
}

async function executeHttpCheck(
  check: HttpCheckConfig,
  region: string
): Promise<CheckResult> {
  const startTime = performance.now();
  const timings: Record<string, number> = {};

  try {
    const response = await fetch(check.url, {
      method: check.method || 'GET',
      headers: check.headers || {},
      signal: AbortSignal.timeout(check.timeoutMs || 30000),
    });

    const responseTime = performance.now() - startTime;
    const body = await response.text();

    const assertions = check.assertions.map(assertion => {
      switch (assertion.type) {
        case 'status_code':
          return {
            type: 'status_code',
            expected: String(assertion.value),
            actual: String(response.status),
            passed: response.status === assertion.value,
          };
        case 'response_time':
          return {
            type: 'response_time',
            expected: `<${assertion.max_ms}ms`,
            actual: `${Math.round(responseTime)}ms`,
            passed: responseTime <= assertion.max_ms,
          };
        case 'body_contains':
          return {
            type: 'body_contains',
            expected: assertion.value,
            actual: body.includes(assertion.value) ? 'found' : 'not found',
            passed: body.includes(assertion.value),
          };
        default:
          return {
            type: assertion.type,
            expected: 'unknown',
            actual: 'unknown',
            passed: false,
          };
      }
    });

    const allPassed = assertions.every(a => a.passed);
    const anyDegraded = assertions.some(
      a => a.type === 'response_time' && !a.passed
    );

    return {
      checkName: check.name,
      region,
      timestamp: new Date(),
      status: allPassed ? 'pass' : anyDegraded ? 'degraded' : 'fail',
      responseTimeMs: Math.round(responseTime),
      assertions,
      dnsTimeMs: timings.dns || 0,
      tcpTimeMs: timings.tcp || 0,
      tlsTimeMs: timings.tls || 0,
      ttfbMs: timings.ttfb || 0,
      downloadTimeMs: timings.download || 0,
    };
  } catch (error) {
    return {
      checkName: check.name,
      region,
      timestamp: new Date(),
      status: 'fail',
      responseTimeMs: performance.now() - startTime,
      assertions: [{
        type: 'connection',
        expected: 'success',
        actual: error instanceof Error ? error.message : 'unknown error',
        passed: false,
      }],
      dnsTimeMs: 0,
      tcpTimeMs: 0,
      tlsTimeMs: 0,
      ttfbMs: 0,
      downloadTimeMs: 0,
    };
  }
}

Intelligent Alerting: Regional Correlation

A single failed check from one region isn't necessarily an incident. We use correlation logic to distinguish between real outages and transient network blips:

# synthetic-monitoring/alerting.py
from dataclasses import dataclass

@dataclass
class RegionalStatus:
    region: str
    consecutive_failures: int
    last_success_seconds_ago: float
    avg_response_time_ms: float

def evaluate_alert_condition(
    check_name: str,
    regional_statuses: list[RegionalStatus],
    config: dict,
) -> dict:
    """
    Determine if an alert should fire based on regional check results.
    
    Rules:
    - Single region failing: wait for 3 consecutive failures before alerting
    - Multiple regions failing: alert after 2 consecutive failures (likely real)
    - All regions failing: alert immediately (global outage)
    """
    failing_regions = [
        r for r in regional_statuses 
        if r.consecutive_failures >= 2
    ]
    
    total_regions = len(regional_statuses)
    failing_count = len(failing_regions)
    
    if failing_count == 0:
        return {"alert": False, "status": "healthy"}
    
    if failing_count == total_regions:
        return {
            "alert": True,
            "severity": "critical",
            "status": "global_outage",
            "message": f"{check_name}: All {total_regions} regions failing",
        }
    
    if failing_count >= total_regions * 0.5:
        return {
            "alert": True,
            "severity": "critical",
            "status": "major_outage",
            "message": f"{check_name}: {failing_count}/{total_regions} regions failing",
            "affected_regions": [r.region for r in failing_regions],
        }
    
    if failing_count >= 2:
        return {
            "alert": True,
            "severity": "warning",
            "status": "partial_outage",
            "message": f"{check_name}: {failing_count} regions degraded",
            "affected_regions": [r.region for r in failing_regions],
        }
    
    # Single region - higher threshold
    single_failure = failing_regions[0]
    if single_failure.consecutive_failures >= 3:
        return {
            "alert": True,
            "severity": "warning",
            "status": "regional_issue",
            "message": f"{check_name}: {single_failure.region} failing consistently",
        }
    
    return {"alert": False, "status": "monitoring"}

Regional Correlation Alert Logic

Browser-Based Checks for Critical Flows

HTTP checks catch availability issues. Browser checks catch functionality issues:

// synthetic-runner/src/browser-checks/checkout-flow.ts
import { chromium, Page } from 'playwright';

interface BrowserCheckStep {
  name: string;
  durationMs: number;
  passed: boolean;
  screenshot?: string;
}

async function checkoutFlowCheck(): Promise<BrowserCheckStep[]> {
  const browser = await chromium.launch({ headless: true });
  const page = await browser.newPage();
  const steps: BrowserCheckStep[] = [];

  try {
    // Step 1: Load pricing page
    let start = Date.now();
    await page.goto('https://www.example.com/pricing');
    await page.waitForSelector('[data-testid="plan-card"]');
    steps.push({
      name: 'pricing_page_loaded',
      durationMs: Date.now() - start,
      passed: true,
    });

    // Step 2: Select a plan
    start = Date.now();
    await page.click('[data-testid="plan-pro-cta"]');
    await page.waitForSelector('#checkout-form');
    steps.push({
      name: 'checkout_form_loaded',
      durationMs: Date.now() - start,
      passed: true,
    });

    // Step 3: Verify payment provider loads
    start = Date.now();
    const stripeFrame = page.frameLocator('iframe[name*="stripe"]');
    await stripeFrame.locator('#card-number').waitFor({ timeout: 10000 });
    steps.push({
      name: 'payment_provider_loaded',
      durationMs: Date.now() - start,
      passed: true,
    });
  } catch (error) {
    steps.push({
      name: 'flow_error',
      durationMs: 0,
      passed: false,
      screenshot: await page.screenshot({ encoding: 'base64' }),
    });
  } finally {
    await browser.close();
  }

  return steps;
}

Results

MetricBeforeAfter
Mean time to detect (regional)23 min47 sec
Mean time to detect (global)4 min12 sec
Incidents detected before user reports34%91%
CDN issues detected proactively0%100%
False positive rate—2.1%
SSL expiry incidents3/year0/year
Third-party outage awarenessMinutes-to-hoursSeconds

Key Takeaways

  1. Synthetic monitoring fills the gaps that server metrics and RUM cannot. CDN failures, DNS issues, and network partitions are invisible to origin-based monitoring.

  2. 12 regions provide global coverage with reasonable cost. You don't need probes everywhere—strategic placement in major user regions catches 99% of regional issues.

  3. Correlate across regions before alerting. A single-region failure needs more evidence than a multi-region failure. This eliminates false positives from transient network issues.

  4. Browser checks catch what HTTP checks miss. A 200 status code doesn't mean the page works. Critical user flows need end-to-end browser-based validation.

  5. Synthetic monitoring is your early warning system. The goal isn't to replace other monitoring—it's to detect problems in the 90-second window between "something broke" and "users noticed."

The best incident response starts before users are impacted. Synthetic monitoring from 12 regions gives you that head start.

Comments

    No comments yet. Be the first to share your thoughts.