Synthetic Monitors From 12 Regions Catching Issues Before Users
How to implement global synthetic monitoring that detects availability and performance degradation from every major region—catching issues minutes before real users are impacted.

The Problem: Blind to Regional Failures
Our application serves users across 6 continents. We had server-side metrics and real user monitoring (RUM), but both share the same blind spot: they only detect problems after real users are affected.
When our CDN provider experienced a partial outage in Southeast Asia, we didn't know for 23 minutes—until a customer in Singapore filed a support ticket. Our server-side metrics showed zero errors because traffic from that region never reached our origin. RUM data was sparse because affected users couldn't load the monitoring script.
After deploying synthetic monitors from 12 global regions, we detect regional issues in under 90 seconds, before any real user reports them. Our mean time to detect dropped from 11 minutes (average across incidents) to 47 seconds for issues with regional scope.
What Synthetic Monitoring Catches That Other Monitoring Misses
| Failure Type | Server Metrics | RUM | Synthetic |
|---|---|---|---|
| Origin server down | Yes | Delayed | Yes |
| CDN regional outage | No | Partial | Yes |
| DNS resolution failure | No | No | Yes |
| SSL certificate issues | No | Partial | Yes |
| Third-party script failure | No | Yes | Yes |
| Regional network partition | No | Sparse | Yes |
| Slow TTFB from specific regions | No | Delayed | Yes |
Architecture: Multi-Region Check Infrastructure
We run synthetic checks from 12 locations using a combination of managed services and self-hosted probes:
# synthetic-monitoring/config.yaml
global:
check_interval: 60s
timeout: 30s
alert_after_consecutive_failures: 2
regions:
- name: us-east-1
provider: self-hosted # ECS Fargate
latitude: 39.0438
longitude: -77.4874
- name: eu-west-1
provider: self-hosted
latitude: 53.3498
longitude: -6.2603
- name: ap-southeast-1
provider: self-hosted
latitude: 1.3521
longitude: 103.8198
- name: ap-northeast-1
provider: self-hosted
latitude: 35.6762
longitude: 139.6503
- name: sa-east-1
provider: self-hosted
latitude: -23.5505
longitude: -46.6333
- name: af-south-1
provider: self-hosted
latitude: -33.9249
longitude: 18.4241
- name: us-west-2
provider: self-hosted
latitude: 45.5155
longitude: -122.6789
- name: eu-central-1
provider: self-hosted
latitude: 50.1109
longitude: 8.6821
- name: ap-south-1
provider: self-hosted
latitude: 19.0760
longitude: 72.8777
- name: me-south-1
provider: self-hosted
latitude: 26.0667
longitude: 50.5577
- name: ap-southeast-2
provider: self-hosted
latitude: -33.8688
longitude: 151.2093
- name: ca-central-1
provider: self-hosted
latitude: 45.5017
longitude: -73.5673
checks:
- name: homepage-availability
type: http
url: https://www.example.com
assertions:
- type: status_code
value: 200
- type: response_time
max_ms: 3000
- type: body_contains
value: "Welcome"
- type: ssl_days_remaining
min: 14
- name: api-health
type: http
url: https://api.example.com/health
method: GET
headers:
Accept: application/json
assertions:
- type: status_code
value: 200
- type: json_path
path: "$.status"
value: "healthy"
- type: response_time
max_ms: 1000
- name: checkout-flow
type: browser
script: checkout-critical-path
assertions:
- type: step_completes
step: "payment_form_loaded"
max_ms: 5000
The Synthetic Check Runner
Each region runs a lightweight check runner that executes HTTP and browser-based checks:
// synthetic-runner/src/check-executor.ts
interface CheckResult {
checkName: string;
region: string;
timestamp: Date;
status: 'pass' | 'fail' | 'degraded';
responseTimeMs: number;
assertions: AssertionResult[];
dnsTimeMs: number;
tcpTimeMs: number;
tlsTimeMs: number;
ttfbMs: number;
downloadTimeMs: number;
}
interface AssertionResult {
type: string;
expected: string;
actual: string;
passed: boolean;
}
async function executeHttpCheck(
check: HttpCheckConfig,
region: string
): Promise<CheckResult> {
const startTime = performance.now();
const timings: Record<string, number> = {};
try {
const response = await fetch(check.url, {
method: check.method || 'GET',
headers: check.headers || {},
signal: AbortSignal.timeout(check.timeoutMs || 30000),
});
const responseTime = performance.now() - startTime;
const body = await response.text();
const assertions = check.assertions.map(assertion => {
switch (assertion.type) {
case 'status_code':
return {
type: 'status_code',
expected: String(assertion.value),
actual: String(response.status),
passed: response.status === assertion.value,
};
case 'response_time':
return {
type: 'response_time',
expected: `<${assertion.max_ms}ms`,
actual: `${Math.round(responseTime)}ms`,
passed: responseTime <= assertion.max_ms,
};
case 'body_contains':
return {
type: 'body_contains',
expected: assertion.value,
actual: body.includes(assertion.value) ? 'found' : 'not found',
passed: body.includes(assertion.value),
};
default:
return {
type: assertion.type,
expected: 'unknown',
actual: 'unknown',
passed: false,
};
}
});
const allPassed = assertions.every(a => a.passed);
const anyDegraded = assertions.some(
a => a.type === 'response_time' && !a.passed
);
return {
checkName: check.name,
region,
timestamp: new Date(),
status: allPassed ? 'pass' : anyDegraded ? 'degraded' : 'fail',
responseTimeMs: Math.round(responseTime),
assertions,
dnsTimeMs: timings.dns || 0,
tcpTimeMs: timings.tcp || 0,
tlsTimeMs: timings.tls || 0,
ttfbMs: timings.ttfb || 0,
downloadTimeMs: timings.download || 0,
};
} catch (error) {
return {
checkName: check.name,
region,
timestamp: new Date(),
status: 'fail',
responseTimeMs: performance.now() - startTime,
assertions: [{
type: 'connection',
expected: 'success',
actual: error instanceof Error ? error.message : 'unknown error',
passed: false,
}],
dnsTimeMs: 0,
tcpTimeMs: 0,
tlsTimeMs: 0,
ttfbMs: 0,
downloadTimeMs: 0,
};
}
}
Intelligent Alerting: Regional Correlation
A single failed check from one region isn't necessarily an incident. We use correlation logic to distinguish between real outages and transient network blips:
# synthetic-monitoring/alerting.py
from dataclasses import dataclass
@dataclass
class RegionalStatus:
region: str
consecutive_failures: int
last_success_seconds_ago: float
avg_response_time_ms: float
def evaluate_alert_condition(
check_name: str,
regional_statuses: list[RegionalStatus],
config: dict,
) -> dict:
"""
Determine if an alert should fire based on regional check results.
Rules:
- Single region failing: wait for 3 consecutive failures before alerting
- Multiple regions failing: alert after 2 consecutive failures (likely real)
- All regions failing: alert immediately (global outage)
"""
failing_regions = [
r for r in regional_statuses
if r.consecutive_failures >= 2
]
total_regions = len(regional_statuses)
failing_count = len(failing_regions)
if failing_count == 0:
return {"alert": False, "status": "healthy"}
if failing_count == total_regions:
return {
"alert": True,
"severity": "critical",
"status": "global_outage",
"message": f"{check_name}: All {total_regions} regions failing",
}
if failing_count >= total_regions * 0.5:
return {
"alert": True,
"severity": "critical",
"status": "major_outage",
"message": f"{check_name}: {failing_count}/{total_regions} regions failing",
"affected_regions": [r.region for r in failing_regions],
}
if failing_count >= 2:
return {
"alert": True,
"severity": "warning",
"status": "partial_outage",
"message": f"{check_name}: {failing_count} regions degraded",
"affected_regions": [r.region for r in failing_regions],
}
# Single region - higher threshold
single_failure = failing_regions[0]
if single_failure.consecutive_failures >= 3:
return {
"alert": True,
"severity": "warning",
"status": "regional_issue",
"message": f"{check_name}: {single_failure.region} failing consistently",
}
return {"alert": False, "status": "monitoring"}
Browser-Based Checks for Critical Flows
HTTP checks catch availability issues. Browser checks catch functionality issues:
// synthetic-runner/src/browser-checks/checkout-flow.ts
import { chromium, Page } from 'playwright';
interface BrowserCheckStep {
name: string;
durationMs: number;
passed: boolean;
screenshot?: string;
}
async function checkoutFlowCheck(): Promise<BrowserCheckStep[]> {
const browser = await chromium.launch({ headless: true });
const page = await browser.newPage();
const steps: BrowserCheckStep[] = [];
try {
// Step 1: Load pricing page
let start = Date.now();
await page.goto('https://www.example.com/pricing');
await page.waitForSelector('[data-testid="plan-card"]');
steps.push({
name: 'pricing_page_loaded',
durationMs: Date.now() - start,
passed: true,
});
// Step 2: Select a plan
start = Date.now();
await page.click('[data-testid="plan-pro-cta"]');
await page.waitForSelector('#checkout-form');
steps.push({
name: 'checkout_form_loaded',
durationMs: Date.now() - start,
passed: true,
});
// Step 3: Verify payment provider loads
start = Date.now();
const stripeFrame = page.frameLocator('iframe[name*="stripe"]');
await stripeFrame.locator('#card-number').waitFor({ timeout: 10000 });
steps.push({
name: 'payment_provider_loaded',
durationMs: Date.now() - start,
passed: true,
});
} catch (error) {
steps.push({
name: 'flow_error',
durationMs: 0,
passed: false,
screenshot: await page.screenshot({ encoding: 'base64' }),
});
} finally {
await browser.close();
}
return steps;
}
Results
| Metric | Before | After |
|---|---|---|
| Mean time to detect (regional) | 23 min | 47 sec |
| Mean time to detect (global) | 4 min | 12 sec |
| Incidents detected before user reports | 34% | 91% |
| CDN issues detected proactively | 0% | 100% |
| False positive rate | — | 2.1% |
| SSL expiry incidents | 3/year | 0/year |
| Third-party outage awareness | Minutes-to-hours | Seconds |
Key Takeaways
-
Synthetic monitoring fills the gaps that server metrics and RUM cannot. CDN failures, DNS issues, and network partitions are invisible to origin-based monitoring.
-
12 regions provide global coverage with reasonable cost. You don't need probes everywhere—strategic placement in major user regions catches 99% of regional issues.
-
Correlate across regions before alerting. A single-region failure needs more evidence than a multi-region failure. This eliminates false positives from transient network issues.
-
Browser checks catch what HTTP checks miss. A 200 status code doesn't mean the page works. Critical user flows need end-to-end browser-based validation.
-
Synthetic monitoring is your early warning system. The goal isn't to replace other monitoring—it's to detect problems in the 90-second window between "something broke" and "users noticed."
The best incident response starts before users are impacted. Synthetic monitoring from 12 regions gives you that head start.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.