AWS WAF Rate Limiting in Production: Protecting APIs Without Blocking Legitimate Traffic

How we configured AWS WAF rate-based rules to stop credential stuffing and API abuse while maintaining 99.99% availability for real users.

#aws#waf#security#rate-limiting
Cover image for the article: AWS WAF Rate Limiting in Production: Protecting APIs Without Blocking Legitimate Traffic

Last quarter, one of our payment APIs started returning 503s during peak hours. CloudWatch showed request counts 40x above baseline. A credential stuffing attack was hitting our authentication endpoint at 12,000 requests per second from a botnet spanning 2,400 unique IPs. Our naive IP-based rate limiting was useless against distributed attacks.

After deploying a layered AWS WAF strategy, we reduced malicious traffic by 97.3% while maintaining a false positive rate below 0.02%. Here is exactly how we built it.

The Problem: Distributed Attacks Bypass Simple Rate Limits

Traditional rate limiting works on a per-IP basis. Set a threshold of 100 requests per 5 minutes, and any single IP exceeding that gets blocked. But modern botnets rotate through thousands of residential proxies. Each individual IP stays well under the threshold while the aggregate traffic overwhelms your backend.

Our initial setup looked like this:

  • CloudFront distribution in front of API Gateway
  • A single WAF rate-based rule: 2,000 requests per 5 minutes per IP
  • No request inspection or pattern matching

The attack sailed right through. Each bot IP sent only 5 requests per second, individually unremarkable, but collectively devastating.

WAF Traffic Analysis

Architecture: Layered WAF Rule Groups

We restructured our WAF ACL into four distinct rule groups evaluated in priority order:

Layer 1 — Known Bad Actors (Priority 0) IP reputation lists and AWS Managed Rules for known botnets. These block before any compute happens.

Layer 2 — Rate-Based Rules (Priority 1) Multiple rate-based rules with different scopes and thresholds.

Layer 3 — Request Pattern Matching (Priority 2) Regex rules targeting specific attack signatures in headers and body.

Layer 4 — Geographic and Token-Based Rules (Priority 3) Country-level restrictions and custom header validation for authenticated endpoints.

Implementation: Multi-Dimensional Rate Limiting

The key insight was combining rate-based rules with scope-down statements. Instead of one global rate limit, we created endpoint-specific rules with different thresholds:

import * as cdk from 'aws-cdk-lib';
import * as wafv2 from 'aws-cdk-lib/aws-wafv2';

const webAcl = new wafv2.CfnWebACL(this, 'ApiWaf', {
  defaultAction: { allow: {} },
  scope: 'REGIONAL',
  visibilityConfig: {
    cloudWatchMetricsEnabled: true,
    metricName: 'ApiWafMetrics',
    sampledRequestsEnabled: true,
  },
  rules: [
    {
      name: 'RateLimitAuthEndpoint',
      priority: 1,
      action: { block: {} },
      statement: {
        rateBasedStatement: {
          limit: 100,
          aggregateKeyType: 'IP',
          scopeDownStatement: {
            byteMatchStatement: {
              fieldToMatch: { uriPath: {} },
              positionalConstraint: 'STARTS_WITH',
              searchString: '/api/v1/auth',
              textTransformations: [{ priority: 0, type: 'LOWERCASE' }],
            },
          },
        },
      },
      visibilityConfig: {
        cloudWatchMetricsEnabled: true,
        metricName: 'RateLimitAuth',
        sampledRequestsEnabled: true,
      },
    },
    {
      name: 'RateLimitByHeaderFingerprint',
      priority: 2,
      action: { block: {} },
      statement: {
        rateBasedStatement: {
          limit: 500,
          aggregateKeyType: 'FORWARDED_IP',
          forwardedIpConfig: {
            headerName: 'X-Forwarded-For',
            fallbackBehavior: 'MATCH',
          },
          scopeDownStatement: {
            notStatement: {
              statement: {
                byteMatchStatement: {
                  fieldToMatch: {
                    singleHeader: { name: 'x-api-key' },
                  },
                  positionalConstraint: 'EXACTLY',
                  searchString: '',
                  textTransformations: [{ priority: 0, type: 'NONE' }],
                },
              },
            },
          },
        },
      },
      visibilityConfig: {
        cloudWatchMetricsEnabled: true,
        metricName: 'RateLimitFingerprint',
        sampledRequestsEnabled: true,
      },
    },
  ],
});

The second rule is critical. It rate-limits by forwarded IP but only applies to requests missing a valid API key header. Legitimate clients always include their key; bots scraping the endpoint rarely do.

Custom Response Bodies for Rate-Limited Clients

When WAF blocks a request, returning a generic 403 frustrates legitimate users who accidentally trigger limits. We configured custom response bodies with retry guidance:

{
  "customResponseBodies": {
    "RateLimitResponse": {
      "contentType": "APPLICATION_JSON",
      "content": "{\"error\":\"rate_limit_exceeded\",\"message\":\"Too many requests. Please retry after the period specified in Retry-After header.\",\"retryAfter\":60,\"documentation\":\"https://docs.api.example.com/rate-limits\"}"
    }
  }
}

This pairs with a Retry-After header so well-behaved clients back off automatically rather than hammering the endpoint.

Benchmarks: Attack Mitigation Results

After deploying the layered WAF configuration, we collected 30 days of data across three separate attack campaigns:

MetricBefore WAFAfter WAFImprovement
Malicious requests blocked0%97.3%—
False positive rateN/A0.018%—
Auth endpoint P99 latency2,340ms89ms96.2% reduction
Backend error rate (5xx)4.7%0.03%99.4% reduction
Monthly WAF cost$0$847—
Prevented fraud losses (est.)$180K/mo$4.8K/mo97.3% reduction

The $847 monthly WAF cost paid for itself many times over in prevented fraud and reduced compute costs from not processing malicious requests.

Monitoring and Adaptive Thresholds

Static thresholds break during traffic spikes like product launches or marketing campaigns. We built an adaptive system using CloudWatch Metric Math:

  • Baseline calculation: rolling 7-day average of requests per endpoint
  • Alert threshold: 3x baseline triggers review, 10x triggers automatic rule tightening
  • Weekly threshold review via Lambda function that adjusts WAF rule limits based on legitimate traffic patterns

The Lambda runs every Monday at 06:00 UTC, pulls the previous week's P95 legitimate traffic per endpoint, and updates WAF rules to set limits at 2x that P95 value. This ensures thresholds track organic growth.

WAF Adaptive Thresholds

Lessons Learned

Start with COUNT mode. Deploy every new WAF rule in COUNT (observe) mode for at least 72 hours before switching to BLOCK. We caught three rules that would have blocked our monitoring service.

Log everything to S3 via Kinesis Firehose. WAF sampled requests only show a subset. Full logging to S3 costs pennies and is invaluable during incident investigation.

Test with realistic load. We run weekly synthetic attack simulations using distributed load generators. If your WAF rules only get tested during real attacks, you will find gaps at the worst possible time.

Coordinate with your CDN. CloudFront and WAF integrate tightly, but ordering matters. Put bot control managed rules before your custom rules to avoid paying for inspection of obviously-bad traffic.

Conclusion

AWS WAF rate limiting is not a checkbox you tick once. It is an evolving defense that requires endpoint-specific thresholds, multi-dimensional aggregation beyond just IP addresses, and continuous monitoring to avoid blocking legitimate traffic during growth periods. The $847/month we spend on WAF saves us over $175K in fraud prevention and keeps our P99 latencies under 100ms even during active attacks. For any production API handling authentication or payments, layered WAF rules are non-negotiable infrastructure.

Comments

    No comments yet. Be the first to share your thoughts.