Multi-Account Strategy for 50+ AWS Accounts with Automated Guardrails

How we designed and implemented an AWS Organizations structure for 50+ accounts with SCPs, automated provisioning, and centralized governance at scale.

#aws#organizations#multi-account#security#governance
Cover image for the article: Multi-Account Strategy for 50+ AWS Accounts with Automated Guardrails

At 12 engineers, we had 3 AWS accounts. At 80 engineers, we had 53. The transition from "a few accounts someone created manually" to "a governed multi-account estate" nearly broke us. IAM policies conflicted, costs were unattributable, and one team's misconfigured S3 bucket exposed another team's data. This is how we built the structure that brought order to the chaos.

Why Multi-Account at All?

Single-account architectures collapse under organizational growth. The breaking points are predictable:

SignalWhat BreaksAccount Boundary Solves
5+ teams sharing VPCsNetwork conflicts, security group sprawlNetwork isolation per workload
IAM policies > 100Permission complexity, audit failuresBlast radius containment
Single bill > $100K/monthCost attribution impossiblePer-account billing clarity
Compliance requirementsPCI/SOC2 scope creepRegulatory boundary isolation
10+ production servicesOne team's incident affects allFailure domain isolation

The AWS Well-Architected Framework recommends multi-account by default. We learned why the hard way.

Organizational Unit Structure

After evaluating several structures, we settled on a workload-centric OU hierarchy:

Root
├── Security OU
│   ├── Log Archive (centralized CloudTrail, Config)
│   ├── Security Tooling (GuardDuty, Security Hub)
│   └── Audit (read-only access for compliance)
├── Infrastructure OU
│   ├── Network Hub (Transit Gateway, DNS)
│   ├── Shared Services (CI/CD, artifact repos)
│   └── Identity (SSO, directory services)
├── Workloads OU
│   ├── Production OU
│   │   ├── Team-A-Prod
│   │   ├── Team-B-Prod
│   │   └── Team-C-Prod
│   ├── Staging OU
│   │   ├── Team-A-Staging
│   │   └── Team-B-Staging
│   └── Development OU
│       ├── Team-A-Dev
│       └── Team-B-Dev
├── Sandbox OU
│   └── Individual developer sandboxes
└── Suspended OU
    └── Accounts pending decommission

AWS Organizations Structure

Service Control Policies: Guardrails That Scale

SCPs are the backbone of governance. We implemented them in layers:

Layer 1: Organization-Wide Denies

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "DenyRegionsOutsideAllowed",
      "Effect": "Deny",
      "Action": "*",
      "Resource": "*",
      "Condition": {
        "StringNotEquals": {
          "aws:RequestedRegion": [
            "us-east-1",
            "us-west-2",
            "eu-west-1",
            "ap-southeast-1"
          ]
        },
        "ArnNotLike": {
          "aws:PrincipalArn": [
            "arn:aws:iam::*:role/OrganizationAdmin"
          ]
        }
      }
    },
    {
      "Sid": "DenyLeaveOrganization",
      "Effect": "Deny",
      "Action": "organizations:LeaveOrganization",
      "Resource": "*"
    },
    {
      "Sid": "DenyDisableCloudTrail",
      "Effect": "Deny",
      "Action": [
        "cloudtrail:StopLogging",
        "cloudtrail:DeleteTrail"
      ],
      "Resource": "*"
    }
  ]
}

Layer 2: Production OU Restrictions

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "DenyPublicS3InProd",
      "Effect": "Deny",
      "Action": [
        "s3:PutBucketPublicAccessBlock"
      ],
      "Resource": "*",
      "Condition": {
        "StringNotEquals": {
          "s3:x-amz-acl": "private"
        }
      }
    },
    {
      "Sid": "RequireIMDSv2",
      "Effect": "Deny",
      "Action": "ec2:RunInstances",
      "Resource": "arn:aws:ec2:*:*:instance/*",
      "Condition": {
        "StringNotEquals": {
          "ec2:MetadataHttpTokens": "required"
        }
      }
    },
    {
      "Sid": "DenyUnencryptedVolumes",
      "Effect": "Deny",
      "Action": "ec2:CreateVolume",
      "Resource": "*",
      "Condition": {
        "Bool": {
          "ec2:Encrypted": "false"
        }
      }
    }
  ]
}

Layer 3: Sandbox OU Permissions

Sandboxes get maximum freedom with cost constraints:

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "DenyExpensiveInstances",
      "Effect": "Deny",
      "Action": "ec2:RunInstances",
      "Resource": "arn:aws:ec2:*:*:instance/*",
      "Condition": {
        "ForAnyValue:StringNotLike": {
          "ec2:InstanceType": [
            "t3.*", "t3a.*", "t4g.*", "m5.large", "m5.xlarge"
          ]
        }
      }
    },
    {
      "Sid": "DenyRDSLargeInstances",
      "Effect": "Deny",
      "Action": "rds:CreateDBInstance",
      "Resource": "*",
      "Condition": {
        "ForAnyValue:StringNotLike": {
          "rds:DatabaseClass": [
            "db.t3.*", "db.t4g.*"
          ]
        }
      }
    }
  ]
}

Account Factory: Automated Provisioning

Manual account creation doesn't scale past 10 accounts. We built an account factory using AWS Control Tower with custom customizations:

// Account vending machine - triggered by Backstage service catalog
import { OrganizationsClient, CreateAccountCommand } from '@aws-sdk/client-organizations';
import { ServiceCatalogClient, ProvisionProductCommand } from '@aws-sdk/client-service-catalog';

interface AccountRequest {
  teamName: string;
  environment: 'production' | 'staging' | 'development' | 'sandbox';
  costCenter: string;
  ownerEmail: string;
  vpcCidr: string;
}

async function provisionAccount(request: AccountRequest): Promise<string> {
  const accountName = `${request.teamName}-${request.environment}`;
  const ouId = getOuForEnvironment(request.environment);

  // Provision via Control Tower Account Factory
  const product = await serviceCatalog.send(new ProvisionProductCommand({
    ProductId: ACCOUNT_FACTORY_PRODUCT_ID,
    ProvisioningArtifactId: LATEST_VERSION_ID,
    ProvisionedProductName: accountName,
    ProvisioningParameters: [
      { Key: 'AccountName', Value: accountName },
      { Key: 'AccountEmail', Value: `aws+${accountName}@company.com` },
      { Key: 'ManagedOrganizationalUnit', Value: ouId },
      { Key: 'SSOUserEmail', Value: request.ownerEmail },
      { Key: 'SSOUserFirstName', Value: 'Account' },
      { Key: 'SSOUserLastName', Value: 'Admin' },
    ],
    Tags: [
      { Key: 'team', Value: request.teamName },
      { Key: 'environment', Value: request.environment },
      { Key: 'cost-center', Value: request.costCenter },
    ],
  }));

  // Post-provisioning: apply baseline via StackSets
  await applyBaseline(accountName, request);

  return product.RecordDetail?.ProvisionedProductId || '';
}

A new account goes from request to fully configured in 18 minutes:

StepTimeAutomated
Account creation5 minYes - Control Tower
OU placement + SCPsInstantYes - Control Tower
VPC + networking baseline4 minYes - StackSets
IAM roles + SSO config2 minYes - StackSets
Security baseline (GuardDuty, Config)3 minYes - StackSets
Cost alerting + tagging2 minYes - StackSets
DNS delegation2 minYes - Lambda
Total18 min100%

Networking: Hub-and-Spoke with Transit Gateway

With 50+ accounts, VPC peering becomes unmanageable (n*(n-1)/2 connections). We use Transit Gateway as a hub:

# Transit Gateway in the Network Hub account
resource "aws_ec2_transit_gateway" "main" {
  description                     = "Organization Transit Gateway"
  default_route_table_association = "disable"
  default_route_table_propagation = "disable"
  dns_support                     = "enable"
  vpn_ecmp_support               = "enable"

  tags = {
    Name = "org-transit-gateway"
  }
}

# Separate route tables for isolation
resource "aws_ec2_transit_gateway_route_table" "production" {
  transit_gateway_id = aws_ec2_transit_gateway.main.id
  tags = { Name = "production-routes" }
}

resource "aws_ec2_transit_gateway_route_table" "non_production" {
  transit_gateway_id = aws_ec2_transit_gateway.main.id
  tags = { Name = "non-production-routes" }
}

# Production cannot reach non-production and vice versa
# Shared services are accessible from both
resource "aws_ec2_transit_gateway_route_table" "shared_services" {
  transit_gateway_id = aws_ec2_transit_gateway.main.id
  tags = { Name = "shared-services-routes" }
}

Cost Governance at Scale

With 53 accounts, cost visibility requires automation:

// Daily cost anomaly detection across all accounts
async function detectCostAnomalies(): Promise<CostAnomaly[]> {
  const anomalies: CostAnomaly[] = [];

  for (const account of await getActiveAccounts()) {
    const dailyCost = await getCostForAccount(account.id, 'today');
    const avgCost = await getAverageCostForAccount(account.id, 'last-30-days');
    const threshold = avgCost * 1.5; // 50% above average

    if (dailyCost > threshold) {
      anomalies.push({
        accountId: account.id,
        accountName: account.name,
        dailyCost,
        averageCost: avgCost,
        percentageIncrease: ((dailyCost - avgCost) / avgCost) * 100,
        ownerEmail: account.tags['owner-email'],
      });
    }
  }

  return anomalies;
}

Monthly cost distribution across our 53 accounts:

OU CategoryAccountsMonthly Spend% of Total
Production12$89,00062%
Infrastructure4$23,00016%
Staging8$14,00010%
Development15$11,0008%
Sandbox14$6,0004%

Cost Distribution by OU

Lessons Learned at 50+ Accounts

1. Start with strong naming conventions. Account names should encode team, environment, and purpose. We use {team}-{env}-{purpose} (e.g., payments-prod-api).

2. Email management is harder than you think. Every AWS account needs a unique email. We use aws+{account-name}@company.com with a shared inbox.

3. SCPs are not IAM policies. SCPs set maximum permissions — they don't grant anything. Teams still need proper IAM roles within their accounts.

4. Plan CIDR ranges upfront. We allocated a /8 and subdivided it systematically. Overlapping CIDRs across accounts make networking impossible.

5. Automate account closure. We have 6 accounts in our Suspended OU awaiting the 90-day closure period. Without automation, these zombie accounts accumulate.

Key Takeaways

  1. Multi-account is not optional at scale — the blast radius, compliance, and cost attribution benefits justify the governance investment.

  2. SCPs are your most powerful security tool — they prevent entire classes of misconfiguration regardless of IAM policies within accounts.

  3. Automate account provisioning from day one — manual account setup creates inconsistency and drift that compounds over time.

  4. Transit Gateway replaces VPC peering — at 5+ accounts, hub-and-spoke networking is the only manageable pattern.

  5. Treat accounts as cattle, not pets — with automated provisioning, spinning up a new account should take minutes, not days.

  6. Governance enables autonomy — teams move faster when guardrails prevent catastrophic mistakes. SCPs let you say "yes" to team freedom because the boundaries are enforced.

The investment in multi-account governance pays dividends in security, compliance, and team velocity. It's infrastructure for your infrastructure — and at 50+ accounts, it's non-negotiable.

Comments

    No comments yet. Be the first to share your thoughts.