Multi-Account Strategy for 50+ AWS Accounts with Automated Guardrails
How we designed and implemented an AWS Organizations structure for 50+ accounts with SCPs, automated provisioning, and centralized governance at scale.

At 12 engineers, we had 3 AWS accounts. At 80 engineers, we had 53. The transition from "a few accounts someone created manually" to "a governed multi-account estate" nearly broke us. IAM policies conflicted, costs were unattributable, and one team's misconfigured S3 bucket exposed another team's data. This is how we built the structure that brought order to the chaos.
Why Multi-Account at All?
Single-account architectures collapse under organizational growth. The breaking points are predictable:
| Signal | What Breaks | Account Boundary Solves |
|---|---|---|
| 5+ teams sharing VPCs | Network conflicts, security group sprawl | Network isolation per workload |
| IAM policies > 100 | Permission complexity, audit failures | Blast radius containment |
| Single bill > $100K/month | Cost attribution impossible | Per-account billing clarity |
| Compliance requirements | PCI/SOC2 scope creep | Regulatory boundary isolation |
| 10+ production services | One team's incident affects all | Failure domain isolation |
The AWS Well-Architected Framework recommends multi-account by default. We learned why the hard way.
Organizational Unit Structure
After evaluating several structures, we settled on a workload-centric OU hierarchy:
Root
├── Security OU
│ ├── Log Archive (centralized CloudTrail, Config)
│ ├── Security Tooling (GuardDuty, Security Hub)
│ └── Audit (read-only access for compliance)
├── Infrastructure OU
│ ├── Network Hub (Transit Gateway, DNS)
│ ├── Shared Services (CI/CD, artifact repos)
│ └── Identity (SSO, directory services)
├── Workloads OU
│ ├── Production OU
│ │ ├── Team-A-Prod
│ │ ├── Team-B-Prod
│ │ └── Team-C-Prod
│ ├── Staging OU
│ │ ├── Team-A-Staging
│ │ └── Team-B-Staging
│ └── Development OU
│ ├── Team-A-Dev
│ └── Team-B-Dev
├── Sandbox OU
│ └── Individual developer sandboxes
└── Suspended OU
└── Accounts pending decommission
Service Control Policies: Guardrails That Scale
SCPs are the backbone of governance. We implemented them in layers:
Layer 1: Organization-Wide Denies
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "DenyRegionsOutsideAllowed",
"Effect": "Deny",
"Action": "*",
"Resource": "*",
"Condition": {
"StringNotEquals": {
"aws:RequestedRegion": [
"us-east-1",
"us-west-2",
"eu-west-1",
"ap-southeast-1"
]
},
"ArnNotLike": {
"aws:PrincipalArn": [
"arn:aws:iam::*:role/OrganizationAdmin"
]
}
}
},
{
"Sid": "DenyLeaveOrganization",
"Effect": "Deny",
"Action": "organizations:LeaveOrganization",
"Resource": "*"
},
{
"Sid": "DenyDisableCloudTrail",
"Effect": "Deny",
"Action": [
"cloudtrail:StopLogging",
"cloudtrail:DeleteTrail"
],
"Resource": "*"
}
]
}
Layer 2: Production OU Restrictions
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "DenyPublicS3InProd",
"Effect": "Deny",
"Action": [
"s3:PutBucketPublicAccessBlock"
],
"Resource": "*",
"Condition": {
"StringNotEquals": {
"s3:x-amz-acl": "private"
}
}
},
{
"Sid": "RequireIMDSv2",
"Effect": "Deny",
"Action": "ec2:RunInstances",
"Resource": "arn:aws:ec2:*:*:instance/*",
"Condition": {
"StringNotEquals": {
"ec2:MetadataHttpTokens": "required"
}
}
},
{
"Sid": "DenyUnencryptedVolumes",
"Effect": "Deny",
"Action": "ec2:CreateVolume",
"Resource": "*",
"Condition": {
"Bool": {
"ec2:Encrypted": "false"
}
}
}
]
}
Layer 3: Sandbox OU Permissions
Sandboxes get maximum freedom with cost constraints:
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "DenyExpensiveInstances",
"Effect": "Deny",
"Action": "ec2:RunInstances",
"Resource": "arn:aws:ec2:*:*:instance/*",
"Condition": {
"ForAnyValue:StringNotLike": {
"ec2:InstanceType": [
"t3.*", "t3a.*", "t4g.*", "m5.large", "m5.xlarge"
]
}
}
},
{
"Sid": "DenyRDSLargeInstances",
"Effect": "Deny",
"Action": "rds:CreateDBInstance",
"Resource": "*",
"Condition": {
"ForAnyValue:StringNotLike": {
"rds:DatabaseClass": [
"db.t3.*", "db.t4g.*"
]
}
}
}
]
}
Account Factory: Automated Provisioning
Manual account creation doesn't scale past 10 accounts. We built an account factory using AWS Control Tower with custom customizations:
// Account vending machine - triggered by Backstage service catalog
import { OrganizationsClient, CreateAccountCommand } from '@aws-sdk/client-organizations';
import { ServiceCatalogClient, ProvisionProductCommand } from '@aws-sdk/client-service-catalog';
interface AccountRequest {
teamName: string;
environment: 'production' | 'staging' | 'development' | 'sandbox';
costCenter: string;
ownerEmail: string;
vpcCidr: string;
}
async function provisionAccount(request: AccountRequest): Promise<string> {
const accountName = `${request.teamName}-${request.environment}`;
const ouId = getOuForEnvironment(request.environment);
// Provision via Control Tower Account Factory
const product = await serviceCatalog.send(new ProvisionProductCommand({
ProductId: ACCOUNT_FACTORY_PRODUCT_ID,
ProvisioningArtifactId: LATEST_VERSION_ID,
ProvisionedProductName: accountName,
ProvisioningParameters: [
{ Key: 'AccountName', Value: accountName },
{ Key: 'AccountEmail', Value: `aws+${accountName}@company.com` },
{ Key: 'ManagedOrganizationalUnit', Value: ouId },
{ Key: 'SSOUserEmail', Value: request.ownerEmail },
{ Key: 'SSOUserFirstName', Value: 'Account' },
{ Key: 'SSOUserLastName', Value: 'Admin' },
],
Tags: [
{ Key: 'team', Value: request.teamName },
{ Key: 'environment', Value: request.environment },
{ Key: 'cost-center', Value: request.costCenter },
],
}));
// Post-provisioning: apply baseline via StackSets
await applyBaseline(accountName, request);
return product.RecordDetail?.ProvisionedProductId || '';
}
A new account goes from request to fully configured in 18 minutes:
| Step | Time | Automated |
|---|---|---|
| Account creation | 5 min | Yes - Control Tower |
| OU placement + SCPs | Instant | Yes - Control Tower |
| VPC + networking baseline | 4 min | Yes - StackSets |
| IAM roles + SSO config | 2 min | Yes - StackSets |
| Security baseline (GuardDuty, Config) | 3 min | Yes - StackSets |
| Cost alerting + tagging | 2 min | Yes - StackSets |
| DNS delegation | 2 min | Yes - Lambda |
| Total | 18 min | 100% |
Networking: Hub-and-Spoke with Transit Gateway
With 50+ accounts, VPC peering becomes unmanageable (n*(n-1)/2 connections). We use Transit Gateway as a hub:
# Transit Gateway in the Network Hub account
resource "aws_ec2_transit_gateway" "main" {
description = "Organization Transit Gateway"
default_route_table_association = "disable"
default_route_table_propagation = "disable"
dns_support = "enable"
vpn_ecmp_support = "enable"
tags = {
Name = "org-transit-gateway"
}
}
# Separate route tables for isolation
resource "aws_ec2_transit_gateway_route_table" "production" {
transit_gateway_id = aws_ec2_transit_gateway.main.id
tags = { Name = "production-routes" }
}
resource "aws_ec2_transit_gateway_route_table" "non_production" {
transit_gateway_id = aws_ec2_transit_gateway.main.id
tags = { Name = "non-production-routes" }
}
# Production cannot reach non-production and vice versa
# Shared services are accessible from both
resource "aws_ec2_transit_gateway_route_table" "shared_services" {
transit_gateway_id = aws_ec2_transit_gateway.main.id
tags = { Name = "shared-services-routes" }
}
Cost Governance at Scale
With 53 accounts, cost visibility requires automation:
// Daily cost anomaly detection across all accounts
async function detectCostAnomalies(): Promise<CostAnomaly[]> {
const anomalies: CostAnomaly[] = [];
for (const account of await getActiveAccounts()) {
const dailyCost = await getCostForAccount(account.id, 'today');
const avgCost = await getAverageCostForAccount(account.id, 'last-30-days');
const threshold = avgCost * 1.5; // 50% above average
if (dailyCost > threshold) {
anomalies.push({
accountId: account.id,
accountName: account.name,
dailyCost,
averageCost: avgCost,
percentageIncrease: ((dailyCost - avgCost) / avgCost) * 100,
ownerEmail: account.tags['owner-email'],
});
}
}
return anomalies;
}
Monthly cost distribution across our 53 accounts:
| OU Category | Accounts | Monthly Spend | % of Total |
|---|---|---|---|
| Production | 12 | $89,000 | 62% |
| Infrastructure | 4 | $23,000 | 16% |
| Staging | 8 | $14,000 | 10% |
| Development | 15 | $11,000 | 8% |
| Sandbox | 14 | $6,000 | 4% |
Lessons Learned at 50+ Accounts
1. Start with strong naming conventions. Account names should encode team, environment, and purpose. We use {team}-{env}-{purpose} (e.g., payments-prod-api).
2. Email management is harder than you think. Every AWS account needs a unique email. We use aws+{account-name}@company.com with a shared inbox.
3. SCPs are not IAM policies. SCPs set maximum permissions — they don't grant anything. Teams still need proper IAM roles within their accounts.
4. Plan CIDR ranges upfront. We allocated a /8 and subdivided it systematically. Overlapping CIDRs across accounts make networking impossible.
5. Automate account closure. We have 6 accounts in our Suspended OU awaiting the 90-day closure period. Without automation, these zombie accounts accumulate.
Key Takeaways
-
Multi-account is not optional at scale — the blast radius, compliance, and cost attribution benefits justify the governance investment.
-
SCPs are your most powerful security tool — they prevent entire classes of misconfiguration regardless of IAM policies within accounts.
-
Automate account provisioning from day one — manual account setup creates inconsistency and drift that compounds over time.
-
Transit Gateway replaces VPC peering — at 5+ accounts, hub-and-spoke networking is the only manageable pattern.
-
Treat accounts as cattle, not pets — with automated provisioning, spinning up a new account should take minutes, not days.
-
Governance enables autonomy — teams move faster when guardrails prevent catastrophic mistakes. SCPs let you say "yes" to team freedom because the boundaries are enforced.
The investment in multi-account governance pays dividends in security, compliance, and team velocity. It's infrastructure for your infrastructure — and at 50+ accounts, it's non-negotiable.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.