Automating AWS IAM Least Privilege with Access Analyzer: From Over-Permissioned to Locked Down
How we reduced our IAM policy permissions by 89% using Access Analyzer policy generation and automated continuous enforcement without breaking production workloads.

Every AWS account accumulates IAM debt. Developers create roles with broad permissions during prototyping, those roles ship to production, and nobody revisits them. Our security audit revealed 340 IAM roles across 12 accounts. 78% had permissions they never used. 23 roles had full *:* access. One Lambda function processing thumbnails had AdministratorAccess.
We built an automated pipeline using IAM Access Analyzer that generated least-privilege policies from actual usage, deployed them through a staged rollout, and continuously monitors for drift. Total permissions across all roles dropped 89% in 90 days with zero production incidents.
The Problem: Permission Sprawl at Scale
Manual IAM policy review does not scale. Our audit found:
- 340 IAM roles across 12 AWS accounts
- Average role had 47 allowed actions; average actual usage was 6 actions
- 23 roles with AdministratorAccess or PowerUserAccess
- 156 roles with wildcard resource permissions (
Resource: "*") - Last policy review: 14 months ago
The risk is not theoretical. A compromised Lambda with s3:* on * gives an attacker access to every bucket in the account. Least privilege is not security theater; it is blast radius containment.
Architecture: The Least Privilege Pipeline
Our automation pipeline has four stages:
- Observation — Access Analyzer monitors CloudTrail for 30 days per role
- Generation — Generate scoped policies from observed API calls
- Validation — Dry-run generated policies against known workload patterns
- Deployment — Staged rollout with automatic rollback on access denied errors
Stage 1: Enabling CloudTrail-Based Policy Generation
Access Analyzer needs CloudTrail data to understand what each role actually does. We ensured organization-wide CloudTrail was logging all management and data events:
import {
AccessAnalyzerClient,
StartPolicyGenerationCommand,
GetGeneratedPolicyCommand,
} from '@aws-sdk/client-accessanalyzer';
interface RoleAnalysis {
roleArn: string;
roleName: string;
accountId: string;
observationDays: number;
}
async function generateLeastPrivilegePolicy(
role: RoleAnalysis
): Promise<string> {
const client = new AccessAnalyzerClient({ region: 'us-east-1' });
const startDate = new Date();
startDate.setDate(startDate.getDate() - role.observationDays);
const startCommand = new StartPolicyGenerationCommand({
policyGenerationDetails: {
principalArn: role.roleArn,
},
cloudTrailDetails: {
trails: [
{
cloudTrailArn: `arn:aws:cloudtrail:us-east-1:${role.accountId}:trail/org-trail`,
regions: ['us-east-1', 'eu-west-1'],
allRegions: false,
},
],
accessRole: `arn:aws:iam::${role.accountId}:role/AccessAnalyzerRole`,
startTime: startDate,
endTime: new Date(),
},
});
const { jobId } = await client.send(startCommand);
// Poll for completion
let status = 'IN_PROGRESS';
let generatedPolicy: string = '';
while (status === 'IN_PROGRESS') {
await new Promise((r) => setTimeout(r, 5000));
const getCommand = new GetGeneratedPolicyCommand({ jobId: jobId! });
const result = await client.send(getCommand);
status = result.jobDetails?.status || 'FAILED';
if (status === 'SUCCEEDED') {
generatedPolicy = JSON.stringify(
result.generatedPolicyResult?.generatedPolicies?.[0]?.policy
);
}
}
return generatedPolicy;
}
Stage 2: Policy Enhancement and Validation
Access Analyzer generates policies based only on observed actions. But some actions are legitimately needed but rarely used — disaster recovery procedures, monthly batch jobs, or seasonal scaling operations. We augment generated policies with declared exceptions:
interface PolicyAugmentation {
roleName: string;
requiredButRareActions: {
action: string;
resource: string;
justification: string;
lastUsed?: string;
reviewDate: string;
}[];
}
function augmentGeneratedPolicy(
generatedPolicy: IAMPolicy,
augmentations: PolicyAugmentation
): IAMPolicy {
const augmentedStatements = augmentations.requiredButRareActions.map(
(aug) => ({
Sid: `RareAction_${aug.action.replace(/[^a-zA-Z]/g, '')}`,
Effect: 'Allow' as const,
Action: [aug.action],
Resource: [aug.resource],
Condition: {
StringEquals: {
'aws:PrincipalTag/team': 'platform-engineering',
},
},
})
);
return {
...generatedPolicy,
Statement: [...generatedPolicy.Statement, ...augmentedStatements],
};
}
Each augmentation requires a justification and a review date. A monthly Lambda flags any augmentation past its review date for re-evaluation.
Stage 3: Staged Rollout with Boundary Policies
We never replace a policy directly. Instead, we use permission boundaries as a safety net:
- Day 1-7: Deploy generated policy as a permission boundary (intersection with existing policy)
- Day 7-14: Monitor for
AccessDeniedevents in CloudTrail - Day 14: If zero denials, replace the inline/attached policy with the generated one
- Day 14-28: Remove the old overly-broad policy, keep monitoring
If any AccessDenied occurs during the boundary phase, we automatically:
- Log the denied action and resource
- Add it to the policy with a review flag
- Alert the owning team for confirmation
Results: 90-Day Transformation
| Metric | Before | After | Change |
|---|---|---|---|
| Total allowed actions (all roles) | 15,980 | 1,734 | -89.1% |
Roles with *:* access | 23 | 0 | -100% |
| Roles with wildcard resources | 156 | 12 | -92.3% |
| Average actions per role | 47 | 5.1 | -89.1% |
| Production access denied incidents | N/A | 0 | — |
| Time to generate policy per role | Manual (hours) | 4.2 min | — |
| Monthly pipeline cost | $0 | $23 | — |
The 12 remaining roles with wildcard resources are legitimate cases (CloudFormation deployment roles that create arbitrary resource types). These are scoped to specific services and tagged for quarterly review.
Continuous Enforcement: Preventing Drift
The hardest part is not the initial remediation — it is preventing regression. We implemented three guardrails:
1. Service Control Policies (SCPs) block the creation of new roles with *:* access at the organization level.
2. IAM Access Analyzer findings are streamed to Security Hub. Any public or cross-account access generates a P1 alert.
3. Weekly drift detection compares current policies against the last-audited baseline:
async function detectPolicyDrift(
roleName: string,
baselinePolicy: IAMPolicy,
currentPolicy: IAMPolicy
): Promise<DriftReport> {
const baselineActions = new Set(
baselinePolicy.Statement.flatMap((s) =>
s.Effect === 'Allow' ? s.Action : []
)
);
const currentActions = new Set(
currentPolicy.Statement.flatMap((s) =>
s.Effect === 'Allow' ? s.Action : []
)
);
const addedActions = [...currentActions].filter(
(a) => !baselineActions.has(a)
);
const removedActions = [...baselineActions].filter(
(a) => !currentActions.has(a)
);
return {
roleName,
driftDetected: addedActions.length > 0 || removedActions.length > 0,
addedActions,
removedActions,
severity: addedActions.some((a) => a.includes('*')) ? 'HIGH' : 'MEDIUM',
detectedAt: new Date().toISOString(),
};
}
Lessons Learned
30 days is the minimum observation window. We initially tried 14 days and missed monthly batch jobs. One role lost access to its end-of-month reporting Lambda. 30 days catches most patterns; 90 days catches quarterly jobs.
Tag roles with ownership from day one. When we needed to contact teams about policy changes, 40% of roles had no ownership tags. We spent two weeks playing detective. Now every role requires a team and service tag at creation via SCP enforcement.
Separate deployment roles from runtime roles. CDK/CloudFormation deployment roles legitimately need broad permissions. Mixing deployment and runtime concerns in a single role makes least privilege impossible. We split every combined role into a deployment role (broad, used only in CI/CD) and a runtime role (minimal, used by the service).
Permission boundaries are your safety net. Never go directly from broad-to-narrow. The boundary phase gives you a risk-free testing period. We caught 14 missing permissions during boundary testing that would have caused production issues.
Conclusion
IAM least privilege at scale requires automation. Manual review of 340 roles would take months of engineering time and be outdated before completion. Our Access Analyzer pipeline generates, validates, and deploys scoped policies in under 5 minutes per role, maintains zero production incidents across 90 days of rollout, and costs $23/month to run. The security posture improvement is dramatic: if any single role is compromised, the blast radius is now limited to exactly the resources that role needs, not the entire AWS account.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.