Managing AWS and GCP from a Single Terraform Codebase with Workspace Isolation
How to structure a multi-cloud Terraform repository using workspaces, provider aliasing, and state isolation for AWS and GCP without losing your sanity.

Most multi-cloud Terraform setups I inherit look the same: two completely separate repositories, divergent module structures, different state backends, and no shared patterns. The teams running them eventually realize they are maintaining two infrastructure platforms with none of the promised portability benefits.
After migrating a fintech platform from single-cloud AWS to a multi-cloud AWS+GCP architecture, I developed a workspace isolation strategy that keeps a single codebase manageable without sacrificing the separation guarantees that production infrastructure demands. Here is the approach, the tradeoffs, and the cost data that justified the effort.
Why Single Codebase Multi-Cloud
The argument for separate repositories is simplicity: each repo owns its cloud, its state, its CI/CD pipeline. The argument against is drift: when you implement a security pattern in AWS and forget to port it to GCP, you have an inconsistent security posture that auditors will find.
| Approach | Code Reuse | Consistency | Blast Radius | Team Overhead |
|---|---|---|---|---|
| Separate repos | None | Low | Contained | 2x review load |
| Monorepo, shared state | High | High | Dangerous | Low |
| Monorepo, isolated workspaces | Medium-High | High | Contained | Medium |
| Terragrunt wrapper | Medium | Medium | Contained | High (tooling) |
The workspace isolation approach gives you consistency guarantees (shared modules, shared policies) with blast radius containment (isolated state per cloud, per environment).
Repository Structure
infrastructure/
├── modules/
│ ├── networking/ # Cloud-agnostic interface
│ │ ├── aws/ # AWS VPC implementation
│ │ │ ├── main.tf
│ │ │ ├── variables.tf
│ │ │ └── outputs.tf
│ │ └── gcp/ # GCP VPC implementation
│ │ ├── main.tf
│ │ ├── variables.tf
│ │ └── outputs.tf
│ ├── compute/
│ │ ├── aws/ # ECS/EKS
│ │ └── gcp/ # Cloud Run/GKE
│ ├── database/
│ │ ├── aws/ # RDS/Aurora
│ │ └── gcp/ # Cloud SQL/AlloyDB
│ └── observability/
│ ├── aws/ # CloudWatch
│ └── gcp/ # Cloud Monitoring
├── environments/
│ ├── aws-prod/
│ │ ├── main.tf
│ │ ├── backend.tf
│ │ └── terraform.tfvars
│ ├── aws-staging/
│ ├── gcp-prod/
│ └── gcp-staging/
├── policies/ # OPA/Sentinel shared policies
│ ├── cost-limits.rego
│ ├── security-baseline.rego
│ └── naming-conventions.rego
└── scripts/
├── plan-all.sh
└── drift-detect.sh
Provider Configuration with Workspace Awareness
The key pattern is using workspace-aware provider configuration with strict backend isolation:
# environments/aws-prod/backend.tf
terraform {
backend "s3" {
bucket = "terraform-state-prod"
key = "aws-prod/terraform.tfstate"
region = "us-east-1"
dynamodb_table = "terraform-locks"
encrypt = true
}
required_providers {
aws = {
source = "hashicorp/aws"
version = "~> 5.40"
}
}
}
provider "aws" {
region = var.aws_region
default_tags {
tags = {
Environment = "production"
ManagedBy = "terraform"
Cloud = "aws"
CostCenter = var.cost_center
Workspace = terraform.workspace
}
}
assume_role {
role_arn = "arn:aws:iam::${var.aws_account_id}:role/TerraformExecution"
}
}
# environments/gcp-prod/backend.tf
terraform {
backend "gcs" {
bucket = "terraform-state-prod-gcp"
prefix = "gcp-prod"
}
required_providers {
google = {
source = "hashicorp/google"
version = "~> 5.20"
}
}
}
provider "google" {
project = var.gcp_project_id
region = var.gcp_region
default_labels = {
environment = "production"
managed-by = "terraform"
cloud = "gcp"
cost-center = var.cost_center
}
}
Cross-Cloud Module Interface Pattern
The power of this approach comes from defining cloud-agnostic interfaces that both implementations satisfy:
# modules/networking/aws/outputs.tf
output "network_id" {
description = "VPC ID for the primary network"
value = aws_vpc.main.id
}
output "private_subnet_ids" {
description = "Private subnet IDs across availability zones"
value = aws_subnet.private[*].id
}
output "private_cidr_blocks" {
description = "CIDR blocks for private subnets"
value = aws_subnet.private[*].cidr_block
}
output "dns_zone_id" {
description = "Private DNS zone ID"
value = aws_route53_zone.private.zone_id
}
# modules/networking/gcp/outputs.tf
output "network_id" {
description = "VPC network self_link"
value = google_compute_network.main.id
}
output "private_subnet_ids" {
description = "Private subnet self_links"
value = [for s in google_compute_subnetwork.private : s.id]
}
output "private_cidr_blocks" {
description = "CIDR blocks for private subnets"
value = [for s in google_compute_subnetwork.private : s.ip_cidr_range]
}
output "dns_zone_id" {
description = "Private DNS zone ID"
value = google_dns_managed_zone.private.id
}
Both modules expose the same output interface, which means the consuming environment code can use consistent variable names regardless of cloud.
State Isolation Strategy
State isolation is non-negotiable. A Terraform plan against GCP production must never be able to affect AWS production resources, even accidentally.
The isolation layers:
- Backend separation: S3 for AWS state, GCS for GCP state. Different credentials, different access policies.
- IAM boundaries: The AWS Terraform role cannot access GCP APIs and vice versa. No ambient credentials.
- Workspace naming convention:
{cloud}-{environment}-{region}(e.g.,aws-prod-us-east-1,gcp-prod-us-central1). - CI/CD pipeline isolation: Separate pipeline stages per cloud with independent approval gates.
# .github/workflows/terraform.yml (simplified)
jobs:
plan-aws-prod:
runs-on: ubuntu-latest
permissions:
id-token: write
steps:
- uses: aws-actions/configure-aws-credentials@v4
with:
role-to-assume: arn:aws:iam::111111111111:role/TerraformPlan
aws-region: us-east-1
- run: |
cd environments/aws-prod
terraform init
terraform plan -out=plan.tfplan
plan-gcp-prod:
runs-on: ubuntu-latest
permissions:
id-token: write
steps:
- uses: google-github-actions/auth@v2
with:
workload_identity_provider: projects/222222/locations/global/workloadIdentityPools/github/providers/github
service_account: terraform-plan@project.iam.gserviceaccount.com
- run: |
cd environments/gcp-prod
terraform init
terraform plan -out=plan.tfplan
Cost Comparison: Dual-Cloud Resource Equivalents
When running equivalent workloads across both clouds, cost normalization becomes critical for capacity planning:
| Resource Category | AWS Service | Monthly Cost | GCP Equivalent | Monthly Cost | Delta |
|---|---|---|---|---|---|
| Compute (8 vCPU, 32GB) | m6i.2xlarge | $280 | n2-standard-8 | $261 | -7% |
| Managed Kubernetes | EKS cluster | $73 | GKE Autopilot | $74 | +1% |
| Managed Postgres (4 vCPU) | RDS db.r6g.xl | $438 | Cloud SQL | $385 | -12% |
| NAT Gateway (100GB) | NAT Gateway | $49 | Cloud NAT | $46 | -6% |
| Load Balancer | ALB | $22 + LCU | Cloud LB | $18 + rules | varies |
| Object Storage (1TB) | S3 Standard | $23 | Cloud Storage | $20 | -13% |
These numbers from our January 2026 billing show GCP running 5-12% cheaper for equivalent compute and storage. However, AWS data transfer costs are lower for inter-region traffic, which can flip the equation for globally distributed workloads.
Shared Policy Enforcement
The consistency win comes from shared OPA policies that apply to both clouds:
# policies/security-baseline.rego
package terraform.security
deny[msg] {
resource := input.resource_changes[_]
resource.type == "aws_s3_bucket"
not has_encryption(resource)
msg := sprintf("S3 bucket %s must have encryption enabled", [resource.address])
}
deny[msg] {
resource := input.resource_changes[_]
resource.type == "google_storage_bucket"
not has_uniform_access(resource)
msg := sprintf("GCS bucket %s must have uniform bucket-level access", [resource.address])
}
deny[msg] {
resource := input.resource_changes[_]
is_database(resource.type)
not has_deletion_protection(resource)
msg := sprintf("Database %s must have deletion protection enabled", [resource.address])
}
is_database(type) {
database_types := {"aws_rds_instance", "aws_rds_cluster", "google_sql_database_instance"}
database_types[type]
}
Drift Detection Across Clouds
Drift detection in a multi-cloud setup requires a unified view. We run hourly drift detection and aggregate results:
#!/bin/bash
# scripts/drift-detect.sh
CLOUDS=("aws-prod" "aws-staging" "gcp-prod" "gcp-staging")
DRIFT_FOUND=0
for env in "${CLOUDS[@]}"; do
echo "Checking drift: ${env}"
cd "environments/${env}"
terraform init -backend=true -input=false > /dev/null 2>&1
PLAN_OUTPUT=$(terraform plan -detailed-exitcode -no-color 2>&1)
EXIT_CODE=$?
if [ $EXIT_CODE -eq 2 ]; then
echo "DRIFT DETECTED in ${env}"
DRIFT_FOUND=1
# Send to monitoring
curl -X POST "$SLACK_WEBHOOK" \
-H 'Content-Type: application/json' \
-d "{\"text\": \"Terraform drift detected in ${env}\"}"
fi
cd ../..
done
exit $DRIFT_FOUND
Lessons from 14 Months in Production
Module versioning is critical. We pin module versions per environment and promote through staging before production. A breaking module change in the networking layer that works for AWS but fails for GCP will ruin your Friday.
Cross-cloud data dependencies require careful handling. When GCP workloads need to reference AWS resources (like an S3 bucket ARN for cross-cloud replication), use Terraform data sources with explicit cross-backend state references or external data documents. Never hardcode.
Plan times grow linearly. With 200+ resources per environment, plan times reached 4-5 minutes. We split into smaller state files by domain (networking, compute, data) while maintaining the shared module structure.
The team skill gap is real. Engineers comfortable with AWS IAM struggle with GCP IAM bindings and vice versa. Invest in cross-training and make the modules self-documenting with extensive variable descriptions.
Key Takeaways
- Isolate state per cloud and environment. Never share a state file across cloud boundaries. The blast radius is unacceptable.
- Define cloud-agnostic output interfaces. Consistent outputs from cloud-specific modules enable reusable patterns upstream.
- Enforce policies centrally. OPA/Sentinel policies that span both clouds are the primary consistency mechanism.
- Automate drift detection aggressively. Multi-cloud drift is twice the surface area; hourly checks are not paranoia, they are necessity.
- Budget for the skills gap. Multi-cloud Terraform fluency takes 3-6 months per engineer. Plan for slower velocity during the transition.
The single-codebase approach works best when you have 3+ engineers working on infrastructure and need audit-grade consistency. Below that team size, separate repos with shared documentation might be more pragmatic. Know your constraints before committing to the pattern.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.