Managing AWS and GCP from a Single Terraform Codebase with Workspace Isolation

How to structure a multi-cloud Terraform repository using workspaces, provider aliasing, and state isolation for AWS and GCP without losing your sanity.

#terraform#multi-cloud#workspaces#aws#gcp
Cover image for the article: Managing AWS and GCP from a Single Terraform Codebase with Workspace Isolation

Most multi-cloud Terraform setups I inherit look the same: two completely separate repositories, divergent module structures, different state backends, and no shared patterns. The teams running them eventually realize they are maintaining two infrastructure platforms with none of the promised portability benefits.

After migrating a fintech platform from single-cloud AWS to a multi-cloud AWS+GCP architecture, I developed a workspace isolation strategy that keeps a single codebase manageable without sacrificing the separation guarantees that production infrastructure demands. Here is the approach, the tradeoffs, and the cost data that justified the effort.

Why Single Codebase Multi-Cloud

The argument for separate repositories is simplicity: each repo owns its cloud, its state, its CI/CD pipeline. The argument against is drift: when you implement a security pattern in AWS and forget to port it to GCP, you have an inconsistent security posture that auditors will find.

ApproachCode ReuseConsistencyBlast RadiusTeam Overhead
Separate reposNoneLowContained2x review load
Monorepo, shared stateHighHighDangerousLow
Monorepo, isolated workspacesMedium-HighHighContainedMedium
Terragrunt wrapperMediumMediumContainedHigh (tooling)

The workspace isolation approach gives you consistency guarantees (shared modules, shared policies) with blast radius containment (isolated state per cloud, per environment).

Repository Structure

infrastructure/
├── modules/
│   ├── networking/           # Cloud-agnostic interface
│   │   ├── aws/             # AWS VPC implementation
│   │   │   ├── main.tf
│   │   │   ├── variables.tf
│   │   │   └── outputs.tf
│   │   └── gcp/            # GCP VPC implementation
│   │       ├── main.tf
│   │       ├── variables.tf
│   │       └── outputs.tf
│   ├── compute/
│   │   ├── aws/            # ECS/EKS
│   │   └── gcp/           # Cloud Run/GKE
│   ├── database/
│   │   ├── aws/           # RDS/Aurora
│   │   └── gcp/          # Cloud SQL/AlloyDB
│   └── observability/
│       ├── aws/           # CloudWatch
│       └── gcp/          # Cloud Monitoring
├── environments/
│   ├── aws-prod/
│   │   ├── main.tf
│   │   ├── backend.tf
│   │   └── terraform.tfvars
│   ├── aws-staging/
│   ├── gcp-prod/
│   └── gcp-staging/
├── policies/               # OPA/Sentinel shared policies
│   ├── cost-limits.rego
│   ├── security-baseline.rego
│   └── naming-conventions.rego
└── scripts/
    ├── plan-all.sh
    └── drift-detect.sh

Provider Configuration with Workspace Awareness

The key pattern is using workspace-aware provider configuration with strict backend isolation:

# environments/aws-prod/backend.tf
terraform {
  backend "s3" {
    bucket         = "terraform-state-prod"
    key            = "aws-prod/terraform.tfstate"
    region         = "us-east-1"
    dynamodb_table = "terraform-locks"
    encrypt        = true
  }

  required_providers {
    aws = {
      source  = "hashicorp/aws"
      version = "~> 5.40"
    }
  }
}

provider "aws" {
  region = var.aws_region

  default_tags {
    tags = {
      Environment  = "production"
      ManagedBy    = "terraform"
      Cloud        = "aws"
      CostCenter   = var.cost_center
      Workspace    = terraform.workspace
    }
  }

  assume_role {
    role_arn = "arn:aws:iam::${var.aws_account_id}:role/TerraformExecution"
  }
}
# environments/gcp-prod/backend.tf
terraform {
  backend "gcs" {
    bucket = "terraform-state-prod-gcp"
    prefix = "gcp-prod"
  }

  required_providers {
    google = {
      source  = "hashicorp/google"
      version = "~> 5.20"
    }
  }
}

provider "google" {
  project = var.gcp_project_id
  region  = var.gcp_region

  default_labels = {
    environment = "production"
    managed-by  = "terraform"
    cloud       = "gcp"
    cost-center = var.cost_center
  }
}

Cross-Cloud Module Interface Pattern

The power of this approach comes from defining cloud-agnostic interfaces that both implementations satisfy:

# modules/networking/aws/outputs.tf
output "network_id" {
  description = "VPC ID for the primary network"
  value       = aws_vpc.main.id
}

output "private_subnet_ids" {
  description = "Private subnet IDs across availability zones"
  value       = aws_subnet.private[*].id
}

output "private_cidr_blocks" {
  description = "CIDR blocks for private subnets"
  value       = aws_subnet.private[*].cidr_block
}

output "dns_zone_id" {
  description = "Private DNS zone ID"
  value       = aws_route53_zone.private.zone_id
}
# modules/networking/gcp/outputs.tf
output "network_id" {
  description = "VPC network self_link"
  value       = google_compute_network.main.id
}

output "private_subnet_ids" {
  description = "Private subnet self_links"
  value       = [for s in google_compute_subnetwork.private : s.id]
}

output "private_cidr_blocks" {
  description = "CIDR blocks for private subnets"
  value       = [for s in google_compute_subnetwork.private : s.ip_cidr_range]
}

output "dns_zone_id" {
  description = "Private DNS zone ID"
  value       = google_dns_managed_zone.private.id
}

Both modules expose the same output interface, which means the consuming environment code can use consistent variable names regardless of cloud.

State Isolation Strategy

State isolation is non-negotiable. A Terraform plan against GCP production must never be able to affect AWS production resources, even accidentally.

Multi-cloud Terraform state isolation diagram

The isolation layers:

  1. Backend separation: S3 for AWS state, GCS for GCP state. Different credentials, different access policies.
  2. IAM boundaries: The AWS Terraform role cannot access GCP APIs and vice versa. No ambient credentials.
  3. Workspace naming convention: {cloud}-{environment}-{region} (e.g., aws-prod-us-east-1, gcp-prod-us-central1).
  4. CI/CD pipeline isolation: Separate pipeline stages per cloud with independent approval gates.
# .github/workflows/terraform.yml (simplified)
jobs:
  plan-aws-prod:
    runs-on: ubuntu-latest
    permissions:
      id-token: write
    steps:
      - uses: aws-actions/configure-aws-credentials@v4
        with:
          role-to-assume: arn:aws:iam::111111111111:role/TerraformPlan
          aws-region: us-east-1
      - run: |
          cd environments/aws-prod
          terraform init
          terraform plan -out=plan.tfplan
      
  plan-gcp-prod:
    runs-on: ubuntu-latest
    permissions:
      id-token: write
    steps:
      - uses: google-github-actions/auth@v2
        with:
          workload_identity_provider: projects/222222/locations/global/workloadIdentityPools/github/providers/github
          service_account: terraform-plan@project.iam.gserviceaccount.com
      - run: |
          cd environments/gcp-prod
          terraform init
          terraform plan -out=plan.tfplan

Cost Comparison: Dual-Cloud Resource Equivalents

When running equivalent workloads across both clouds, cost normalization becomes critical for capacity planning:

Resource CategoryAWS ServiceMonthly CostGCP EquivalentMonthly CostDelta
Compute (8 vCPU, 32GB)m6i.2xlarge$280n2-standard-8$261-7%
Managed KubernetesEKS cluster$73GKE Autopilot$74+1%
Managed Postgres (4 vCPU)RDS db.r6g.xl$438Cloud SQL$385-12%
NAT Gateway (100GB)NAT Gateway$49Cloud NAT$46-6%
Load BalancerALB$22 + LCUCloud LB$18 + rulesvaries
Object Storage (1TB)S3 Standard$23Cloud Storage$20-13%

These numbers from our January 2026 billing show GCP running 5-12% cheaper for equivalent compute and storage. However, AWS data transfer costs are lower for inter-region traffic, which can flip the equation for globally distributed workloads.

Shared Policy Enforcement

The consistency win comes from shared OPA policies that apply to both clouds:

# policies/security-baseline.rego
package terraform.security

deny[msg] {
  resource := input.resource_changes[_]
  resource.type == "aws_s3_bucket"
  not has_encryption(resource)
  msg := sprintf("S3 bucket %s must have encryption enabled", [resource.address])
}

deny[msg] {
  resource := input.resource_changes[_]
  resource.type == "google_storage_bucket"
  not has_uniform_access(resource)
  msg := sprintf("GCS bucket %s must have uniform bucket-level access", [resource.address])
}

deny[msg] {
  resource := input.resource_changes[_]
  is_database(resource.type)
  not has_deletion_protection(resource)
  msg := sprintf("Database %s must have deletion protection enabled", [resource.address])
}

is_database(type) {
  database_types := {"aws_rds_instance", "aws_rds_cluster", "google_sql_database_instance"}
  database_types[type]
}

Drift Detection Across Clouds

Drift detection in a multi-cloud setup requires a unified view. We run hourly drift detection and aggregate results:

#!/bin/bash
# scripts/drift-detect.sh

CLOUDS=("aws-prod" "aws-staging" "gcp-prod" "gcp-staging")
DRIFT_FOUND=0

for env in "${CLOUDS[@]}"; do
  echo "Checking drift: ${env}"
  cd "environments/${env}"
  
  terraform init -backend=true -input=false > /dev/null 2>&1
  PLAN_OUTPUT=$(terraform plan -detailed-exitcode -no-color 2>&1)
  EXIT_CODE=$?
  
  if [ $EXIT_CODE -eq 2 ]; then
    echo "DRIFT DETECTED in ${env}"
    DRIFT_FOUND=1
    # Send to monitoring
    curl -X POST "$SLACK_WEBHOOK" \
      -H 'Content-Type: application/json' \
      -d "{\"text\": \"Terraform drift detected in ${env}\"}"
  fi
  
  cd ../..
done

exit $DRIFT_FOUND

Lessons from 14 Months in Production

Module versioning is critical. We pin module versions per environment and promote through staging before production. A breaking module change in the networking layer that works for AWS but fails for GCP will ruin your Friday.

Cross-cloud data dependencies require careful handling. When GCP workloads need to reference AWS resources (like an S3 bucket ARN for cross-cloud replication), use Terraform data sources with explicit cross-backend state references or external data documents. Never hardcode.

Plan times grow linearly. With 200+ resources per environment, plan times reached 4-5 minutes. We split into smaller state files by domain (networking, compute, data) while maintaining the shared module structure.

The team skill gap is real. Engineers comfortable with AWS IAM struggle with GCP IAM bindings and vice versa. Invest in cross-training and make the modules self-documenting with extensive variable descriptions.

Key Takeaways

  1. Isolate state per cloud and environment. Never share a state file across cloud boundaries. The blast radius is unacceptable.
  2. Define cloud-agnostic output interfaces. Consistent outputs from cloud-specific modules enable reusable patterns upstream.
  3. Enforce policies centrally. OPA/Sentinel policies that span both clouds are the primary consistency mechanism.
  4. Automate drift detection aggressively. Multi-cloud drift is twice the surface area; hourly checks are not paranoia, they are necessity.
  5. Budget for the skills gap. Multi-cloud Terraform fluency takes 3-6 months per engineer. Plan for slower velocity during the transition.

The single-codebase approach works best when you have 3+ engineers working on infrastructure and need audit-grade consistency. Below that team size, separate repos with shared documentation might be more pragmatic. Know your constraints before committing to the pattern.

Comments

    No comments yet. Be the first to share your thoughts.