Migrating from Terraform to OpenTofu: Hard-Won Lessons from 2,400 State Files
A practical guide to migrating enterprise Terraform infrastructure to OpenTofu, covering state file compatibility, provider registry changes, and CI/CD pipeline updates.

When HashiCorp switched Terraform to the BSL license, we had 2,400 state files across 14 AWS accounts, 340 modules in our private registry, and 28 CI/CD pipelines depending on terraform CLI commands. The migration to OpenTofu took us four months — not because the tooling was hard, but because enterprise IaC has tentacles everywhere. Here is our complete migration playbook.
Why We Migrated
Let me be direct: this was not an ideological decision. Our legal team flagged BSL compliance risks for our consulting arm, where we deploy infrastructure for clients using our tooling. Under BSL, offering Terraform-based managed services that compete with HashiCorp products creates legal ambiguity. OpenTofu's MPL-2.0 license eliminates that risk entirely.
The secondary driver was feature velocity. OpenTofu 1.8 shipped client-side state encryption — a feature our security team had been requesting for two years. Provider-defined functions and the improved for_each semantics also accelerated our decision.
Pre-Migration Audit
Before touching any configuration, we needed a complete inventory. We built a scanner that crawled our GitLab instance and S3 state backends:
import boto3
import json
from pathlib import Path
from dataclasses import dataclass
@dataclass
class StateFileAudit:
account_id: str
bucket: str
key: str
terraform_version: str
provider_count: int
resource_count: int
serial: int
def scan_state_backends(accounts: list[str]) -> list[StateFileAudit]:
"""Scan all S3 backends for Terraform state files."""
results = []
for account_id in accounts:
session = boto3.Session(profile_name=f"org-{account_id}")
s3 = session.client('s3')
# Convention: all state buckets follow this pattern
bucket = f"terraform-state-{account_id}"
paginator = s3.get_paginator('list_objects_v2')
for page in paginator.paginate(Bucket=bucket, Suffix='.tfstate'):
for obj in page.get('Contents', []):
state_data = s3.get_object(Bucket=bucket, Key=obj['Key'])
state = json.loads(state_data['Body'].read())
results.append(StateFileAudit(
account_id=account_id,
bucket=bucket,
key=obj['Key'],
terraform_version=state.get('terraform_version', 'unknown'),
provider_count=len(state.get('resources', [])),
resource_count=sum(
len(r.get('instances', []))
for r in state.get('resources', [])
),
serial=state.get('serial', 0),
))
return results
The audit revealed:
- 2,400 state files managing 89,000 resources
- 12 distinct Terraform versions in use (0.14 through 1.6)
- 47 unique providers, 8 of which were community-maintained
- 340 private modules in our GitLab-hosted registry
Phase 1: Binary Swap (Week 1-2)
OpenTofu 1.6+ is state-file compatible with Terraform 1.6. The state format is identical — same JSON schema, same serial numbering. The simplest migration path is literally replacing the binary:
#!/bin/bash
# migration-swap.sh - Binary replacement with validation
set -euo pipefail
TOFU_VERSION="1.8.2"
WORKDIR="${1:-.}"
# Download and verify OpenTofu binary
curl -sSL "https://github.com/opentofu/opentofu/releases/download/v${TOFU_VERSION}/tofu_${TOFU_VERSION}_linux_amd64.zip" -o /tmp/tofu.zip
echo "sha256sum_here /tmp/tofu.zip" | sha256sum --check
unzip -o /tmp/tofu.zip -d /usr/local/bin/
# Validate state compatibility
cd "$WORKDIR"
tofu init -upgrade
tofu validate
# Run plan to confirm no drift
PLAN_OUTPUT=$(tofu plan -detailed-exitcode 2>&1) || {
EXIT_CODE=$?
if [ $EXIT_CODE -eq 2 ]; then
echo "WARNING: Plan shows changes. Review before proceeding."
echo "$PLAN_OUTPUT"
exit 1
fi
echo "ERROR: Plan failed"
echo "$PLAN_OUTPUT"
exit 1
}
echo "Migration validated: no changes detected"
The critical point: run tofu plan immediately after switching. If the plan shows zero changes, the state is compatible. If it shows drift, you have a pre-existing issue to resolve before migrating.
Phase 2: Provider Registry Migration (Week 3-4)
This is where things get interesting. Terraform uses registry.terraform.io by default. OpenTofu uses registry.opentofu.org. Most providers are mirrored, but the resolution order differs:
# Before: implicit HashiCorp registry
terraform {
required_providers {
aws = {
source = "hashicorp/aws"
version = "~> 5.40"
}
}
}
# After: explicit source still works — OpenTofu resolves hashicorp/*
# through its own registry mirror. No change needed for major providers.
terraform {
required_providers {
aws = {
source = "hashicorp/aws"
version = "~> 5.40"
}
}
}
For most teams, provider blocks require zero changes. OpenTofu's registry mirrors all providers that were published before the BSL switch. However, we hit issues with three proprietary providers that only publish to the HashiCorp registry. Solution: network mirror with a local filesystem mirror.
Phase 3: State Encryption (Week 5-8)
This was the feature that justified the migration timeline to leadership. OpenTofu 1.8 supports client-side state encryption with multiple key providers:
terraform {
encryption {
method "aes_gcm" "primary" {
keys = key_provider.aws_kms.state_key
}
key_provider "aws_kms" "state_key" {
kms_key_id = "arn:aws:kms:us-east-1:123456789:key/mrk-abc123"
region = "us-east-1"
key_spec = "AES_256"
}
state {
method = method.aes_gcm.primary
enforced = true
}
plan {
method = method.aes_gcm.primary
enforced = true
}
}
}
With enforced = true, OpenTofu refuses to read or write unencrypted state. We rolled this out per-account, starting with our production accounts that manage PII-adjacent resources.
Phase 4: CI/CD Pipeline Updates (Week 9-12)
Our GitLab CI pipelines needed surgical updates. The terraform command becomes tofu, but wrapper scripts, caching layers, and artifact paths all needed attention:
# .gitlab-ci.yml - OpenTofu pipeline
stages:
- validate
- plan
- apply
variables:
TOFU_VERSION: "1.8.2"
TF_PLUGIN_CACHE_DIR: "${CI_PROJECT_DIR}/.terraform.d/plugin-cache"
.tofu-base:
image: ghcr.io/opentofu/opentofu:${TOFU_VERSION}
before_script:
- tofu --version
- export TF_IN_AUTOMATION=true
cache:
key: "${CI_COMMIT_REF_SLUG}-providers"
paths:
- .terraform.d/plugin-cache/
plan:
extends: .tofu-base
stage: plan
script:
- tofu init -backend-config="key=${CI_PROJECT_PATH}/${CI_ENVIRONMENT_NAME}.tfstate"
- tofu plan -out=plan.cache -input=false
artifacts:
paths:
- plan.cache
expire_in: 1 hour
apply:
extends: .tofu-base
stage: apply
script:
- tofu init -backend-config="key=${CI_PROJECT_PATH}/${CI_ENVIRONMENT_NAME}.tfstate"
- tofu apply -input=false plan.cache
when: manual
only:
- main
Phase 5: Module Registry (Week 13-16)
Our private module registry was the hardest piece. We were using Terraform Cloud's private registry. Options:
- GitLab's built-in Terraform Module Registry — supports OpenTofu natively
- Artifactory — generic module hosting with version constraints
- Git sources directly — simplest but loses version discovery
We chose GitLab's registry since our modules were already in GitLab repos:
module "vpc" {
source = "gitlab.internal.company.com/infrastructure/modules/aws-vpc"
version = "~> 3.2"
cidr_block = var.vpc_cidr
availability_zones = var.azs
enable_nat_gateway = true
}
Benchmarks: OpenTofu vs Terraform Performance
We measured plan/apply times across our largest configurations (1,200+ resources):
| Operation | Terraform 1.6 | OpenTofu 1.8 | Delta |
|---|---|---|---|
| Init (cold) | 34s | 31s | -9% |
| Init (cached) | 2.1s | 1.8s | -14% |
| Plan (1200 resources) | 48s | 42s | -12% |
| Apply (50 changes) | 3m 12s | 2m 58s | -7% |
| State pull (encrypted) | N/A | 1.2s | — |
OpenTofu is marginally faster due to parallelism improvements in the graph walker. Not a migration driver, but a nice bonus.
What Broke
Sentinel policies don't exist in OpenTofu. We had 84 Sentinel policies. We migrated to OPA (Open Policy Agent) with conftest, which actually gave us more flexibility:
# policy/cost_limits.rego
package terraform.cost
deny[msg] {
resource := input.resource_changes[_]
resource.type == "aws_instance"
resource.change.after.instance_type == "x1e.32xlarge"
msg := sprintf("Instance type %s requires VP approval", [resource.change.after.instance_type])
}
Terraform Cloud remote state data sources. Any terraform_remote_state referencing a Terraform Cloud workspace needed rerouting to S3 backends. This was our biggest refactoring effort.
Third-party CI integrations. Tools like Spacelift and env0 had varying OpenTofu support timelines. We self-hosted our CI to avoid vendor lock-in.
Conclusion
Migrating 2,400 state files sounds terrifying, but the state format compatibility made it mechanical. The real work was in the ecosystem — registries, CI/CD, policy engines, and team training. My advice: start with non-production accounts, validate with plan -detailed-exitcode, and treat the migration as a four-phase project, not a flag day.
The investment paid for itself in three ways: eliminated BSL licensing risk, gained state encryption (a security audit requirement), and unified our IaC toolchain under a truly open-source license. Four months of migration effort for years of reduced risk.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.