s☁
silvio.cloud
Back to all articles
• 4 min read • Silvio Silva

Terraform in Production (Part 5): Rebuilding Code Without Rebuilding Infrastructure

The ultimate test of IaC maturity: refactoring massive codebases without triggering a single resource recreation. State blast-radius reduction and zero-churn refactoring.

Terraform in Production (Part 5): Rebuilding Code Without Rebuilding Infrastructure

Terraform in Production (Part 5): Rebuilding Code Without Rebuilding Infrastructure

One of the defining hallmarks of a senior infrastructure engineer is the ability to drastically restructure, refactor, and modernize an existing codebase with zero infrastructure churn.

How do you know you have succeeded?

[!IMPORTANT] The Golden Rule of IaC Refactoring:
A pure code refactoring pull request must ALWAYS produce:
Plan: 0 to add, 0 to change, 0 to destroy.

If a plan shows unexpected resource recreation (~ forces replacement), you have not refactored—you are about to cause an outage.

Here is the exact playbook to refactor production Terraform safely.


1. Surgical Use of lifecycle

Cloud resources are often modified dynamically outside of Terraform. For example:

  • An Auto Scaling Group adjusting its desired_capacity based on real-time CPU alarms.
  • An external deployment controller updating image tags on an ECS task definition.
  • A managed service appending runtime metadata tags.

If you don’t account for this, Terraform plans will constantly try to revert external changes. Use lifecycle { ignore_changes = [...] } surgically:

resource "aws_autoscaling_group" "workers" {
  name             = "production-worker-asg"
  min_size         = 2
  max_size         = 10
  desired_capacity = 2

  lifecycle {
    # Allow external autoscaling metrics to adjust capacity without plan noise
    ignore_changes = [desired_capacity]
  }
}

2. Taming the Monolith State File

Never put your entire company's infrastructure into a single root state file. If your state file tracks 500+ resources:

  • terraform plan takes 5 minutes just to refresh.
  • A lock on the state file blocks other teams.
  • A single syntax error halts all deployments across the organization.

Divide your state by rate of change and operational domain:

infrastructure/
├── 01-networking/     # VPC, Transit Gateways, Subnets (changes quarterly)
├── 02-security/       # IAM Roles, KMS Keys, Vault (changes monthly)
├── 03-compute/        # EKS Clusters, ALB, ECS (changes weekly)
└── 04-apps/           # Service deployments, DNS records (changes daily)

Decoupling Layers via Data Sources

Instead of tightly coupling stacks with brittle terraform_remote_state outputs, wire cross-layer dependencies using data source lookups governed by consistent tagging conventions:

# In compute stack: look up the VPC dynamically by tag
data "aws_vpc" "production" {
  filter {
    name   = "tag:Environment"
    values = ["production"]
  }
}

This decouples your deployment lifecycles completely. The networking team can upgrade route tables without holding a lock on the application stack.


3. The 4-Step Refactoring Checklist

Before merging any structural refactor PR:

  1. Verify State Alignment: Run terraform plan in your isolated CI runner. Every renamed block must have a corresponding moved block.
  2. Reconcile Provider Defaults: If the plan shows trivial changes (~ update in-place) for default attributes (e.g. ipv6_cidr_block = null), align your HCL to match provider defaults until the diff is completely clean.
  3. Validate the Golden Rule: Confirm that the final plan output is:
    Plan: 0 to add, 0 to change, 0 to destroy.
    
  4. Merge and Apply: Merge the PR. The apply step should complete in seconds with zero infrastructure downtime.

Complete Series Index

Congratulations on completing the Terraform in Production series! Here is the full index for reference:

  1. Part 1: Why HCL Beats Cloud-Native IaC — Overcoming the single-cloud fallacy and multi-provider graphs.
  2. Part 2: Air-Gapped CI/CD with GitHub Actions & Private Runners — Hardening pipelines with OIDC and zero inbound ports.
  3. Part 3: The Module Dilemma — Encapsulation vs. Over-Engineering — Escaping the leaky abstraction trap and God modules.
  4. Part 4: Declarative State Surgery with moved and import Blocks — Git-audited state migrations and automated HCL generation.
  5. Part 5: Rebuilding Code Without Rebuilding Infrastructure — Achieving 0 to destroy and state blast-radius hygiene.

Keep your modules lean, your runners private, and your plans at 0 to destroy!

SS
Silvio Silva

Cloud & Systems Engineer · silvio.cloud