Changing Live Infrastructure Safely
Refactoring and operating Terraform against real resources: moved and import blocks, lifecycle rules, drift, environments, and reviewing plans in CI before anything is applied.
SPACED REPETITION Β· 15 practice questions
Make this lesson stick.
Try 3 questions now. No account needed. Sample answers aren't saved.
or sign in to practice all 15Reading a Plan: In-Place Updates, Replacement, and What Each Symbol Costs You
You have just changed one line of HCL β a log level, a timeout, a tag β and you are about to run terraform apply. Between that keystroke and either a two-second config update or a paged incident sits a screen of text you either read or skimmed.
terraform plan computes a diff β a list of proposed actions against real resources β from three inputs: your configuration (the .tf files you wrote), the state file (Terraform's record of what it created last time, including each resource's real AWS ID), and a refresh of the live resources (Terraform calls AWS and reads each resource's current attributes). Every plan line answers one question: what will Terraform do to this account right now?
config/*.tf (what you wrote)
+
terraform.tfstate (what Terraform recorded at the last apply)
+
live AWS resources (what refresh just read back via the API)
β
terraform plan β proposed actions, each with an action symbol
(That is the base picture; -target narrows it, as a later section covers.)
The action symbols
Every planned block opens with a symbol naming the action.
| Symbol | Action | Rough meaning |
|---|---|---|
+ |
create | resource does not exist yet |
- |
destroy | resource is being removed |
~ |
update in place | same resource, attributes changed |
-/+ |
replace | destroy, then create a new one |
<= |
read | data source lookup, no resource change |
In place means Terraform sends an API call (for example UpdateFunctionConfiguration) and the resource keeps its identity β its ARN, its ID, and its data. Replacement means Terraform destroys the existing resource and builds a fresh one; nothing travels with it. When a changed attribute forces the replacement, the -/+ block carries a # forces replacement annotation naming the exact attribute that triggered it (a tainted resource, one you asked to replace with -replace, or one replaced through replace_triggered_by shows -/+ without it), and the summary line states the cost: Plan: 1 to add, 0 to change, 1 to destroy.
<= marks a data source, a read-only lookup such as data "aws_ssm_parameter" that reads a value instead of owning a resource. Data sources Terraform can read during planning do not appear in the list of actions at all; one whose arguments are not yet known, or that depends on a resource with changes pending, shows as a <= block headed will be read during apply.
Why replacement happens: ForceNew attributes
The AWS provider ships a schema: a description of every resource type's arguments, including which ones AWS cannot mutate after creation. Those are marked ForceNew β "changing this requires building a new resource." The plan does not guess; it reads the schema, marks the whole block -/+, and annotates the attribute.
The instructive cases are attributes of the same service landing on opposite sides:
resource "aws_lambda_function" "orders" {
function_name = "orders" # ForceNew
role = aws_iam_role.lambda_exec.arn
handler = "index.handler"
runtime = "nodejs22.x"
filename = "orders.zip"
source_code_hash = filebase64sha256("orders.zip")
environment {
variables = {
LOG_LEVEL = "DEBUG" # mutable, in place
}
}
}
environment.variables is patched in place, because AWS's UpdateFunctionConfiguration accepts new environment variables. function_name is ForceNew: a Lambda's name is part of its identity and AWS has no rename operation. The same split recurs across the stack β aws_dynamodb_table.hash_key is ForceNew (you cannot repartition a live table), while an aws_iam_policy's policy document updates in place, because AWS versions policy documents rather than replacing them.
Worked trace: one line, then two
Suppose the only change in your config is LOG_LEVEL = "INFO" β "DEBUG". The plan shows a single ~ block and nothing else:
Terraform will perform the following actions:
# aws_lambda_function.orders will be updated in-place
~ resource "aws_lambda_function" "orders" {
id = "orders"
# (other attributes unchanged)
~ environment {
~ variables = {
~ "LOG_LEVEL" = "INFO" -> "DEBUG"
}
}
}
Plan: 0 to add, 1 to change, 0 to destroy.
Read the three numbers: zero adds, zero destroys, one change. source_code_hash is untouched, so no new zip is uploaded β this is an environment-variable patch, and the function keeps its ARN and its concurrency settings.
Now suppose the same commit also renames function_name = "orders" to "orders-v2". The same resource, which was ~, becomes -/+:
# aws_lambda_function.orders must be replaced
-/+ resource "aws_lambda_function" "orders" {
~ function_name = "orders" -> "orders-v2" # forces replacement
# (other attributes unchanged)
}
Plan: 1 to add, 0 to change, 1 to destroy.
The word must is the signal, and the arrow next to # forces replacement points at the culprit. The decision that follows is yours, not Terraform's. A renamed function gets a new ARN: every permission that references it is recreated, every event source mapping is re-pointed in place, and anything still invoking the old name fails once the old function is deleted. Whether a tidier name is worth that window is a judgment call the plan just handed you; the ripple to dependents is covered in "Destructive Changes and Cascades: Triage, Sequencing, and Minimizing Downtime".
What each symbol actually costs
~costs an API call. The resource keeps its data. Log levels, timeouts, capacity units, policy documents, and tags generally live here.-/+on a stateless resource (Lambda, ECS task definition) costs scheduling: a service gap, dropped in-flight work, recreated dependents.-/+on a stateful resource costs data. A replacedaws_dynamodb_tablestarts empty β deletion protection and point-in-time recovery help only if you enabled them before the apply, and PITR restores to an earlier point, not instantly. A replacedaws_sqs_queueloses every message still in the queue; if the replaced queue is a dead-letter queue, that is its whole backlog of failed messages.
Reading a plan costs seconds. Not reading it costs an incident.
Practice: classify three diffs
You run terraform plan and see three changed attributes. For each, decide in place or replacement, and identify which change discards data.
visibility_timeout_seconds = 30 -> 60onaws_sqs_queue.orders_dlqname = "orders" -> "orders-v2"onaws_sqs_queue.ordersread_capacity = 5 -> 25on a PROVISIONEDaws_dynamodb_table.orders
Check your answer
- In place (
~). SQS lets you change a visibility timeout on a live queue withSetQueueAttributes; the queue, its ARN, and its messages survive. No data loss. - Replacement (
-/+). A queue'snameis its identity, and there is no rename API. Terraform deletesordersand createsorders-v2. This is the one that discards data: every message sitting in the old queue disappears, so drain or pause producers first. Dependents that referenced the old queue are updated or recreated too (an event source mapping is recreated, because itsevent_source_arnforces replacement). - In place (
~). Provisioned capacity is a mutable setting on a PROVISIONED table (UpdateTable); items are untouched. Note the precondition: on aPAY_PER_REQUESTtable these arguments are not used at all, so there is nothing to change.
The shared rule: read the # forces replacement annotation, and when it is absent, ask what the resource is. Identity fields (names, keys, types) are usually ForceNew; tuning fields (timeouts, capacities, variables, policy documents) usually are not. Trust the plan over your intuition β the schema decides, and it occasionally disagrees with what looks mutable.
Next comes the case where a plan surprises you even though nobody edited the config β see "Drift: When the Real Account No Longer Matches State or Config".
Drift: When the Real Account No Longer Matches State or Config
Someone opens the DynamoDB console for the orders table, turns off the stream because a consumer was erroring, and moves on. Your config still says stream_enabled = true. Nothing in the repository changed and nothing in the plan's inputs changed β but the account did.
That gap is drift: a live resource whose real attributes have diverged from what state records or config declares. Drift isn't an error or a failed apply; it's the normal consequence of an account more than one path can change. Console clicks are the obvious source, but drift on an AWS backend also arrives from other automation β a cost-allocation tool that stamps tags on every table, an Application Auto Scaling policy that moves an ECS Fargate service's desired_count, a security scanner that edits a rule β or from an apply that died halfway.
Drift is invisible until Terraform reads the resource
State is a record, not a live connection. Terraform doesn't poll AWS; it learns current attributes by refreshing β calling the AWS API to read each resource in state, then comparing what came back against config. That read happens at the start of nearly every plan:
terraform plan
ββ 1. refresh: read current attributes from AWS for every resource in state
ββ 2. compare refreshed reality against config
ββ 3. different? β plan line (~ in-place, -/+ replace)
same? β silence
Here's the config, trimmed to what matters:
resource "aws_dynamodb_table" "orders" {
name = "orders"
billing_mode = "PAY_PER_REQUEST"
hash_key = "order_id"
attribute {
name = "order_id"
type = "S"
}
stream_enabled = true
stream_view_type = "NEW_AND_OLD_IMAGES"
tags = {
Environment = "prod"
Owner = "payments"
}
}
After the console click, a normal terraform plan opens with a section headed Objects have changed outside of Terraform and then proposes the fix:
# aws_dynamodb_table.orders will be updated in-place
~ resource "aws_dynamodb_table" "orders" {
id = "orders"
name = "orders"
~ stream_arn = "arn:aws:dynamodb:us-east-1:123456789012:table/orders/stream/2026-01-10T08:00:00.000" -> (known after apply)
~ stream_enabled = false -> true
# (12 unchanged attributes hidden)
}
Plan: 0 to add, 1 to change, 0 to destroy.
Read the arrow: left is the refreshed real value, right is what config demands. Drift never announces itself β it appears as an unexpected ~ line in a plan you expected to be empty. That ~ is in-place, not replacement; for how to tell those apart, see the earlier section Reading a Plan. Note the stream_arn line: re-enabling a stream creates a new stream with a new ARN, so an event source mapping that reads the old stream would be replaced in the same plan.
Catching drift on a schedule
The cheapest drift detector is a plan whose exit code you actually read. terraform plan -detailed-exitcode returns a distinct code when there is anything to do:
terraform plan -detailed-exitcode -lock-timeout=5m
# 0 β nothing to change
# 1 β the plan itself failed (bad config, expired SSO session, denied API call)
# 2 β changes or drift exist
A nightly CI job that assumes a role via OIDC needs only terraform plan and read access to your resources β refreshing reads them but changes nothing. The exception is the state backend: plan takes the state lock by default, so with the S3 backend's use_lockfile the role also needs s3:GetObject, s3:PutObject and s3:DeleteObject on the .tflock object, plus s3:ListBucket on the bucket and s3:GetObject on the state file. Treat 1 and 2 differently: 2 means "someone changed the account, go look," 1 means "the check itself is broken." A job that lumps them together reports green over a red reality.
Looking without touching anything
Three read-only moves answer "what does Terraform currently believe?"
terraform state list
# aws_dynamodb_table.orders
# aws_lambda_function.orders
terraform state show aws_dynamodb_table.orders
# stream_enabled = true β what STATE records, not what AWS says
β οΈ state show reads the state file; it does not refresh. It will happily print true for a stream that is actually off β and only a refresh reveals that discrepancy.
For a drift-only view, use refresh-only mode, added in Terraform 0.15.4:
terraform plan -refresh-only # show drift, propose no config changes
terraform apply -refresh-only # write reality into state, change nothing in AWS
plan -refresh-only reports what moved and stops. apply -refresh-only records those real values in state so state and reality agree β it never calls a write API on your resources; the only thing it writes is the state file. (The older standalone terraform refresh command is deprecated in favor of this.) That matters: of the two applies here, refresh-only is the one that cannot take your infrastructure down.
Two ways to make them agree
Once drift is visible you have exactly two endings, and choosing between them is the actual skill. Reconcile reality means accepting the out-of-band change: apply refresh-only so state records it, then edit config so Terraform stops fighting it. Reconcile the account means making config win: a normal apply that writes your declared values back over whatever is out there.
| Drift source | Direction | First command |
|---|---|---|
| Console change was a mistake | Config wins | terraform apply |
| Change was intended | Reality wins, then codify | apply -refresh-only |
| Automation will repeat it | Codify | apply -refresh-only |
The cost tool's tag is the clearest reality-wins case. Suppose it added CostCenter = "cc-4471" to every table, and your config lists only Environment and Owner:
~ resource "aws_dynamodb_table" "orders" {
~ tags = {
- "CostCenter" = "cc-4471" -> null
# (2 unchanged elements hidden)
}
}
A normal apply deletes that tag β and the tool adds it back next run, buying you permanent noise and a fight you lose weekly. Better: apply -refresh-only to accept it, then add CostCenter to the provider's default_tags so every resource declares it. Now it's both real and in config, and the plan goes quiet.
β οΈ The opposite decision is dangerous. A normal apply that "corrects" drift writes your config's version of reality β if an on-call engineer raised a Lambda's memory during an incident, applying your 512 MB config re-breaks production. When a plan's ~ lines touch a resource someone was just debugging, stop and ask before applying.
The mistake that costs you: tolerating the noise
The failure mode isn't a wrong command, it's habituation. When every plan shows four ~ lines from known sources, readers learn to skim the block β and a real change hides inside it. Two concrete costs. First, -target (covered in Applying with Control) refreshes only the part of the graph you targeted, so drift in untargeted resources stays unexamined and the plan you reviewed was never the whole picture. Second, apply -refresh-only silently promotes whatever is in the account to the source of truth β a hand-resized SQS queue, a hand-widened IAM policy, a raised maxReceiveCount in a queue's redrive policy β so the next reviewer approves it as if it had been designed. A scheduled drift check with a named owner beats a permanent shrug.
Guided practice: three symptoms, three decisions
Your orders stack declares aws_lambda_function.orders with memory_size = 512, the DynamoDB table above with an explicit tags block, and an aws_security_group fronting a Fargate service. Three things happened this week. For each, decide refresh-only apply or normal apply, then say what you'd change in config afterward.
- During a load test someone set the Lambda's memory to 1024 in the console. It's still 1024; nobody decided 1024 is the new normal.
- A cost-allocation tool added
CostCenter = "cc-4471"to every DynamoDB table in the account. - A developer added an ingress rule
0.0.0.0/0 β 5432to the Fargate service's security group to debug a connection. It's still there three days later.
Check your answer
1. Normal apply β config wins. The plan shows ~ memory_size = 1024 -> 512 and the apply resets it. Nothing changes in config unless the load test proved 512 is too small; then edit config to 1024 and run the same normal apply, because the apply refreshes first and finds config and reality already agreeing. You don't need refresh-only to adopt a value.
2. Refresh-only apply first, then codify. terraform apply -refresh-only records the tag in state; then add CostCenter to the provider's default_tags (or the table's tags). Deciding factor: the tool will re-add it, so a normal apply only produces noise you re-review every week.
3. Normal apply β after checking Terraform can even see the rule. If the rules are inline ingress blocks on aws_security_group, the plan shows the rule being removed and a normal apply revokes it. If you use standalone aws_vpc_security_group_ingress_rule resources β the pattern the AWS provider now recommends β that hand-added rule was never in state, so neither a normal nor a refresh-only apply will touch it; Terraform is silent. Remove it out of band, then decide whether the rule belongs in code with a narrow source (a narrow CIDR, a managed prefix list, or a referenced security group) instead of 0.0.0.0/0. Either way, follow up on why a developer needed console access in production.
The through-line: refresh tells you the truth, and you choose which side β state or the account β you correct. Never let an unfamiliar ~ line ride along in an apply you reviewed for something else.
Applying with Control: Plan Files, State Locking, Targeting, and Ordering
A plan you skimmed is a promise; an apply is the delivery. Everything that happens between those two moments β the file that pins the diff, the lock that keeps a second apply out, the flags that narrow scope β either keeps the promise or quietly breaks it. This section covers the mechanics that make terraform apply do exactly what you reviewed, and the failure mode each mechanic introduces.
Pin the diff with a plan file
Running terraform plan alone is a rehearsal. Add -out and you get a saved plan file: a serialized record of the diff Terraform computed, which apply can execute without recomputing anything.
terraform plan -out=tfplan
terraform apply tfplan
The plan ends with output like this:
Plan: 0 to add, 1 to change, 0 to destroy.
Saved the plan to: tfplan
Most plan runs are not saved, and that matters. A normal plan refreshes first β it reads the current attributes of every resource in state from AWS before diffing. If someone adds a tag, disables the DynamoDB stream on aws_dynamodb_table.orders, or changes an SQS visibility timeout while your PR sits in review, an unsaved terraform apply refreshes again and can build a different diff than the one you approved. Same config, different account, different actions. Applying the saved file executes the reviewed diff instead.
Two properties of plan files are worth committing to memory:
- A plan file is bound to the state serial it was created from. If the state has moved since β another apply landed β Terraform refuses with a stale plan error rather than applying a mismatched diff. That refusal is protection, not a bug.
- A plan file embeds the planned attribute values, including sensitive ones: SSM parameter values, Lambda environment variables, generated passwords. Treat
tfplanas a secret. Don't commit it; in CI, pass it between jobs as a short-lived artifact.
State locking: one writer at a time
State locking is the mechanism that stops two applies from writing the same state file at once. With the S3 backend, Terraform 1.10 and later can lock natively in S3 using use_lockfile = true; the older pattern, deprecated since Terraform 1.11, points at a DynamoDB table with the dynamodb_table argument.
terraform {
backend "s3" {
bucket = "acme-tfstate"
key = "orders/prod/terraform.tfstate"
region = "us-east-1"
use_lockfile = true # Terraform 1.10+ S3-native locking
}
}
The lock is taken at the start of an operation and released when it ends. A locked state is a safety mechanism, not an error β Terraform reports it with a Lock Info block naming the lock ID, the operation, who took it, and when:
Error: Error acquiring the state lock
Lock Info:
ID: a1b2c3d4-...
Path: acme-tfstate/orders/prod/terraform.tfstate
Operation: OperationTypeApply
Who: ci-runner@build-4821
Created: 2026-02-14 09:12:44 +0000
That block is a witness statement. When a run is interrupted mid-apply β CI job killed, laptop closed, Ctrl-C during a slow ECS deployment β the lock outlives the process. The reflex is terraform force-unlock, and β οΈ it is the wrong reflex. Read Who and Created first, confirm no apply is actually running (check the CI job named in Who), and remember that an interrupted apply may have already created half its resources. Only then clear it with the ID from the error:
terraform force-unlock -force a1b2c3d4-...
Clearing a lock while another apply is genuinely running gives you two writers, and the loser's resources end up orphaned or re-created β precisely the unplanned divergence that the section on drift ('Drift: When the Real Account No Longer Matches State or Config') teaches you to detect and repair.
Narrowing scope: -target and -parallelism
-target restricts an operation to one resource and its dependencies. It is an emergency tool for a broken resource blocking everything else, not a workflow:
# Emergency only: apply just this Lambda, then reconcile fully
terraform apply -target=aws_lambda_function.orders
What it skips is the point. Dependents are not updated, so aws_lambda_event_source_mapping.orders (the resource that tells Lambda to poll a queue or stream) may still reference the old function configuration, and changes you made to other resources are left out of this run; Terraform only warns that the plan "may not represent all of the changes requested by the current configuration". Config and state are now temporarily inconsistent. Always follow a targeted apply with a full terraform plan and terraform apply.
-parallelism sets how many resources Terraform walks the graph with concurrently; the default is 10. Lower it β -parallelism=1 β when an AWS API is rate-limiting you. Raising it does not make anything more correct; it widens the blast radius of a partial failure, because more operations are in flight when one fails.
Ordering comes from the graph, not from file order
Terraform builds a dependency graph (a map of which resources must exist before others) from the references in your configuration. Splitting resources across files, or writing them in a certain order, changes nothing.
For a queue-driven Lambda, the graph resolves in waves:
Wave 1 (parallel): aws_sqs_queue.orders
aws_dynamodb_table.orders (stream_enabled = true)
β
Wave 2: aws_lambda_function.orders
β
Wave 3: aws_lambda_event_source_mapping.orders
resource "aws_lambda_function" "orders" {
function_name = "orders"
role = aws_iam_role.orders.arn
filename = "orders.zip"
handler = "index.handler"
runtime = "nodejs22.x"
# ".id" reference creates a graph edge: the queue is created first
environment {
variables = {
QUEUE_URL = aws_sqs_queue.orders.id
STREAM_ARN = aws_dynamodb_table.orders.stream_arn
}
}
}
resource "aws_lambda_event_source_mapping" "orders" {
event_source_arn = aws_sqs_queue.orders.arn
function_name = aws_lambda_function.orders.function_name
}
Here is the trap: event_source_arn = aws_sqs_queue.orders.arn creates an edge, but event_source_arn = "arn:aws:sqs:us-east-1:123456789012:orders" β the same value written as a literal string β creates no edge at all. Terraform cannot see intent, only references. Those silent edges are where ordering bugs hide; when a literal string is unavoidable, depends_on = [aws_sqs_queue.orders] states the edge explicitly. Use that sparingly, as when a Lambda needs an IAM policy to be attached before it is created or invoked.
Habits with concrete rules
| Rule | Why it holds |
|---|---|
| One logical change per PR | Plan review is human; small diffs are reviewable |
-auto-approve only in CI |
The pipeline already printed the plan |
| Snapshot state before large applies | terraform state pull > backup.tfstate |
| Revert config, apply forward | Terraform has no rollback command |
That last row deserves emphasis: there is no terraform undo. Recovery from a bad apply is editing the configuration back and applying again, which means the next apply creates a new set of changes. A state snapshot β terraform state pull prints current state to stdout, so redirect it to a file β is a cheap way to keep a record of what Terraform's state held before (versioning on the state bucket also keeps earlier copies).
Practice
Task 1. Your CI apply was interrupted mid-run. You re-run it and get the state-lock error with a Lock Info block whose Who is ci-runner@build-4821. Write the check-and-recover sequence before anyone touches the state.
Check your answer
- Check whether job
build-4821is still running β the lock may be legitimate. - If it is dead, the apply stopped partway; resources may have been created but not recorded in state.
- Clear with
terraform force-unlock -force <ID>using the ID from the error β never delete the S3 lock object by hand. - Run
terraform planafter clearing and read the diff carefully for+creates of resources the interrupted run may already have built; applying those typically fails with an "already exists" error. - Follow with a full plan; expect the recovery apply to create what the interrupted run missed or to import what it orphaned.
The reasoning: force-unlock clears the lock, not the half-finished state of the account. The dangerous step is skipping the plan between unlocking and applying.
Task 2. A config declares four resources: the DynamoDB table with streaming enabled, aws_lambda_function.orders, aws_lambda_event_source_mapping.orders, and aws_sqs_queue.orders. The Lambda's environment variables reference both the queue's id and the table's stream_arn; the mapping references the queue's arn and the Lambda's function_name. Write the order the graph actually uses.
Check your answer
Wave 1 (parallel): aws_sqs_queue.orders + aws_dynamodb_table.orders
Wave 2: aws_lambda_function.orders
Wave 3: aws_lambda_event_source_mapping.orders
The queue and table have no reference between them, so Terraform creates them concurrently β the order is not linear. The Lambda sits in wave 2 because it references the queue's id and the table's stream_arn. The mapping sits in wave 3 because it references the Lambda. If you wrote it as a strict four-step line, you missed that wave 1 runs in parallel.
The through-line: every control here β the saved plan, the lock, the target flag, the graph β either binds an apply to what you reviewed or deliberately loosens that binding. Loosen it on purpose, and always with a full plan after.
Destructive Changes and Cascades: Triage, Sequencing, and Minimizing Downtime
A plan lands in review and one line jumps out:
# aws_lambda_function.orders must be replaced
-/+ resource "aws_lambda_function" "orders" {
~ function_name = "orders" -> "orders-v2" # forces replacement
...
}
Skim to the summary β "1 to add, 0 to change, 1 to destroy" β and you approve it. Then the resources that point at it change too, because replacement (the -/+ action, destroy-then-create) is contagious.
One Replacement Is Never One Replacement
Every AWS resource has an identity: the attributes AWS uses to name it, such as a Lambda's function_name, a queue's name, or a role's name. When an identity attribute changes, the AWS provider annotates it # forces replacement, meaning AWS cannot mutate the resource in place and Terraform must build a new one. Everything that referenced the old identity now points at something gone.
Trace the Lambda rename: aws_lambda_permission and aws_lambda_alias (a named pointer to a version, used for weighted traffic shifts) both take function_name as ForceNew β an attribute that forces a new resource instead of an in-place update β so each is destroyed and recreated. aws_lambda_event_source_mapping (the resource that lets Lambda poll a queue or stream) is the exception: its function_name updates in place, but its event_source_arn is ForceNew, so renaming the queue would replace it.
ECS arrives there by a different route. An aws_ecs_task_definition (the blueprint β image, CPU, environment β that ECS Fargate runs) is immutable in AWS, so every edit registers a new revision and changes the identity string. The aws_ecs_service referencing it rolls a new deployment.
β οΈ Read the whole plan, not the first block and the summary count. The cascading -/+ blocks are where the outage lives.
Triage: Stateful or Stateless?
For every -/+, ask one question first: does the destroyed resource hold data that config cannot rebuild?
| Resource | Holds data? | Replacement costs |
|---|---|---|
aws_dynamodb_table |
Yes β items | Irrecoverable loss |
aws_sqs_queue |
Yes β messages | Lost in-flight work |
aws_ssm_parameter |
Only if its value is set outside config | Lost value |
aws_lambda_function |
No β code lives in the artifact | A scheduling gap |
aws_ecs_task_definition |
No | A rolling deploy |
aws_iam_role |
No | Brief permission gap |
Stateful replacement is a data-loss decision; stateless replacement is a scheduling decision. That split sets your urgency. A stateless -/+ needs a quiet window and a rollback story. A stateful one needs a migration plan, a backup you have actually tested, and a written answer to "what happens to the data?" β because no meta-argument moves rows between DynamoDB tables.
(This is a triage heuristic, not a law β a Lambda mid-batch has transient cost. What matters is where the irreplaceable loss sits.)
Sequencing Beats Cleverness
You cannot fix a destructive change with a cleverer diff. You fix it with ordering.
1. Create the replacement (new name / new ARN)
β
2. Cut consumers over to the new identity
β
3. Verify traffic, drain the old resource
β
4. Destroy the old resource β a separate, later apply
Terraform's default for -/+ is destroy-then-create. The create_before_destroy = true lifecycle meta-argument flips that β but only when both resources can coexist, which means the names must differ. It is a pointed tool with a hard limit: a renamed DynamoDB table cannot preserve its items whatever lifecycle rules you attach. prevent_destroy = true is a seatbelt that makes Terraform reject any plan that would destroy the resource, buying time to think β but it blocks your staged plan too, so it is a guardrail, not a strategy. (Lifecycle meta-arguments get dedicated coverage elsewhere; reach for them with a reason.)
Step 4 being its own apply is the point. A single-apply -/+ does steps 4 and 1 in one pass, so anything still queued on the old resource dies there. That staging is why plan files matter β each stage is a separately reviewed artifact, as "Applying with Control: Plan Files, State Locking, Targeting, and Ordering" covers.
Cross-Stack Consumers: A Staged Migration
The hard variant: a DynamoDB name change forces replacement, and Lambda functions in other state files read that table. Terraform cannot create-before-destroy into the same name and cannot coordinate a cutover across states. You stage it.
- Verify protections. Confirm
deletion_protection_enabled = trueandpoint_in_time_recovery { enabled = true }on the live table. Deletion protection is an AWS-level flag that rejects delete calls; point-in-time recovery (PITR) keeps continuous backups that restore to any second in a rolling window. - Provision the new table alongside the old. A PITR restore always creates a new table, so one route is to create it from a restore of the old table (the provider's
restore_source_nameandrestore_to_latest_timearguments do this). - Migrate the data. A restore cannot fill a table that already exists, and its latest restorable point is typically about five minutes old, so catch up the writes made after it. The other route is an empty table backfilled by a copy job while the application dual-writes. DynamoDB Streams keep only 24 hours of changes, so they can catch a table up but cannot copy the items it already holds.
- Cut readers over β dual reads (prefer new) or a short maintenance window.
- Retire the old table once the new one serves traffic and reconciles counts.
What protects you and what does not: PITR restores to a point in time, into a new table, not instantly β it is not an undo button. A plan file prevents surprise diffs, not data loss. And the job is unfinished until a full apply of the final configuration runs; until then config and account disagree.
Practice: The Shared Queue Rename
A plan shows this and nothing else:
# aws_sqs_queue.orders must be replaced
-/+ resource "aws_sqs_queue" "orders" {
~ name = "orders" -> "orders-v2" # forces replacement
}
# aws_lambda_event_source_mapping.orders_esm must be replaced
-/+ resource "aws_lambda_event_source_mapping" "orders_esm" {
~ event_source_arn = "arn:aws:sqs:us-east-1:111122223333:orders" -> (known after apply) # forces replacement
}
The queue feeds a Lambda through orders_esm and has a dead-letter queue (DLQ) β a separate queue that receives messages the consumer failed to process after its retries. Write a minimal-downtime sequence, name every recreated resource, and say what breaks if you apply the whole plan at once.
Check your answer
Recreated: aws_sqs_queue.orders plus the cascade aws_lambda_event_source_mapping.orders_esm. The Lambda function and the DLQ are unchanged β the DLQ is a distinct queue that was not renamed.
Minimal-downtime sequence:
- Ask whether the rename is necessary. If not, revert the config β nothing here preserves queued messages, so the cheapest safe plan is no plan.
- If required, do not rename in place. Add
orders-v2as a secondaws_sqs_queueresource (with the same redrive policy to the DLQ) plus a second event source mapping, and apply: both queues exist and Lambda polls both. (create_before_destroy = trueon the existing resource would not help: the same apply createsorders-v2and then destroysorders, before anything has drained.) - Update producers to the new queue URL and let the old queue drain.
- Once the old queue is empty, remove the old queue and its mapping from config and run that apply on its own.
Applying all at once: Terraform destroys orders first by default, discarding every queued message and failing every producer still posting to the old URL. The ESM is torn down with it, so Lambda stops consuming during the gap, and the old queue name is gone permanently β that is a rename, not a migration.
The Triage Habit
Two questions collapse this whole topic: is the destroyed resource stateful? decides whether you are protecting data or scheduling a window, and what else references its identity? decides how far the -/+ reaches. Answer both before you approve anything.
Independent Transfer: Go/No-Go Review of a Real Plan Before Apply
A pull request titled tune orders pipeline arrives with a reassuring description: bump the Lambda's LOG_LEVEL and timeout, add a global secondary index to the orders table, tighten the IAM inline policy on the exec role, and refresh the redrive policy on the SQS dead-letter queue. Four small edits, none of them flagged as risky. Your job is to decide whether the apply runs β and the only evidence that counts is the plan text, not the description.
The plan you must review
Terraform will perform the following actions:
# aws_dynamodb_table.orders will be updated in-place
~ resource "aws_dynamodb_table" "orders" {
id = "orders"
name = "orders"
+ global_secondary_index {
+ hash_key = "status"
+ name = "gsi_status"
+ projection_type = "ALL"
}
# (8 unchanged attributes hidden)
}
# aws_iam_role_policy.orders_exec will be updated in-place
~ resource "aws_iam_role_policy" "orders_exec" {
id = "orders-role:orders-exec"
name = "orders-exec"
~ policy = jsonencode(
~ {
~ Statement = [
+ {
+ Action = "dynamodb:Query"
+ Effect = "Allow"
+ Resource = "arn:aws:dynamodb:us-east-1:123456789012:table/orders/index/gsi_status"
},
]
}
)
}
# aws_lambda_function.orders will be updated in-place
~ resource "aws_lambda_function" "orders" {
id = "orders"
~ environment {
~ variables = {
~ "LOG_LEVEL" = "info" -> "debug"
# (1 unchanged element hidden)
}
}
~ timeout = 30 -> 60
# (12 unchanged attributes hidden)
}
# aws_sqs_queue.orders_dlq must be replaced
-/+ resource "aws_sqs_queue" "orders_dlq" {
~ arn = "arn:aws:sqs:us-east-1:123456789012:orders-dlq" -> (known after apply)
~ id = "https://sqs.us-east-1.amazonaws.com/123456789012/orders-dlq" -> (known after apply)
~ name = "orders-dlq" -> "orders-dlq-v2" # forces replacement
~ url = "https://sqs.us-east-1.amazonaws.com/123456789012/orders-dlq" -> (known after apply)
# (7 unchanged attributes hidden)
}
# aws_sqs_queue.orders will be updated in-place
~ resource "aws_sqs_queue" "orders" {
id = "https://sqs.us-east-1.amazonaws.com/123456789012/orders"
~ redrive_policy = jsonencode(
~ {
~ deadLetterTargetArn = "arn:aws:sqs:us-east-1:123456789012:orders-dlq" -> "arn:aws:sqs:us-east-1:123456789012:orders-dlq-v2"
~ maxReceiveCount = 5 -> 3
}
)
# (8 unchanged attributes hidden)
}
Plan: 1 to add, 4 to change, 1 to destroy.
Do the analysis yourself first
Produce all four of these before you scroll:
- A table with one row per block: resource, action symbol, plain meaning, and whether any data is at risk.
- Every line carrying
# forces replacement, and the resource it sits under. - Which changes can share one apply and which must be staged separately.
- A short go/no-go note: apply order, the exact command sequence, and the single change you would split out β with the reason.
Check your answer
Classification. Four of the five resources are edited in place; one is destroyed and recreated.
| Block | Symbol | Plain meaning | Data at risk? |
|---|---|---|---|
aws_dynamodb_table.orders |
~ + new index |
in-place edit | No β index is built online |
aws_iam_role_policy.orders_exec |
~ |
in-place edit | No |
aws_lambda_function.orders |
~ |
in-place edit | No |
aws_sqs_queue.orders_dlq |
-/+ |
destroy-then-create | Yes β queued messages |
aws_sqs_queue.orders |
~ |
in-place edit | No |
Forces replacement. Exactly one line: name on aws_sqs_queue.orders_dlq. SQS queue names are fixed at creation, so the provider marks that attribute as requiring a new resource (ForceNew). Note the symbol: -/+ is destroy then create. Since Terraform 0.12 the plan distinguishes the two orderings β a create_before_destroy = true lifecycle rule on the queue would flip this to +/-, creating orders-dlq-v2 before the old queue is destroyed.
The description lied by omission. Nothing in the PR text says the dead-letter queue is renamed, yet a rename is precisely why the redrive policy below it is changing at all. A plan review that trusted the summary would have shipped a data-loss change as a config tweak.
Why the other three are safe together. The Lambda change patches function configuration only β no package redeploy, no -/+. The IAM inline policy update is in place and the new dynamodb:Query statement references the index ARN as a literal string, so there is no dependency edge to the table; the three edits are independent and their order does not matter. The fourth in-place edit, the redrive_policy on orders, is not: it points at orders-dlq-v2, so it travels with the DLQ replacement. In-place edits are also trivially reversible: revert the config and apply forward, since Terraform has no rollback.
Why the DLQ must be staged. It is the only stateful destroy in the plan. Bundling it with the three safe edits means a mid-apply failure leaves you unable to tell from the log which safe edits landed and which did not, while the destroyed queue is gone for good. Worse, the redrive policy on orders still points at the old queue name until the new DLQ exists, so during that window a message exceeding maxReceiveCount has nowhere to go.
The deliverable: a go/no-go note
GO for PR 1, GO for PR 2 as a separate apply. NO for applying both plans together.
Apply A (safe, reversible): DynamoDB GSI, IAM inline policy, Lambda env + timeout.
Apply B (stateful destroy): SQS DLQ replacement + the matching redrive_policy update
on the source queue, drained and reviewed on its own.
Split reason: Apply A's edits all show `~`. Apply B's `-/+` destroys
orders_dlq and every message in it; no rollback exists for that destroy.
The command sequence never skips the review step:
# Apply A - PR 1 only (DLQ rename reverted): non-destructive edits
terraform plan -out=tfplan-a
terraform show tfplan-a # render the saved diff for the reviewer
terraform apply tfplan-a # applies exactly the reviewed diff
# Apply B - DLQ replacement, after draining the old queue
terraform plan -out=tfplan-b
terraform show tfplan-b # confirm it still reads -/+ on orders_dlq
terraform apply tfplan-b
terraform plan -out=tfplan-a
β
terraform show tfplan-a (review the pinned diff)
β
terraform apply tfplan-a (exactly what was reviewed)
Two properties make this sequence trustworthy. First, terraform apply tfplan-a applies the saved diff rather than re-planning, so the thing you reviewed is the thing that runs. Second, the apply takes the state lock in the S3 backend for its duration, and if anything changed state between plan and apply, Terraform rejects the file as stale instead of applying a diff whose baseline moved β the stale plan behaviour covered in Applying with Control: Plan Files, State Locking, Targeting, and Ordering. In CI this is the only shape that is safe: the pipeline applies a plan file it already published for review, never a freshly computed plan.
Your turn: a variation
A second PR touches the same stack. The plan reads:
# aws_dynamodb_table.orders must be replaced
-/+ resource "aws_dynamodb_table" "orders" {
~ hash_key = "order_id" -> "order_id_v2" # forces replacement
# (10 unchanged attributes hidden)
}
# aws_lambda_function.orders will be updated in-place
~ resource "aws_lambda_function" "orders" {
~ source_code_hash = "e3b0c442..." -> "9f86d081..."
# (14 unchanged attributes hidden)
}
Plan: 1 to add, 1 to change, 1 to destroy.
Decide go/no-go, name the data-loss risk, and state the first command you would run after opening the PR.
Check your answer
NO-GO as a single apply. hash_key is ForceNew, so the table is destroyed and recreated β and a DynamoDB table's items are destroyed with it. The Lambda line is a different animal: source_code_hash changing is an in-place code update (~), safe and independently revertible.
The distinction that drives the decision is stateful versus stateless replacement, the triage rule from Destructive Changes and Cascades: Triage, Sequencing, and Minimizing Downtime: stateful replacement is a data-loss decision, stateless replacement is a scheduling decision. Here you have one of each, which is why the two lines must not travel in the same apply.
What protects you, and what does not. Point-in-time recovery lets you restore the old table's contents into a new table to an earlier moment β it is not an in-place undo, and it does not make the destroy safe. A reviewed plan file prevents surprise diffs, not data loss.
First command: terraform plan -out=tfplan if you need a pinned copy, but the real first move is a comment on the PR saying no, with the plan line quoted. The technical sequence, if the key change is genuinely needed, is: verify PITR and deletion protection on the existing table, provision a new table under the new key, backfill or dual-write, cut readers over, then retire the old table in a later apply.
Outcome checklist
- I can read
+,~,-, and-/+, and I know-/+means destroy then create. - I can find every
# forces replacementline and name the resource it belongs to. - I can tell stateful from stateless replacement before I decide anything.
- I can choose
-refresh-onlyfor drift versus a normal apply β reality wins with refresh-only, config wins with a normal apply β per Drift: When the Real Account No Longer Matches State or Config. - I can pin a reviewed diff with
-outandapply <file>, and explain what the state lock and a stale plan file each protect against. - I can sequence a destructive change β drain, replace, cut over β and say which single change deserves its own apply.