The Plan and Apply Loop

What Terraform does on init, plan and apply; desired configuration versus recorded state versus real infrastructure; reading a plan line by line (+, ~, -, -/+), 'known after apply', and what forces replacement.

Last generated

Lesson 3 of 6 available15 practice questions

SPACED REPETITION Β· 15 practice questions

Make this lesson stick.

Try 3 questions now. No account needed. Sample answers aren't saved.

Tracing init, plan and apply on a small SQS stack

You already know what a queue with a dead-letter queue (DLQ, the queue that receives messages that failed too many times) looks like in the AWS console. Suppose a teammate asks you to build the same pair in a fresh account, and then again in staging, and then to change one setting safely. Clicking works once. Terraform makes the second and third time boring, but only if you can predict what each command will do before you run it. This section walks one complete first run so that prediction becomes routine.

The stack we will trace

Here is the whole configuration, the set of .tf files that describes what you want. Save it as main.tf in an empty directory.

terraform {
  required_version = ">= 1.9"
  required_providers {
    aws = {
      source  = "hashicorp/aws"
      version = "~> 6.0" # pin to the current major version of the provider
    }
  }
}

provider "aws" {
  region = "eu-west-1"
}

resource "aws_sqs_queue" "dlq" {
  name = "orders-dlq"
}

resource "aws_sqs_queue" "main" {
  name                       = "orders"
  visibility_timeout_seconds = 60
  redrive_policy = jsonencode({
    deadLetterTargetArn = aws_sqs_queue.dlq.arn
    maxReceiveCount     = 5
  })
}

Four terms appear here that the rest of the lesson relies on:

  • A provider is the plugin that translates Terraform's requests into API calls for one platform. hashicorp/aws talks to AWS. Note it contains only a region: no keys, as we'll see below.
  • A resource address is the unique name Terraform gives each thing it manages: resource type plus your label, such as aws_sqs_queue.main. Plans, state and error messages all use addresses.
  • An argument is a value you set inside a resource block, such as name or visibility_timeout_seconds.
  • An attribute is a value AWS reports back after creating the thing, such as arn, id or url. You can read any attribute from another block, which is what aws_sqs_queue.dlq.arn does. (Many arguments are also readable as attributes.)

required_version = ">= 1.9" makes Terraform refuse to run on an older CLI. The redrive_policy is the JSON that tells SQS where failed messages go. jsonencode turns an HCL object into that JSON string.

terraform init: prepare the working directory

$ terraform init
...
Terraform has been successfully initialized!

init prepares the working directory. It is not purely local: it downloads providers (and any modules) over the network and, with a remote backend, connects to that backend. For this stack it does three things:

  1. Downloads the provider plugin into a hidden .terraform/ directory.
  2. Configures the backend, the place where state is stored. With no backend block, the backend is local and state will live in terraform.tfstate beside your files. Team setups use a remote backend instead; the loop is identical.
  3. Writes .terraform.lock.hcl, recording the exact provider version and checksums chosen. Commit this file so teammates and CI install the same provider build. Do not commit .terraform/.

init never creates, changes or reads the AWS resources your configuration manages. Re-run it after you add a provider or module, change the backend configuration, or clone the repo fresh. Terraform will tell you when it is needed.

Signing in without static keys

The provider block has no credentials because it picks them up from the standard AWS credential chain. With IAM Identity Center:

$ aws sso login --profile dev
$ export AWS_PROFILE=dev
$ terraform plan

aws sso login opens a browser and caches an IAM Identity Center access token locally; the AWS CLI and the Terraform provider exchange it for short-lived credentials. AWS_PROFILE tells both the AWS CLI and the Terraform provider which profile to use. When the SSO session expires, plan fails while the provider validates credentials, with a message along the lines of failed to refresh cached credentials mentioning an expired SSO token. The exact wording varies by provider version. The fix is always the same: run aws sso login --profile dev again. Do not paste access keys into the provider block.

Three things Terraform compares

Terraform tracks three separate things:

  • Configuration: your .tf files, the desired state.
  • State: Terraform's own record, a map from each resource address to the real ID and last-known attributes of what it created.
  • Real infrastructure: what AWS actually holds right now.

terraform plan works in two steps:

1. Refresh : ask the AWS API about every resource in state, update state's view
2. Diff    : compare configuration against that refreshed state
   Result  : a plan (create / change / destroy) shown to you, nothing changed yet

On the first run the state is empty, so every resource in the configuration is a difference.

Reading the first plan

$ terraform plan

Terraform will perform the following actions:

  # aws_sqs_queue.dlq will be created
  + resource "aws_sqs_queue" "dlq" {
      + arn  = (known after apply)
      + id   = (known after apply)
      + name = "orders-dlq"
      + url  = (known after apply)
      # (other attributes omitted)
    }

  # aws_sqs_queue.main will be created
  + resource "aws_sqs_queue" "main" {
      + arn                        = (known after apply)
      + name                       = "orders"
      + redrive_policy             = (known after apply)
      + visibility_timeout_seconds = 60
      # (other attributes omitted)
    }

Plan: 2 to add, 0 to change, 0 to destroy.

(Real output lists more attributes; they are trimmed here.) Each + means create. Arguments you wrote show their values; attributes AWS will assign show (known after apply). The redrive_policy is also unknown even though you wrote it, because it embeds the DLQ's ARN, which does not exist yet.

That reference is also how Terraform orders the work. It builds a dependency graph: because main reads aws_sqs_queue.dlq.arn, dlq must be created first. Resources with no reference between them have no ordering constraint and run in parallel.

Applying and proving idempotence

$ terraform apply
...
Do you want to perform these actions?
  Enter a value: yes

aws_sqs_queue.dlq: Creating...
aws_sqs_queue.dlq: Creation complete after 1s [id=https://sqs.eu-west-1.amazonaws.com/123456789012/orders-dlq]
aws_sqs_queue.main: Creating...
aws_sqs_queue.main: Creation complete after 1s [id=https://sqs.eu-west-1.amazonaws.com/123456789012/orders]

Apply complete! Resources: 2 added, 0 changed, 0 destroyed.

apply shows the plan again and nothing happens until you type exactly yes. It then walks the graph, creating dlq and only afterwards main, and records each resource in state as it completes. The state now maps aws_sqs_queue.main to its real queue URL and attributes. Run terraform plan again:

No changes. Your infrastructure matches the configuration.

This is idempotence: re-running the loop on an unchanged configuration does nothing. The refresh found AWS matching state, and the diff found state matching configuration.

Your turn: predict before you plan

You now extend the configuration. Assume aws_lambda_function.worker is already declared elsewhere in the configuration and was applied last week, so it is in state. Nothing from the SQS stack has been applied yet. You add:

resource "aws_dynamodb_table" "orders" {
  name         = "orders"
  billing_mode = "PAY_PER_REQUEST"
  hash_key     = "order_id"

  attribute {
    name = "order_id"
    type = "S"
  }
}

resource "aws_lambda_event_source_mapping" "worker" {
  event_source_arn = aws_sqs_queue.main.arn
  function_name    = aws_lambda_function.worker.arn
}

Task: Before running anything, write down (1) the number in Plan: N to add, and (2) which resources must be created before which, using only the references in the code.

Check your answer

(1) Plan: 4 to add, 0 to change, 0 to destroy. The new resources are dlq, main, the table and the mapping. The function is already in state and matches the configuration, so it is not in the count.

(2) Ordering from references:

  • main reads aws_sqs_queue.dlq.arn, so dlq comes before main.
  • The mapping reads aws_sqs_queue.main.arn, so it comes after main (and therefore after dlq). Its reference to the function adds no wait, because the function already exists.
  • The DynamoDB table references nothing and nothing references it, so it can be created in parallel with dlq at the start.

If the function were new in the same plan, the count would be 5, and the mapping would wait for both main and the function.

If your count was off, look for the resource you forgot was already in state: the plan counts only differences, not everything in the configuration. Next, '[Reading a plan line by line: symbols, in-place updates and known after apply]' covers the other markers beyond +.

Reading a plan line by line: symbols, in-place updates and 'known after apply'

A teammate asks you to raise a queue's visibility timeout and add a cost tag. The plan scrolls past with a dozen lines you have never seen. Which of them could interrupt message processing, and which are bookkeeping? The plan is the only preview Terraform gives you, so the first skill is reading it quickly and accurately.

The symbols

Every resource in a plan starts with a header comment (# aws_sqs_queue.main will be updated in-place) that names its resource address and the action. The line below it starts with a symbol that gives the same action in compact form.

Symbol Action Serverless example
+ create new aws_sqs_queue.audit
~ update in place Lambda timeout 30 -> 60
- destroy deleted aws_sqs_queue.old block
-/+ destroy, then create (replace) DynamoDB hash_key change
+/- create, then destroy same replace, with create_before_destroy = true
<= read a data source data.aws_iam_policy_document read during apply

A data source (a read-only lookup of something that already exists) normally runs during plan and stays out of the output. It shows as <= with will be read during apply only when its read has to wait: usually because one of its arguments is not yet known, which the known after apply section below explains, or because it depends on a resource with changes pending.

The last line is the summary: Plan: 1 to add, 3 to change, 1 to destroy. A replacement counts as one add and one destroy. So 0 to destroy is the fastest safety check you have.

Worked example: a timeout and a tag

The queue aws_sqs_queue.main has visibility_timeout_seconds = 30 and tags = { Env = "prod" }. You edit the configuration to 90 and add Owner = "payments". The plan, trimmed to the interesting lines:

  # aws_sqs_queue.main will be updated in-place
  ~ resource "aws_sqs_queue" "main" {
        id                         = "https://sqs.eu-west-1.amazonaws.com/111122223333/orders"
        name                       = "orders"
      ~ tags                       = {
          + "Owner" = "payments"
            # (1 unchanged element hidden)
        }
      ~ tags_all                   = {
          + "Owner" = "payments"
            # (1 unchanged element hidden)
        }
      ~ visibility_timeout_seconds = 30 -> 90
        # (unchanged attributes hidden)
    }

Plan: 0 to add, 1 to change, 0 to destroy.

Read it in order:

  1. Header and symbol: will be updated in-place and ~. The queue keeps its identity.
  2. Attribute lines: old -> new shows the change. + inside a map adds a key. Unmarked lines are context.
  3. No # forces replacement anywhere, and the summary shows 0 to destroy. The queue URL and its messages survive.

⚠️ In place does not mean harmless. A Lambda reading this queue should have a function timeout comfortably below the visibility timeout (AWS recommends the queue value be at least six times the function timeout, plus any batching window). Lowering visibility_timeout_seconds below the function timeout would also plan as a calm ~ line, yet it would make messages reappear while they are still being processed. Lambda compares the two timeouts only when an event source mapping is created or updated, so a change to the queue alone is not caught. The symbol tells you the mechanism. Only the attribute tells you the risk.

(known after apply)

Some attributes are assigned by AWS only once the resource exists: arn, id, url. For a new resource the plan cannot print them:

  # aws_sqs_queue.audit will be created
  + resource "aws_sqs_queue" "audit" {
      + arn  = (known after apply)
      + id   = (known after apply)
      + name = "audit"
      + url  = (known after apply)
    }

  # aws_lambda_event_source_mapping.audit will be created
  + resource "aws_lambda_event_source_mapping" "audit" {
      + event_source_arn = (known after apply)
      + function_name    = "audit-worker"
    }

The unknown value spreads: anything that references aws_sqs_queue.audit.arn is unknown too. That is normal, and the graph orders the creation for you. On a replacement you will also see ~ arn = "arn:..." -> (known after apply), a quiet hint that dependents will receive a new value.

The unknown becomes an error when Terraform needs it at plan time to decide how many copies of a resource exist. A for_each key or a count must be known during planning:

resource "aws_ssm_parameter" "url" {
  # BROKEN when aws_sqs_queue.main is new in this plan: its url is not known yet
  for_each = toset([aws_sqs_queue.main.url])
  name     = "/app/queue_url"
  type     = "String"
  value    = each.key
}
Error: Invalid for_each argument

  on main.tf line 3, in resource "aws_ssm_parameter" "url":
   3:   for_each = toset([aws_sqs_queue.main.url])

The "for_each" set includes values derived from resource attributes that
cannot be determined until apply, and so Terraform cannot determine the
full set of keys that will identify the instances of this resource.

The fix is to key on values you write yourself and let only the value be unknown:

locals {
  queues = toset(["orders", "billing"])
}

resource "aws_sqs_queue" "q" {
  for_each = local.queues
  name     = each.key
}

resource "aws_ssm_parameter" "url" {
  for_each = local.queues # keys are known at plan time
  name     = "/app/${each.key}/queue_url"
  type     = "String"
  value    = aws_sqs_queue.q[each.key].url # unknown now, fine
}

Unknown count values fail the same way, with Invalid count argument.

Noise versus signal

  • tags_all churn. If you change default_tags in the provider block, every taggable resource shows ~ tags_all, which can mean dozens of "to change" with no real edit. Confirm each changed resource differs only in tags_all, then discount them.
  • (sensitive value) masks a secret in the plan output. A ~ line with this text tells you something changed but not what. The real value is still stored in state, which is why state must be protected.
  • jsonencode policies show as a structured diff, not one long string, so you can see the exact statement that changed:
  # aws_iam_role_policy.worker will be updated in-place
  ~ resource "aws_iam_role_policy" "worker" {
      ~ policy = jsonencode(
          ~ {
              ~ Statement = [
                  ~ {
                      ~ Action   = [
                            "sqs:ReceiveMessage",
                          + "sqs:DeleteMessage",
                        ]
                        # (2 unchanged attributes hidden)
                    },
                ]
                # (1 unchanged attribute hidden)
            }
        )
        # (unchanged attributes hidden)
    }

Reading order: (1) the summary, especially the destroy count; (2) every header whose symbol is not + or ~; (3) the old -> new lines under each ~, because policies and timeouts live there; (4) tags_all noise last. This is a triage heuristic, not a substitute for reading what each attribute does.

Guided attempt

Here are the headers and key lines of a plan for a Lambda function, a queue and an IAM role policy:

  # aws_iam_role_policy.worker must be replaced
-/+ resource "aws_iam_role_policy" "worker" {
      ~ id   = "worker-role:worker" -> (known after apply)
      ~ name = "worker" -> "worker-v2" # forces replacement
    }

  # aws_lambda_function.worker will be updated in-place
  ~ resource "aws_lambda_function" "worker" {
      ~ timeout = 30 -> 60
    }

  # aws_sqs_queue.retry will be created
  + resource "aws_sqs_queue" "retry" {
      + arn = (known after apply)
    }

Task: give each resource's action, then write the summary line.

Check your answer

The Lambda function is an in-place update (1 change). The retry queue is a create (1 add). The role policy is a replace: 1 add and 1 destroy.

Summary: Plan: 2 to add, 1 to change, 1 to destroy. The destroy count is nonzero because of the -/+ line, so that is the line to read first. The # forces replacement comment on name names the cause.

Independent practice

The worker Lambda has timeout = 30 and is triggered by aws_sqs_queue.main.

  # aws_cloudwatch_log_group.worker will be updated in-place
  ~ resource "aws_cloudwatch_log_group" "worker" {
      ~ retention_in_days = 14 -> 30
    }

  # aws_dynamodb_table.orders will be updated in-place
  ~ resource "aws_dynamodb_table" "orders" {
      ~ tags     = { + "CostCenter" = "1042" }
      ~ tags_all = { + "CostCenter" = "1042" }
    }

  # aws_sqs_queue.audit will be created
  + resource "aws_sqs_queue" "audit" {
      + arn = (known after apply)
      + url = (known after apply)
    }

  # aws_sqs_queue.main will be updated in-place
  ~ resource "aws_sqs_queue" "main" {
      ~ visibility_timeout_seconds = 90 -> 5
    }

Tasks: (1) List each resource's action and write the summary line. (2) Which single change affects live message processing, and how? (3) What would differ if aws_sqs_queue.main showed -/+ instead of ~?

Check your answer
  1. Log group: update. DynamoDB table: update (tags only). audit queue: create. main queue: update. Summary: Plan: 1 to add, 3 to change, 0 to destroy.
  2. The main queue's visibility_timeout_seconds = 90 -> 5. The function can run for up to 30 seconds, but a received message becomes visible again after 5, so other consumers can pick up the same message while it is still being processed and duplicates result. The log retention, the tags and the new audit queue touch nothing live.
  3. A -/+ would delete the queue and create a new one. Waiting and in-flight messages would be lost, and the new queue would get a new ARN and URL (known after apply) that the event source mapping and any producers must pick up. The summary would read 2 to add, 2 to change, 1 to destroy, because the replace adds one add and one destroy. If the event source mapping is managed in this configuration and reads aws_sqs_queue.main.arn, it is replaced too (event_source_arn forces replacement), adding one more add and one more destroy. The same edit as ~ keeps the messages, the URL and the ARN.

Next, '# forces replacement' gets its own treatment in 'What forces replacement and how to avoid destroying live data'.

What forces replacement and how to avoid destroying live data

You change one line in a DynamoDB table definition, run terraform plan, and the summary reads 1 to add, 0 to change, 1 to destroy. The production table holds every order your system has taken. Terraform did not make a mistake: AWS cannot edit that argument on an existing table, so the only way to match your configuration is to delete the table and build a new one. This section teaches you to see that coming and stop it.

The annotation that names the culprit

A replacement means Terraform destroys the real resource and creates a new one in its place. The plan shows it with the -/+ marker, and it puts the comment # forces replacement on the exact argument that caused it. Take this table:

resource "aws_dynamodb_table" "orders" {
  name         = "orders"
  billing_mode = "PAY_PER_REQUEST"
  hash_key     = "orderId" # the partition key: the attribute DynamoDB uses to place items

  attribute {
    name = "orderId"
    type = "S"
  }
}

Edit hash_key to "order_id" (and the matching attribute name) and the plan reads:

  # aws_dynamodb_table.orders must be replaced
-/+ resource "aws_dynamodb_table" "orders" {
      ~ arn      = "arn:aws:dynamodb:eu-west-1:111122223333:table/orders" -> (known after apply)
      ~ hash_key = "orderId" -> "order_id" # forces replacement
      ~ id       = "orders" -> (known after apply)
        name     = "orders"
        # (other attributes hidden)
    }

Plan: 1 to add, 0 to change, 1 to destroy.

The header says must be replaced, and the # forces replacement comment points at hash_key. Everything else with known after apply is just the new table getting a fresh ARN.

Compare that with edits AWS can apply to a live resource:

Change Plan marker Why
Lambda timeout ~ Function configuration is updatable
DynamoDB billing_mode ~ AWS switches modes on a live table
SQS visibility_timeout_seconds ~ Queue attribute, editable
DynamoDB hash_key or name -/+ Key schema and name are fixed at creation
SQS fifo_queue or name -/+ Queue type and name are fixed at creation
IAM role name -/+ IAM cannot rename a role

This table is a starting point, not the authority. The authority is the plan for your exact change, backed by the argument notes in the AWS provider documentation for that resource. Always read the plan.

What a replacement costs

Resource What you lose
DynamoDB table All items
SQS queue In-flight and queued messages; the URL changes
IAM role The ARN changes; dependents must pick it up
ECS service A gap in service while tasks stop and start

For the IAM role, Terraform handles references that go through the attribute, such as role = aws_iam_role.worker.arn. A Lambda function's role then shows as an in-place update, while an inline aws_iam_role_policy, whose role argument forces replacement, is replaced along with the role. A hardcoded ARN in a string will not follow. For the SQS queue, producers that hardcode the old URL keep sending to a queue that no longer exists.

The rename trap

The most dangerous replacements involve no AWS argument at all. Suppose you tidy the label:

resource "aws_dynamodb_table" "orders_v2" { # was "orders"
  name         = "orders"
  billing_mode = "PAY_PER_REQUEST"
  hash_key     = "orderId"

  attribute {
    name = "orderId"
    type = "S"
  }
}

The label is part of the resource address, which is the key Terraform uses to find the real table in state. State has aws_dynamodb_table.orders, and the configuration no longer does, so the plan shows aws_dynamodb_table.orders will be destroyed, followed by (because aws_dynamodb_table.orders is not in configuration), and aws_dynamodb_table.orders_v2 will be created. This is the same real table, planned as delete plus create. Since the new block still says name = "orders", the create can also collide with the table that still exists.

The fix is a moved block (Terraform 1.1 or later), which tells Terraform that the thing at the old address is now at the new one:

moved {
  from = aws_dynamodb_table.orders
  to   = aws_dynamodb_table.orders_v2
}

The plan should now say # aws_dynamodb_table.orders has moved to aws_dynamodb_table.orders_v2 and the summary should show 0 to destroy. (When the moved resource also has other changes, the move appears instead as a # (moved from aws_dynamodb_table.orders) line under its update or replace header.) If it still shows a destroy, stop. Either the from address is mistyped or another argument is forcing replacement. Keep the moved block until every environment has applied it.

Guardrails

⚠️ prevent_destroy makes the plan fail instead of destroying:

resource "aws_dynamodb_table" "orders" {
  name                        = "orders"
  billing_mode                = "PAY_PER_REQUEST"
  hash_key                    = "orderId"
  deletion_protection_enabled = true # AWS itself refuses DeleteTable

  attribute {
    name = "orderId"
    type = "S"
  }

  lifecycle {
    prevent_destroy = true # Terraform refuses to plan a destroy
  }
}

With it in place, a hash_key edit produces Error: Instance cannot be destroyed ... has lifecycle.prevent_destroy set, but the plan calls for this resource to be destroyed. The error blocks the whole plan, which is the point. Its limit is that it lives in the configuration. If you delete the whole resource block, the guard goes with it.

deletion_protection_enabled works at a different layer: it is a real DynamoDB table setting, so even a console user or a different tool cannot delete the table. It only takes effect after it has been applied, and a replacement plan still looks fine because the refusal happens at apply time. Use both guards on tables you cannot rebuild.

create_before_destroy reverses the order to +/-: the new resource is built first, then the old one is removed. It suits resources whose replacement gets a new name. A fixed queue name cannot give you two queues, because SQS allows only one queue per name in an account and Region: CreateQueue either rejects the request (when the attributes differ) or returns the existing queue (when they match), and the destroy step would then delete that very queue. The usual workaround is name_prefix, which lets AWS-side names differ:

resource "aws_sqs_queue" "jobs" {
  name_prefix = "jobs-"

  lifecycle {
    create_before_destroy = true
  }
}

The URL still changes, so this protects availability, not producers that hardcode the URL.

Replacing on purpose

Sometimes you want a replacement, for example to reset a broken ECS service. Use terraform apply -replace="aws_ecs_service.api" (since 0.15.2), which supersedes the old terraform taint command. It still shows -/+ because it is a real replacement, but the header reads # aws_ecs_service.api will be replaced, as requested and no argument carries # forces replacement. Run terraform plan -replace=... first to review it. If you only need new tasks, an in-place redeployment avoids the replacement entirely: with force_new_deployment = true on the service, any in-place update other than a tags-only change, such as a changed value in its triggers map, starts a new deployment.

Exercise: fix the plan

Context. Production has a table created as aws_dynamodb_table.orders with hash_key = "orderId". A teammate's branch contains:

resource "aws_dynamodb_table" "orders_v2" {
  name         = "orders"
  billing_mode = "PAY_PER_REQUEST"
  hash_key     = "order_id"

  attribute {
    name = "order_id"
    type = "S"
  }
}

The plan shows the old table destroyed and a new one created. Task: decide among reverting, adding a moved block, or building a new table with a data migration. Write the configuration whose plan destroys nothing. Success criterion: the plan shows has moved to and 0 to destroy.

Check your answer

There are two separate causes. The label change orders to orders_v2 is the rename trap, fixed with a moved block. The hash_key and attribute change from orderId to order_id forces replacement on its own, and a moved block cannot fix it, because DynamoDB cannot change a key schema. Without reverting it, the plan would still show -/+ after the move.

resource "aws_dynamodb_table" "orders_v2" {
  name         = "orders"
  billing_mode = "PAY_PER_REQUEST"
  hash_key     = "orderId" # reverted

  attribute {
    name = "orderId" # reverted
    type = "S"
  }

  lifecycle {
    prevent_destroy = true
  }
}

moved {
  from = aws_dynamodb_table.orders
  to   = aws_dynamodb_table.orders_v2
}

The plan should show has moved to and 0 to destroy. If the team truly needs the order_id key, the only safe route is a new table with a different name under a new label. Copy the data across, switch the application to it, and retire the old table later as a separate, reviewed step.

When reality disagrees: drift, refresh-only and half-finished applies

A teammate is debugging poison messages and, in the SQS console, raises the main queue's maxReceiveCount (how many times a message is received before it moves to the dead-letter queue) from 3 to 10. Nobody touches the Terraform files. On Monday you change an unrelated tag and run terraform plan. What does Terraform do about the console edit? Terraform compares configuration, state and real AWS, so it notices. What happens next depends on a decision you make per attribute.

Drift: configuration wins by default

Drift means real infrastructure no longer matches what Terraform last recorded. The plan's first step is a refresh, which reads current attributes from the AWS API and compares them with state. Here is the relevant configuration, followed by an illustrative, abridged plan and the output of terraform plan -refresh-only. They are not captured from a run, and exact formatting varies by provider version.

resource "aws_sqs_queue" "main" {
  name = "orders"
  redrive_policy = jsonencode({
    deadLetterTargetArn = aws_sqs_queue.dlq.arn
    maxReceiveCount     = 3
  })
}
$ terraform plan
  # aws_sqs_queue.main will be updated in-place
  ~ resource "aws_sqs_queue" "main" {
      ~ redrive_policy = jsonencode(
          ~ { ~ maxReceiveCount = 10 -> 3 }
        )
    }

Plan: 0 to add, 1 to change, 0 to destroy.

$ terraform plan -refresh-only
Note: Objects have changed outside of Terraform

  # aws_sqs_queue.main has changed
  ~ resource "aws_sqs_queue" "main" {
      ~ redrive_policy = jsonencode(
          ~ { ~ maxReceiveCount = 3 -> 10 }
        )
    }

The two outputs point in opposite directions. The ordinary plan shows only the proposed action, reality versus configuration (10 back to 3), and nothing in it says the 10 came from a console edit: it looks like any edit to the file. Since Terraform 1.2 an ordinary plan prints the Note: Objects have changed outside of Terraform section only for outside changes that feed another planned change (for example, an attribute that a changing resource reads), and never when the plan has no changes. The refresh-only plan shows the note, state versus reality (3 became 10). It is information and changes nothing. So when a ~ line surprises you, run terraform plan -refresh-only to see what changed outside Terraform. Configuration is the desired state, so unless you decide otherwise, apply reverts the console edit.

Accept, revert or ignore

Decide for each drifted attribute separately.

Choice What you do Right when
Accept Edit the configuration to the new value The change was wanted
Revert Apply as proposed The change was a mistake or temporary
Ignore lifecycle { ignore_changes } Another system owns the value

To accept, set maxReceiveCount = 10 in the file. The next plan prints no note and reports No changes., because configuration now matches reality.

Ignore is for values something else legitimately manages. An ECS service under target-tracking autoscaling has its desired_count changed by AWS all day. Without an ignore, every plan proposes resetting it to the configured number, and a careless apply would shrink a loaded service:

# Fragment: the other arguments of aws_ecs_service.api are omitted.
resource "aws_ecs_service" "api" {
  desired_count = 2 # used only at creation; autoscaling owns it afterwards

  lifecycle {
    ignore_changes = [desired_count]
  }
}

⚠️ ignore_changes is a permanent blind spot. Terraform stops reporting that attribute whether the cause is autoscaling or a bad console edit, and it also ignores your own later edits to it in the file. Ignore single attributes, never all, and write a comment naming the system that owns the value.

Refresh-only: updating state without touching AWS

terraform plan -refresh-only shows what drifted and proposes only state updates. terraform apply -refresh-only (both since Terraform 0.15.4) writes them. No infrastructure changes either way.

It is easy to misread as "accept the drift". It only updates the state file. If your configuration still says 3, the next ordinary plan still proposes reverting to 3. Refresh-only is right after you have already edited the configuration to match, or when you want state to reflect a resource deleted outside Terraform. It merely records a mistake when the console edit was wrong: state then holds 10, so not even plan -refresh-only shows the drift any more, and the ordinary plan's 10 -> 3 looks like any other edit.

Inspecting state

terraform state list                      # every address Terraform tracks
terraform state show aws_sqs_queue.main   # recorded attributes of one resource
terraform show                            # the whole state, readable

Use these to confirm what Terraform believes, for example which queue URL an address maps to. ⚠️ State holds attribute values in plain text, including values you marked sensitive, such as a database password passed to a resource. Restrict who can read the state backend, and never paste state show output into a ticket.

A half-finished apply

Suppose a new change adds three resources: aws_iam_role.worker, then aws_lambda_function.worker (which references the role), then aws_lambda_event_source_mapping.orders (which references the function and the queue). Apply runs:

aws_iam_role.worker: Creation complete after 2s [id=worker-role]
aws_lambda_function.worker: Creating...

Error: creating Lambda Function (worker): AccessDeniedException:
  User: .../AWSReservedSSO_Developer_... is not authorized to perform:
  iam:PassRole on resource: arn:aws:iam::111122223333:role/worker-role

Nothing rolls back. State records the role, which succeeded. The function failed, and the mapping was never attempted because it depends on the function. Fix the missing iam:PassRole permission, which lives in your permission set or CI role, not in this configuration. Then run terraform plan. It reads state, sees the role exists, and proposes only the remainder: Plan: 2 to add, 0 to change, 0 to destroy. Apply again.

⚠️ Read the error and the new plan before re-running. If a resource was created but failed a later step, Terraform marks it tainted, which means flagged as untrustworthy. The next plan then shows -/+ under the header is tainted, so must be replaced, so check that replacing it is safe, as covered in 'What forces replacement and how to avoid destroying live data'.

Practice: two drift notes and one real change

The illustrative, abridged output below is for your stack: first terraform plan -refresh-only, then terraform plan. The configuration file already says maxReceiveCount = 3, desired_count = 2 and memory_size = 512.

$ terraform plan -refresh-only
Note: Objects have changed outside of Terraform
  # aws_ecs_service.api has changed
      ~ desired_count = 2 -> 5        (target-tracking autoscaling)
  # aws_sqs_queue.main has changed
      ~ maxReceiveCount = 3 -> 10     (ticket OPS-1, closed: "temporary, while debugging poison messages")

$ terraform plan
  # aws_ecs_service.api will be updated in-place
      ~ desired_count = 5 -> 2
  # aws_lambda_function.worker will be updated in-place
      ~ memory_size = 256 -> 512
  # aws_sqs_queue.main will be updated in-place
      ~ maxReceiveCount = 10 -> 3

Plan: 0 to add, 3 to change, 0 to destroy.

Task: For each of the three changes, choose accept, revert, ignore or "apply as intended", and justify it. Then state what the summary line should say before you apply.

Check your answer
  • maxReceiveCount: revert. The ticket says the change was temporary and is closed, and the configuration already holds the intended value, so applying restores 3.
  • desired_count: ignore. Autoscaling legitimately owns the value. Add ignore_changes = [desired_count] to aws_ecs_service.api before applying. Otherwise this apply would cut a service that autoscaling scaled up to 5 back to 2 under load.
  • memory_size: apply as intended. This is not drift. You edited the configuration, and the plan shows an in-place update.

After adding the ignore, re-plan. The ECS change disappears, and the summary should read Plan: 0 to add, 2 to change, 0 to destroy., covering the queue revert and the Lambda edit.

Running the loop safely and gating a risky plan

You can now read a plan. The remaining risk is procedural: the plan you reviewed is not necessarily the change that gets applied. Between your review and terraform apply, a teammate may have merged something, or someone may have edited a resource in the console. This section turns review into a fixed step: a saved plan, a CI shape around it, and an ordered go/no-go checklist.

Saved plans: apply exactly what was reviewed

terraform plan -out=tfplan   # write the computed actions to a file
terraform show tfplan        # human-readable review of that file
terraform apply tfplan       # perform those actions; no re-plan, no yes prompt

A saved plan is a file holding the exact actions Terraform computed. Plain terraform apply re-plans first, so what you approved and what runs can differ. terraform apply tfplan skips the re-plan, and the file itself is the approval, which is why there is no prompt. Two properties matter:

  • Outdated plans are rejected. The file records the state it was built from. If state changed afterwards, Terraform refuses with a "Saved plan is stale" error. You re-plan and re-review. A console edit made after the plan changes no state, so it is not detected: the saved actions are applied as reviewed.
  • Plan files can hold secrets. They can contain resolved variable and attribute values, including sensitive ones, and they are not encrypted. Never commit one. Keep it as an access-restricted CI artifact and delete it after apply.

Predict: you save a plan at 10:00. A teammate applies a different change at 10:30. At 11:00 you run terraform apply tfplan. What happens?

Check your answer

Terraform rejects the plan as stale and changes nothing. The teammate's apply updated state after your plan was made, so your reviewed actions may no longer be correct. Re-run terraform plan -out=tfplan and review again. The rejection is the safeguard working.

The CI shape

CI automates the same loop. No workflow YAML is needed to see the shape:

pull request opened
   ↓
CI assumes the AWS role through OIDC (no static keys)
   ↓
terraform init -input=false
   ↓
terraform plan -input=false -lock-timeout=5m -detailed-exitcode -out=tfplan
   ↓
plan posted on the PR; a human applies the checklist below
   ↓
merge
   ↓
terraform apply -input=false -lock-timeout=5m tfplan

The provider block stays region only. Locally the credentials come from your AWS_PROFILE after aws sso login. In CI they come from the role the job assumed. The flags:

  • -input=false makes Terraform fail instead of waiting at a prompt for a missing variable. An unattended job should never hang.
  • -lock-timeout=5m waits up to five minutes for the state lock (held by another run) instead of failing immediately. It only helps if the backend locks: the S3 backend locks only when you configure it (use_lockfile = true, Terraform 1.10 or later; the older dynamodb_table setting is deprecated).
  • -detailed-exitcode makes plan report its result through the exit code:
Exit code Meaning CI reaction
0 Success, no changes Skip apply
1 Error Fail the job
2 Success, changes present Post plan, require review

πŸ’‘ Pro Tip: a shell step that stops on any non-zero code treats 2 as a failure. Capture the code in a variable and branch on it.

A plan saved on the PR can go stale if another PR merges first. Teams either re-plan after merge and re-approve, or pass the PR's plan file along and let the stale-plan rejection force a fresh review. This is a simplified flow: production setups add approval rules and per-environment state.

The go/no-go checklist, in order

Read in this order, because each step catches the most damaging problems first:

  1. Destroy count in the summary. Any non-zero "to destroy" needs an explanation before you read further.
  2. Every -/+ and every # forces replacement. The annotation names the argument to fix.
  3. Unexpected resource addresses. A resource you did not touch means a rename, a changed module, or a mistaken for_each.
  4. (known after apply) on values that matter. Check what depends on them: a queue URL, an ARN inside an IAM policy, a name another system hardcodes.
  5. tags_all noise. Discount it when default_tags changed; it is not signal.

Here is the checklist applied to a small plan:

  # aws_dynamodb_table.orders must be replaced
-/+ resource "aws_dynamodb_table" "orders" {
      ~ hash_key = "id" -> "orderId" # forces replacement
    }

Plan: 1 to add, 0 to change, 1 to destroy.

Step 1 stops you: one destroy. Step 2 shows a -/+ on a table because hash_key changed. Replacing the table deletes its items, so the verdict is no-go until the edit is reverted or a migration is planned.

Independent task: gate this plan

A developer wants to rename the orders queue, give the worker Lambda the queue URL, and let it update items. The plan is abbreviated, with unchanged-attribute noise trimmed. The Lambda reads the queue through an event source mapping (the resource that connects the queue to the function).

  # aws_iam_role_policy.worker_ddb will be updated in-place
  ~ resource "aws_iam_role_policy" "worker_ddb" {
        id     = "orders-worker:ddb"
      ~ policy = jsonencode(
          ~ {
              ~ Statement = [
                  ~ {
                      ~ Action   = [
                          + "dynamodb:UpdateItem",
                            "dynamodb:PutItem",
                        ]
                      ~ Resource = "arn:aws:dynamodb:eu-west-1:111122223333:table/orders" -> "*"
                        # (2 unchanged attributes hidden)
                    },
                ]
            }
        )
    }

  # aws_lambda_event_source_mapping.jobs must be replaced
-/+ resource "aws_lambda_event_source_mapping" "jobs" {
      ~ event_source_arn = "arn:aws:sqs:eu-west-1:111122223333:orders-queue" -> (known after apply) # forces replacement
      ~ id               = "5c1e8f0a-example" -> (known after apply)
        # (4 unchanged attributes hidden)
    }

  # aws_lambda_function.worker will be updated in-place
  ~ resource "aws_lambda_function" "worker" {
        id = "orders-worker"
      ~ environment {
          ~ variables = {
              + "QUEUE_URL"  = (known after apply)
                "TABLE_NAME" = "orders"
            }
        }
        # (22 unchanged attributes hidden)
    }

  # aws_sqs_queue.jobs must be replaced
-/+ resource "aws_sqs_queue" "jobs" {
      ~ arn  = "arn:aws:sqs:eu-west-1:111122223333:orders-queue" -> (known after apply)
      ~ id   = "https://sqs.eu-west-1.amazonaws.com/111122223333/orders-queue" -> (known after apply)
      ~ name = "orders-queue" -> "orders-jobs" # forces replacement
        # (9 unchanged attributes hidden)
    }

Plan: 2 to add, 2 to change, 2 to destroy.

Task: issue go or no-go, name every hazard, and write the minimal configuration fix. State what the revised summary should read.

Check your answer

No-go. Hazards, in checklist order:

  1. Destroy count is 2. The queue and the mapping are both destroyed.
  2. name forces replacement of the queue. The SQS queue name cannot change in place. Replacing it discards in-flight messages and gives a new URL. The mapping is replaced as a consequence, because event_source_arn changes, so consumption pauses.
  3. The IAM policy change is not the intended tweak. Adding dynamodb:UpdateItem is right, but Resource widens from one table to "*", which grants access to every DynamoDB table in the account.
  4. QUEUE_URL = (known after apply) is harmless in itself, but it exists only because the queue is being replaced. It would point at the new URL.

The minimal fix is to restore the old name and the table ARN reference (excerpts):

resource "aws_sqs_queue" "jobs" {
  name = "orders-queue" # original name restored; renaming forces replacement
}

resource "aws_iam_role_policy" "worker_ddb" {
  name = "ddb"
  role = aws_iam_role.worker.id
  policy = jsonencode({
    Version = "2012-10-17"
    Statement = [{
      Effect   = "Allow"
      Action   = ["dynamodb:PutItem", "dynamodb:UpdateItem"]
      Resource = aws_dynamodb_table.orders.arn
    }]
  })
}

The revised plan should read Plan: 0 to add, 2 to change, 0 to destroy. The two changes are the policy gaining UpdateItem and the Lambda gaining QUEUE_URL, now a known value because the queue is untouched.

Harder variation: producers hardcode the queue URL

Suppose the rename is mandatory, but the services that send messages hardcode the old URL and cannot change this week. Decide which options work.

Option Works?
create_before_destroy Check below
Keep the name Check below
moved block Check below
Check your answer
  • create_before_destroy: does not work. It only reverses the order, so the new queue is created before the old one is destroyed. The new queue has a different URL, and the old URL still dies at the end of the apply. Producers fail anyway. With the same name it would not give you a second queue either: SQS allows one queue per name, so the create fails or hands back the old queue, which the destroy step then deletes.
  • Keep the name: works. The URL contains the queue name, so keeping name = "orders-queue" means no replacement and an unchanged URL. Rename later, when producers can move.
  • moved block: does not work. It changes only the Terraform address in state (for example aws_sqs_queue.jobs to aws_sqs_queue.orders). It never changes the real queue's name, so the name edit would still force replacement. It is useful only if you rename the resource label while keeping the name.

The constraint is on the real queue name, so only leaving the name alone satisfies it.

Closing checklist

You have finished this lesson when you can do each of these without looking:

  • Run aws sso login, then init, plan and apply from your profile, with a provider block that holds only region.
  • Read +, ~, -, -/+ and +/-, and the summary line, and tell an in-place edit from a replacement.
  • Find the argument behind # forces replacement and decide: revert, moved block, or migrate.
  • Classify drift as accept, revert, or ignore, and know when -refresh-only applies.
  • Gate a plan before apply: saved plan, destroy count, replacements, unexpected addresses, unknown values, and a verdict you can justify.