Terraform Core
How Terraform works and the language you write it in: configuration, state, plan and apply; resources, variables, locals, outputs and expressions; providers and authentication to AWS.
SPACED REPETITION Β· 15 practice questions
Make this lesson stick.
Try 3 questions now. No account needed. Sample answers aren't saved.
or sign in to practice all 15New to Terraform? This lesson assumes you have already run
terraform init,planandapplyon a small configuration. If you have not, start with the lesson Start Here: What Terraform Is in this roadmap.
Resource Identity and State: What a Plan Is Really Comparing
You rename one line in a Terraform file β the name of a DynamoDB table β and Terraform announces it will destroy your table. Nothing about the change looked dangerous. The table still exists, the data is still there, you never asked for a deletion. The explanation is that terraform plan (the proposed list of actions Terraform would take if you ran terraform apply) is not a diff between your file and your AWS account. It is a diff between your file and Terraform's own record of what it created last time, refreshed with a read of live AWS attributes but keyed by that record.
Trace how a configuration block becomes a state entry and back into a plan, and you can look at any plan before applying and say which lines create, update, replace, or destroy β and where the data loss is hiding.
The three inputs to a plan
State is Terraform's record of what it manages: which configuration block maps to which real AWS object, plus the attribute values it last read. By default it lives in a local terraform.tfstate file; in a team setup it lives in a remote backend, Terraform's word for wherever state is stored (on AWS usually an S3 bucket). That is a different thing from the application backend this course builds.
A resource instance address is the key inside that record β for example aws_dynamodb_table.orders or aws_lambda_function.consumer (resource type plus your local block name), and module.messaging.aws_sqs_queue.orders once the block lives inside a module. A resource block is the HCL declaration itself; its attributes are the fields on it (name, billing_mode).
1. Configuration (HCL): what you want
|
v
2. State file: what Terraform recorded last time
|
v
3. Refresh: read current attribute values from AWS
|
v
4. Diff: desired vs recorded + refreshed
|
v
5. Plan symbols: + ~ -/+ - <=
The middle step is the refresh: before planning, Terraform calls AWS to update the stored attribute values, so out-of-band changes show up as drift (a difference between state and reality). What refresh does not do is invent identity. If an object exists in AWS but has no address in state, Terraform plans to create it β because the state-as-source-of-truth model has never heard of it. That gap is exactly what import closes later in this lesson.
The provider β here the hashicorp/aws plugin that knows how to call AWS APIs β supplies the schema that tells Terraform which attributes are writable and which are ForceNew (short for "a change forces a new resource": the attribute cannot be changed on the live object, so it must be destroyed and recreated).
Worked trace: one table, one rename
resource "aws_dynamodb_table" "orders" {
name = "orders"
billing_mode = "PAY_PER_REQUEST"
hash_key = "id"
attribute {
name = "id"
type = "S"
}
}
After the first terraform apply (the command that executes the plan against AWS), state holds something like what terraform state show aws_dynamodb_table.orders prints:
# aws_dynamodb_table.orders:
resource "aws_dynamodb_table" "orders" {
arn = "arn:aws:dynamodb:eu-west-1:111122223333:table/orders"
billing_mode = "PAY_PER_REQUEST"
hash_key = "id"
id = "orders"
name = "orders"
}
The address is the key; id and arn are the real AWS identifiers and derived attributes.
Now change one line: name = "orders-v2". The plan you should expect:
# aws_dynamodb_table.orders must be replaced
-/+ resource "aws_dynamodb_table" "orders" {
~ arn = "arn:aws:dynamodb:eu-west-1:111122223333:table/orders" -> (known after apply)
~ id = "orders" -> (known after apply)
~ name = "orders" -> "orders-v2" # forces replacement
# (other attributes unchanged)
}
Plan: 1 to add, 0 to change, 1 to destroy.
name is ForceNew because the DynamoDB API has no rename operation for a table. Terraform therefore plans -/+, which means destroy then create. The data goes with the old table. Reverting the name, or planning a data migration, are the only ways to keep it β a lifecycle rule can block the apply but cannot make this configuration safe, as we will see in 'Lifecycle Rules: Controlling Forced Replacement and Drift'.
Reading plan symbols before you read anything else
| Symbol | Meaning | Typical cause |
|---|---|---|
+ |
create | new address in config, none in state |
~ |
update in place | writable attribute changed |
-/+ |
replace (destroy, then create) | ForceNew attribute changed |
- |
destroy | address in state, gone from config |
<= |
read data source | a read-only lookup whose read had to wait until apply; nothing is changed |
A data source is a read-only lookup, e.g. data "aws_caller_identity" "current"; Terraform normally reads it during planning, so it does not appear among the plan's actions; it shows as <= only when its read has to wait until apply (for example, because an argument depends on a resource that is still changing). Either way it never modifies anything.
The symbol alone does not tell you whether to panic. -/+ aws_iam_policy.consumer is usually low-risk: an IAM policy is a stateless definition, so recreating it costs a few seconds (during which whatever relies on it lacks those permissions). -/+ on a stateful resource β DynamoDB tables, SQS queues holding messages, EFS filesystems β is a data-loss signal, because the contents are destroyed with the container. (A replaced SSM parameter is re-created from its configured value; what is lost is its version history and any value written outside Terraform.) That distinction, not the symbol count, is what drives your next move.
Practice: predicting a plan for a queue and its DLQ
State holds two SQS queues: aws_sqs_queue.orders with name = "orders", visibility_timeout_seconds = 30, message_retention_seconds = 345600 (AWS's 4-day default), and aws_sqs_queue.orders_dlq with name = "orders-dlq", message_retention_seconds = 345600.
You make three edits: orders gets name = "orders-prod" and visibility_timeout_seconds = 60; orders_dlq gets message_retention_seconds = 1209600 (the 14-day maximum).
Task: write the plan lines you expect for both resources and classify each as in-place update, replacement, or destroy. Then say which one threatens data.
Check your answer
-/+ aws_sqs_queue.orders # forces replacement
name: "orders" -> "orders-prod" # forces replacement
visibility_timeout_seconds: 30 -> 60
~ aws_sqs_queue.orders_dlq
message_retention_seconds: 345600 -> 1209600
Plan: 1 to add, 1 to change, 1 to destroy.
(Shorthand: a real plan prints each block as # aws_sqs_queue.orders must be replaced / -/+ resource "aws_sqs_queue" "orders" {, with attributes as ~ name = "orders" -> "orders-prod", and lists resources in address order, as here.)
Why: message_retention_seconds and visibility_timeout_seconds are writable through the queue attribute API, so they plan as ~. name is ForceNew on aws_sqs_queue, so the address aws_sqs_queue.orders plans as -/+ and the changed name is the attribute that caused it. The replacement is the dangerous line: every unconsumed message in orders is deleted with the queue, and anything referencing its ARN (a Lambda event source mapping, for instance) is rebuilt against the new queue.
Harder variation. The plan shows two separate lines: - aws_sqs_queue.orders and + aws_sqs_queue.main_queue, and the renamed block still has name = "orders". Same queue, same data β but you renamed the block. Terraform sees two unrelated addresses: one in state with no config, one in config with no state. Since they are independent, ordering is not guaranteed. If the delete runs first, the new queue comes up empty (SQS also makes you wait 60 seconds before reusing a deleted queue's name). If the create runs first, SQS can hand back the existing queue (CreateQueue does that when the name and attributes match), and the delete then removes the very queue the new address points at. The safest reading is "this is an address change, not a resource change," and the repair is a moved block covered in 'Safe State Surgery: Import, Moved, and Removed Blocks'.
The pitfall that burns people
Renaming a block, or moving it into a module (changing aws_sqs_queue.orders to module.messaging.aws_sqs_queue.orders), changes the state address while the real AWS resource is untouched. Terraform consequently plans - for the old address and + for the new one, even though the queue is right there. This is not a bug; it is the identity model working exactly as specified. It is also the single most common way a tidy refactor becomes an outage.
Cost analysis: scan for -/+ and - first
Before reading a single attribute diff, scan the plan for -/+ and - and ask what each one is attached to. Replacement of a stateful AWS resource costs you three things at once: the data, the downtime while the replacement is created, and any dependent resources that must be recreated because they referenced the destroyed object β dependency order is the subject of the next section, 'The Dependency Graph: Ordering Creates, Updates, and Destroys'.
Checklist for this section: every plan line maps to a state address you can name; every -/+ and - is explained by a specific attribute or address; and you have decided, before applying, whether that action is acceptable. If you cannot explain a -/+ yet, do not apply.
The Dependency Graph: Ordering Creates, Updates, and Destroys
A fresh stack plans "8 to add", you approve, and the apply sits on the event source mapping (Still creating...) while Lambda keeps rejecting it, because the function's role cannot read the queue yet. Nothing in the diff looks wrong. The cause is invisible in the plan text itself: Terraform builds a graph of what must exist before what, and one genuine dependency never made it into that graph.
Nodes, edges, and what a reference creates
Terraform turns configuration into a directed graph β nodes joined by one-way arrows that mean "this must exist first." Each resource block becomes a node (expanded into one node per instance when it uses count or for_each). A data block, a data source (a read-only lookup of something already in AWS), is also a node. Every node has a resource address, the label Terraform uses in plans and state, like aws_sqs_queue.orders.
The arrows are edges. Terraform derives most of them from references β any expression that reads another node's attribute:
resource "aws_lambda_function" "consumer" {
function_name = "orders-consumer"
role = aws_iam_role.lambda_exec.arn
# runtime, handler and the deployment package omitted here
}
aws_iam_role.lambda_exec.arn reads the role's ARN (Amazon Resource Name, the identifier that names an AWS resource unambiguously; most resource types have one). That read creates the edge aws_iam_role.lambda_exec β aws_lambda_function.consumer. Meta-arguments β special arguments Terraform accepts on any resource (and mostly on module blocks), such as depends_on, for_each and lifecycle β can add edges too: depends_on declares one directly, and a reference inside for_each or count creates one like any other reference. Module boundaries connect nodes across scopes. Terraform walks the graph running independent nodes concurrently up to -parallelism (default 10), and waiting wherever an edge points in.
Worked example: the orders pipeline
resource "aws_sqs_queue" "orders_dlq" {
name = "orders-dlq"
}
resource "aws_sqs_queue" "orders" {
name = "orders"
redrive_policy = jsonencode({
deadLetterTargetArn = aws_sqs_queue.orders_dlq.arn
maxReceiveCount = 5
})
}
resource "aws_iam_role" "lambda_exec" {
name = "orders-consumer-role"
assume_role_policy = jsonencode({
Version = "2012-10-17"
Statement = [{
Effect = "Allow"
Action = "sts:AssumeRole"
Principal = { Service = "lambda.amazonaws.com" }
}]
})
}
resource "aws_iam_role_policy" "consume" {
name = "consume-orders"
role = aws_iam_role.lambda_exec.id
policy = jsonencode({
Version = "2012-10-17"
Statement = [{
Effect = "Allow"
Action = ["sqs:ReceiveMessage", "sqs:DeleteMessage", "sqs:GetQueueAttributes"]
Resource = "arn:aws:sqs:us-east-1:123456789012:orders"
}]
})
}
resource "aws_lambda_function" "consumer" {
function_name = "orders-consumer"
role = aws_iam_role.lambda_exec.arn
runtime = "nodejs24.x"
handler = "index.handler"
filename = "consumer.zip"
source_code_hash = filebase64sha256("consumer.zip")
}
resource "aws_lambda_event_source_mapping" "orders" {
event_source_arn = aws_sqs_queue.orders.arn
function_name = aws_lambda_function.consumer.arn
batch_size = 10
}
resource "aws_cloudwatch_event_rule" "tick" {
name = "orders-tick"
schedule_expression = "rate(5 minutes)"
}
resource "aws_cloudwatch_event_target" "to_orders" {
rule = aws_cloudwatch_event_rule.tick.name
arn = aws_sqs_queue.orders.arn
}
Quick glosses: the dead-letter queue (DLQ) is a second queue that receives messages the consumer fails to process repeatedly; the event source mapping is the AWS resource that tells Lambda to poll a queue and invoke the function; an IAM role is the identity the function assumes, and an IAM policy is a JSON document listing allowed actions on named resources. To keep the graph small, the example leaves out the aws_sqs_queue_policy a working EventBridge-to-SQS target also needs: the queue must allow events.amazonaws.com to call sqs:SendMessage.
Read the references and the graph layers itself:
Layer 0 (all parallel β no unmet dependency):
aws_iam_role.lambda_exec
aws_sqs_queue.orders_dlq
aws_cloudwatch_event_rule.tick
Layer 1:
aws_iam_role_policy.consume <- role.id
aws_sqs_queue.orders <- orders_dlq.arn
aws_lambda_function.consumer <- role.arn
Layer 2:
aws_cloudwatch_event_target.to_orders <- rule.name + orders.arn
aws_lambda_event_source_mapping.orders <- orders.arn + consumer.arn
Notice what is missing: neither aws_lambda_function.consumer nor the event source mapping waits for aws_iam_role_policy.consume. The role argument reads only the role's ARN, which exists the moment the role does. Lambda checks the role's SQS permissions when the mapping is created, so if the policy is still in flight it rejects the mapping with an InvalidParameterValueException saying the execution role does not have permissions. The AWS provider retries that error for up to five minutes, which usually turns the race into a slow apply instead of a failure β but that is luck, not design. Write the ordering explicitly on the mapping: depends_on = [aws_iam_role_policy.consume].
Hidden dependencies: the edge inside a string
Look again at the policy's Resource. It is a literal ARN string, not aws_sqs_queue.orders.arn, so Terraform sees no edge from the queue to the policy. This is the trap: "${aws_sqs_queue.orders.arn}" inside an interpolation does create the edge, but a pasted literal does not.
IAM itself accepts a policy naming a queue that does not exist yet, so a missing edge here often applies "successfully" and bites later. Other resources turn the same gap into a race. If the EventBridge target above used a literal ARN instead of aws_sqs_queue.orders.arn, nothing would stop Terraform from creating the target before the queue exists. The PutTargets API reference does not say it checks that the queue exists, so the symptom can be a failed apply or a rule whose first events are not delivered β timing-dependent, not deterministic.
Prefer the reference. When you genuinely must keep a literal (a cross-account ARN, a value from a local), add the edge by hand: depends_on = [aws_sqs_queue.orders].
Cycles, and how to break them
Two Lambdas that reference each other's names produce a graph with no valid start:
# lambda A env: PEER = aws_lambda_function.b.function_name
# lambda B env: PEER = aws_lambda_function.a.function_name
Terraform stops before planning (terraform validate catches it too) with Error: Cycle: followed by the nodes in the loop, here aws_lambda_function.a and aws_lambda_function.b. Break it by deleting one direct edge in favour of a value that does not point back at the graph β an input variable, an SSM Parameter Store entry (AWS's managed config store), or a data source. For example, let a read b's name through an SSM parameter while b takes a's name from var.peer_a_name. Edges now run one way.
What the graph costs you
Your longest dependency chain sets the floor on apply wall time; adding parallel-friendly nodes costs nothing but widening. Two consequences: needless depends_on edges serialize work that could run together, and missing edges cause retries or outright failures. On destroy, Terraform walks the same graph in reverse, which is why a tidy graph also makes terraform destroy predictable.
β οΈ for_each keys must be known before apply. Keying a map off a resource attribute fails with Invalid for_each argument ... keys derived from resource attributes that cannot be determined until apply. Use stable literals:
resource "aws_lambda_event_source_mapping" "queues" {
for_each = toset(["orders", "returns"])
event_source_arn = aws_sqs_queue.main[each.key].arn
function_name = aws_lambda_function.consumer.arn
}
Practice: find the missing edge
resource "aws_sqs_queue" "orders" {
name = "orders"
}
resource "aws_iam_role_policy" "producer" {
role = aws_iam_role.producer.id
policy = jsonencode({
Statement = [{
Effect = "Allow"
Action = ["sqs:SendMessage"]
Resource = "arn:aws:sqs:us-east-1:123456789012:orders"
}]
})
}
resource "aws_cloudwatch_event_target" "to_orders" {
rule = aws_cloudwatch_event_rule.tick.name
arn = "arn:aws:sqs:us-east-1:123456789012:orders"
}
- Mark the missing edges in the graph.
- Predict the failure symptom on a first apply (assume
aws_cloudwatch_event_rule.tickalready exists). - Write the corrected target block.
- State the resulting creation order for the queue and the target.
Check your answer
Missing edges: aws_sqs_queue.orders β aws_iam_role_policy.producer and aws_sqs_queue.orders β aws_cloudwatch_event_target.to_orders. Both ARNs are literals, so neither node has a dependency on the queue.
Symptom: a race on the event target. Nothing orders to_orders after the queue, so Terraform may create the target while the queue does not exist yet. EventBridge's PutTargets reference does not document checking the queue, so expect either a failed apply or early events that are never delivered, with a result that can change from run to run. The IAM policy usually applies without complaint, because IAM does not check that the queue a policy names exists β that edge is defensive, not the immediate cause.
Corrected target:
resource "aws_cloudwatch_event_target" "to_orders" {
rule = aws_cloudwatch_event_rule.tick.name
arn = aws_sqs_queue.orders.arn # a real reference replaces the literal
# depends_on = [aws_sqs_queue.orders] # only if a literal ARN is unavoidable
}
Order: aws_sqs_queue.orders completes first, then aws_cloudwatch_event_target.to_orders. Add depends_on = [aws_sqs_queue.orders] to the producer policy as well if you want the graph to state that intent explicitly. Because the target now reads orders.arn, no depends_on is needed there.
One limit of this mental model: it captures creation and destruction ordering, but not every real-world sequence. Provisioners, IAM's eventual consistency, and out-of-band changes live outside the graph β the last one is drift, which the lifecycle section handles with ignore_changes.
Lifecycle Rules: Controlling Forced Replacement and Drift
A plan line like -/+ aws_dynamodb_table.orders means Terraform intends to delete the table and build a new one: every order in it gone. You already know how to read that symbol and how the dependency graph decides when each action runs. What you have not yet seen are the controls that let you steer those actions without hand-editing state: the lifecycle block. The block is a meta-argument β a setting on a resource block that changes how Terraform treats the resource rather than describing a real AWS property β and its entries are lifecycle rules.
create_before_destroy: build the new one first
By default a replacement is destroy-then-create. For an ECS task definition that means the old revision is deregistered (marked INACTIVE) before its successor is registered. The service keeps running on an INACTIVE revision, but you cannot run new tasks from it or point a service at it, so for a moment neither the old nor the new revision is usable (other ACTIVE revisions of the family, if any, still are).
default replacement order:
old resource
| destroy
v
(nothing)
| create
v
new resource
create_before_destroy = true:
old resource
| create (both exist briefly)
v
new + old
| destroy old
v
new resource
resource "aws_ecs_task_definition" "api" {
family = "orders-api"
requires_compatibilities = ["FARGATE"]
network_mode = "awsvpc"
cpu = "512"
memory = "1024"
execution_role_arn = aws_iam_role.task_execution.arn
container_definitions = jsonencode([
{
name = "api"
image = "123456789012.dkr.ecr.eu-west-1.amazonaws.com/orders-api:1.4.2"
essential = true
}
])
lifecycle {
create_before_destroy = true
}
}
The new revision registers before the old one is deregistered, so a usable revision exists at every moment. In the plan this replacement shows as +/- (create replacement, then destroy) instead of -/+.
β οΈ Two ways this bites:
- Unique names. ECS service names are unique per cluster, IAM role names are account-unique, DynamoDB table names are region-unique. Put
create_before_destroyonaws_ecs_serviceand Terraform tries to create a second service calledorders-apiwhile the first still exists β the API rejects it with a name conflict and the apply fails. Use it where the replacement gets a fresh identity, not where it reuses a unique name. - Hard limits and duplicate capacity. Old and new coexist for a moment, so anything governed by a quota or a fixed-size pool can fail or double.
prevent_destroy: the seatbelt on stateful resources
resource "aws_dynamodb_table" "orders" {
name = "orders"
billing_mode = "PAY_PER_REQUEST"
hash_key = "orderId"
attribute {
name = "orderId"
type = "S"
}
lifecycle {
prevent_destroy = true
}
}
Any plan that would destroy the table β including terraform destroy β now fails with Instance cannot be destroyed, and the apply stops. Apply the same guard to an SQS queue holding messages, an SSM Parameter Store value, or any table you cannot rebuild.
β οΈ Three limits:
- It protects; it does not fix. If
namechanged, the plan still shows-/+; you simply cannot apply it. The configuration is still wrong. - It lasts only as long as the block does. Delete the resource block (or just the
lifecycleblock) and nothing objects β Terraform plans a destroy without complaint. - It is a plan-time check read from configuration, not stored in state. It cannot protect a resource Terraform no longer knows about.
ignore_changes: stop fighting the scaler
ECS service autoscaling writes desired_count (how many tasks the service should keep running). Your configuration says 2; the scaler pushed it to 7; the next plan shows ~ desired_count = 7 -> 2. Terraform scales back down, the scaler scales up, and the two fight forever.
resource "aws_ecs_service" "api" {
name = "orders-api"
cluster = aws_ecs_cluster.main.id
task_definition = aws_ecs_task_definition.api.arn
desired_count = 2
launch_type = "FARGATE"
# network_configuration { ... } omitted here; Fargate (awsvpc) services need it
lifecycle {
ignore_changes = [desired_count]
}
}
Terraform now leaves that one attribute alone and still tracks everything else on the service. Names inside the list are the block's own argument names, so it is desired_count, not the AWS API's desiredCount. To ignore drift in a single tag key, index the map: ignore_changes = [tags["LastModified"]]. The bare ignore_changes = [tags] ignores all tag drift.
β οΈ Do not ignore security-bearing arguments. ignore_changes = [policy] on an IAM policy or [network_configuration] on an ECS service means that when someone loosens a policy or opens a group out of band, the plan stays clean and you never see it. Ignore operational noise β counts, timestamps, scaler-owned values β not security posture.
β οΈ ignore_changes is permanent until you remove it. While it is in place Terraform will never reconcile that argument, even if you later change it deliberately in your configuration.
replace_triggered_by: replace when a referenced value moves
Some resources are effectively immutable β an ECS task definition revision is never edited, only replaced β and must be rebuilt when a value they consume changes, even though that value is not one of their arguments. A Lambda function is not one of them: its code and settings update in place, so put this on a function only when you really want a new one. replace_triggered_by (Terraform 1.2+) watches a reference and forces replacement when it changes.
resource "aws_ssm_parameter" "config" {
name = "/orders/api/config-version"
type = "String"
value = "2026-04-11"
}
resource "aws_ecs_task_definition" "api" {
family = "orders-api"
# ... container_definitions read /orders/api/config-version at runtime ...
lifecycle {
replace_triggered_by = [aws_ssm_parameter.config]
}
}
Referencing the whole resource triggers on any change to it. Referencing aws_ssm_parameter.config.version narrows that to version bumps only. Choose deliberately: the broad form is easy to reason about, the narrow form avoids surprise rebuilds.
β οΈ It does not cascade. It replaces this resource and nothing else; a dependent changes only if it references the value that changed β for example a service that reads the new task definition ARN. Nothing downstream is rebuilt for free.
precondition and postcondition: refuse to apply nonsense
Both sit inside lifecycle (Terraform 1.2+) and both fail the run with your own message when a condition is false.
resource "aws_lambda_function" "consumer" {
function_name = "orders-consumer"
runtime = var.lambda_runtime
handler = "handler.main"
role = aws_iam_role.consumer.arn
filename = "build/consumer.zip"
lifecycle {
precondition {
condition = contains(["python3.12", "python3.13"], var.lambda_runtime)
error_message = "lambda_runtime must be python3.12 or python3.13, got ${var.lambda_runtime}."
}
postcondition {
condition = can(regex("^arn:aws:lambda:eu-west-1:", self.arn))
error_message = "consumer Lambda was created outside eu-west-1."
}
}
}
A precondition runs before Terraform does work on the resource: a bad runtime never gets created. A postcondition runs afterwards and reads the resource's own attributes through self. β οΈ Postcondition failure does not roll the resource back β it exists in AWS and in state β but Terraform halts the rest of the apply and reports the failure. These are tripwires, not transactions; Terraform has no automatic rollback.
Worked decision: the DynamoDB -/+
# aws_dynamodb_table.orders must be replaced
-/+ resource "aws_dynamodb_table" "orders" {
~ arn = "arn:aws:dynamodb:eu-west-1:...:table/orders" -> (known after apply)
~ name = "orders" -> "orders-v2" # forces replacement
# (11 unchanged attributes hidden)
}
Plan: 1 to add, 0 to change, 1 to destroy.
Work the decision in this order:
- Address change or attribute change? If only the block was renamed or moved into a module, the plan would show no
-/+at all but a-for the old address (with# (because ... is not in configuration)) and a+for the new one, and the fix is amovedblock β see "Safe State Surgery: Import, Moved, and Removed Blocks". Herenamegenuinely differs, so this is not that case. - Is the new value intentional? If not, revert the configuration. That is the fix, and the plan returns to
No changes. - Intentional? Replacement is unavoidable. Add
prevent_destroy = trueto block the apply while you plan a migration (new table, copy, cut over), then remove it and migrate deliberately.
β οΈ prevent_destroy here yields Instance cannot be destroyed and stops the apply β a seatbelt, not a repair. Your configuration still disagrees with reality.
Practice: three plan diffs, choose the control
~ environment { "LOG_LEVEL" = "info" -> "debug" }on a Lambda function.-/+ resource "aws_dynamodb_table" "orders"with~ name = "orders" -> "orders-v2" # forces replacement.~ desired_count = 7 -> 2on an ECS service, where autoscaling set it to 7.
For each: name the action, write the lifecycle rule if one applies, or explain why none does. Then state what the next plan shows.
Check your answer
- In-place update (
~). No lifecycle rule applies β Terraform should apply it. Expected plan:Plan: 0 to add, 1 to change, 0 to destroy. - Replacement (
-/+) of a stateful resource. No lifecycle rule makes this safe: revertnameif the change was accidental (expected plan:No changes.), usemovedif only the address changed, otherwise accept replacement only alongside a data migration.lifecycle { prevent_destroy = true }optionally blocks the apply until that migration exists β it does not make the configuration correct. - Drift Terraform wants to revert. Add
lifecycle { ignore_changes = [desired_count] }to the service. Expected plan afterwards, assuming nothing else drifted:No changes.βdesired_countno longer appears.
Harder variation: pin a Lambda to a parameter version
resource "aws_lambda_function" "consumer" {
function_name = "orders-consumer"
runtime = "python3.12"
handler = "handler.main"
role = aws_iam_role.consumer.arn
filename = "build/consumer.zip"
lifecycle {
replace_triggered_by = [aws_ssm_parameter.config.version]
postcondition {
condition = can(regex("^arn:aws:lambda:eu-west-1:", self.arn))
error_message = "consumer Lambda was created outside eu-west-1."
}
}
}
Bumping the SSM parameter's version now replaces the function (destroy, then create, so for a moment it does not exist), while .version keeps unrelated edits to the parameter from triggering a redeploy. The postcondition catches the case where a provider alias accidentally points the function at the wrong region β a mistake that silently breaks every ARN-based permission you granted elsewhere.
Checklist before you apply
- Every
-/+and-line is explained: address change, immutable attribute, or genuine removal? - No stateful resource (DynamoDB, SQS, or an SSM parameter whose value is written outside Terraform) shows
-/+or-. - Scaler-owned or externally-written attributes that should not be reverted have
ignore_changes. ignore_changesdoes not coverpolicy,network_configuration, or any other security-bearing argument.- Any
precondition/postconditionyou rely on has anerror_messagethat tells the next reader what to fix.
Safe State Surgery: Import, Moved, and Removed Blocks
You rename a Lambda block, tidy your files, run terraform plan β and Terraform announces it will destroy aws_lambda_function.consumer and create module.consumer.aws_lambda_function.this. Same function, same orders-consumer name in AWS, one dead-end plan. Terraform compares addresses, not vibes, which is exactly the mechanism Section 1 traced. This section is about the three operations that change an address, adopt an outsider, or retire a resource, each of which can work without touching the real AWS object: moved, import, and removed.
moved: change the address, not the resource
A moved block is a declaration that a resource now lives at a different configuration address than the one recorded in state. It was introduced in Terraform 1.1 and is the safe way to rename a resource or extract it into a module.
moved {
from = aws_lambda_function.consumer
to = module.consumer.aws_lambda_function.this
}
Before the block, state and configuration disagree about the address:
config says: module.consumer.aws_lambda_function.this
state says: aws_lambda_function.consumer -> id=orders-consumer
Terraform reads that as "destroy one function, create another," and the plan says so. The second comment line is the clue that the block left the configuration rather than the function changing:
# aws_lambda_function.consumer will be destroyed
# (because aws_lambda_function.consumer is not in configuration)
# module.consumer.aws_lambda_function.this will be created
Plan: 1 to add, 0 to change, 1 to destroy.
With the moved block in place, Terraform moves the state entry to the new address before planning, refreshes the same function there, and the plan records the move (the new address is saved to state when you apply):
module.consumer.aws_lambda_function.this: Refreshing state... [id=orders-consumer]
# aws_lambda_function.consumer has moved to module.consumer.aws_lambda_function.this
resource "aws_lambda_function" "this" {
id = "orders-consumer"
# (unchanged attributes hidden)
}
Plan: 0 to add, 0 to change, 0 to destroy.
Keep the block in configuration until every workspace that owns a copy of that state has applied it once. After that, the address is recorded and the block can be deleted β but deleting it early in a workspace that hasn't applied yet reintroduces the destroy/create plan. If from has a typo, Terraform finds nothing at that address and silently ignores the block: the plan still shows the destroy and the create, with no warning. Look for the has moved to line before you apply.
import: adopt what already exists
An import block (Terraform 1.5+) adopts a resource that exists in AWS but not in your state, declaratively rather than as a one-off command. You write the block and a matching resource block, run plan, and Terraform reads the real object.
resource "aws_lambda_function" "consumer" {
function_name = "orders-consumer"
role = aws_iam_role.consumer.arn
handler = "app.handler"
runtime = "nodejs22.x" # must match the deployed function
filename = "build/consumer.zip"
source_code_hash = filebase64sha256("build/consumer.zip")
}
import {
to = aws_lambda_function.consumer
id = "orders-consumer" # Lambda's import ID is the function name
}
The import ID is the provider-specific string Terraform uses to find the real object; it is not always the same thing as the resource's name attribute:
| Resource | Import ID |
|---|---|
aws_lambda_function |
function name |
aws_dynamodb_table |
table name |
aws_sqs_queue |
full queue URL |
β οΈ Write the block and run plan before apply. The first plan prints the import attempt and then the diff between the real object and your configuration:
aws_lambda_function.consumer: Preparing import... [id=orders-consumer]
aws_lambda_function.consumer: Refreshing state... [id=orders-consumer]
Plan: 1 to import, 0 to add, 1 to change, 0 to destroy.
That 1 to change is expected for a function deployed from a zip: AWS cannot return your local filename or your source_code_hash, so the provider imports them empty and the first plan sets both in place. Applying it uploads your local zip as the function's code, so build that zip from the source that is actually deployed. For most resources the goal is 0 to change, adoption plus a no-op, and any other ~ here deserves the same suspicion. If runtime, role, or memory_size differ from the deployed function, you get ~ in-place updates and should correct the config rather than apply a surprise change. If function_name differs from the ID, you get -/+ β function_name is ForceNew for this resource. Terraform still requires a resource block to import into; when you don't have one, terraform plan -generate-config-out=generated.tf writes a draft for you (shipped alongside import blocks in 1.5).
removed: stop managing without deleting
A removed block (Terraform 1.7+) says a resource is no longer managed by this configuration. Use it when handing a queue to another team's stack, or when deleting configuration while keeping the AWS object alive.
removed {
from = aws_sqs_queue.legacy
lifecycle {
destroy = false # forget it; do not delete it in AWS
}
}
β οΈ The destroy argument inside a removed block's lifecycle controls whether Terraform deletes the object, and its default is true β a bare removed block plans a - destroy. Always read the action line: you want # aws_sqs_queue.legacy will no longer be managed by Terraform, but will not be destroyed, not a destroy. The same lifecycle vocabulary appeared in Section 3, but here it is scoped to the removal itself, not to the resource's normal operation.
The imperative escape hatch: state mv, rm, list
Before these blocks existed, you did the same work by editing state directly, and the commands are still useful for one-off repairs. terraform state mv rewrites an address in state. terraform state rm removes an entry, leaving the resource unmanaged in AWS. terraform state list reads back the addresses currently recorded, which is how you verify either operation:
terraform state rm aws_lambda_function.old
terraform state list | grep lambda # confirm the address is gone
Two differences from the block forms matter. First, these commands are immediate β there is no plan to review. You can preview with -dry-run (both state mv and state rm accept it), and on local state state rm first writes a timestamped backup (terraform.tfstate.<timestamp>.backup), but nothing goes through the review a plan gets. Second, they leave no trace in configuration, so the next plan re-derives intent from your .tf files: if aws_lambda_function.old still has a resource block, Terraform plans to create it again, and the create fails because a function with that name already exists. State commands take the backend's lock when the backend has locking enabled, so a state command on a laptop cannot interleave with a CI run. The S3 backend locks only if you configure it: use_lockfile = true (added in Terraform 1.10, generally available in 1.11) or the older dynamodb_table setting, which is now deprecated. Authentication is unchanged: SSO on your laptop, OIDC-assumed role in CI, no static keys.
Practice
Task 1. Your aws_dynamodb_table.orders block was moved into module.data, and the plan now shows a destroy (-) for aws_dynamodb_table.orders plus a create for module.data.aws_dynamodb_table.orders. Write the moved block.
Check your answer
moved {
from = aws_dynamodb_table.orders
to = module.data.aws_dynamodb_table.orders
}
After applying once, the plan should report no changes for the table. If the table name itself also changed, moved is not enough β Section 3 covers the -/+ decision for a ForceNew attribute.
Task 2. An SQS queue named orders exists in AWS at https://sqs.eu-west-1.amazonaws.com/123456789012/orders, but no block manages it. Write the import block and name the first three plan lines you check before applying.
Check your answer
import {
to = aws_sqs_queue.orders
id = "https://sqs.eu-west-1.amazonaws.com/123456789012/orders"
}
Expected lines:
aws_sqs_queue.orders: Preparing import... [id=https://sqs.eu-west-1.amazonaws.com/123456789012/orders]
aws_sqs_queue.orders: Refreshing state... [id=https://sqs.eu-west-1.amazonaws.com/123456789012/orders]
Plan: 1 to import, 0 to add, 0 to change, 0 to destroy.
Check the address matches your resource block, the ID is the full queue URL (not just orders), and the summary says 0 to change β any ~ or -/+ means a configuration attribute disagrees with the live queue.
Outcome checklist
movedfor address changes;importfor resources that already exist;removedfor stopped management.- You can state the import ID form for Lambda, DynamoDB, and SQS.
- You always read the plan's action line for a
removedblock before applying. - You use
terraform state listto verify any imperativestateoperation.
Independent Transfer: Diagnose and Repair a Drifting AWS Backend
You now hold every tool this lesson promised: you can read a plan, predict it from state plus configuration, control replacement with lifecycle, and repair addresses with moved/import/removed. This final section hands you a plan you did not write and asks you to do what a real reviewer does β explain every symbol before anyone types apply.
The plan as it lands
Since the last apply, the configuration was edited (a table name, a zip path, one block moved into a module) and an autoscaler changed the ECS service. terraform plan returns:
Terraform will perform the following actions:
# aws_dynamodb_table.orders must be replaced
-/+ resource "aws_dynamodb_table" "orders" {
~ name = "orders-v2" -> "orders" # forces replacement
}
# aws_ecs_service.api will be updated in-place
~ resource "aws_ecs_service" "api" {
~ desired_count = 4 -> 2
}
# aws_lambda_function.consumer will be updated in-place
~ resource "aws_lambda_function" "consumer" {
~ filename = "consumer.zip" -> "build/consumer.zip"
~ source_code_hash = "Ah8S..." -> "Bz9T..."
}
# aws_sqs_queue.orders_dlq will be destroyed
# (because aws_sqs_queue.orders_dlq is not in configuration)
- resource "aws_sqs_queue" "orders_dlq" {
- name = "orders-dlq" -> null
}
# module.queues.aws_sqs_queue.this will be created
+ resource "aws_sqs_queue" "this" {
+ name = "orders-dlq"
}
Plan: 2 to add, 2 to change, 2 to destroy.
Two things should stop you before anything else. The DynamoDB table is -/+: name is a ForceNew attribute β one that forces a destroy-and-recreate rather than an in-place update β and it is the table's identity, so replacement means an empty table. Then read the SQS pair carefully: a destroy at aws_sqs_queue.orders_dlq and a create at module.queues.aws_sqs_queue.this both carry name = "orders-dlq". That is the fingerprint of a resource that moved address without moving in AWS: one real queue, one state entry still at the old address, and a plan whose destroy deletes the queue holding your undelivered messages. The destroy and the create run in no guaranteed order, so you end up with an empty new queue or with no queue at all (see 'Harder variation' in Section 1); the messages are lost either way.
Classify every action first
Four labels cover this plan. Assign one to each line before choosing a fix, because the label picks the tool.
| Address | Symbol | Class | Trigger |
|---|---|---|---|
aws_dynamodb_table.orders |
-/+ |
Immutable attribute change | name |
aws_lambda_function.consumer |
~ |
In-place update (code deploy) | filename, source_code_hash |
aws_ecs_service.api |
~ |
Drift | desired_count |
...orders_dlq -> module...this |
- / + |
Address change | block moved to module |
Notice that no line here is a genuine config removal. Nothing was deliberately deleted from the configuration; every destroy is a side effect of something changeable. That difference β accidental destroy versus intentional handoff β is what decides between a repair and a removed block.
Root causes and the smallest safe fix
DynamoDB -/+ β configuration no longer matches state. State records name = "orders-v2"; the config now says "orders". One of those edits is the mistake. The smallest safe fix is to restore the name in configuration to orders-v2, leaving the live table untouched. If the rename is genuinely wanted, that is a data migration β new table, backfill, cutover β not a one-line change.
SQS -/+ β a missing moved block. The DLQ block was cut into module.queues. To Terraform, an address is an identity, so the old address looks deleted and the new one looks new. Add:
moved {
from = aws_sqs_queue.orders_dlq
to = module.queues.aws_sqs_queue.this
}
Now the same queue is tracked under its new address and both lines collapse to nothing. This is the safe-refactor pattern from "Safe State Surgery: Import, Moved, and Removed Blocks" applied to a module extraction.
Lambda ~ β a code deploy, not a replacement. filename points at a different zip and source_code_hash changed with it. Neither argument is ForceNew on aws_lambda_function: the provider uploads the new package to the existing function, which keeps its name, ARN and triggers. The only question is whether that zip is the build you mean to ship; revert the path if the move was incidental.
ECS ~ β drift owned by someone else. The autoscaler set desired_count to 4; configuration still says 2. Terraform wants to scale it back down, which fights the scaler on every plan. Hand ownership to the scaler:
resource "aws_ecs_service" "api" {
name = "api"
desired_count = 2
lifecycle {
ignore_changes = [desired_count]
}
}
Reserve ignore_changes for fields another system legitimately owns. Applying it to task_definition or security-group membership would silently hide real drift.
Repair in order, then verify
1 moved block -> address matches -> DLQ plans as no-op
|
v
2 ignore_changes -> drift stops -> ECS plans as no-op
|
v
3 revert config value -> ForceNew matches -> DynamoDB no-op
|
v
4 terraform plan -> only intended change remains
Address changes come first because they are pure bookkeeping β no AWS impact, instant to verify. Attribute and drift fixes follow. Then re-run terraform plan and read the summary: DynamoDB and every SQS queue must show no -/+ and no -. The Lambda code update may still show as ~; that is fine. ECS should not, once desired_count is ignored, unless something else drifted. Success is not "the plan is empty" β it is that you can explain every symbol that remains.
Three wrong answers, diagnosed
"DynamoDB replacement is unavoidable." No β the config value is wrong, not the resource. Reverting name removes the replacement. Accepting -/+ on a table you never intended to delete is how data disappears.
"Add prevent_destroy to be safe." It blocks the apply but leaves the bad configuration in place, so every future plan fails identically. It is a smoke alarm, not a repair.
"Import the queue to its new address." An import block adopts a resource Terraform does not track; here the queue is tracked, just under the old address. Importing leaves two state entries for one queue, and the old address still plans -: applying it deletes the queue you just imported.
Your turn: a fresh drift report
A staging stack returns this plan:
-/+ aws_dynamodb_table.sessions {
~ hash_key = "sessionId" -> "session_id" # forces replacement
}
-/+ aws_sqs_queue.web {
~ name = "web" -> "web-queue" # forces replacement
}
~ aws_lambda_function.web {
~ environment { "QUEUE_URL" = "..." -> "..." }
}
~ aws_ecs_service.web {
~ desired_count = 3 -> 1
}
Classify all four lines and give the smallest safe fix for each. Then answer the harder question: which of these must not be "fixed" with a moved block, and why? Finish by writing the plan summary you expect once your fixes are in.
Check your answer
All four are the same three classes as before β no address change appears this time.
aws_dynamodb_table.sessionsβ immutable attribute change.hash_keyis the partition key and ForceNew; replacement destroys every session. The plan's "old" valuesessionIdis the live truth, so revert the config.movedcannot help: it changes addresses, not attributes. If the new key is intentional, this becomes a migration project.aws_sqs_queue.webβ immutable attribute change.nameis ForceNew; revert config toweb.prevent_destroywould only block the apply.aws_lambda_function.webβ in-place environment update, safe by itself. Check the cause: ifQUEUE_URLchanged because of the queue rename, fixing line 2 may erase this line too.aws_ecs_service.webβ drift. If an autoscaler ownsdesired_count, addignore_changes = [desired_count]. If nothing else owns the count, decide which value is intended: keep the configuration at 1 and let the apply scale the service down, or set it to 3 to keep the live count. Decide by asking who is supposed to own the field.
The trap: lines 1 and 2 are not address changes, so moved would be a no-op fix β it would rewrite an address that never changed and leave both -/+ actions in the plan.
Expected summary after the two reverts and ignore_changes, if the Lambda line came from the queue rename: No changes. Your infrastructure matches the configuration. If you keep the configuration at 1 and apply the scale-down instead, expect 0 to add, 1 to change, 0 to destroy β a single in-place ECS update and nothing else.
Outcome checklist
- State addresses match configuration β no module or rename churn.
- Every
-/+and-is classified and intentional. - Lifecycle rules protect stateful resources;
ignore_changescovers only externally-owned fields. - State operations (
moved,import,removed) are used instead of recreating live resources. - You can name the trigger for every remaining plan symbol before applying.
The habit that carries over: open a plan by scanning for -/+ and - on DynamoDB tables, SQS queues and SSM parameters, then explain each before choosing a tool. Repair the address when the address is wrong, the configuration when the value is wrong, and the lifecycle when another system owns the field.