State and Modules
Where Terraform's memory lives and how configurations are packaged for reuse: remote state on S3 with locking, what state contains, modules and their interfaces, and repeating resources with for_each and count.
SPACED REPETITION Β· 15 practice questions
Make this lesson stick.
Try 3 questions now. No account needed. Sample answers aren't saved.
or sign in to practice all 15Reading State to Predict Plans: Addresses, Attributes, and Drift
You run terraform plan and it answers with three lines: one ~, one -/+, and Plan: 1 to add, 1 to change, 1 to destroy. Nobody asked for a destroy. Terraform is not being dramatic β it is doing arithmetic on two documents: your .tf configuration, and a file called terraform.tfstate. Learn to read both together and you can predict the plan before you run it.
Terraform's state is its private inventory β a JSON file recording which real AWS resources it manages, keyed by address (the unique name a resource has, like aws_lambda_function.processor), storing each resource's AWS ID and last-known attribute values. A plan is the diff between what your configuration asks for and what state says already exists. Drift is the gap that opens when real AWS stops matching state.
The state file is an inventory, not a config
Here is the configuration we will trace, a Lambda consuming an SQS queue:
# main.tf β root module
variable "memory_size" {
type = number
default = 128
}
resource "aws_sqs_queue" "dlq" {
name = "processor-dlq"
}
resource "aws_iam_role" "processor" {
name = "processor-role"
assume_role_policy = jsonencode({
Version = "2012-10-17"
Statement = [{
Effect = "Allow"
Action = "sts:AssumeRole"
Principal = { Service = "lambda.amazonaws.com" }
}]
})
}
resource "aws_lambda_function" "processor" {
function_name = "processor"
role = aws_iam_role.processor.arn
handler = "index.handler"
runtime = "nodejs22.x"
filename = "processor.zip"
memory_size = var.memory_size
}
After terraform apply, the state file holds one entry per managed resource. Stripped down, it looks like this:
{
"version": 4,
"terraform_version": "1.9.5",
"serial": 12,
"lineage": "b1f0c3a2-...",
"resources": [
{
"mode": "managed",
"type": "aws_lambda_function",
"name": "processor",
"instances": [
{
"attributes": {
"id": "processor",
"memory_size": 128,
"runtime": "nodejs22.x"
},
"dependencies": ["aws_iam_role.processor"]
}
]
}
]
}
Four header fields matter. version is the state format version (4 in current Terraform) β unrelated to your CLI version. terraform_version records which CLI last wrote the file. An older 1.x CLI can often still read it, but HashiCorp does not guarantee that, so pin one Terraform version for the team and CI. serial is a counter bumped on every write; terraform state push uses it to refuse a stale copy, one whose serial is lower than the destination's (unless you add -force). lineage is a UUID tied to this state's history, so a state file from another environment can't be silently swapped in.
Each resource entry is keyed by address and carries the AWS ID, the attributes Terraform last observed, and dependencies β the addresses this resource depended on at its last apply. Creation order comes from the references in configuration; this copy exists so that, if you later delete both blocks, Terraform can still destroy the Lambda before its IAM role. (One module call does not split the file into separate states; that packaging detail belongs to "Module Boundaries as State Namespaces and Interfaces".) The plan pipeline is:
.tf configuration + terraform.tfstate
β
refresh: read the real AWS objects
β
diff: desired vs. refreshed state
β
plan: create | update | replace | destroy
Worked trace: one variable, one plan line
Change the variable default and watch the arithmetic:
variable "memory_size" {
type = number
default = 256 # was 128
}
Terraform refreshes aws_lambda_function.processor, so the "current" value becomes whatever AWS actually reports. Then it diffs against configuration: same address, memory_size 128 β 256, everything else equal. Because memory_size is an updateable argument β one the AWS UpdateFunctionConfiguration API can change on a live function β the plan is an in-place update:
Terraform will perform the following actions:
# aws_lambda_function.processor will be updated in-place
~ resource "aws_lambda_function" "processor" {
~ memory_size = 128 -> 256
}
Plan: 0 to add, 1 to change, 0 to destroy.
Memorize the tilde: ~ means "same resource, different attribute." Contrast it with -/+, which means Terraform will destroy the object and build a new one.
Addresses decide match, replace, or destroy
An address has a grammar, and moving a resource between two positions in that grammar is a destroy/create even when the real object never changes.
| Address shape | Example |
|---|---|
| root resource | aws_lambda_function.processor |
| inside a module | module.messaging.aws_sqs_queue.dlq |
count index |
aws_sqs_queue.work[0] |
for_each key |
aws_sqs_queue.work["orders"] |
Terraform matches a configuration block to a state entry by address string, not by the queue URL or resource name inside it. So renaming a block:
resource "aws_sqs_queue" "dead_letter" { # was "dlq"
name = "processor-dlq"
}
produces two actions β destroy aws_sqs_queue.dlq, create aws_sqs_queue.dead_letter β because the old address has no configuration block and the new address has no state entry. The name argument is unchanged, and Terraform does not order the destroy against the create. Destroy first: the old queue and its messages are deleted, and the AWS provider retries through SQS's 60-second name-reuse block until it creates an empty queue. Create first: SQS returns the existing queue's URL when the attributes match (or an error when they differ), and the destroy then deletes that very queue, leaving state pointing at nothing. Either way the original queue and its messages are gone. The fix is a moved block, a declarative rename Terraform reads from configuration β see "State Operations for Safe Repair".
The other replacement trigger is an immutable argument, one the AWS API cannot change after creation. Rename a DynamoDB table or change an argument that only exists at creation time and the plan shows -/+. Attribute updates are cheap; address and immutable-argument changes are where stateful resources die.
Drift: when AWS is edited without you
Someone bumps the Lambda's memory to 512 in the console. Configuration says 128; state says 128. Run terraform plan and refresh reads AWS first, then the plan proposes to undo it: ~ memory_size = 512 -> 128, Plan: 0 to add, 1 to change, 0 to destroy. The 512 on the left is the refreshed value. Here a normal plan prints no separate drift notice: since Terraform 1.2 it lists "Objects have changed outside of Terraform" only for outside changes that feed into another proposed change, and nothing reads memory_size. terraform plan -refresh-only, below, lists it either way.
That is the crucial mental model: state is a cache, not the source of truth. Configuration is the source of truth for what should exist; real AWS is the source of truth for what does exist; state is the last-known reading, refreshed before every plan. If the console change was intentional, accept it by editing configuration β or add lifecycle { ignore_changes = [memory_size] } to stop managing that one attribute. If you only want to see drift, terraform plan -refresh-only reports what changed outside Terraform and produces a plan that touches state only, never resources. That flag cannot propose destroying anything, which makes it the safer first look.
The one rule: never hand-edit the state file
State is a schema-versioned JSON document, and text-editing it breaks in two specific ways. A malformed or schema-invalid file makes Terraform refuse to plan; a parse that drops a resource entry does not delete the AWS resource β it orphans it, leaving a live queue or table Terraform no longer knows about. And the id field is what Terraform hands to the AWS API during refresh, so forging it makes Terraform read a different real object and then plan to update β or replace β a stateful resource such as a DynamoDB table, taking the data with it. Use the terraform state subcommands and moved blocks instead; those are the next section's subject.
The second cost is confidentiality. State stores attribute values in plaintext, including anything sensitive your configuration writes: a random_password.result fed to a database, or a secret passed through aws_ssm_parameter. Marking an output sensitive = true only redacts terminal display; the value is still in the file. With an S3 backend, the bucket policy and IAM permissions on the state object are part of your security boundary β consistent with a no-static-keys stack, but only if read access is scoped as tightly as the roles that can assume into the account.
Practice: predict the plan
You have this abbreviated state:
{
"serial": 12,
"resources": [
{ "type": "aws_lambda_function", "name": "processor",
"instances": [{ "attributes": { "id": "processor", "memory_size": 128 } }] },
{ "type": "aws_sqs_queue", "name": "dlq",
"instances": [{ "attributes": { "id": "https://sqs.../processor-dlq", "name": "processor-dlq" } }] },
{ "type": "aws_dynamodb_table", "name": "sessions",
"instances": [{ "attributes": { "id": "sessions", "billing_mode": "PAY_PER_REQUEST" } }] }
]
}
Three configuration edits, made together:
var.memory_sizechanges from 128 to 256.- The block
aws_sqs_queue.dlqis renamed toaws_sqs_queue.dead_letter; itsname = "processor-dlq"is unchanged. - The
aws_dynamodb_table.sessionsblock gainstags = { Team = "payments" }.
For each, predict create, update in place, replace, or destroy, and justify it from the address and attribute rules above. Then state the total count line.
Check your answer
- Update in place (
~). Address unchanged, andmemory_sizeis updateable onaws_lambda_function. - Destroy + create (
-for the old address,+for the new one). The address changed, so Terraform will not match them. Becausenamedid not change, applying this deletes the existing queue and its messages whichever half runs first. - Update in place (
~).tagsis updateable onaws_dynamodb_table; adding one does not replace the table. Had you renamed the table instead, that would be-/+and the sessions data would be gone.
Totals: Plan: 1 to add, 2 to change, 1 to destroy.
Every plan line is now derivable before you run it: read the address first (match or replace?), then the attribute (updateable or immutable?), then ask whether refresh found something you never wrote.
State Operations for Safe Repair: Inspect, Move, Remove, Import, and Refactor
You rename aws_dynamodb_table.sessions so it lives inside a new module, run terraform plan, and read this:
# aws_dynamodb_table.sessions will be destroyed
# (because aws_dynamodb_table.sessions is not in configuration)
- resource "aws_dynamodb_table" "sessions" { ... }
# module.data.aws_dynamodb_table.sessions will be created
+ resource "aws_dynamodb_table" "sessions" { ... }
Plan: 1 to add, 0 to change, 1 to destroy.
The sessions table is live. Terraform is not confused about AWS β it is confused about bookkeeping. Nothing in AWS moved; only the address changed, the path Terraform uses to key one resource inside state (aws_dynamodb_table.sessions, module.data.aws_dynamodb_table.sessions). Building on the state model from "Reading State to Predict Plans", a plan is a diff of configuration against state, so an address present in one and missing from the other looks exactly like a delete plus a create. This section covers the four bookkeeping operations β inspect, move, forget, adopt β that let you repair the diff instead of the infrastructure.
Inspect before you touch: state list and state show
Both commands are read-only: they print what state already contains and change nothing in AWS or in state.
$ terraform state list
aws_dynamodb_table.sessions
aws_lambda_function.processor
aws_sqs_queue.dlq
module.messaging.aws_sqs_queue.this
$ terraform state show aws_dynamodb_table.sessions
# aws_dynamodb_table.sessions:
resource "aws_dynamodb_table" "sessions" {
...
billing_mode = "PAY_PER_REQUEST"
...
hash_key = "pk"
id = "sessions"
name = "sessions"
...
}
terraform state list answers "which addresses does Terraform believe it owns?" terraform state show <address> answers "what attributes did Terraform record for this one?" β and those recorded attributes are the left-hand side of the next diff.
β οΈ state show prints attribute values, and state stores them in plaintext, so its output can contain credentials, connection strings, or ARNs you would not paste into a ticket.
Make this a habit: before any plan that proposes destruction, run terraform state list and check whether the address that "will be destroyed" and the address that "will be created" are two names for the same live resource. If they are, you have an address problem, not an infrastructure problem.
Moving an address: terraform state mv
terraform state mv rewrites state so one address becomes another, without touching AWS. In the rename above the configuration already declares the new address, so one command fixes it β move the state entry to match:
terraform state mv aws_dynamodb_table.sessions module.data.aws_dynamodb_table.sessions
# Successfully moved 1 object(s).
Now state and configuration agree, and the plan reads:
No changes. Your infrastructure matches the configuration.
That no-op is the proof. Had you moved the state entry before editing the configuration, the plan would flip β it proposes destroying the new address and creating the old one. Both halves must agree.
β οΈ state mv needs the state lock (the mechanism that stops two writers, your laptop and CI, from editing state at once). If Terraform reports a lock error, someone may be mid-apply. Do not reach for terraform force-unlock until you have confirmed no apply is running β stealing a live apply's lock can corrupt state. Older S3 setups lock through a DynamoDB lock table; Terraform 1.10 added native S3 locking with use_lockfile = true; it became generally available in 1.11, which also deprecated the DynamoDB lock arguments.
Declarative moves: the moved block
state mv rewrites one state immediately and leaves no trace in code. A moved block records the same rename in code (added in Terraform 1.1):
moved {
from = aws_dynamodb_table.sessions
to = module.data.aws_dynamodb_table.sessions
}
With this block committed, terraform plan reports the move explicitly instead of a destroy/create pair:
# aws_dynamodb_table.sessions has moved to module.data.aws_dynamodb_table.sessions
resource "aws_dynamodb_table" "sessions" { ... }
Plan: 0 to add, 0 to change, 0 to destroy.
| Aspect | state mv |
moved block |
|---|---|---|
| Runs in | local CLI | plan and apply |
| Applies to | the one state you ran it on | every state the code is applied to |
| Shows in plan | no plan run | explicit move line |
| Undone by | rerunning in reverse | editing the block |
Use moved when the refactor is real and permanent β it is versioned, reviewable, and everyone who applies gets the same result. Use state mv for one-off local repairs, or when state is drifted enough that you want to fix it before committing anything. Keep moved blocks in the code after applying; they are harmless once applied and rescue any state that has not caught up.
Forget and adopt: state rm and import
Sometimes the resource outlives your management of it. terraform state rm <address> forgets a resource β removing its entry from state without deleting anything in AWS:
terraform state rm aws_sqs_queue.dlq
# Successfully removed 1 resource instance(s).
Terraform now has amnesia about that queue; the queue keeps receiving messages. Reach for rm when another team is taking ownership, or when you intend to re-adopt the resource under a different address with hand-written configuration. Terraform 1.7 and later also offer a declarative form you can commit: delete the resource block and add a removed block with lifecycle { destroy = false }, and the next apply forgets the queue without deleting it.
terraform import does the reverse: it attaches an existing AWS resource to an address you have already declared in configuration. The command takes the address and the resource's AWS ID β for SQS, the queue URL:
terraform import aws_sqs_queue.dlq https://sqs.us-east-1.amazonaws.com/123456789012/processor-dlq
# aws_sqs_queue.dlq: Importing from ID "https://sqs.us-east-1.amazonaws.com/123456789012/processor-dlq"...
# Import successful!
β οΈ Import writes state only. It does not generate configuration. If your hand-written resource "aws_sqs_queue" "dlq" block does not match the real queue exactly, the next plan proposes to change β or replace β the resource you just adopted. On aws_sqs_queue, changing name forces replacement: by default Terraform deletes the old queue and then creates a new one, discarding every message still sitting in the DLQ. A no-op plan after import is your only proof the configuration matches.
Choosing the repair
| Symptom | What is really wrong | Tool |
|---|---|---|
| Address differs, resource identical | state keys, not AWS | state mv / moved |
| Live resource, absent from state | nothing owned yet | import |
| In state, owned by another team | should not be managed | state rm |
| Attribute edited in the console | live value drifted from configuration | refresh-only plan, then edit config or apply |
The last row belongs to "Reading State to Predict Plans": when someone changed a live value that configuration declares, a refresh-only plan shows the drift without changing infrastructure; you then either edit configuration to keep the console value or run a normal apply to restore the configured one. Import and rm are for identity problems, not value problems.
The cost asymmetry breaks ties. Replacing an SQS DLQ loses undelivered messages; replacing a DynamoDB table loses every item in it. Importing or moving avoids both, because neither operation deletes anything in AWS. When a plan proposes replacement of a stateful resource, stop and ask whether the real change is a rename you can fix with bookkeeping.
Practice: undo a module rename
A refactor renamed a module and an ECS service. State holds aws_dynamodb_table.sessions and aws_ecs_service.worker; configuration declares module.data.aws_dynamodb_table.sessions and module.app.aws_ecs_service.worker. The plan reads:
# aws_dynamodb_table.sessions will be destroyed
# (because aws_dynamodb_table.sessions is not in configuration)
- resource "aws_dynamodb_table" "sessions" { ... }
# aws_ecs_service.worker will be destroyed
# (because aws_ecs_service.worker is not in configuration)
- resource "aws_ecs_service" "worker" { ... }
# module.app.aws_ecs_service.worker will be created
+ resource "aws_ecs_service" "worker" { ... }
# module.data.aws_dynamodb_table.sessions will be created
+ resource "aws_dynamodb_table" "sessions" { ... }
Plan: 2 to add, 0 to change, 2 to destroy.
Write the sequence that makes plan report no changes, and name one situation in this stack where import is the right tool instead.
Check your answer
Move both addresses, then plan:
terraform state mv aws_dynamodb_table.sessions module.data.aws_dynamodb_table.sessions
terraform state mv aws_ecs_service.worker module.app.aws_ecs_service.worker
Either command alone fixes only half; both must match configuration before the plan is a no-op. A committed pair of moved blocks achieves the same result for the whole team:
moved {
from = aws_dynamodb_table.sessions
to = module.data.aws_dynamodb_table.sessions
}
moved {
from = aws_ecs_service.worker
to = module.app.aws_ecs_service.worker
}
import is right when the resource exists in AWS but was never in state at all β for example, an EventBridge rule a teammate created by hand in the console that the configuration already declares. state mv cannot help, because there is no old address to move from; you need terraform import aws_cloudwatch_event_rule.orders orders-rule.
Module Boundaries as State Namespaces and Interfaces
Rename a module block and nothing in AWS changes β yet terraform plan may offer to destroy an SQS queue and a DynamoDB table. The cause is the one behind the queue rename you traced in "Reading State to Predict Plans": Terraform keys every state entry by address, the dotted path such as module.messaging.aws_sqs_queue.this that names exactly one resource instance. A module call inserts itself into that path, so the module boundary is really a state namespace. Read it correctly and you predict replacement before it happens.
A module is a namespace, not a resource
A module is a directory of .tf files that Terraform reads as a child configuration. The module "messaging" { ... } block is not a resource and never appears as one in state β it only declares that the child's contents exist, under a name. An input (variable "queue_name") is a parameter the module accepts; an output (output "queue_arn") is a named value it hands back to the caller.
# root/main.tf
module "messaging" {
source = "./modules/messaging"
queue_name = "orders"
}
# modules/messaging/main.tf
variable "queue_name" {
type = string
}
resource "aws_sqs_queue" "dlq" {
name = "${var.queue_name}-dlq"
}
resource "aws_sqs_queue" "this" {
name = var.queue_name
redrive_policy = jsonencode({
deadLetterTargetArn = aws_sqs_queue.dlq.arn
maxReceiveCount = 5
})
}
output "queue_arn" {
value = aws_sqs_queue.this.arn
}
output "dlq_url" {
value = aws_sqs_queue.dlq.url
}
The resources have short local addresses β aws_sqs_queue.this, aws_sqs_queue.dlq β but the state keys are module-qualified:
terraform.tfstate
βββ module.messaging.aws_sqs_queue.this
βββ module.messaging.aws_sqs_queue.dlq
Child-module outputs such as queue_arn are not stored in state: Terraform re-evaluates them on every plan, which is how one module's output feeds another module's input. Only root-module outputs are saved. β οΈ By default all module addresses live in the root module's single state file. Splitting code into separate module directories does not split state, and a backend block inside a child module is ignored with a warning; only a separate root configuration, planned and applied on its own, gets its own state β the pattern that belongs with environment isolation.
Predict: the child declares aws_sqs_queue.this. What does terraform state list print, and what does the plan show if the root call is renamed to module.queues?
Check your answer
terraform state list prints module.messaging.aws_sqs_queue.this and module.messaging.aws_sqs_queue.dlq β module-qualified, one state file.
Rename the call to module.queues with nothing else changed and the addresses become module.queues.aws_sqs_queue.this and module.queues.aws_sqs_queue.dlq. Terraform finds no state entry at those new keys, so the plan destroys the two recorded queues and creates two new ones β same queue names in AWS, different state keys. Messages still in the old queue are gone.
Inputs and outputs are the contract
The root wires modules together by referencing outputs:
# root/main.tf (continued)
module "worker" {
source = "./modules/worker"
event_source_arn = module.messaging.queue_arn
dlq_url = module.messaging.dlq_url
}
Two rules follow. First, the output names are the contract. Rename queue_arn to sqs_queue_arn without updating the caller and the plan stops before it starts, with Error: Unsupported attribute pointing at module.messaging.queue_arn; terraform validate catches it too. That failure is loud and safe β treat it like a compile error.
Second, the values are not protected by the boundary. Suppose the messaging module changes how it derives the queue name to "${var.queue_name}-v2". name on aws_sqs_queue is a replacement-only (ForceNew) argument, so the queue is destroyed and recreated. The new ARN flows out through queue_arn into module.worker, and because event_source_arn on aws_lambda_event_source_mapping is also ForceNew, the consumer's trigger is replaced as well. A change inside one module propagated into a replacement in another.
Provider inheritance
A provider is the plugin that talks to an API β here hashicorp/aws. A provider alias is a second, named configuration of the same provider. A child module inherits the root's default provider configuration: same region, same credentials (IAM Identity Center SSO locally, an OIDC-assumed role in CI), same default tags. To point a module at another region or account, pass the alias explicitly:
provider "aws" {
region = "us-east-1"
}
provider "aws" {
alias = "replica"
region = "us-west-2"
}
module "messaging" {
source = "./modules/messaging"
providers = { aws = aws, aws.replica = aws.replica }
queue_name = "orders"
}
The module has to request that alias β it cannot invent a second region on its own:
# modules/messaging/versions.tf
terraform {
required_providers {
aws = {
source = "hashicorp/aws"
version = ">= 5.0"
configuration_aliases = [aws.replica]
}
}
}
β οΈ Do not declare a provider "aws" block inside a reusable module. Every caller then gets a duplicate provider configuration it cannot influence, and Terraform refuses to use such a module with count, for_each or depends_on β a limitation that surfaces exactly when you want one module call per environment. Terraform has discouraged in-module provider blocks since 0.13; configuration_aliases is the supported alternative.
Source and version changes rewrite internal addresses
Changing source from ./modules/messaging to a registry module, or bumping a pinned version, swaps the internal resource blocks for a different set. Terraform still matches by address, so if the old version named its queue aws_sqs_queue.queue and the new one names it aws_sqs_queue.this, the plan destroys module.messaging.aws_sqs_queue.queue and creates module.messaging.aws_sqs_queue.this. Two mitigations, and good modules ship both:
- Stable internal addresses β names like
this,main,dlqthat survive version bumps. movedblocks inside the module, so the module declares its own rename and callers never see it:
# modules/messaging/moved.tf
moved {
from = aws_sqs_queue.queue
to = aws_sqs_queue.this
}
Addresses inside a child module are relative to that module, so this one block applies to every caller at once: module.messaging.aws_sqs_queue.queue becomes module.messaging.aws_sqs_queue.this in every state that uses the module, and the plan reports a move rather than a destroy.
β οΈ The relative-address rule cuts both ways: a moved block inside a module can only move its own resources and those of its child modules. A move from one module call to a different module call must be written in the module that calls both β the root, for top-level calls. The one-off command-line alternative, terraform state mv, is covered in "State Operations for Safe Repair"; prefer the versioned moved block when a team shares the code, because it travels with the repository.
Sensitive outputs still live in state
sensitive = true on an output is a display control, not a storage control:
output "consumer_secret" {
value = random_password.consumer.result
sensitive = true
}
The CLI prints <sensitive> instead of the string (plans show (sensitive value)), but the plaintext still lands in terraform.tfstate, in the result attribute of random_password.consumer (and under outputs too when this is a root-module output) β and that file is what your S3 backend uploads. For AWS work, avoid routing secrets through outputs at all: keep the value in SSM Parameter Store or Secrets Manager, pass Lambda only the parameter name as an environment variable, and let the function fetch the value at runtime. Pulling the value with data "aws_ssm_parameter" and injecting it into a resource argument puts the plaintext straight back into state. β οΈ State leakage is credential leakage: whoever can read the state bucket reads every secret the configuration has touched.
Practice: split one module into two
State currently holds:
module.messaging.aws_sqs_queue.this
module.messaging.aws_sqs_queue.dlq
module.messaging.aws_sns_topic.events
module.messaging.aws_dynamodb_table.offsets
You split modules/messaging into modules/queues (queue, DLQ, offsets table) and modules/topics (the SNS topic). The new modules keep the same internal resource names, and the root now calls module "queues" and module "topics" instead of module "messaging".
(a) With no refactoring help, which resources does the plan replace, and which replacement loses data?
(b) Write moved blocks β and say where they live β so the plan reports moves instead of destroys.
Check your answer
(a) Every address changed its prefix, so all four are replaced. Replacing aws_dynamodb_table.offsets loses the offsets data, and replacing the SQS queue discards any messages still sitting in it. Replacing the DLQ also discards any failed messages waiting in it. The SNS topic stores no messages, but deleting it also deletes all its subscriptions.
(b) All four moves cross module boundaries, so they belong in the root module, not inside modules/queues or modules/topics:
# root/moved.tf
moved {
from = module.messaging.aws_sqs_queue.this
to = module.queues.aws_sqs_queue.this
}
moved {
from = module.messaging.aws_sqs_queue.dlq
to = module.queues.aws_sqs_queue.dlq
}
moved {
from = module.messaging.aws_dynamodb_table.offsets
to = module.queues.aws_dynamodb_table.offsets
}
moved {
from = module.messaging.aws_sns_topic.events
to = module.topics.aws_sns_topic.events
}
Expected plan: 0 to add, 0 to change, 0 to destroy, with four move notices. A two-block chain works too: one block with from = module.messaging and to = module.queues moves all four, and a second moves module.queues.aws_sns_topic.events on to module.topics.aws_sns_topic.events. Two limits. If the new modules had renamed their internal resources β say aws_sqs_queue.main instead of aws_sqs_queue.this β each to would need the new internal name. And if the split genuinely wants a different AWS object, such as a brand-new topic rather than the existing one, a move is the wrong tool: you want the create, adopting a pre-existing resource with import only if it already exists.
Before you merge a module change, check three things: (1) every address in the new plan already exists in the old state or is a deliberate create β no unexplained destroys; (2) moved blocks sit in the module that owns both addresses (inside for internal renames, root for cross-module moves); (3) no secret reaches an output, and the state bucket's access policy still matches who may read it.
Composing Modules for AWS Backends Without Creating State Hazards
Terraform applies resources in dependency order β but it only knows the order you tell it. Two modules that look obviously related, a queue and the Lambda that consumes it, have no relationship in Terraform's graph unless something in one is referenced by the other. That single fact explains most broken compositions: a Lambda event source mapping (the resource that tells Lambda to poll an SQS queue) created before the queue exists, an IAM policy attached after AWS already rejected the poll, and plans that fail outright with a dependency cycle β a loop in the graph where A waits on B and B waits on A.
Wiring by reference: how Terraform learns the order
An implicit dependency exists when one resource or module reads an attribute of another. Consider a messaging module (SQS queue plus dead-letter queue) and a worker module (Lambda plus its event source mapping):
module "messaging" {
source = "./modules/messaging"
queue_name = "orders-${var.environment}"
}
module "worker" {
source = "./modules/worker"
function_name = "orders-processor-${var.environment}"
event_source_arn = module.messaging.queue_arn # <- the dependency
}
module.messaging exposes queue_arn from aws_sqs_queue.this.arn β the input/output contract described in "Module Boundaries as State Namespaces and Interfaces". Passing that value into event_source_arn produces this graph:
aws_sqs_queue.dlq (dead-letter queue)
β
aws_sqs_queue.this (main queue, redrive policy -> dlq)
β queue_arn output
aws_lambda_event_source_mapping.this
Because the ARN cannot be read until the queue exists, Terraform creates the DLQ, then the queue, then the mapping. You declared no ordering; the reference is the ordering.
When no attribute can express it: depends_on
Some dependencies have nothing to reference. A classic: the worker's execution role needs sqs:ReceiveMessage permission attached before the event source mapping is created, or AWS rejects the mapping because the role cannot poll yet. If module.worker only references module.iam.role_arn, Terraform knows to create the role first β but not the separate aws_iam_role_policy_attachment inside that module, which the output does not depend on.
module "worker" {
source = "./modules/worker"
function_name = "orders-processor-${var.environment}"
role_arn = module.iam.role_arn
event_source_arn = module.messaging.queue_arn
depends_on = [module.iam] # all of module.iam, not just the role
}
The distinction matters: an attribute reference pins only what produces that value; depends_on on a module pins every resource the module creates. Reach for depends_on when the ordering is real but invisible in the data flow β not as a reflex.
β οΈ A data source (data "aws_ssm_parameter" ..., data "terraform_remote_state" ...) is a poor place for depends_on. With depends_on, Terraform defers its read to apply time whenever anything it depends on has a pending change, so the plan fills with (known after apply) and you lose the ability to review what will change. This is revisited in the practice below.
Reaching across state files
A terraform_remote_state data source reads the outputs of another state file β nothing else. It exists for values a different team or stack owns, such as the VPC and private subnet IDs an ECS Fargate service needs:
# Read-only: this data source lists the bucket and reads the state object.
data "terraform_remote_state" "network" {
backend = "s3"
config = {
bucket = "acme-tfstate"
key = "network/terraform.tfstate"
region = "eu-west-1"
}
}
module "api" {
source = "./modules/ecs-service"
cluster_id = data.terraform_remote_state.network.outputs.cluster_id
subnet_ids = data.terraform_remote_state.network.outputs.private_subnet_ids
environment = var.environment
}
Grant the reading role s3:ListBucket on the bucket and s3:GetObject on that key, and treat the outputs as a contract you validate. That access exposes the whole state file, secrets included, not just its outputs. The cost is coupling: you depend on an entire state file that other people apply whenever they like. If they rename private_subnet_ids, your plan fails with Unsupported attribute, and because remote state exposes all outputs, a change anywhere in that file is a change to your blast radius.
For stable AWS values, prefer an SSM Parameter Store lookup β a narrower interface with its own lifecycle and IAM permissions:
data "aws_ssm_parameter" "private_subnets" {
name = "/network/${var.environment}/private_subnet_ids"
}
Use remote state when you genuinely need another stack's own Terraform-computed outputs; use SSM when the value is simply "the current subnet IDs," which is most of the time.
Keeping environments apart
Two isolation strategies, with different failure modes:
- Separate root directories β
envs/dev/,envs/prod/, each with its own backend block andterraform.tfvars. State boundaries are explicit on disk; you cannot run the prod plan from the dev folder. - Terraform workspaces β one configuration, and each workspace gets its own state under a shared key pattern. Modules receive
environment = terraform.workspaceand name resources with it.
Workspaces are convenient and dangerous: everything shares one backend configuration and one variable set, so terraform workspace select is a one-word step between planning dev and applying to prod. If any resource name forgets to include terraform.workspace, two environments collide on the same AWS name. For backends that hold DynamoDB tables and SQS queues, the directory-per-environment layout is the safer default.
Module lifecycle and the plan you must read
Pin registry modules with version, upgrade one environment first, and read the plan for # forces replacement on anything stateful. Replacing aws_dynamodb_table loses data; replacing aws_sqs_queue discards undelivered messages; replacing aws_ecs_service or an API Gateway stage can drop traffic, and replacing the REST API itself changes its invoke URL, which embeds the API ID. Addresses live here too: if an upgrade replaces the aws_eip behind a NAT gateway, your outbound public IP changes and any partner allowlist pinned to it breaks.
Diagnosing a broken composition
Two failures dominate. First, a cycle: module A passes an output to module B while B passes one back to A, and each of those outputs is computed from what the module received. (Mutual references alone are fine: Terraform orders individual variables and outputs, not whole modules.) Terraform prints Error: Cycle: listing every variable, output and resource in the loop, and no plan is produced. Fix it by removing one edge β read the shared value from a data source instead, or split the interface so values flow one way only.
Second, a null output from a conditional resource:
resource "aws_sqs_queue" "dlq" {
count = var.create_dlq ? 1 : 0
name = "orders-dlq-${var.environment}"
}
output "dlq_arn" {
value = try(aws_sqs_queue.dlq[0].arn, null)
}
With create_dlq = false, aws_sqs_queue.dlq is an empty tuple, so aws_sqs_queue.dlq[0].arn raises an invalid-index error; try swallows it and yields null. A consumer that writes that null into a redrive policy gets a broken configuration rather than a clear failure. An empty string is no better; keep null as the "no DLQ" signal and have the consumer branch on it, e.g. redrive_policy = var.dlq_arn == null ? null : jsonencode({ ... }), or make the optional piece of the interface a separate module.
Practice: compose the stack, then judge one depends_on
Build a Lambda + SQS + DLQ + DynamoDB stack from three modules (messaging, worker, iam) plus a terraform_remote_state data source that supplies the DynamoDB table name and ARN from the data team's state. Write the module calls. Then state one place where depends_on is required and one where it would be wrong, with the reason.
Check your answer
data "terraform_remote_state" "data" {
backend = "s3"
config = {
bucket = "acme-tfstate"
key = "data/terraform.tfstate"
region = "eu-west-1"
}
}
module "messaging" {
source = "./modules/messaging"
queue_name = "orders-${var.environment}"
}
module "iam" {
source = "./modules/iam"
role_name = "orders-processor-${var.environment}"
}
module "worker" {
source = "./modules/worker"
function_name = "orders-processor-${var.environment}"
role_arn = module.iam.role_arn
event_source_arn = module.messaging.queue_arn
table_name = data.terraform_remote_state.data.outputs.sessions_table_name
table_arn = data.terraform_remote_state.data.outputs.sessions_table_arn
# Required: the execution role must already carry the SQS read permissions
# (sqs:ReceiveMessage, sqs:DeleteMessage, sqs:GetQueueAttributes), or AWS
# rejects the event source mapping at create time.
depends_on = [module.iam]
}
The depends_on on module.worker is required because the policy attachment inside module.iam produces no value worker references β the role ARN alone does not order it. It would be wrong on data.terraform_remote_state.data: it adds no real ordering (the read is just an API call) and whenever its dependency has a pending change it defers the read to apply time, blanking the table names to (known after apply) in exactly the plans you most need to review. The correct answer to "worker needs the queue first" is the existing event_source_arn reference, not another depends_on.
Transfer task and checklist
Take the configuration above and add a fourth module that creates an EventBridge rule targeting your queue. Decide, before writing anything, whether the rule's target ARN should come from module.messaging (a reference), from SSM, or from remote state β and justify it by blast radius.
- Attribute references order most dependencies;
depends_oncovers the rest, and rarely belongs on a data source. terraform_remote_statereads outputs only, from another state file; prefer SSM for stable AWS values.- Separate root directories make environment state boundaries explicit; workspaces share one backend key pattern.
- Read the plan for
# forces replacementon stateful or costly resources before merging a module upgrade. - Name the cycle or the null output before attempting a fix.
Independent Transfer: Diagnose and Safely Change a Serverless Backend
A colleague refactored the Terraform for your order-processing backend β the DynamoDB table moved under a new module, the worker module's version was bumped, and a Lambda function was renamed. terraform plan now proposes to destroy your sessions table. Nobody has run apply yet. That gap between plan and apply is where this lesson lives.
Your job is not to recall syntax. It is to read the plan, classify each proposed action by its cause, and pick exactly one of three tools for each problem: a moved block (a declarative instruction in configuration that says "this resource lives at a new address now"), terraform import (a command that attaches an existing AWS resource to an address Terraform is not yet tracking), or plain apply. The commands themselves were covered in "State Operations for Safe Repair"; here you are choosing between them with a loaded plan on screen.
The state you inherited
$ terraform state list
aws_dynamodb_table.sessions
module.messaging.aws_sqs_queue.orders
module.worker.aws_lambda_event_source_mapping.processor
module.worker.aws_lambda_function.processor
Three facts about this environment:
- Configuration now declares the table at
module.data.aws_dynamodb_table.sessions. The real table in AWS is untouched. - The module bump changed
function_namefromorder-processortoorder-processor-v2. That argument is immutable β AWS has no API to rename a function, so Terraform must destroy and rebuild it. - Someone deleted the DLQ (dead-letter queue β the queue that catches messages a consumer failed to process) in the console; a colleague then recreated it by hand under the same name. The queue exists in AWS but appears nowhere in the state list above.
Task 1 β Read the plan
$ terraform plan
Terraform will perform the following actions:
# aws_dynamodb_table.sessions will be destroyed
# (because aws_dynamodb_table.sessions is not in configuration)
- resource "aws_dynamodb_table" "sessions" { ... }
# module.data.aws_dynamodb_table.sessions will be created
+ resource "aws_dynamodb_table" "sessions" { ... }
# module.messaging.aws_sqs_queue.dlq will be created
+ resource "aws_sqs_queue" "dlq" { ... }
# module.worker.aws_lambda_event_source_mapping.processor will be updated in-place
~ resource "aws_lambda_event_source_mapping" "processor" {
~ batch_size = 10 -> 25
}
# module.worker.aws_lambda_function.processor must be replaced
-/+ resource "aws_lambda_function" "processor" {
~ function_name = "order-processor" -> "order-processor-v2" # forces replacement
# (3 unchanged attributes hidden)
}
Plan: 3 to add, 1 to change, 2 to destroy.
| Plan line | Symbol | Root cause |
|---|---|---|
aws_dynamodb_table.sessions β module.data... |
- then + |
Address change (same table) |
module.worker...processor |
-/+ |
Immutable argument: function_name |
module.messaging...dlq |
+ only |
Drift: gone from state |
...event_source_mapping.processor |
~ |
Updateable argument: batch_size |
The diagnostic rule that separates the two dangerous lines:
-/+on one address means the resource itself is being replaced β look for the# forces replacementcomment; it names the immutable argument (function_nameon a Lambda,nameon a queue or table).-on address A and+on a different address B means only the address changed; the destroy carries a# (because β¦ is not in configuration)line. The AWS resource is fine. This is the case you can fix without touching AWS at all.+with no matching-means Terraform has no state entry: a resource you just added, one deleted out-of-band (drift that refresh dropped from state), or a live resource nobody ever tracked.
Task 2 β Stop the DynamoDB destruction
The table is stateful: destroying it destroys every session. Between moved, terraform state mv, and import, the right tool here is a move β you are not adopting a foreign resource, only telling Terraform that an already-tracked one changed address. Because the fix belongs in version control where the whole team and CI see it, prefer the declarative form:
# root module β declares that the table's address changed
moved {
from = aws_dynamodb_table.sessions
to = module.data.aws_dynamodb_table.sessions
}
Expected plan output β the destroy/create pair is replaced by a single note, and the counts drop by one destroy and one add:
# aws_dynamodb_table.sessions has moved to module.data.aws_dynamodb_table.sessions
resource "aws_dynamodb_table" "sessions" { ... }
Plan: 2 to add, 1 to change, 1 to destroy.
The two remaining additions are the DLQ and the replacement Lambda. If the move were the only change in the configuration, the summary would read Plan: 0 to add, 0 to change, 0 to destroy. β that is the test for a clean refactor. Exact move rendering varies slightly by version; the count is the part to trust. terraform state mv aws_dynamodb_table.sessions module.data.aws_dynamodb_table.sessions rewrites the shared state at once but leaves no record in code: every other state that holds the table (staging, prod) needs the same command by hand, and until your branch merges, anyone planning from the old configuration sees the table destroyed and re-created.
Not a conflict: running
terraform state mvand committing amovedblock is safe. Once a state already holds thetoaddress, Terraform ignores the block, and the block still moves the table in every state nobody fixed by hand.
Task 3 β Repair the DLQ: import or recreate?
The plan's bare + for module.messaging.aws_sqs_queue.dlq is ambiguous by itself β deleted-for-real and never-tracked look identical. Check AWS before deciding:
$ aws sqs get-queue-url --queue-name order-dlq
{
"QueueUrl": "https://sqs.eu-west-1.amazonaws.com/123456789012/order-dlq"
}
The queue is alive. Import it. A plain apply would not build a second queue: SQS rejects a create for an existing name whose attributes differ, and when they match it returns the existing queue's URL, so Terraform would adopt the queue without ever showing you a diff against it:
$ terraform import module.messaging.aws_sqs_queue.dlq \
https://sqs.eu-west-1.amazonaws.com/123456789012/order-dlq
On Terraform 1.5 or later you can commit the equivalent instead, which survives into CI:
import {
to = module.messaging.aws_sqs_queue.dlq
id = "https://sqs.eu-west-1.amazonaws.com/123456789012/order-dlq"
}
Import writes state only β it never edits your configuration (on 1.5+, terraform plan -generate-config-out=FILE can draft a block from an import block for you to review). The plan is a no-op only if every argument in your aws_sqs_queue block matches what the hand-recreated queue actually has. If the console recreation left message_retention_seconds at the AWS default instead of your configured value, the first plan after import shows an in-place update, not a replacement. Bring those attributes into your block (or accept the update) before you apply.
When would recreating be right instead? If aws sqs get-queue-url returns an error, the queue genuinely does not exist, there is nothing to import, and apply builds a fresh one. Justify that choice on message-loss cost and whether the URL is referenced elsewhere: because the SQS URL is derived from account, region, and queue name, a same-name rebuild yields the same URL and same ARN, so a redrive policy (the main queue's setting naming which DLQ receives failures) written as an attribute reference re-resolves cleanly. What the out-of-band deletion already cost you is the undelivered messages, any attributes tuned in the console (the rebuild uses configuration's values), and a window where the redrive target was absent.
Task 4 β The module upgrade edge case
A new version of modules/api renamed its internal ECS service address (the ECS service, the long-running definition that keeps N copies of a Fargate task alive) from aws_ecs_service.main to aws_ecs_service.api. Left alone, every caller of that module gets a destroy-and-recreate on next apply. The fix ships inside the module version, so callers inherit it for free:
# modules/api/moved.tf β addresses here are relative to the module itself
moved {
from = aws_ecs_service.main
to = aws_ecs_service.api
}
Inside a module, you write addresses without the module.api. prefix. If instead the module call was renamed in the root (module.api β module.orders_api), the moved block lives in the root and does carry the prefix. Before merging the version bump, run terraform plan against a real environment and require that the ECS service line shows a move and 0 to destroy. An ECS replacement forces a new deployment and can drop in-flight requests; for a stateful resource the same mistake loses data outright.
Your turn β an independent variation
You inherit this state and configuration:
# state list
aws_dynamodb_table.audit
module.api.aws_lambda_function.authorizer
# configuration (new layout)
module "storage" { source = "./modules/storage" } # declares aws_dynamodb_table.audit
module "api" { source = "./modules/api" } # declares aws_lambda_function.jwt_authorizer
Both real resources still exist in AWS. Write the exact configuration or commands needed, then predict the plan summary.
Check your answer
The DynamoDB table exists in state at aws_dynamodb_table.audit and in AWS β only its address changed. Move it; do not import a table you already track.
moved {
from = aws_dynamodb_table.audit
to = module.storage.aws_dynamodb_table.audit
}
The Lambda's address stayed inside module.api, so the rename is decidable locally by the module. Two valid options: put a moved block inside modules/api, or use terraform state mv module.api.aws_lambda_function.authorizer module.api.aws_lambda_function.jwt_authorizer. The block is preferable because CI and teammates see it.
Predicted summary: Plan: 0 to add, 0 to change, 0 to destroy. If the plan instead shows 1 to change, the imported or moved resource has an attribute mismatch β a tag or retention value drifted, and the fix is to reconcile the configuration, not to re-run the move.
Why not import? Import is for a resource that exists in AWS but has no state entry. Importing the table at its new address would succeed and leave two state entries for one table; the old entry, with no configuration left, would then plan a destroy that deletes the real table. state rm plus import would work, but it is two risky steps where one move suffices.
Checklist before you touch state
- Read the plan symbol first:
-/+means an immutable argument on the same address;-on A with+on B means an address change. - Inspect safely with
terraform state listandterraform state showbefore any operation that proposes destruction β and remembershowcan print secrets in plaintext. - Prefer a committed
movedblock over a localterraform state mvwhen a team or CI shares the state. - Import only what AWS actually still has; verify with a read-only API call, and expect a no-op plan solely when configuration matches the live attributes.
- Use
terraform state rmto stop managing a resource, never to delete it. - Treat any plan that destroys a DynamoDB table, an SQS queue, a database, or a secrets-bearing resource as a stop-and-diagnose signal, not a review comment.