A Serverless AWS Stack in Terraform

The services of a typical event-driven backend written as Terraform: Lambda with SQS triggers, DynamoDB, IAM, ECS on Fargate, API Gateway, EventBridge and Parameter Store, plus how to test and review the result.

Last generated

Lesson 6 of 6 available15 practice questions

SPACED REPETITION Β· 15 practice questions

Make this lesson stick.

Try 3 questions now. No account needed. Sample answers aren't saved.

Reading the Plan of a Whole Stack Before You Change Anything

You have built this backend by hand before: click the function, paste the queue ARN into an environment variable, hope you remembered the table name. Terraform promises the second time is cheaper β€” but the price is that one terraform apply can touch a dozen AWS resources, and the only warning you get is a wall of text that scrolls past in two seconds and ends with Plan: 8 to add, 0 to change, 0 to destroy.

That wall of text is the most valuable artifact in the entire workflow. This section teaches you to read it: what each symbol means, what the printed order does and does not tell you, and how to predict the plan for an eight-resource serverless stack before you run the command. Every later technique in this lesson β€” references, replacements, moved blocks β€” is judged by how it changes this output.

Five words to fix before the first line of output

Configuration is the set of .tf files you edit: plain HCL declaring what should exist. State is Terraform's own record, usually terraform.tfstate, of the resources it last applied and the attribute values it saw at that moment β€” it is how Terraform knows that the DynamoDB table in front of it is the same table it created last week rather than a new one. The plan is the diff among three inputs: your configuration, that state file, and the real cloud, read back through AWS API calls.

Two more. A resource address is the name you see in plans and state: aws_lambda_function.worker means "resource of type aws_lambda_function, named worker in this configuration"; inside a module it gains a prefix, as in module.app.aws_lambda_function.worker. And the provider is the plugin β€” hashicorp/aws here β€” that knows how to turn a resource "aws_lambda_function" block into a CreateFunction call. The required_providers block constrains which provider versions are acceptable; the dependency lock file .terraform.lock.hcl records the exact version terraform init selected.

An eight-resource stack you can hold in your head

Here is the whole thing in one file. Read the expressions, not just the argument names β€” those expressions are the subject of this section.

terraform {
  required_version = ">= 1.9"

  required_providers {
    aws = {
      source  = "hashicorp/aws"
      version = "~> 6.0"
    }
  }
}

resource "aws_ssm_parameter" "worker_config" {
  name  = "/worker/config"
  type  = "String"
  value = "max-batch=10"
}

resource "aws_iam_role" "worker" {
  name = "worker-role"

  assume_role_policy = jsonencode({
    Version = "2012-10-17"
    Statement = [{
      Effect    = "Allow"
      Action    = "sts:AssumeRole"
      Principal = { Service = "lambda.amazonaws.com" }
    }]
  })
}

resource "aws_dynamodb_table" "jobs" {
  name         = "jobs"
  billing_mode = "PAY_PER_REQUEST"
  hash_key     = "id"

  attribute {
    name = "id"
    type = "S"
  }
}

resource "aws_sqs_queue" "jobs" {
  name = "jobs"
}

resource "aws_lambda_function" "worker" {
  function_name = "worker"
  role          = aws_iam_role.worker.arn
  runtime       = "nodejs24.x"
  handler       = "index.handler"
  filename      = "worker.zip"

  environment {
    variables = {
      QUEUE_URL  = aws_sqs_queue.jobs.url
      TABLE_NAME = aws_dynamodb_table.jobs.name
    }
  }
}

resource "aws_lambda_event_source_mapping" "worker" {
  event_source_arn = aws_sqs_queue.jobs.arn
  function_name    = aws_lambda_function.worker.arn
  batch_size       = 10
}

resource "aws_apigatewayv2_api" "http" {
  name          = "jobs-api"
  protocol_type = "HTTP"
}

resource "aws_apigatewayv2_stage" "default" {
  api_id      = aws_apigatewayv2_api.http.id
  name        = "$default"
  auto_deploy = true
}

Each aws_X.Y.attribute expression is a reference β€” it names another resource's attribute, and Terraform treats it as both a value and an ordering constraint. Read the references and you already have the graph: the role's arn, the table's name and the queue's url feed the function; the queue's arn and the function's arn feed the event source mapping; the API's id feeds the stage. The SSM parameter feeds nothing, so nothing waits for it.

(Simplified on purpose: no IAM policy attachment, no API route or integration, no DLQ, and worker.zip must exist on disk when the function is created: the provider reads it during apply, so the plan still prints and a missing file fails the apply with reading ZIP file (worker.zip). Each addition would create more resources but would not change how you read the plan.)

What the plan prints

Symbol Meaning
+ create
~ update in place
- destroy
-/+ destroy and recreate
+/- create the replacement, then destroy (with create_before_destroy)
<= read a data source

A first apply shows only + blocks, and it looks like this (trimmed β€” five more blocks follow the same shape):

Terraform will perform the following actions:

  # aws_apigatewayv2_api.http will be created
  + resource "aws_apigatewayv2_api" "http" {
      + api_endpoint  = (known after apply)
      + name          = "jobs-api"
      + protocol_type = "HTTP"
    }

  # aws_dynamodb_table.jobs will be created
  + resource "aws_dynamodb_table" "jobs" {
      + arn          = (known after apply)
      + billing_mode = "PAY_PER_REQUEST"
      + hash_key     = "id"
      + name         = "jobs"
    }

  # aws_lambda_function.worker will be created
  + resource "aws_lambda_function" "worker" {
      + filename      = "worker.zip"
      + function_name = "worker"
      + role          = (known after apply)
      + runtime       = "nodejs24.x"
      + environment {
          + variables = {
              + "QUEUE_URL"  = (known after apply)
              + "TABLE_NAME" = "jobs"
            }
        }
    }

Plan: 8 to add, 0 to change, 0 to destroy.

Read (known after apply) as "AWS computes this at apply": a resource's own arn or api_endpoint prints that way, and so does any value that flows in through a reference, which is why, inside a consumer's block, it fingerprints a reference edge. TABLE_NAME prints as "jobs" because you set the name literally, so Terraform knows it now; QUEUE_URL and role are values AWS computes at creation time, so they cannot be known yet. When you later change something, ~ marks an in-place update, and if the changed attribute is immutable for that service, the block flips to -/+ with an annotation on the offending line β€” for example ~ name = "jobs" -> "jobs-v2" # forces replacement. That annotation is how you find the reason for a surprise destroy without hunting through provider docs.

The printed order is not the order of work

Terraform prints changes in resource address order, with creates, updates and destroys interleaved rather than grouped by action, which for our eight creates is alphabetical: api, stage, table, role, event source mapping, function, queue, parameter. That ordering is cosmetic. The real schedule comes from the dependency graph, which Terraform derives from your references, and independent nodes run concurrently, up to Terraform's own -parallelism limit (10 operations by default).

wave 1 (all at once)
  aws_iam_role.worker
  aws_dynamodb_table.jobs
  aws_sqs_queue.jobs
  aws_apigatewayv2_api.http
  aws_ssm_parameter.worker_config
        ↓
wave 2
  aws_lambda_function.worker      (needs role.arn, table.name, queue.url)
  aws_apigatewayv2_stage.default  (needs api.id)
        ↓
wave 3
  aws_lambda_event_source_mapping.worker  (needs function.arn, queue.arn)

To see the graph itself, terraform graph prints it in DOT format (noisy once you have a hundred nodes). To get an unambiguous, machine-readable version of a plan, save it β€” terraform plan -out=tfplan, then terraform show tfplan β€” and note that terraform show -json tfplan lists an explicit action for every resource change, so you never have to infer one from column alignment.

Your turn: predict the plan

Using only the configuration above, write out (a) the eight plan lines in the order you believe Terraform prints them, (b) the list of resources that cannot start until another has finished, and (c) whether any line would carry # forces replacement. Commit to an answer before opening the check.

Check your answer

Printed order (address order within the create group): aws_apigatewayv2_api.http, aws_apigatewayv2_stage.default, aws_dynamodb_table.jobs, aws_iam_role.worker, aws_lambda_event_source_mapping.worker, aws_lambda_function.worker, aws_sqs_queue.jobs, aws_ssm_parameter.worker_config.

Real dependency edges, and the reference that creates each:

  1. aws_iam_role.worker.arn β†’ aws_lambda_function.worker.role
  2. aws_dynamodb_table.jobs.name β†’ function's TABLE_NAME
  3. aws_sqs_queue.jobs.url β†’ function's QUEUE_URL
  4. aws_sqs_queue.jobs.arn β†’ aws_lambda_event_source_mapping.worker.event_source_arn
  5. aws_lambda_function.worker.arn β†’ mapping's function_name
  6. aws_apigatewayv2_api.http.id β†’ aws_apigatewayv2_stage.default.api_id

Five resources have no inbound edges and can be created simultaneously; the function and the stage wait for wave 1; the mapping waits for both the function and the queue. That is why the printed sequence cannot be the execution order: the queue, printed seventh, must finish before the function printed sixth can start, and the mapping printed fifth is the last thing created.

Forces replacement: none. On a create-only plan there is no old resource to destroy, so # forces replacement cannot appear β€” it only ever annotates an attribute that already exists in state.

Three ways this plan misleads people

Treating the printed order as the build order. If you debug a failure by assuming aws_lambda_event_source_mapping.worker is attempted fifth because it prints fifth, you will look in the wrong place. It runs last, and only after the queue and function both exist.

Reading only the summary line. Plan: 8 to add, 0 to change, 0 to destroy is a count, not a verdict. A single -/+ buried in the body of a large plan means a table gets deleted and recreated; the summary would still read 2 to add, 1 to change, 1 to destroy and look perfectly calm. Scan the body for - and -/+ before you approve anything.

Assuming an empty plan means config matches the cloud. An empty plan means Terraform found nothing to change among the resources it manages. It says nothing about resources Terraform does not manage (a queue someone created by hand never appears in it), about attributes listed in ignore_changes, or about values your code looks up for itself at runtime. Deploy-time versus runtime resolution of a value is exactly the distinction we take up next in "Wiring Services Together: References, Data Flow and Contracts."

Wiring Services Together: References, Data Flow and Contracts

A Lambda in one corner of your stack needs three things from other corners: the SQS queue it polls, the DynamoDB table it writes, and an ARN to drop into an IAM policy. None of those strings are knowable in advance β€” the queue URL contains your account ID, the table ARN contains the region. So the stack has to carry them from where they are created to where they are used. There are exactly two ways to do that, and they leave very different fingerprints in terraform plan.

Deploy-time wiring vs runtime lookup

A deploy-time reference is a resource attribute used inside another resource's argument, like aws_sqs_queue.jobs.url inside a Lambda's environment block. Terraform reads the value and writes it into the function's configuration. That single expression does two jobs at once: it moves the value, and it creates an implicit dependency β€” an ordering edge Terraform derives automatically from the reference, so the queue is created before the function.

A runtime lookup is when Terraform stores only a pointer (a parameter name, a hostname) and the application resolves the actual value itself when it runs β€” calling SSM at cold start, or following a DNS name.

Deploy-time reference             Runtime lookup
────────────────────              ─────────────────
aws_sqs_queue.jobs.url            aws_ssm_parameter.db.name
        β”‚ writes the value                 β”‚ writes only the name
        β–Ό                                  β–Ό
Lambda env: QUEUE_URL="<url>"     Lambda env: DB_PARAM="/app/db"
        β”‚                                  β”‚
  baked into config                 app calls SSM at runtime

Only the deploy-time form shows up in the plan as an update when the source changes. That single sentence is the whole section.

Deploy-time reference Runtime lookup
Resolved by Terraform your code
In plan when source changes ~ update nothing
Repoint without redeploy no yes
Ordering edge created yes no

Worked example: one parameter, two consumers

The crispest way to feel the difference is one SSM parameter consumed two ways.

resource "aws_ssm_parameter" "db_url" {
  name  = "/app/db_url"
  type  = "String"
  value = "postgres://db-old.internal:5432/app"
}

# Consumer A: deploy-time -- the value is baked into the function config.
resource "aws_lambda_function" "a" {
  function_name = "consumer-a"
  # ... role, filename ...
  environment {
    variables = { DB_URL = aws_ssm_parameter.db_url.value }
  }
}

# Consumer B: runtime -- only the parameter NAME is baked in.
resource "aws_lambda_function" "b" {
  function_name = "consumer-b"
  # ... role, filename ...
  environment {
    variables = { DB_URL_PARAM = aws_ssm_parameter.db_url.name }
  }
}

Now edit value to postgres://db-new.internal:5432/app and run terraform plan:

  # aws_lambda_function.a will be updated in-place
  ~ resource "aws_lambda_function" "a" {
      ~ environment {
          ~ variables = {
              ~ "DB_URL" = (sensitive value)
            }
        }
    }

  # aws_ssm_parameter.db_url will be updated in-place
  ~ resource "aws_ssm_parameter" "db_url" {
      ~ value   = (sensitive value)
      ~ version = 1 -> (known after apply)
    }

Two resources marked ~, and the changed value is redacted in both: the provider always marks a parameter's value as sensitive, and Terraform hides anything derived from it. Consumer A is redeployed because its config literally contained the old string. Consumer B produces no function diff at all: its DB_URL_PARAM is still /app/db_url, which did not change. B picks up the new URL on its next cold start with zero redeploy.

The data-source variant, and its footgun

A data "aws_ssm_parameter" block reads the parameter at refresh time:

data "aws_ssm_parameter" "db_url" {
  name = "/app/db_url"
}

resource "aws_lambda_function" "c" {
  environment {
    variables = { DB_URL = data.aws_ssm_parameter.db_url.value }
  }
}

Because a refresh runs on every default plan, a change to /app/db_url does surface as a ~ on function c. So it is not invisible β€” but it carries no ordering edge to whatever writes the parameter. If this same config also creates aws_ssm_parameter.db_url, the data source reads the value that existed before the apply, so the apply that changes the parameter still wires the old value into c and a second apply is required to converge; on a fresh stack, where the parameter does not exist yet, the plan fails outright because there is nothing to read. A saved plan file (-out) freezes the read entirely, so a change made between plan and apply will not appear until you re-plan. Use the resource reference whenever the parameter lives in the same config; reach for a data source only to read something another team owns, and check whether a real ordering need forces depends_on.

value, name, arn β€” and id, url, arn

Most wiring bugs are not about deploy-time vs runtime; they are about picking the wrong attribute of the right resource. Every AWS resource exposes several strings that look interchangeable and are not.

SSM parameter Meaning Feeds
name the path /app/db_url an env var your code re-reads
value the string itself an env var used directly
arn arn:aws:ssm:<region>:<acct>:parameter/app/db_url an IAM policy Resource

Two failures recur. First, feeding an ARN where a value is expected:

# BROKEN: the SDK will try to connect to a string that starts with "arn:".
variables = { DB_URL = aws_ssm_parameter.db_url.arn }

# FIXED
variables = { DB_URL = aws_ssm_parameter.db_url.value }

Second, feeding a URL where an ARN is expected β€” IAM identifies resources by ARN, so a URL in Resource never names the queue:

# BROKEN: this is the queue URL, not its ARN; it never names the queue.
Resource = aws_sqs_queue.jobs.url

# FIXED
Resource = aws_sqs_queue.jobs.arn

For SQS, id and url hold the same string (the queue URL) β€” the SDK sends to it, so it belongs in the function's env var; the arn belongs in policies. There is no rule that id is an ARN; it is whatever the provider chose, so read the attribute list rather than guessing.

Implicit edges, and depends_on as the fallback

Every reference is an edge in the dependency graph, which is why deploy-time wiring orders your stack for free. A runtime lookup removes that edge β€” nothing in the function's config mentions the table β€” so Terraform no longer knows the table must exist first. That is exactly when the depends_on meta-argument earns its place:

resource "aws_lambda_function" "worker" {
  function_name = "worker"
  # ... role, filename ...
  environment {
    variables = { TABLE_PARAM = aws_ssm_parameter.table_name.name }
  }

  # The function reads the table name from SSM at cold start;
  # no argument references the table, so state the ordering by hand.
  depends_on = [aws_dynamodb_table.jobs]
}

Another honest use is an S3 bucket notification that must not be created until the Lambda permission letting S3 invoke the function exists β€” the two resources share no attribute, but order matters.

⚠️ Reach for depends_on sparingly. It hides the real data flow (a reader cannot see why), it waits for the whole dependency to finish rather than just the part you need, and it can serialize resources that could otherwise apply in parallel. Prefer an explicit reference that carries a value; use depends_on only when you genuinely cannot.

Guided repair: replace hardcoded strings with references

This function and policy pass literals instead of references. The table is named jobs; the queue URL is correct; the topic ARN is correct.

resource "aws_lambda_function" "worker" {
  function_name = "worker"
  # ... role, filename ...
  environment {
    variables = {
      QUEUE_URL  = "https://sqs.us-east-1.amazonaws.com/123456789012/jobs"
      TABLE_NAME = "jobs-v1"
      TOPIC_ARN  = "arn:aws:sns:us-east-1:123456789012:alerts"
    }
  }
}

resource "aws_iam_policy" "worker" {
  name = "worker"
  policy = jsonencode({
    Statement = [{
      Effect   = "Allow"
      Action   = "dynamodb:PutItem"
      Resource = "arn:aws:dynamodb:us-east-1:123456789012:table/jobs"
    }]
  })
}

Replace each literal with a reference to aws_sqs_queue.jobs, aws_dynamodb_table.jobs, and aws_sns_topic.alerts. Then state which resources the next plan shows as ~ and which new graph edges appear.

Check your answer
variables = {
  QUEUE_URL  = aws_sqs_queue.jobs.url
  TABLE_NAME = aws_dynamodb_table.jobs.name
  TOPIC_ARN  = aws_sns_topic.alerts.arn
}
# policy Resource
Resource = aws_dynamodb_table.jobs.arn

The trap is the plan. Replacing a literal with a reference changes the desired value only if they differ. QUEUE_URL and TOPIC_ARN matched reality, so those attributes produce no diff β€” the plan looks unchanged for them. TABLE_NAME pointed at jobs-v1 while the table is jobs, so the function shows ~ on that variable. The new edges are the real win and they appear regardless of the diff: function β†’ queue, function β†’ table, function β†’ topic, and policy β†’ table. Those edges exist even when the plan is a no-op, which is precisely why hardcoding "works" right up until it doesn't.

Pitfalls

⚠️ Hardcoding ARNs or names breaks the moment you build a second environment: the account ID and region are wrong there. It also drifts silently β€” if Terraform replaces the queue (say, after a name change), the live URL changes but the function keeps the stale hand-typed one and every send to the stale URL fails, because that queue no longer exists (QueueDoesNotExist). The plan will not warn you; the literal never referenced the queue, so there is no edge to surface the change.

⚠️ Cycles. Two resources that reference each other's attributes have no valid build order:

resource "aws_security_group" "a" {
  name = "a"
  ingress {
    from_port       = 5432
    to_port         = 5432
    protocol        = "tcp"
    security_groups = [aws_security_group.b.id] # a depends on b
  }
}

resource "aws_security_group" "b" {
  name = "b"
  ingress {
    from_port       = 5432
    to_port         = 5432
    protocol        = "tcp"
    security_groups = [aws_security_group.a.id] # b depends on a  -> cycle
  }
}

Terraform refuses, already at terraform validate, with an Error: Cycle: diagnostic that lists both aws_security_group.a and aws_security_group.b. The fix is to move the rules into separate aws_vpc_security_group_ingress_rule resources so the groups no longer reference each other; each rule then depends on both groups with no back edge. If you hit a cycle, look for the pair of references that point at each other and break one of them β€” do not paper over it with depends_on, which cannot resolve a cycle either.

Wire services by reference so the value and the ordering edge travel together; keep runtime lookups for values that must change without a redeploy, and add depends_on only when a genuine ordering need has no reference to express it.

Judging Which Changes Are Safe: Update, Replace, or Destroy

A team changes one line β€” the name of a DynamoDB table β€” and the plan ends with Plan: 1 to add, 0 to change, 1 to destroy. The count is accurate. It is also the wrong thing to have looked at, because that single destroy takes every row in the table with it. Reading a plan safely means reading its body, not its summary, and classifying each change before you type apply.

The plan is Terraform's diff between three things: the configuration (the .tf HCL files you wrote), the state (Terraform's record of the resources it last applied), and the real cloud as reported by the provider (the plugin that talks to AWS).

One attribute decides the verdict

Every attribute in the provider's schema is either updatable in place (AWS exposes an API call that mutates the existing resource) or immutable (no such API exists β€” the only path is to build a new resource and delete the old one). Immutable attributes are what the plan annotates with # forces replacement, and when one of them changes, the resource header flips from ~ to -/+.

change an attribute in the configuration
        ↓
is "# forces replacement" printed on that attribute's line?
        ↓
no   β†’  '~'    update in place: same id, same arn, same data
        ↓
yes  β†’  '-/+'  destroy, then create: a new object (new id and arn, unless both are built from a name that stays the same), empty of everything the old one held

The arn (Amazon Resource Name) is AWS's globally unique identifier for a resource. Your job is not to memorize which attributes are immutable β€” it is to read the annotation, and, when you want to know before writing the configuration, to look for "Forces replacement" or "Forces new resource" in the provider docs for that argument, while treating the docs as incomplete: the aws_dynamodb_table page does not mark name that way, yet changing it forces replacement.

Three plans, three verdicts

1. In-place update β€” verdict: safe. Changing a Lambda function's environment variable, or a DynamoDB table's capacity on a table using billing_mode = "PROVISIONED", calls an update API. Nothing is recreated:

resource "aws_lambda_function" "worker" {
  function_name = "worker"
  # ... role, handler, filename omitted
  environment {
    variables = {
      LOG_LEVEL = var.log_level # "info" -> "debug"
    }
  }
}
  # aws_lambda_function.worker will be updated in-place
  ~ resource "aws_lambda_function" "worker" {
        id = "worker"
      ~ environment {
          ~ variables = {
              ~ "LOG_LEVEL" = "info" -> "debug"
            }
        }
        # (12 unchanged attributes hidden)
    }

Plan: 0 to add, 1 to change, 0 to destroy.

2. Harmless replacement β€” verdict: safe, but only after you check dependents. An SQS queue's name is immutable, so a rename produces -/+:

  # aws_sqs_queue.jobs must be replaced
-/+ resource "aws_sqs_queue" "jobs" {
      ~ arn  = "arn:aws:sqs:us-east-1:111122223333:jobs" -> (known after apply)
      ~ id   = "https://sqs.us-east-1.amazonaws.com/111122223333/jobs" -> (known after apply)
      ~ name = "jobs" -> "jobs-v2" # forces replacement
        # (5 unchanged attributes hidden)
    }

This is safe if nothing else in the configuration consumes jobs' arn, id or url, and if the queue holds nothing you need β€” a fresh queue with no producers yet, for example. "Nothing depends on it" is a claim you verify, not assume.

3. Destructive replacement β€” verdict: unsafe. Rename the table instead:

  # aws_dynamodb_table.orders must be replaced
-/+ resource "aws_dynamodb_table" "orders" {
      ~ arn  = "arn:aws:dynamodb:us-east-1:111122223333:table/orders" -> (known after apply)
      ~ name = "orders" -> "orders-v2" # forces replacement
        # (9 unchanged attributes hidden)
    }

Same shape as the queue, radically different consequence: the old table is deleted with its rows, and the new table starts empty. The API Gateway stage in this stack is the other classic case β€” a stage's name is immutable, and a named stage's name is the first path segment of every URL your callers hit (.../prod/jobs), so replacing the stage silently moves the endpoint. (The $default stage in this stack is served from the base URL with no stage segment.)

Lifecycle meta-arguments: exact effects

A lifecycle block is a nested block inside a resource that changes how Terraform performs an action. Three meta-arguments matter for replacements:

Meta-argument Exact effect on the plan
prevent_destroy A plan that would destroy the resource fails with an error
create_before_destroy New resource is created before the old one is removed
ignore_changes Diff for the listed attributes is suppressed

prevent_destroy is a tripwire, not a fix: it stops the apply, then leaves you deciding what to do. It cannot protect data if you remove the meta-argument to get past it, and it does not stop a destroy once the resource block itself is gone from the configuration, which includes a block renamed or moved into a module without a moved block: the setting goes with the block.

create_before_destroy flips the default order β€” Terraform normally destroys first, which is why a plain -/+ has a downtime window. It gives you zero-downtime replacement within provider limits: it cannot work when the replacement must keep a name that is unique per account and Region, because the new one cannot be created while the old still holds the name: a DynamoDB table replaced because its hash_key changed, for example, fails with ResourceInUseException. A rename is different: the new name is free, so creating first works.

ignore_changes = [tags] tells Terraform to stop comparing that attribute between configuration and state. It is appropriate for a value something else owns. It is dangerous because it also suppresses the diff that was reporting drift β€” the cloud diverging from what the state records β€” so an out-of-band change becomes invisible forever after.

resource "aws_dynamodb_table" "orders" {
  name         = "orders"
  billing_mode = "PAY_PER_REQUEST"
  hash_key     = "order_id"

  lifecycle {
    prevent_destroy = true # this plan errors instead of deleting the table
  }
}

What a replacement actually costs

Four consequences, in rough order of how badly they bite: irreversible data loss (the table), a downtime window (destroy-then-create unless you set create_before_destroy), longer apply times (a new resource plus any data migration), and blast radius β€” the reach of the change into other resources.

Blast radius is where the line-by-line reading fails. Read the plan as a graph:

aws_sqs_queue.jobs  replaced  (new arn)
        ↓ referenced by
aws_ssm_parameter.jobs_url                  β†’ '~' update in place
aws_lambda_event_source_mapping.jobs        β†’ '-/+' replaced

The SSM parameter merely stores a new string, so it updates in place. The event source mapping pins the queue's arn in an argument that is itself immutable, so it gets replaced too β€” a second -/+ you did not write. Five dependents do not automatically mean five replacements; most become ~ updates. But each one must be checked, because the ones that do cascade are invisible until you look. The dependency graph that produces this ordering was covered in "Reading the Plan of a Whole Stack Before You Change Anything".

Practice: classify three excerpts

Read each excerpt, then decide safe-update, safe-replace, or unsafe. For the unsafe one, choose a mitigation β€” a moved block, prevent_destroy, create_before_destroy, or a snapshot first β€” and state what it changes about the plan.

Excerpt A

  # aws_dynamodb_table.orders will be updated in-place
  ~ read_capacity  = 5 -> 25
  ~ write_capacity = 5 -> 25
Plan: 0 to add, 1 to change, 0 to destroy.

Excerpt B

  # aws_sqs_queue.dlq must be replaced
-/+ resource "aws_sqs_queue" "dlq" {
      ~ name = "orders-dlq" -> "orders-dlq-v2" # forces replacement
    }

Context: the DLQ is empty, and no other resource in this configuration references its arn or url.

Excerpt C

  # aws_dynamodb_table.orders must be replaced
-/+ resource "aws_dynamodb_table" "orders" {
      ~ name = "orders" -> "orders_v2" # forces replacement
    }
  ~ environment { ~ variables = { ~ TABLE_NAME = "orders" -> "orders_v2" } }
Check your answer

A β€” safe-update. ~ with no # forces replacement means the provider calls an update API. The table keeps its arn and its rows; capacity changes take effect without a recreate.

B β€” safe-replace. A -/+ header, and the named attributes (arn, url) are either unreferenced or read at runtime from Parameter Store rather than wired in by reference. The empty DLQ loses nothing. create_before_destroy would work here, because orders-dlq-v2 is a different name from orders-dlq, but it adds nothing: the DLQ is empty and nothing points at it.

C β€” unsafe. The table's name is immutable, so the old table and all its rows are destroyed. A snapshot first β€” a DynamoDB on-demand backup taken before the apply β€” is the mitigation that changes the recoverability of the plan without changing the plan itself. prevent_destroy is the complementary guardrail: it makes this plan fail outright, which is usually the correct outcome until you have decided that the new physical table and a data migration are genuinely what you want. A moved block does not apply here: it fixes a change of resource address (the label aws_dynamodb_table.orders), not a change of an attribute like name. Using moved against an attribute-driven -/+ changes nothing in the plan.

Pitfalls

⚠️ Reading only the summary. Plan: 0 to add, 1 to change, 0 to destroy is a count of resource instances, and a -/+ is counted as one destroy and one add elsewhere. Scan the body for -/+ and forces replacement before you look at the totals.

⚠️ Using ignore_changes to silence a diff. If a plan keeps proposing to update an attribute that someone changed in the console, that diff is a report, not a nuisance. Suppressing it stops Terraform from correcting the drift and stops you from seeing it.

⚠️ Assuming a replacement is safe because it was safe last time. The queue rename that was harmless in staging is not harmless in the stack where a Lambda event source mapping pins its arn. Safety is a property of the resource and its consumers β€” which is exactly why the plan has to be read as a graph.

When the fix is a genuine rename rather than an attribute change, the repair moves from the resource's attributes into state editing territory β€” moved and import blocks β€” which "Evolving the Stack Without Destroying It" treats in full.

Evolving the Stack Without Destroying It: Move, Import and State Boundaries

Terraform does not recognise a resource by its name argument or by its ARN. It recognises it by its address β€” the dotted path you write on the left of a block, like aws_lambda_function.worker or module.app.aws_lambda_function.worker. State is a map from addresses to real cloud resources, and that map is what lets Terraform say "update this Lambda" instead of "create a new Lambda." So the moment you reorganise the text β€” rename a block, wrap a group of resources in a module, move a file into a subdirectory β€” you have changed the address, and Terraform's default reading of the change is brutal: the old address is gone (destroy), a new address has appeared (create). The code describes the same Lambda, the same queue, the same DynamoDB table with all its rows, but the plan shows a destroy (-) for the old address and a create (+) for the new one.

This section covers the three tools that stop that destroy β€” moved, import, and state boundaries β€” plus the manual fallback when you must touch state by hand.

Addresses Are Identity: Why a Rename Looks Like a Destroy

Say you start with a Lambda inline in the root of your configuration:

resource "aws_lambda_function" "worker" {
  function_name = "order-worker"
  role          = aws_iam_role.worker.arn
  # ... handler, runtime, filename
}

Its address is aws_lambda_function.worker. Later you extract it into a module so the stack reads more cleanly, and inside that module the same resource now has the address module.app.aws_lambda_function.worker. Nothing about function_name = "order-worker" changed. But terraform plan compares two things: state (which knows aws_lambda_function.worker β†’ the real, live function) and configuration (which now only declares module.app.aws_lambda_function.worker). Neither side matches, so:

  # aws_lambda_function.worker will be destroyed
  - resource "aws_lambda_function" "worker" { ... }

  # module.app.aws_lambda_function.worker will be created
  + resource "aws_lambda_function" "worker" { ... }

Plan: 1 to add, 0 to change, 1 to destroy.

The live function is deleted and a fresh one created. For a stateless Lambda that costs a cold start; for a queue or a table it means lost messages or lost rows, and for an API stage it changes the URL callers already use β€” the same cascade we classified in "Judging Which Changes Are Safe."

Config rename
     ↓
Address changes in configuration only
     ↓
State still holds the old address
     ↓
Default plan: destroy old + create new
     ↓
moved block re-keys state -> no-op plan

moved Blocks: Telling Terraform the Resource Didn't Change

A moved block is a declaration that an address you used to have is the same thing as an address you now have. Introduced in Terraform 1.1, it exists for exactly this problem:

moved {
  from = aws_lambda_function.worker
  to   = module.app.aws_lambda_function.worker
}

Place it where the unit is applied β€” typically next to the module you extracted. The plan now reads:

Terraform will perform the following actions:

  # aws_lambda_function.worker has moved to module.app.aws_lambda_function.worker
    resource "aws_lambda_function" "worker" {
        id            = "order-worker"
        function_name = "order-worker"
        # ...
    }

Plan: 0 to add, 0 to change, 0 to destroy.

No - line, no + line. Terraform re-keyed the state entry from old address to new and recognised the resource as the same object. That is the no-op plan you want: 0 to add, 0 to change, 0 to destroy. Apply once and state records the new address. A moved block only fires if the addresses match reality: from must be the address currently in state, to the address in your new configuration. Get either wrong and the plan still destroys and creates.

⚠️ Pitfall, front and center: order matters. If you rename first and apply before adding the moved block, you have performed the destroy for real, and no moved block brings it back. Add the moved block and the renamed configuration in the same change, then read the plan before applying. The plan is your proof nothing will be destroyed.

import Blocks: Adopting Resources Terraform Didn't Create

The mirror-image problem: a resource exists in AWS but was never created by Terraform, so there is no state entry. Terraform's default is to try to create it β€” and for a table with a fixed name, that fails or duplicates. An import block attaches an existing cloud resource (one the provider already knows about, but which Terraform does not manage) to a configuration address. Introduced in Terraform 1.5, it is declarative:

import {
  to = aws_dynamodb_table.orders
  id = "orders-table" # the resource's real identifier in AWS
}

resource "aws_dynamodb_table" "orders" {
  name         = "orders-table"
  billing_mode = "PAY_PER_REQUEST"
  hash_key     = "order_id"
  attribute {
    name = "order_id"
    type = "S"
  }
}

The id is the identifier the provider uses to find the resource β€” for DynamoDB, the table name. On the first plan, Terraform reports it will import that resource, then shows the diff between your configuration and the attributes read back from AWS. That follow-up diff is the part people skip: importing does not write your config for you, it records what is really there. If your billing_mode or an attribute block differs from reality, the plan shows an update, and you decide whether the config or the cloud is right. terraform apply performs the import (a state write) and then reconciles per the plan.

The older imperative form runs once per resource:

terraform import aws_dynamodb_table.orders orders-table

It does the same state write without the declarative safety net (there is no plan to review first), and it requires the resource block to exist already. Generating a starting configuration with terraform plan -generate-config-out=generated.tf (Terraform 1.5+, still marked experimental) works only with import blocks. Either form: an import is a state operation, not a cloud operation. It creates, modifies and deletes nothing in AWS by itself β€” only the apply that follows can, and only to reconcile attributes.

Where to Draw State Boundaries

A state boundary is a decision about which resources share a state file β€” really a blast-radius decision. The default, one state for the whole tightly coupled stack, is often right: one plan to read, one apply to run (not a transaction: if it fails midway, whatever was already created stays and is recorded in state), and Terraform orders creation for you because it sees every reference.

You split when resources have genuinely different change frequencies or lifecycles. The classic split puts long-lived data (DynamoDB tables, S3 buckets) in one state and frequently redeployed compute (Lambda, ECS services, the API) in another, so a bad compute deploy cannot touch the data state. The cost is real and paid in three places: cross-state references (the compute config reads the table ARN through a terraform_remote_state data source or a plain data "aws_dynamodb_table" lookup), no single apply (two applies, two chances to leave a half-updated stack), and manual ordering β€” the data stack must apply before the compute stack that consumes its outputs.

Decision One state Split states
Apply One apply (not atomic) Several applies
Cross-refs Direct reference Output or data source
Blast radius Whole stack Per boundary
Ordering Terraform infers You manage

Treat it as a heuristic, not a law: split only where the lifecycle difference is real, because every boundary adds a reference you must keep wired.

The Manual Fallback: terraform state mv and state rm

When you cannot or do not want to re-apply, the imperative commands edit the state map by hand:

terraform state mv aws_lambda_function.worker module.app.aws_lambda_function.worker
terraform state rm aws_dynamodb_table.orders

terraform state mv renames an entry β€” the manual equivalent of a moved block, but it lives in your shell history, not the configuration, so the next person to run plan has no record of why. terraform state rm removes the entry entirely, which leaves the resource unmanaged: Terraform forgets it exists, and if the resource block is still in your configuration the next apply tries to create it again: for a table or function whose name is taken that create fails with an already-exists error, and for resources whose names need not be unique you get a real duplicate. Use state rm when you are genuinely handing a resource off (for instance, importing it into another state), not to "clean up." Since Terraform 1.7 the declarative alternative is a removed block with lifecycle { destroy = false }, which shows the removal in a plan before anything happens.

⚠️ Two risks. First, editing state while a plan is pending invalidates that plan β€” the plan was computed against the old state contents, so Terraform refuses to apply that saved plan (Error: Saved plan is stale); always re-run plan after any state command. Second, these commands touch state only; like imports, they never touch the cloud, which is why they feel safe and can still leave you a duplicate on the next apply.

Practice: Turn a Destroy-and-Create Plan Back into a No-op

You open a plan and see:

  # aws_dynamodb_table.customer_orders will be created
  + resource "aws_dynamodb_table" "customer_orders" {
      + arn  = (known after apply)
      + name = "customer_orders"
    }

  # aws_dynamodb_table.orders will be destroyed
  # (because aws_dynamodb_table.orders is not in configuration)
  - resource "aws_dynamodb_table" "orders" {
      - arn  = "arn:aws:dynamodb:us-east-1:...:table/orders" -> null
      - name = "orders" -> null
    }

Plan: 1 to add, 0 to change, 1 to destroy.

The table was created by Terraform and is in state under its old address; someone renamed the block from orders to customer_orders in the config, and changed the name argument too. Decide: moved block or import? Then write the exact block, and say what you check after apply.

Check your answer

This is a moved block, not an import. The resource is already in state β€” the problem is only that its address changed, which is exactly what moved fixes. An import would be wrong: it is for resources Terraform has never tracked, and this table is tracked.

moved {
  from = aws_dynamodb_table.orders
  to   = aws_dynamodb_table.customer_orders
}

But the name argument changed too. With only the moved block added, the next plan reads # aws_dynamodb_table.customer_orders must be replaced, # (moved from aws_dynamodb_table.orders), with ~ name = "orders" -> "customer_orders" # forces replacement. Renaming the block address is free; renaming the actual DynamoDB table is not β€” name is immutable, so changing it forces a replace regardless of the moved block. To get a true no-op you must also revert the name argument back to "orders". The moved block fixes the address; only leaving name alone keeps the data. If renaming the table is genuinely required, the honest plan is -/+, and you snapshot first, as covered in "Judging Which Changes Are Safe."

After apply (with name reverted), check three things: the summary says 0 to add, 0 to change, 0 to destroy; the table's CreationDateTime (and TableId) in AWS are unchanged (a table re-created under the same name keeps the same ARN, so only these show that no new table was made); and the row count is what you expect, e.g. aws dynamodb scan --table-name orders --select COUNT. The block being called customer_orders while the table is orders is fine β€” the block name is a local label; the name argument is what exists in AWS.

That generalises: moved and import correct identity, not values. Moving an address never changes an attribute, and importing never changes a resource in the cloud. If a destroy still appears after you add the block, either the addresses don't match or a real attribute change is doing the destroying β€” read the # forces replacement line to tell which.

This is the last piece before you review a full multi-service change end to end in "Independent Transfer: Reviewing a Multi-Service Change End to End," where renames, imports and state strategy all appear in one change request.

Independent Transfer: Reviewing a Multi-Service Change End to End

Sections 1–4 gave you one move at a time: read a plan, wire by reference, classify a change, refactor without destroying. This section hands you a pull request with no scaffolding. Nothing here is new machinery β€” the difficulty is choosing which of those tools apply, and defending the choice in writing, the way a reviewer would read it on a real PR.

The change request

The stack you have been building lives in a single prod directory: an SSM parameter, an IAM role, a DynamoDB table jobs, an SQS queue jobs with a dead-letter queue, a Lambda worker with an event source mapping, and an HTTP API with stage prod. The pull request does three things:

  1. Extracts the table, queue, function, event source mapping and stage into a module "app", re-addressing them (with moved blocks).
  2. Adds var.environment, prefixes the name argument on the table and queue with it, and renames the stage.
  3. Adds a staging/ directory with its own backend key (the path of the state object inside the bucket β€” what makes two directories two states).

You have two deliverables to write before you run anything.

Step 1 β€” Choose a state boundary and defend it

Three candidate layouts, and the choice is not stylistic:

Strategy Blast radius Drift risk Runtime difference
Per-directory root, own backend key staging plan cannot reach prod state drift stays in one env full: names, sizes, lifecycle
Workspaces one config, one backend, state per workspace wrong terraform workspace select hits prod only values you thread through terraform.workspace
Separate state per env (same shape) isolated isolated full

The deciding constraint in this stack is that the environments must differ structurally, not just in values. prevent_destroy and the other lifecycle settings are written literally in the block, so a single shared config cannot give prod prevent_destroy = true on the table and staging false β€” you would need conditionals HCL does not offer there. Separate directories keep that guardrail honest, provided the table block itself is written in each environment's directory: a lifecycle literal inside a shared module is the same for every caller.

Your task: in three sentences, pick a layout and justify it on blast radius, drift risk and runtime difference. Success criteria: your justification names one concrete resource whose config must differ between environments, and one failure a wrong command could cause.

Check your answer

Model answer, to compare against β€” not to copy. "I'd use a shared child module called from prod/ and staging/, each with its own backend key, and declare the table in each directory rather than in the module. Prod's table keeps prevent_destroy = true as a literal in its own directory, which a workspace could not express, and staging can run 5 RCU where prod runs 20. With workspaces, terraform workspace select left on default and a -auto-approve apply would reach prod state from a staging habit."

Step 2 β€” Deliverable one: the change plan

Ask for two things, in writing: the classification of every affected resource, and the guardrail attached to each.

Task: write a table with one row per affected resource address and four columns β€” action (create / update / replace / move-only), what causes it, dependency-order consequence, guardrail. Then hand it to a reviewer before touching prod.

Check your answer
Address Action Cause Consequence Guardrail
module.app.aws_lambda_function.worker move-only address changed none on AWS moved from aws_lambda_function.worker
module.app.aws_lambda_event_source_mapping.worker move-only address changed none on AWS moved
module.app.aws_dynamodb_table.jobs move-only address changed; name unchanged none on AWS moved + prevent_destroy = true (inside the shared module it guards staging's table too; to differ per environment, declare the table per directory as in Step 1)
module.app.aws_sqs_queue.jobs move-only address changed none on AWS moved
module.app.aws_apigatewayv2_stage.default move-only address changed none on AWS moved + keep name = "prod"

The point of the exercise is row three: the only way the module extraction is safe is if every name argument stays byte-identical, which means var.environment may prefix names in staging/ but must not change them in prod/. If you wrote "replace" on any row, re-read Judging Which Changes Are Safe on which attributes are immutable.

Step 3 β€” Deliverable two: review the plan

A colleague parameterized the names globally and ran terraform plan in prod/. The output below lists every resource that changes, trimmed to the lines that matter (other computed attributes such as id and url also turn (known after apply)):

Terraform will perform the following actions:

  # module.app.aws_apigatewayv2_stage.default must be replaced
-/+ resource "aws_apigatewayv2_stage" "default" {
      ~ invoke_url = "https://a1b2c3.execute-api.us-east-1.amazonaws.com/prod" -> (known after apply)
      ~ name       = "prod" -> "prod-v2" # forces replacement
        # (8 unchanged attributes hidden)
    }

  # module.app.aws_dynamodb_table.jobs must be replaced
-/+ resource "aws_dynamodb_table" "jobs" {
      ~ arn  = "arn:aws:dynamodb:us-east-1:111122223333:table/jobs" -> (known after apply)
      ~ name = "jobs" -> "prod-jobs" # forces replacement
        # (9 unchanged attributes hidden)
    }

  # module.app.aws_lambda_event_source_mapping.worker must be replaced
-/+ resource "aws_lambda_event_source_mapping" "worker" {
      ~ event_source_arn = "arn:aws:sqs:us-east-1:111122223333:jobs" -> (known after apply) # forces replacement
        # (6 unchanged attributes hidden)
    }

  # module.app.aws_lambda_function.worker will be updated in-place
  ~ resource "aws_lambda_function" "worker" {
      ~ environment {
          ~ variables = {
              ~ "QUEUE_URL"  = "https://sqs.us-east-1.amazonaws.com/111122223333/jobs" -> (known after apply)
              ~ "TABLE_NAME" = "jobs" -> "prod-jobs"
            }
        }
      ~ memory_size = 512 -> 1024
        # (6 unchanged attributes hidden)
    }

  # module.app.aws_sqs_queue.jobs must be replaced
-/+ resource "aws_sqs_queue" "jobs" {
      ~ arn  = "arn:aws:sqs:us-east-1:111122223333:jobs" -> (known after apply)
      ~ name = "jobs" -> "prod-jobs" # forces replacement
        # (10 unchanged attributes hidden)
    }

Plan: 4 to add, 1 to change, 4 to destroy.

Task: flag the disqualifying lines. For each, state the exact attribute that forces replacement, whether the row is a root cause or a consequence, and the minimal config edit that removes the line. Then name the rows you would not edit, and say why.

Check your answer

Model review β€” compare, don't copy.

  1. module.app.aws_dynamodb_table.jobs β€” disqualifying, root cause. Cause: name changing to "prod-jobs", annotated # forces replacement. A DynamoDB table's name is immutable, so the old table is destroyed and its rows are gone; backups help only within retention and never make the address match again. Minimal edit: in prod/, keep name = "jobs" (or a prod-only variable whose value is "jobs"). The prefix belongs in staging/, where the table is new and the name change is a create.

  2. module.app.aws_sqs_queue.jobs β€” disqualifying, root cause. Same cause, same # forces replacement. Queues hold messages that consumers have not yet read; destroying the queue discards them, and any producer holding the old URL starts failing its sends. Minimal edit: same β€” do not rename the prod queue.

  3. module.app.aws_lambda_event_source_mapping.worker β€” disqualifying, but a consequence. Cause: event_source_arn changing, because it references the replaced queue's ARN. Do not edit this resource; fix row 2 and it disappears. Proposing a moved block here would be wrong: the resource is not re-addressed, it depends on a replaced ARN.

  4. module.app.aws_apigatewayv2_stage.default β€” disqualifying, root cause. Cause: name "prod" -> "prod-v2". The stage name is the last path segment of the invoke URL, so the URL callers already hold changes from /prod to /prod-v2 and every caller gets a 404 β€” the API id is untouched, which is exactly why this one is easy to miss. Minimal edit: keep name = "prod"; give staging/ its own stage name.

  5. module.app.aws_lambda_function.worker β€” not a flag. memory_size 512 β†’ 1024 is an in-place update. The QUEUE_URL variable moving to (known after apply) is a consequence of row 2, and TABLE_NAME changing to "prod-jobs" a consequence of row 1, so both resolve without an edit.

The cascade, in one view:

queue name changes
        ↓ forces replacement
aws_sqs_queue.jobs  -/+   (unread messages discarded)
        ↓ arn becomes (known after apply)
event_source_mapping  -/+   (re-created)
        ↓ url reference in env vars
lambda worker  ~            (QUEUE_URL rewritten)

Note the summary line is honest but not diagnostic: 4 to add, 4 to destroy tells you four resources are being re-created but not that two of them hold data (the table's rows, the queue's unread messages). The body tells you. (A real plan body never lists resources that have no changes, so nothing else is hiding in it.)

What a green apply does not prove

Apply complete! means Terraform reconciled configuration and state. It does not mean the change was safe, because the plan was the only artifact that could have told you. Check these invariants against the environment, not the exit code:

Invariant How to check
The table still holds its rows item count before vs after
The invoked URL is unchanged compare caller config to the stage's invoke_url
Nothing was replaced that should have updated scan the apply log for -/+
The queue still has its DLQ wiring redrive policy on the live queue

The deeper per-service testing routine β€” smoke tests per Lambda, API contract tests, review workflow and promotion gates β€” belongs to its own lesson; this section ends with the state and plan discipline that has to be right first.

Closing checklist

  • I can say what a plan will do to state (Terraform's record of what it last applied) and to the cloud before running apply.
  • I can spot forces-replacement from one annotated line rather than the summary counts.
  • I wire services by reference (a value interpolated from another resource) rather than hardcoded ARNs or URLs, so a second environment does not require editing literals.
  • I refactor with moved and import and can say what each one leaves untouched: moved changes only the address in state; import adds an existing resource to state without touching the cloud.
  • I can pick a state boundary and defend it on blast radius, drift risk and runtime difference β€” and I know which guardrails that choice makes possible.