Terragrunt: One Configuration, Many Environments

Terragrunt as a thin wrapper that runs Terraform or OpenTofu for you, for a developer who already knows plan and apply, S3 remote state and modules. The problem it solves: the same backend, provider and module wiring repeated in every environment directory. Units (one terragrunt.hcl per state), the terraform block with a pinned module source, inputs, and include with a shared root file (root.hcl) found with find_in_parent_folders; generating backend.tf and provider.tf with remote_state and generate blocks, and a distinct S3 state key per unit from path_relative_to_include; locals and read_terragrunt_config for account, Region and environment values; dependency blocks that pass one unit's outputs to another, with mock_outputs so plan works before the upstream unit exists; running many units in dependency order with run --all, reading the run queue, and why apply across units is riskier than a single plan; the .terragrunt-cache directory and what Terragrunt actually executes; stacks (terragrunt.stack.hcl) in brief. Use the current Terragrunt CLI and file names, and name the older forms a reader still meets in existing repositories (run-all, a root terragrunt.hcl, TERRAGRUNT_ environment variables instead of TG_). AWS authentication with IAM Identity Center profiles locally and OIDC role assumption in CI, no static keys. End with when plain Terraform with directories and tfvars is enough and Terragrunt is not worth adding.

Last generated

Lesson 7 of 7 available15 practice questions

SPACED REPETITION Β· 15 practice questions

Make this lesson stick.

Try 3 questions now. No account needed. Sample answers aren't saved.

From Copied Directories to a Unit: Your First terragrunt.hcl

Open your dev, staging and prod directories side by side and diff them. In many Terraform-on-AWS repositories you will find three copies of backend.tf, three copies of provider.tf, and three main.tf files calling the same module, with a few strings changed. Every backend fix then has to be made three times, and one copy gets forgotten.

Terragrunt is a thin wrapper that runs Terraform (or OpenTofu) for you. It does not replace state, plans or modules. It writes the repeated files on your behalf and then runs the commands you already know. The building block is the unit: one directory containing one terragrunt.hcl, which maps to exactly one Terraform state. This section converts one directory into a unit and traces what Terragrunt really executes.

The worked problem: what actually differs?

Here is the dev directory. It calls a module sqs-with-dlq that creates a queue, a dead-letter queue (DLQ) and the redrive policy linking them. Lines marked DIFFERS change per environment; everything else is copy-paste.

# dev/orders-queue/backend.tf
terraform {
  backend "s3" {
    bucket  = "acme-tf-state"                       # copy-paste
    key     = "dev/orders-queue/terraform.tfstate"  # DIFFERS
    region  = "eu-west-1"                           # copy-paste (the bucket's Region)
    encrypt = true                                  # copy-paste
  }
}

# dev/orders-queue/provider.tf
provider "aws" {
  region = "eu-west-1"                              # DIFFERS (prod uses eu-central-1)
}

# dev/orders-queue/main.tf
module "queue" {
  source                     = "git::git@github.com:acme/tf-modules.git//sqs-with-dlq?ref=v1.4.0"  # copy-paste
  queue_name                 = "orders-dev"         # DIFFERS
  max_receive_count          = 3                    # DIFFERS (retry count before a message goes to the DLQ)
  visibility_timeout_seconds = 60                   # copy-paste
}

Out of roughly a dozen meaningful lines, only four differ: the state key, the provider Region, the queue name and the retry count. The rest is repetition that Terragrunt is meant to absorb.

The unit file

The new dev directory holds only this file. Read each block before moving on.

# live/dev/orders-queue/terragrunt.hcl
terraform {
  # Pinned to a Git tag. The double slash selects a sub-folder of the repo.
  source = "git::git@github.com:acme/tf-modules.git//sqs-with-dlq?ref=v1.4.0"
}

# Tells Terragrunt to write backend.tf for us.
remote_state {
  backend = "s3"
  generate = {
    path      = "backend.tf"
    if_exists = "overwrite_terragrunt" # replace the file only if Terragrunt wrote it
  }
  config = {
    bucket  = "acme-tf-state"
    key     = "dev/orders-queue/terraform.tfstate"
    region  = "eu-west-1"
    encrypt = true
  }
}

# Tells Terragrunt to write provider.tf. No access keys: credentials come from the environment.
generate "provider" {
  path      = "provider.tf"
  if_exists = "overwrite_terragrunt"
  contents  = <<EOF
provider "aws" {
  region = "eu-west-1"
}
EOF
}

# Values for the module's variables.
inputs = {
  queue_name                 = "orders-dev"
  max_receive_count          = 3
  visibility_timeout_seconds = 60
}

The ?ref=v1.4.0 is the safety feature. If every environment pointed at a shared local path such as ../../modules/sqs-with-dlq, or at a floating branch like ?ref=main, one module edit would change the next plan of every environment at once. A tag lets you bump dev first, read the plan, and promote the tag to staging and prod later. Tags can be moved by force-pushing, so treat them as immutable by convention.

What Terragrunt runs on your behalf

When you run terragrunt plan, it does this:

terragrunt plan  (in live/dev/orders-queue)
  ↓
1. Download the module at ?ref=v1.4.0 into .terragrunt-cache/<hash>/<hash>/sqs-with-dlq/
  ↓
2. Copy this directory's own files over it
  ↓
3. Write the generated backend.tf and provider.tf next to the module files
  ↓
4. Export each input as a TF_VAR_ environment variable (TF_VAR_queue_name=orders-dev ...)
  ↓
5. Run terraform init, then terraform plan, inside that cache folder

The module becomes the root module (the directory Terraform runs in), and the generated files complete it. Three consequences follow:

  • terragrunt plan, terragrunt apply and terragrunt init mirror the Terraform commands and run init automatically. Only the common commands have such shortcuts (among them validate, output, destroy, import, state and show). Since v0.88.0 Terragrunt rejects any other Terraform command typed directly after terragrunt, so terraform providers becomes terragrunt run -- providers.
  • Terraform ignores any TF_VAR_ value that matches no declared variable. A misspelled input name silently falls back to the variable's default, so check the plan, or run terragrunt hcl validate --inputs, which lists inputs that match no variable.
  • Current Terragrunt releases look for the tofu binary first and use terraform only when tofu is not on your PATH. To be sure of Terraform, as this course assumes, set TG_TF_PATH=terraform. To run OpenTofu explicitly, set TG_TF_PATH=tofu. Older scripts use TERRAGRUNT_TFPATH.

Add .terragrunt-cache to .gitignore. Deleting it is always safe, because it holds only downloaded and generated files. State lives in S3, and the next run rebuilds the cache.

Prediction task

A new staging unit has never been applied. Its file contains source = "...//sqs-with-dlq?ref=v1.4.0" and inputs = { queue_name = "orders-staging", max_receive_count = 5, visibility_timeout_seconds = 60 }, plus the same remote_state and generate blocks with the key staging/orders-queue/terraform.tfstate.

Before opening the answer, predict: (1) what sits in .terragrunt-cache, (2) which variables the module receives, (3) the plan summary line.

Check your answer
  1. A <hash>/<hash>/sqs-with-dlq/ folder holding the module's files (main.tf, variables.tf, outputs.tf), the generated backend.tf and provider.tf, and a .terraform folder after init.
  2. Three: queue_name, max_receive_count and visibility_timeout_seconds, passed as TF_VAR_ variables.
  3. Three new resources, with an empty state behind the key:
  # aws_sqs_queue.dlq will be created
  # aws_sqs_queue.main will be created
  # aws_sqs_queue_redrive_policy.main will be created

Plan: 3 to add, 0 to change, 0 to destroy.

The failure to name: a quiet edit that replaces a live queue

Suppose you rename the dev queue by editing one input, queue_name = "orders-development". An SQS queue's name cannot change in place, so the plan reads:

  # aws_sqs_queue.main must be replaced
-/+ resource "aws_sqs_queue" "main" {
      ~ arn  = "arn:aws:sqs:eu-west-1:111122223333:orders-dev" -> (known after apply)
      ~ name = "orders-dev" -> "orders-development" # forces replacement
      ...
    }

Plan: 3 to add, 0 to change, 3 to destroy.

The DLQ name is derived from the queue name, and the redrive policy is attached to the main queue's URL, so all three are replaced. Messages in the old queues are gone, and producers still pointing at the old URL fail. A ?ref= bump can do the same if the new module version renames a resource or a name template. Before any apply, scan for -/+, must be replaced, will be destroyed, and a non-zero destroy count in the summary.

Guided attempt: convert dev and get No changes

The dev directory already has live state, so the goal is a first plan that says No changes.

  1. Sign in without static keys: aws sso login --profile dev, then export AWS_PROFILE=dev.
  2. Back up the state from the old directory: terraform state pull > dev-backup.tfstate.
  3. Create terragrunt.hcl with the same bucket, key, Region and module tag the old directory used. Delete the old main.tf, backend.tf and provider.tf, because Terragrunt copies your directory's files into the cache and they would clash with the module.
  4. Run terragrunt plan. Expect a surprise: Plan: 3 to add, 0 to change, 3 to destroy. The old resources lived at module.queue.aws_sqs_queue.main, but the module is now the root module, so the addresses lost the module.queue. prefix. Terraform sees three unknown resources to create and three state entries with no configuration.
  5. Move the addresses (the first plan was only a plan, so nothing was changed):
terragrunt state mv 'module.queue.aws_sqs_queue.main' 'aws_sqs_queue.main'
terragrunt state mv 'module.queue.aws_sqs_queue.dlq' 'aws_sqs_queue.dlq'
terragrunt state mv 'module.queue.aws_sqs_queue_redrive_policy.main' 'aws_sqs_queue_redrive_policy.main'
  1. Plan again. The expected result is No changes. Your infrastructure matches the configuration. Any -/+ line means an input or the module tag differs from what the old directory used.

Now list what is still duplicated across dev, staging and prod units: the whole remote_state block (only key differs, and you typed it by hand), the whole provider generate block (only the Region differs), and the repeated source line. Removing those repeats is the job of the shared root file, covered in the next section, "One Root File for Many Environments: include, remote_state, generate and Environment Values".

One Root File for Many Environments: include, remote_state, generate and Environment Values

The previous section left you with a unit (one directory, one terragrunt.hcl, one state) that still carries its own backend and provider settings. Copy that unit into five directories and you have rebuilt the duplication Terragrunt was meant to remove. The fix is one shared file at the top of the repository that every unit includes, meaning Terragrunt merges that file's settings into the unit's own configuration.

Layout and the include block

The layout puts all per-environment facts in the directory path. The live/<account>/<region>/<unit> shape gets one extra level here, <env>, so that dev and prod can share an account:

repo/
  root.hcl                      <- shared backend, provider, locals
  live/
    main/                       <- account directory
      account.hcl
      eu-west-1/
        region.hcl
        dev/
          env.hcl
          queue/terragrunt.hcl
          worker/terragrunt.hcl
          table/terragrunt.hcl

Each unit now shrinks to a module source, a few inputs and an include:

# live/main/eu-west-1/dev/queue/terragrunt.hcl
include "root" {
  path = find_in_parent_folders("root.hcl")
}

terraform {
  source = "git::https://github.com/acme/tf-modules.git//sqs-with-dlq?ref=v1.4.0"
}

inputs = {
  queue_name        = "orders"
  max_receive_count = 5
}

find_in_parent_folders("root.hcl") is a Terragrunt function that starts in the unit's directory and walks up, one folder at a time, until it finds a file with that name. Included inputs are merged with the unit's own inputs, and the unit wins on a clash.

⚠️ Older repositories name the root file terragrunt.hcl and write find_in_parent_folders() with no argument, which finds the nearest parent terragrunt.hcl. They may also use an unlabeled include { ... } block. These older forms still work, but Terragrunt now treats them as deprecated. Bare includes are deprecated since v0.81.0 (labeled includes are also faster). A root named terragrunt.hcl makes Terragrunt print a warning that recommends root.hcl, a name that also avoids confusing a root file with a runnable unit.

The root file: remote_state, generate and locals

# root.hcl
locals {
  account_vars = read_terragrunt_config(find_in_parent_folders("account.hcl"))
  region_vars  = read_terragrunt_config(find_in_parent_folders("region.hcl"))
  env_vars     = read_terragrunt_config(find_in_parent_folders("env.hcl"))

  account_id  = local.account_vars.locals.account_id
  aws_region  = local.region_vars.locals.aws_region
  environment = local.env_vars.locals.environment
}

remote_state {
  backend = "s3"
  config = {
    bucket       = "acme-tfstate-${local.account_id}"
    key          = "${path_relative_to_include()}/terraform.tfstate"
    region       = "eu-west-1" # where the state bucket lives, not the unit's Region
    encrypt      = true
    use_lockfile = true
  }
  generate = {
    path      = "backend.tf"
    if_exists = "overwrite_terragrunt"
  }
}

generate "provider" {
  path      = "provider.tf"
  if_exists = "overwrite_terragrunt"
  contents  = <<EOF
provider "aws" {
  region = "${local.aws_region}"

  default_tags {
    tags = {
      Environment = "${local.environment}"
      ManagedBy   = "terragrunt"
    }
  }
}
EOF
}

inputs = {
  environment = local.environment # the module declares variable "environment"
}

Three things are happening. First, the generate attribute inside remote_state makes Terragrunt write a backend.tf into the working folder (the .terragrunt-cache copy from the previous section). Second, the separate generate "provider" block writes provider.tf. default_tags is the AWS provider setting that tags every taggable resource, and here it receives the environment name. if_exists = "overwrite_terragrunt" lets Terragrunt replace a file it generated earlier, and it refuses to overwrite a hand-written file with the same name. Third, the generated provider holds a Region and tags and no access keys; credentials come from the environment, as covered in 'Running Many Units: run --all, the Run Queue, Apply Risk and Credentials'.

One key per unit

path_relative_to_include() returns the path from the directory holding root.hcl to the unit that included it. The state key is the object path inside the bucket, so each unit gets its own state file:

Unit directory Generated key
live/main/eu-west-1/dev/queue live/main/eu-west-1/dev/queue/terraform.tfstate
live/main/eu-west-1/dev/worker live/main/eu-west-1/dev/worker/terraform.tfstate
live/main/eu-west-1/dev/table live/main/eu-west-1/dev/table/terraform.tfstate

Locking stops two runs from writing one state at once. Terraform 1.10 introduced use_lockfile = true as an experimental option, and 1.11 made it generally available (prefer 1.11 or later); it keeps a lock object beside the state in S3. Existing repositories often still carry dynamodb_table = "..." instead, and Terraform has deprecated that setting. If you see it, leave it until you plan a deliberate migration. Setting both during the changeover is allowed.

Tracing one environment value

env.hcl:  locals { environment = "dev" }
   ↓ read_terragrunt_config(find_in_parent_folders("env.hcl"))
root.hcl: local.env_vars.locals.environment
   ↓ local.environment
   β”œβ†’ provider.tf: default_tags Environment = "dev"
   β””β†’ inputs.environment = "dev" β†’ module variable

find_in_parent_folders("env.hcl") runs from the unit's directory even though it is written in root.hcl, so each unit finds its own environment's file. One file edit changes tags and module input together.

The mistake that creates duplicates

The key comes from the directory path, so renaming or moving a unit directory changes its state key. Rename dev/queue to dev/orders-queue, and the next terragrunt plan initialises against .../dev/orders-queue/terraform.tfstate, which is empty. Terraform sees no resources and plans to create everything again:

Plan: 3 to add, 0 to change, 0 to destroy.

The three are the queue, its DLQ and the redrive policy, all of which already exist. Applying this either fails midway with already-exists errors or leaves two states claiming the same objects. A "0 to change" plan full of adds on a live environment is the warning sign.

The safe order is to copy the state before the first plan:

aws s3 cp \
  s3://acme-tfstate-111111111111/live/main/eu-west-1/dev/queue/terraform.tfstate \
  s3://acme-tfstate-111111111111/live/main/eu-west-1/dev/orders-queue/terraform.tfstate
terragrunt plan
# expected: No changes. Your infrastructure matches the configuration.

Only after that plan passes should you delete the old key, since leaving it invites someone to use a stale copy. Copy only the .tfstate object, not any lock object next to it.

Guided task: add prod

Task. Add prod in the same account and Region with only two new files: live/main/eu-west-1/prod/env.hcl containing locals { environment = "prod" }, and live/main/eu-west-1/prod/queue/terragrunt.hcl, a copy of the dev queue unit with max_receive_count = 3. Before running anything, predict (a) the generated backend key and (b) the provider's default_tags.

Check your answer

(a) The path relative to root.hcl is live/main/eu-west-1/prod/queue, so the key is live/main/eu-west-1/prod/queue/terraform.tfstate, in the same bucket as dev but a different object. (b) Environment = "prod" and ManagedBy = "terragrunt". The account and Region files are found by walking up, so nothing else changes. The module input environment is also "prod", and the first plan should show creates only, because prod's state is genuinely empty. Check that the plan's name values contain prod and not dev.

Independent fix-it: one hardcoded key

A teammate wrote this in root.hcl:

config = {
  bucket = "acme-tfstate-111111111111"
  key    = "terraform.tfstate"
  region = "eu-west-1"
}

The dev/queue unit is already applied. Task. Explain what dev/worker's plan and apply would do to the queue's resources, then correct the file.

Check your answer

Every unit now points at one state object, which already records the queue, DLQ and redrive policy. The worker's module declares none of those, so its plan shows them as to destroy alongside the worker's own to add resources. Applying would delete the live queue. The fix is key = "${path_relative_to_include()}/terraform.tfstate". Because the existing state holds the queue's resources, copy it to the queue's new key first. Then the queue plan should say No changes, and the worker starts with an empty state of its own. Also delete the shared old key once both units pass, so the worker cannot be pointed at it again.

Passing Outputs Between Units with dependency and mock_outputs

The queue unit from the earlier sections creates an SQS queue and its dead-letter queue (DLQ, where messages go after too many failed receives). Now a second unit, the worker, deploys a Lambda function that consumes that queue. The worker needs the queue's ARN twice: as the event_source_arn of its event source mapping (the Lambda feature that polls SQS and invokes the function) and in the IAM policy that allows sqs:ReceiveMessage. The ARN only exists after the queue is created. How does one unit get a value that another unit's state owns?

Three ways to get the ARN

Approach Where the ARN comes from Weak point
Hardcoded string Pasted into inputs Typos, and it goes stale if the queue is replaced
Data source (data "aws_sqs_queue") Looked up in AWS at plan time The queue must already exist, and Terragrunt cannot see the ordering
dependency block The queue unit's outputs, read from its state Queue must be applied first (solved for plan by mocks, below)

A hardcoded ARN works until someone renames the queue and forgets the worker. A data source is plain Terraform and fine inside a single state, but across units it hides the fact that the worker depends on the queue, so run --all has no reason to apply the queue first. A dependency block states the relationship where Terragrunt can see it.

The dependency block

A dependency block names an upstream unit by its directory and exposes that unit's Terraform outputs. Given live/main/eu-west-1/dev/queue/ and a sibling worker/, the worker's terragrunt.hcl looks like this:

include "root" {
  path = find_in_parent_folders("root.hcl")
}

terraform {
  source = "git::https://github.com/example-org/tf-modules.git//sqs-worker?ref=v1.4.0"
}

dependency "queue" {
  config_path = "../queue" # directory of the upstream unit
}

inputs = {
  function_name = "orders-worker"
  queue_arn     = dependency.queue.outputs.queue_arn
}

The label ("queue") is the name you use in expressions: dependency.<label>.outputs.<output_name>. The queue_arn output must be declared with an output block in the queue unit's module; Terragrunt can only read what the upstream module exports.

queue unit  --apply-->  S3 state (output: queue_arn)
     |
     v   Terragrunt reads the outputs (like terraform output -json)
worker unit: inputs.queue_arn  -->  module variable queue_arn

Because Terragrunt reads the upstream unit's state, two conditions follow. The queue must already be applied, and your current credentials must be able to read the queue unit's state bucket and key. For another account, that means the right SSO profile or CI role, not just the worker's.

What a fresh environment does

In a brand-new environment nobody has applied the queue yet, and you run terragrunt plan in worker/. Terragrunt tries to read the queue's outputs, finds no state with outputs, and stops before Terraform runs. The message is along the lines of:

ERROR: ../queue/terragrunt.hcl is a dependency of worker/terragrunt.hcl but detected no outputs.
Either the target module has not been applied yet, or the module has no outputs.

The exact wording varies by version, but the cause does not: no outputs in the upstream state. You could apply the queue first, but that blocks a pull request from showing the worker's plan. That is what mocks are for.

mock_outputs: placeholders for plan only

mock_outputs are placeholder values Terragrunt uses when the upstream has no real outputs. Pair them with mock_outputs_allowed_terraform_commands, which limits the commands where the placeholders may be used.

dependency "queue" {
  config_path = "../queue"

  mock_outputs = {
    queue_arn = "arn:aws:sqs:eu-west-1:000000000000:mock-queue"
  }
  # Mocks may feed validate and plan only. apply with no real outputs fails.
  mock_outputs_allowed_terraform_commands = ["validate", "plan"]
}

Once the queue is applied, real outputs take over and the mock is ignored. Before that, terragrunt plan in the worker runs and shows something like:

  # aws_lambda_event_source_mapping.this will be created
  + resource "aws_lambda_event_source_mapping" "this" {
      + event_source_arn = "arn:aws:sqs:eu-west-1:000000000000:mock-queue"
      + function_name    = (known after apply)
      ...
    }

Plan: 4 to add, 0 to change, 0 to destroy.

The resource count depends on your module. What matters is how to read it: the structure is trustworthy, but the ARN is fake. A plan on mocks proves the worker's configuration is valid and shows which resources would be created. It does not prove the final values.

⚠️ Never allow apply in the list. The IAM policy would be created successfully against a queue that does not exist, silently granting nothing. The event source mapping would fail or point at the wrong target. The failure shows up later as a Lambda that never receives messages. With apply excluded, Terragrunt refuses to run, which is the safe outcome.

Pitfalls

  • Mock shapes must match real types. Terraform does not always catch a mismatch. If the module declares queue_arn as string and your mock is a list, Terragrunt passes the list as JSON text and the plan succeeds with that text as the value. If the real output is an object, the mock needs the same attribute names. Mock only the outputs you actually use.
  • Renaming an upstream output breaks downstream units. Rename queue_arn to arn in the queue module and every dependency.queue.outputs.queue_arn stops resolving. Treat output names as an interface: add the new one, migrate consumers, then remove the old one.
  • Cycles cannot be ordered. If queue depends on worker and worker depends on queue, there is no valid run order and Terragrunt reports a cycle. Break it by moving the shared value into a third unit.
  • dependencies orders without passing values. dependencies { paths = ["../network"] } only says "run the network unit first". Use it when the worker needs the network applied but takes no outputs from it. Use dependency when you need a value.

Alternative: publish to SSM Parameter Store

The queue unit can write queue_arn to an SSM parameter (Parameter Store, AWS's key-value configuration service). The worker then reads it with a data "aws_ssm_parameter" source. The two approaches compare like this:

Cue dependency block SSM parameter
Coupling Tight, by directory path Loose, by parameter name
Ordering visible to Terragrunt Yes No
Plan before queue exists Mocks Data source errors
Works across teams or tools Needs state access Needs only SSM read access

Choose SSM when the consumer belongs to another team or does not use Terragrunt. Choose dependency inside one repository, where you want run --all to apply the queue first without anyone remembering to.

Practice

Exercise 1. The worker module takes queue_arn (string). The queue module exports queue_arn and dlq_arn. Write the dependency block and inputs for the worker with mocks that allow plan but not apply. Then predict the worker's plan (a) before the queue is applied and (b) after, and say which values are real in each.

Check your answer
dependency "queue" {
  config_path = "../queue"
  mock_outputs = {
    queue_arn = "arn:aws:sqs:eu-west-1:000000000000:mock-queue"
  }
  mock_outputs_allowed_terraform_commands = ["validate", "plan"]
}

inputs = {
  queue_arn = dependency.queue.outputs.queue_arn
}

(a) Before: the plan succeeds with everything shown as to-add. Any non-dependency values, such as literal inputs or module defaults, are real; queue_arn in the event source mapping and IAM policy is the mock string. (b) After: the resource list is the same, but queue_arn is the real ARN from the queue's state. Only dlq_arn is left out of the mock because the worker does not use it.

Harder variation. Add a table unit (a DynamoDB table) exporting table_arn and table_name. The worker's IAM policy uses table_arn, and its environment uses table_name. Write the second dependency, decide which commands may use mocks, and say what goes wrong if apply were allowed.

Check your answer
dependency "table" {
  config_path = "../table"
  mock_outputs = {
    table_arn  = "arn:aws:dynamodb:eu-west-1:000000000000:table/mock-table"
    table_name = "mock-table"
  }
  mock_outputs_allowed_terraform_commands = ["validate", "plan"]
}

Both dependencies allow only validate and plan. If apply were allowed with the table unapplied, the IAM policy would be created for a nonexistent table ARN and the Lambda would receive mock-table as its table name. Apply would succeed, then every write would fail with AccessDenied or ResourceNotFoundException at runtime. Because each dependency has its own mock list, a missing output in one still fails an apply even if the other is real.

Running Many Units: run --all, the Run Queue, Apply Risk and Credentials

A dev environment for this backend is four units: network, queue (SQS with its DLQ), worker (the Lambda consumer) and api (API Gateway). Each unit is its own directory and its own state. Running terragrunt plan four times, in the right order, by hand, does not scale. Terragrunt can run every unit under a directory, but one command that touches many states needs careful reading before you trust it.

The current command and the older forms

terragrunt run --all <command> runs a Terraform command in every unit under the current directory, in dependency order. Existing repositories mostly use the older spelling, so you need to read both.

Older form you will meet Current form
terragrunt run-all plan terragrunt run --all plan
TERRAGRUNT_TFPATH TG_TF_PATH
TERRAGRUNT_PARALLELISM / --terragrunt-parallelism TG_PARALLELISM / --parallelism
TERRAGRUNT_NON_INTERACTIVE / --terragrunt-non-interactive TG_NON_INTERACTIVE / --non-interactive
--terragrunt-include-dir --queue-include-dir
--terragrunt-exclude-dir --queue-exclude-dir
--terragrunt-no-auto-approve --no-auto-approve

The pattern is that TERRAGRUNT_ becomes TG_ and --terragrunt- flags lose their prefix. The include and exclude flags also gain queue-, and in current releases both are aliases for the newer --filter flag. How much of the older column still works depends on your Terragrunt version. On Terragrunt 1.1.6, run-all and every --terragrunt- flag are rejected with an error (those flags were removed in v0.85.0), while the TERRAGRUNT_ variables still work and print a deprecation warning. Pin the version in CI and check the release notes before a bulk rename.

Reading the run queue

Before running anything, Terragrunt builds the run queue: the list of units it will process, ordered by their dependencies. A unit starts only when every unit it depends on has finished, and units that do not depend on each other may run in parallel. A handy way to predict the order is to sort the units into groups: group 1 has no dependencies, and each later group depends only on earlier ones. Your task is to predict the queue from these dependency blocks alone:

Unit Its dependency blocks point at
network none
queue none
worker network, queue
api worker, network

Write down the groups, and also the order for destroy.

Check your answer

Group 1 is network and queue, since neither waits for anything. Group 2 is worker, which needs both. Group 3 is api, which needs worker; its dependency on network is already satisfied. Current Terragrunt (1.1.6 here) prints the queue as a tree, with each unit listed under the units it waits for:

The following units will be run, starting with dependencies and then their dependents:
.
β”œβ”€β”€ network
β”‚   β”œβ”€β”€ worker
β”‚   ╰── api
╰── queue
    ╰── worker
        ╰── api

Older versions printed the same queue as a list headed Group 1, Group 2, Group 3.

destroy reverses the order: api first, then worker, then network and queue together. A consumer must disappear before the thing it consumes from.

Scoping a run

run --all acts on the directory tree you are standing in. From live/main/eu-west-1/dev it covers only dev. From the repository root it covers every account, which is almost never what you want. You can narrow further with --queue-include-dir or --queue-exclude-dir. An included unit's dependencies are not added to the run (on Terragrunt 1.1.6, --queue-include-dir ./worker runs worker alone), so always read the printed queue before you apply, and check that it lists only the intended units. If many units hit the same AWS API limits, cap the fan-out with --parallelism 2.

Why apply across units is riskier than one plan

  • Many states, one command. A failure in worker leaves network and queue applied and api untouched: a half-applied environment. Terraform cannot roll that back.
  • Downstream plans can be provisional. If queue is not applied yet, worker plans against mock_outputs. Its fake ARN will be replaced later, so the plan is not the final one.
  • Dependencies read state, not plans. Say queue shows a replace in the plan. worker still reads the old queue ARN from state, so it shows no change. It only shows one after queue is applied.
  • Apply re-plans. run --all apply runs terraform apply in each unit, which computes a fresh plan. The plan you read earlier is not the one that executes, unless you use saved plan files (the --out-dir option collects one per unit; confirm it exists in your pinned version).
  • Auto-approve is implied. run --all apply asks one yes/no question for the whole queue (none at all with --non-interactive) and then passes -auto-approve to every unit, so there is no per-unit prompt. Opt out with --no-auto-approve. Otherwise your reviewed plan is the only gate.

The plan-first routine

  1. terragrunt run --all plan from the environment directory, after reading the queue.
  2. Scan each unit's summary for -/+ (replace) and destroy lines. Plan: 3 to add, 0 to change, 3 to destroy on the queue unit means a live queue is being swapped, and messages inside it are lost.
  3. Apply safe changes with run --all apply. For risky ones (queue or table renames), cd into that unit and run terragrunt apply alone, then re-plan the downstream units.

⚠️ A plan across all units is only complete after the upstream units are applied. Plan again afterwards, and expect new changes downstream.

Credentials without static keys

No key pair appears in root.hcl, env files or CI secrets.

Locally, each account gets an IAM Identity Center (SSO) profile. Sign in, then choose it:

aws sso login --profile dev-admin
export AWS_PROFILE=dev-admin
terragrunt run --all plan

You can instead keep a profile name in env.hcl and use it in root.hcl, which stops dev code running with prod credentials. But a profile name does not exist on a CI runner, so AWS_PROFILE is the simpler choice.

In CI, the job assumes an IAM role through OIDC (OpenID Connect, which lets the CI provider prove its identity to AWS) and receives short-lived credentials before Terragrunt starts. A GitHub Actions job looks like this:

permissions:
  id-token: write   # lets the job request an OIDC token
  contents: read
steps:
  - uses: aws-actions/configure-aws-credentials@v4
    with:
      role-to-assume: arn:aws:iam::111111111111:role/tg-plan-readonly
      aws-region: eu-west-1
  - run: terragrunt run --all --non-interactive -- plan -lock=false

A role's trust policy decides who may assume it. This one, for the apply role, allows only the main branch of one repository (it assumes the account's GitHub OIDC provider resource already exists):

data "aws_iam_policy_document" "apply_trust" {
  statement {
    actions = ["sts:AssumeRoleWithWebIdentity"]
    principals {
      type        = "Federated"
      identifiers = [aws_iam_openid_connect_provider.github.arn]
    }
    condition {
      test     = "StringEquals"
      variable = "token.actions.githubusercontent.com:aud"
      values   = ["sts.amazonaws.com"]
    }
    condition {
      test     = "StringEquals"
      variable = "token.actions.githubusercontent.com:sub"
      values   = ["repo:my-org/backend-infra:ref:refs/heads/main"]
    }
  }
}

Use two roles. The plan role is read-only and its sub value is repo:my-org/backend-infra:pull_request. These values use the format of repositories created before July 15, 2026. A repository created, renamed or transferred after that date, or one that has opted in, gets owner and repository IDs in the claim (for example repo:octo-org@123456/octo-repo@456789:ref:refs/heads/main), so check which format your repository issues before writing the condition. The apply role can write and trusts only main. A read-only role cannot write the S3 lock object, which is why the PR plan above passes -lock=false; this is safe only because plan changes nothing. The -- in that command matters: Terragrunt rejects a Terraform flag such as -lock=false written after run --all plan without it.

Exercise: predict and place

Five units: network, queue and table have no dependencies. worker depends on queue and table. api depends on worker and network. A pull request bumps the module ?ref= in queue/terragrunt.hcl, and the new version renames the DLQ.

  1. Write the queue groups.
  2. Which units show changes in run --all plan, and which after queue is applied?
  3. Which commands belong in CI on the pull request, and which only after merge?
Check your answer
  1. Group 1 is network, queue, table. Group 2 is worker. Group 3 is api.
  2. Only queue shows changes in the first plan, probably a -/+ on the DLQ. worker and api read the queue outputs from the old state, so they show nothing. After queue is applied and its outputs change, worker may show changes wherever it uses a queue output that changed (the redrive policy itself belongs to the queue unit), and api may change if it uses the worker's outputs.
  3. The pull request runs run --all plan with the read-only role, showing the -/+ line for review. After merge, the apply role applies queue on its own, then run --all plan again to see the downstream changes, and then applies the rest. Nothing on a pull request ever runs apply.

Stacks in Brief and Choosing Between Terragrunt and Plain Terraform

By now dev and prod each hold a queue/terragrunt.hcl and a worker/terragrunt.hcl. The files differ only in a few values, yet you still maintain the same two-unit tree twice. This section does two things. First it shows how Terragrunt can remove that repetition. Then it asks the harder question: is Terragrunt worth having at all for your repository?

Stacks: declare the tree once

A stack is a terragrunt.stack.hcl file that lists the units an environment needs. Each unit block has four parts:

  • a name, the label on the block;
  • a source, where the unit's terragrunt.hcl comes from (a Git URL with ?ref= or a local path);
  • a path, the directory to create for it;
  • values, a map of settings the unit reads.

Terragrunt generates the unit directories for you.

# live/main/eu-west-1/dev/terragrunt.stack.hcl  (prod copies this file with different values)
unit "queue" {
  source = "git::git@github.com:acme/infra-catalog.git//units/queue?ref=v1.2.0"
  path   = "queue"
  values = {
    environment       = "dev"
    max_receive_count = 3
  }
}

unit "worker" {
  source = "git::git@github.com:acme/infra-catalog.git//units/worker?ref=v1.2.0"
  path   = "worker"
  values = {
    environment = "dev"
    queue_path  = "../queue"
  }
}

The catalog's worker unit is an ordinary unit file that reads values.<key> instead of hardcoding anything. Terragrunt writes each unit's values into a terragrunt.values.hcl file next to the generated terragrunt.hcl:

# units/worker/terragrunt.hcl in the catalog repository
include "root" {
  path = find_in_parent_folders("root.hcl")
}

terraform {
  source = "git::git@github.com:acme/infra-catalog.git//modules/sqs-worker?ref=v1.2.0"
}

dependency "queue" {
  config_path  = values.queue_path
  mock_outputs = { queue_arn = "arn:aws:sqs:eu-west-1:123456789012:mock" }
  mock_outputs_allowed_terraform_commands = ["validate", "plan"]  # never apply
}

inputs = {
  environment = values.environment
  queue_arn   = dependency.queue.outputs.queue_arn
}

Generating the stack produces this tree:

live/main/eu-west-1/dev/
  terragrunt.stack.hcl
  .terragrunt-stack/
    queue/terragrunt.hcl     (generated)
    worker/terragrunt.hcl    (generated)

The generated files are ordinary units, so include, dependency and run --all work on them as before. Two cautions. Because the units now live under .terragrunt-stack, path_relative_to_include returns a path containing that segment, so adopting stacks for units that already have state moves their state keys. Copy the state object first and confirm No changes, exactly as for a moved directory. Also, stacks are newer than units and include, so confirm the current commands (such as terragrunt stack generate) and file layout in the Terragrunt documentation for your pinned version. Treat stacks as optional: everything before this section works without them.

The plain Terraform version of the same system

Here is the same queue-and-worker system with no wrapper: one shared module and one directory per environment. Terraform's backend block cannot use variables, so it stays partial (empty) and each environment supplies its values at init time.

modules/orders/            (queue, DLQ, worker; tagged v1.2.0)
envs/dev/main.tf  backend.hcl  terraform.tfvars
envs/prod/main.tf backend.hcl  terraform.tfvars
# envs/dev/main.tf  (identical in prod)
terraform {
  required_version = ">= 1.10"
  backend "s3" {}
  required_providers {
    aws = { source = "hashicorp/aws", version = "~> 6.0" }
  }
}

variable "region"            { type = string }
variable "environment"       { type = string }
variable "max_receive_count" { type = number }

provider "aws" {
  region = var.region
  default_tags {
    tags = { environment = var.environment }
  }
}

module "orders" {
  source            = "git::git@github.com:acme/infra-catalog.git//modules/orders?ref=v1.2.0"
  environment       = var.environment
  max_receive_count = var.max_receive_count
}
# envs/dev/backend.hcl
bucket       = "acme-tfstate-dev"
key          = "orders/terraform.tfstate"
region       = "eu-west-1"
encrypt      = true
use_lockfile = true   # S3 native locking (experimental in Terraform 1.10, stable from 1.11)
# envs/dev/terraform.tfvars  (auto-loaded from the working directory)
region            = "eu-west-1"
environment       = "dev"
max_receive_count = 3
terraform -chdir=envs/dev init -backend-config=backend.hcl
terraform -chdir=envs/dev plan

For values that cross a state boundary, either publish them as SSM Parameter Store parameters and read them with a data "aws_ssm_parameter" block, or read another state's outputs with terraform_remote_state. Ordering between states is then your job, typically as separate CI steps.

Measuring what remains duplicated: the terraform and provider stanzas and the module call, roughly 25 lines copied per environment, plus a 5-line backend.hcl and a 3-line tfvars. The risk is drift. If someone edits only envs/prod/main.tf's provider block, nothing flags it. A CI step that diffs the shared stanzas across directories narrows that gap.

The cost ledger

Terragrunt cost Where it bites
Extra binary Pin it locally and in CI, alongside Terraform
Generated files Errors point into .terragrunt-cache
mock_outputs A mock allowed on apply reaches real IAM
run --all One command touches many states
Onboarding Two tools and two vocabularies to learn
Terragrunt benefit Why it holds
No backend drift One remote_state block generates every backend.tf
No provider drift One generate block writes every provider.tf
Manageable matrix Units Γ— environments Γ— accounts stay as small files
Ordered runs dependency blocks drive run --all

Authentication is not a differentiator. Both tools run after IAM Identity Center login locally or OIDC role assumption in CI, so neither needs static keys.

Decision cues, not thresholds

These are heuristics, not a formula. Plain Terraform is enough when you have few states, one or two accounts, and rare backend edits. Terragrunt earns its place when you have many units, several accounts, dependency chains, and you keep editing backend or provider blocks in many places. Switching is not a one-way door: Terragrunt keeps ordinary Terraform state in S3, so you can later run Terraform directly against the same keys, or adopt Terragrunt on top of existing directories.

Practice: three scenario cards

Choose plain Terraform or Terragrunt for each, and justify with the cues.

  1. A two-person team runs dev and prod in one AWS account, with three states each.
  2. A platform team owns six units (network, table, queue, worker, api, events) deployed to four accounts, with dependencies between units.
  3. A repository has one shared state for a single Lambda stack, with no environments.
Check your answer
  1. Plain Terraform. Six states, one account, and a few lines of duplication per directory. The wrapper's cost exceeds the drift risk.
  2. Terragrunt. That is 24 unit-instances, with a dependency chain to order and backend/provider blocks that would otherwise be edited in 24 places. Account, Region and environment values become small files read by a single root.hcl.
  3. Plain Terraform. One state means nothing is repeated, so there is nothing for Terragrunt to remove.

Independent transfer task

A backend has SQS queues with DLQs, Lambda workers, a DynamoDB table and an API Gateway. It runs in two AWS accounts (dev, prod). CI assumes roles through OIDC. Choose one:

  • Plain path: write the init and plan command lines for prod, and name the files that hold the state key and the environment values.
  • Terragrunt path: sketch the root.hcl and unit layout, say which unit depends on which, and say which commands may use mocks.

Whichever you choose, state in one sentence why the cues support it.

Check your answer

Plain path (defensible with one state per account, two states in total):

terraform -chdir=envs/prod init -backend-config=backend.hcl
terraform -chdir=envs/prod plan -out=tfplan

envs/prod/backend.hcl holds the bucket in the prod account and the key. envs/prod/terraform.tfvars holds the Region and environment values. Duplication stays small and nothing needs cross-state ordering.

Terragrunt path (more defensible if you split into separate states per component):

root.hcl                      (remote_state + generate provider)
live/dev/account.hcl
live/dev/eu-west-1/region.hcl
live/dev/eu-west-1/table/terragrunt.hcl
live/dev/eu-west-1/queue/terragrunt.hcl
live/dev/eu-west-1/worker/terragrunt.hcl   (dependencies: queue, table)
live/dev/eu-west-1/api/terragrunt.hcl      (dependency: worker)
live/prod/...                               (same tree)

The keys derive from path_relative_to_include, so they are unique per unit. Mocks on queue_arn and the table ARN are allowed only for validate and plan. Either answer is acceptable if the justification matches the cues. The plain path wins at this size, and the Terragrunt path wins once components, accounts or edits multiply.

Final checklist

  • Unit traced: I can say what each unit's cache folder runs, and where its generated files come from.
  • State keys unique: every unit or directory maps to a different S3 key.
  • Mocks limited to plan: mock_outputs_allowed_terraform_commands excludes apply.
  • Run queue read: I checked the listed units and their order before any run --all apply.
  • No static keys: local runs use SSO profiles and CI uses OIDC role assumption.