Terragrunt: One Configuration, Many Environments
Terragrunt as a thin wrapper that runs Terraform or OpenTofu for you, for a developer who already knows plan and apply, S3 remote state and modules. The problem it solves: the same backend, provider and module wiring repeated in every environment directory. Units (one terragrunt.hcl per state), the terraform block with a pinned module source, inputs, and include with a shared root file (root.hcl) found with find_in_parent_folders; generating backend.tf and provider.tf with remote_state and generate blocks, and a distinct S3 state key per unit from path_relative_to_include; locals and read_terragrunt_config for account, Region and environment values; dependency blocks that pass one unit's outputs to another, with mock_outputs so plan works before the upstream unit exists; running many units in dependency order with run --all, reading the run queue, and why apply across units is riskier than a single plan; the .terragrunt-cache directory and what Terragrunt actually executes; stacks (terragrunt.stack.hcl) in brief. Use the current Terragrunt CLI and file names, and name the older forms a reader still meets in existing repositories (run-all, a root terragrunt.hcl, TERRAGRUNT_ environment variables instead of TG_). AWS authentication with IAM Identity Center profiles locally and OIDC role assumption in CI, no static keys. End with when plain Terraform with directories and tfvars is enough and Terragrunt is not worth adding.
SPACED REPETITION Β· 15 practice questions
Make this lesson stick.
Try 3 questions now. No account needed. Sample answers aren't saved.
or sign in to practice all 15From Copied Directories to a Unit: Your First terragrunt.hcl
Open your dev, staging and prod directories side by side and diff them. In many Terraform-on-AWS repositories you will find three copies of backend.tf, three copies of provider.tf, and three main.tf files calling the same module, with a few strings changed. Every backend fix then has to be made three times, and one copy gets forgotten.
Terragrunt is a thin wrapper that runs Terraform (or OpenTofu) for you. It does not replace state, plans or modules. It writes the repeated files on your behalf and then runs the commands you already know. The building block is the unit: one directory containing one terragrunt.hcl, which maps to exactly one Terraform state. This section converts one directory into a unit and traces what Terragrunt really executes.
The worked problem: what actually differs?
Here is the dev directory. It calls a module sqs-with-dlq that creates a queue, a dead-letter queue (DLQ) and the redrive policy linking them. Lines marked DIFFERS change per environment; everything else is copy-paste.
# dev/orders-queue/backend.tf
terraform {
backend "s3" {
bucket = "acme-tf-state" # copy-paste
key = "dev/orders-queue/terraform.tfstate" # DIFFERS
region = "eu-west-1" # copy-paste (the bucket's Region)
encrypt = true # copy-paste
}
}
# dev/orders-queue/provider.tf
provider "aws" {
region = "eu-west-1" # DIFFERS (prod uses eu-central-1)
}
# dev/orders-queue/main.tf
module "queue" {
source = "git::git@github.com:acme/tf-modules.git//sqs-with-dlq?ref=v1.4.0" # copy-paste
queue_name = "orders-dev" # DIFFERS
max_receive_count = 3 # DIFFERS (retry count before a message goes to the DLQ)
visibility_timeout_seconds = 60 # copy-paste
}
Out of roughly a dozen meaningful lines, only four differ: the state key, the provider Region, the queue name and the retry count. The rest is repetition that Terragrunt is meant to absorb.
The unit file
The new dev directory holds only this file. Read each block before moving on.
# live/dev/orders-queue/terragrunt.hcl
terraform {
# Pinned to a Git tag. The double slash selects a sub-folder of the repo.
source = "git::git@github.com:acme/tf-modules.git//sqs-with-dlq?ref=v1.4.0"
}
# Tells Terragrunt to write backend.tf for us.
remote_state {
backend = "s3"
generate = {
path = "backend.tf"
if_exists = "overwrite_terragrunt" # replace the file only if Terragrunt wrote it
}
config = {
bucket = "acme-tf-state"
key = "dev/orders-queue/terraform.tfstate"
region = "eu-west-1"
encrypt = true
}
}
# Tells Terragrunt to write provider.tf. No access keys: credentials come from the environment.
generate "provider" {
path = "provider.tf"
if_exists = "overwrite_terragrunt"
contents = <<EOF
provider "aws" {
region = "eu-west-1"
}
EOF
}
# Values for the module's variables.
inputs = {
queue_name = "orders-dev"
max_receive_count = 3
visibility_timeout_seconds = 60
}
The ?ref=v1.4.0 is the safety feature. If every environment pointed at a shared local path such as ../../modules/sqs-with-dlq, or at a floating branch like ?ref=main, one module edit would change the next plan of every environment at once. A tag lets you bump dev first, read the plan, and promote the tag to staging and prod later. Tags can be moved by force-pushing, so treat them as immutable by convention.
What Terragrunt runs on your behalf
When you run terragrunt plan, it does this:
terragrunt plan (in live/dev/orders-queue)
β
1. Download the module at ?ref=v1.4.0 into .terragrunt-cache/<hash>/<hash>/sqs-with-dlq/
β
2. Copy this directory's own files over it
β
3. Write the generated backend.tf and provider.tf next to the module files
β
4. Export each input as a TF_VAR_ environment variable (TF_VAR_queue_name=orders-dev ...)
β
5. Run terraform init, then terraform plan, inside that cache folder
The module becomes the root module (the directory Terraform runs in), and the generated files complete it. Three consequences follow:
terragrunt plan,terragrunt applyandterragrunt initmirror the Terraform commands and runinitautomatically. Only the common commands have such shortcuts (among themvalidate,output,destroy,import,stateandshow). Since v0.88.0 Terragrunt rejects any other Terraform command typed directly afterterragrunt, soterraform providersbecomesterragrunt run -- providers.- Terraform ignores any
TF_VAR_value that matches no declared variable. A misspelled input name silently falls back to the variable's default, so check the plan, or runterragrunt hcl validate --inputs, which lists inputs that match no variable. - Current Terragrunt releases look for the
tofubinary first and useterraformonly whentofuis not on yourPATH. To be sure of Terraform, as this course assumes, setTG_TF_PATH=terraform. To run OpenTofu explicitly, setTG_TF_PATH=tofu. Older scripts useTERRAGRUNT_TFPATH.
Add .terragrunt-cache to .gitignore. Deleting it is always safe, because it holds only downloaded and generated files. State lives in S3, and the next run rebuilds the cache.
Prediction task
A new staging unit has never been applied. Its file contains source = "...//sqs-with-dlq?ref=v1.4.0" and inputs = { queue_name = "orders-staging", max_receive_count = 5, visibility_timeout_seconds = 60 }, plus the same remote_state and generate blocks with the key staging/orders-queue/terraform.tfstate.
Before opening the answer, predict: (1) what sits in .terragrunt-cache, (2) which variables the module receives, (3) the plan summary line.
Check your answer
- A
<hash>/<hash>/sqs-with-dlq/folder holding the module's files (main.tf,variables.tf,outputs.tf), the generatedbackend.tfandprovider.tf, and a.terraformfolder after init. - Three:
queue_name,max_receive_countandvisibility_timeout_seconds, passed asTF_VAR_variables. - Three new resources, with an empty state behind the key:
# aws_sqs_queue.dlq will be created
# aws_sqs_queue.main will be created
# aws_sqs_queue_redrive_policy.main will be created
Plan: 3 to add, 0 to change, 0 to destroy.
The failure to name: a quiet edit that replaces a live queue
Suppose you rename the dev queue by editing one input, queue_name = "orders-development". An SQS queue's name cannot change in place, so the plan reads:
# aws_sqs_queue.main must be replaced
-/+ resource "aws_sqs_queue" "main" {
~ arn = "arn:aws:sqs:eu-west-1:111122223333:orders-dev" -> (known after apply)
~ name = "orders-dev" -> "orders-development" # forces replacement
...
}
Plan: 3 to add, 0 to change, 3 to destroy.
The DLQ name is derived from the queue name, and the redrive policy is attached to the main queue's URL, so all three are replaced. Messages in the old queues are gone, and producers still pointing at the old URL fail. A ?ref= bump can do the same if the new module version renames a resource or a name template. Before any apply, scan for -/+, must be replaced, will be destroyed, and a non-zero destroy count in the summary.
Guided attempt: convert dev and get No changes
The dev directory already has live state, so the goal is a first plan that says No changes.
- Sign in without static keys:
aws sso login --profile dev, thenexport AWS_PROFILE=dev. - Back up the state from the old directory:
terraform state pull > dev-backup.tfstate. - Create
terragrunt.hclwith the same bucket, key, Region and module tag the old directory used. Delete the oldmain.tf,backend.tfandprovider.tf, because Terragrunt copies your directory's files into the cache and they would clash with the module. - Run
terragrunt plan. Expect a surprise:Plan: 3 to add, 0 to change, 3 to destroy.The old resources lived atmodule.queue.aws_sqs_queue.main, but the module is now the root module, so the addresses lost themodule.queue.prefix. Terraform sees three unknown resources to create and three state entries with no configuration. - Move the addresses (the first plan was only a plan, so nothing was changed):
terragrunt state mv 'module.queue.aws_sqs_queue.main' 'aws_sqs_queue.main'
terragrunt state mv 'module.queue.aws_sqs_queue.dlq' 'aws_sqs_queue.dlq'
terragrunt state mv 'module.queue.aws_sqs_queue_redrive_policy.main' 'aws_sqs_queue_redrive_policy.main'
- Plan again. The expected result is
No changes. Your infrastructure matches the configuration.Any-/+line means an input or the module tag differs from what the old directory used.
Now list what is still duplicated across dev, staging and prod units: the whole remote_state block (only key differs, and you typed it by hand), the whole provider generate block (only the Region differs), and the repeated source line. Removing those repeats is the job of the shared root file, covered in the next section, "One Root File for Many Environments: include, remote_state, generate and Environment Values".
One Root File for Many Environments: include, remote_state, generate and Environment Values
The previous section left you with a unit (one directory, one terragrunt.hcl, one state) that still carries its own backend and provider settings. Copy that unit into five directories and you have rebuilt the duplication Terragrunt was meant to remove. The fix is one shared file at the top of the repository that every unit includes, meaning Terragrunt merges that file's settings into the unit's own configuration.
Layout and the include block
The layout puts all per-environment facts in the directory path. The live/<account>/<region>/<unit> shape gets one extra level here, <env>, so that dev and prod can share an account:
repo/
root.hcl <- shared backend, provider, locals
live/
main/ <- account directory
account.hcl
eu-west-1/
region.hcl
dev/
env.hcl
queue/terragrunt.hcl
worker/terragrunt.hcl
table/terragrunt.hcl
Each unit now shrinks to a module source, a few inputs and an include:
# live/main/eu-west-1/dev/queue/terragrunt.hcl
include "root" {
path = find_in_parent_folders("root.hcl")
}
terraform {
source = "git::https://github.com/acme/tf-modules.git//sqs-with-dlq?ref=v1.4.0"
}
inputs = {
queue_name = "orders"
max_receive_count = 5
}
find_in_parent_folders("root.hcl") is a Terragrunt function that starts in the unit's directory and walks up, one folder at a time, until it finds a file with that name. Included inputs are merged with the unit's own inputs, and the unit wins on a clash.
β οΈ Older repositories name the root file terragrunt.hcl and write find_in_parent_folders() with no argument, which finds the nearest parent terragrunt.hcl. They may also use an unlabeled include { ... } block. These older forms still work, but Terragrunt now treats them as deprecated. Bare includes are deprecated since v0.81.0 (labeled includes are also faster). A root named terragrunt.hcl makes Terragrunt print a warning that recommends root.hcl, a name that also avoids confusing a root file with a runnable unit.
The root file: remote_state, generate and locals
# root.hcl
locals {
account_vars = read_terragrunt_config(find_in_parent_folders("account.hcl"))
region_vars = read_terragrunt_config(find_in_parent_folders("region.hcl"))
env_vars = read_terragrunt_config(find_in_parent_folders("env.hcl"))
account_id = local.account_vars.locals.account_id
aws_region = local.region_vars.locals.aws_region
environment = local.env_vars.locals.environment
}
remote_state {
backend = "s3"
config = {
bucket = "acme-tfstate-${local.account_id}"
key = "${path_relative_to_include()}/terraform.tfstate"
region = "eu-west-1" # where the state bucket lives, not the unit's Region
encrypt = true
use_lockfile = true
}
generate = {
path = "backend.tf"
if_exists = "overwrite_terragrunt"
}
}
generate "provider" {
path = "provider.tf"
if_exists = "overwrite_terragrunt"
contents = <<EOF
provider "aws" {
region = "${local.aws_region}"
default_tags {
tags = {
Environment = "${local.environment}"
ManagedBy = "terragrunt"
}
}
}
EOF
}
inputs = {
environment = local.environment # the module declares variable "environment"
}
Three things are happening. First, the generate attribute inside remote_state makes Terragrunt write a backend.tf into the working folder (the .terragrunt-cache copy from the previous section). Second, the separate generate "provider" block writes provider.tf. default_tags is the AWS provider setting that tags every taggable resource, and here it receives the environment name. if_exists = "overwrite_terragrunt" lets Terragrunt replace a file it generated earlier, and it refuses to overwrite a hand-written file with the same name. Third, the generated provider holds a Region and tags and no access keys; credentials come from the environment, as covered in 'Running Many Units: run --all, the Run Queue, Apply Risk and Credentials'.
One key per unit
path_relative_to_include() returns the path from the directory holding root.hcl to the unit that included it. The state key is the object path inside the bucket, so each unit gets its own state file:
| Unit directory | Generated key |
|---|---|
live/main/eu-west-1/dev/queue |
live/main/eu-west-1/dev/queue/terraform.tfstate |
live/main/eu-west-1/dev/worker |
live/main/eu-west-1/dev/worker/terraform.tfstate |
live/main/eu-west-1/dev/table |
live/main/eu-west-1/dev/table/terraform.tfstate |
Locking stops two runs from writing one state at once. Terraform 1.10 introduced use_lockfile = true as an experimental option, and 1.11 made it generally available (prefer 1.11 or later); it keeps a lock object beside the state in S3. Existing repositories often still carry dynamodb_table = "..." instead, and Terraform has deprecated that setting. If you see it, leave it until you plan a deliberate migration. Setting both during the changeover is allowed.
Tracing one environment value
env.hcl: locals { environment = "dev" }
β read_terragrunt_config(find_in_parent_folders("env.hcl"))
root.hcl: local.env_vars.locals.environment
β local.environment
ββ provider.tf: default_tags Environment = "dev"
ββ inputs.environment = "dev" β module variable
find_in_parent_folders("env.hcl") runs from the unit's directory even though it is written in root.hcl, so each unit finds its own environment's file. One file edit changes tags and module input together.
The mistake that creates duplicates
The key comes from the directory path, so renaming or moving a unit directory changes its state key. Rename dev/queue to dev/orders-queue, and the next terragrunt plan initialises against .../dev/orders-queue/terraform.tfstate, which is empty. Terraform sees no resources and plans to create everything again:
Plan: 3 to add, 0 to change, 0 to destroy.
The three are the queue, its DLQ and the redrive policy, all of which already exist. Applying this either fails midway with already-exists errors or leaves two states claiming the same objects. A "0 to change" plan full of adds on a live environment is the warning sign.
The safe order is to copy the state before the first plan:
aws s3 cp \
s3://acme-tfstate-111111111111/live/main/eu-west-1/dev/queue/terraform.tfstate \
s3://acme-tfstate-111111111111/live/main/eu-west-1/dev/orders-queue/terraform.tfstate
terragrunt plan
# expected: No changes. Your infrastructure matches the configuration.
Only after that plan passes should you delete the old key, since leaving it invites someone to use a stale copy. Copy only the .tfstate object, not any lock object next to it.
Guided task: add prod
Task. Add prod in the same account and Region with only two new files: live/main/eu-west-1/prod/env.hcl containing locals { environment = "prod" }, and live/main/eu-west-1/prod/queue/terragrunt.hcl, a copy of the dev queue unit with max_receive_count = 3. Before running anything, predict (a) the generated backend key and (b) the provider's default_tags.
Check your answer
(a) The path relative to root.hcl is live/main/eu-west-1/prod/queue, so the key is live/main/eu-west-1/prod/queue/terraform.tfstate, in the same bucket as dev but a different object. (b) Environment = "prod" and ManagedBy = "terragrunt". The account and Region files are found by walking up, so nothing else changes. The module input environment is also "prod", and the first plan should show creates only, because prod's state is genuinely empty. Check that the plan's name values contain prod and not dev.
Independent fix-it: one hardcoded key
A teammate wrote this in root.hcl:
config = {
bucket = "acme-tfstate-111111111111"
key = "terraform.tfstate"
region = "eu-west-1"
}
The dev/queue unit is already applied. Task. Explain what dev/worker's plan and apply would do to the queue's resources, then correct the file.
Check your answer
Every unit now points at one state object, which already records the queue, DLQ and redrive policy. The worker's module declares none of those, so its plan shows them as to destroy alongside the worker's own to add resources. Applying would delete the live queue. The fix is key = "${path_relative_to_include()}/terraform.tfstate". Because the existing state holds the queue's resources, copy it to the queue's new key first. Then the queue plan should say No changes, and the worker starts with an empty state of its own. Also delete the shared old key once both units pass, so the worker cannot be pointed at it again.
Passing Outputs Between Units with dependency and mock_outputs
The queue unit from the earlier sections creates an SQS queue and its dead-letter queue (DLQ, where messages go after too many failed receives). Now a second unit, the worker, deploys a Lambda function that consumes that queue. The worker needs the queue's ARN twice: as the event_source_arn of its event source mapping (the Lambda feature that polls SQS and invokes the function) and in the IAM policy that allows sqs:ReceiveMessage. The ARN only exists after the queue is created. How does one unit get a value that another unit's state owns?
Three ways to get the ARN
| Approach | Where the ARN comes from | Weak point |
|---|---|---|
| Hardcoded string | Pasted into inputs |
Typos, and it goes stale if the queue is replaced |
Data source (data "aws_sqs_queue") |
Looked up in AWS at plan time | The queue must already exist, and Terragrunt cannot see the ordering |
dependency block |
The queue unit's outputs, read from its state | Queue must be applied first (solved for plan by mocks, below) |
A hardcoded ARN works until someone renames the queue and forgets the worker. A data source is plain Terraform and fine inside a single state, but across units it hides the fact that the worker depends on the queue, so run --all has no reason to apply the queue first. A dependency block states the relationship where Terragrunt can see it.
The dependency block
A dependency block names an upstream unit by its directory and exposes that unit's Terraform outputs. Given live/main/eu-west-1/dev/queue/ and a sibling worker/, the worker's terragrunt.hcl looks like this:
include "root" {
path = find_in_parent_folders("root.hcl")
}
terraform {
source = "git::https://github.com/example-org/tf-modules.git//sqs-worker?ref=v1.4.0"
}
dependency "queue" {
config_path = "../queue" # directory of the upstream unit
}
inputs = {
function_name = "orders-worker"
queue_arn = dependency.queue.outputs.queue_arn
}
The label ("queue") is the name you use in expressions: dependency.<label>.outputs.<output_name>. The queue_arn output must be declared with an output block in the queue unit's module; Terragrunt can only read what the upstream module exports.
queue unit --apply--> S3 state (output: queue_arn)
|
v Terragrunt reads the outputs (like terraform output -json)
worker unit: inputs.queue_arn --> module variable queue_arn
Because Terragrunt reads the upstream unit's state, two conditions follow. The queue must already be applied, and your current credentials must be able to read the queue unit's state bucket and key. For another account, that means the right SSO profile or CI role, not just the worker's.
What a fresh environment does
In a brand-new environment nobody has applied the queue yet, and you run terragrunt plan in worker/. Terragrunt tries to read the queue's outputs, finds no state with outputs, and stops before Terraform runs. The message is along the lines of:
ERROR: ../queue/terragrunt.hcl is a dependency of worker/terragrunt.hcl but detected no outputs.
Either the target module has not been applied yet, or the module has no outputs.
The exact wording varies by version, but the cause does not: no outputs in the upstream state. You could apply the queue first, but that blocks a pull request from showing the worker's plan. That is what mocks are for.
mock_outputs: placeholders for plan only
mock_outputs are placeholder values Terragrunt uses when the upstream has no real outputs. Pair them with mock_outputs_allowed_terraform_commands, which limits the commands where the placeholders may be used.
dependency "queue" {
config_path = "../queue"
mock_outputs = {
queue_arn = "arn:aws:sqs:eu-west-1:000000000000:mock-queue"
}
# Mocks may feed validate and plan only. apply with no real outputs fails.
mock_outputs_allowed_terraform_commands = ["validate", "plan"]
}
Once the queue is applied, real outputs take over and the mock is ignored. Before that, terragrunt plan in the worker runs and shows something like:
# aws_lambda_event_source_mapping.this will be created
+ resource "aws_lambda_event_source_mapping" "this" {
+ event_source_arn = "arn:aws:sqs:eu-west-1:000000000000:mock-queue"
+ function_name = (known after apply)
...
}
Plan: 4 to add, 0 to change, 0 to destroy.
The resource count depends on your module. What matters is how to read it: the structure is trustworthy, but the ARN is fake. A plan on mocks proves the worker's configuration is valid and shows which resources would be created. It does not prove the final values.
β οΈ Never allow apply in the list. The IAM policy would be created successfully against a queue that does not exist, silently granting nothing. The event source mapping would fail or point at the wrong target. The failure shows up later as a Lambda that never receives messages. With apply excluded, Terragrunt refuses to run, which is the safe outcome.
Pitfalls
- Mock shapes must match real types. Terraform does not always catch a mismatch. If the module declares
queue_arnasstringand your mock is a list, Terragrunt passes the list as JSON text and the plan succeeds with that text as the value. If the real output is an object, the mock needs the same attribute names. Mock only the outputs you actually use. - Renaming an upstream output breaks downstream units. Rename
queue_arntoarnin the queue module and everydependency.queue.outputs.queue_arnstops resolving. Treat output names as an interface: add the new one, migrate consumers, then remove the old one. - Cycles cannot be ordered. If queue depends on worker and worker depends on queue, there is no valid run order and Terragrunt reports a cycle. Break it by moving the shared value into a third unit.
dependenciesorders without passing values.dependencies { paths = ["../network"] }only says "run the network unit first". Use it when the worker needs the network applied but takes no outputs from it. Usedependencywhen you need a value.
Alternative: publish to SSM Parameter Store
The queue unit can write queue_arn to an SSM parameter (Parameter Store, AWS's key-value configuration service). The worker then reads it with a data "aws_ssm_parameter" source. The two approaches compare like this:
| Cue | dependency block | SSM parameter |
|---|---|---|
| Coupling | Tight, by directory path | Loose, by parameter name |
| Ordering visible to Terragrunt | Yes | No |
| Plan before queue exists | Mocks | Data source errors |
| Works across teams or tools | Needs state access | Needs only SSM read access |
Choose SSM when the consumer belongs to another team or does not use Terragrunt. Choose dependency inside one repository, where you want run --all to apply the queue first without anyone remembering to.
Practice
Exercise 1. The worker module takes queue_arn (string). The queue module exports queue_arn and dlq_arn. Write the dependency block and inputs for the worker with mocks that allow plan but not apply. Then predict the worker's plan (a) before the queue is applied and (b) after, and say which values are real in each.
Check your answer
dependency "queue" {
config_path = "../queue"
mock_outputs = {
queue_arn = "arn:aws:sqs:eu-west-1:000000000000:mock-queue"
}
mock_outputs_allowed_terraform_commands = ["validate", "plan"]
}
inputs = {
queue_arn = dependency.queue.outputs.queue_arn
}
(a) Before: the plan succeeds with everything shown as to-add. Any non-dependency values, such as literal inputs or module defaults, are real; queue_arn in the event source mapping and IAM policy is the mock string. (b) After: the resource list is the same, but queue_arn is the real ARN from the queue's state. Only dlq_arn is left out of the mock because the worker does not use it.
Harder variation. Add a table unit (a DynamoDB table) exporting table_arn and table_name. The worker's IAM policy uses table_arn, and its environment uses table_name. Write the second dependency, decide which commands may use mocks, and say what goes wrong if apply were allowed.
Check your answer
dependency "table" {
config_path = "../table"
mock_outputs = {
table_arn = "arn:aws:dynamodb:eu-west-1:000000000000:table/mock-table"
table_name = "mock-table"
}
mock_outputs_allowed_terraform_commands = ["validate", "plan"]
}
Both dependencies allow only validate and plan. If apply were allowed with the table unapplied, the IAM policy would be created for a nonexistent table ARN and the Lambda would receive mock-table as its table name. Apply would succeed, then every write would fail with AccessDenied or ResourceNotFoundException at runtime. Because each dependency has its own mock list, a missing output in one still fails an apply even if the other is real.
Running Many Units: run --all, the Run Queue, Apply Risk and Credentials
A dev environment for this backend is four units: network, queue (SQS with its DLQ), worker (the Lambda consumer) and api (API Gateway). Each unit is its own directory and its own state. Running terragrunt plan four times, in the right order, by hand, does not scale. Terragrunt can run every unit under a directory, but one command that touches many states needs careful reading before you trust it.
The current command and the older forms
terragrunt run --all <command> runs a Terraform command in every unit under the current directory, in dependency order. Existing repositories mostly use the older spelling, so you need to read both.
| Older form you will meet | Current form |
|---|---|
terragrunt run-all plan |
terragrunt run --all plan |
TERRAGRUNT_TFPATH |
TG_TF_PATH |
TERRAGRUNT_PARALLELISM / --terragrunt-parallelism |
TG_PARALLELISM / --parallelism |
TERRAGRUNT_NON_INTERACTIVE / --terragrunt-non-interactive |
TG_NON_INTERACTIVE / --non-interactive |
--terragrunt-include-dir |
--queue-include-dir |
--terragrunt-exclude-dir |
--queue-exclude-dir |
--terragrunt-no-auto-approve |
--no-auto-approve |
The pattern is that TERRAGRUNT_ becomes TG_ and --terragrunt- flags lose their prefix. The include and exclude flags also gain queue-, and in current releases both are aliases for the newer --filter flag. How much of the older column still works depends on your Terragrunt version. On Terragrunt 1.1.6, run-all and every --terragrunt- flag are rejected with an error (those flags were removed in v0.85.0), while the TERRAGRUNT_ variables still work and print a deprecation warning. Pin the version in CI and check the release notes before a bulk rename.
Reading the run queue
Before running anything, Terragrunt builds the run queue: the list of units it will process, ordered by their dependencies. A unit starts only when every unit it depends on has finished, and units that do not depend on each other may run in parallel. A handy way to predict the order is to sort the units into groups: group 1 has no dependencies, and each later group depends only on earlier ones. Your task is to predict the queue from these dependency blocks alone:
| Unit | Its dependency blocks point at |
|---|---|
network |
none |
queue |
none |
worker |
network, queue |
api |
worker, network |
Write down the groups, and also the order for destroy.
Check your answer
Group 1 is network and queue, since neither waits for anything. Group 2 is worker, which needs both. Group 3 is api, which needs worker; its dependency on network is already satisfied. Current Terragrunt (1.1.6 here) prints the queue as a tree, with each unit listed under the units it waits for:
The following units will be run, starting with dependencies and then their dependents:
.
βββ network
β βββ worker
β β°ββ api
β°ββ queue
β°ββ worker
β°ββ api
Older versions printed the same queue as a list headed Group 1, Group 2, Group 3.
destroy reverses the order: api first, then worker, then network and queue together. A consumer must disappear before the thing it consumes from.
Scoping a run
run --all acts on the directory tree you are standing in. From live/main/eu-west-1/dev it covers only dev. From the repository root it covers every account, which is almost never what you want. You can narrow further with --queue-include-dir or --queue-exclude-dir. An included unit's dependencies are not added to the run (on Terragrunt 1.1.6, --queue-include-dir ./worker runs worker alone), so always read the printed queue before you apply, and check that it lists only the intended units. If many units hit the same AWS API limits, cap the fan-out with --parallelism 2.
Why apply across units is riskier than one plan
- Many states, one command. A failure in
workerleavesnetworkandqueueapplied andapiuntouched: a half-applied environment. Terraform cannot roll that back. - Downstream plans can be provisional. If
queueis not applied yet,workerplans againstmock_outputs. Its fake ARN will be replaced later, so the plan is not the final one. - Dependencies read state, not plans. Say
queueshows a replace in the plan.workerstill reads the old queue ARN from state, so it shows no change. It only shows one afterqueueis applied. - Apply re-plans.
run --all applyrunsterraform applyin each unit, which computes a fresh plan. The plan you read earlier is not the one that executes, unless you use saved plan files (the--out-diroption collects one per unit; confirm it exists in your pinned version). - Auto-approve is implied.
run --all applyasks one yes/no question for the whole queue (none at all with--non-interactive) and then passes-auto-approveto every unit, so there is no per-unit prompt. Opt out with--no-auto-approve. Otherwise your reviewed plan is the only gate.
The plan-first routine
terragrunt run --all planfrom the environment directory, after reading the queue.- Scan each unit's summary for
-/+(replace) anddestroylines.Plan: 3 to add, 0 to change, 3 to destroyon the queue unit means a live queue is being swapped, and messages inside it are lost. - Apply safe changes with
run --all apply. For risky ones (queue or table renames),cdinto that unit and runterragrunt applyalone, then re-plan the downstream units.
β οΈ A plan across all units is only complete after the upstream units are applied. Plan again afterwards, and expect new changes downstream.
Credentials without static keys
No key pair appears in root.hcl, env files or CI secrets.
Locally, each account gets an IAM Identity Center (SSO) profile. Sign in, then choose it:
aws sso login --profile dev-admin
export AWS_PROFILE=dev-admin
terragrunt run --all plan
You can instead keep a profile name in env.hcl and use it in root.hcl, which stops dev code running with prod credentials. But a profile name does not exist on a CI runner, so AWS_PROFILE is the simpler choice.
In CI, the job assumes an IAM role through OIDC (OpenID Connect, which lets the CI provider prove its identity to AWS) and receives short-lived credentials before Terragrunt starts. A GitHub Actions job looks like this:
permissions:
id-token: write # lets the job request an OIDC token
contents: read
steps:
- uses: aws-actions/configure-aws-credentials@v4
with:
role-to-assume: arn:aws:iam::111111111111:role/tg-plan-readonly
aws-region: eu-west-1
- run: terragrunt run --all --non-interactive -- plan -lock=false
A role's trust policy decides who may assume it. This one, for the apply role, allows only the main branch of one repository (it assumes the account's GitHub OIDC provider resource already exists):
data "aws_iam_policy_document" "apply_trust" {
statement {
actions = ["sts:AssumeRoleWithWebIdentity"]
principals {
type = "Federated"
identifiers = [aws_iam_openid_connect_provider.github.arn]
}
condition {
test = "StringEquals"
variable = "token.actions.githubusercontent.com:aud"
values = ["sts.amazonaws.com"]
}
condition {
test = "StringEquals"
variable = "token.actions.githubusercontent.com:sub"
values = ["repo:my-org/backend-infra:ref:refs/heads/main"]
}
}
}
Use two roles. The plan role is read-only and its sub value is repo:my-org/backend-infra:pull_request. These values use the format of repositories created before July 15, 2026. A repository created, renamed or transferred after that date, or one that has opted in, gets owner and repository IDs in the claim (for example repo:octo-org@123456/octo-repo@456789:ref:refs/heads/main), so check which format your repository issues before writing the condition. The apply role can write and trusts only main. A read-only role cannot write the S3 lock object, which is why the PR plan above passes -lock=false; this is safe only because plan changes nothing. The -- in that command matters: Terragrunt rejects a Terraform flag such as -lock=false written after run --all plan without it.
Exercise: predict and place
Five units: network, queue and table have no dependencies. worker depends on queue and table. api depends on worker and network. A pull request bumps the module ?ref= in queue/terragrunt.hcl, and the new version renames the DLQ.
- Write the queue groups.
- Which units show changes in
run --all plan, and which afterqueueis applied? - Which commands belong in CI on the pull request, and which only after merge?
Check your answer
- Group 1 is
network,queue,table. Group 2 isworker. Group 3 isapi. - Only
queueshows changes in the first plan, probably a-/+on the DLQ.workerandapiread the queue outputs from the old state, so they show nothing. Afterqueueis applied and its outputs change,workermay show changes wherever it uses a queue output that changed (the redrive policy itself belongs to thequeueunit), andapimay change if it uses the worker's outputs. - The pull request runs
run --all planwith the read-only role, showing the-/+line for review. After merge, the apply role appliesqueueon its own, thenrun --all planagain to see the downstream changes, and then applies the rest. Nothing on a pull request ever runsapply.
Stacks in Brief and Choosing Between Terragrunt and Plain Terraform
By now dev and prod each hold a queue/terragrunt.hcl and a worker/terragrunt.hcl. The files differ only in a few values, yet you still maintain the same two-unit tree twice. This section does two things. First it shows how Terragrunt can remove that repetition. Then it asks the harder question: is Terragrunt worth having at all for your repository?
Stacks: declare the tree once
A stack is a terragrunt.stack.hcl file that lists the units an environment needs. Each unit block has four parts:
- a
name, the label on the block; - a
source, where the unit'sterragrunt.hclcomes from (a Git URL with?ref=or a local path); - a
path, the directory to create for it; values, a map of settings the unit reads.
Terragrunt generates the unit directories for you.
# live/main/eu-west-1/dev/terragrunt.stack.hcl (prod copies this file with different values)
unit "queue" {
source = "git::git@github.com:acme/infra-catalog.git//units/queue?ref=v1.2.0"
path = "queue"
values = {
environment = "dev"
max_receive_count = 3
}
}
unit "worker" {
source = "git::git@github.com:acme/infra-catalog.git//units/worker?ref=v1.2.0"
path = "worker"
values = {
environment = "dev"
queue_path = "../queue"
}
}
The catalog's worker unit is an ordinary unit file that reads values.<key> instead of hardcoding anything. Terragrunt writes each unit's values into a terragrunt.values.hcl file next to the generated terragrunt.hcl:
# units/worker/terragrunt.hcl in the catalog repository
include "root" {
path = find_in_parent_folders("root.hcl")
}
terraform {
source = "git::git@github.com:acme/infra-catalog.git//modules/sqs-worker?ref=v1.2.0"
}
dependency "queue" {
config_path = values.queue_path
mock_outputs = { queue_arn = "arn:aws:sqs:eu-west-1:123456789012:mock" }
mock_outputs_allowed_terraform_commands = ["validate", "plan"] # never apply
}
inputs = {
environment = values.environment
queue_arn = dependency.queue.outputs.queue_arn
}
Generating the stack produces this tree:
live/main/eu-west-1/dev/
terragrunt.stack.hcl
.terragrunt-stack/
queue/terragrunt.hcl (generated)
worker/terragrunt.hcl (generated)
The generated files are ordinary units, so include, dependency and run --all work on them as before. Two cautions. Because the units now live under .terragrunt-stack, path_relative_to_include returns a path containing that segment, so adopting stacks for units that already have state moves their state keys. Copy the state object first and confirm No changes, exactly as for a moved directory. Also, stacks are newer than units and include, so confirm the current commands (such as terragrunt stack generate) and file layout in the Terragrunt documentation for your pinned version. Treat stacks as optional: everything before this section works without them.
The plain Terraform version of the same system
Here is the same queue-and-worker system with no wrapper: one shared module and one directory per environment. Terraform's backend block cannot use variables, so it stays partial (empty) and each environment supplies its values at init time.
modules/orders/ (queue, DLQ, worker; tagged v1.2.0)
envs/dev/main.tf backend.hcl terraform.tfvars
envs/prod/main.tf backend.hcl terraform.tfvars
# envs/dev/main.tf (identical in prod)
terraform {
required_version = ">= 1.10"
backend "s3" {}
required_providers {
aws = { source = "hashicorp/aws", version = "~> 6.0" }
}
}
variable "region" { type = string }
variable "environment" { type = string }
variable "max_receive_count" { type = number }
provider "aws" {
region = var.region
default_tags {
tags = { environment = var.environment }
}
}
module "orders" {
source = "git::git@github.com:acme/infra-catalog.git//modules/orders?ref=v1.2.0"
environment = var.environment
max_receive_count = var.max_receive_count
}
# envs/dev/backend.hcl
bucket = "acme-tfstate-dev"
key = "orders/terraform.tfstate"
region = "eu-west-1"
encrypt = true
use_lockfile = true # S3 native locking (experimental in Terraform 1.10, stable from 1.11)
# envs/dev/terraform.tfvars (auto-loaded from the working directory)
region = "eu-west-1"
environment = "dev"
max_receive_count = 3
terraform -chdir=envs/dev init -backend-config=backend.hcl
terraform -chdir=envs/dev plan
For values that cross a state boundary, either publish them as SSM Parameter Store parameters and read them with a data "aws_ssm_parameter" block, or read another state's outputs with terraform_remote_state. Ordering between states is then your job, typically as separate CI steps.
Measuring what remains duplicated: the terraform and provider stanzas and the module call, roughly 25 lines copied per environment, plus a 5-line backend.hcl and a 3-line tfvars. The risk is drift. If someone edits only envs/prod/main.tf's provider block, nothing flags it. A CI step that diffs the shared stanzas across directories narrows that gap.
The cost ledger
| Terragrunt cost | Where it bites |
|---|---|
| Extra binary | Pin it locally and in CI, alongside Terraform |
| Generated files | Errors point into .terragrunt-cache |
mock_outputs |
A mock allowed on apply reaches real IAM |
run --all |
One command touches many states |
| Onboarding | Two tools and two vocabularies to learn |
| Terragrunt benefit | Why it holds |
|---|---|
| No backend drift | One remote_state block generates every backend.tf |
| No provider drift | One generate block writes every provider.tf |
| Manageable matrix | Units Γ environments Γ accounts stay as small files |
| Ordered runs | dependency blocks drive run --all |
Authentication is not a differentiator. Both tools run after IAM Identity Center login locally or OIDC role assumption in CI, so neither needs static keys.
Decision cues, not thresholds
These are heuristics, not a formula. Plain Terraform is enough when you have few states, one or two accounts, and rare backend edits. Terragrunt earns its place when you have many units, several accounts, dependency chains, and you keep editing backend or provider blocks in many places. Switching is not a one-way door: Terragrunt keeps ordinary Terraform state in S3, so you can later run Terraform directly against the same keys, or adopt Terragrunt on top of existing directories.
Practice: three scenario cards
Choose plain Terraform or Terragrunt for each, and justify with the cues.
- A two-person team runs
devandprodin one AWS account, with three states each. - A platform team owns six units (network, table, queue, worker, api, events) deployed to four accounts, with dependencies between units.
- A repository has one shared state for a single Lambda stack, with no environments.
Check your answer
- Plain Terraform. Six states, one account, and a few lines of duplication per directory. The wrapper's cost exceeds the drift risk.
- Terragrunt. That is 24 unit-instances, with a dependency chain to order and backend/provider blocks that would otherwise be edited in 24 places. Account, Region and environment values become small files read by a single
root.hcl. - Plain Terraform. One state means nothing is repeated, so there is nothing for Terragrunt to remove.
Independent transfer task
A backend has SQS queues with DLQs, Lambda workers, a DynamoDB table and an API Gateway. It runs in two AWS accounts (dev, prod). CI assumes roles through OIDC. Choose one:
- Plain path: write the
initandplancommand lines forprod, and name the files that hold the state key and the environment values. - Terragrunt path: sketch the
root.hcland unit layout, say which unit depends on which, and say which commands may use mocks.
Whichever you choose, state in one sentence why the cues support it.
Check your answer
Plain path (defensible with one state per account, two states in total):
terraform -chdir=envs/prod init -backend-config=backend.hcl
terraform -chdir=envs/prod plan -out=tfplan
envs/prod/backend.hcl holds the bucket in the prod account and the key. envs/prod/terraform.tfvars holds the Region and environment values. Duplication stays small and nothing needs cross-state ordering.
Terragrunt path (more defensible if you split into separate states per component):
root.hcl (remote_state + generate provider)
live/dev/account.hcl
live/dev/eu-west-1/region.hcl
live/dev/eu-west-1/table/terragrunt.hcl
live/dev/eu-west-1/queue/terragrunt.hcl
live/dev/eu-west-1/worker/terragrunt.hcl (dependencies: queue, table)
live/dev/eu-west-1/api/terragrunt.hcl (dependency: worker)
live/prod/... (same tree)
The keys derive from path_relative_to_include, so they are unique per unit. Mocks on queue_arn and the table ARN are allowed only for validate and plan. Either answer is acceptable if the justification matches the cues. The plain path wins at this size, and the Terragrunt path wins once components, accounts or edits multiply.
Final checklist
- Unit traced: I can say what each unit's cache folder runs, and where its generated files come from.
- State keys unique: every unit or directory maps to a different S3 key.
- Mocks limited to plan:
mock_outputs_allowed_terraform_commandsexcludesapply. - Run queue read: I checked the listed units and their order before any
run --all apply. - No static keys: local runs use SSO profiles and CI uses OIDC role assumption.