Start Here: What Terraform Is
A true beginner's introduction for someone who has never used Terraform, before any internals: what infrastructure as code means and why teams use it instead of clicking in the AWS console; how Terraform compares with CloudFormation and the AWS CDK in one short table; installing Terraform and pinning its version; what a provider is (the AWS provider as a plugin that talks to the AWS APIs); a ten-line first configuration that creates one S3 bucket, run with terraform init, plan, apply and destroy, showing and explaining each command's output; what the state file is, in plain words, and why it matters; which files to commit and which to ignore (.terraform/, state files, the lock file); and where to go next in this roadmap. Assume no prior Terraform knowledge; define every term at first use; keep the AWS side to one bucket.
SPACED REPETITION Β· 15 practice questions
Make this lesson stick.
Try 3 questions now. No account needed. Sample answers aren't saved.
or sign in to practice all 15Why Infrastructure as Code, and Where Terraform Fits
You built an SQS queue, a DynamoDB table and an IAM role by clicking through the console. It worked. Now a teammate has to rebuild the same thing in a second AWS account. You open the console and try to remember which settings you changed, and you can't. Nobody reviewed those clicks, nothing recorded them, and the console has no "do it again" button.
Infrastructure as code (IaC) fixes this by describing cloud resources in text files that a tool applies for you. Terraform is one such tool. This section covers what IaC fixes, what it does not fix, and why this roadmap uses Terraform.
What goes wrong with console clicking
- No review. Nobody sees a change before it happens. A pull request lets a teammate catch a mistake first.
- No history. Git can answer "who changed the table's capacity mode, and why?" The console keeps no reviewed diff and no reason. AWS CloudTrail does record who made which API call, including console clicks, but its built-in event history covers only the last 90 days and never says why.
- No repeatability. A second account means a second round of clicking, with a second chance to forget a setting.
- Silent drift. Drift means the real resource no longer matches what anyone remembers. A queue's timeout gets edited during an incident and nobody writes it down.
Declarative versus imperative
You could script the clicks with the aws CLI. This is the imperative style: you list the steps.
set -e # stop at the first failing command
aws sqs create-queue --queue-name orders
aws dynamodb create-table --table-name orders \
--attribute-definitions AttributeName=id,AttributeType=S \
--key-schema AttributeName=id,KeyType=HASH \
--billing-mode PAY_PER_REQUEST
aws iam create-role --role-name orders-worker \
--assume-role-policy-document file://trust.json
Suppose trust.json has a typo and the last line fails. The queue and table now exist. You fix the file and rerun the script. It stops at the table step, because create-table fails with ResourceInUseException when the table already exists. The script cannot tell "already done" from "broken", so you end up editing it by hand.
Terraform is declarative: you describe the end state you want, and Terraform works out the steps. Your files are the configuration, written in HCL (HashiCorp Configuration Language, Terraform's file syntax). Each real thing you describe is a resource, for example one table:
resource "aws_dynamodb_table" "orders" {
name = "orders"
billing_mode = "PAY_PER_REQUEST"
hash_key = "id"
attribute {
name = "id"
type = "S"
}
}
Before touching anything, Terraform produces a plan: a preview of the changes it would make. For this resource on a fresh account, the plan ends with Plan: 1 to add, 0 to change, 0 to destroy. Once the table exists, the same configuration gives No changes. Rerunning is safe because Terraform compares the desired state with what exists. If a run fails halfway, the next run creates only what is still missing.
configuration (.tf files) -- what you want
β
plan -- preview: what would change
β
apply -- Terraform calls the AWS APIs
β
real AWS resources
The later sections walk through running this cycle for real, starting with 'Your First Configuration: One S3 Bucket Through init, plan, apply and destroy'.
Terraform, CloudFormation and the AWS CDK
Terraform is not the only way to do IaC on AWS. Here are the two alternatives you will meet most often. Two words in the table come before their proper sections: a tool's state is its record of which real resources it created and manages, and a provider is a Terraform plugin that talks to one service's API (AWS, GitHub, Datadog and so on). Both are explained properly later in this lesson.
| Tool | Language | State lives | Beyond AWS | Preview |
|---|---|---|---|---|
| Terraform | HCL | State file you store (local or remote) | Yes, via providers | terraform plan |
| CloudFormation | YAML or JSON | Inside AWS, per stack | Mostly no (third-party resource types through its registry) | Change sets (a preview of a stack update) |
| AWS CDK | TypeScript, Python and others | CloudFormation stacks | AWS-focused | cdk diff (compares your code with the deployed stack) |
Honest cues for each:
- CloudFormation fits teams that are AWS-only and want AWS to hold the state, with no files to store.
- CDK fits teams that want loops, classes and tests in a general-purpose language. It generates CloudFormation underneath.
- Terraform fits teams that want one tool and one readable syntax, possibly across AWS and other services, with a plan you can read in a pull request.
This roadmap uses Terraform with HCL only. The reader's job is to read a plan and change infrastructure safely, and a single declarative language keeps that one skill in focus. (This table is a simplified view. CDK and CloudFormation have their own ways of previewing and storing state, and each tool has edge cases.)
Prediction task: what does IaC remove?
For each scenario, decide whether IaC removes the problem, only exposes it, or leaves it open. Answer before opening the check.
- A teammate changes a security setting directly in the console.
- A pull request changes the DynamoDB table definition.
- A new account needs the same stack as the existing one.
Check your answer
- Exposed, not prevented. Terraform cannot stop someone from clicking. The next plan shows the difference as a change that would revert the setting. IaC turns silent drift into something you can see. A team still needs a habit of making changes only through the code.
- Review is removed, risk is not. The diff of the
.tffile is reviewable, and Git keeps the history. But a small-looking edit can still make the plan show a replacement (destroy then create) of a table that holds data. Someone has to read the plan, not just the diff. - Repeatability is removed. The same configuration runs against the other account. You still have to handle values that differ per account, and globally unique names such as S3 buckets cannot be copied verbatim. Parameterizing those comes later in the roadmap, under variables.
Pitfall: one owner per resource
Suppose Terraform created the table, and then someone edits it in the console. The next terraform plan shows changes nobody expected, because the real table no longer matches the configuration. Terraform will try to undo the console edit or, in some cases, replace the resource. The reverse also happens: changing in code something someone tuned by hand overwrites their work. The rule is that each resource has exactly one owner. Either Terraform manages it and all changes go through the code, or Terraform does not manage it.
Installing Terraform, Pinning Versions, and Meeting the AWS Provider
Two teammates check out the same Terraform project. One has a Terraform release from last year, the other installed whatever was newest this morning. Each runs the same command and gets different behavior. Worse, HashiCorp does not promise that an older release can read a state file (Terraform's record of the real resources it manages, covered in a later section) that a newer one wrote. This section removes that class of problem: you install Terraform, pin which versions are allowed, connect it to AWS without a single access key, and check which account it will touch.
Installing the CLI and reading terraform version
The Terraform CLI is one executable that reads your .tf files and talks to AWS. Install it in one of two ways:
- Package manager (one version at a time): on macOS,
brew tap hashicorp/tapthenbrew install hashicorp/tap/terraform. On Linux, add HashiCorp's apt repository (Ubuntu, Debian) or its yum/dnf repository (RHEL, CentOS, Fedora, Amazon Linux) and install theterraformpackage; the exact commands for each distribution are on developer.hashicorp.com/terraform/install. On Windows, that page offers a zip download instead: unzip it and putterraform.exein a folder that is on yourPATH. - Version manager (several versions side by side, switching per project): tools such as
tfenvormise. This suits you if you maintain several projects that pin different versions.
Then confirm the install:
$ terraform version
Terraform v1.9.8
on darwin_arm64
Your version of Terraform is out of date! The latest version
is 1.X.Y. You can update by downloading from https://www.terraform.io/downloads.html
The first line is the exact CLI version. The second is the platform (operating system and CPU), which matters when a provider build is missing for your machine. The last paragraph appears only when a newer release exists. It is a notice, not an error, and you do not need to upgrade just because you see it. Upgrading is a team decision, which is why the next step exists.
Pinning: required_version and the ~> operator
A version constraint is a rule in your configuration saying which versions may run it. The terraform block is where you write the rules:
terraform {
required_version = "~> 1.9"
}
The pessimistic operator ~> means "this version or newer, but only by raising the last number you wrote." The last component you write is the one allowed to rise:
| Constraint | Allows | Blocks |
|---|---|---|
~> 1.9 |
1.9.0 up to below 2.0.0 | 2.0.0 and above, 1.8.x |
~> 1.9.0 |
1.9.0 up to below 1.10.0 | 1.10.0 and above |
= 1.9.8 |
exactly 1.9.8 | everything else |
If someone runs a CLI outside the range, Terraform stops with an error before touching anything. The bound also protects your state. HashiCorp promises that you can upgrade within 1.x, not that you can go back: a newer release may store state in a format that older releases cannot read. Keeping everyone in one range avoids finding that out the hard way.
What a provider is
Terraform itself knows nothing about AWS. A provider is a plugin, a separate program Terraform downloads, that talks to one service's API. The AWS provider translates a block like "an S3 bucket named X" into the matching AWS API calls. Two separate blocks configure it:
terraform {
required_version = "~> 1.9"
required_providers {
aws = {
source = "hashicorp/aws" # which plugin to download (publisher/name)
version = "~> 6.0" # which versions of that plugin are acceptable
}
}
}
provider "aws" {
region = "eu-west-1" # which region this provider acts in
}
required_providers answers which plugin and which versions, and it lives inside the terraform block. The provider "aws" block answers how that plugin behaves, such as the region. Keep the two ideas apart: one is about downloading, the other about configuring.
Worked example and your turn
Read the block above. The CLI constraint ~> 1.9 allows Terraform 1.9.0 up to, but not including, 2.0.0. The provider constraint ~> 6.0 allows AWS provider 6.0.0 up to, but not including, 7.0.0, so 6.1 or 6.40 qualify and 7.0 does not. The exact provider version chosen inside that range is recorded the first time you run terraform init, which the next section covers.
Your task: edit both constraints so that only patch updates (the third number) are allowed. Then say whether Terraform 1.10.0 and provider 6.0.7 would be accepted.
Check your answer
terraform {
required_version = "~> 1.9.0"
required_providers {
aws = {
source = "hashicorp/aws"
version = "~> 6.0.0"
}
}
}
Writing three components makes the third one the movable part. ~> 1.9.0 means at least 1.9.0 and below 1.10.0, so 1.10.0 is rejected. ~> 6.0.0 means at least 6.0.0 and below 6.1.0, so 6.0.7 is accepted. Tight pins like this give maximum repeatability, at the cost of upgrading deliberately by editing the constraint.
Credentials without static keys
The provider needs permission to call AWS. This roadmap never uses static access keys, long-lived key pairs that work until someone deletes them. Engineers use IAM Identity Center (AWS's single sign-on service, SSO), which hands out short-lived credentials. (CI gets the same effect by assuming a role through OIDC, covered later in the roadmap.) Set up once with aws configure sso (a command of the AWS CLI version 2, which you install separately), which writes a named profile into ~/.aws/config. Each working session then looks like this:
aws sso login --profile sandbox # opens a browser, you approve, a temporary session starts
export AWS_PROFILE=sandbox # tell the CLI and Terraform which profile to use
The AWS provider uses the standard AWS credential lookup, which understands SSO profiles, so you add nothing to your .tf files. When the session expires, Terraform fails with an expired-token error, and you run aws sso login again.
β οΈ Never write access_key or secret_key into a .tf file, and never commit a key to Git. Git history is permanent, and those files are copied to every clone. Also unset any old AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY or AWS_SESSION_TOKEN variables in your shell so there is no doubt which identity is in use.
Quick check before the first apply
A profile name is only a label. Verify the real identity:
$ aws sts get-caller-identity
{
"UserId": "AROAEXAMPLEID:alice@example.com",
"Account": "111122223333",
"Arn": "arn:aws:sts::111122223333:assumed-role/AWSReservedSSO_Developer_0123abcd/alice@example.com"
}
Account is where resources will be created. The Arn shows the role you are acting as, here the role IAM Identity Center created for the Developer permission set (AWSReservedSSO_Developer_...). Terraform will act as exactly this identity.
Your task: the sandbox account is 111122223333 and production is 999988887777. You run the command and see "Account": "999988887777". What do you do before any Terraform command?
Check your answer
Stop. You are signed in to production. Run export AWS_PROFILE=sandbox (after aws sso login --profile sandbox if needed), then run aws sts get-caller-identity again and confirm the account is 111122223333. Checking costs seconds, while an apply against the wrong account changes real infrastructure.
Your First Configuration: One S3 Bucket Through init, plan, apply and destroy
You know what an S3 bucket is. Now you will create one without opening the console, change it, and delete it, all from a text file. The goal is to learn the loop you will repeat for every Lambda, queue and table later: write, preview, apply, verify. Before you start, make sure you can run terraform version and that aws sts get-caller-identity shows the account you intend to use (both covered in the previous section).
The configuration
Create an empty directory and put this in a file called main.tf:
terraform {
required_providers {
aws = { source = "hashicorp/aws", version = "~> 6.0" }
}
}
provider "aws" {
region = "eu-west-1"
}
resource "aws_s3_bucket" "example" {
bucket = "tf-first-demo-jsmith-4821" # replace with your own unique name
}
The file has 11 non-blank lines, and four of them are just closing braces. Every Terraform file is built from blocks, which are a keyword, optional labels, and a { } body. This file has three:
- The
terraformblock holds settings about Terraform itself. Here it declares which provider plugin to download and which versions are acceptable. - The
providerblock configures that plugin. Here it sets the AWS region. - The
resourceblock declares one real thing you want to exist.
The resource block has two labels. aws_s3_bucket is the resource type, which says what kind of thing this is and which provider owns it (the aws_ prefix). example is the local name, a label you choose that is only meaningful inside your configuration. Inside the body, bucket = "..." is an argument: a setting you supply, written as name = value. Together, type and local name form the reference address aws_s3_bucket.example, which is how every other part of Terraform refers to this block, including plan output and the state file. So the address must be unique: a second resource "aws_s3_bucket" "example" block in the same directory is rejected with Error: Duplicate resource "aws_s3_bucket" configuration.
β οΈ By default, a bucket name must be unique across all AWS accounts in all standard AWS Regions, and must be 3 to 63 characters of lowercase letters, digits, hyphens and dots (best avoided), starting and ending with a letter or digit. A name like my-bucket is almost certainly taken, so add your name and some digits.
Check the file before running anything
terraform fmt # rewrites files to canonical style (spacing, alignment); prints names of files it changed
terraform validate # checks syntax and argument names against the provider; touches nothing in AWS
fmt is cosmetic. validate is a first correctness check, but it needs the provider's schema (the list of valid arguments), so it only works after terraform init, described next. Misspell bucket as buckett and validate reports:
Error: Unsupported argument
on main.tf line 12, in resource "aws_s3_bucket" "example":
12: buckett = "tf-first-demo-jsmith-4821" # replace with your own unique name
An argument named "buckett" is not expected here. Did you mean "bucket"?
The message gives the file, line, block and often a suggestion. Validate cannot know whether a bucket name is already taken. Only AWS can tell you that, at apply time.
terraform init
terraform init
Initializing the backend...
Initializing provider plugins...
- Finding hashicorp/aws versions matching "~> 6.0"...
- Installing hashicorp/aws v6.x.y...
- Installed hashicorp/aws v6.x.y (signed by HashiCorp)
Terraform has created a lock file .terraform.lock.hcl to record the provider
selections it made above. Include this file in your version control repository...
Terraform has been successfully initialized!
Read it top to bottom. "Backend" is where state is stored; the default is a local file, so nothing happens yet. Terraform then finds the newest AWS provider release matching ~> 6.0 and downloads it into a hidden .terraform/ directory. Finally it writes .terraform.lock.hcl, which records the exact version chosen (your v6.x.y will differ) plus checksums, so a teammate or CI job running init later gets the same plugin. The constraint says what is allowed; the lock file says what was picked.
terraform plan
Before you run it, predict: this is a new directory with one resource block. How many resources will the summary say are added, changed, destroyed?
terraform plan
Terraform used the selected providers to generate the following execution
plan. Resource actions are indicated with the following symbols:
+ create
Terraform will perform the following actions:
# aws_s3_bucket.example will be created
+ resource "aws_s3_bucket" "example" {
+ arn = (known after apply)
+ bucket = "tf-first-demo-jsmith-4821"
+ id = (known after apply)
# ... other attributes omitted
}
Plan: 1 to add, 0 to change, 0 to destroy.
Three things to read: the + marker means create; (known after apply) marks values Terraform cannot know yet, like the ARN (Amazon Resource Name, the unique identifier of an AWS object), which cannot exist until the bucket does; and the summary line (1 add, 0 change, 0 destroy) is the line to check before you do anything else. A plan only reads; it changes nothing in AWS.
terraform apply
terraform apply
Apply shows the same plan again, then asks:
Do you want to perform these actions?
Terraform will perform the actions described above.
Only 'yes' will be accepted to approve.
Enter a value: yes
aws_s3_bucket.example: Creating...
aws_s3_bucket.example: Creation complete after 2s [id=tf-first-demo-jsmith-4821]
Apply complete! Resources: 1 added, 0 changed, 0 destroyed.
Only the exact word yes proceeds. The id in brackets is the real bucket's identifier, which for S3 is its name. Confirm from outside Terraform with aws s3 ls, where the name should appear.
Now run terraform apply again:
aws_s3_bucket.example: Refreshing state... [id=tf-first-demo-jsmith-4821]
No changes. Your infrastructure matches the configuration.
Terraform has compared your real infrastructure against your configuration
and found no differences, so no changes are needed.
Apply complete! Resources: 0 added, 0 changed, 0 destroyed.
Refreshing state... is Terraform re-reading the real bucket before it compares. This is idempotence: running the same operation repeatedly leaves the same result, and the second run does nothing. A shell script with aws s3 mb would fail the second time: in eu-west-1, S3 rejects the repeated create with BucketAlreadyOwnedByYou.
Guided attempt: predict a change
Task. Add a tags argument inside the resource block:
tags = {
Project = "terraform-intro"
}
Before running terraform plan, predict the symbol next to the resource and the summary line. Then run it and compare.
Check your answer
The symbol is ~, meaning update in place, because AWS can add tags to an existing bucket without recreating it. Expect a + "Project" = "terraform-intro" line inside the tags map, and the summary Plan: 0 to add, 1 to change, 0 to destroy. You may also see a tags_all entry change; that is a computed copy of the tags that also includes any provider-level default tags. Apply it, then run aws s3api get-bucket-tagging --bucket <your-name> to see the tag.
terraform destroy
terraform destroy
# aws_s3_bucket.example will be destroyed
- resource "aws_s3_bucket" "example" {
- bucket = "tf-first-demo-jsmith-4821" -> null
# ... other attributes omitted
}
Plan: 0 to add, 0 to change, 1 to destroy.
Do you really want to destroy all resources?
Terraform will destroy all your managed infrastructure, as shown above.
There is no undo. Only 'yes' will be accepted to confirm.
Enter a value: yes
aws_s3_bucket.example: Destroying... [id=tf-first-demo-jsmith-4821]
aws_s3_bucket.example: Destruction complete after 1s
Destroy complete! Resources: 1 destroyed.
The - marker means delete, and -> null shows each value ceasing to exist. Destroy is just a plan whose target is an empty configuration, so the same read-before-you-type habit applies.
β οΈ If the bucket contains objects, destroy fails with a BucketNotEmpty error. Setting force_destroy = true on the resource makes Terraform empty the bucket first, which permanently deletes the data. The setting is stored in state, so it must be applied before the destroy takes effect. Leave it off for any bucket that holds anything you would miss.
Cost and safety
An empty bucket costs essentially nothing, so this exercise is safe. The habit to build now is cheaper to learn on a bucket than on a database: read the Plan: summary line, and for anything with destroy or replace in it, stop and ask what data lives there, before typing yes.
Practice: read and diagnose
Task. A teammate runs terraform plan (which succeeds) and then terraform apply on a new configuration with aws_s3_bucket.logs, and the apply fails with Error: creating S3 Bucket (logs): BucketAlreadyExists, even though validate passed. (1) Why did validate pass? (2) What do you change? (3) What do you expect plan to say after the fix?
Check your answer
- Validate checks only syntax and argument names against the provider's schema. It never asks AWS whether a name is taken. 2. Change
bucketto a globally unique name, for example with your account or team name and some digits appended. 3.Plan: 1 to add, 0 to change, 0 to destroy.The failed apply created nothing, so the state has no record of the bucket and plan again proposes a create.
Checklist. You can: explain each of block, resource type, local name, argument and reference address on aws_s3_bucket.example; say what init, plan, apply and destroy do in one sentence each; point to the summary line of a plan; and explain why force_destroy is a data-loss setting.
State: What Terraform Remembers, and What Belongs in Git
In the previous section, terraform destroy deleted exactly one bucket, the one you created, and nothing else in your AWS account. Your .tf file only says aws_s3_bucket.example, and that name means nothing to AWS. Something must have recorded which real bucket belongs to that block. That record is the state.
State in plain words
State is Terraform's record of which real resources belong to which block in your configuration. With the default local setup it is a JSON file named terraform.tfstate, written into your working directory after apply. Here is an abbreviated excerpt for the bucket from the last section (real state has more fields):
{
"version": 4,
"terraform_version": "1.9.8",
"resources": [
{
"mode": "managed",
"type": "aws_s3_bucket",
"name": "example",
"instances": [
{
"attributes": {
"id": "tf-first-demo-jsmith-4821",
"bucket": "tf-first-demo-jsmith-4821",
"arn": "arn:aws:s3:::tf-first-demo-jsmith-4821"
}
}
]
}
]
}
Read it as a lookup table: the block address (type + name, which gives aws_s3_bucket.example) maps to the real bucket's ID and attributes. Every plan uses three sources of truth:
Your .tf files (what you want)
β
terraform.tfstate (what Terraform believes it owns)
β
Real AWS (what actually exists; Terraform re-reads it during plan)
β
Plan = the smallest set of changes that makes AWS match your .tf files
When the three agree, you get "No changes." When a block has no entry in state, Terraform assumes the resource does not exist and plans to create it.
Inspect it safely
You never need to open the JSON. Two read-only commands show what Terraform tracks:
terraform state list
# aws_s3_bucket.example
terraform show
# # aws_s3_bucket.example:
# resource "aws_s3_bucket" "example" {
# arn = "arn:aws:s3:::tf-first-demo-jsmith-4821"
# bucket = "tf-first-demo-jsmith-4821"
# id = "tf-first-demo-jsmith-4821"
# ...
# }
state list prints one address per tracked resource. show prints the recorded attributes (abbreviated here). Both only read.
β οΈ Never hand-edit terraform.tfstate. The file has internal bookkeeping, and a typo can make Terraform lose track of a live resource. Terraform also has dedicated terraform state subcommands for safe repairs, which come later in this roadmap.
β οΈ Treat state like a credential. State stores the attributes of your resources in plaintext. Later in this roadmap that will include things like database passwords and other secrets passed to resources, even if you hide them in plan output. Anyone who can read the file can read them.
Thought experiment: delete the state
Suppose the bucket still exists in AWS, but you delete terraform.tfstate and run terraform plan. Predict the plan before reading on.
The plan shows:
# aws_s3_bucket.example will be created
+ resource "aws_s3_bucket" "example" {
+ bucket = "tf-first-demo-jsmith-4821"
...
}
Plan: 1 to add, 0 to change, 0 to destroy.
The block has no entry in state, so Terraform concludes it owns nothing and offers + create for a bucket that already exists. On apply, the create call fails because the name is already taken, by you: in eu-west-1 S3 answers BucketAlreadyOwnedByYou. For resource types without a unique name, such as an EC2 instance, the same situation would silently create a duplicate. Terraform cannot find existing resources by name alone. It knows only what state says it owns. (Bringing an existing resource back under management is called importing, and it is covered later.)
Why local state breaks for a team
Local state is fine for a solo exercise like yours. For a team it fails in two ways:
- No shared copy. Your laptop holds the state. A teammate running
planfrom their own checkout has no state and sees+ createfor everything, exactly as in the experiment above. - No locking between people. Locking means only one
applycan modify the state at a time. Local state is locked only against other runs on your own machine. If a team shares one state file by copying it around, two people can apply at once and overwrite each other's record.
The fix is remote state with locking: state stored in a shared location, with a lock taken during changes. This roadmap teaches it in the lesson "State and Modules". For now, remember that state must not live only on one laptop, and it must not live in Git either.
What belongs in Git
Your project directory now holds files of very different kinds. The rule: commit what describes the infrastructure, ignore what is generated, secret, or machine-specific.
| File | Git? | Reason |
|---|---|---|
*.tf |
Commit | Your configuration |
.terraform.lock.hcl |
Commit | Pins provider version and checksums |
.terraform/ |
Ignore | Downloaded plugins; init recreates it |
*.tfstate, *.tfstate.backup |
Ignore | Secrets, and wrong place for shared state |
crash.log |
Ignore | Debug dump; may include values |
*.tfvars |
Ignore | Input values; may hold secrets |
Two terms: .terraform.lock.hcl is the lock file init wrote, recording the exact AWS provider version and its checksums, so a teammate or CI job installs the identical plugin. Do not confuse it with state locking above. A *.tfvars file supplies values for input variables, which a later lesson covers. Commit it only if you are certain it holds nothing sensitive.
β οΈ Some .gitignore templates found online ignore .terraform.lock.hcl too. Remove that line. Without the committed lock file, teammates can get different provider versions.
Exercise: classify and write the .gitignore
Here is the directory listing of a small project:
infra/
βββ main.tf
βββ versions.tf
βββ outputs.tf
βββ terraform.tfvars
βββ terraform.tfstate
βββ terraform.tfstate.backup
βββ crash.log
βββ .terraform.lock.hcl
βββ .terraform/
β βββ providers/
βββ README.md
Task: mark each entry commit or ignore, then write the .gitignore that excludes the ignored ones. Success means a fresh git add . stages only files you can safely share.
Check your answer
Commit: main.tf, versions.tf, outputs.tf, .terraform.lock.hcl, README.md.
Ignore: terraform.tfvars, terraform.tfstate, terraform.tfstate.backup, crash.log, .terraform/.
# Downloaded providers; recreated by terraform init
.terraform/
# State and its backup (contain plaintext resource attributes)
*.tfstate
*.tfstate.*
# Crash dumps
crash.log
crash.*.log
# Variable files that may hold secrets
*.tfvars
The .gitignore itself is also committed. *.tfstate.* catches terraform.tfstate.backup.
What leaks: a committed state exposes every attribute of every resource, including account IDs, ARNs and any secret value passed to a resource, and it stays in Git history for everyone with repository access. A committed *.tfvars exposes whatever you put in it, such as passwords or API tokens. Adding a file to .gitignore after committing it does not remove it. You must run git rm --cached <file>, and because the history still contains it, treat the exposed values as compromised and rotate them.
Reading a Plan and Fixing a Configuration Before You Apply
Two plans can both end with Plan: 1 to add, 0 to change, 1 to destroy. One is a harmless new bucket plus a cleanup. The other plans to delete a live bucket and rebuild it empty. The summary line alone cannot tell them apart, so this section practises reading the symbols that can.
The symbol key
| Symbol | Meaning | Data at risk? |
|---|---|---|
+ |
create | No |
~ |
update in place | Usually no |
- |
destroy | Yes |
-/+ |
replace: destroy, then create | Yes |
A replace means Terraform cannot change the real resource to match your edit, so it deletes the resource and builds a new one. Terraform marks the argument that caused it with # forces replacement.
The contrast on our bucket starts from this configuration. The tags line is the edit:
resource "aws_s3_bucket" "example" {
bucket = "orders-logs-a1b2c3"
tags = {
Env = "dev" # added
}
}
The plan (trimmed; real output lists more attributes):
# aws_s3_bucket.example will be updated in-place
~ resource "aws_s3_bucket" "example" {
id = "orders-logs-a1b2c3"
~ tags = {
+ "Env" = "dev"
}
}
Plan: 0 to add, 1 to change, 0 to destroy.
Now change bucket to "orders-logs-prod" instead. A bucket's name is its identity in S3 and cannot be changed in place:
# aws_s3_bucket.example must be replaced
-/+ resource "aws_s3_bucket" "example" {
~ arn = "arn:aws:s3:::orders-logs-a1b2c3" -> (known after apply)
~ bucket = "orders-logs-a1b2c3" -> "orders-logs-prod" # forces replacement
}
Plan: 1 to add, 0 to change, 1 to destroy.
On a live bucket, applying this tries to delete the old bucket first, then create an empty one. With the default force_destroy = false, the delete fails with BucketNotEmpty while the bucket holds objects. Only if force_destroy = true was already applied does Terraform delete every object, permanently. Either way, the objects never move to the new name.
Reading order before typing yes:
- Read the summary line.
- Search for
must be replacedandforces replacement. - For every
-and-/+, ask: does this resource hold data I cannot recreate?
Guided attempt: predict the summary
You edited a configuration with four bucket blocks. The plan headers read:
# aws_s3_bucket.logs must be replaced
-/+ resource "aws_s3_bucket" "logs" { ... }
# aws_s3_bucket.assets will be updated in-place
~ resource "aws_s3_bucket" "assets" { ... }
# aws_s3_bucket.archive will be created
+ resource "aws_s3_bucket" "archive" { ... }
# aws_s3_bucket.old will be destroyed
# (because aws_s3_bucket.old is not in configuration)
- resource "aws_s3_bucket" "old" { ... }
Task: write the summary line, then name which resources need a "does it hold data?" check.
Check your answer
Plan: 2 to add, 1 to change, 2 to destroy.
A replace counts on both sides: it adds one (the new logs) and destroys one (the old logs). So add = archive + new logs = 2. Change = assets = 1. Destroy = old logs + old = 2. The data check applies to logs and old. Notice that old is destroyed simply because you deleted its block, with no forces replacement annotation involved.
Fix the configuration
A teammate hands you this and says it "should just work":
terraform {
required_version = ">= 1.0"
required_providers {
aws = {
version = "~> 6.0"
}
}
}
provider "aws" {
region = "eu-west-1"
}
resource "aws_s3_bucket" "logs" {
bucket = "Orders_Logs"
}
Task: find three faults. For each, say which command would catch it, if any, and write the repair. Mentally run init, validate and plan.
Check your answer
| Fault | Caught by | Repair |
|---|---|---|
No source for aws |
Nothing here | source = "hashicorp/aws" |
Orders_Logs |
Not validate or plan; the AWS provider, at apply |
lowercase, hyphens |
>= 1.0 |
Nothing | "~> 1.9" |
- Missing source. For a HashiCorp provider, Terraform infers
hashicorp/aws, so nothing fails here. Write it explicitly anyway. The inference breaks for providers from other publishers, and a reader should not have to guess where a plugin comes from. - Bucket name. S3 names must be lowercase letters, digits and hyphens (dots are allowed too, but avoid them).
validatechecks syntax and argument names, not S3's naming rules, andplandoes not check them either. In eu-west-1 the AWS provider rejects the name whenapplystarts to create the bucket (validating S3 Bucket (Orders_Logs) name: only lowercase alphanumeric characters and hyphens allowed in "Orders_Logs"), before it sends any request to S3. Nothing gets created, but you learn late. - Version.
>= 1.0accepts any future major version, so a teammate on a very different release behaves differently from you.~> 1.9allows 1.9 and above but stays below 2.0.
terraform {
required_version = "~> 1.9"
required_providers {
aws = {
source = "hashicorp/aws"
version = "~> 6.0"
}
}
}
provider "aws" {
region = "eu-west-1"
}
resource "aws_s3_bucket" "logs" {
bucket = "orders-logs-a1b2c3"
}
The lesson: validate passing does not mean the configuration is right.
Harder variation: two problems at once
A teammate's terraform plan behaves differently from yours, and you find they run a Terraform release from three years ago. Looking around, you also find terraform.tfstate committed to the shared Git repository. The repo has no required_version.
Task: order these three fixes and justify each: (a) pin required_version, (b) remove the file from Git, (c) treat the exposed state as sensitive.
Check your answer
One defensible order is (b), then (c), then (a).
- Remove from Git first, but keep the local file. Run
git rm --cached terraform.tfstateand add*.tfstateto.gitignore. Every new clone or push spreads the exposure. Do not delete the file itself: with local state it is Terraform's only record of what it owns. - Treat it as sensitive. Removing the file does not remove it from Git history, and anyone with repo access may have copied it. Inspect it with
terraform showfor secrets and rotate any you find. Consider history rewriting as a separate decision. - Pin the version. Version drift is a nuisance, not yet damage, but HashiCorp does not guarantee that an older release can read state a newer one wrote. Set
required_version = "~> 1.9"(any range that includes the version the team actually runs), commit it, and upgrade the teammate. From then on, a CLI outside the range stops withError: Unsupported Terraform Core versionbefore doing anything.
The real cure for the shared-file problem is remote state with locking, which the lesson "State and Modules" covers.
Transfer task
A pull request changes bucket = "orders-data" to "orders-data-v2" on a bucket that holds production exports. The plan ends with Plan: 1 to add, 0 to change, 1 to destroy.
Task: state what will happen, whether you approve, and what you ask the author.
Check your answer
This is a replace, not an unrelated add plus destroy, and you can confirm it because bucket carries # forces replacement. Terraform will try to delete the old bucket and create an empty one: with the default force_destroy = false the delete fails with BucketNotEmpty, and if force_destroy = true is already applied the objects are permanently deleted. Either way the exports do not move to the new name. Do not approve. Ask the author to revert the rename, or to explain how the data will be copied first. The counts alone would look identical to a safe change, so the annotation and the data question are what decide it.
Self-check
- Can you say each command in one sentence?
initdownloads providers and writes the lock file;planpreviews changes without touching AWS;applymakes the changes after you confirm;destroydeletes everything Terraform manages here. - Can you name the three things never to commit:
.terraform/,*.tfstate(and its backup), and*.tfvarsfiles that may hold secrets? - Can you explain what
-/+means, and which question to ask before approving it?
Where the roadmap goes next
- Terraform Core and The Plan and Apply Loop: how configuration, state and plan fit together, including variables, so names and settings are not hard-coded.
- State and Modules: remote state with locking, so a team shares one safe copy.
- A Serverless AWS Stack in Terraform: IAM, Lambda, SQS and DynamoDB resources, where the same plan-reading habit protects real data and live traffic.