CloudWatch Beyond Metrics

Finding and reading logs first: log groups, log streams and events; where Lambda, ECS on Fargate (the awslogs driver) and API Gateway write their logs; retention; tailing with aws logs tail; filter patterns; CloudWatch Logs Insights queries (fields, filter, parse, stats, sort, limit) to trace one request ID across services; structured JSON logging and what it enables. Then metric filters, custom metrics and the embedded metric format, alarms and composite alarms, anomaly detection, and what each of these costs.

Last generated

Lesson 24 of 31 available15 practice questions

SPACED REPETITION Β· 15 practice questions

Make this lesson stick.

Try 3 questions now. No account needed. Sample answers aren't saved.

Finding and Reading Logs: Where Each Service Writes and How to Tail Them

You deploy a function, send a request, and get a 500. You open CloudWatch to find out why, and there is nothing there. No error and no trace of the request. Before you can debug anything, you have to know where each AWS service is supposed to write its logs and which setting or permission stops it from doing so.

This section builds that map for Lambda, ECS on Fargate and API Gateway. It then shows how to read the logs quickly from the terminal with aws logs tail and filter patterns, and how to stop retention defaults from quietly inflating your bill.

Three nouns that organize everything

  • A log event is a timestamp plus a message: one line your code printed, or one record a service emitted.
  • A log stream is the ordered sequence of events from one source, such as one Lambda execution environment (the sandbox that runs your function code) or one container.
  • A log group is the container of streams. Retention, encryption (KMS key) and access permissions are set on the group, so every stream inside it shares them.
log group  /aws/lambda/orders-api      <- retention, KMS, IAM live here
   ↓
log stream  one per execution environment
   ↓
log event   timestamp + message

Prediction task: where do the logs land?

Before reading further, commit to answers. You have three things deployed:

  1. A Lambda function named orders-api.
  2. An ECS task (one running copy of a task definition) on Fargate (AWS-managed container compute). Its container is named worker, and it uses the awslogs log driver, which ships the container's stdout and stderr to CloudWatch Logs. The driver options are awslogs-group=/ecs/orders-worker and awslogs-stream-prefix=orders.
  3. A REST API in API Gateway with a stage named prod. A stage is a named deployment of the API. The API's ID is a1b2c3. Nobody has touched its logging settings.

Task: For each one, write the log group name and the stream name pattern. One of the three will have no log group at all. Say which and why.

Check your answer
  • Lambda: group /aws/lambda/orders-api. Streams are named like date/[version]id, one per execution environment (for example .../[$LATEST]3f1c9...).
  • ECS: group /ecs/orders-worker. Streams follow prefix/container-name/task-id, so orders/worker/<task-id>.
  • API Gateway: nothing yet. Neither execution logs nor access logs exist until you enable them on the stage. Once execution logging is on, a REST API writes to API-Gateway-Execution-Logs_a1b2c3/prod. Access logs go to a group you choose.

The pattern: Lambda logs automatically, while ECS and API Gateway only log where you configured them to.

Lambda: automatic, if the role allows it

Lambda's runtime captures what your code prints and sends it, by default, to /aws/lambda/<function-name> (a function's logging configuration can name a different group). Each execution environment writes to its own stream, so concurrent invocations scatter across several streams. This is why you search at the group level rather than guessing a stream.

The usual cause of "no logs" is that the function's execution role (the IAM role the function assumes when it runs) lacks logs:CreateLogGroup, logs:CreateLogStream and logs:PutLogEvents. The managed policy AWSLambdaBasicExecutionRole grants exactly these. Lambda does not fail your invocation when log delivery is denied, so the function succeeds silently and logs nothing.

ECS on Fargate: you wire it up

Nothing is logged until the container definition says where. Here is the relevant fragment of a task definition:

"logConfiguration": {
  "logDriver": "awslogs",
  "options": {
    "awslogs-group": "/ecs/orders-worker",
    "awslogs-region": "us-east-1",
    "awslogs-stream-prefix": "orders"
  }
}

The stream name becomes orders/worker/<task-id>, where worker is the container's name in the task definition. On Fargate the prefix option is required.

The permissions belong on the task execution role, which the ECS agent uses to pull images and write logs. They do not belong on the task role, which your application code uses for its own AWS calls. The execution role needs logs:CreateLogStream and logs:PutLogEvents. If you set awslogs-create-group to "true", it also needs logs:CreateLogGroup. Putting the permission on the wrong role is a classic mistake because both roles look similar in the console.

API Gateway: two log types, both off by default

A REST API has two separate kinds of logs:

  • Execution logs record what API Gateway did internally (request validation, integration calls, errors). You switch them on per stage by setting a log level (ERROR or INFO). They go to the auto-named group API-Gateway-Execution-Logs_<api-id>/<stage>.
  • Access logs record one line per request in a format you define, written to a log group you create and name.

Both need the account-level CloudWatch Logs role: a setting in API Gateway (per region) that holds the ARN of an IAM role allowing API Gateway to write to CloudWatch. Until it is set, enabling logging on a REST stage fails or produces nothing. HTTP APIs (the lighter, cheaper API type) offer access logs only, with no execution log level.

Tailing from the CLI

The console is slow when you are chasing a live problem. In AWS CLI v2, aws logs tail reads a whole group, merging all its streams:

# Live view of the last 15 minutes, then keep following
aws logs tail /aws/lambda/orders-api --follow --since 15m

# Only events containing the term ERROR
aws logs tail /aws/lambda/orders-api --since 15m --filter-pattern ERROR

# Narrow to streams of one container (prefix match on stream name)
aws logs tail /ecs/orders-worker --since 1h --log-stream-name-prefix orders/worker

A filter pattern is a small matching language the service applies server-side, so you only download matching events. Patterns are case-sensitive.

Goal Pattern Notes
One term ERROR Matches events containing it
Exact phrase "connection refused" Quotes keep words together
Either term ?ERROR ?Exception ? means any of
JSON field { $.level = "ERROR" } Needs JSON log lines
Numeric JSON { $.status >= 500 } Compares numbers

In a shell, wrap patterns in single quotes so the inner double quotes survive. Filter patterns match events; they do not compute anything. They cannot count errors per minute or average a latency, which is the job of Logs Insights (covered in the next section, 'Structured JSON Logs and Logs Insights: Tracing One Request ID Across Services').

Guided attempt: find a 5xx burst without the console

Assume the prod stage writes JSON access logs to /apigw/orders-prod/access with an unquoted numeric status field.

  1. Start wide: aws logs tail /apigw/orders-prod/access --since 30m --filter-pattern '{ $.status >= 500 }' --format short. Note the timestamps. If the lines bunch around one minute, that is the burst.
  2. Zoom in: rerun with --since 5m and --follow to see whether it is still happening.
  3. Your turn: the access lines show 502 responses. Which log group do you tail next, and with which pattern, to see whether the Lambda behind it raised errors?
Check your answer

Tail /aws/lambda/orders-api with the same time window and --filter-pattern '?ERROR ?Task'. Lambda logs Task timed out on a timeout, and a 502 typically follows an integration failure like that. Matching the two timestamp clusters tells you whether the problem is in the function or in API Gateway itself.

Retention: the cost you set once

New log groups never expire, including those that Lambda or awslogs-create-group creates for you. You pay for ingestion (data sent in) and storage (data kept), and ingestion is usually the larger share. Retention only controls the storage part, but a never-expiring group makes that part grow forever.

aws logs put-retention-policy --log-group-name /aws/lambda/orders-api --retention-in-days 14

Pick per group: around 14 days for debug-heavy services, longer where compliance demands it. Create groups yourself (ideally in your infrastructure-as-code) with retention declared, instead of letting services auto-create them.

Independent variation: Fargate shows no logs after deploy

A new orders-worker service deploys, but /ecs/orders-worker shows nothing. Here is what you are given.

"executionRoleArn": "arn:aws:iam::111122223333:role/orders-exec-role",
"taskRoleArn": "arn:aws:iam::111122223333:role/orders-task-role",
"logConfiguration": {
  "logDriver": "awslogs",
  "options": {
    "awslogs-group": "/ecs/orders-worker",
    "awslogs-region": "us-east-1",
    "awslogs-stream-prefix": "orders"
  }
}

orders-exec-role allows only ecr:GetAuthorizationToken, ecr:BatchGetImage and ecr:GetDownloadUrlForLayer. orders-task-role allows logs:CreateLogStream and logs:PutLogEvents. Nobody created the log group.

Task: name the two possible causes and justify your fix.

Check your answer
  1. Missing permission on the right role. The logs permissions sit on the task role, but the awslogs driver uses the execution role. Add logs:CreateLogStream and logs:PutLogEvents (scoped to the group's ARN) to orders-exec-role.
  2. The group does not exist, and awslogs-create-group is not set. Either create /ecs/orders-worker ahead of time or set awslogs-create-group to "true" and also grant logs:CreateLogGroup on the execution role.

Pre-creating the group with a retention policy is the better fix, since an auto-created group never expires. When a task fails this way it usually stops early, so aws ecs describe-tasks and its stopped reason are the next place to look.

Checklist: Can you name the group and stream pattern for each service? Can you say which role or setting causes missing logs for each? Can you tail a group with a filter pattern? Have you set retention on every group you own?

Structured JSON Logs and Logs Insights: Tracing One Request ID Across Services

A customer says order r-42 "just spun and failed." The request passed through API Gateway, a Lambda function and a container service, so the evidence sits in three log groups (named containers of log streams that share retention and access settings). With the tail-and-filter skills from the previous section, you can search each group one at a time. This section shows how to follow one request through all three in a single query.

The naive approach and where it stops

Searching free-text lines works inside one service. You run aws logs tail with --filter-pattern '"r-42"' (double quotes inside the single quotes, because the term contains a hyphen) and read what comes back. It breaks down across services for three reasons:

  • Each service writes the ID in a different place and format (id=r-42, RequestId: r-42, [r-42]).
  • You must run three searches and merge the results by eye, comparing timestamps yourself.
  • Text has no schema, so you cannot ask for "latency by route" without hand-parsing every line.

The fix is to make the logs queryable by construction.

Structured logging: one JSON object per line

Structured logging means each log event is a single JSON object with consistent field names. A line from the Lambda function might look like this:

{"level":"INFO","message":"charge ok","service":"orders-api","requestId":"r-42","route":"POST /orders","latencyMs":87}

CloudWatch Logs Insights is the query tool for log groups. It auto-discovers top-level JSON fields, so filter level = 'ERROR' works with no extraction step. Nested fields use dot notation, such as ctx.userId. Plain-text events expose only built-in fields such as @timestamp, @message, @logStream and @log (the source log group). Lambda's own report lines also carry discovered fields such as @requestId, @duration, @billedDuration, @memorySize and @maxMemoryUsed. Anything else must be pulled out with parse, covered below.

Here is a minimal Python helper in the Lambda that keeps the schema consistent. It also forwards the ID downstream:

import json
import time
import urllib.request

def log(level, message, **fields):
    # One JSON object per line: Lambda sends stdout to the function's log group
    print(json.dumps({'level': level, 'message': message, 'service': 'orders-api', **fields}))

def handler(event, context):
    # The edge-assigned ID; present in REST and HTTP API proxy events
    request_id = event['requestContext']['requestId']
    start = time.time()
    log('INFO', 'received order', requestId=request_id)
    req = urllib.request.Request(
        'http://payments.internal/charge', data=b'{}',
        headers={'X-Request-Id': request_id, 'Content-Type': 'application/json'},
    )
    urllib.request.urlopen(req, timeout=3)  # pass the ID downstream
    log('INFO', 'charge ok', requestId=request_id, route='POST /orders',
        latencyMs=int((time.time() - start) * 1000))
    return {'statusCode': 200, 'body': '{}'}

Correlation ID propagation

A correlation ID is one identifier attached to every log line produced while serving a single request. Take it at the edge: API Gateway generates $context.requestId, and you can also honor an incoming X-Request-Id header. Then do two things at every hop: put it in every log line, and forward it to the next service.

Client
  ↓
API Gateway   assigns $context.requestId = r-42 (access log)
  ↓
Lambda        reads event.requestContext.requestId, logs it
  ↓  header X-Request-Id: r-42
Container service   reads header, logs it
  ↓  SQS message attribute requestId = r-42
Worker        must read the attribute and log it

The API Gateway access log format is itself JSON. This example uses HTTP API variable names, and REST APIs use others such as $context.resourcePath:

{"service":"api-gateway","message":"access","requestId":"$context.requestId","route":"$context.routeKey","status":$context.status,"latencyMs":$context.responseLatency}

⚠️ The chain usually breaks at queues and async calls. An SQS consumer or an event-driven function receives only what the producer put in the message, so the ID must travel as a message attribute or body field. It does not travel automatically.

Worked trace: rebuild one request's timeline

In Logs Insights, select all three log groups (/aws/apigateway/orders-access, /aws/lambda/orders-api, /ecs/payments) and a narrow time window. Then run:

fields @timestamp, service, level, message
| filter requestId = 'r-42'
| sort @timestamp asc
| limit 50
@timestamp service level message
10:00:01.050 orders-api INFO received order
10:00:01.090 payments INFO charge started
10:00:03.080 payments ERROR card processor timeout
10:00:03.100 orders-api ERROR charge failed
10:00:03.120 api-gateway access

The story is now clear. The payments service waited about two seconds on the card processor. Notice that the gateway row comes last even though the gateway received the request first, because an access log is written when the response completes. Sorting by @timestamp shows when each line was written, not when the request entered each system.

The commands you will use most often:

Command What it does
fields Choose columns to display
filter Keep only matching events
parse Extract new fields from text
stats Aggregate (count, avg, pct)
sort Order the rows
limit Cap the rows returned

Commands chain with |, and each one receives the previous one's output. The same query can run from the CLI. start-query returns a query ID, and get-query-results fetches the rows:

# date -d is GNU date (Linux, CloudShell); on macOS use $(date -v-1H +%s) for the start time
aws logs start-query \
  --log-group-names /aws/lambda/orders-api /ecs/payments \
  --start-time $(date -d '-1 hour' +%s) --end-time $(date +%s) \
  --query-string "fields @timestamp, service, message | filter requestId = 'r-42' | sort @timestamp asc"

Guided attempt: p95 latency by route

Task: Every orders-api line that carries latencyMs also has a route field. Write a query that shows p95 latency and request count per route in 5-minute bins. Then write a second query for error count per route.

Check your answer
filter service = 'orders-api' and ispresent(latencyMs)
| stats pct(latencyMs, 95) as p95, count(*) as requests by bin(5m), route

bin(5m) buckets each event's timestamp into 5-minute windows, and by makes one output row per bin and route combination. ispresent drops lines such as "received order" that have no latency. For errors, filter first and then count:

filter service = 'orders-api' and level = 'ERROR'
| stats count(*) as errors by bin(5m), route

The filter runs before stats, so one stats clause cannot mix "all requests" and "only errors" in this simple form. Run two queries.

Parsing unstructured lines

Legacy services often log plain text such as 2026-03-02T10:15:03Z INFO order created id=r-42 took 87ms. parse extracts fields from @message at query time. It supports a glob pattern, where * is a capture slot, or a regular expression with named groups:

parse @message 'id=* took *ms' as requestId, ms
| filter requestId = 'r-42'
parse @message /id=(?<requestId>[\w-]+)/
| filter requestId = 'r-42'

parse works, but it is fragile. If someone rewrites the message to order created (req r-42), the pattern stops matching and the field is silently empty, with no error. The extraction is also repeated in every query by every person. Logging JSON from the start moves that work to write time, where the schema is defined once.

Cost and pitfalls

  • ⚠️ Insights bills by data scanned. Choose the narrowest time range and only the log groups you need before you run anything.
  • ⚠️ limit does not reduce scanning. It trims the output after the scan, so limit 10 over 30 days costs the same as without it. (The separate limit any 10 form does stop scanning early, but it returns an arbitrary, unordered sample.)
  • Multi-line stack traces are often split into one event per line (the awslogs driver does this by default; on Lambda it depends on the runtime and how you log), and then only the first line carries your JSON, so the others have no requestId. Log the stack as a single JSON string field (newlines escaped) instead.
  • High-cardinality or huge fields, such as full request payloads, inflate ingestion and every later scan. Log IDs, sizes and status codes instead. (Cardinality, the number of distinct values a field takes, matters even more for metrics, as we'll see in the next section.)

Independent variation: the worker that dropped the ID

The API, a queue worker and a database proxy all handled order o-77. You are given these events, already in time order:

[api]     {"service":"api","requestId":"r-42","orderId":"o-77","message":"order accepted, enqueued"}
[worker]  {"service":"worker","orderId":"o-77","message":"charge failed: timeout"}
[proxy]   slow query 2400ms tag=o-77

Task: Find where the requestId chain breaks. Then write one query over all three groups that rebuilds the timeline for r-42 anyway.

Check your answer

The chain breaks at the worker, the queue hop, because the producer did not put requestId in the message and the worker did not log it. The proxy only ever saw a text tag=. The API line is the bridge, since it contains both requestId and orderId, so use orderId as a secondary key:

fields @timestamp, service, message
| parse @message /tag=(?<tagId>[\w-]+)/
| filter requestId = 'r-42' or orderId = 'o-77' or tagId = 'o-77'
| sort @timestamp asc

The parse is needed only for the plain-text proxy line, and it yields an empty tagId on JSON lines, which is harmless. The durable fix is to send requestId as an SQS message attribute and log it in the worker.

From Logs to Numbers: Metric Filters, Custom Metrics and Embedded Metric Format

Logs Insights can answer almost any question about past events, but an alarm (a rule that watches a number and notifies someone) cannot watch a log search. It needs a metric, a named series of numeric datapoints over time. So the question for each signal is how to turn application behavior into a metric without paying for a metric you never graph. Three routes exist, and they differ in who does the work and what each unit of detail costs.

Route 1  app logs JSON ──→ log group ──→ metric filter ──→ metric
Route 2  app code ──→ PutMetricData API call ───────────→ metric
Route 3  app logs JSON with _aws block ──→ log group ──→ metric (extracted automatically)
                                            └──→ raw event stays searchable in Insights

Route 1: Metric filters on logs you already have

A metric filter is a rule attached to a log group. It uses the same pattern syntax as the filter patterns from the first section. Each time a newly ingested event matches, CloudWatch adds a datapoint to a metric. This turns the JSON error logs into an ErrorCount metric:

aws logs put-metric-filter \
  --log-group-name /aws/lambda/orders-api \
  --filter-name ErrorCount \
  --filter-pattern '{ $.level = "ERROR" }' \
  --metric-transformations \
    metricName=ErrorCount,metricNamespace=OrdersApp,metricValue=1,defaultValue=0

metricValue=1 means every matching event contributes 1. defaultValue=0 publishes a 0 for each one-minute period in which the group receives log events but none of them match, so a service that is logging without errors shows a flat zero instead of gaps. If the group receives no events at all in a minute, nothing is published, and an alarm still sees that minute as missing data.

Predict before reading on: you create this filter at 10:00. Errors were logged at 9:50. What does ErrorCount show for 9:50?

Check your answer

Nothing. A metric filter only evaluates events ingested after it exists, and there is no backfill. For the 9:50 errors you would run a Logs Insights query over that time range instead.

A filter can also pull a number out of the event and set dimensions. A dimension is a name/value label that splits one metric into separate series, such as Route=/orders. This filter publishes latencyMs as the value and splits by route:

aws logs put-metric-filter \
  --log-group-name /aws/lambda/orders-api \
  --filter-name RequestLatency \
  --filter-pattern '{ $.latencyMs = * }' \
  --metric-transformations '[{"metricName":"RequestLatency","metricNamespace":"OrdersApp","metricValue":"$.latencyMs","unit":"Milliseconds","dimensions":{"Route":"$.route"}}]'

Because the metric keeps real values, you can graph p95 or p99 of it later. The pattern { $.latencyMs = * } matches any event that has the field. Filters are a quick win because they need no code change. Their limits are that you can only extract what the log line already contains, and a filter that sets dimensions cannot also use defaultValue.

Route 2: Calling PutMetricData directly

The PutMetricData API publishes a datapoint you compute yourself. It is the right tool when no log line exists to derive the number from, for example a scheduled function that reads a vendor's queue depth every minute:

import boto3
cloudwatch = boto3.client("cloudwatch")

def publish_depth(depth):
    cloudwatch.put_metric_data(
        Namespace="OrdersApp",
        MetricData=[{
            "MetricName": "VendorQueueDepth",
            "Dimensions": [{"Name": "Queue", "Value": "billing"}],
            "Value": depth,
            "Unit": "Count",
        }],
    )

For request-path metrics this route has real costs. Every call is a network round trip that adds latency to the request. It needs an extra IAM permission (cloudwatch:PutMetricData) and its own error handling. Under load it can be throttled, and you can lose datapoints unless you batch and retry. That is a whole new code path to maintain. Use it for metrics with no log line behind them, and prefer Route 3 for everything an invocation or request already knows.

Route 3: Embedded metric format (EMF)

The embedded metric format is a JSON log line with a reserved _aws block that tells CloudWatch which fields to turn into metrics. CloudWatch extracts the metrics as the line is ingested, and the full event stays in the log group. Here is a Lambda handler emitting OrderPlaced and OrderLatency:

import json, time

def handler(event, context):
    start = time.time()
    order = place_order(event)  # your business logic
    latency_ms = (time.time() - start) * 1000

    print(json.dumps({
        "_aws": {
            "Timestamp": int(time.time() * 1000),
            "CloudWatchMetrics": [{
                "Namespace": "OrdersApp",
                "Dimensions": [["Service", "Route"]],
                "Metrics": [
                    {"Name": "OrderPlaced", "Unit": "Count"},
                    {"Name": "OrderLatency", "Unit": "Milliseconds"}
                ]
            }]
        },
        "Service": "orders-api",
        "Route": "POST /orders",
        "OrderPlaced": 1,
        "OrderLatency": latency_ms,
        "requestId": context.aws_request_id,
        "userId": order["userId"],
        "level": "INFO",
        "message": "order placed"
    }))

The _aws block is the instruction sheet. Dimensions lists which top-level fields become labels, here Service and Route together. Metrics names which top-level fields are values, OrderPlaced and OrderLatency. Every other field, such as userId and requestId, is only a property: it stays in the log event, searchable in Insights, but never becomes a metric label. This example prints one line per invocation, so it costs one log event and no API call. Later you can still run filter userId = "u-123" in Insights against the same line.

Getting EMF to CloudWatch by runtime
Runtime Route Note
Lambda Print the JSON to stdout Lambda already ships stdout to the function's log group
ECS/Fargate or EC2 Client library, or stdout via the log driver Libraries typically send to a CloudWatch agent endpoint; verify the library's environment setting

The point is that EMF is extracted from whatever lands in a log group. In a container, either run the CloudWatch agent as a sidecar (a helper container in the same task) and point the EMF library at it, or configure the library to write to stdout and let the awslogs driver ship it.

The cost trap: cardinality

Cardinality means the number of distinct values a field takes. CloudWatch bills each unique combination of metric name and dimension values as its own custom metric, and metric filters that set dimensions follow the same rule. Take OrderLatency with a customerId dimension. The dollar figures use an illustrative tiered rate of $0.30 per metric per month for the first 10,000 metrics and $0.10 for the next 240,000 (the rate drops again above 250,000); check current regional pricing.

  • 5 customer IDs: 5 metrics Γ— $0.30 = $1.50/month.
  • 50,000 customer IDs: 10,000 Γ— $0.30 = $3,000, plus 40,000 Γ— $0.10 = $4,000, giving about $7,000/month.

The code did not change, only the data did. Listing several dimension sets, such as [["Service"],["Service","Route"]], multiplies the count again. ⚠️ Log ingestion is billed on top of whichever route you pick, and these are metric costs only. A dimension should have a small, bounded set of values you would actually put on a dashboard axis.

Decision task

For each case, choose metric filter, EMF or logs-only (Insights queries, no metric), and justify it by cost and query needs.

  1. Count 5xx responses from an access log your load balancer already writes.
  2. Per-route latency percentiles for a new Lambda with 12 routes.
  3. Usage per tenant across 20,000 tenants, mostly for monthly reports and occasional support questions.
Check your answer
  1. Metric filter. The log line already exists and you only need a count, so no code changes. Remember it counts only from the moment the filter exists.
  2. EMF. It is new code, 12 routes is a bounded dimension, and you want p95/p99 plus the raw event in Insights, without an API call per request.
  3. Logs-only. Tenant-level metrics would be about 20,000 metrics, roughly $4,000/month at the illustrative rate for each metric name. Log tenantId as a field and run stats count() by tenantId in Insights when someone asks. Optionally emit an aggregate metric with no tenant dimension.

Independent variation

A teammate's EMF design for checkout failures is: Dimensions: [["Service","userId"]], Metrics: [{"Name":"CheckoutFailed"}], with 200,000 active users. Rewrite it so a support engineer can still investigate one user's failures without per-user metrics.

Check your answer

Change it to Dimensions: [["Service"]] (or [["Service","Route"]] if routes are bounded), and keep userId as a top-level property in the same log line. The metric stays at a handful of series, and the alarm watches the aggregate. To investigate a user, run an Insights query such as filter userId = "u-123" and CheckoutFailed = 1 | sort @timestamp desc, scoped to the relevant log group and time range. The detail is kept, but you pay for it in log storage and query scanning, not in 200,000 metrics.

Checklist: I can say what a metric filter will not count; I can choose between the three routes; I can read an _aws block; and I can estimate what a dimension costs before shipping it.

Alarms That Page the Right People: Alarms, Composite Alarms and Anomaly Detection

At 03:12 your phone buzzes: CPU is at 82% on a web fleet. Customers are fine and the nightly batch job just started. By Friday the team has muted the channel, and the first real outage goes unnoticed. This section is about building alarms that make the opposite trade: fewer pages, each one meaning a human should act. You will use the metrics you created in 'From Logs to Numbers: Metric Filters, Custom Metrics and Embedded Metric Format' as alarm inputs.

Alarm anatomy

A CloudWatch alarm watches one metric (or a metric-math expression, a formula over metrics such as errors divided by invocations) and moves between three states: OK (not breaching), ALARM (breaching) and INSUFFICIENT_DATA (not enough datapoints to decide). Five settings decide when it moves:

  • Statistic: how raw samples are summarized (Sum, Average, p99).
  • Period: the length in seconds of each datapoint (60 means one datapoint per minute).
  • Evaluation periods (N): how many of the most recent datapoints form the window.
  • Datapoints to alarm (M): how many datapoints in that window must breach. This is the M of N rule, and the breaching datapoints need not be consecutive.
  • Threshold plus a comparison operator, which define what 'breaching' means.
# Lambda 'checkout': alarm when Errors >= 5 in at least 3 of the last 5 minutes
aws cloudwatch put-metric-alarm \
  --alarm-name checkout-lambda-errors \
  --namespace AWS/Lambda --metric-name Errors \
  --dimensions Name=FunctionName,Value=checkout \
  --statistic Sum --period 60 \
  --evaluation-periods 5 --datapoints-to-alarm 3 \
  --threshold 5 --comparison-operator GreaterThanOrEqualToThreshold \
  --treat-missing-data notBreaching \
  --alarm-actions arn:aws:sns:us-east-1:111122223333:oncall-page

This uses a raw count to keep the flags simple. An error-rate alarm replaces the single metric with a metric-math expression such as Errors divided by Invocations times 100, supplied through --metrics.

Worked trace: 3 of 5 on an error rate

Setup: error rate above 5% counts as breaching, the period is 1 minute, N=5 and M=3. The four minutes before minute 1 were healthy, so the window is full from the start. An incident then hits.

Minute Error rate Breaching in window State
1 9% 1 of 5 OK
2 12% 2 of 5 OK
3 15% 3 of 5 ALARM (fires)
4 2% 3 of 5 (minutes 1-3) ALARM
5 1% 3 of 5 (minutes 1-3) ALARM
6 1% 2 of 5 (minutes 2-3) OK (recovers)

The alarm fires at minute 3, once the third bad datapoint arrives. It recovers at minute 6, three minutes after the last bad datapoint, because the window must slide past the bad minutes. M of N therefore has memory. That memory suppresses flapping, which is repeated flipping between OK and ALARM.

Now hold a continuous outage constant and vary only the period. These times are approximate, since metric publishing adds a delay of a minute or so.

Period M of N Datapoints needed Time to page
1 min 3 of 5 3 about 3 min
5 min 3 of 5 3 about 15 min
5 min 1 of 1 1 about 5 min

The 5-minute 3-of-5 setting is calm but slow, and a short sharp outage can average below the threshold inside a 5-minute bucket. The 5-minute 1-of-1 setting is faster, but a single noisy bucket pages someone. Choose by asking how many minutes of customer pain you can tolerate before a human knows.

Missing data: when the metric goes quiet

Metrics are only published when something happens. A Lambda function that is not invoked emits no Invocations datapoint, and an alarm needs a rule for those gaps. The treat missing data setting has four options:

  • missing (the default): the alarm evaluates only datapoints that exist, and goes INSUFFICIENT_DATA if the whole window is empty.
  • notBreaching: a gap counts as a healthy datapoint.
  • breaching: a gap counts as a bad datapoint.
  • ignore: the alarm keeps its current state.

The right choice depends on what silence means. For an error alarm on a low-traffic function, silence at 4 a.m. is fine, so notBreaching keeps it OK instead of INSUFFICIENT_DATA. For a heartbeat alarm such as 'Invocations < 1' on a function a schedule should trigger every 5 minutes, silence is the failure, so you need breaching. ⚠️ With the default, a dead scheduled job produces no data, the alarm goes INSUFFICIENT_DATA, and nobody is paged. An INSUFFICIENT_DATA alarm only notifies you if you attached an action to that state.

Alarm actions and what deserves one

An alarm can run actions on any state transition. The usual action is publishing to an SNS topic (a pub/sub channel that fans out to email, chat or a paging tool). Alarms can also trigger an Auto Scaling policy or an EC2 action (stop, terminate, reboot or recover an instance).

That split suggests a rule of thumb. Cause metrics like CPU belong in scaling actions, where a machine reacts to a machine. Symptom metrics are what users feel, so they should page people. Typical symptoms are 5xx error rate, p99 latency (the time within which 99% of requests finish) and queue age (how long the oldest message has waited). A CPU alarm that pages with no user impact is the 03:12 problem from the opening.

Composite alarms: pages only on co-occurrence

A composite alarm has no metric of its own. It evaluates a rule over other alarms, using ALARM(name), OK(name), AND, OR and NOT. The child alarms carry no actions, so only the composite notifies anyone.

# Children (checkout-p99-latency, checkout-5xx-rate) are created WITHOUT --alarm-actions
aws cloudwatch put-composite-alarm \
  --alarm-name checkout-slow-and-failing \
  --alarm-rule 'ALARM(checkout-p99-latency) AND ALARM(checkout-5xx-rate)' \
  --alarm-actions arn:aws:sns:us-east-1:111122223333:oncall-page

Now a latency blip alone, or a handful of errors alone, stays silent, and the page means users are both waiting and failing. NOT lets you mute a page during known work, for example ... AND NOT ALARM(maintenance-window-flag). A separate actions suppressor setting can also silence a composite's actions while a named alarm is firing. ⚠️ A composite hides things: if the rule is too strict, a real single-signal outage never pages. Keep a plain static alarm for the failures that must always page, such as a very high 5xx rate.

Anomaly detection: a learned band

An anomaly detection alarm trains a model on a metric's history, draws an expected band around it, and alarms when the metric leaves the band. You set how wide the band is. It suits seasonal metrics that rise and fall on daily or weekly cycles, like request count: 40 requests per minute is normal at 03:00 and alarming at 14:00.

It works poorly in three cases:

  • No history. A new metric has nothing to learn from.
  • Hard limits. Disk at 90% is bad no matter what the band says.
  • Flat lines. The band hugs the line, so tiny wiggles fire constantly.

Bad periods inside the training data teach the model that outages are normal. You can exclude time ranges from training.

What it costs

Exact prices change, so read the current pricing page. The shape is stable:

Alarm type Billing shape
Standard per alarm metric, per month
High-resolution (under 60 s) higher per-alarm rate
Metric-math alarm can count each metric it references
Composite separate, flat per composite
Anomaly detection multiple of a standard alarm

In units of one standard alarm, 40 anomaly alarms cost about three times as much as 40 standard ones, and only some metrics benefit. A sound plan caps the paging alarms per service, keeps most alarms static, spends anomaly detection on a few seasonal traffic metrics, and uses a composite where it removes pages. Silent child alarms still cost money, so count them.

Guided attempt: the checkout API

Task. The checkout path is ALB (load balancer) β†’ Lambda β†’ DynamoDB. Choose 4 alarms that can notify a human (silent children of a composite are allowed). For each, say static, composite or anomaly, and what it pages for. Available metrics: ALB HTTPCode_Target_5XX_Count and RequestCount, ALB TargetResponseTime p99, Lambda ConcurrentExecutions and Errors, DynamoDB ThrottledRequests. Think about which signals users feel and which have hard limits or seasonality.

Check your answer
  1. Static, pages: ALB 5xx rate above about 2%, 3 of 5 one-minute datapoints, notBreaching. Customers are seeing errors.
  2. Composite, pages: p99 latency high AND DynamoDB throttles high (both children silent). Checkout is slow because of database capacity, so the runbook says to raise capacity. Neither signal alone justifies waking someone.
  3. Anomaly, notifies a team channel (no page): RequestCount below its band. Traffic is seasonal, so a static floor would be wrong at some hour. A silent drop can mean a broken upstream or DNS problem.
  4. Static, ticket: Lambda ConcurrentExecutions near the reserved concurrency limit. It is a hard limit, so a learned band adds nothing.

CPU is absent because nothing here pages on a cause. (Other reasonable designs exist; check yours against user impact and hard-limit versus seasonal reasoning.)

Independent variation: 30 pages a week

Task. A web fleet in an Auto Scaling group has an alarm: average CPU > 70%, period 60 s, 1 of 1, paging on-call. It pages about 30 times a week, mostly around 02:00 when a batch job runs. During those pages p99 latency stays under 400 ms and the 5xx rate is about 0.1%. Redesign the alarms. Say what happens to the CPU alarm, which new alarms you add, and whether any is a composite.

Check your answer
  • Remove paging from CPU. Move CPU to a target-tracking scaling policy, which is the machine-reacts-to-machine use. If you keep a CPU alarm for visibility, give it no page action and loosen it, for example above 85% for 10 of 15 minutes.
  • Add symptom alarms with silent children: p99 latency above 1 s (3 of 5) and 5xx rate above 2% (3 of 5).
  • Add a composite: ALARM(cpu-high) AND ALARM(p99-latency) pages for saturation that hurts users. Keep the plain 5xx alarm paging on its own so that a hard failure with low CPU still reaches someone.

The batch-time CPU spikes no longer page because latency never breaches. Real saturation still does, so the pages that remain mean customers are affected.

Design Task: Cost-Aware Observability for One Service

You are handed a CloudWatch bill and an on-call rotation that has stopped trusting its pager. Fixing both means deciding, signal by signal, what to log, how long to keep it, what to turn into a number, and what deserves to wake someone. This task uses only ideas from the earlier sections: log groups and retention, structured JSON with a request ID, metric filters and EMF, and alarms.

The scenario

The order service has three parts. An API Gateway REST API (stage prod) calls a Lambda function orders-api, which validates the order and puts a message on an SQS queue (a managed message queue). An ECS Fargate service, orders-worker, consumes that queue. Current state:

  • orders-api runs at DEBUG level and logs every request body. The worker logs payloads and a line for every health check.
  • No log group has a retention setting.
  • A latency metric is published with a customerId dimension (300 customers).
  • There are 25 alarms. The loudest is orders-worker-cpu-high, which fires several times a week with no user impact. On-call has muted it, and with it every other alarm.

Step 1: Build the cost worksheet

The rates below are illustrative round numbers for the worksheet. Check the CloudWatch pricing page for your Region before quoting a real bill.

Driver Volume per month Illustrative rate
Ingestion: orders-api 120 GB $0.50 / GB
Ingestion: orders-worker 60 GB $0.50 / GB
Ingestion: API access logs 20 GB $0.50 / GB
Storage (never expires) 1,200 GB held $0.03 / GB-month
Insights data scanned 2,000 GB $0.005 / GB
Custom metrics 1 name x 300 customers $0.30 / metric
Standard alarms 25 $0.10 / alarm

Task: compute each driver's monthly cost and rank them. Predict which lever saves the most before you open the answer.

Check your answer
  1. Ingestion: 200 GB x $0.50 = $100
  2. Custom metrics: 300 x $0.30 = $90. Each customerId value is billed as its own metric.
  3. Storage: 1,200 x $0.03 = $36
  4. Insights: 2,000 x $0.005 = $10
  5. Alarms: 25 x $0.10 = $2.50

The total is about $238.50. The surprise is that the dimension mistake nearly equals all log ingestion, and the alarm bill is trivial. Alarms cost little. Their real cost is lost trust, which the money figures hide.

Step 2: Apply the log cost levers

Ingestion is the largest log line item, so cut it at the source:

  • Log level per environment. Set LOG_LEVEL=INFO in production and DEBUG only in dev, read from an environment variable.
  • Remove payload dumps. Log the order ID, tenant and size instead of the body. Drop health-check lines in the worker.
  • Retention per group. Fourteen days for the Lambda and worker groups. Thirty for access logs, assuming that window suits your audit needs.
  • Infrequent Access (IA) log class. This class costs less to ingest. It suits logs you rarely query, such as access logs. It does not support metric filters, subscription filters (rules that stream events elsewhere), embedded metric format or Live Tail. It also does not support the GetLogEvents and FilterLogEvents operations, so aws logs tail (which calls FilterLogEvents) cannot read it: you read an IA group through Logs Insights. The class is fixed when the group is created. Never use it for a group you derive metrics from or need to tail.

Step 3: Decide what each signal becomes

For every signal, ask three questions in this order:

Question Log only Derive a metric Alarm
How often is it read? Rarely, ad hoc Dashboards, trends Continuously
How many distinct values? Unbounded is fine Small and fixed Small and fixed
Must a human wake up? No No Yes, and it is user-visible

A request ID or tenant ID is unbounded, so it stays a log field. "Orders per minute" is bounded, so it becomes a metric. "Customers cannot check out" is a symptom, so it becomes an alarm. Worker CPU fails the third question, so it is a dashboard line at most. This is a starting heuristic. A rare signal can still deserve an alarm when missing it is costly.

Step 4: The harder variation

New constraint: three premium tenants (t-101, t-102, t-103) need per-tenant latency visibility. The custom metric budget cannot grow. Your design currently publishes OrderLatency by Route (4 routes), plus 1 worker error metric, for 5 custom metrics.

Task: choose between EMF with a tenant dimension for the premium tenants only, and Insights queries for everyone else. Then explain how the metric count stays at 5.

Check your answer

Use both. EMF handles the three premium tenants with a bounded, known list. Insights handles the other tenants, because their IDs are unbounded and queried rarely.

To keep the count flat, drop Route from the dimensions (4 metrics become 1) and spend the savings on the 3 tenant metrics. 1 + 3 + 1 worker metric = 5. Route and TenantId stay as plain properties on the log line, so Insights can still query them.

import json
import time

PREMIUM = {"t-101", "t-102", "t-103"}

def emit_latency(tenant_id, route, latency_ms, request_id):
    # Service-only metric for everyone; tenant dimension for premium only
    dims = [["Service"]]
    if tenant_id in PREMIUM:
        dims.append(["Service", "TenantId"])
    print(json.dumps({
        "_aws": {
            "Timestamp": int(time.time() * 1000),
            "CloudWatchMetrics": [{
                "Namespace": "Orders",
                "Dimensions": dims,
                "Metrics": [{"Name": "OrderLatency", "Unit": "Milliseconds"}]
            }]
        },
        "Service": "orders-api",
        "TenantId": tenant_id,   # a plain property for non-premium tenants
        "Route": route,
        "OrderLatency": latency_ms,
        "requestId": request_id
    }))

For any other tenant, a scoped query answers the question without a metric. Select only /aws/lambda/orders-api, set the time range to the last 24 hours, and run:

filter TenantId = "t-207"
| stats pct(OrderLatency, 95) as p95, count(*) by bin(1h)

Step 5: Deliverable and model design

Write a one-page design containing:

  1. Log groups with retention.
  2. The structured log schema, including the request ID field.
  3. Three metrics, each with how it is produced.
  4. Three alarms, one of them composite.
  5. A monthly cost estimate using the Step 1 rates.

Do this before opening the model below. Your alarms and metrics may legitimately differ.

Check your answer

Log groups:

  • /aws/lambda/orders-api: 14 days
  • /ecs/orders-worker: 14 days
  • /aws/apigateway/orders-prod-access: 30 days, IA class. This is a trade-off. aws logs tail does not work on an IA group, and the Logs Insights console queries one log class at a time (Standard by default), so you search the access log in its own query with Log class set to Infrequent Access, not alongside the two Standard groups. AWS's pages disagree on whether JSON fields are auto-discovered in IA groups, so if filter requestId = ... finds nothing there, extract the field with parse first. If you want the single three-group trace from the second section, keep this group in Standard and count its ingestion at the full rate ($10 instead of $5).

Schema: one JSON object per line, shared by all three services.

{"level":"INFO","service":"orders-worker","requestId":"c6af9ac6-7b61","message":"order stored","latencyMs":84,"tenantId":"t-207"}

requestId comes from API Gateway's $context.requestId in the access log format, from event.requestContext.requestId in Lambda, and from an SQS message attribute in the worker.

Metrics:

  • API 5XXError: built-in, free.
  • OrderLatency by Route: EMF from Lambda, 4 metrics.
  • WorkerErrors: a metric filter { $.level = "ERROR" } on the worker group, 1 metric. The filter counts only future events.

Alarms:

  • Api5xxHigh: 5xx sum at or above a chosen threshold, 3 of 5 one-minute datapoints, missing data treated as notBreaching. No action.
  • OrderLatencyHigh: p95 above your latency target for POST /orders, 3 of 5 datapoints. No action.
  • OrdersUserImpact: composite, ALARM(Api5xxHigh) OR ALARM(OrderLatencyHigh). It pages once, and child notifications are suppressed. An AND rule would miss an outage that fails fast with low latency.

Worker CPU becomes a dashboard line, and the other 22 alarms are removed or demoted.

Estimate: ingestion is 30 GB Lambda ($15) + 25 GB worker ($12.50) + 20 GB access logs in IA (assume half rate, $5) = $32.50. Add storage of about $1, Insights of about $0.50 (roughly 100 GB scanned), metrics 5 x $0.30 = $1.50, and alarms 2 x $0.10 + $0.50 = $0.70. The total is about $36 against $238.50.

Self-check rubric

  • The request ID appears in every service's logs, including across the queue.
  • No unbounded dimensions.
  • Retention is set on every group.
  • Every alarm maps to a user-visible symptom.
  • Insights queries are scoped by time range and log group.

Pitfalls to review inside this design

  • API Gateway logging is off until you enable it. Check the stage's access log setting, and the account-level CloudWatch role, which a REST API needs for either log type. Without them, edge failures have no request IDs to trace.
  • Metric filters count only events ingested after creation. Create WorkerErrors before the incident, and do not expect it to backfill history.
  • Missing data. A worker error count goes quiet when nothing fails. Choose notBreaching deliberately, and add a separate alarm if silence from a service that should always emit is itself a failure.
  • Insights across every log group. The scan bill is the sum of all selected groups. Select only the order service's groups and a narrow time range.

Final check

Can you rank a bill's drivers by cost? Can you justify each signal as log, metric or alarm? Can you keep tenant visibility without per-tenant metrics for everyone? Can you explain why each of your alarms deserves a page?