API Gateway Basics
Basic concepts only, for a developer who has to read and debug an existing setup: what API Gateway does in front of Lambda; REST APIs versus HTTP APIs versus WebSocket APIs and when each is used; routes, integrations, stages and deployments; Lambda proxy integration and the event and response shape; authorization options in brief (IAM, Lambda authorizer, JWT). WebSocket APIs: the $connect, $disconnect and $default routes, the route selection expression, connection IDs, and sending a message back to a client through the @connections management API with the permission it needs. Timeouts, payload limits and throttling as they stand today; where access and execution logs go; common errors (403, 502 from a malformed Lambda response, 504 on timeout, 410 Gone for a closed WebSocket connection).
SPACED REPETITION Β· 15 practice questions
Make this lesson stick.
Try 3 questions now. No account needed. Sample answers aren't saved.
or sign in to practice all 15Reading an API Gateway Setup: API Types, Routes, Integrations, Stages
Someone hands you an AWS account and says, 'the orders endpoint returns 403 since yesterday.' You open the Lambda console and the function looks fine, and its log group is empty. The request never reached it. Something sits in front of your code, and until you can read that layer, you are debugging blind. That layer is Amazon API Gateway, and this section teaches you to trace one request through it.
We will cover routes, integrations, stages and deployments, then how to choose among REST, HTTP and WebSocket APIs, and finally the 'I changed it and nothing happened' trap.
The managed front door
API Gateway is a managed service that accepts HTTPS requests on your behalf. It terminates TLS (decrypts the connection), matches the request to a route, can check credentials, applies throttling (rejecting traffic above a rate limit), writes logs, and then calls your backend. With Lambda behind it, the backend call is an invocation: API Gateway packs the request into a JSON event and hands it to your function.
Client
β HTTPS request: GET /prod/orders/42
API Gateway: TLS, stage lookup, route match, auth, throttling
β invokes the integration with a JSON event
Lambda function (knows nothing about HTTP except what the event says)
β returns JSON
API Gateway turns it back into an HTTP response
Your function never sees a socket or a URL, only the event. (The exact event shape is the subject of the next section, 'Lambda Proxy Integration: Event and Response Shape'.)
Prediction task: which route fires?
A route is a method plus a path pattern that API Gateway matches, such as GET /orders/{id}. The {id} part is a path parameter: a named slot whose value is captured and passed to the function. Here is an HTTP API's route table. All three routes call Lambda functions:
| Route | Integration |
|---|---|
GET /orders/{id} |
GetOrderFn |
POST /orders |
CreateOrderFn |
$default |
FallbackFn |
$default is the catch-all: it receives any request that matches no other route. Before reading on, predict the route, the function, and the pathParameters for each request:
GET /orders/42POST /ordersGET /ordersDELETE /orders/42
Check your answer
- Route
GET /orders/{id}βGetOrderFn, withpathParameters = {"id": "42"}. - Route
POST /ordersβCreateOrderFn. No path parameters (the key is absent or null, depending on API type). GET /ordershas no matching route: theGETroute needs an id segment andPOST /ordershas the wrong method. It falls to$defaultβFallbackFn, with no path parameters.DELETE /orders/42matches the path pattern but not the method, so it also goes to$defaultβFallbackFn, with no path parameters. A route is method and path, so a matching path alone is not enough.
The rule is most specific match wins: a literal route beats a parameterized one, which beats a greedy one like /{proxy+} (which swallows any deeper path), which beats $default. Without a $default route, requests 3 and 4 would fail instead. An HTTP API answers 404 Not Found. A REST API, which has no $default route at all, answers an unmatched request with a confusing 403 Missing Authentication Token, which in that case means 'no route matched', not 'you forgot a token'.
Four words, one running example
Take the URL https://abc123.execute-api.us-east-1.amazonaws.com/prod/orders/42.
- Route:
GET /orders/{id}, the method and path pattern that is matched. - Integration: the backend the route calls, here the Lambda function
GetOrderFn. Routes and integrations are separate objects: a route points at an integration. - Stage: a named, addressable version of the API.
prodis the stage, and it appears as the first path segment of the invoke URL. API Gateway strips it before route matching, so the route still reads/orders/{id}. - Deployment: a snapshot of the API's configuration published to a stage. Clients only ever see what has been deployed.
Stage variables are key-value settings attached to a stage and referenced as ${stageVariables.name}. A common use is pointing the prod stage at a Lambda alias and dev at another, using the same API definition.
Choosing the API type
| Cue | REST API | HTTP API | WebSocket API |
|---|---|---|---|
| Traffic shape | Request/response | Request/response | Two-way, server pushes |
| Standout features | Usage plans, API keys, request validation, caching, mapping templates | Built-in JWT authorizers, lower cost and latency | Persistent connections, message routes |
| Typical pick | Feature-rich or partner APIs | Simple, cheap Lambda backends | Chat, live dashboards |
Three scenarios:
- A partner API needs a per-customer quota of 1,000 calls a day, with keys you can revoke. Usage plans and API keys exist on REST APIs, so choose REST.
- A mobile app already signs users in with an OpenID Connect provider, and the backend is five CRUD functions. An HTTP API's JWT authorizer validates the tokens without custom code, and it is cheaper, so choose HTTP.
- A dashboard must update the moment an order ships, without polling. The server has to push, which needs a persistent connection, so choose WebSocket.
(This table is a starting heuristic. Feature sets shift over time, so confirm a specific feature in the current documentation before committing.) WebSocket routing gets its own section later in this lesson.
Pitfall: 'I changed the route and nothing happened'
A REST API serves only what was deployed to a stage. Edit a route in the console or in a template-less workflow and the live stage keeps the old behavior until you create a new deployment to it. An HTTP API stage can have auto-deploy on, so changes publish themselves. Check that setting on the stage first. (Tools such as SAM usually create deployments for you when you run a deploy, which hides this until someone edits in the console.)
The second most common mistake is the URL: calling /orders/42 instead of /prod/orders/42 on a named stage fails before any route is matched. An HTTP API answers 404 Not Found. A REST API reads orders as a stage name that does not exist and answers 403 with a plain Forbidden (not 'Missing Authentication Token'). The exception is the stage named $default on an HTTP API, which has no stage segment in the URL.
Guided attempt: label this setup
Resources:
OrdersApi:
Type: AWS::Serverless::HttpApi
Properties:
StageName: prod
GetOrderFn:
Type: AWS::Serverless::Function
Properties:
Runtime: python3.12
Handler: app.handler
CodeUri: src/
Events:
GetOrder:
Type: HttpApi
Properties:
ApiId: !Ref OrdersApi
Method: GET
Path: /order/{id}
The console's HTTP API page shows: Routes: GET /order/{id} β integration GetOrderFn (Lambda). Stages: prod, auto-deploy on. A client calls GET https://abc123.execute-api.us-east-1.amazonaws.com/prod/orders/42 and receives {"message":"Not Found"}.
Task: label the route, integration, stage and API type, then say which piece explains the symptom.
Check your answer
- API type: HTTP API (
AWS::Serverless::HttpApi, and the console says so). - Route:
GET /order/{id}, built fromMethodandPath. - Integration: the Lambda function
GetOrderFn, created implicitly by theHttpApievent. - Stage:
prod, which is why the client's URL contains/prod/.
The symptom comes from the route. The template says /order/{id} (singular) but the client calls /orders/42 (plural), so nothing matches and, with no $default route, an HTTP API answers 404. The stage in the URL is correct, and auto-deploy is on, so it is not a deployment problem. Fix the path to /orders/{id} (or fix the client). Because the function was never invoked, its log group will be empty, which is itself a clue that the failure is in the front door.
Lambda Proxy Integration: Event and Response Shape
A client calls your API and gets {"message": "Internal server error"} with status 502. You open the function's logs and see a clean run: START, your print statements, END, REPORT. Nothing failed, yet the call did. The explanation is that API Gateway and your function communicate through a strict JSON contract, and API Gateway rejected what the function returned. This section covers that contract.
One JSON in, one JSON out
With Lambda proxy integration (integration type AWS_PROXY), API Gateway does no reshaping. It packs the whole HTTP request into a single JSON event (the dictionary passed as the first argument to your handler) and expects the function to return one JSON object describing the HTTP response.
Client HTTP request
β
API Gateway packs method, path, headers, query, body into one JSON event
β
handler(event, context)
β
returns {statusCode, headers, body}
β
API Gateway checks the shape β HTTP response to client (or 502 if the shape is wrong)
Older REST API setups may use non-proxy (custom) integration (type AWS). There, mapping templates (small scripts in the Velocity Template Language, VTL) transform the request into whatever the function expects and transform the function's output into an HTTP response. If the console shows templates under "Integration Request" or "Integration Response", the function may receive a custom shape rather than the one below. Read the template before assuming anything about event.
Reading the event
Here is a trimmed event from a REST API (called payload format 1.0) for POST /orders/42?expand=items:
{
"httpMethod": "POST",
"path": "/orders/42",
"resource": "/orders/{id}",
"headers": {"Content-Type": "application/json"},
"queryStringParameters": {"expand": "items"},
"pathParameters": {"id": "42"},
"body": "{\"sku\": \"A1\"}",
"isBase64Encoded": false,
"requestContext": {"stage": "prod", "requestId": "c6af9ac6-...", "authorizer": {"claims": {"sub": "u-123"}}}
}
Notice that body is a string, not an object: the JSON you parse is JSON inside JSON. When isBase64Encoded is true, that string is base64 text and must be decoded first. pathParameters holds the {id} values from the route, and requestContext carries metadata API Gateway added: the stage, a unique requestId (useful for correlating logs), and whatever the authorizer produced.
HTTP APIs default to payload format 2.0, which is similar but organized differently (HTTP APIs can also be set to 1.0, so check the integration):
| Field | REST (1.0) | HTTP API (2.0) |
|---|---|---|
| Method | httpMethod |
requestContext.http.method |
| Path | path (+ resource) |
rawPath (+ routeKey) |
| Query | queryStringParameters |
same, plus rawQueryString |
| Header names | as the client sent them | lowercased |
| Cookies | in the Cookie header |
cookies list |
| Authorizer data | authorizer.claims (Cognito) |
authorizer.jwt.claims |
In 2.0, routeKey looks like POST /orders/{id}, which tells you which route matched.
The response contract
import json
def respond(status, payload):
return {
"statusCode": status, # required integer
"headers": {"Content-Type": "application/json"},
"body": json.dumps(payload), # must be a STRING
}
def handler(event, context):
params = event.get("pathParameters") or {}
order_id = params.get("id")
if order_id is None:
return respond(400, {"error": "id is required"})
return respond(200, {"id": order_id, "status": "shipped"})
Trace: GET /orders/42 matches /orders/{id}, so pathParameters is {"id": "42"}. The handler returns statusCode 200 and a body string. The client sees HTTP 200, a Content-Type: application/json header, and the text {"id": "42", "status": "shipped"}. The statusCode is not part of the body: API Gateway uses it as the HTTP status line. For binary output (an image, say), return base64 text in body with "isBase64Encoded": true; on a REST API, the API must also be configured to treat that content type as binary (its binaryMediaTypes list).
Debugging a 502 "Malformed Lambda proxy response"
API Gateway validates what the function returns; if it cannot turn the result into an HTTP response, it answers 502. Three causes cover most cases:
| Cause | Function log group shows | Fix |
|---|---|---|
body is a dict |
Clean run, no error | json.dumps(...) |
No statusCode |
Clean run, no error | Add integer statusCode |
| Crash / unhandled exception | Traceback or error line, then REPORT | Fix the bug, catch and return 4xx/5xx |
Open the function's own log group, /aws/lambda/<function-name>, and find the stream for the failing time. A Traceback (Python's printed stack of an unhandled exception) means a crash; Lambda then returns an error payload instead of your response object. A clean run with no error means the function succeeded and the bug is the shape you returned. Add a print(json.dumps(result)) just before returning to see exactly what API Gateway received. If REST API execution logging is on, it will also say "Malformed Lambda proxy response" (see the logging discussion in 'Limits, Throttling, Logs, and an Error Triage Capstone'). If the function log group has no stream for that request, the function was never invoked, which points at routing or authorization, covered in 'Authorization Options and Diagnosing 401/403 Responses'.
β οΈ Two behaviors to separate: the "no statusCode" failure is for REST APIs. An HTTP API on payload 2.0 is lenient: if the function returns valid JSON without statusCode, it assumes 200 and uses the returned JSON as the body. The same handler can therefore work on one API type and fail on the other.
Pitfalls
- Null body and null parameters. On a REST API GET with no body,
event["body"]isNone(on HTTP API payload 2.0 the key may be missing altogether), sojson.loads(event["body"])raisesTypeErrorand the result is a 502. Likewise, with no path or query parameters,pathParametersandqueryStringParameterscan beNone, soevent["pathParameters"]["id"]crashes. Usejson.loads(event.get("body") or "{}")and(event.get("pathParameters") or {}). - Header casing. REST 1.0 passes names as the client sent them (
Content-Typeorcontent-type); HTTP API 2.0 lowercases them. Normalize once:headers = {k.lower(): v for k, v in (event.get("headers") or {}).items()}. - CORS. CORS (the browser rule requiring the server to say which origins may call it) is your function's job in proxy mode for REST APIs: include
Access-Control-Allow-Originin every response, errors too. A 502 from a crash carries no such header, so the browser reports a misleading CORS error while the real problem is the crash. Check the Network tab status first. (HTTP APIs can also configure CORS on the API itself.) - Errors as 200s. Catching an exception and returning
statusCode: 200with{"error": ...}hides the failure from clients, retries, and the API's 4XX/5XX metrics. Return 400 for bad input and 500 for your own faults.
Practice
Assume a REST API with Lambda proxy integration. For each handler, predict the client-visible status, then write the fix before opening the answer.
A
def handler(event, context):
return {"statusCode": 200, "body": {"ok": True}}
B
import json
def handler(event, context):
return {"headers": {"Content-Type": "application/json"},
"body": json.dumps({"ok": True})}
C (the client posts a body for which isBase64Encoded arrives as true, for example a content type configured as binary)
import json
def handler(event, context):
data = json.loads(event["body"])
return {"statusCode": 200, "body": json.dumps({"name": data["name"]})}
Check your answer
A: 502. body is a dict, not a string. Fix: "body": json.dumps({"ok": True}). The function log looks perfectly healthy, which is the clue that the problem is the shape.
B: 502 on a REST API (an HTTP API on payload 2.0 would return 200). statusCode is missing. Fix: add "statusCode": 200.
C: 502. The body is base64 text, so json.loads raises JSONDecodeError, an unhandled exception. The log group shows the traceback. Fix:
import base64
import json
def handler(event, context):
body = event.get("body") or ""
if event.get("isBase64Encoded"):
body = base64.b64decode(body).decode("utf-8")
try:
data = json.loads(body)
except json.JSONDecodeError:
return {"statusCode": 400, "body": json.dumps({"error": "invalid JSON"})}
return {"statusCode": 200, "body": json.dumps({"name": data.get("name")})}
Decoding first fixes the normal case; the try/except turns a client mistake into a 400 instead of a 502.
Authorization Options and Diagnosing 401/403 Responses
A client calls your API and gets a 403. You open the Lambda function's log group and find nothing, not even a START line. That silence is the first clue. Authorization runs before the integration, so a request rejected at the gate never invokes the backend function. An empty function log therefore means "look at the gate", not "the handler is broken".
Client request
β
TLS + route match (no route β 403/404 here)
β
WAF / resource policy (REST APIs: can block here)
β
Authorization (IAM, Lambda authorizer, or JWT)
β
Integration β your Lambda (only now does the function log anything)
Three mechanisms cover most existing setups. Your job when debugging is to recognize which one is attached to the route and what it checks.
IAM authorization: the caller is an AWS principal
With IAM authorization, the caller signs each request with SigV4 (Signature Version 4, AWS's request-signing scheme that uses the caller's credentials). API Gateway checks the signature and then evaluates the caller's IAM policies for the execute-api:Invoke permission. It fits service-to-service or internal callers that already have an IAM role, such as another Lambda, an EC2 service or a CI job. Browsers and mobile apps rarely use it.
The policy names the exact route in the Resource ARN, whose shape is arn:aws:execute-api:REGION:ACCOUNT:API_ID/STAGE/METHOD/PATH:
{
"Version": "2012-10-17",
"Statement": [{
"Effect": "Allow",
"Action": "execute-api:Invoke",
"Resource": "arn:aws:execute-api:us-east-1:123456789012:a1b2c3d4e5/prod/GET/orders/*"
}]
}
This lets the caller do GET /orders/{anything} on the prod stage and nothing else. A POST /orders from the same role is denied because the ARN does not match.
Lambda authorizer: the caller uses a custom token scheme
A Lambda authorizer (formerly "custom authorizer") is a function that API Gateway calls first. It receives the token (or the whole request) and returns an allow or deny decision. In REST APIs the decision is an IAM policy document containing a Resource ARN. In HTTP APIs you can instead return a simple response, such as {"isAuthorized": true}, with no policy to build.
Lambda authorizers can cache their result for a TTL (time to live, how long the decision is reused). HTTP API Lambda authorizers can cache too, keyed on their identity sources. On a REST API the default is 300 seconds and the maximum 3600, set by AuthorizerResultTtlInSeconds in CloudFormation or by ReauthorizeEvery in a SAM authorizer's Identity block. That produces a classic bug, which is the guided attempt below.
JWT authorizer: the caller is an end user with an identity provider
A JWT authorizer (JSON Web Token, a signed token issued by a login provider such as Cognito, Auth0 or Okta) is built into HTTP APIs. You configure the issuer (who minted the token) and the audience (who the token is for). API Gateway verifies the signature, expiry, issuer and audience with no custom code. REST APIs have no JWT authorizer: for Cognito tokens they use a Cognito user pool authorizer, and for any other provider's tokens a Lambda authorizer that validates the JWT.
Failure cues for HTTP APIs: the client gets a 401, usually with a www-authenticate header containing invalid_token. Check three things:
- Expired token: compare
expwith the current time. - Issuer mismatch: the
issclaim differs from the configured issuer, often by a trailing slash or the wrong user pool region. - Audience mismatch: the
audclaim (orclient_idfor Cognito access tokens) is not in the configured audience list.
You can decode the token's payload (not the signature) at any JWT viewer or with base64 to compare these claims directly.
Choosing: who is the caller?
| Caller | Pick | Why |
|---|---|---|
| AWS principal (role, user) | IAM | Reuses existing credentials |
| End user with an identity provider | JWT / Cognito | No code to maintain |
| Custom token, API keys in a database, legacy scheme | Lambda authorizer | Arbitrary logic |
This is a starting heuristic. Real APIs sometimes mix types per route.
The 403 triage checklist
Read the response body message first, then the logs. The wording separates causes that look identical as a bare status code.
| Message or symptom | Likely cause | First check |
|---|---|---|
| "Missing Authentication Token" (REST) | No such path or method on that stage (or not yet deployed), or an unsigned call to an IAM-protected method | Path, method, deployment |
| "User is not authorized to access this resource with an explicit deny" | Authorizer returned Deny | Authorizer logic |
| "User is not authorized to access this resource" (no "explicit deny") | Authorizer returned an Allow that does not cover this method ARN, often a cached policy | Policy Resource, authorizer cache |
| "User: ... is not authorized to perform: execute-api:Invoke" | IAM policy misses the ARN, or the API's resource policy does not allow the caller | Resource ARN, API resource policy |
401 with invalid_token |
JWT expired or wrong aud/iss | Token claims |
| Plain "Forbidden" (REST) | Stage name in the URL does not exist, missing or invalid API key, or a WAF block | URL stage segment, x-api-key and usage plan, WAF logs |
β οΈ "Missing Authentication Token" is misleading. On a REST API it usually means the path or method does not exist on that stage (for example calling /prod/order/42 when the resource is /orders/{id}), or that the change was never deployed. In those cases the token was never the problem. Two relatives: a stage name that does not exist gives a plain "Forbidden" instead, and a wrong path sent together with an Authorization header can give a message starting "Authorization header requires 'Credential' parameter". HTTP APIs return 404 "Not Found" for the same mistake. Resource policies and WAF attachment apply to REST APIs; HTTP APIs do not support them directly.
To confirm which layer rejected the call, compare logs. If the access log shows the status but the integration's function log is empty, the gate rejected it. An access-log format that includes $context.authorizer.error helps for Lambda authorizers, and the authorizer function's own log group shows whether it ran and what it returned.
Guided attempt: the 403 that only hits the second route
A REST API uses a TOKEN-type Lambda authorizer with a 300-second TTL. Here is the authorizer:
def handler(event, context):
effect = "Allow" if is_valid(event["authorizationToken"]) else "Deny"
return {
"principalId": "user-123",
"policyDocument": {
"Version": "2012-10-17",
"Statement": [{
"Action": "execute-api:Invoke",
"Effect": effect,
"Resource": event["methodArn"],
}],
},
}
Symptom: a user's GET /orders/42 succeeds. Ten seconds later, with the same token, GET /orders/43 returns 403 with the message "User is not authorized to access this resource", and the backend function's log shows no invocation. The authorizer log shows only one invocation.
Task: explain why, and give two fixes.
Check your answer
event["methodArn"] is the ARN of the specific request, here .../prod/GET/orders/42. The token is valid, so the policy says Allow, but only for that one ARN. With a TOKEN authorizer the cache key is the token alone, so for the next 300 seconds API Gateway reuses the cached policy for every request carrying that token. GET /orders/43 does not match the cached Resource, so it is denied, and the authorizer is not re-run (hence only one authorizer invocation).
Fix 1: widen the Resource to the stage and all routes the user may call:
parts = event["methodArn"].split("/") # ["arn:...:a1b2c3d4e5", "prod", "GET", "orders", "42"]
resource = "/".join(parts[:2]) + "/*/*" # arn:...:a1b2c3d4e5/prod/*/*
Fix 2: set AuthorizerResultTtlInSeconds to 0 so the decision is computed for every request, at the cost of an authorizer invocation each time.
If different routes need different permissions, a wildcard grants too much. Use TTL 0, or a REQUEST-type authorizer whose identity sources include the method and path so they become part of the cache key.
WebSocket APIs: Routes, Connection IDs, and Pushing Messages Back
A chat app has a problem that ordinary request/response APIs cannot solve: when Alice posts a message, Bob's browser must receive it without asking. A normal HTTP request ends when the response is sent, so the server has no way to speak first. A WebSocket API fixes this by keeping one long-lived, two-way connection open between the client and API Gateway.
The new idea: the connection is to API Gateway, not Lambda
Lambda functions are short-lived and cannot hold a socket open. So API Gateway holds the connection and invokes your function once per event: one invocation when the client connects, one for each message, one when it disconnects. These are separate invocations, possibly on different execution environments, and they share no memory.
Client <==== persistent WebSocket ====> API Gateway
| (one invocation per event)
v
Lambda -> reads/writes connection IDs in DynamoDB
|
v
post_to_connection(connectionId, data)
|
v
API Gateway pushes data down the open socket
API Gateway gives each connection an opaque string, the connectionId. It is your only handle to that client. If a later invocation should message someone, an earlier invocation must have saved their connectionId somewhere durable. DynamoDB is the usual choice: a table with connectionId as the partition key (the attribute DynamoDB uses to locate an item).
The three special routes
A WebSocket route is a key that decides which integration (here, which Lambda function) handles an event. Three keys have built-in meaning:
| Route | Runs when | Typical job |
|---|---|---|
$connect |
Client opens the connection | Authorize, store connectionId |
$disconnect |
Connection closes | Delete connectionId (best effort) |
$default |
No other route matches | Fallback or error reply |
On $connect, a successful 200 response accepts the connection and an error response such as 403 rejects it, so this is where WebSocket authorization happens (of the options in the authorization section, WebSocket APIs support IAM authorization and Lambda authorizers of the REQUEST type, not JWT or Cognito authorizers, and only on $connect). $disconnect is best effort: API Gateway tries to invoke it, but an abrupt network drop can mean it never runs. Stored IDs can therefore go stale, which matters shortly.
Route selection: which handler runs for a message?
Messages that are not connect or disconnect are routed by the API's route selection expression, a rule that extracts a value from each message to use as the route key. The common setting is $request.body.action, meaning: parse the body as JSON and read its action field.
| Message body | Expression yields | Route that runs |
|---|---|---|
{"action":"sendmessage","text":"hi"} |
sendmessage | sendmessage |
{"action":"ping"} |
ping | ping if defined, else $default |
{"text":"hi"} |
nothing | $default |
β οΈ When the wrong handler runs, check three things: (1) the expression configured on the API, since it lives on the API and not on a route; (2) whether the key is spelled and cased exactly like the route key, because matching is exact; (3) whether the client sends valid JSON containing that field. A malformed message silently lands on $default, which looks like "my route never fires."
Reading the event
Every WebSocket invocation carries its routing facts in requestContext:
{
"requestContext": {
"routeKey": "sendmessage",
"eventType": "MESSAGE",
"connectionId": "abc123=",
"domainName": "a1b2c3d4e5.execute-api.us-east-1.amazonaws.com",
"stage": "prod"
},
"body": "{\"action\":\"sendmessage\",\"text\":\"hi\"}",
"isBase64Encoded": false
}
routeKey says which route fired; eventType is CONNECT, MESSAGE or DISCONNECT; connectionId identifies the sender. The body is a string you must parse yourself, as with proxy events. Joining domainName and stage gives the management endpoint, the URL your code uses to talk back: https://a1b2c3d4e5.execute-api.us-east-1.amazonaws.com/prod. (Behind a custom domain, domainName differs and the stage may not belong in the path. Pass the endpoint in an environment variable instead of deriving it.)
A minimal handler that stores and removes IDs:
import os
import boto3
table = boto3.resource("dynamodb").Table(os.environ["TABLE_NAME"])
def handler(event, context):
ctx = event["requestContext"]
conn_id = ctx["connectionId"]
if ctx["routeKey"] == "$connect":
table.put_item(Item={"connectionId": conn_id})
elif ctx["routeKey"] == "$disconnect":
table.delete_item(Key={"connectionId": conn_id})
return {"statusCode": 200} # on $connect, an error response would reject the client
Pushing data back with @connections
The @connections management API is a separate HTTPS API on the same domain that lets your code send data to, inspect or close a connection. In boto3 it is the apigatewaymanagementapi client, and its endpoint must be the management endpoint above:
client = boto3.client(
"apigatewaymanagementapi",
endpoint_url=f"https://{ctx['domainName']}/{ctx['stage']}",
)
client.post_to_connection(ConnectionId=conn_id, Data=b'{"text":"hi"}')
Data is bytes. The call is signed with the function's role, so that role needs execute-api:ManageConnections on the @connections resource ARN. Symptoms of getting this wrong:
- Missing permission: a 403 / access-denied error from
post_to_connection. - Wrong stage or API ID in the endpoint: a 403 or a connection/endpoint error, often reading as "it can't find my API."
- Function inside a VPC with no route to the internet or an interface endpoint: the call hangs, then times out.
410 Gone: stale connection IDs
When the client has left but $disconnect never cleaned up, post_to_connection raises GoneException (HTTP 410). The same error comes back if you post before the connection is established, for example from inside the $connect handler. For a client that left, this is expected, not a bug. The handler should delete that ID and keep broadcasting. Letting the exception escape means the loop stops partway, later clients get nothing, and the dead ID stays in the table to fail again next time.
Practice
1. Complete the skeleton. A sendmessage route should send the message text to every stored connection. Fill in blanks 1-5.
def broadcast(event, context):
ctx = event["requestContext"]
client = boto3.client("apigatewaymanagementapi", endpoint_url=___1___)
text = json.loads(event["body"])["text"]
for item in table.scan()["Items"]:
cid = item["connectionId"]
try:
client.post_to_connection(ConnectionId=___2___, Data=___3___)
except client.exceptions.___4___:
table.___5___(Key={"connectionId": cid})
return {"statusCode": 200}
2. Write the minimal IAM statement letting this function do the sending, for API a1b2c3d4e5, stage prod, region us-east-1, account 111122223333.
3. The table holds connection IDs A, B and C in that order. B's client left without $disconnect running. Predict what each client receives, what the table holds afterward, and what happens if the except clause is removed.
Check your answer
1. Blank 1: f"https://{ctx['domainName']}/{ctx['stage']}". Blank 2: cid. Blank 3: text.encode() (bytes). Blank 4: GoneException. Blank 5: delete_item.
2.
{
"Effect": "Allow",
"Action": "execute-api:ManageConnections",
"Resource": "arn:aws:execute-api:us-east-1:111122223333:a1b2c3d4e5/prod/POST/@connections/*"
}
The resource names the stage, the POST method that post_to_connection uses, and @connections/* for any connection ID.
3. A receives the message. B raises GoneException; the handler deletes B and moves on. C receives the message. The table afterward holds A and C, and the function returns 200. Without the except, the exception ends the loop at B: C never gets the message, the invocation fails, and B stays in the table to fail again on every later broadcast.
One limit of this skeleton: scan returns at most one page (up to 1 MB) per call, so a large table needs pagination, and very large audiences need a different fan-out design.
Limits, Throttling, Logs, and an Error Triage Capstone
A POST that worked in testing returns a 504 in production, a second endpoint returns 502 for no obvious reason, and a third returns 403 right after someone added a route. Each failure has a different cause, and each leaves evidence in a different place. This section covers the limits that produce these failures and the logs and metrics that show which one you are looking at. Then you diagnose a mixed set on your own.
Limits belong to the integration and the API type
A limit is a hard ceiling that API Gateway or Lambda enforces regardless of your code. The ceiling that applies is the smallest one along the request's path.
| Limit | Applies to | Value in October 2026 (verify) |
|---|---|---|
| Integration timeout | REST, HTTP, WebSocket | 29 s on REST and WebSocket, 30 s on HTTP |
| API payload size | REST, HTTP | 10 MB |
| Sync Lambda payload | request and response | 6 MB each |
| WebSocket message size | WebSocket | 128 KB, in frames of at most 32 KB |
| WebSocket connection life | WebSocket | 2 hours; closed after 10 minutes idle |
The integration timeout is how long API Gateway waits for the backend before giving up. On Regional and private REST APIs the 29-second default can be raised through a quota increase (AWS may lower the account's throttle quota in exchange); on edge-optimized REST APIs, HTTP APIs and WebSocket APIs it cannot. Treat a longer timeout as something to check, not assume. (The table describes the default buffered mode. A REST API proxy integration can instead be set to response streaming, transfer mode STREAM, which can send more than 10 MB and keep streaming for up to 15 minutes.) These numbers have changed over the years. Read the real values in Service Quotas (the AWS console page listing each quota and whether it can be raised) and the current API Gateway documentation, and don't memorize them. The practical consequence is that a 10 MB upload can pass API Gateway and still fail at a Lambda behind it, because the smaller 6 MB synchronous cap applies there.
504 worked example: the function outlives the API
Predict first. A REST API route calls a function with Timeout: 60. The work takes 40 seconds. What does the client see, and does the work finish?
Check your answer
After about 29 seconds the client gets a 504 Gateway Timeout. API Gateway cannot cancel the function, so it keeps running, finishes at 40 seconds with no error, and its result is discarded. Side effects still happen, so a client that retries the 504 may do the work twice. The function log shows a clean REPORT line with a duration of about 40 s, while the API metric IntegrationLatency sits near the timeout.
Two fixes work together:
- Align timeouts. For synchronous routes, set the Lambda timeout a few seconds below the API timeout. The function then fails loudly in its own log instead of succeeding after the client has gone.
- Move long work off the request path. Return 202 Accepted ("received, not finished") and process in the background, here through an SQS queue:
import json
import os
import boto3
sqs = boto3.client("sqs")
def handler(event, context):
req_id = event["requestContext"]["requestId"]
job = {"requestId": req_id, "body": event.get("body")}
# Hand the slow work to a worker; this call returns in milliseconds
sqs.send_message(QueueUrl=os.environ["QUEUE_URL"], MessageBody=json.dumps(job))
# Log the API request ID so this line can be matched to the access log
print(json.dumps({"apiRequestId": req_id, "msg": "queued"}))
return {"statusCode": 202, "body": json.dumps({"status": "accepted", "jobId": req_id})}
A second function reads the queue and does the 40 seconds of work. The client polls a status route using jobId, or is notified some other way.
Throttling: two different buckets
API Gateway throttles with a token bucket. The bucket holds up to burst tokens and refills at rate tokens per second. Every request takes a token, and an empty bucket returns 429 Too Many Requests. With rate 100/s and burst 200, a sudden spike of 300 requests lets roughly 200 through immediately and rejects the rest. Limits exist at the account level, the stage level and per route, and REST APIs add usage plans. A usage plan ties an API key (an identifier for one client) to its own rate, burst and quota, so one noisy client can't use up everyone's tokens.
Lambda has a separate limit, concurrency (how many copies of a function run at once). When it is exhausted, Lambda throttles the invocation. API Gateway reports this as a 5XX, not as its own token-bucket 429, and it is typically a 500, though the exact code can vary by API type. Tell them apart with metrics:
- API Gateway throttle: the API's 4XX metric and
Countrise, but the function'sInvocationsdo not follow. Requests were rejected before reaching Lambda. (REST metrics are4XXError/5XXError; HTTP APIs name them4xx/5xx.) - Lambda throttle: the function's
Throttlesmetric is above zero and the API's 5XX rises. - Slow rather than throttled: compare
Latency(total) withIntegrationLatency(time spent in the backend). A large gap points at API Gateway itself. A small gap means the time went into the function.
Where the evidence lives
- Access logs write one line per request in a format you choose with
$contextvariables, for example$context.requestId $context.status $context.httpMethod $context.resourcePath $context.responseLatency $context.integrationLatency. HTTP APIs have no$context.resourcePath; log$context.routeKey(method plus route pattern) there. For the backend's error text, HTTP APIs offer$context.integrationErrorMessageand REST APIs$context.integration.error. Variable names differ slightly by API type, so check the list for yours. You choose the destination log group. - REST API execution logs are detailed, per-step traces. They are opt-in: set a log level (ERROR or INFO) on the stage, and API Gateway writes them to a log group it manages,
API-Gateway-Execution-Logs_{rest-api-id}/{stage-name}. Both kinds of REST API logging need an account-level CloudWatch role registered once per Region in API Gateway settings. HTTP APIs have access logs only, no execution logs. Never enable full request and response body logging in production, because it writes tokens and personal data into logs. - The function's own log group,
/aws/lambda/<function-name>, holds yourprintoutput, stack traces and theREPORTline with duration.
The IDs differ across these places: the function's awsRequestId is not the API's requestId. Log event["requestContext"]["requestId"] yourself (as in the handler above), then search for it in both groups, for example with CloudWatch Logs Insights using filter @message like "<requestId>". An access-log line with no matching function line also tells you something: the request never reached the function.
Error triage table
| Symptom | First check | Evidence |
|---|---|---|
| 403 | Authorizer, policy, path | Empty function log |
| 502 | Response shape or crash | Function log |
| 504 | Timeout vs IntegrationLatency | Latency near cap |
| 429 | Throttle and usage plan | Invocations below Count |
| 410 | Stale WebSocket connectionId | GoneException |
For 403, also read the response message body. "Missing Authentication Token" often means a wrong path or method, not a missing token, and a plain "Forbidden" on a REST API can mean the stage name in the URL does not exist.
Capstone: diagnose three failing requests
A REST API with a Lambda token authorizer (cache TTL 300 s). Request IDs are shortened.
Api:
Type: AWS::Serverless::Api
Properties:
StageName: prod
Auth:
DefaultAuthorizer: TokenAuth
Authorizers:
TokenAuth:
FunctionArn: !GetAtt AuthFn.Arn
Identity:
ReauthorizeEvery: 300
ReportFn:
Type: AWS::Serverless::Function
Properties:
Timeout: 60
Events:
Create: {Type: Api, Properties: {RestApiId: !Ref Api, Path: /reports, Method: post}}
OrderFn:
Type: AWS::Serverless::Function
Properties:
Timeout: 10
Events:
Get: {Type: Api, Properties: {RestApiId: !Ref Api, Path: '/orders/{id}', Method: get}}
Recent: {Type: Api, Properties: {RestApiId: !Ref Api, Path: /orders/recent, Method: get}}
Access log (requestId status method resource responseLatency integrationLatency, in ms):
r-aaa1 504 POST /reports 29004 29001
r-bbb2 502 GET /orders/{id} 41 12
r-ccc3 403 GET /orders/recent 9 -
Function and authorizer logs:
ReportFn: {"apiRequestId":"r-aaa1","msg":"report start"}
ReportFn: REPORT RequestId: 5e21... Duration: 41873.12 ms Billed Duration: 41874 ms
OrderFn: REPORT RequestId: 9b07... Duration: 11.80 ms (no ERROR lines)
OrderFn code: return {"statusCode": 200, "body": {"id": order_id}}
AuthFn (earlier, same token): allow arn:aws:execute-api:us-east-1:111122223333:abc123/prod/GET/orders/42
AuthFn: no invocation recorded for r-ccc3
OrderFn: no log lines for r-ccc3
Task: for each request, name the evidence that identifies the cause and propose a fix. Try it before opening the answers.
Check your answer: r-aaa1 (504)
responseLatency and integrationLatency both sit at about 29 s, the default integration timeout (Regional and private REST APIs can have it raised by quota request; HTTP APIs and edge-optimized REST APIs cannot). The function's REPORT shows 41.9 s with no error, because its 60 s timeout is longer than the API's. The work finishes after the client has already received the 504. Fix: use the 202-plus-queue pattern. If a synchronous path must remain, set the Lambda timeout a few seconds below the API timeout.
Check your answer: r-bbb2 (502)
integrationLatency is 12 ms and the function log has no errors, so this is neither a timeout nor a crash. The handler returns body as a dict, not a string, which is a malformed proxy response. Fix: "body": json.dumps({"id": order_id}).
Check your answer: r-ccc3 (403)
The integration latency is - and OrderFn has no log lines, so the request was rejected before reaching the function. The authorizer was not invoked either, meaning a cached decision was reused. That cached policy allowed only GET /orders/42, so the new /orders/recent route is implicitly denied for the same token. Fix: have the authorizer return a wildcard Resource such as .../prod/GET/orders/* (or a scope that fits your rules), or disable authorizer result caching in your authorizer configuration.
Self-check
- I can trace a request through route, integration and stage in an existing setup.
- I can check a proxy response for a string
bodyand astatusCode. - I can tell whether an authorizer or a policy rejected a call, using the empty function log as a clue.
- I can explain how
$connect,$default, the connectionId and@connectionsfit together, and what a 410 means. - I can choose which log to open first: the access log for status and latency, the function log for crashes and shape, the execution log for step-by-step detail.