Skip to main content

Keeping SQS and Lambda Timeouts in Sync with Terraform

A Lambda consumer can succeed on a message and the same message can still run twice. Often nothing malfunctioned: two numbers that were never set together are interacting, the queue's visibility timeout and the function's timeout. The reasoning behind how those numbers should relate (including the recommendation of six times the function timeout plus the batching window) is covered in the earlier post, SQS and Lambda: Choosing Visibility Timeout, Retries and Dead-Letter Queue Settings. This one stays on configuration: how to make sure the numbers cannot quietly drift apart.

Three settings, edited on different days

The function timeout lives on the function. The visibility timeout lives on the queue. The batching window lives on the event source mapping. In a console-driven setup they get changed by different people on different days, each change reasonable in isolation. Lambda does check one relationship: the function timeout must be less than or equal to the queue's visibility timeout, and the check runs when you create or update the event source mapping. That catches the most obvious mistake and nothing more.

Infrastructure as code helps because it puts all three values in one reviewable place, but only if they are written as one decision rather than three independent numbers.

Derive the lease instead of typing it

The simplest defense is to compute the visibility timeout from the function's timeout in the same configuration. The numbers below are illustrative, a 30-second function timeout and a 5-second batching window, not values from any particular system.

variable "batching_window" {
  type    = number
  default = 5
}

resource "aws_lambda_function" "worker" {
  function_name = "jobs-worker"
  timeout       = 30
  # runtime, handler, role and package omitted
}

resource "aws_sqs_queue" "dlq" {
  name                      = "jobs-dlq"
  message_retention_seconds = 1209600 # 14 days, longer than the source queue
}

resource "aws_sqs_queue" "jobs" {
  name                       = "jobs"
  visibility_timeout_seconds = aws_lambda_function.worker.timeout * 6 + var.batching_window
  message_retention_seconds  = 345600 # 4 days

  redrive_policy = jsonencode({
    deadLetterTargetArn = aws_sqs_queue.dlq.arn
    maxReceiveCount     = 5
  })
}

resource "aws_lambda_event_source_mapping" "jobs" {
  event_source_arn                   = aws_sqs_queue.jobs.arn
  function_name                      = aws_lambda_function.worker.arn
  batch_size                         = 10
  maximum_batching_window_in_seconds = var.batching_window
  function_response_types            = ["ReportBatchItemFailures"]
}

Raising the function timeout now moves the lease with it: 6 × 30 + 5 = 185 seconds today, 6 × 60 + 5 = 365 tomorrow, with no second edit to forget.

One caveat: if the function needs the queue's URL in an environment variable, referencing the function from the queue creates a dependency cycle. In that case, move the timeout into a shared local value or variable and reference it from both resources instead.

Guard the relationship at plan time

Deriving the value protects you while the formula stays in place. Someone can still replace the expression with a hard-coded number during an unrelated change. A precondition turns the relationship into a rule that fails the plan instead of relying on reviewers to notice. Terraform supports preconditions in a resource's lifecycle block from version 1.2.

resource "aws_lambda_event_source_mapping" "jobs" {
  # arguments as above

  lifecycle {
    precondition {
      condition = (
        aws_sqs_queue.jobs.visibility_timeout_seconds >=
        aws_lambda_function.worker.timeout * 6 + var.batching_window
      )
      error_message = "Visibility timeout must be at least 6x the function timeout plus the batching window."
    }
  }
}

The six-times factor is a recommendation, not a hard platform limit, so the guard is a policy you are choosing to enforce. If you have a good reason to go tighter for a very short, predictable function, change the rule deliberately and record why, rather than letting it erode.

Drift: what the documentation does and does not say

The documentation describes the timeout check at the moment an event source mapping is created or updated. I did not find a statement saying it is re-evaluated when someone later edits only the queue or only the function through another path. So whether a manual change gets caught, and when, is something to test rather than assume.

Terraform gives you two practical tools regardless:

  • Plan compares reality with state. If someone edits the queue's visibility timeout in the console, the next plan shows the difference from the configuration.
  • Scheduled plans find drift early. Running terraform plan -detailed-exitcode on a schedule returns exit code 2 when changes are pending, which a pipeline can turn into an alert instead of a surprise during an unrelated deploy.

Test the drift cases in a sandbox

A throwaway environment is enough. Apply the configuration, then change one value at a time outside Terraform and watch what happens:

Change made outside Terraform What to observe
Raise the function timeout above the queue's visibility timeoutWhether AWS accepts it, and what the next mapping update does
Lower the queue's visibility timeout below the function timeoutWhether the queue edit is accepted and whether the consumer starts reprocessing messages
Change the batching window onlyWhether the lease still covers the window, and whether your precondition fails on the next plan
Revert everything with terraform applyThat the plan shows the drift clearly and apply restores the intended values

Record the results next to the configuration. Behavior can change over time, so repeat the test when you upgrade the provider or change how the consumer is deployed.

What this does not solve

Synchronized numbers reduce premature redelivery. They do not make a handler idempotent, and they do not guarantee a message is processed exactly once. Standard queues deliver at least once, so each side effect still needs its own duplicate protection. Partial batch responses (ReportBatchItemFailures above) limit replay to the records that failed, and AWS notes that Powertools for AWS Lambda includes a batch processor, available for .NET, that implements the response format. Check the current documentation before hand-rolling it.

A sensible order of work

  1. Move the function timeout, batching window and queue settings into one Terraform configuration.
  2. Derive the visibility timeout from the function timeout instead of typing a number.
  3. Add a precondition that fails the plan when the relationship is broken.
  4. Make the DLQ retention longer than the source queue's, and decide who watches it.
  5. Run the drift cases in a sandbox and write down what AWS and Terraform actually do.
  6. Schedule a plan with -detailed-exitcode so out-of-band changes raise an alert.

The aim is not a perfect number but a configuration where the numbers cannot change independently without someone noticing. Whether the six-times rule suits your workload is a question for your own failure-path tests.

Comments