Skip to main content

SQS and Lambda: Choosing Visibility Timeout, Retries and Dead-Letter Queue Settings

Picture a Lambda function with a 60-second timeout reading from an SQS queue. How long should the queue hide a message once the function has picked it up? The intuitive answer is 60 seconds, maybe a little more. The documented answer is considerably larger, and understanding why matters more than copying the number. This guide walks through the reasoning behind each setting that connects a queue to its consumer, and the failure modes each one is meant to prevent.

Three questions drive the decisions. Can a message be processed twice without harm? How long should a genuinely failed message stay hidden before it is retried? And who looks at the dead-letter queue (DLQ), and how long do they have to act? The defaults answer none of these for you.

A visibility timeout is a lease, not a limit on your code

When a consumer receives a message, SQS hides it for the length of the visibility timeout. If the message is not deleted within that window, it becomes visible again. SQS cannot distinguish a slow consumer from a dead one, so it assumes the worst. If your function is merely slow, a second invocation can pick up the same message while the first is still running, which is how duplicate writes and duplicate side effects happen.

Lambda protects you from the most obvious mistake. When you create or update an event source mapping, it checks that the function timeout is less than or equal to the queue's visibility timeout and rejects the configuration otherwise. That check only covers the case where the lease is shorter than the function itself. It says nothing about whether the lease is long enough for everything else that happens between receiving a message and deleting it.

Why "equal to the function timeout" is still too short

The Lambda documentation recommends a visibility timeout of at least six times the function timeout, plus the value of MaximumBatchingWindowInSeconds if you use a batch window. The extra headroom exists so Lambda can retry a batch when the function is throttled while processing an earlier one.

The reason is timing. The lease starts when the poller receives the message, not when your code begins to run. Time spent waiting for a batch to fill, or waiting out throttling, comes out of the same lease. A visibility timeout equal to the function timeout leaves no room for any of it.

The price of this rule is worth calculating. With a 60-second function timeout and a 5-second batch window, the recommendation is 6 × 60 + 5 = 365 seconds. If an invocation crashes outright, that message stays hidden for over six minutes before anything retries it. With a five-minute function timeout the same rule gives thirty minutes. That trade-off is the real design decision, and it is a reason to shorten the function timeout rather than inflate the queue's. If the work cannot finish inside a tight timeout, it may need a different shape: split into smaller steps, or moved to a long-running worker.

Resources:
  WorkDlq:
    Type: AWS::SQS::Queue
    Properties:
      MessageRetentionPeriod: 1209600   # 14 days, longer than the source queue

  WorkQueue:
    Type: AWS::SQS::Queue
    Properties:
      MessageRetentionPeriod: 345600    # 4 days
      VisibilityTimeout: 365            # 6 x 60s function timeout + 5s batch window
      RedrivePolicy:
        deadLetterTargetArn: !GetAtt WorkDlq.Arn
        maxReceiveCount: 5

  WorkMapping:
    Type: AWS::Lambda::EventSourceMapping
    Properties:
      EventSourceArn: !GetAtt WorkQueue.Arn
      FunctionName: !Ref WorkFunction   # Timeout: 60
      BatchSize: 10
      MaximumBatchingWindowInSeconds: 5
      FunctionResponseTypes:
        - ReportBatchItemFailures

The visibility timeout above is a derived value, not a magic number. If someone later raises the function timeout, the queue setting has to move with it, otherwise the two drift apart and the mapping fails validation the next time it is recreated. Keeping the arithmetic in a comment, or computing it in your infrastructure code, is cheap insurance.

One bad message should not replay the whole batch

By default, if Lambda hits an error while processing a batch, every message in that batch becomes visible again, including those that were handled successfully. A batch of ten with one poison message replays nine good ones on every attempt, multiplying work and side effects.

The fix is to enable ReportBatchItemFailures on the event source mapping, as in the template above, and have the handler return only the IDs that failed:

{
  "batchItemFailures": [
    { "itemIdentifier": "the-failed-messageId" }
  ]
}

A few details are worth knowing:

  • The setting must be on the mapping. A handler that returns this shape without it gets no partial-batch behavior.
  • A thrown exception fails the whole batch. Reserve exceptions for problems with the entire invocation, and catch errors per message otherwise.
  • FIFO queues need extra care. To preserve ordering, stop processing after the first failure and report the remaining unprocessed messages as failed too.
  • Returning failures is gentler than throwing. Lambda reacts to function errors by adjusting how aggressively it polls, so a steady trickle of bad messages is better reported than thrown. Confirm the current scaling behavior in the documentation.

In practice this shapes the handler: wrap each message in its own try/catch, log the message ID alongside the error, and return the list of failures at the end.

The redrive policy is a retry budget

maxReceiveCount is how many times a message may be delivered before SQS moves it to the DLQ. Set it too low and a few transient failures send healthy messages straight to the DLQ. A value of 1 means a single failed receive is enough. The Lambda documentation suggests at least 5 so the function gets several chances.

Combined with the lease, this gives the worst-case time before a persistently failing message reaches the DLQ. With a 365-second lease and a count of 5, a message that fails by crashing spends roughly half an hour in hidden windows. Failures reported quickly through partial batch responses cycle much faster.

Throttling adds an uncertainty. Lambda backs off when throttled, but the documentation does not clearly state how throttled attempts affect the receive count, so treat the half-hour figure as an estimate and test the behavior in your own account.

The DLQ is a clock, not a bucket

For standard queues, message expiry is always calculated from the original enqueue timestamp, and that timestamp does not change when the message is moved to the DLQ. A message that spent one day in a source queue and then lands in a DLQ with four-day retention is deleted three days later, not four. This is why the DLQ retention period should be longer than the source queue's, and why the template above uses 14 days against 4. FIFO queues behave differently, since the timestamp resets on the move.

A DLQ that nobody watches is just delayed deletion. Alarm on any message arriving, keep enough logging to understand what failed, and once the consumer is fixed, use the redrive feature to send messages back to the source queue. That only works safely if the consumer can handle the same message twice.

Assume duplicates anyway

None of these settings gives exactly-once delivery. Standard queues guarantee at-least-once delivery, so duplicates can occur even with a perfectly tuned lease. Pair the configuration with a deduplication key, such as the message ID or a business identifier, and make writes conditional or upserts. Redrive is a deliberate bulk replay, so a consumer that is safe to replay is what turns a DLQ into a real recovery path.

The settings at a glance

Setting Guideline What it protects against
VisibilityTimeout6 × function timeout + batch windowMessages reappearing while still being processed
ReportBatchItemFailuresEnable and return only failed IDsReplaying successful messages in a failed batch
maxReceiveCount5 or higher, chosen deliberatelyHealthy messages reaching the DLQ after transient errors
DLQ retentionLonger than the source queue (14 days is the maximum)Failed messages expiring before anyone looks
Idempotent handlerDeduplication key plus conditional writesDuplicate side effects from retries and redrive

A sensible order of work

  1. Decide the function timeout first, keeping it as short as the work allows.
  2. Derive the visibility timeout from it (six times, plus any batch window) and record the derivation next to the value.
  3. Enable ReportBatchItemFailures and test that a handler returning one failed ID replays only that message.
  4. Set maxReceiveCount deliberately, starting at 5 or higher, and consider what a short outage does to the count.
  5. Give the DLQ longer retention than the source queue and alarm on any depth above zero.
  6. Replay a DLQ message into a non-production queue to prove the handler is idempotent before you need it for real.
  7. Check how throttling affects receive counts in your own account.

The right answer depends on the workload. If jobs are short, idempotent and well below their timeout, the six-times rule plus a watched DLQ is cheap and sufficient. If individual jobs routinely run for minutes, the lease makes failure recovery slow, and a queue-fed container worker that can extend the lease with ChangeMessageVisibility as it works may be a better fit. Either way, set these values on purpose, because the defaults will not do it for you.

Comments