I discovered the hard way that EventBridge Scheduler behaves differently from the event-driven services most teams are used to. When a scheduled Lambda invocation fails, the retry logic is not the same as SQS, Step Functions, or even direct Lambda invocations through the console. The target configuration, the dead-letter queue setup, and the context attributes passed to your function all require deliberate decisions that are easy to overlook until your automation silently stops working.
This post walks through the three pitfalls I encountered while building scheduled automation on AWS Lambda, the configuration gaps that caused failures, and the design patterns that made scheduled invocations reliable. If you are migrating from CloudWatch Events or adding scheduled tasks to a serverless application, these details will save you from silent failures and duplicate work.
Context Attributes Are Not Automatic
When EventBridge Scheduler invokes a Lambda function, it does not automatically pass the same context you get from an event-driven trigger like SQS or API Gateway. The input payload is whatever you define in the schedule target configuration, and if you need the schedule name, the scheduled time, or the attempt number, you must explicitly configure those fields.
I assumed the Lambda event would include metadata about the schedule, similar to how CloudWatch Events passes rule name and timestamp. It does not. If your function needs to know which schedule triggered it, or if you want to log the scheduled versus actual execution time for observability, you must add those values to the input template. This matters when you run multiple schedules that invoke the same function with different parameters, or when you need idempotency keys derived from the schedule identity.
The input configuration is a JSON template defined in the schedule target. You can reference schedule attributes using the <aws.scheduler.*> placeholders. Here is an example that includes the schedule name, ARN, and scheduled time:
{
"scheduleArn": "<aws.scheduler.schedule-arn>",
"scheduleName": "<aws.scheduler.schedule-name>",
"scheduledTime": "<aws.scheduler.scheduled-time>",
"taskType": "daily-summary"
}
The Lambda function receives this as the event payload. The scheduledTime field is useful for comparing intended versus actual execution time, which helps surface scheduler lag or throttling issues. The scheduleArn becomes part of the idempotency key if your function writes to a database or external API and needs to guarantee at-most-once semantics across retries.
Without these attributes, debugging a failed invocation requires correlating CloudWatch Logs timestamps with the schedule configuration, which is brittle and slow during an incident. Including them in the input makes your logs self-describing and your metrics easier to query.
Retry Policies and When They Stop
EventBridge Scheduler supports a retry policy on each schedule target, but the defaults are conservative and the behavior is not the same as SQS redrive. The maximum number of retries is configurable from zero to 185, and the maximum event age controls how long the scheduler will attempt retries before giving up. If your Lambda function throws an exception, the scheduler will retry according to this policy, but once the event age or retry count is exceeded, the event is dropped unless you configure a dead-letter queue.
The retry policy lives in the schedule target configuration, not in the Lambda function settings. This is a critical distinction. If you rely on the Lambda asynchronous invocation retry behavior, you will not get it. The scheduler invokes Lambda synchronously and applies its own retry logic based on the target configuration.
I encountered this when a Lambda function started failing due to a downstream API timeout. The function had been tested with asynchronous invocation from SQS, where retries and a dead-letter queue were configured at the event source mapping. When the same function was invoked by EventBridge Scheduler, the retries stopped after two attempts, the event was discarded, and I had no visibility into the failure until I noticed missing records in a downstream report.
The fix was to add an explicit retry policy and dead-letter queue to the schedule target. Here is the relevant CloudFormation configuration:
Target:
Arn: !GetAtt MyLambdaFunction.Arn
RoleArn: !GetAtt SchedulerExecutionRole.Arn
RetryPolicy:
MaximumRetryAttempts: 3
MaximumEventAgeInSeconds: 3600
DeadLetterConfig:
Arn: !GetAtt FailureQueue.Arn
The MaximumRetryAttempts value controls how many times the scheduler retries after the initial invocation. Setting it to three means four total attempts. The MaximumEventAgeInSeconds value is a backstop: if the event is still failing after one hour, the scheduler stops retrying and routes the event to the dead-letter queue if configured.
These settings should match your downstream dependencies. If your Lambda function calls an external API with a known recovery time, set the maximum event age to exceed that window. If the API has rate limits, set retry attempts conservatively to avoid compounding throttling errors.
Dead-Letter Queues Are Optional But Essential
EventBridge Scheduler treats dead-letter configuration as optional. If you do not configure a dead-letter queue, failed events after retry exhaustion are silently dropped. There is no automatic fallback to CloudWatch Logs or SNS. This is a silent failure mode that will cause data loss if your schedule is responsible for any critical workflow.
The dead-letter queue must be an SQS queue, not an SNS topic. The scheduler writes the original input payload and metadata about the failure into the queue, where you can inspect it, replay it, or route it to an alerting system. I use a standard SQS queue with a message retention period of 14 days, which gives enough time to investigate and fix the root cause without losing the event.
The IAM role used by the scheduler must have sqs:SendMessage permission on the dead-letter queue. This is a separate role from the Lambda execution role, and forgetting to grant this permission results in a silent failure to write to the dead-letter queue, which defeats the entire purpose of the configuration.
Here is the IAM policy I attach to the scheduler execution role:
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Action": "lambda:InvokeFunction",
"Resource": "arn:aws:lambda:us-east-1:123456789012:function:MyLambdaFunction"
},
{
"Effect": "Allow",
"Action": "sqs:SendMessage",
"Resource": "arn:aws:sqs:us-east-1:123456789012:FailureQueue"
}
]
}
The first statement allows the scheduler to invoke the target Lambda function. The second allows it to write to the dead-letter queue. Without the second statement, retries will fail silently, and events will be lost after the retry policy is exhausted.
I monitor the dead-letter queue depth with a CloudWatch alarm. Any message in the queue indicates a failure that needs investigation. The alarm triggers an SNS topic, which routes to my on-call notification system. This gives visibility into schedule failures without waiting for a downstream report to break.
Designing for Observability and Recovery
Once the schedule target configuration includes context attributes, a retry policy, and a dead-letter queue, the next question is how to verify that the schedule is working and how to recover from failures without manual intervention. I use three mechanisms: structured logging, CloudWatch metrics, and a replay function that reads from the dead-letter queue.
The Lambda function logs the schedule ARN, the scheduled time, and the attempt number at the start of each invocation. This makes it possible to query logs by schedule name and correlate failures across invocations. I also emit a custom CloudWatch metric for each successful and failed invocation, tagged with the schedule name and task type. This powers a dashboard that shows schedule health over time and surfaces patterns like intermittent downstream timeouts or rate-limit errors.
The replay function is a separate Lambda that reads messages from the dead-letter queue, validates the payload, and re-invokes the original function. This is useful for transient failures that have been resolved, such as a downstream API that was temporarily unavailable. The replay function includes a manual approval step using a Step Functions workflow, so I can inspect the failure reason before replaying the event. This prevents replaying events that failed due to bad input data or configuration errors that would fail again.
When EventBridge Scheduler Is the Right Choice
EventBridge Scheduler is a good fit when you need a large number of independent schedules with different cadences, or when you need to manage schedules dynamically through an API without deploying infrastructure changes. It is not a good fit if your scheduled tasks require complex orchestration, branching logic, or human approval steps. In those cases, Step Functions with a scheduled trigger is a better choice.
The pitfalls I described—missing context attributes, silent retry exhaustion, and the optional dead-letter queue—are not bugs. They are design choices that trade simplicity for explicit configuration. The scheduler assumes you know what context your function needs, how many retries are appropriate, and whether you care about failed events. If you do not configure these explicitly, the scheduler will quietly drop events, and you will not know until something breaks downstream.
Here is a checklist I use when adding a new schedule:
- Include schedule ARN, name, and scheduled time in the input template for observability and idempotency.
- Configure a retry policy with maximum attempts and event age that match downstream recovery expectations.
- Add a dead-letter queue and verify the scheduler execution role has permission to write to it.
- Emit custom CloudWatch metrics for successful and failed invocations, tagged by schedule name.
- Monitor dead-letter queue depth with a CloudWatch alarm that routes to on-call notifications.
- Build a replay mechanism for transient failures, with manual approval to prevent replaying bad events.
EventBridge Scheduler is a reliable service once you configure it correctly, but the defaults are not safe for production workloads. Treat the target configuration as an architecture decision, not a deployment detail, and test failure paths before you rely on the schedule for critical workflows.
Comments
Post a Comment