Skip to main content

Secrets Manager Rotation and AWS Lambda: Designing for Credentials That Change

Most Lambda functions read a secret once during initialization, keep it in a static field and never think about it again. That works until someone turns on automatic rotation. From that moment the value in memory can become outdated while the function is still running, and the symptom is rarely a Secrets Manager error. It shows up as a 401 or a database login failure somewhere downstream, which is easy to mistake for a configuration problem.

This guide explains where the problem comes from, which strategies handle it, and what the rotation side needs to get right. The short version: rotation is a design decision for both the code that reads the secret and the code that changes it, not a switch you flip afterwards.

Why cached secrets and rotation clash

Lambda reuses an execution environment for many invocations, and code outside the handler runs once per environment. That is the right place for SDK clients and the wrong place for a credential that is meant to change. An environment can live for a long time, so a value read at startup may be used for minutes or hours after Secrets Manager has replaced it.

What makes this hard to diagnose is that nothing tells the function its secret changed. The only evidence is a rejected request. If the code retries with the same cached value, the retry fails for the same reason, and the problem lasts until the environment is recycled by a deployment, a scale event or idle time.

Strategy 1: read the secret on every invocation

The simplest fix is to fetch the secret inside the handler instead of during initialization. Create the Secrets Manager client once, because clients are cheap to reuse, but request the value each time.

public class Function
{
    // The client is reused; the secret value is not
    private static readonly IAmazonSecretsManager Secrets = new AmazonSecretsManagerClient();

    public async Task<APIGatewayProxyResponse> FunctionHandler(
        APIGatewayProxyRequest request, ILambdaContext context)
    {
        var secret = await Secrets.GetSecretValueAsync(
            new GetSecretValueRequest { SecretId = "my-service/api-key" });

        var apiKey = secret.SecretString;
        // call the downstream API with apiKey
        return new APIGatewayProxyResponse { StatusCode = 200 };
    }
}

Every invocation now sees whatever value is current. The cost is one extra network call per invocation, which adds latency and a small per-request charge. For functions with modest traffic or relaxed latency targets, that is usually an acceptable price for never having to reason about stale values.

Strategy 2: cache with a TTL and refresh on failure

When traffic is high enough that an extra call per invocation matters, cache the value but give it a lifetime, and add a recovery path for the case where it goes bad earlier than expected. Two rules make this safe:

  • Keep the TTL well below the rotation interval. A cache that lives longer than the time between rotations will regularly hold an outdated value.
  • Refresh once on authentication failure. If the downstream service rejects the credential, fetch a fresh copy and retry a single time before giving up.
private static string? _cachedSecret;
private static DateTime _cachedAt;
private static readonly TimeSpan Ttl = TimeSpan.FromMinutes(5);

private static async Task<string> GetSecretAsync(bool forceRefresh = false)
{
    if (!forceRefresh && _cachedSecret is not null && DateTime.UtcNow - _cachedAt < Ttl)
        return _cachedSecret;

    var response = await Secrets.GetSecretValueAsync(
        new GetSecretValueRequest { SecretId = "my-service/api-key" });

    _cachedSecret = response.SecretString;
    _cachedAt = DateTime.UtcNow;
    return _cachedSecret;
}

private static async Task<HttpResponseMessage> CallApiAsync(HttpClient http, string url)
{
    var response = await SendAsync(http, url, await GetSecretAsync());

    if (response.StatusCode == HttpStatusCode.Unauthorized)
    {
        response.Dispose();
        response = await SendAsync(http, url, await GetSecretAsync(forceRefresh: true));
    }

    return response;
}

Limit the retry to one attempt. If the second call also fails, the cause is probably not rotation, and retrying further only hides a real problem and adds latency. Also be careful about which status codes you treat as a signal: a 403 often means missing permissions rather than a bad credential, so retrying on it can be wasted effort.

If you would rather not write this yourself, AWS provides a caching client for .NET and a Lambda extension that caches secrets outside your code. Both give you a configurable TTL, but check the current documentation for their defaults and limits before relying on them.

Choosing between the two

Approach Best when Trade-off
Read on every invocationModerate traffic, simplicity matters more than a few millisecondsExtra call and cost per invocation
TTL cache with one retryHigh traffic, latency-sensitive pathsMore code; brief window where an old value may be used
Cache forever (init only)Secrets that never rotateBreaks as soon as rotation is enabled

When rotation is worth it

Rotation pays off most for long-lived credentials that cannot easily be replaced by short-lived ones: database passwords, third-party API keys and service account secrets. A leaked credential of that kind stays useful to an attacker indefinitely unless something changes it, and rotation puts an upper bound on that window.

It adds less for credentials that already expire quickly, such as temporary credentials issued by AWS STS or access tokens that last an hour. If you can use those instead of a static secret, you may not need rotation at all. Also check how the provider behaves: some OAuth providers issue refresh tokens that stay valid until revoked and do not change on use, so rotating them means working through the provider's own token management rather than a simple replace.

Designing the rotation function

Secrets Manager runs a rotation Lambda function in four steps: createSecret, setSecret, testSecret and finishSecret. During the process the new value is stored under the AWSPENDING staging label, and only after it has been tested does it become AWSCURRENT, with the previous value moving to AWSPREVIOUS. Your code should respect that flow:

  • Test before promoting. The testSecret step should use the new credential against the real target. If it fails, the current value must remain untouched so the application keeps working.
  • Make each step safe to repeat. Secrets Manager can call a step more than once, so running setSecret twice should not leave the target in a broken state.
  • Never overwrite the working value early. Writing a new credential into the secret before verifying it removes your fallback.

For databases, there are two common strategies. A single-user strategy changes the password of one account, which means connections opened with the old password may stop working at the moment of change. An alternating-users strategy keeps two accounts and switches between them, so the previous credential stays valid for a while after rotation. That overlap is what gives cached values time to expire safely, which is why it is the better fit for services that cache. Confirm the details for your database engine in the AWS documentation.

Seeing rotation problems before users do

Because rotation failures surface somewhere else, log the events that let you connect the two sides:

  • In the rotation function, log each step with the secret identifier and version, so a failed rotation can be traced.
  • In the service, log when a secret is fetched, when a downstream call is rejected, and when a refresh-and-retry happens. A burst of retries that lines up with a rotation time is a clear signal.
  • Outside both, watch rotation events through CloudTrail and EventBridge, and alert on rotations that fail or fall overdue. Check the current documentation for the exact events and metrics available.

A sensible order of work

  1. Decide whether the secret needs rotation at all, or whether a short-lived credential would do.
  2. Remove any read-once-at-startup pattern for secrets that will rotate.
  3. Start with a per-invocation read; move to a TTL cache only if measurements show it is needed.
  4. If you cache, set the TTL below the rotation interval and add a single refresh-and-retry on authentication failure.
  5. Write the rotation function so it tests the new credential before promoting it, and prefer an overlap strategy where the target supports it.
  6. Add logging and alerts on both sides, then trigger a test rotation and watch what the service does.

The aim is not a service that never sees a rejected credential. It is one that notices quickly, recovers on its own and leaves enough evidence to explain what happened. Run a manual rotation in a test environment before you trust it in production.

Comments