Skip to main content

Lambda Cold Starts in Serverless .NET: What Actually Helps Versus Folklore Fixes

Cold starts in AWS Lambda are one of the most discussed and least understood performance problems in serverless .NET applications. I have built and maintained AWS Lambda services for years, and I have watched teams apply folklore fixes that do not address the actual latency contributors. Some optimizations help. Others add complexity without measurable improvement. The difference comes down to understanding what happens during a cold start and which parts of that process you can actually control.

A Lambda cold start occurs when AWS provisions a new execution environment to handle a request. The runtime must initialize, the .NET assembly must load, and your application code must execute its startup logic before the first request completes. For .NET 6 and later on Lambda, this sequence includes downloading the deployment package, extracting it, loading the Common Language Runtime, JIT-compiling code, and running static constructors and dependency injection configuration. Each step contributes latency, but not all steps respond to the same optimizations.

The Latency Budget You Cannot Change

AWS controls the infrastructure provisioning time. You cannot optimize the time it takes AWS to allocate compute, attach network interfaces, or download your deployment package from S3. This baseline accounts for a significant portion of cold start latency, particularly for functions deployed in a VPC. VPC-attached functions incur additional delay as AWS creates and attaches elastic network interfaces. I have measured VPC cold starts at 8 to 12 seconds for .NET 6 Lambda functions, compared to 2 to 4 seconds for functions without VPC attachment. If your function does not need to access VPC resources such as RDS or private APIs, removing VPC configuration is the single most effective cold start optimization.

The deployment package size affects download time, but the relationship is not linear. A 10 MB package and a 50 MB package produce similar cold start latency if the limiting factor is provisioning time, not download time. I have seen teams spend days trimming dependencies to reduce package size from 40 MB to 25 MB, only to observe no measurable change in cold start duration. Package size matters more when you cross the Lambda layer limit or when the function uses container images larger than 500 MB, but for most .NET applications, package size is not the dominant contributor to cold start latency.

Assembly Loading and JIT Compilation

.NET Lambda functions must load the CLR and JIT-compile IL to native code on every cold start. This process is deterministic and tied to the size and complexity of your code, not just the deployment package size. A function that references dozens of NuGet packages but only uses a small fraction of their APIs still loads those assemblies and pays the JIT cost for types instantiated during startup.

The optimization that consistently reduces this latency is limiting the dependency graph. Remove unused NuGet packages. Avoid large frameworks that initialize eagerly. Lazy-load services that are not required for every invocation. I replaced a dependency on a full JSON schema validation library with a lightweight manual validator for the specific schemas my function handled. The cold start improved by 400 milliseconds because the runtime no longer loaded and JIT-compiled hundreds of unused types.

Ahead-of-time compilation using ReadyToRun can reduce JIT latency by pre-compiling IL to native code during the build. However, ReadyToRun increases deployment package size, and the tradeoff is not always favorable. I tested ReadyToRun on a Lambda function with moderate dependency complexity and observed a 200-millisecond cold start improvement at the cost of a 15 MB larger package. For functions invoked frequently enough that cold starts are rare, the larger package is not worth the marginal gain. For functions with unpredictable traffic and frequent cold starts, ReadyToRun can help, but you must measure it in your actual workload rather than assuming it will always improve latency.

Startup Logic and Dependency Injection

The code your function executes during initialization is under your control and often the most significant opportunity for cold start optimization. Dependency injection containers, configuration providers, HTTP clients, and SDK clients all add startup latency if you initialize them on every cold start. The Lambda execution model reuses environments across multiple invocations, so you can initialize expensive resources once and reuse them.

Move initialization code outside the handler method into static constructors or lazy singletons. For example, I moved AWS SDK client initialization from the handler to a static field. The SDK client initializes once per execution environment, not once per request. This reduced cold start latency by 300 milliseconds and eliminated redundant SDK configuration on every invocation.

public class Function
{
    private static readonly AmazonS3Client S3Client = new AmazonS3Client();

    public async Task<APIGatewayProxyResponse> FunctionHandler(
        APIGatewayProxyRequest request, ILambdaContext context)
    {
        // S3Client is already initialized
        var response = await S3Client.GetObjectAsync("my-bucket", "key");
        // handle response
    }
}

The static readonly field ensures the S3 client initializes once per execution environment. Subsequent invocations reuse the same client instance without re-initializing the SDK or its internal HTTP connection pool. This pattern applies to any expensive resource: database connections, HTTP clients, serializers, configuration objects, or cryptographic providers.

Avoid initializing resources you do not need for every request. If your function handles multiple event types and only some require database access, initialize the database client lazily when the first request that needs it arrives. Use Lazy<T> to defer initialization until the resource is accessed:

private static readonly Lazy<SqlConnection> DbConnection = 
    new Lazy<SqlConnection>(() => 
        new SqlConnection(Environment.GetEnvironmentVariable("DB_CONNECTION_STRING")));

This defers the connection initialization cost to the first invocation that calls DbConnection.Value, avoiding the penalty for requests that do not need the database.

Provisioned Concurrency and When It Makes Sense

Provisioned concurrency keeps a fixed number of execution environments initialized and ready to handle requests. AWS charges for the provisioned capacity regardless of whether requests arrive, so this is a cost-versus-latency tradeoff. Provisioned concurrency eliminates cold starts for the provisioned instances, but it does not eliminate cold starts entirely. If traffic exceeds provisioned capacity, AWS creates new on-demand instances that experience normal cold starts.

I use provisioned concurrency only when cold start latency violates a strict latency SLA and traffic patterns are predictable enough to justify the cost. For example, a customer-facing API with a 200-millisecond p99 latency requirement and steady traffic during business hours benefits from provisioned concurrency. A background processing function that runs every hour and tolerates multi-second cold starts does not.

Provisioned concurrency is not a substitute for optimizing cold start latency. A function with 5-second cold starts is still expensive to keep warm. Reduce cold start latency through the techniques above before adding provisioned concurrency. You will either avoid the need for provisioned capacity or reduce the number of instances you must keep warm.

Folklore Fixes That Do Not Help

I have seen several cold start optimizations repeated across blog posts and AWS forums that either do not work or produce negligible improvement. Testing them in your workload is the only way to know, but these rarely justify the effort:

  • Increasing memory allocation to improve cold start time: Memory configuration affects CPU allocation during execution, but it does not change provisioning time or assembly loading. I tested memory settings from 512 MB to 3008 MB and observed no consistent cold start improvement. Memory affects execution duration and cost, not cold start latency.
  • Warming functions with scheduled pings: This works only if the ping keeps every execution environment warm. Lambda scales horizontally, so a single scheduled invocation every few minutes keeps one environment warm. If traffic creates additional environments, those still cold start. Provisioned concurrency is a more reliable solution if you need guaranteed warm capacity.
  • Splitting a function into smaller functions to reduce package size: This adds operational complexity and does not address the actual latency contributors. If the original function has a large dependency graph, each smaller function likely still references most of the same dependencies. The result is more functions with similar cold start characteristics and more complex deployment and monitoring.
  • Using Lambda layers to share dependencies: Layers reduce deployment package size, but they do not reduce assembly loading or JIT compilation time. The runtime still loads the same assemblies and compiles the same code. Layers are useful for sharing common dependencies across multiple functions and reducing deployment size limits, but they are not a cold start optimization.

A Framework for Cold Start Optimization

Start by measuring cold start latency in your actual workload. Use AWS X-Ray or CloudWatch Logs Insights to identify cold start invocations and calculate p50, p95, and p99 latency. Compare cold start latency to warm invocation latency to understand the cost. If cold starts represent a small fraction of traffic and do not violate latency SLAs, optimization may not be necessary.

If cold starts are a problem, apply optimizations in order of impact. Remove VPC attachment if the function does not need VPC resources. Trim unused dependencies and lazy-load expensive resources. Test ReadyToRun if JIT compilation is a significant contributor. Move SDK clients and configuration to static initialization. Measure each change to confirm it improves latency in your environment.

Consider provisioned concurrency only after optimizing the cold start itself. A function with 1-second cold starts is cheaper to keep warm than a function with 5-second cold starts. If provisioned concurrency is necessary, start with the minimum capacity that meets your latency SLA and scale up based on observed traffic patterns.

Cold start optimization is a tradeoff between latency, cost, and complexity. The goal is not zero cold starts, but cold starts that are fast enough and infrequent enough that they do not degrade user experience or violate SLAs. Focus on the changes that address the actual latency contributors in your workload, and measure the results instead of applying folklore fixes.

Comments