Skip to main content

Right-Sizing AWS Lambda Memory: Matching Configuration to Observed Workload Instead of Guessing

When you deploy an AWS Lambda function, you choose a memory allocation. AWS scales CPU and network proportionally to that memory setting. Most teams pick a number—512 MB, 1024 MB, maybe 2048 MB—and move on. The function works, so the setting never changes. A year later, you are paying for memory the function never uses, or the function runs slower than it should because you chose conservatively and never revisited the decision.

I have done this. I have deployed Lambda functions with 1024 MB of memory because it felt safe, only to discover months later that the function used 200 MB at peak and completed in half the time when I reduced the allocation to 512 MB. I have also under-provisioned functions, then spent time investigating timeout errors that disappeared when I doubled the memory and gave the function more CPU.

Right-sizing Lambda memory is not a one-time configuration decision. It is an ongoing operational practice that requires measurement, not intuition. The cost difference between an over-provisioned function and a correctly sized one compounds across invocations. The performance difference between an under-provisioned function and one with adequate CPU shows up as latency, retries, and wasted developer time.

Why Memory Configuration Controls More Than Memory

AWS Lambda charges by gigabyte-seconds: the product of memory allocation and execution duration. If you allocate 1024 MB and the function runs for 500 milliseconds, you pay for 512 MB-seconds. If you reduce memory to 512 MB and the function still runs for 500 milliseconds, you pay half as much. But the function will not necessarily run for the same duration, because memory allocation also controls CPU.

AWS does not publish the exact CPU scaling curve, but the relationship is linear. A function with 1024 MB gets twice the CPU of a function with 512 MB. A function with 1792 MB gets a full vCPU. Beyond that threshold, you get additional vCPUs proportionally. If your function is CPU-bound—parsing JSON, running cryptographic operations, or transforming data—it will run faster with more memory, even if it never uses the extra RAM.

This creates a tradeoff. You can allocate less memory and pay a lower rate per millisecond, but the function may run longer. You can allocate more memory and pay a higher rate, but the function may finish faster and consume fewer total gigabyte-seconds. The optimal allocation is not the smallest amount of memory the function can run in. It is the allocation that minimizes total cost while meeting latency requirements.

Measuring Actual Memory and Duration

CloudWatch Logs includes memory usage in the report line at the end of each invocation. It looks like this:

REPORT RequestId: abc123 Duration: 482.34 ms Billed Duration: 483 ms Memory Size: 1024 MB Max Memory Used: 203 MB

The Max Memory Used value is the peak memory consumption during that invocation. If you see 203 MB used with 1024 MB allocated, the function has 821 MB of unused memory. That suggests you can reduce the allocation, but it does not tell you whether reducing memory will increase duration enough to offset the cost savings.

I prefer to capture this data over a sample of invocations, not a single run. Lambda functions that process variable payloads—reading objects from S3, calling external APIs, transforming datasets—can have different memory profiles depending on input size. I export CloudWatch Logs to S3 and query the report lines, or I use CloudWatch Logs Insights to aggregate statistics:

fields @timestamp, @billedDuration, @memorySize, @maxMemoryUsed
| filter @type = "REPORT"
| stats avg(@maxMemoryUsed) as avgMemory, max(@maxMemoryUsed) as peakMemory, avg(@billedDuration) as avgDuration by @memorySize

This query groups invocations by memory size and shows average and peak memory usage alongside average duration. If the function is already deployed with multiple memory settings—perhaps you tested 512 MB and 1024 MB in different environments—you can compare the duration and compute the cost per invocation for each configuration.

Testing Memory Allocations to Find the Optimal Setting

I change the memory allocation in the Lambda configuration and invoke the function with representative workloads. I do not test with synthetic payloads that are smaller or simpler than production traffic, because that will underestimate both memory and duration. I test with recorded production payloads or realistic samples.

For a .NET Lambda function that processes API requests and calls downstream services, I tested four memory settings: 512 MB, 1024 MB, 1536 MB, and 2048 MB. I invoked the function 100 times at each setting with the same set of requests and recorded the average billed duration and memory usage:

  • 512 MB: 680 ms average duration, 420 MB peak memory
  • 1024 MB: 480 ms average duration, 430 MB peak memory
  • 1536 MB: 460 ms average duration, 435 MB peak memory
  • 2048 MB: 450 ms average duration, 440 MB peak memory

The function never used more than 440 MB of memory, so every allocation above 512 MB was paying for unused RAM. But the duration dropped significantly from 512 MB to 1024 MB, then flattened. The function was CPU-bound at 512 MB, and giving it more CPU reduced execution time. Beyond 1024 MB, additional CPU had diminishing returns.

I calculated cost per invocation using AWS Lambda pricing for the region. At 512 MB, the function cost $0.0000057 per invocation. At 1024 MB, it cost $0.0000050 per invocation, despite the higher memory rate, because the shorter duration reduced total gigabyte-seconds. At 1536 MB and 2048 MB, the cost increased because the duration savings did not offset the higher rate.

The optimal setting was 1024 MB: lowest cost per invocation and acceptable latency. I deployed that configuration to production and set a CloudWatch alarm to alert if max memory usage exceeded 850 MB, which would indicate the workload had changed and I needed to reevaluate.

When Memory Allocation Affects Cold Start Latency

Cold starts—initializing a new execution environment when no warm instance is available—take longer at lower memory allocations. The Lambda service provisions CPU proportional to memory during initialization, so functions with more memory complete the init phase faster. For .NET Lambda functions running in a container image, the init phase includes pulling the image, extracting layers, and starting the runtime. That work is CPU-intensive.

I measured cold start latency for the same .NET function at different memory settings by forcing a cold start and recording the init duration from CloudWatch Logs:

  • 512 MB: 2,100 ms init duration
  • 1024 MB: 1,400 ms init duration
  • 1536 MB: 1,200 ms init duration
  • 2048 MB: 1,150 ms init duration

Doubling memory from 512 MB to 1024 MB reduced cold start time by 33%. Going higher than 1024 MB had smaller improvements. If the function serves user-facing requests and cold start latency matters, allocating 1024 MB instead of 512 MB is worth the marginal cost increase. If the function runs on a schedule or processes queue messages where cold starts are rare and latency is not critical, 512 MB may be acceptable.

Revisiting the Setting When the Workload Changes

The optimal memory allocation is not static. If the function starts processing larger payloads, calling additional downstream services, or performing more compute-intensive transformations, memory usage and duration will change. I treat memory configuration as an operational metric, not a deployment constant.

I monitor max memory usage in CloudWatch and set an alarm that triggers when usage exceeds 80% of the allocated memory. That threshold gives me time to test a higher allocation before the function runs out of memory and fails. I also track the P95 and P99 duration in production and alert if duration increases without a corresponding change in payload size or downstream latency. That signal suggests the function is CPU-bound and would benefit from more memory.

When I add significant new logic to a Lambda function—such as integrating a new API, adding data validation, or introducing a caching layer—I retest memory allocations in a staging environment before deploying to production. I do not assume the previous setting is still optimal.

A Practical Framework for Right-Sizing Lambda Memory

Start with a reasonable default, measure actual usage, test alternatives, and choose the allocation that minimizes cost while meeting latency requirements. Do not guess. Do not set memory once and forget it. Do not allocate more memory than you need because it feels safer.

Use CloudWatch Logs to capture memory usage and duration over a representative sample of invocations. Test at least three memory settings—your current allocation, one step lower, and one step higher—with realistic workloads. Calculate cost per invocation using the billed duration and memory allocation. Choose the setting that delivers acceptable latency at the lowest cost, and set alarms to detect when the workload changes.

If your function is CPU-bound, increasing memory will reduce duration and may reduce total cost. If your function is memory-bound, reducing memory allocation will lower cost without increasing duration. If your function is I/O-bound—waiting for network calls, database queries, or external APIs—changing memory allocation will have little effect on duration, and you should choose the smallest allocation that keeps memory usage below 80%.

Right-sizing Lambda memory is not a one-time task. It is a decision you revisit when the workload changes, when you refactor the function, or when cost becomes a visible operational metric. Measurement replaces guesswork. Observability makes the tradeoff clear.

Comments