Skip to main content

Posts

Right-Sizing AWS Lambda Memory: Matching Configuration to Observed Workload Instead of Guessing

When you deploy an AWS Lambda function, you choose a memory allocation. AWS scales CPU and network proportionally to that memory setting. Most teams pick a number—512 MB, 1024 MB, maybe 2048 MB—and move on. The function works, so the setting never changes. A year later, you are paying for memory the function never uses, or the function runs slower than it should because you chose conservatively and never revisited the decision. I have done this. I have deployed Lambda functions with 1024 MB of memory because it felt safe, only to discover months later that the function used 200 MB at peak and completed in half the time when I reduced the allocation to 512 MB. I have also under-provisioned functions, then spent time investigating timeout errors that disappeared when I doubled the memory and gave the function more CPU. Right-sizing Lambda memory is not a one-time configuration decision. It is an ongoing operational practice that requires measurement, not intuition. The cost differen...
Recent posts

Characterization Tests: Writing Tests for Legacy Code Before You Understand It

When you inherit a legacy .NET system that needs modernization, the instinct is to read the code until you understand it, then write tests. That sequence is backward. If the system works in production, you need to lock down its behavior before you refactor anything, and you need to do it without knowing why the code does what it does. Characterization tests solve this problem. They document observed behavior, not intended behavior, and they catch regressions while you're still learning the system. I've used characterization tests to modernize several legacy .NET systems where the original developers were gone, documentation was sparse, and the only source of truth was production. The goal was not to prove the code was correct. The goal was to prevent silent breakage while I replaced tightly coupled classes, swapped out WCF endpoints for REST APIs, and introduced dependency injection. Characterization tests are the safety net that makes incremental modernization possible. T...

EventBridge Scheduler Pitfalls: Context Attributes, Retries, and Dead-Letter Design

I discovered the hard way that EventBridge Scheduler behaves differently from the event-driven services most teams are used to. When a scheduled Lambda invocation fails, the retry logic is not the same as SQS, Step Functions, or even direct Lambda invocations through the console. The target configuration, the dead-letter queue setup, and the context attributes passed to your function all require deliberate decisions that are easy to overlook until your automation silently stops working. This post walks through the three pitfalls I encountered while building scheduled automation on AWS Lambda, the configuration gaps that caused failures, and the design patterns that made scheduled invocations reliable. If you are migrating from CloudWatch Events or adding scheduled tasks to a serverless application, these details will save you from silent failures and duplicate work. Context Attributes Are Not Automatic When EventBridge Scheduler invokes a Lambda function, it does not automatical...

Strangler-Fig Migrations: Replacing a Legacy .NET Monolith One Seam at a Time

When you inherit a legacy .NET system that runs a production business, you face a choice: rewrite it in one risky cutover or replace it incrementally while features keep shipping. The strangler-fig pattern—wrapping the old system and routing new behavior to new services while the monolith continues to handle existing traffic—sounds appealing in architecture diagrams. In practice, it forces you to answer a harder question: which seam do you split first, and how do you prove the replacement works before you remove the old code? I have modernized legacy .NET systems using this approach. The pattern works, but only if you choose boundaries based on technical isolation rather than organizational convenience, and only if you accept that both systems will run in production for months or years. This post walks through the decision framework, the technical mechanisms that make the split safe, and the operational consequences that matter more than the architecture slides suggest. Why the F...

Least-Privilege IAM and Secret Handling as Architecture Decisions, Not Compliance Cleanup

When I review cloud architecture in modernization projects, I often find IAM policies and secret-handling patterns treated as post-deployment security tasks rather than foundational design decisions. Teams grant broad permissions during development, defer secret rotation until a compliance audit surfaces the risk, and treat least-privilege access as a hardening step that happens after the application works. This approach consistently produces architectures that are expensive to secure later and difficult to operate safely at scale. I learned this lesson the hard way. In one AWS data pipeline, early prototypes used a single service role with wide S3, Redshift, and Lambda permissions. The role worked across environments, simplified local testing, and allowed rapid iteration. When the pipeline moved to production and began processing regulated data, tightening those permissions required rewriting deployment scripts, breaking existing integrations, and coordinating changes across multip...

Building Resilient Data Connectors When External APIs Fail, Drift, or Throttle

External APIs break in predictable ways. They rate-limit your requests during peak hours. They change field names without warning. They return HTTP 200 with an error object nested three levels deep in JSON. They timeout after exactly 29 seconds, one second before your own timeout fires. If your data connector treats these as exceptional cases, your pipeline stops working the day your business depends on it. I have built data connectors for regulated environments where the external API was a legacy SOAP service, a partner REST endpoint with unpredictable availability, and a SaaS vendor that changed response schemas without versioning. The lesson from all three is the same: resilience is not a retry loop wrapped around an HTTP client. It is a design decision about failure ownership, state management, and operational visibility. Failure Modes You Must Design For, Not Catch Later Most data connector failures fall into three categories: transient errors you can retry, permanent error...

Why CI/CD Still Fails When the Pipeline Turns Green

A green build does not mean your deployment worked. I learned this the hard way while modernizing a legacy .NET system that had spent years accumulating manual release steps outside the automation. The CI/CD pipeline would pass, the artifact would land in the environment, and then someone would remember the configuration flag that never made it into the repository, or the database migration script that lived in a wiki, or the API key rotation that happened through a support ticket. The pipeline said success, but the application was broken in production. The problem was not the CI/CD tool. We were using GitHub Actions, which is capable and flexible. The problem was that CI/CD had been treated as a checkbox—something you set up once to automate the build and maybe run a few tests—rather than a system designed to make deployments smaller, safer, and repeatable. The hidden manual steps were not documented anywhere the pipeline could see them, so they became failure points every time we ...