When you inherit a legacy .NET system that runs a production business, you face a choice: rewrite it in one risky cutover or replace it incrementally while features keep shipping. The strangler-fig pattern—wrapping the old system and routing new behavior to new services while the monolith continues to handle existing traffic—sounds appealing in architecture diagrams. In practice, it forces you to answer a harder question: which seam do you split first, and how do you prove the replacement works before you remove the old code?
I have modernized legacy .NET systems using this approach. The pattern works, but only if you choose boundaries based on technical isolation rather than organizational convenience, and only if you accept that both systems will run in production for months or years. This post walks through the decision framework, the technical mechanisms that make the split safe, and the operational consequences that matter more than the architecture slides suggest.
Why the First Seam Matters More Than the Migration Plan
The strangler-fig pattern depends on finding a seam—a boundary where you can route some traffic to a new service and compare its behavior against the monolith without breaking production. The first seam sets the precedent for every subsequent split. Choose wrong, and you spend months untangling shared state, duplicate business logic, and data-consistency failures that only appear under load.
The best first seam is a vertical slice with clear input and output contracts, minimal shared state, and low coupling to the rest of the monolith. In .NET systems, this often means a background job, a reporting module, or an API endpoint that reads from a bounded dataset and does not trigger cascading writes across the system. The worst first seam is a foundational layer—like authentication, logging, or a shared data-access object—that every other component depends on. Splitting a dependency first creates a migration path where you must update every caller before you can remove the old code.
When I modernized a legacy .NET system with tightly coupled WCF services and stored procedures, the first seam was a nightly batch job that aggregated reporting data. The job had a clear schedule, known inputs from SQL Server, and a single output table. We replaced it with a .NET 6 console application running in a Docker container on a schedule, wrote the same output to the same table, and ran both implementations in parallel for two weeks. The new service used dependency injection, nUnit tests, and Moq to isolate external dependencies—patterns the monolith did not have. When the outputs matched, we deleted the old job and moved to the next boundary.
Running Both Implementations in Parallel Without Corrupting State
The strangler-fig pattern requires running the old and new systems simultaneously, which means you must prevent them from corrupting shared state. If both write to the same database table without coordination, you introduce race conditions, duplicate records, or conflicting updates. If the new service reads stale data that the monolith has not yet written, you observe inconsistencies that only appear in production.
The safest parallel execution strategy is to run the new service in shadow mode: it processes the same inputs as the monolith, writes to a separate output location, and does not affect production traffic. You compare the outputs offline, measure differences, and iterate until they converge. Only then do you switch production traffic to the new service and retire the old implementation.
For the reporting job, we added a shadow output table with the same schema as the production table. The new service wrote to the shadow table, and a SQL query compared row counts, sums, and sample records every morning. When we found discrepancies, we traced them to rounding differences in decimal calculations and time-zone handling in date filters—bugs that existed in the monolith but had never been visible. We fixed the new service to match the existing behavior, even when the existing behavior was wrong, because changing output format during a migration introduces a second variable into the comparison.
Once the outputs matched for two weeks without manual intervention, we updated the new service to write to the production table and disabled the old job. The monolith still ran; we had only replaced one small piece. The next seam was an API endpoint that generated PDF reports, then a data-import connector, then a background worker that sent email notifications. Each split followed the same pattern: isolate, shadow, compare, switch, delete.
Routing Traffic to the New Service Without Breaking Existing Clients
Splitting a public API endpoint requires a routing layer that can direct traffic to the old or new implementation based on a feature flag, a request header, a tenant identifier, or a percentage rollout. The routing layer must preserve the original API contract so that clients do not know a migration is happening. If the new service returns a different HTTP status code, a different error message format, or a different JSON schema, you have introduced a breaking change that will surface as client errors in production.
In the .NET modernization, we placed an API Gateway in front of the monolith and the new microservices. The gateway inspected a custom HTTP header—initially set only in our internal test client—and routed requests with that header to the new service. All other traffic continued to the monolith. The new service implemented the same REST API contract as the monolith: identical URL paths, query parameters, request and response schemas, and error codes. We validated the contract using integration tests that called both services with the same inputs and asserted identical outputs.
[Test]
public async Task GetReport_ShouldMatchLegacyService()
{
var request = new ReportRequest { StartDate = new DateTime(2024, 1, 1), EndDate = new DateTime(2024, 1, 31) };
var legacyResponse = await _legacyClient.GetReportAsync(request);
var newResponse = await _newClient.GetReportAsync(request);
Assert.AreEqual(legacyResponse.StatusCode, newResponse.StatusCode);
Assert.AreEqual(legacyResponse.TotalRecords, newResponse.TotalRecords);
CollectionAssert.AreEqual(legacyResponse.Data, newResponse.Data);
}
This test ran in CI/CD on every deployment. If the new service diverged from the monolith, the build failed before the change reached production. The test was not exhaustive—we could not enumerate every possible input—but it caught schema drift, null-handling bugs, and changes in default sorting.
After the header-based routing worked in staging, we replaced the header check with a percentage-based rollout: 5% of production traffic went to the new service, then 25%, then 50%, then 100%. We monitored error rates, latency percentiles, and database query counts at each step. When the new service showed higher latency, we traced it to a missing SQL index that the monolith's query plan had relied on. We added the index, verified the query plan, and continued the rollout.
When to Delete the Old Code and When to Leave It
The hardest part of a strangler-fig migration is knowing when to delete the old implementation. Delete too early, and you lose the ability to roll back when the new service fails in ways your tests did not predict. Delete too late, and you carry the operational cost of running two systems, maintaining two codebases, and reasoning about two failure modes for every incident.
I wait until the new service has handled 100% of production traffic for at least two weeks without a rollback, without a bug that required a hotfix, and without an operational question that could only be answered by inspecting the old code. At that point, the new service has survived realistic load, realistic input variation, and realistic failure modes. The old code is no longer a safety net; it is technical debt.
We deleted each replaced component from the monolith after the two-week soak period. We did not delete the database tables or stored procedures immediately—those stayed as read-only archives until we confirmed no downstream system depended on them. We did remove the old API routes, the old background jobs, and the old WCF service contracts. Each deletion reduced the surface area of the monolith and made the next seam easier to identify.
The Operational Cost of Running Two Systems
Strangler-fig migrations are not free. You run two implementations of the same feature, which means two sets of logs to search, two sets of metrics to monitor, two deployment pipelines, and two places where an incident can start. You must design the new services to coexist with the monolith's assumptions about database transactions, shared caches, and background job schedules. If the monolith locks a table during a nightly batch process, your new service must respect that lock or risk deadlocks.
The operational cost is highest in the middle of the migration, when half the system is modern .NET 6 microservices with CI/CD and structured logging, and half is the legacy monolith with manual deployments and logs written to text files. Incidents span both worlds. A slow API response might be caused by the new service, the monolith, the shared database, or the network between them. You need distributed tracing to connect the request across both systems, but the monolith does not emit trace context, so you instrument the API Gateway to inject trace IDs and pass them to both sides.
The cost justifies itself only if you use the migration as an opportunity to improve the architecture, not just rewrite the same code in a newer framework. We added dependency injection to decouple business logic from infrastructure, wrote unit tests with Moq to isolate external dependencies, introduced CI/CD with GitHub Actions to automate deployments, and designed the new services to fail gracefully when the database is unavailable. The monolith had none of those properties. The migration was slower than a rewrite, but it delivered incremental value, reduced risk, and left us with a system we could evolve without another multi-year project.
A Framework for Choosing the Next Seam
After the first seam, every subsequent split follows the same questions:
- Does this component have a clear input-output contract, or does it depend on shared mutable state?
- Can I run the old and new implementations in parallel without corrupting production data?
- Can I route traffic to the new service and roll back without downtime?
- Does this split reduce coupling in the system, or does it just move complexity to a network boundary?
- Will this seam let me delete old code within a month, or will it create a permanent integration layer?
If the answer to any of these is no, choose a different seam. The goal is not to split the monolith as fast as possible; it is to split it in a way that improves the system's operability, testability, and deployability. Some boundaries are worth creating. Others just turn one hard-to-change monolith into five hard-to-change microservices with network latency between them.
Strangler-fig migrations work when you treat them as a series of small, reversible decisions rather than a grand architecture transformation. Split one seam, prove it works, delete the old code, and move to the next boundary. The migration is never finished—there is always another component to modernize—but each step leaves the system better than you found it, and each step ships to production without pausing feature delivery.
Comments
Post a Comment