A green build does not mean your deployment worked. I learned this the hard way while modernizing a legacy .NET system that had spent years accumulating manual release steps outside the automation. The CI/CD pipeline would pass, the artifact would land in the environment, and then someone would remember the configuration flag that never made it into the repository, or the database migration script that lived in a wiki, or the API key rotation that happened through a support ticket. The pipeline said success, but the application was broken in production.
The problem was not the CI/CD tool. We were using GitHub Actions, which is capable and flexible. The problem was that CI/CD had been treated as a checkbox—something you set up once to automate the build and maybe run a few tests—rather than a system designed to make deployments smaller, safer, and repeatable. The hidden manual steps were not documented anywhere the pipeline could see them, so they became failure points every time we shipped.
The Hidden Manual Steps That Break Deployments
Manual steps accumulate because they start as exceptions. Someone needs to update a connection string in production, so they log into the server and edit a config file. A feature flag needs to toggle before the next release, so it goes into a shared spreadsheet. A database migration script is too risky to automate, so it gets run by hand during a maintenance window. Each of these steps is reasonable in isolation, but together they create a deployment process that cannot be reproduced by the pipeline.
When I joined the .NET modernization effort, the team had a detailed release checklist that lived in Confluence. It had twenty-three steps. The CI/CD pipeline covered six of them. The rest were manual: apply migrations, restart specific services in a specific order, verify API keys, update feature flags, clear caches, and notify downstream systems. If you forgot one, the deployment would fail in a way that was hard to diagnose because the pipeline had no visibility into what had actually happened in the environment.
The first deployment I participated in took four hours and failed twice. The pipeline ran in twelve minutes. The rest of the time was spent coordinating manual steps, troubleshooting configuration drift, and waiting for someone with production access to fix a permission issue that only surfaced after the deployment.
Smaller Deployments Reduce the Blast Radius
The most effective change we made was not a better tool or a more sophisticated pipeline. It was a policy: deploy smaller changes more frequently. Instead of batching a week of work into a single release, we started deploying individual pull requests as soon as they were reviewed and tested. This forced us to confront the hidden manual steps because they became bottlenecks we hit multiple times per day instead of once per week.
Smaller deployments also reduced the risk of each release. When a deployment contains one feature or one bug fix, it is obvious what broke if something goes wrong. When a deployment contains twenty changes, the root cause could be any of them, or an interaction between several of them. Debugging a failed deployment with twenty changes is expensive. Debugging a failed deployment with one change is fast.
We broke down the large .NET monolith into independently deployable services where the boundaries made sense. Not every boundary was worth creating—splitting too aggressively would have created needless coordination overhead—but the core application had clear seams where business logic, data access, and external integrations could be separated. Each service had its own CI/CD pipeline that built, tested, and deployed to a staging environment automatically. Production deployments required approval, but the deployment itself was automated.
Automated Rollbacks and the Deployment Manifest
A useful CI/CD pipeline must make rollbacks as easy as deployments. If rolling back requires manual coordination or tribal knowledge, you will hesitate to roll back when something goes wrong. That hesitation costs time and increases the impact of the failure.
We implemented rollbacks by treating each deployment as an immutable artifact with a manifest. The manifest included the application version, the configuration version, the database migration version, and the feature flag state. When a deployment failed, rolling back meant deploying the previous manifest. The pipeline knew how to apply the previous configuration, run the reverse migrations if needed, and toggle the feature flags back to their prior state.
Here is a simplified example of the deployment step in our GitHub Actions workflow that applied the configuration and deployed the artifact:
- name: Deploy to staging
run: |
aws s3 cp config/${{ github.sha }}.json s3://app-config-bucket/staging/config.json
aws lambda update-function-code \
--function-name my-service-staging \
--image-uri ${{ env.ECR_REGISTRY }}/my-service:${{ github.sha }}
aws lambda wait function-updated \
--function-name my-service-staging
env:
AWS_REGION: us-east-1
The configuration file was versioned by commit SHA and stored in S3. The Lambda function pulled its configuration from S3 at startup. If the deployment failed, we could redeploy the previous commit SHA, and the pipeline would apply the previous configuration automatically. No manual coordination required.
This approach only worked because we eliminated configuration drift. Every environment's configuration was stored in the repository and applied by the pipeline. If someone needed to change a setting, they opened a pull request. The change was reviewed, tested in staging, and then promoted to production through the same pipeline. This removed the shared spreadsheets, the Confluence pages, and the manual edits that had caused so many deployment failures.
The Cost of Invisible State Changes
Database migrations were the hardest manual step to automate because they carried real risk. A bad migration could corrupt data or lock tables in production. The team had been running migrations by hand during maintenance windows, which meant deployments were slow and had to be scheduled in advance.
We automated migrations by making them part of the deployment artifact and enforcing strict rules:
- Every migration must be reversible. If the deployment fails, the pipeline must be able to roll back the schema change.
- Migrations must be idempotent. Running the same migration twice must be safe.
- Migrations must not lock tables for more than a few seconds. Long-running schema changes were broken into smaller steps or deferred to background jobs.
- Migrations were tested in staging with production-like data volume before they ran in production.
We used Entity Framework Core migrations in the .NET services, and the pipeline applied them automatically during deployment. If a migration failed, the deployment failed, and the pipeline would not promote the new code. This made migrations visible to the CI/CD system instead of invisible state changes that happened outside the pipeline.
What Actually Made CI/CD Useful
Useful CI/CD is not about the sophistication of the pipeline. It is about eliminating the manual steps that the pipeline cannot see. Every manual step is a potential failure point and a reason the pipeline cannot be trusted. If the deployment process requires a checklist, the checklist should be encoded in the pipeline, not stored in a wiki.
The changes that mattered most were organizational, not technical:
- We moved all configuration into the repository and applied it through the pipeline.
- We automated database migrations and enforced rules that made them safe to run without human intervention.
- We deployed smaller changes more frequently, which reduced the risk of each deployment and forced us to confront the hidden manual steps.
- We made rollbacks as easy as deployments by treating each deployment as an immutable artifact with a manifest.
The technical implementation mattered, but the real work was changing how the team thought about deployments. CI/CD is not a tool you install. It is a discipline that requires eliminating the manual steps, the hidden state changes, and the tribal knowledge that make deployments fragile. If your pipeline turns green but your deployments still fail, the problem is not the pipeline. The problem is the manual steps the pipeline does not know about.
Comments
Post a Comment