Skip to main content

The Production Cost Traps Behind VPC Networking, NAT Gateways, and Data Transfer

I spent three months optimizing an AWS data pipeline that processed billions of records through Apache Spark and loaded them into Amazon Redshift. The job ran reliably. Query performance was acceptable. The engineering team had moved on to the next feature. Then the monthly bill arrived, and a single line item stopped the conversation: $4,200 for NAT Gateway data processing charges.

The pipeline itself was efficient. We had tuned the Spark partitions, compressed the intermediate files, and designed the Redshift distribution keys to minimize broadcast joins. But we had placed every Lambda function, every Spark worker, and every API client inside a VPC with private subnets, and we had routed all outbound traffic through a NAT Gateway in each availability zone. The architecture diagram looked secure. The cost structure was unsustainable.

VPC networking decisions feel like infrastructure hygiene when you make them. They become budget traps when you forget that AWS charges for data movement, not just for compute time or storage. Understanding where those charges accumulate and which alternatives actually reduce cost without sacrificing isolation requires looking past the security theater and into the specific traffic patterns your workload creates.

Where NAT Gateway Costs Accumulate

AWS charges for NAT Gateway usage in two dimensions: an hourly fee per gateway (about $0.045 per hour, or roughly $32 per month) and a data processing fee of $0.045 per gigabyte that passes through it. For a pipeline processing billions of records, the hourly fee is negligible. The per-gigabyte charge becomes the dominant cost when your workload generates sustained outbound traffic.

In our case, the pipeline pulled data from external REST APIs hosted outside AWS, transformed it with Spark, and wrote the results to Redshift. Every API call, every dependency download during Lambda cold starts, every log stream sent to a third-party observability service, and every SDK call to an AWS service that did not support VPC endpoints passed through the NAT Gateway. We were paying $0.045 per gigabyte to reach the public internet, and we were moving terabytes every month.

The cost breakdown looked like this:

  • Lambda cold starts downloading .NET runtime dependencies and NuGet packages from public repositories: approximately 800 GB per month across all invocations
  • API polling from external SaaS providers with paginated JSON responses: approximately 1,200 GB per month
  • Spark workers pulling compressed CSV files from an S3 bucket through an implicit route to the NAT Gateway before we configured the S3 VPC endpoint: approximately 2,400 GB per month
  • CloudWatch Logs and metrics sent to a third-party aggregation service: approximately 300 GB per month

Each of these traffic patterns had an alternative that avoided the NAT Gateway entirely, but the default VPC configuration made the expensive path invisible until the bill arrived.

VPC Endpoints and the Services That Support Them

AWS offers VPC endpoints for many of its own services, allowing resources in a private subnet to reach S3, DynamoDB, Secrets Manager, and others without routing through a NAT Gateway or an internet gateway. These endpoints come in two types: gateway endpoints (for S3 and DynamoDB, with no hourly charge) and interface endpoints (for most other services, billed at approximately $0.01 per hour per availability zone plus $0.01 per gigabyte of data processed).

We had configured the VPC with private subnets but had forgotten to add the S3 gateway endpoint. Spark workers reading source files from S3 were routing through the NAT Gateway, paying $0.045 per gigabyte instead of zero. Adding the gateway endpoint required one Terraform resource and a route table association:

resource "aws_vpc_endpoint" "s3" {
  vpc_id       = aws_vpc.main.id
  service_name = "com.amazonaws.us-east-1.s3"
  route_table_ids = [
    aws_route_table.private.id
  ]
}

This single change eliminated 2,400 GB of NAT Gateway traffic per month, saving approximately $108. The S3 gateway endpoint has no hourly charge and no per-gigabyte fee for traffic between S3 and resources in the same region.

For Secrets Manager, which the Lambda functions called to retrieve API credentials at startup, we added an interface endpoint. The hourly cost was $7.20 per month (one endpoint in one availability zone), and the data processing fee was negligible because secret retrieval involves kilobytes, not gigabytes. The alternative was continuing to route secret fetches through the NAT Gateway, which cost more and added latency.

When Private Subnets Are the Wrong Default

The instinct to place everything in a private subnet comes from enterprise network design, where the assumption is that direct internet access is a vulnerability. In AWS, that assumption often leads to paying for isolation you do not need. If a Lambda function only calls AWS services that support VPC endpoints, it does not need to be in a VPC at all. If it does need to reach the public internet, placing it in a public subnet with a restrictive security group is often cheaper and simpler than maintaining a NAT Gateway.

We moved the API polling Lambda functions out of the VPC entirely. They did not access any resources that required private network connectivity. They called external HTTPS endpoints, wrote results to S3, and logged to CloudWatch. Running them outside the VPC eliminated the NAT Gateway data processing charges for that traffic (1,200 GB per month, or $54 in savings) and removed the cold-start latency caused by the elastic network interface attachment that VPC-connected Lambda functions require.

The security model did not degrade. The functions still used IAM roles with least-privilege policies, still validated TLS certificates on outbound calls, and still wrote to S3 buckets with encryption at rest and restrictive bucket policies. The private subnet had provided the appearance of isolation, not the substance.

Data Transfer Charges and Cross-Region Pitfalls

NAT Gateway costs are one category of AWS networking charges. Data transfer between regions, between availability zones, and out to the internet creates additional line items. For our pipeline, the Redshift cluster and the S3 buckets were in the same region and the same availability zone, so data transfer charges were minimal. But during an earlier prototype, we had accidentally placed the Redshift cluster in us-east-1 and the S3 bucket in us-west-2. The cross-region data transfer fee is $0.02 per gigabyte in each direction, and loading a few terabytes would have cost hundreds of dollars per job.

The failure mode is easy to create. Terraform modules default to the provider region, but if you configure multiple providers or inherit a module from another team, the region can differ from your expectation. The fix is to make region and availability zone explicit in every resource configuration and to add a validation check in your infrastructure tests:

# Simplified example: verify S3 bucket and Redshift cluster regions match
def test_data_pipeline_region_alignment():
    s3_region = get_bucket_region("source-data-bucket")
    redshift_region = get_cluster_region("analytics-cluster")
    assert s3_region == redshift_region, \
        f"Region mismatch: S3 in {s3_region}, Redshift in {redshift_region}"

This kind of check surfaces the mistake during deployment, not after the bill arrives.

Observability and the Cost of Logging to Third Parties

We routed CloudWatch Logs to a third-party observability platform using a subscription filter and a Lambda forwarder. The forwarder ran inside the VPC and sent logs over HTTPS to the vendor's API. Every gigabyte of logs passed through the NAT Gateway at $0.045 per gigabyte, then left AWS to the internet at $0.09 per gigabyte (the standard data transfer out rate for the first 10 TB per month). The combined cost was $0.135 per gigabyte.

The pipeline generated approximately 300 GB of logs per month. Sending them to the third party cost $40.50 in NAT Gateway and data transfer fees, on top of the vendor's ingestion fee. Moving the forwarder Lambda out of the VPC eliminated the NAT Gateway charge, reducing the cost to $27 per month (just the data transfer out fee). That change required updating the Lambda function's IAM role to allow CloudWatch Logs access and removing the VPC configuration block from the Terraform resource.

The deeper question is whether you need to send all logs to a third party. We reduced log volume by 60 percent by filtering out debug-level entries and by aggregating repetitive messages at the application layer before they reached CloudWatch. The cost savings compounded: fewer gigabytes written to CloudWatch, fewer gigabytes forwarded to the vendor, and a smaller vendor bill.

A Framework for VPC and NAT Gateway Decisions

Start with the assumption that Lambda functions and containers do not need to be in a VPC unless they access a resource that requires private network connectivity, such as an RDS database, an ElastiCache cluster, or an on-premises system over VPN or Direct Connect. If the workload only calls AWS APIs or public HTTPS endpoints, run it outside the VPC.

When you do need a VPC, add gateway endpoints for S3 and DynamoDB immediately. They have no cost and eliminate a common source of NAT Gateway traffic. For other AWS services, evaluate whether the interface endpoint cost (roughly $7 to $22 per month depending on availability zone coverage) is lower than the NAT Gateway data processing cost for the traffic volume you expect.

Before adding a NAT Gateway, list every outbound destination and estimate the monthly data volume. If the total is more than a few hundred gigabytes, look for alternatives: VPC endpoints, moving workloads out of the VPC, using S3 Transfer Acceleration for large uploads, or reducing log verbosity. If you need a NAT Gateway, deploy one per availability zone only if your availability requirements justify the cost. A single NAT Gateway in one availability zone costs $32 per month; three gateways for high availability cost $96 per month plus tripled data processing fees if traffic is not balanced.

Track VPC networking costs as a separate budget line item. AWS Cost Explorer allows you to filter by usage type, including NAT Gateway hours, NAT Gateway data processing, and data transfer. Set a budget alert at a threshold that reflects your expected traffic. When the alert fires, investigate before the variance becomes a sustained cost.

VPC networking is not inherently expensive, but the default configurations and the hidden per-gigabyte charges make it easy to build a system that works correctly and costs more than it should. The fix is not to avoid VPCs entirely, but to make networking decisions based on the specific traffic patterns and access requirements of each workload, and to verify the cost implications before the architecture becomes load-bearing.

Comments