When I review cloud architecture in modernization projects, I often find IAM policies and secret-handling patterns treated as post-deployment security tasks rather than foundational design decisions. Teams grant broad permissions during development, defer secret rotation until a compliance audit surfaces the risk, and treat least-privilege access as a hardening step that happens after the application works. This approach consistently produces architectures that are expensive to secure later and difficult to operate safely at scale. I learned this lesson the hard way. In one AWS data pipeline, early prototypes used a single service role with wide S3, Redshift, and Lambda permissions. The role worked across environments, simplified local testing, and allowed rapid iteration. When the pipeline moved to production and began processing regulated data, tightening those permissions required rewriting deployment scripts, breaking existing integrations, and coordinating changes across multip...
External APIs break in predictable ways. They rate-limit your requests during peak hours. They change field names without warning. They return HTTP 200 with an error object nested three levels deep in JSON. They timeout after exactly 29 seconds, one second before your own timeout fires. If your data connector treats these as exceptional cases, your pipeline stops working the day your business depends on it. I have built data connectors for regulated environments where the external API was a legacy SOAP service, a partner REST endpoint with unpredictable availability, and a SaaS vendor that changed response schemas without versioning. The lesson from all three is the same: resilience is not a retry loop wrapped around an HTTP client. It is a design decision about failure ownership, state management, and operational visibility. Failure Modes You Must Design For, Not Catch Later Most data connector failures fall into three categories: transient errors you can retry, permanent error...