External APIs break in predictable ways. They rate-limit your requests during peak hours. They change field names without warning. They return HTTP 200 with an error object nested three levels deep in JSON. They timeout after exactly 29 seconds, one second before your own timeout fires. If your data connector treats these as exceptional cases, your pipeline stops working the day your business depends on it.
I have built data connectors for regulated environments where the external API was a legacy SOAP service, a partner REST endpoint with unpredictable availability, and a SaaS vendor that changed response schemas without versioning. The lesson from all three is the same: resilience is not a retry loop wrapped around an HTTP client. It is a design decision about failure ownership, state management, and operational visibility.
Failure Modes You Must Design For, Not Catch Later
Most data connector failures fall into three categories: transient errors you can retry, permanent errors you must route elsewhere, and silent errors you will not notice until the data is missing. A transient error is a 429 rate limit or a 503 service unavailable. A permanent error is a 401 authentication failure or a 404 when the resource has been deleted upstream. A silent error is a 200 status code with an empty result set when you expected records, or a schema change that makes your deserializer ignore new fields.
The mistake I see most often is treating all HTTP errors as transient and retrying them with exponential backoff. That works for 503, but it wastes time on 401 and hides the real problem. If your API key has expired, retrying for five minutes before raising an alert just delays the fix. If the upstream service has removed an endpoint, your retry logic will never succeed. You need to distinguish between errors that warrant a retry and errors that warrant a dead-letter queue or an immediate page.
Schema drift is harder to detect because the API returns success. The upstream team adds a new required field, renames an existing field, or changes a string to an integer. Your deserializer silently ignores the new field or throws a parse exception that looks like a transient error. I have seen connectors lose days of data because the error was logged but not routed to a dead-letter queue, and the team assumed the API was healthy because the HTTP status was 200.
Idempotency and State: Why Retries Alone Are Not Enough
Retries require idempotency. If your connector fetches a page of records, writes them to a database, then crashes before updating the cursor, the next run will fetch the same page again. If your insert logic does not handle duplicates, you will have duplicate records. If your logic does handle duplicates but the API does not guarantee stable ordering, you might miss records that appeared between the first and second fetch.
I handle this with a state table that tracks the last successful cursor, timestamp, or page token before writing any records. The connector reads the state, fetches the next batch, writes the records in a transaction, and updates the state in the same transaction. If the process crashes, the next run starts from the last committed state. If the API returns duplicates, the database constraint or upsert logic deduplicates them. This pattern works for REST APIs with pagination tokens, SOAP services with timestamp filters, and webhook deliveries with sequence numbers.
public async Task<int> FetchAndStoreRecords(string connectorId, CancellationToken cancellationToken)
{
var state = await _stateRepository.GetLastCursor(connectorId);
var request = new ApiRequest { Cursor = state.Cursor, PageSize = 100 };
var response = await _apiClient.GetRecords(request, cancellationToken);
using var transaction = await _dbContext.Database.BeginTransactionAsync(cancellationToken);
foreach (var record in response.Records)
{
await _recordRepository.Upsert(record);
}
await _stateRepository.UpdateCursor(connectorId, response.NextCursor);
await transaction.CommitAsync(cancellationToken);
return response.Records.Count;
}
The transaction ensures that records and state are committed together. If the API call succeeds but the database write fails, the next run will retry the same batch. If the API call fails, the state is unchanged and the next run starts from the same cursor. The upsert logic in the record repository handles duplicates using a unique constraint on the upstream record ID.
Rate Limiting and Throttling: Backoff That Respects API Contracts
Rate limits are not errors. They are a contract. If the API returns HTTP 429 with a Retry-After header, your connector should sleep for that duration and retry the request. If the API does not provide a Retry-After header, you need exponential backoff with jitter to avoid a thundering herd when multiple connectors retry at the same time.
I have worked with APIs that enforce rate limits per API key, per IP address, and per tenant. The failure mode depends on the limit. A per-key limit means one connector can exhaust the quota for all tenants. A per-IP limit means scaling horizontally with more workers does not help. A per-tenant limit means you need separate credentials or request queues for each tenant. You cannot design a resilient connector without understanding the rate limit scope.
For APIs with tight rate limits, I use a token bucket or leaky bucket algorithm to spread requests over time. The connector calculates the maximum request rate from the API contract, then schedules fetches to stay below that rate. If the API returns 429 anyway, the connector respects the Retry-After header and adjusts the bucket refill rate. This prevents the connector from wasting retries and reduces the risk of getting blocked.
Dead-Letter Queues and Observability: Failures You Can Act On
When a record cannot be processed after all retries, it must go somewhere you can inspect it. A dead-letter queue is a durable store for failed records with enough context to diagnose the failure: the original request, the response, the error message, the retry count, and the timestamp. Without this, you are guessing why records are missing.
I route permanent errors and exhausted retries to a dead-letter queue implemented as a database table or an SQS queue. Each entry includes the connector ID, the upstream record ID, the request payload, the last response, and the error type. A separate process monitors the queue and raises alerts when the failure count exceeds a threshold. For schema errors, the dead-letter queue is the only way to identify which records were skipped and why.
Observability starts with structured logs and metrics. Every API call logs the endpoint, the response status, the duration, and the record count. Every failure logs the error type, the retry count, and whether the record was routed to the dead-letter queue. Metrics track the request rate, the error rate, the retry rate, and the queue depth. These signals tell you whether the connector is healthy, whether the API is degraded, and whether the dead-letter queue needs attention.
Design Checklist for Production Data Connectors
Based on the connectors I have built and maintained, these are the design decisions that matter most:
- Distinguish transient, permanent, and silent errors. Retry transient errors with exponential backoff and jitter. Route permanent errors to a dead-letter queue. Detect silent errors with record count checks and schema validation.
- Make every fetch idempotent. Store state before writing records. Use transactions to commit state and records together. Handle duplicates with upsert logic or unique constraints.
- Respect rate limits as contracts, not obstacles. Read the
Retry-Afterheader. Use token buckets to spread requests over time. Understand whether the limit is per key, per IP, or per tenant. - Route failures to a dead-letter queue with diagnostic context. Store the request, response, error, and retry count. Monitor the queue depth and alert when failures accumulate.
- Instrument every API call with structured logs and metrics. Track request rate, error rate, retry rate, and latency. Use these signals to detect degradation before it becomes an outage.
- Test failure modes, not just success paths. Simulate 429, 503, timeouts, and schema changes. Verify that state is preserved, retries are bounded, and failures are routed correctly.
When the API Is the Constraint
You cannot make an unreliable API reliable, but you can design a connector that fails gracefully and recovers automatically. The difference is whether your pipeline stops working the first time the API changes, or whether it logs the failure, routes the bad records to a queue, and keeps processing the good ones.
The connectors that survive production are the ones that treat external APIs as unreliable dependencies, not trusted services. They assume rate limits will be hit, schemas will drift, and timeouts will happen. They handle these as expected failure modes with well-defined recovery paths, not exceptions that crash the process. If your connector does not have a dead-letter queue, idempotent state management, and structured retry logic, it is not ready for production.
Comments
Post a Comment