What this prompt does
This prompt makes the AI design an error-handling framework for ETL pipelines, treating failure as the normal case rather than an afterthought. You specify [pipeline_type], the [data_volume] processed daily, and the [tech_stack], and the model builds a classification system that separates transient errors (timeouts, rate limits) from permanent ones (schema mismatch, invalid data) and partial failures, each with its own handling strategy.
The structure works because it forces the unhappy paths to be designed before the happy path. It specifies a retry mechanism with exponential backoff ([initial_delay], [max_retries]) and a circuit breaker that opens after [circuit_threshold] failures and resets after [circuit_reset] minutes, a dead letter queue in [dlq_storage] with full error context, quarantine tables with reason codes, tiered alerting to [alert_channels] keyed off [dlq_threshold], a replay/recovery workflow with idempotency, and observability metrics to [monitoring_tool]. Classifying transient versus permanent up front is what stops a pipeline from retrying itself into an outage.
When to use it
- You're building a batch pipeline that will inevitably hit bad data and want resilient paths designed in.
- You need to stop retrying permanent errors while still backing off on transient ones.
- You want failed records captured in a DLQ with enough context to debug, not silently dropped.
- You need a quarantine path so a few bad rows don't block the whole load.
- You want tiered alerting that distinguishes a routine retry from a circuit-breaker-open emergency.
- You need a tested replay workflow to reprocess DLQ records after fixing the root cause.
Example output
Expect a framework design: an error taxonomy with handling rules per class, retry-with-backoff and circuit-breaker logic with your thresholds, a DLQ schema in [dlq_storage] capturing payload, error message, timestamp, retry count, and stage, quarantine table definitions with reason codes, an alerting tier table (INFO/WARNING/CRITICAL) mapped to [alert_channels], an idempotent replay script, and the metrics exported to [monitoring_tool].
Pro tips
- Get the transient-versus-permanent split right — retrying a schema mismatch wastes cycles and can amplify an outage, while not retrying a timeout drops recoverable data.
- Tune
[circuit_threshold]and[circuit_reset]so the breaker trips before a failing dependency causes a retry storm, but not so eagerly that one blip halts the pipeline. - Capture full context in the
[dlq_storage]records (original payload, stage, retry count); a DLQ without the payload is hard to replay. - Make the replay workflow genuinely idempotent — reprocessing DLQ records must not double-load rows that partially succeeded the first time.
- Set
[dlq_threshold]based on normal noise; alerting on a single DLQ entry will train the team to ignore the alerts. - Wire the metrics (error rate, retry count, DLQ depth, latency) to
[monitoring_tool]early so you can see a slow degradation before it becomes a halt.