What this prompt does
This prompt turns a generic "break things on purpose" idea into a structured, safe chaos engineering test plan for a specific [system_type] running on [infrastructure] with declared [dependencies]. It walks the AI through seven concrete failure experiments — network partition between [partition_services], database connection exhaustion, downstream latency injection of [latency_ms]ms, instance kills, disk exhaustion, DNS failure, and memory pressure up to [memory_limit] — and forces a per-experiment hypothesis, blast-radius control, abort conditions, and rollback procedure.
The structure works because it mirrors how real chaos engineering is run: scientifically, not recklessly. By naming [chaos_tool] and starting in [environment] before graduating to production, the output stays grounded in your actual stack rather than producing abstract theory. The hypothesis-and-abort framing is the safety rail — it converts "let's kill a pod" into a controlled experiment with a known expected outcome and a clear stop button.
When to use it
- Before a big launch, when you need evidence that the system degrades gracefully under failure
- When migrating critical services to Kubernetes and want to validate auto-recovery
- After a production incident, to reproduce the failure mode in a controlled experiment
- When establishing a recurring game-day practice for an on-call team
- To pressure-test timeout and retry handling for a flaky downstream
[dependencies]API - When you need to justify resilience work to stakeholders with concrete experiment results
Example output
You get a structured test plan: a numbered list of experiments, each with its own hypothesis, blast-radius scope, abort conditions, expected-versus-actual behavior table, and rollback steps. The network-partition experiment between [partition_services], for instance, comes with its own degradation hypothesis and a stop condition keyed to error-rate thresholds, while the memory-pressure experiment ramps gradually toward [memory_limit] so the OOM behavior is observed rather than triggered abruptly. It closes with a runbook template and notes on integrating findings into your incident-response process. Overall it reads closer to an executable game-day document, scoped to [infrastructure] and your [chaos_tool], than a one-paragraph summary you would still have to flesh out.
Pro tips
- Be specific with
[dependencies]and[partition_services]— naming the real services (e.g. order-service and payment-service) makes the blast-radius analysis far more useful than generic placeholders - Always start
[environment]at staging; I never let the AI default to production for the first run of a new experiment - Tune
[latency_ms]to just above your configured timeout so you actually exercise the timeout path rather than passing it cleanly - Set
[memory_limit]as a percentage of the container limit so the OOM behavior reflects your real resource caps - Pick a
[chaos_tool]you can actually operate; the plan's abort conditions are only as good as the tooling that enforces them - Treat the generated hypotheses as a draft — refine each one until failing it would teach you something specific about recovery