Skip to main content

Claude Prompt to Build a Chaos Engineering Test Plan

Design chaos engineering experiments: inject network, latency, and resource failures, then measure recovery and build production confidence.

Fill in the placeholders

Edit the values, then copy your finished prompt.

Your Prompt
prompt.txt
Design a chaos engineering test plan for a microservices e-commerce platform running on Kubernetes on AWS EKS. Service dependencies: PostgreSQL, Redis, Elasticsearch, Stripe API, SendGrid. Experiments: 1) Network partition between order-service and payment-service — verify graceful degradation, 2) Database connection exhaustion — test connection pool recovery, 3) Downstream API latency injection (3000ms) — verify timeout handling, 4) Pod/instance kill — test auto-recovery and load redistribution, 5) Disk space exhaustion — verify alerts and graceful handling, 6) DNS failure — test fallback resolution, 7) Memory pressure — gradual increase to 90% of container limit to test OOM behavior. For each experiment: hypothesis, blast radius control, abort conditions, expected vs actual behavior, rollback procedure. Use Litmus Chaos. Schedule: start in staging before graduating to production. Include runbook template and incident response integration.

What this prompt does

This prompt turns a generic "break things on purpose" idea into a structured, safe chaos engineering test plan for a specific [system_type] running on [infrastructure] with declared [dependencies]. It walks the AI through seven concrete failure experiments — network partition between [partition_services], database connection exhaustion, downstream latency injection of [latency_ms]ms, instance kills, disk exhaustion, DNS failure, and memory pressure up to [memory_limit] — and forces a per-experiment hypothesis, blast-radius control, abort conditions, and rollback procedure.

The structure works because it mirrors how real chaos engineering is run: scientifically, not recklessly. By naming [chaos_tool] and starting in [environment] before graduating to production, the output stays grounded in your actual stack rather than producing abstract theory. The hypothesis-and-abort framing is the safety rail — it converts "let's kill a pod" into a controlled experiment with a known expected outcome and a clear stop button.

When to use it

  • Before a big launch, when you need evidence that the system degrades gracefully under failure
  • When migrating critical services to Kubernetes and want to validate auto-recovery
  • After a production incident, to reproduce the failure mode in a controlled experiment
  • When establishing a recurring game-day practice for an on-call team
  • To pressure-test timeout and retry handling for a flaky downstream [dependencies] API
  • When you need to justify resilience work to stakeholders with concrete experiment results

Example output

You get a structured test plan: a numbered list of experiments, each with its own hypothesis, blast-radius scope, abort conditions, expected-versus-actual behavior table, and rollback steps. The network-partition experiment between [partition_services], for instance, comes with its own degradation hypothesis and a stop condition keyed to error-rate thresholds, while the memory-pressure experiment ramps gradually toward [memory_limit] so the OOM behavior is observed rather than triggered abruptly. It closes with a runbook template and notes on integrating findings into your incident-response process. Overall it reads closer to an executable game-day document, scoped to [infrastructure] and your [chaos_tool], than a one-paragraph summary you would still have to flesh out.

Pro tips

  • Be specific with [dependencies] and [partition_services] — naming the real services (e.g. order-service and payment-service) makes the blast-radius analysis far more useful than generic placeholders
  • Always start [environment] at staging; I never let the AI default to production for the first run of a new experiment
  • Tune [latency_ms] to just above your configured timeout so you actually exercise the timeout path rather than passing it cleanly
  • Set [memory_limit] as a percentage of the container limit so the OOM behavior reflects your real resource caps
  • Pick a [chaos_tool] you can actually operate; the plan's abort conditions are only as good as the tooling that enforces them
  • Treat the generated hypotheses as a draft — refine each one until failing it would teach you something specific about recovery

Frequently Asked Questions

Is this prompt safe to run against production directly?
The plan deliberately starts in your chosen `[environment]`, typically staging, before graduating to production. Each experiment includes abort conditions and a rollback procedure, but you should run it in lower environments first and only promote experiments you have validated.
Does it work with chaos tools other than the default?
Yes. The `[chaos_tool]` variable lets you specify Litmus Chaos, Chaos Mesh, Gremlin, or any tool you use. The experiment design stays the same; only the implementation details for injecting failures change to match your tooling.
Will it generate the actual chaos manifests or just the plan?
It focuses on the test plan: hypotheses, blast radius, abort conditions, and rollback steps. It can describe how to configure experiments in your `[chaos_tool]`, but you should treat any generated config as a starting draft and validate it carefully before running.
How do I make the experiments specific to my system?
Fill in `[system_type]`, `[infrastructure]`, and `[dependencies]` with your real stack, and name concrete services in `[partition_services]`. The more accurate these are, the more the blast-radius and degradation analysis reflects your actual failure surface.
Engr Mejba Ahmed

Need this built for real?

Engr Mejba Ahmed

AI Developer · Software Engineer

I'm Mejba — I design and ship production AI systems, automations, and full-stack apps. If you want this turned into a working solution for your team, let's talk.

More in Testing & QA Automation Prompts

Engr Mejba Ahmed

Engr Mejba Ahmed

AI assistant · trained on my work

👋

Hey there!

Quick Actions

WhatsApp Direct line to me

Chat on WhatsApp

+880 1723 741224 · Replies within the hour on working days

Popular Questions

Engr Mejba Ahmed is connected
Engr Mejba Ahmed is typing...
Engr Mejba Ahmed avatar

✉ Want me to follow up? Drop your email

Engr Mejba Ahmed avatar

📞 Connect Directly

Choose how you'd like to reach me

WhatsApp

+880 1723 741224

Email

mejba.13@gmail.com

✓ Details sent! I'll get back to you shortly.

Powered by OpenAI

335+

Blog Posts

25

AI Courses

63

Projects

Services & Expertise

Pricing & Process

Learning & Resources

Connect & Support