Skip to main content

Claude/ChatGPT Prompt to Write an On-Call Runbook for a Critical Service

Write an actionable on-call runbook an engineer paged at 3am can follow: SLOs, alerts, triage, mitigations, escalation, rollback.

Fill in the placeholders

Edit the values, then copy your finished prompt.

Your Prompt
prompt.txt

                                

What this prompt does

This prompt writes an on-call runbook built for a tired human at 3am — short bullets and exact commands, not prose. It casts the assistant as a senior DevOps engineer and supplies four context variables: [service_name], [platform], [datastore], and [tooling]. The deliverables are a health summary with the service's SLOs, common alerts and what each actually means, first-five-minutes triage in order, the top five incident scenarios each with step-by-step mitigation commands, the escalation path with who to page and when, and the rollback procedure with links to the dashboards and logs that matter.

The structure works because a runbook written for a calm afternoon is useless during the incident it was meant for. By demanding copy-paste commands inline, the prompt removes the worst failure mode: someone half-awake reconstructing a kubectl or psql invocation from memory under pressure. Naming [tooling] (PagerDuty, Grafana, CloudWatch) makes the alerts and escalation steps concrete to your stack, and the first-five-minutes section imposes an ordered triage so responders don't freeze or flail.

When to use it

  • The day a critical service goes live, before its first real page
  • When an existing runbook is vague prose instead of actionable commands
  • When on-call engineers keep paging seniors for steps that should be documented
  • When you want a consistent triage order across the whole on-call rotation
  • After an incident, to fold the lessons into a refined runbook
  • When onboarding new on-call staff to a service they don't know deeply

Example output

You get a scannable runbook: a health-and-SLO summary, an alerts table explaining what each alert means, an ordered first-five-minutes triage checklist, five incident scenarios each with inline mitigation commands, an escalation path naming who to page and when, and a rollback procedure with dashboard and log links in place. The format favors short bullets and copy-paste commands over explanation, so a responder can act without reading paragraphs. The ordering matters as much as the content: the triage checklist is sequenced so the first checks rule out the most common causes, and the escalation path is explicit about the threshold at which you stop debugging alone and pull in a second person.

Pro tips

  • Set [tooling] to your real alerting stack so the alert meanings and escalation steps match what responders actually see
  • Name [datastore] precisely so mitigation commands target the right database and cache
  • Insist every mitigation has exact commands inline; a runbook that says "restart the service" without the command fails at 3am
  • Draft it the day the service launches, then refine after each real page rather than waiting for the perfect version
  • Replace placeholder dashboard links with real URLs before the runbook goes into rotation
  • Have someone unfamiliar with the service dry-run the triage steps to catch missing context

Frequently Asked Questions

Why does the prompt insist on commands instead of explanations?
Because the runbook is meant to be used by a tired responder mid-incident, not read leisurely. Inline copy-paste commands remove the worst failure mode, where someone half-awake reconstructs a complex invocation from memory and makes it worse under pressure.
How many incident scenarios does it cover?
It covers the top five incident scenarios, each with step-by-step mitigation commands. Five is a deliberate scope that captures the most common failures without producing a document too long to scan during an actual page. Add service-specific scenarios as you learn them.
Will the dashboard links be real?
The prompt asks for dashboard and log links in place, but the model can only insert placeholders or derivable URLs from the context you give it. Replace any placeholders with your real Grafana or CloudWatch links before the runbook goes into the on-call rotation.
Should I treat the first draft as final?
No. Draft it when the service launches, then refine it after each real page. A runbook written before any incident inevitably misses edge cases, so the post-incident refinement loop is where it becomes genuinely reliable.
Engr Mejba Ahmed

Need this built for real?

Engr Mejba Ahmed

AI Developer · Software Engineer

I'm Mejba — I design and ship production AI systems, automations, and full-stack apps. If you want this turned into a working solution for your team, let's talk.

More in DevOps & Cloud Prompts

Engr Mejba Ahmed

Engr Mejba Ahmed

Claude Code Expert · Online

👋

Hey there!

Quick Actions

WhatsApp Instant reply

Chat on WhatsApp

+880 1723 741224 · Instant reply

Popular Questions

Engr Mejba Ahmed is connected
Engr Mejba Ahmed is typing...
Engr Mejba Ahmed avatar

✉ Want me to follow up? Drop your email

Engr Mejba Ahmed avatar

📞 Connect Directly

Choose how you'd like to reach me

WhatsApp

+880 1723 741224

Email

[email protected]

✓ Details sent! I'll get back to you shortly.

Powered by OpenAI

335+

Blog Posts

25

AI Courses

63

Projects

Services & Expertise

Pricing & Process

Learning & Resources

Connect & Support