Skip to main content

Incident Post-Mortem Report Writer

Write blameless incident post-mortems with a clear timeline, 5-Whys root cause, impact assessment, and a short list of action items that actually get done.

Fill in the placeholders

Edit the values, then copy your finished prompt.

Your Prompt
prompt.txt
You are an SRE lead who writes exemplary incident post-mortems. Generate a blameless post-mortem report for the following incident.

**Incident Summary:**
Production database connection pool exhaustion caused 502 errors for 45 minutes during peak traffic. Approximately 30% of API requests failed.

**Timeline (what we know):**
14:30 - Alerts fire for elevated 502 rates. 14:35 - On-call engineer acknowledges. 14:45 - Identified DB connection pool at 100%. 14:50 - Attempted restart of PHP-FPM. 15:00 - Found long-running query from new report feature. 15:05 - Killed query, pool recovered. 15:15 - Error rates back to normal.

**Impact:**
Approximately 12,000 users experienced errors. Estimated 450 failed checkout attempts. No data loss.

**Environment:** AWS: 3x EC2 instances, RDS MySQL (db.r5.xlarge), ElastiCache Redis, behind ALB

**Generate This Post-Mortem Report:**

---
## Incident Post-Mortem: <Title>
**Date:** [Incident Date]
**Severity:** [P1/P2/P3/P4]
**Duration:** [Total duration]
**Author:** <Author>
**Status:** Draft — pending review

---

### Executive Summary
3 sentences: what happened, how long it lasted, what was the impact. Written for non-technical leadership.

### Impact Assessment
- **Users affected:** [number and percentage]
- **Revenue impact:** [if applicable]
- **Data integrity:** Was any data lost or corrupted?
- **SLA impact:** Did we breach any SLA commitments?
- **Customer communication:** What did we tell customers and when?

### Timeline
| Time (UTC) | Event | Actor |
|------------|-------|-------|
| HH:MM | <Event> | <Person/System> |

Include: detection time, first response, escalations, mitigation attempts (including failed ones), resolution, and all-clear.

### Root Cause Analysis
Use the **5 Whys** method:
1. Why did <symptom> happen? → Because [cause 1]
2. Why did [cause 1] happen? → Because [cause 2]
3. Why did [cause 2] happen? → Because [cause 3]
4. Why did [cause 3] happen? → Because [cause 4]
5. Why did [cause 4] happen? → Because [root cause]

**Contributing Factors:**
- Technical factors (what broke)
- Process factors (what process gaps allowed this)
- Detection factors (why did it take N minutes to detect)

### What Went Well
- Things that worked during incident response
- Systems or processes that limited the blast radius
- People who went above and beyond

### What Went Wrong
- Detection: [how long until we knew]
- Response: [gaps in the response process]
- Communication: [internal/external communication issues]
- Tooling: [what tools were missing or inadequate]
- The new report feature was not load-tested before deployment

### Action Items
**CRITICAL — These prevent recurrence:**

| # | Action | Owner | Priority | Due Date | Status |
|---|--------|-------|----------|----------|--------|
| 1 | <Specific, measurable action> | <Owner> | P1 | <Date> | Open |

Rules for action items:
- Each must be specific and measurable (not "improve monitoring")
- Each must have a single owner (not a team)
- Include both immediate fixes AND systemic improvements
- Limit to 8 items — too many means none get done
- Categorize: Detection | Prevention | Response | Process

### Lessons Learned
- What would we do differently next time?
- What assumption was wrong?
- What documentation needs updating?

---
**Blameless Culture Note:** This post-mortem focuses on systems and processes, not individuals. People made the best decisions they could with the information available.

What this prompt does

This prompt casts the model as an SRE lead and turns raw incident notes into a structured, blameless post-mortem. You feed it four inputs — [incident_summary], [timeline_data], [impact_description] and [environment] — and it returns a full report: executive summary, impact assessment, a UTC timeline table, a 5-Whys root cause analysis, what went well, what went wrong, an action-item table, and lessons learned.

The structure is what makes it useful. By forcing the 5 Whys, the prompt pushes past the first symptom toward a real root cause instead of stopping at "the database fell over." The [timeline_data] you paste becomes the backbone of the timeline table — detection, escalation, failed mitigations, and resolution all get rows — so the narrative stays factual. The [max_actions] cap is deliberate: it limits action items so they actually get finished rather than rotting in a backlog.

When to use it

  • Right after a production incident, while the timeline is still fresh in everyone's memory.
  • When you need a leadership-readable summary and an engineering deep-dive in the same document.
  • To enforce blameless culture on a team that tends to point fingers.
  • When postmortems on your team keep producing vague action items like "improve monitoring."
  • To standardize the format across many incidents so they're comparable over time.

Example output

You get a Markdown report with a header block (date, severity, duration, status), a three-sentence executive summary for non-technical readers, a bulleted impact assessment, a Markdown timeline table in UTC, a numbered 5-Whys chain ending in the root cause, "what went well / wrong" lists, and an action-item table with columns for owner, priority, due date, status, and category. It closes with a lessons-learned section and a blameless-culture note.

Pro tips

  • Put real timestamps in [timeline_data] — the more granular, the better the timeline table. Include failed mitigation attempts, not just the fix that worked.
  • Quantify [impact_description]: users affected, failed transactions, revenue. Numbers make the executive summary land with leadership.
  • Keep [max_actions] low (5-8). I cap it because a post-mortem with twenty action items produces zero completed ones.
  • Describe [environment] precisely (instance types, datastore versions, load balancer) so the root cause analysis reasons about the actual architecture.
  • After the first draft, ask the model to convert each action item into a tracked ticket title — specific, measurable, single-owner — before you paste it into your issue tracker.
  • Treat the 5-Whys chain as a starting point and correct any link that doesn't match what your logs actually show; the model infers causes it can't verify.

Frequently Asked Questions

Does this prompt keep the post-mortem blameless?
Yes. It explicitly frames the report around systems and processes rather than individuals, and ends with a blameless-culture note. You should still review the wording to make sure no phrasing accidentally singles out a person.
What is the 5 Whys method it uses?
It chains five 'why' questions from the visible symptom down to the underlying root cause, so each answer becomes the question for the next step. This pushes past surface explanations, but you must validate each link against your real logs.
How many action items will it generate?
It limits action items to the number you set in the `[max_actions]` variable, defaulting to eight. The cap is intentional: too many action items means none of them actually get finished after the incident.
Will the timeline be accurate?
The timeline reflects whatever you paste into `[timeline_data]`, so accuracy depends on your notes. Give it precise UTC timestamps for detection, escalation, mitigation attempts, and resolution, then verify the generated table against your monitoring.
Can I use this for a minor incident, not just a P1?
Yes. Set the severity in the header to P3 or P4 and keep the inputs short. The same structure works for smaller incidents, though you can trim sections like revenue impact when they don't apply.
Engr Mejba Ahmed

Need this built for real?

Engr Mejba Ahmed

AI Developer · Software Engineer

I'm Mejba — I design and ship production AI systems, automations, and full-stack apps. If you want this turned into a working solution for your team, let's talk.

More in AI Writing for Developers

Engr Mejba Ahmed

Engr Mejba Ahmed

AI assistant · trained on my work

👋

Hey there!

Quick Actions

WhatsApp Direct line to me

Chat on WhatsApp

+880 1723 741224 · Replies within the hour on working days

Popular Questions

Engr Mejba Ahmed is connected
Engr Mejba Ahmed is typing...
Engr Mejba Ahmed avatar

✉ Want me to follow up? Drop your email

Engr Mejba Ahmed avatar

📞 Connect Directly

Choose how you'd like to reach me

WhatsApp

+880 1723 741224

Email

mejba.13@gmail.com

✓ Details sent! I'll get back to you shortly.

Powered by OpenAI

335+

Blog Posts

25

AI Courses

63

Projects

Services & Expertise

Pricing & Process

Learning & Resources

Connect & Support