Skip to main content

Claude Prompt to Design a Monitoring & Alerting Stack

Design an observability stack with Prometheus, Grafana, and PagerDuty: metrics, dashboards, tiered alert rules, and on-call runbooks.

Fill in the placeholders

Edit the values, then copy your finished prompt.

Your Prompt
prompt.txt
You are a Site Reliability Engineer. Design a complete monitoring and alerting stack for a multi-service web application with API backend running on Kubernetes on AWS EKS.

**Application Context:**
- Services: API gateway, auth service, core API, payment service, notification service, PostgreSQL, Redis, RabbitMQ
- Traffic: 10K requests/minute, 500K daily active users
- SLA Target: 99.95% availability (21.9 minutes downtime/month)

**1. Metrics Collection Strategy:**

Define the metrics to collect at each layer:

**Infrastructure Metrics:**
- CPU, memory, disk I/O, network throughput per host/container
- Container restart counts, OOM kills
- Node availability and resource pressure

**Application Metrics (RED Method):**
- **Rate:** requests per second by endpoint and status code
- **Errors:** error rate by type (4xx client, 5xx server, timeout)
- **Duration:** latency percentiles (p50, p95, p99) by endpoint

**Business Metrics:**
- Signup conversion rate, subscription activation rate, API quota usage per customer
- Revenue-impacting events (failed payments, cart abandonment)
- User experience signals (page load time, Core Web Vitals)

**Database Metrics:**
- Query duration percentiles, slow query count
- Connection pool utilization, replication lag
- Table lock wait time, deadlock count

**2. Prometheus Configuration:**
- Scrape configs for each service
- Recording rules for pre-computed aggregations
- Retention and storage sizing based on 10K requests/minute, 500K daily active users
- Service discovery configuration for Kubernetes on AWS EKS

**3. Grafana Dashboards:**
Design 4 dashboards:
- **Overview:** SLA status, error budget remaining, top-level health
- **Service Detail:** per-service deep dive with latency heatmaps
- **Infrastructure:** host/container resource utilization
- **Business:** revenue metrics, user funnel, feature adoption

For each dashboard: panels, queries, thresholds, and variable filters.

**4. Alert Rules (tiered):**

| Severity | Condition | Response Time | Channel |
|----------|-----------|---------------|---------|
| P1 Critical | <define> | 5 min | PagerDuty page |
| P2 High | <define> | 30 min | PagerDuty + Slack |
| P3 Medium | <define> | 4 hours | Slack |
| P4 Low | <define> | Next business day | Email digest |

Include alert rules that AVOID false positives:
- Use multi-window burn rate for SLO-based alerts
- Require sustained conditions (not single-spike triggers)
- Alert on symptoms, not causes

**5. On-Call Runbooks:**
For each P1/P2 alert, generate a runbook:
- What the alert means
- First 3 diagnostic commands to run
- Common root causes and fixes
- Escalation criteria
- SSL certificate expiry monitoring and automatic renewal alerts

What this prompt does

This prompt makes the AI a Site Reliability Engineer that designs a complete monitoring and alerting stack for a [app_type] on [infrastructure]. You provide the [services], the [traffic_volume], and the [sla_target], and it produces five parts: a metrics collection strategy, a Prometheus configuration, four Grafana dashboards, tiered alert rules, and on-call runbooks. The output is an observability blueprint built around the RED method (Rate, Errors, Duration) rather than a vague list of things to graph.

The structure works because it ties alerts to symptoms, not noise. The metrics layer spans infrastructure, application, business ([business_metrics]), and database signals. The alert rules are tiered P1 through P4 with response times and channels, and they explicitly avoid false positives by using multi-window burn-rate rules and requiring sustained conditions instead of single spikes. Every P1 and P2 alert gets a runbook with the first three diagnostic commands and common fixes, plus whatever you add in [additional_concern], so on-call stays sane.

When to use it

  • You are standing up observability for a multi-service app and need a coherent plan
  • Your current alerts are noisy and you want symptom-based, burn-rate alerting instead
  • You need Prometheus scrape configs and recording rules sized to your traffic
  • You want four purpose-built Grafana dashboards, not one cluttered catch-all
  • You need an SLA-aligned alert tiering with the right channels per severity
  • You want runbooks so a paged engineer knows the first commands to run

Example output

You get a metrics catalog organized by layer, Prometheus scrape and recording-rule configs with retention sized to [traffic_volume], four Grafana dashboard specs (Overview, Service Detail, Infrastructure, Business) with panels and queries, a tiered P1-P4 alert table with conditions and channels, and per-alert runbooks listing the first diagnostic commands and common root causes.

Pro tips

  • List every service in [services] precisely so each gets its own scrape config and dashboard panels
  • Set [traffic_volume] realistically; retention and storage sizing depend directly on it
  • Make [sla_target] exact (99.95% with the minutes/month) so burn-rate alert math is meaningful
  • Use [business_metrics] to surface revenue-impacting signals, not just infrastructure health
  • Add the easily-forgotten check via [additional_concern], like SSL certificate expiry alerts
  • Treat the alert thresholds as starting points and tune them against real traffic to kill remaining false positives

Frequently Asked Questions

Does it help avoid noisy, false-positive alerts?
Yes. The alert rules use multi-window burn-rate logic for SLO-based alerts, require sustained conditions rather than single spikes, and alert on symptoms instead of causes. You should still tune the thresholds against your real traffic to eliminate the last false positives.
What monitoring tools does it assume?
It is built around Prometheus for metrics, Grafana for dashboards, and PagerDuty plus Slack for alert routing. It generates scrape configs, recording rules, four dashboard specs, and tiered alert rules for that stack.
How are the alerts prioritized?
Alerts are tiered P1 through P4, each with a defined condition, response time, and channel, from a PagerDuty page for critical issues down to an email digest for low-priority ones, so the right people are woken for the right reasons.
Does it include runbooks for on-call?
Yes. Every P1 and P2 alert gets a runbook explaining what the alert means, the first three diagnostic commands to run, common root causes and fixes, and escalation criteria, so a paged engineer is not starting from zero.
Can it track business metrics, not just infrastructure?
Yes. The metrics strategy includes a business layer driven by the `[business_metrics]` variable, covering revenue-impacting events like failed payments and cart abandonment alongside the infrastructure and application signals.
Engr Mejba Ahmed

Need this built for real?

Engr Mejba Ahmed

AI Developer · Software Engineer

I'm Mejba — I design and ship production AI systems, automations, and full-stack apps. If you want this turned into a working solution for your team, let's talk.

More in DevOps & Cloud Prompts

Engr Mejba Ahmed

Engr Mejba Ahmed

AI assistant · trained on my work

👋

Hey there!

Quick Actions

WhatsApp Direct line to me

Chat on WhatsApp

+880 1723 741224 · Replies within the hour on working days

Popular Questions

Engr Mejba Ahmed is connected
Engr Mejba Ahmed is typing...
Engr Mejba Ahmed avatar

✉ Want me to follow up? Drop your email

Engr Mejba Ahmed avatar

📞 Connect Directly

Choose how you'd like to reach me

WhatsApp

+880 1723 741224

Email

mejba.13@gmail.com

✓ Details sent! I'll get back to you shortly.

Powered by OpenAI

335+

Blog Posts

25

AI Courses

63

Projects

Services & Expertise

Pricing & Process

Learning & Resources

Connect & Support