Skip to main content

Design a Metrics & Monitoring System

Design a distributed metrics collection, aggregation, and alerting system for monitoring infrastructure and application health at scale.

Fill in the placeholders

Edit the values, then copy your finished prompt.

Your Prompt
prompt.txt
You are a site reliability engineer. Design a metrics and monitoring system for a cloud-native microservices infrastructure with 5,000 servers generating 2 million metrics per second.

Step 1: Define the metrics taxonomy. Categorize metrics into infrastructure (CPU, memory, disk, network), application (request rate, error rate, latency percentiles), business (orders per minute, revenue, conversion rate), and custom metrics from 200 applications. Define the data model: metric name, tags/labels (key-value pairs for dimensions), timestamp, and value. Support counter, gauge, histogram, and summary metric types.

Step 2: Design the collection pipeline. Deploy a lightweight agent on each server that scrapes local metrics every 15 seconds and forwards them to regional aggregators. Application metrics are emitted via a Prometheus client client library using either push (StatsD protocol) or pull (Prometheus-compatible /metrics endpoint). Regional aggregators pre-aggregate high-cardinality metrics before forwarding to the central Thanos (Prometheus long-term).

Step 3: Architect the storage layer using Thanos (Prometheus long-term) optimized for time-series write patterns. Implement a tiered retention policy: full resolution for 7 days, 1-minute rollups for 90 days, and 1-hour rollups for 2 years in cheap object storage. Design the schema to support efficient queries by metric name, tag filters, and time range. Estimate the total storage needed for 2 million metrics/second over the retention periods.

Step 4: Build the query engine that supports common monitoring queries: current value, rate of change, percentile calculations (p50, p95, p99), aggregation across dimensions (sum, avg, max by tag), and cross-metric formulas (error rate = errors / total requests). Optimize for dashboard queries that refresh every 10 seconds with sub-second response time. Implement query caching for popular dashboards.

Step 5: Design the alerting system. Define alert rules using a PromQL expression language (e.g., "avg(cpu_usage{service=api}) > 80 for 5m"). Implement a multi-stage evaluation pipeline: rule evaluation every 30 seconds, state machine (OK, PENDING, FIRING, RESOLVED), and notification routing to PagerDuty, Slack, email. Build alert deduplication, grouping (batch related alerts), and silencing (suppress during maintenance windows).

Step 6: Create the dashboard and visualization layer. Design a template system for standard dashboards (service overview, infrastructure health, SLO tracking). Implement SLO monitoring with error budget tracking and burn-rate alerting. Build an anomaly detection module that uses seasonal decomposition with dynamic thresholds to automatically detect unusual patterns without manually configured thresholds.

What this prompt does

This prompt makes the AI a site reliability engineer designing a metrics and monitoring system for [infrastructure_type] infrastructure with [server_count] servers generating [metrics_per_second] metrics per second. It covers the metrics taxonomy and data model, the collection pipeline, tiered storage, the query engine, alerting, and dashboards with SLO tracking. Agents scrape each server every [scrape_interval], regional aggregators pre-aggregate high-cardinality metrics, and everything lands in [time_series_db].

The structure works because monitoring at scale is dominated by two costs people underestimate: cardinality and retention. The tiered policy keeps full resolution for [hot_retention], one-minute rollups for [warm_retention], and one-hour rollups for [cold_retention] in cheap object storage, which keeps storage growth from [metrics_per_second] under control while preserving queryable history. Alerting uses an [alert_language] expression engine with a proper OK→PENDING→FIRING→RESOLVED state machine plus deduplication, grouping of related alerts, and silencing during maintenance, so on-call isn't flooded. SLO burn-rate alerting then catches slow degradations before the error budget is gone, which a static threshold would miss entirely.

When to use it

  • You're standing up monitoring and want cardinality and retention cost controlled from day one
  • You need a clear metrics taxonomy and data model across many applications
  • You want a collection pipeline that pre-aggregates high-cardinality metrics before central storage
  • You need tiered retention to keep [time_series_db] storage affordable
  • You want alerting with proper state, dedup, grouping, and maintenance silencing
  • You need SLO tracking with error budgets and burn-rate alerts
  • You want anomaly detection via [anomaly_method] instead of hand-tuned thresholds everywhere

Example output

Expect a design document: a metrics taxonomy with the name/tags/timestamp/value model and the counter, gauge, histogram, and summary types, a collection pipeline with per-server agents scraping every [scrape_interval] and regional aggregators feeding [time_series_db], a tiered storage section with retention math across the hot, warm, and cold tiers, a query engine covering rates, percentiles, dimensional aggregation, and cross-metric formulas, an alerting pipeline with the state machine and routing to [notification_channels], and a dashboard layer with SLO error-budget tracking, burn-rate alerting, and [anomaly_method] detection. It's a reasoned architecture, not code.

Pro tips

  • Estimate storage for [metrics_per_second] across all retention tiers early; it's the number that surprises teams most
  • Pre-aggregate high-cardinality metrics at the regional aggregators before they hit [time_series_db]
  • Keep [hot_retention] short and lean on rollups; full resolution for months is rarely worth the cost
  • Insist on SLO burn-rate alerting, not just threshold alerts — it catches slow degradations a static threshold misses
  • Use silencing during maintenance windows so planned work doesn't page anyone via [notification_channels]
  • Pre-aggregate at the regional layer rather than shipping raw high-cardinality series straight to the central store
  • Be careful adding tags; every new label dimension multiplies cardinality and the cost behind it

Frequently Asked Questions

Why does this design emphasize cardinality and retention so much?
Cardinality (the number of unique tag combinations) and retention (how long data is kept) are the two costs that grow fastest and surprise teams. The prompt pre-aggregates high-cardinality metrics and uses tiered rollups so storage for `[metrics_per_second]` stays affordable rather than ballooning over time.
How does tiered retention reduce storage cost?
It keeps full-resolution data only for `[hot_retention]`, then downsamples to one-minute rollups for `[warm_retention]` and one-hour rollups for `[cold_retention]` in cheap object storage. Older data stays queryable at coarser granularity, so you keep history without paying to store every raw datapoint forever.
What is burn-rate alerting and why prefer it over thresholds?
Burn-rate alerting tracks how fast you're consuming an SLO's error budget rather than firing on a fixed value. It catches slow, sustained degradations that a static threshold would miss and avoids noise from brief spikes, which is why the prompt treats it as a production requirement alongside SLO tracking.
Does the alerting handle noisy or duplicate alerts?
Yes. The alerting pipeline includes a state machine, deduplication, grouping of related alerts into a single notification, and silencing during maintenance windows. Combined with routing to `[notification_channels]`, this keeps on-call from being flooded by repeated or expected alerts.
Engr Mejba Ahmed

Need this built for real?

Engr Mejba Ahmed

AI Developer · Software Engineer

I'm Mejba — I design and ship production AI systems, automations, and full-stack apps. If you want this turned into a working solution for your team, let's talk.

More in System Design Prompts

Engr Mejba Ahmed

Engr Mejba Ahmed

AI assistant · trained on my work

👋

Hey there!

Quick Actions

WhatsApp Direct line to me

Chat on WhatsApp

+880 1723 741224 · Replies within the hour on working days

Popular Questions

Engr Mejba Ahmed is connected
Engr Mejba Ahmed is typing...
Engr Mejba Ahmed avatar

✉ Want me to follow up? Drop your email

Engr Mejba Ahmed avatar

📞 Connect Directly

Choose how you'd like to reach me

WhatsApp

+880 1723 741224

Email

mejba.13@gmail.com

✓ Details sent! I'll get back to you shortly.

Powered by OpenAI

335+

Blog Posts

25

AI Courses

63

Projects

Services & Expertise

Pricing & Process

Learning & Resources

Connect & Support