What this prompt does
This prompt makes the AI a site reliability engineer designing a metrics and monitoring system for [infrastructure_type] infrastructure with [server_count] servers generating [metrics_per_second] metrics per second. It covers the metrics taxonomy and data model, the collection pipeline, tiered storage, the query engine, alerting, and dashboards with SLO tracking. Agents scrape each server every [scrape_interval], regional aggregators pre-aggregate high-cardinality metrics, and everything lands in [time_series_db].
The structure works because monitoring at scale is dominated by two costs people underestimate: cardinality and retention. The tiered policy keeps full resolution for [hot_retention], one-minute rollups for [warm_retention], and one-hour rollups for [cold_retention] in cheap object storage, which keeps storage growth from [metrics_per_second] under control while preserving queryable history. Alerting uses an [alert_language] expression engine with a proper OK→PENDING→FIRING→RESOLVED state machine plus deduplication, grouping of related alerts, and silencing during maintenance, so on-call isn't flooded. SLO burn-rate alerting then catches slow degradations before the error budget is gone, which a static threshold would miss entirely.
When to use it
- You're standing up monitoring and want cardinality and retention cost controlled from day one
- You need a clear metrics taxonomy and data model across many applications
- You want a collection pipeline that pre-aggregates high-cardinality metrics before central storage
- You need tiered retention to keep
[time_series_db]storage affordable - You want alerting with proper state, dedup, grouping, and maintenance silencing
- You need SLO tracking with error budgets and burn-rate alerts
- You want anomaly detection via
[anomaly_method]instead of hand-tuned thresholds everywhere
Example output
Expect a design document: a metrics taxonomy with the name/tags/timestamp/value model and the counter, gauge, histogram, and summary types, a collection pipeline with per-server agents scraping every [scrape_interval] and regional aggregators feeding [time_series_db], a tiered storage section with retention math across the hot, warm, and cold tiers, a query engine covering rates, percentiles, dimensional aggregation, and cross-metric formulas, an alerting pipeline with the state machine and routing to [notification_channels], and a dashboard layer with SLO error-budget tracking, burn-rate alerting, and [anomaly_method] detection. It's a reasoned architecture, not code.
Pro tips
- Estimate storage for
[metrics_per_second]across all retention tiers early; it's the number that surprises teams most - Pre-aggregate high-cardinality metrics at the regional aggregators before they hit
[time_series_db] - Keep
[hot_retention]short and lean on rollups; full resolution for months is rarely worth the cost - Insist on SLO burn-rate alerting, not just threshold alerts — it catches slow degradations a static threshold misses
- Use silencing during maintenance windows so planned work doesn't page anyone via
[notification_channels] - Pre-aggregate at the regional layer rather than shipping raw high-cardinality series straight to the central store
- Be careful adding tags; every new label dimension multiplies cardinality and the cost behind it