Skip to main content

Claude/ChatGPT Prompt to Set Up AWS CloudWatch Observability

Set up AWS CloudWatch observability: custom business metrics, dashboards, alarms, and Log Insights queries to track and troubleshoot your live infrastructure.

Vul de plaatshouders in

Edit the values, then copy your finished prompt.

Jouw Prompt
prompt.txt

                                

What this prompt does

This prompt sets up comprehensive CloudWatch observability for an [application_type] running across [services_list]. It instruments [application_name] to publish [business_metrics] using the embedded metric format from [compute_service], builds operations dashboards with [dashboard_sections], and configures [alarm_count] alarms covering [alarm_scenarios]. The goal is to actually see what a system is doing before something breaks, rather than reconstructing it afterward from raw logs.

The structure works because it pairs signal with noise reduction. It defines composite alarms combining [composite_conditions] so on-call only pages when conditions like high error rate AND elevated latency coincide, writes saved Log Insights queries for [log_query_scenarios], enables anomaly detection on [anomaly_metrics], and even estimates CloudWatch cost against [monitoring_budget]. Business metrics plus composite alarms cut the alert fatigue that gets monitoring ignored, and the saved queries pay off the first time you debug an incident at speed.

When to use it

  • You are standing up monitoring for a new [application_type] and want it instrumented before launch
  • Your alerts are noisy and you need composite alarms to escalate only on meaningful [composite_conditions]
  • You want business metrics like [business_metrics], not just CPU and memory, on your dashboards
  • You need a library of saved Log Insights queries for [log_query_scenarios] ready before the next incident
  • You want anomaly detection on [anomaly_metrics] to catch drift that static thresholds miss
  • You need to keep CloudWatch spend under a [monitoring_budget] with sampling and aggregation

Example output

Expect an observability plan: instrumentation guidance for publishing your business metrics via embedded metric format, dashboard definitions with appropriate widget types for each section, a set of alarms with evaluation periods and thresholds, composite-alarm definitions, saved Log Insights queries with regex and stats aggregation, anomaly-detection configuration, cross-account sharing setup, and a cost estimate with reduction strategies to stay under budget.

Pro tips

  • List [business_metrics] that map to outcomes you care about — orders per minute and payment success rate tell you more during an incident than raw infrastructure counters
  • Use [composite_conditions] to combine signals before paging; requiring error rate AND latency together kills most false alarms from a single transient spike
  • Make [log_query_scenarios] specific, like "error rate by endpoint" or "slow query identification," so the saved queries are ready to run mid-incident
  • Give anomaly detection on [anomaly_metrics] a real training period before trusting its bands; it needs baseline data to avoid noisy alerts
  • Keep an eye on [monitoring_budget] — high-cardinality custom metrics and verbose logs add up fast, so apply the sampling strategies the prompt suggests
  • The model proposes thresholds, but tune them against your real traffic; a latency threshold that fits one workload will flap on another
  • Pick widget types in [dashboard_sections] deliberately — gauges suit single health values, while line charts suit latency percentiles over time
  • Set up cross-account sharing from [source_accounts] to a central monitoring account so on-call has one place to look during an incident instead of hopping between accounts

Frequently Asked Questions

How do composite alarms reduce alert noise?
They only fire when multiple conditions you define in `[composite_conditions]` are true at once, such as high error rate combined with elevated latency. A single transient spike on one metric no longer pages on-call, which cuts the false alarms that lead teams to start ignoring their monitoring entirely.
Can I track custom business metrics, not just infrastructure metrics?
Yes. That is a core part of the prompt. You define `[business_metrics]` like orders per minute or payment success rate, and it instruments them using CloudWatch embedded metric format from your `[compute_service]`. These outcome metrics are often more useful during an incident than CPU or memory alone.
Will this help me stay within a CloudWatch budget?
It includes a step to estimate CloudWatch costs and recommend sampling and aggregation strategies to stay under your `[monitoring_budget]`. High-cardinality custom metrics and verbose logging drive cost quickly, so apply those strategies. Treat the estimate as a guide and confirm against your actual billing once metrics flow.
Do I need to tune the alarm thresholds it generates?
Yes. The prompt proposes evaluation periods and thresholds for your `[alarm_scenarios]`, but appropriate values depend on your real traffic patterns. A latency or error-rate threshold that fits one workload will flap on another, so validate and adjust each alarm against observed baselines before relying on it for paging.
Engr Mejba Ahmed

Need this built for real?

Engr Mejba Ahmed

AI Developer · Software Engineer

I'm Mejba — I design and ship production AI systems, automations, and full-stack apps. If you want this turned into a working solution for your team, let's talk.

Meer in AWS & Cloud Architecture Prompts

Engr Mejba Ahmed

Engr Mejba Ahmed

Claude Code Expert · Online

👋

Hey there!

Quick Actions

WhatsApp Instant reply

Chat on WhatsApp

+880 1723 741224 · Instant reply

Popular Questions

Engr Mejba Ahmed is connected
Engr Mejba Ahmed is typing...
Engr Mejba Ahmed avatar

✉ Want me to follow up? Drop your email

Engr Mejba Ahmed avatar

📞 Connect Directly

Choose how you'd like to reach me

WhatsApp

+880 1723 741224

Email

[email protected]

✓ Details sent! I'll get back to you shortly.

Powered by OpenAI

335+

Blog Posts

25

AI Courses

63

Projects

Services & Expertise

Pricing & Process

Learning & Resources

Connect & Support