Skip to main content

Infrastructure Drift Detection System

Build an automated infrastructure drift detection system that identifies manual changes, classifies severity, and triggers remediation.

Fill in the placeholders

Edit the values, then copy your finished prompt.

Your Prompt
prompt.txt
Build an infrastructure drift detection system for AWS infrastructure (networking, compute, databases, IAM) managed with Terraform. The infrastructure spans 800+ resources across 5 AWS accounts accounts/projects. Design: 1) Scheduled drift detection using terraform plan -detailed-exitcode in CI with JSON output — run Terraform plan in a read-only CI pipeline every every 4 hours to detect differences between desired state and actual state. Parse the plan output to extract: resource address, change type (create/update/delete), and changed attributes. 2) Classification engine: categorize drift as critical, high, medium, low, informational based on rules — security-relevant changes (security group ingress changes, IAM policy modifications, S3 bucket policy changes, encryption disabled) are critical, configuration drift (instance type changes, scaling policy modifications, alarm threshold changes) is medium, tag-only changes are low. 3) Notification system: send drift reports to Slack #infra-drift and email to infrastructure team with severity, affected resources, diff details, and links to the relevant Terraform code. Critical drift triggers PagerDuty alert to on-call SRE. 4) A drift dashboard showing: current drift count by severity, drift trend over time, most frequently drifting resources, and time-to-remediation metrics. Store data in PostgreSQL with Grafana visualization. 5) Automated remediation for tag drift, non-security configuration drift in non-production environments — trigger a Terraform apply for low-risk drift with proper approval gates. Flag high-risk drift for manual review. 6) Root cause tracking: correlate drift with AWS CloudTrail to identify who made the manual change and why. 7) Prevention measures: SCPs blocking console changes to managed resources, mandatory IaC PR process, education on IaC workflows to reduce drift occurrence over time.

What this prompt does

This prompt hands the AI a full specification for an infrastructure drift detection system rather than asking a vague "how do I detect drift" question. You set [infrastructure_scope] and [iac_tool] so the design targets a real estate, then [resource_count] and [account_count] give it the scale to design for. The model returns a layered system: scheduled detection running every [scan_interval] using [detection_method], a classification engine that sorts findings into [severity_levels], notifications to [alert_channels], a dashboard backed by [dashboard_storage], and automated remediation for [auto_remediate_categories].

The structure works because drift is not one problem but several. The [security_drift_examples] need to scream while [config_drift_examples] only need a ticket, and tag-only changes can often self-heal. By forcing the AI to separate detection, classification, escalation via [escalation_path], root-cause tracking against [audit_source], and prevention via [prevention_strategies], you get a pipeline that distinguishes signal from noise instead of one giant alert firehose. The [resource_count] and [account_count] inputs also matter, because a system that scans 50 resources in one account looks nothing like one sweeping hundreds across several, and the design adapts the batching and parallelism accordingly.

When to use it

  • You manage IaC-defined infrastructure where people still occasionally make console changes
  • You want drift caught on a schedule, not discovered during the next failed plan
  • You need to justify severity tiers so on-call is paged only for real risk
  • You are building a drift dashboard and need a data model and metrics to track
  • You want to wire [audit_source] correlation in so each drift event names who changed what
  • You are designing prevention controls like SCPs and want them scoped sensibly
  • You need time-to-remediation metrics to show leadership drift is actually shrinking

Example output

Expect a structured design document: numbered sections matching the seven parts of the prompt, a classification rule table mapping change types to [severity_levels], a suggested schema for the [dashboard_storage] backend, and sample notification payloads for [alert_channels]. It typically includes a sketch of the CI pipeline steps that run [detection_method] and parse the output, plus a remediation decision tree showing which categories auto-apply and which get flagged for review. You usually also get a description of the dashboard panels — drift count by severity, trend over time, and the most frequently drifting resources — so the metrics layer is concrete rather than aspirational.

Pro tips

  • Make [detection_method] concrete — naming terraform plan -detailed-exitcode with JSON output gets you parsing logic instead of hand-waving
  • Tune [scan_interval] to your change velocity; every 4 hours is reasonable, but hourly on a hot estate or daily on a stable one both have a place
  • Keep [auto_remediate_categories] conservative at first — tag and non-prod config drift only, never security-relevant changes
  • Spell out [security_drift_examples] precisely, since this is the list that decides what pages [escalation_path] at 3am
  • Pair detection with real [prevention_strategies]; a system that only reports drift you keep re-introducing is treating the symptom
  • Route [alert_channels] so medium and low drift lands somewhere quiet and only critical drift interrupts people, or alert fatigue sets in fast
  • Iterate by feeding back the actual plan JSON shape your tool emits so the parsing logic matches reality, not a generic example

Frequently Asked Questions

Does this prompt write the actual drift detection code or just the design?
It produces a design with concrete structure: pipeline steps, classification rules, schema, and notification formats. You will still need to implement the parsing and remediation logic, but the output is detailed enough to code against rather than abstract advice.
Can I use it with Pulumi or OpenTofu instead of Terraform?
Yes. Set `[iac_tool]` to your tool and adjust `[detection_method]` to that tool's plan or preview command. The classification, escalation, and prevention layers are tool-agnostic, so only the detection step changes meaningfully.
How does it decide what counts as critical drift?
You define the rules through `[security_drift_examples]`, `[config_drift_examples]`, and `[severity_levels]`. The model builds a classification engine around those, treating security group, IAM, and encryption changes as critical and tag-only changes as low by default.
Is automated remediation safe to enable?
Only for the narrow `[auto_remediate_categories]` you specify, ideally tag drift and non-production config drift behind approval gates. Auto-applying security or production changes is risky, so the prompt is designed to flag those for manual review instead.
Engr Mejba Ahmed

Need this built for real?

Engr Mejba Ahmed

AI Developer · Software Engineer

I'm Mejba — I design and ship production AI systems, automations, and full-stack apps. If you want this turned into a working solution for your team, let's talk.

More in Terraform & Infrastructure as Code Prompts

Engr Mejba Ahmed

Engr Mejba Ahmed

AI assistant · trained on my work

👋

Hey there!

Quick Actions

WhatsApp Direct line to me

Chat on WhatsApp

+880 1723 741224 · Replies within the hour on working days

Popular Questions

Engr Mejba Ahmed is connected
Engr Mejba Ahmed is typing...
Engr Mejba Ahmed avatar

✉ Want me to follow up? Drop your email

Engr Mejba Ahmed avatar

📞 Connect Directly

Choose how you'd like to reach me

WhatsApp

+880 1723 741224

Email

mejba.13@gmail.com

✓ Details sent! I'll get back to you shortly.

Powered by OpenAI

335+

Blog Posts

25

AI Courses

63

Projects

Services & Expertise

Pricing & Process

Learning & Resources

Connect & Support