What this prompt does
This prompt right-sizes Kubernetes CPU and memory from real Prometheus data and returns concrete patches, not generic capacity advice. It casts the assistant as a senior DevOps engineer and takes four context variables: [namespace], [cluster], [goal], and a [metrics_snapshot] of the last 14 days. Per deployment it returns the current CPU/memory request and limit, p95 actual usage over the window, a suggested request sized to real usage with headroom, a suggested limit that protects neighbours without throttling normal load, expected cost or node savings, and a flag on any deployment at risk of OOM. The output is a recommendations table plus a ready-to-run kubectl patch per change.
The structure works because it forces decisions off real usage instead of vibes. Requests and limits get guessed once and never revisited, so half the nodes sit idle while a couple of pods quietly OOM. Feeding [metrics_snapshot] grounds every recommendation in p95 numbers. The [goal] variable ("cut node cost without raising OOM risk") tells the model which direction to optimize, and separating request advice from limit advice matters because getting requests right is what actually changes scheduling and cost.
When to use it
- Quarterly capacity reviews of a cluster whose requests were guessed and forgotten
- After a traffic shift that changed real usage patterns
- When nodes sit idle while a few pods quietly OOM
- When you want to cut node cost without raising OOM risk
- When request and limit values were copied between services without thought
- When you need kubectl patches you can apply, not abstract sizing guidance
Example output
You get a per-deployment table: current request and limit, p95 actual CPU and memory over the 14-day window, a suggested request sized to real usage with headroom, a suggested limit that protects neighbours, and an estimated cost or node saving. Deployments at OOM risk under current settings are flagged so you address the dangerous cases before chasing savings. Each recommended change comes with a ready-to-run kubectl patch command, so you can review the table and apply changes one at a time. Seeing current and suggested values side by side makes the reasoning auditable rather than a black box, which matters when you're about to change scheduling behavior across a production namespace.
Pro tips
- Provide a full 14-day
[metrics_snapshot]covering at least one peak cycle, since p95 from a quiet window will undersize requests - State
[goal]clearly so the model knows whether to favor cost savings or headroom - Move requests before limits — getting requests right is what actually changes scheduling and node cost
- Review the OOM-risk flags first and address those deployments before chasing savings elsewhere
- Apply the kubectl patches one deployment at a time and watch behavior, rather than batching every change
- Re-run after the next traffic shift; right-sizing is a recurring task, not a one-time fix