Skip to main content

Claude/ChatGPT Prompt to Right-Size Kubernetes CPU and Memory

Analyse Prometheus metrics to right-size Kubernetes CPU/memory requests and limits, flag OOM risk, and emit kubectl patches.

Fill in the placeholders

Edit the values, then copy your finished prompt.

Your Prompt
prompt.txt
You are a senior DevOps engineer. Return concrete recommendations and patch commands, not generic capacity advice.

Context:
- Namespace: production
- Cluster context: AWS EKS prod-us-east-1
- Optimisation goal: cut node cost without raising OOM risk
- Prometheus metrics (last 14 days): <paste Prometheus usage metrics>

Deliverables:
1. Per deployment: current CPU/memory request and limit.
2. p95 actual usage over the window for CPU and memory.
3. A suggested request sized to real usage with headroom.
4. A suggested limit that protects neighbours without throttling normal load.
5. Expected cost or node savings from the change.
6. A flag on any deployment at risk of OOM under current settings.

Output: a table of recommendations plus a ready-to-run kubectl patch command per change.

What this prompt does

This prompt right-sizes Kubernetes CPU and memory from real Prometheus data and returns concrete patches, not generic capacity advice. It casts the assistant as a senior DevOps engineer and takes four context variables: [namespace], [cluster], [goal], and a [metrics_snapshot] of the last 14 days. Per deployment it returns the current CPU/memory request and limit, p95 actual usage over the window, a suggested request sized to real usage with headroom, a suggested limit that protects neighbours without throttling normal load, expected cost or node savings, and a flag on any deployment at risk of OOM. The output is a recommendations table plus a ready-to-run kubectl patch per change.

The structure works because it forces decisions off real usage instead of vibes. Requests and limits get guessed once and never revisited, so half the nodes sit idle while a couple of pods quietly OOM. Feeding [metrics_snapshot] grounds every recommendation in p95 numbers. The [goal] variable ("cut node cost without raising OOM risk") tells the model which direction to optimize, and separating request advice from limit advice matters because getting requests right is what actually changes scheduling and cost.

When to use it

  • Quarterly capacity reviews of a cluster whose requests were guessed and forgotten
  • After a traffic shift that changed real usage patterns
  • When nodes sit idle while a few pods quietly OOM
  • When you want to cut node cost without raising OOM risk
  • When request and limit values were copied between services without thought
  • When you need kubectl patches you can apply, not abstract sizing guidance

Example output

You get a per-deployment table: current request and limit, p95 actual CPU and memory over the 14-day window, a suggested request sized to real usage with headroom, a suggested limit that protects neighbours, and an estimated cost or node saving. Deployments at OOM risk under current settings are flagged so you address the dangerous cases before chasing savings. Each recommended change comes with a ready-to-run kubectl patch command, so you can review the table and apply changes one at a time. Seeing current and suggested values side by side makes the reasoning auditable rather than a black box, which matters when you're about to change scheduling behavior across a production namespace.

Pro tips

  • Provide a full 14-day [metrics_snapshot] covering at least one peak cycle, since p95 from a quiet window will undersize requests
  • State [goal] clearly so the model knows whether to favor cost savings or headroom
  • Move requests before limits — getting requests right is what actually changes scheduling and node cost
  • Review the OOM-risk flags first and address those deployments before chasing savings elsewhere
  • Apply the kubectl patches one deployment at a time and watch behavior, rather than batching every change
  • Re-run after the next traffic shift; right-sizing is a recurring task, not a one-time fix

Frequently Asked Questions

What data do I need to provide?
You paste a Prometheus metrics snapshot covering the last 14 days for the namespace. The model bases its p95 usage figures and recommendations on that data, so a window missing your peak traffic will undersize requests and skew the suggestions.
Does it change requests, limits, or both?
It recommends both, but separates them deliberately. Getting requests right is what actually changes scheduling and node cost, so prioritize those, while limits are tuned to protect neighbouring pods without throttling normal load.
Will it warn me about pods that might OOM?
Yes. It flags any deployment at risk of out-of-memory under current settings. Review those flags first and address them before chasing cost savings elsewhere, since an OOM kill is more disruptive than an idle node.
Are the kubectl patches safe to apply blindly?
Treat them as reviewed recommendations, not guaranteed-safe commands. Apply them one deployment at a time and watch behavior, because right-sizing based on a snapshot can still miss bursts that fall outside your 14-day window.
Engr Mejba Ahmed

Need this built for real?

Engr Mejba Ahmed

AI Developer · Software Engineer

I'm Mejba — I design and ship production AI systems, automations, and full-stack apps. If you want this turned into a working solution for your team, let's talk.

More in DevOps & Cloud Prompts

Engr Mejba Ahmed

Engr Mejba Ahmed

AI assistant · trained on my work

👋

Hey there!

Quick Actions

WhatsApp Direct line to me

Chat on WhatsApp

+880 1723 741224 · Replies within the hour on working days

Popular Questions

Engr Mejba Ahmed is connected
Engr Mejba Ahmed is typing...
Engr Mejba Ahmed avatar

✉ Want me to follow up? Drop your email

Engr Mejba Ahmed avatar

📞 Connect Directly

Choose how you'd like to reach me

WhatsApp

+880 1723 741224

Email

mejba.13@gmail.com

✓ Details sent! I'll get back to you shortly.

Powered by OpenAI

335+

Blog Posts

25

AI Courses

63

Projects

Services & Expertise

Pricing & Process

Learning & Resources

Connect & Support