Claude Cursor GitHub Copilot Skill

rightsize-recommendation

Emit pod CPU and memory request/limit recommendations from user-pasted p50, p95, and p99 utilization metrics over a 7-14 day window. Outputs recommended requests at p95 plus 20% headroom, limits at p99 plus 30%, estimated monthly savings, and Karpenter consolidation eligibility f

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download vincentchuwaichow-vanguard-frontier-agentic-skills_finops_rightsize-recommendation-febe32a.zip · 8 KB
Part of vincentchuwaichow/vanguard-frontier-agentic — 293 skills

Install

skills CLI npx skills add https://github.com/VincentChuWaiChow/vanguard-frontier-agentic/tree/master/skills/finops/rightsize-recommendation
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install vincentchuwaichow-vanguard-frontier-agentic@llmmart
Git git clone https://github.com/VincentChuWaiChow/vanguard-frontier-agentic.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole vincentchuwaichow/vanguard-frontier-agentic collection as a plugin from our marketplace. Git is the plain clone.

README

Rightsize Recommendation

A FinOps skill that emits Kubernetes pod CPU and memory request/limit recommendations from user-supplied utilization percentile metrics. Read-only, stateless math — no cluster connection, no kubectl.

Purpose

Given p50, p95, and p99 CPU and memory utilization metrics from a 7-14 day window, compute recommended resource requests (p95 + 20% headroom) and limits (p99 + 30% headroom), estimate monthly savings if a unit price is provided, and assess Karpenter consolidation eligibility.

Allowed tools

Read Grep Glob

Usage

Single pod: Paste the pod name, namespace, current CPU/memory requests and limits, and p50/p95/p99 metrics. Optionally include $/vCPU-hour and $/GiB-hour for a savings estimate. The skill returns recommended requests and limits with headroom rationale and a consolidation eligibility flag.

Batch: Paste metrics for multiple pods or workloads in any consistent format (table, YAML snippet, or list). The skill returns a recommendation block per pod.

Trust posture

Read-only. No cloud credentials, billing account IDs, or tenant data accepted. No cluster connection is made. All outputs are labeled inference (computed from caller inputs) or assumed (where defaults were applied). Monthly savings estimates are excluded unless unit price is provided by the caller.

See SKILL.md for the full methodology, required input format, and response shape.

Skill manifest

Rightsize Recommendation

Purpose

Produce CPU and memory rightsizing recommendations for Kubernetes pods based on user-supplied utilization percentile metrics. All math is performed on the inputs provided by the caller; no cluster connection or live metric fetch is performed.

No kubectl. No WebFetch. No cluster credentials accepted.

When to use

Use this skill when:

  • The user has pasted or attached p50, p95, and p99 CPU and memory utilization metrics for one or more pods or workloads
  • The user wants actionable request and limit recommendations with explicit headroom rationale
  • The user wants an estimated monthly savings figure (requires unit price from the caller)
  • The user wants to know whether a pod is eligible for Karpenter node consolidation

Operating rules

  • Inputs only. All calculations are performed on the metrics and pod specs supplied by the caller. This skill does not fetch live metrics, connect to Prometheus, or query the Kubernetes API.
  • No credentials accepted. No cloud credentials, billing account IDs, tenant data, kubeconfig, bearer tokens, or service account JWTs are accepted or required.
  • Metric window requirement. The caller must supply metrics from a 7-14 day representative window. Metrics from shorter windows are accepted but the recommendation must be labeled with a reduced-confidence warning.
  • Provenance labels mandatory. Every numeric output must carry one label:
    • inference — computed from caller-supplied inputs using the documented methodology
    • assumed — derived from a default assumption where caller data was missing (state the assumption)
    • excluded — cost savings figure excluded because unit price was not supplied by the caller
  • FOCUS column mapping. Where a cost savings estimate is produced, note: BilledCost (current), EffectiveCost (projected after rightsizing), ChargeCategory = Usage, ServiceCategory = Containers.
  • Load references only when needed.

Required input

The caller must supply, per pod or workload:

  • Pod name and namespace (or Deployment/StatefulSet name)
  • Current CPU request (millicores) and current memory request (MiB)
  • Current CPU limit (millicores) and current memory limit (MiB); state no-limit if unset
  • p50 CPU (millicores), p95 CPU (millicores), p99 CPU (millicores)
  • p50 memory (MiB), p95 memory (MiB), p99 memory (MiB)
  • Metric window (days, must be 7-14; shorter accepted with warning)
  • Unit price per vCPU-hour and per GiB-hour (optional; required only for $/mo saved estimate)

If any field is missing, ask one clarifying question per gap. Do not fabricate metric values.

Recommendation methodology

CPU

Output Formula Rationale
Recommended CPU request p95 CPU + 20% p95 covers normal burst; 20% headroom absorbs measurement noise and brief spikes
Recommended CPU limit p99 CPU + 30% p99 covers rare spikes; 30% headroom prevents throttling during outlier bursts

Memory

Output Formula Rationale
Recommended memory request p95 memory + 20% Same headroom logic as CPU; memory pressure leads to OOMKill rather than throttling
Recommended memory limit p99 memory + 30% Tighter limits risk OOMKill; 30% buffer is conservative for stateful workloads

Round all recommendations up to the nearest 10 millicores (CPU) or 16 MiB (memory) for cleaner manifests.

Monthly savings estimate

If the caller supplies unit price ($/vCPU-hour and $/GiB-hour):

cpu_savings_per_month = (current_cpu_request - recommended_cpu_request) / 1000
                       × unit_price_per_vcpu_hour × 730

memory_savings_per_month = (current_memory_request - recommended_memory_request) / 1024
                          × unit_price_per_gib_hour × 730

Label: inference. Note: savings are approximate; they assume the freed capacity is not replaced by new workloads and that the node pool scales down proportionally.

If unit price is not supplied, state the savings formula and mark the $/mo cell as excluded — unit price not provided by caller.

Karpenter consolidation eligibility

A pod is flagged as consolidation-eligible when all of the following are true:

  • No PodDisruptionBudget with maxUnavailable: 0 or minAvailable: 100% applies to the pod.
  • No pod-anti-affinity or topologySpreadConstraints rule that would prevent the pod from co-locating with other pods of the same owner.
  • No hostPath or local PersistentVolume mount.
  • No nodeName node selector pinning the pod to a specific node.
  • The pod's request fits within a standard node SKU that Karpenter would provision (the caller must confirm the NodePool SKU list).

If the caller has not supplied enough information to confirm all five conditions, state which conditions could not be verified and flag as not-verified — [missing conditions]. Do not assume eligibility when data is incomplete; unknown blockers present a consolidation risk.

Response shape

Return, per pod or workload:

Pod / Workload: <name> (<namespace>)

  CPU
    Current request:     <current> m
    Recommended request: <p95 × 1.20, rounded> m  [inference]
    Current limit:       <current> m  (or "no-limit")
    Recommended limit:   <p99 × 1.30, rounded> m  [inference]

  Memory
    Current request:     <current> MiB
    Recommended request: <p95 × 1.20, rounded> MiB  [inference]
    Current limit:       <current> MiB  (or "no-limit")
    Recommended limit:   <p99 × 1.30, rounded> MiB  [inference]

  Monthly savings estimate
    CPU:    $<amount>/mo  [inference]  or  [excluded — unit price not provided]
    Memory: $<amount>/mo  [inference]  or  [excluded — unit price not provided]
    Total:  $<amount>/mo

  Karpenter consolidation
    Eligible: <Yes | No | Not-verified>
    Blockers: <list any confirmed blockers, or list missing data preventing verification, or "None confirmed">

  Metric window: <N days>  [inference: window meets 7-14 day requirement]
  Confidence:   <Normal | Reduced — window < 7 days>

References

Load these only when needed:

  • Metric sources — how to gather p50/p95/p99 from Prometheus, Cloud Monitoring, and Azure Monitor for use as input to this skill.
  • Karpenter consolidation — what makes a pod consolidation-eligible and known blockers.
Files (vanguard-frontier-agentic)
  • references
    • karpenter-consolidation.md 4.7 KB
      # Karpenter Consolidation Eligibility
      
      ## What consolidation does
      
      Karpenter consolidation (enabled via `consolidationPolicy: WhenUnderutilized` or `WhenEmpty` in a NodePool) automatically removes underutilized nodes and reschedules their pods to better-packed nodes. This reduces idle node cost without requiring manual intervention.
      
      For a pod to be a candidate for consolidation, Karpenter must be able to disrupt it safely and reschedule it onto a smaller or fewer nodes.
      
      ## Eligibility conditions
      
      A pod is consolidation-eligible when all five conditions are met:
      
      ### 1. No blocking PodDisruptionBudget
      
      A PodDisruptionBudget (PDB) blocks consolidation if it would prevent Karpenter from evicting the pod at the time of consolidation.
      
      Blocking PDB configurations:
      - `maxUnavailable: 0` — zero replicas can be disrupted; Karpenter cannot evict any pod.
      - `minAvailable` equal to the current replica count — effectively blocks all evictions.
      - `maxUnavailable: 0%` — same as `maxUnavailable: 0` in percentage form.
      
      Non-blocking PDB configurations:
      - `maxUnavailable: 1` or higher — at least one replica can be disrupted.
      - `minAvailable` less than the current replica count — some replicas can be disrupted.
      
      To check for blocking PDBs: the caller should inspect `kubectl get pdb -n <namespace>` output and confirm whether any PDB covers the workload.
      
      ### 2. No consolidation-blocking pod anti-affinity
      
      Pod anti-affinity rules that spread replicas across nodes by `kubernetes.io/hostname` (host anti-affinity) can prevent Karpenter from co-locating pods that are currently on separate nodes.
      
      If every replica of a Deployment has a `requiredDuringSchedulingIgnoredDuringExecution` anti-affinity rule against other replicas on the same host, Karpenter cannot pack them together, blocking node consolidation.
      
      `preferredDuringSchedulingIgnoredDuringExecution` anti-affinity is soft and does not block consolidation.
      
      ### 3. No topologySpreadConstraints with WhenUnsatisfiable: DoNotSchedule
      
      `topologySpreadConstraints` with `whenUnsatisfiable: DoNotSchedule` and a spread domain of `kubernetes.io/hostname` prevents Karpenter from scheduling multiple pods onto the same node, blocking consolidation of node pairs.
      
      `whenUnsatisfiable: ScheduleAnyway` is soft and does not block consolidation.
      
      ### 4. No local storage (hostPath or local PV)
      
      Pods that use `hostPath` volumes or PersistentVolumes backed by a `local` StorageClass are pinned to the node where the volume exists. Karpenter cannot reschedule these pods to a different node without data loss.
      
      Pods using network-attached storage (EBS, Azure Disk, GCE PD, OCI Block Volume) are not pinned and are eligible for consolidation — the volume can be re-attached to the new node.
      
      Note: EBS volumes in AWS have an availability-zone constraint; consolidation targets must be in the same AZ as the volume.
      
      ### 5. No nodeName selector pinning
      
      A pod with `spec.nodeName` set to a specific node name is pinned to that node and cannot be rescheduled. Karpenter will not consolidate the host node as long as pinned pods remain on it.
      
      Similarly, a `nodeSelector` that matches only one specific node (e.g., `kubernetes.io/hostname: <exact-node-name>`) effectively pins the pod.
      
      ## Summary eligibility table
      
      | Condition | Eligible when | Blocks consolidation when |
      |---|---|---|
      | PodDisruptionBudget | No PDB, or PDB with `maxUnavailable >= 1` | PDB with `maxUnavailable: 0` or `minAvailable == replica count` |
      | Pod anti-affinity | No anti-affinity, or `preferredDuring...` only | `requiredDuring...` host anti-affinity between replicas |
      | TopologySpreadConstraints | No constraints, or `whenUnsatisfiable: ScheduleAnyway` | `whenUnsatisfiable: DoNotSchedule` on hostname topology |
      | Storage | Network-attached PV only | `hostPath` volume or `local` StorageClass PV |
      | Node selector | No `nodeName` and no single-node `nodeSelector` | `spec.nodeName` set or `nodeSelector` matching one node |
      
      ## Consolidation policy modes
      
      | Policy | Behavior |
      |---|---|
      | `WhenEmpty` | Consolidates only fully empty nodes (all pods evicted or completed). Conservative. |
      | `WhenUnderutilized` | Consolidates nodes when doing so would reduce total node count without violating pod scheduling constraints. Aggressive. |
      
      For FinOps purposes, `WhenUnderutilized` provides the greater cost reduction but requires that pods pass all five eligibility conditions above.
      
      ## Disruption budget in Karpenter v0.33+
      
      Karpenter v0.33 introduced a `disruption.budgets` field in the NodePool spec that limits the rate of node replacement across all consolidation actions. A low disruption budget (e.g., `nodes: "10%"`) slows consolidation speed but reduces service disruption risk.
      
      The disruption budget does not affect eligibility; it affects when Karpenter will act on an eligible pod.
      
    • metric-sources.md 4.9 KB
      # Metric Sources
      
      This reference describes how to gather p50, p95, and p99 CPU and memory utilization metrics from common observability platforms for use as input to the rightsize-recommendation skill. These are recipes and query patterns, not commands to be executed by the skill.
      
      ## Why percentile metrics matter
      
      Using only the average (mean) of CPU or memory utilization understates peaks and leads to under-resourced pods. Using the p99 alone for requests leads to over-provisioning for workloads with rare spikes. The combination of p50, p95, and p99 over a 7-14 day window captures normal load, burst behavior, and outlier events.
      
      - p50: median utilization; represents steady-state load.
      - p95: captures the top 5% of utilization events; used for request sizing.
      - p99: captures the top 1% of utilization events; used for limit sizing.
      
      ## Prometheus (self-hosted or managed)
      
      ### CPU utilization per pod (millicores)
      
      ```
      # p50 CPU (millicores) over 14 days
      quantile_over_time(0.50,
        rate(container_cpu_usage_seconds_total{
          namespace="<namespace>",
          pod=~"<pod-name-prefix>.*",
          container!="POD",
          container!=""
        }[5m])[14d:5m]
      ) * 1000
      
      # p95 CPU (millicores) over 14 days
      quantile_over_time(0.95, ...) * 1000
      
      # p99 CPU (millicores) over 14 days
      quantile_over_time(0.99, ...) * 1000
      ```
      
      Replace `<namespace>` and `<pod-name-prefix>` with your values. Adjust the window (`14d`) and step (`5m`) as needed. For a Deployment or StatefulSet, aggregate across all replica pods by including `pod=~"<deployment-name>-.*"`.
      
      ### Memory working set per pod (MiB)
      
      ```
      # p95 memory working set (MiB) over 14 days
      quantile_over_time(0.95,
        container_memory_working_set_bytes{
          namespace="<namespace>",
          pod=~"<pod-name-prefix>.*",
          container!="POD",
          container!=""
        }[14d:5m]
      ) / 1048576
      ```
      
      Use `container_memory_working_set_bytes` (not `container_memory_usage_bytes`) because the working set excludes file-backed pages that can be evicted by the kernel without causing an OOMKill. This gives a more accurate picture of the memory the container actually needs to retain.
      
      ## Google Cloud Monitoring (GKE)
      
      Use the Cloud Monitoring Metrics Explorer or the Metrics Query Language (MQL) console.
      
      Relevant metric types:
      
      - CPU: `kubernetes.io/container/cpu/request_utilization` (fraction of requested CPU used) or `kubernetes.io/container/cpu/core_usage_time` (cumulative core-seconds)
      - Memory: `kubernetes.io/container/memory/request_utilization` (fraction of requested memory used) or `kubernetes.io/container/memory/used_bytes`
      
      To compute p95 over 14 days using MQL:
      
      ```
      fetch k8s_container
      | metric 'kubernetes.io/container/memory/used_bytes'
      | filter (resource.namespace_name == '<namespace>'
            && resource.container_name != 'POD')
      | group_by [resource.pod_name, resource.container_name],
          percentile(value.used_bytes, 95)
      | within 14d
      ```
      
      Divide the result by 1,048,576 to convert bytes to MiB before passing to the skill.
      
      ## Azure Monitor (AKS)
      
      Use Container Insights or the Azure Monitor Metrics blade.
      
      Relevant metrics:
      
      - CPU: `cpuUsageNanoCores` (container CPU usage in nanocores; divide by 1,000,000 to convert to millicores)
      - Memory: `memoryWorkingSetBytes` (working set in bytes; divide by 1,048,576 for MiB)
      
      To query p95 over 14 days using Azure Monitor Logs (KQL):
      
      ```kql
      InsightsMetrics
      | where Namespace == "container.azm.ms/cpuUsageNanoCores"
      | where Tags contains '"namespace":"<namespace>"'
      | summarize p95 = percentile(Val, 95) by bin(TimeGenerated, 5m)
      | where TimeGenerated > ago(14d)
      | summarize p95_cpu_mc = percentile(p95, 95) / 1000000
      ```
      
      ## OpenCost (in-cluster)
      
      If OpenCost is deployed, it exposes a REST API that returns pre-computed allocation data including max usage over a window. The skill cannot call this API directly, but the caller can retrieve the data using:
      
      ```
      curl -G http://<opencost-service>:9003/allocation \
        --data-urlencode 'window=14d' \
        --data-urlencode 'aggregate=pod' \
        --data-urlencode 'namespace=<namespace>'
      ```
      
      From the response, extract `cpuCoreRequestAverage`, `cpuCoreUsageAverage`, `ramByteRequestAverage`, and `ramByteUsageAverage`. Note: OpenCost exposes averages, not percentiles. If percentile data is required, use Prometheus directly.
      
      ## Vertical Pod Autoscaler (VPA) recommendations as a starting point
      
      If VPA is deployed in recommendation mode (not auto), it exposes `LowerBound`, `Target`, `UpperBound`, and `UncappedTarget` for CPU and memory. These can serve as an approximate p50-to-p99 range:
      
      - `LowerBound` ≈ p50 equivalent (conservative minimum)
      - `Target` ≈ p95 equivalent (VPA recommendation)
      - `UpperBound` ≈ p99 equivalent (headroom ceiling)
      
      These are VPA's internal heuristics, not directly equivalent to Prometheus percentiles. Label them as `assumed` when using VPA output as a substitute for percentile metrics.
      
      To retrieve VPA recommendations without modifying the cluster:
      
      ```
      # Output from: kubectl describe vpa <vpa-name> -n <namespace>
      # Paste the "Container Recommendations:" block into the skill input
      ```
      
  • metadata.json 1.1 KB
    {
      "id": "rightsize-recommendation",
      "name": "Rightsize Recommendation",
      "type": "skill",
      "provider": "kubernetes",
      "harnesses": [
        "codex",
        "claude-code",
        "cursor",
        "gemini",
        "kiro",
        "other"
      ],
      "summary": "Emit pod CPU and memory request/limit recommendations from user-pasted p50/p95/p99 utilization metrics. Outputs recommended requests at p95 plus 20% headroom, limits at p99 plus 30%, estimated monthly savings, and Karpenter consolidation eligibility. Read-only, no kubectl.",
      "source_type": "original",
      "official_docs": [
        "https://karpenter.sh/docs/",
        "https://kubernetes.io/docs/tasks/run-application/vertical-pod-autoscaler/",
        "https://www.opencost.io/docs/"
      ],
      "security_notes": "No cluster credentials, kubeconfig, bearer tokens, service account JWTs, or cloud IAM credentials are accepted or required. All calculations are performed on user-supplied metric inputs only. No live cluster or metric API connection is made.",
      "last_verified": "2026-05-13",
      "path": "skills/finops/rightsize-recommendation",
      "author": "github: VincentChuWaiChow",
      "version": "0.1.2",
      "lifecycle": "experimental"
    }
    
  • README.md 1.4 KB
    # Rightsize Recommendation
    
    A FinOps skill that emits Kubernetes pod CPU and memory request/limit recommendations from user-supplied utilization percentile metrics. Read-only, stateless math — no cluster connection, no kubectl.
    
    ## Purpose
    
    Given p50, p95, and p99 CPU and memory utilization metrics from a 7-14 day window, compute recommended resource requests (p95 + 20% headroom) and limits (p99 + 30% headroom), estimate monthly savings if a unit price is provided, and assess Karpenter consolidation eligibility.
    
    ## Allowed tools
    
    `Read` `Grep` `Glob`
    
    ## Usage
    
    **Single pod:** Paste the pod name, namespace, current CPU/memory requests and limits, and p50/p95/p99 metrics. Optionally include $/vCPU-hour and $/GiB-hour for a savings estimate. The skill returns recommended requests and limits with headroom rationale and a consolidation eligibility flag.
    
    **Batch:** Paste metrics for multiple pods or workloads in any consistent format (table, YAML snippet, or list). The skill returns a recommendation block per pod.
    
    ## Trust posture
    
    Read-only. No cloud credentials, billing account IDs, or tenant data accepted. No cluster connection is made. All outputs are labeled `inference` (computed from caller inputs) or `assumed` (where defaults were applied). Monthly savings estimates are `excluded` unless unit price is provided by the caller.
    
    See [SKILL.md](SKILL.md) for the full methodology, required input format, and response shape.
    
  • SKILL.md 6.8 KB
    ---
    name: rightsize-recommendation
    description: Emit pod CPU and memory request/limit recommendations from user-pasted p50, p95, and p99 utilization metrics over a 7-14 day window. Outputs recommended requests at p95 plus 20% headroom, limits at p99 plus 30%, estimated monthly savings, and Karpenter consolidation eligibility flag. Read-only, no kubectl.
    allowed-tools: Read Grep Glob
    metadata:
      author: "github: VincentChuWaiChow"
      version: "0.1.2"
      updated: "2026-05-13"
      category: finops
      lifecycle: experimental
    ---
    
    # Rightsize Recommendation
    
    ## Purpose
    
    Produce CPU and memory rightsizing recommendations for Kubernetes pods based on user-supplied utilization percentile metrics. All math is performed on the inputs provided by the caller; no cluster connection or live metric fetch is performed.
    
    No kubectl. No WebFetch. No cluster credentials accepted.
    
    ## When to use
    
    Use this skill when:
    
    - The user has pasted or attached p50, p95, and p99 CPU and memory utilization metrics for one or more pods or workloads
    - The user wants actionable request and limit recommendations with explicit headroom rationale
    - The user wants an estimated monthly savings figure (requires unit price from the caller)
    - The user wants to know whether a pod is eligible for Karpenter node consolidation
    
    ## Operating rules
    
    - **Inputs only.** All calculations are performed on the metrics and pod specs supplied by the caller. This skill does not fetch live metrics, connect to Prometheus, or query the Kubernetes API.
    - **No credentials accepted.** No cloud credentials, billing account IDs, tenant data, kubeconfig, bearer tokens, or service account JWTs are accepted or required.
    - **Metric window requirement.** The caller must supply metrics from a 7-14 day representative window. Metrics from shorter windows are accepted but the recommendation must be labeled with a reduced-confidence warning.
    - **Provenance labels mandatory.** Every numeric output must carry one label:
      - `inference` — computed from caller-supplied inputs using the documented methodology
      - `assumed` — derived from a default assumption where caller data was missing (state the assumption)
      - `excluded` — cost savings figure excluded because unit price was not supplied by the caller
    - **FOCUS column mapping.** Where a cost savings estimate is produced, note: `BilledCost` (current), `EffectiveCost` (projected after rightsizing), `ChargeCategory = Usage`, `ServiceCategory = Containers`.
    - **Load references only when needed.**
    
    ## Required input
    
    The caller must supply, per pod or workload:
    
    - **Pod name and namespace** (or Deployment/StatefulSet name)
    - **Current CPU request** (millicores) and **current memory request** (MiB)
    - **Current CPU limit** (millicores) and **current memory limit** (MiB); state `no-limit` if unset
    - **p50 CPU** (millicores), **p95 CPU** (millicores), **p99 CPU** (millicores)
    - **p50 memory** (MiB), **p95 memory** (MiB), **p99 memory** (MiB)
    - **Metric window** (days, must be 7-14; shorter accepted with warning)
    - **Unit price per vCPU-hour and per GiB-hour** (optional; required only for $/mo saved estimate)
    
    If any field is missing, ask one clarifying question per gap. Do not fabricate metric values.
    
    ## Recommendation methodology
    
    ### CPU
    
    | Output | Formula | Rationale |
    |---|---|---|
    | Recommended CPU request | p95 CPU + 20% | p95 covers normal burst; 20% headroom absorbs measurement noise and brief spikes |
    | Recommended CPU limit | p99 CPU + 30% | p99 covers rare spikes; 30% headroom prevents throttling during outlier bursts |
    
    ### Memory
    
    | Output | Formula | Rationale |
    |---|---|---|
    | Recommended memory request | p95 memory + 20% | Same headroom logic as CPU; memory pressure leads to OOMKill rather than throttling |
    | Recommended memory limit | p99 memory + 30% | Tighter limits risk OOMKill; 30% buffer is conservative for stateful workloads |
    
    Round all recommendations up to the nearest 10 millicores (CPU) or 16 MiB (memory) for cleaner manifests.
    
    ### Monthly savings estimate
    
    If the caller supplies unit price ($/vCPU-hour and $/GiB-hour):
    
    ```
    cpu_savings_per_month = (current_cpu_request - recommended_cpu_request) / 1000
                           × unit_price_per_vcpu_hour × 730
    
    memory_savings_per_month = (current_memory_request - recommended_memory_request) / 1024
                              × unit_price_per_gib_hour × 730
    ```
    
    Label: `inference`. Note: savings are approximate; they assume the freed capacity is not replaced by new workloads and that the node pool scales down proportionally.
    
    If unit price is not supplied, state the savings formula and mark the $/mo cell as `excluded — unit price not provided by caller`.
    
    ### Karpenter consolidation eligibility
    
    A pod is flagged as consolidation-eligible when all of the following are true:
    
    - No `PodDisruptionBudget` with `maxUnavailable: 0` or `minAvailable: 100%` applies to the pod.
    - No `pod-anti-affinity` or `topologySpreadConstraints` rule that would prevent the pod from co-locating with other pods of the same owner.
    - No `hostPath` or `local` PersistentVolume mount.
    - No `nodeName` node selector pinning the pod to a specific node.
    - The pod's request fits within a standard node SKU that Karpenter would provision (the caller must confirm the NodePool SKU list).
    
    If the caller has not supplied enough information to confirm all five conditions, state which conditions could not be verified and flag as `not-verified — [missing conditions]`. Do not assume eligibility when data is incomplete; unknown blockers present a consolidation risk.
    
    ## Response shape
    
    Return, per pod or workload:
    
    ```
    Pod / Workload: <name> (<namespace>)
    
      CPU
        Current request:     <current> m
        Recommended request: <p95 × 1.20, rounded> m  [inference]
        Current limit:       <current> m  (or "no-limit")
        Recommended limit:   <p99 × 1.30, rounded> m  [inference]
    
      Memory
        Current request:     <current> MiB
        Recommended request: <p95 × 1.20, rounded> MiB  [inference]
        Current limit:       <current> MiB  (or "no-limit")
        Recommended limit:   <p99 × 1.30, rounded> MiB  [inference]
    
      Monthly savings estimate
        CPU:    $<amount>/mo  [inference]  or  [excluded — unit price not provided]
        Memory: $<amount>/mo  [inference]  or  [excluded — unit price not provided]
        Total:  $<amount>/mo
    
      Karpenter consolidation
        Eligible: <Yes | No | Not-verified>
        Blockers: <list any confirmed blockers, or list missing data preventing verification, or "None confirmed">
    
      Metric window: <N days>  [inference: window meets 7-14 day requirement]
      Confidence:   <Normal | Reduced — window < 7 days>
    ```
    
    ## References
    
    Load these only when needed:
    
    - [Metric sources](references/metric-sources.md) — how to gather p50/p95/p99 from Prometheus, Cloud Monitoring, and Azure Monitor for use as input to this skill.
    - [Karpenter consolidation](references/karpenter-consolidation.md) — what makes a pod consolidation-eligible and known blockers.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related