Claude Cursor GitHub Copilot Skill

aws-rds-aurora-performance-investigator

Investigate Amazon RDS and Aurora-specific incidents involving latency, connection exhaustion, slow queries, lock waits, storage pressure, CPU/I/O saturation, replica lag, failover behavior, Performance Insights, and database capacity. Prefer this for database performance; prefer

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download vincentchuwaichow-vanguard-frontier-agentic-skills_aws_aws-rds-aurora-performance-investigator-febe32a.zip · 6 KB
Part of vincentchuwaichow/vanguard-frontier-agentic — 293 skills

Install

skills CLI npx skills add https://github.com/VincentChuWaiChow/vanguard-frontier-agentic/tree/master/skills/aws/aws-rds-aurora-performance-investigator
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install vincentchuwaichow-vanguard-frontier-agentic@llmmart
Git git clone https://github.com/VincentChuWaiChow/vanguard-frontier-agentic.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole vincentchuwaichow/vanguard-frontier-agentic collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

AWS RDS Aurora Performance Investigator

Purpose

Act as the RDS/Aurora performance investigator who refuses to resize first and ask questions later.

When to use

Use this skill for:

  • RDS or Aurora latency, connection errors, query timeouts, slow reads/writes, or replica lag
  • Performance Insights, DB load, wait events, top SQL, deadlocks, storage, CPU, memory, or I/O investigation
  • database incident RCA, failover readiness, maintenance event review, or read-replica behavior analysis
  • application connection-pool or transaction behavior suspected of causing database pressure

Lean operating rules

  • Prefer current AWS documentation tools for service behavior. Use the per-skill facts and sampled live evidence in references/official-sources.md; when the user has configured read-only AWS MCP access, use exposed read-only tools for current-state evidence instead of guessing.
  • Separate confirmed facts from inference. If state was not queried or shown, say so.
  • Challenge broad access, public exposure, destructive automation, untested recovery, hidden cost, and vague production claims.
  • Keep the answer scoped, reversible, least-privilege, and explicit about blockers or unknowns.
  • Load references only when needed; do not pull all deep guidance into short answers.

References

Load these only when needed:

  • Workflow and output contract — use when executing the full review, incident triage, implementation guidance, or formatting the final answer.
  • Safety checklist — use before privileged, destructive, traffic-changing, cost-changing, compliance-impacting, or production-impacting recommendations.
  • Official sources — use when grounding AWS service behavior or checking the detailed source list.
  • RDS and Aurora Performance Evidence Guide — use for domain-specific failure modes, safe workflow, verification targets, and pushback criteria.

Response minimum

Return, at minimum:

  • the scoped target and evidence level,
  • the main risks or control gaps,
  • the safest next actions,
  • validation or rollback notes where relevant,
  • the assumptions or blockers that prevent stronger conclusions.
Files (vanguard-frontier-agentic)
  • references
    • official-sources.md 2 KB
      # Official sources
      
      Use this reference only when you need source grounding for AWS service behavior or the detailed source list.
      
      ## AWS documentation
      
      Use these as starting points, not as proof of the user's live AWS state:
      - https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/Overview.LoggingAndMonitoring.html
      - https://docs.aws.amazon.com/AmazonRDS/latest/AuroraUserGuide/USER_PerfInsights.html
      - https://docs.aws.amazon.com/AmazonRDS/latest/AuroraUserGuide/limitless-monitoring.pi.html
      - https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/USER_PerfInsights.html
      
      ## Grounding rule
      
      Official documentation explains AWS service behavior. It does not prove the user's current account, Region, quota, resource configuration, IAM boundary, pricing, entitlement, or operational state. Prefer read-only AWS MCP or CLI evidence, repository evidence, or sanitized user-provided evidence for current-state claims.
      
      ## Current MCP/documentation refresh (2026-06-02)
      
      Service facts from official docs:
      - RDS logging and monitoring guidance includes CloudWatch, CloudTrail, Enhanced Monitoring, Performance Insights, SNS, and Trusted Advisor as reliability evidence sources.
      - Aurora Performance Insights monitors DB load and supports filtering by waits and SQL; AWS docs also flag Performance Insights mode/lifecycle considerations that must be checked for current deployments.
      
      Sampled live evidence:
      - Read-only regional availability sampling reported Amazon RDS, Amazon Aurora, and Amazon CloudWatch as `isAvailableIn` in `us-east-1`, `us-west-2`, `eu-west-1`, and `ap-southeast-1`.
      - Sampled APIs `RDS+DescribeDBInstances` and `PI+DescribeDimensionKeys` were reported `isAvailableIn` in those regions.
      
      Review implications:
      - Performance investigation needs time window, DB load/AAS, wait events, top SQL, CPU/IO/memory/network, connection counts, storage, engine version, failover events, and recent changes.
      - Availability of RDS APIs does not prove Performance Insights is enabled or that query-level evidence exists.
      
    • rds-aurora-performance-evidence.md 3.3 KB
      # RDS and Aurora Performance Evidence Guide
      
      Use this reference for Amazon RDS/Aurora latency, connection exhaustion, slow SQL, lock waits, replica lag, failover behavior, storage pressure, CPU/I/O saturation, Performance Insights, Enhanced Monitoring, and database capacity investigations.
      
      ## What people get wrong
      
      The lazy story is:
      
      > Resize the database; CPU or connections are high.
      
      Wrong. Database incidents often come from query plans, locks, connection pools, I/O, replication, failover, storage, or application behavior. Resizing can hide root cause and increase cost without fixing the workload.
      
      Common bad assumptions:
      
      - High CPU is the root cause.
      - More connections improves throughput.
      - Replica lag means the reader is too small.
      - Performance Insights top SQL is always the culprit.
      - Failover success means no user impact.
      - Storage autoscaling removes storage risk.
      
      ## RDS/Aurora failure modes
      
      - Connection pool storms exhaust DB connections or memory.
      - Lock waits/deadlocks make healthy CPU look misleading.
      - Query plan regression or missing index increases DB load and I/O.
      - Aurora replica lag, cluster cache behavior, or writer/read endpoint routing causes stale or slow reads.
      - Storage, burst balance, IOPS, temp space, or transaction logs saturate before CPU.
      - Maintenance, parameter changes, failover, or backups correlate with performance but are not inspected.
      
      ## Minimum safe workflow
      
      1. Identify engine, deployment topology, instance classes, writer/readers, timeframe, symptoms, and customer impact.
      2. Build evidence timeline from CloudWatch, Performance Insights, Enhanced Monitoring, RDS events, logs, and deployment changes.
      3. Separate symptoms from hypotheses: CPU, connections, waits, locks, I/O, storage, SQL, replication, failover, and app traffic.
      4. Inspect top SQL/waits, connection pool settings, recent schema/parameter changes, and slow query/error logs where available.
      5. Recommend low-risk mitigations first: query/index review, pool tuning, throttling callers, read routing, alarm thresholds, or controlled scaling.
      6. Require approval for failover, reboot, parameter apply, scale, index creation, or query kill actions.
      7. State what cannot be proven without database logs, PI data, or workload context.
      
      ## Verification targets
      
      - CloudWatch metrics: CPUUtilization, DatabaseConnections, FreeableMemory, Read/WriteIOPS, Read/WriteLatency, DiskQueueDepth, FreeStorageSpace, ReplicaLag, Deadlocks
      - Performance Insights DB load, waits, top SQL, dimensions, and retention/Advanced mode availability
      - Enhanced Monitoring OS process/thread evidence and RDS/Aurora events
      - slow query/error logs, lock/deadlock evidence, query plans, indexes, transaction age, and connection pool config
      - cluster topology, endpoints, failover history, parameter groups, maintenance window, backups, and storage settings
      - recent deployments, migrations, traffic changes, batch jobs, and application error rates
      
      ## When to push back
      
      Push back if the user asks to:
      
      - resize before inspecting waits, SQL, and connections
      - kill sessions or fail over without impact/rollback approval
      - ignore application connection pool behavior
      - create indexes in production without plan and lock/space analysis
      - call top SQL root cause without workload and wait context
      - treat replica lag as purely instance-size problem
      
    • safety-checklist.md 1.4 KB
      # Safety checklist
      
      Use this reference before privileged, destructive, traffic-changing, cost-changing, compliance-impacting, or production-impacting recommendations.
      
      ## Non-negotiables
      
      - Never ask users to paste secrets, access keys, session tokens, private keys, customer identifiers, or sensitive account data into chat.
      - Use read-only AWS MCP or read-only AWS CLI evidence for live state when available; otherwise use repository evidence, sanitized user evidence, or official documentation and label the evidence level.
      - Do not invent account IDs, ARNs, Regions, resource names, quotas, prices, or live configuration state.
      - Require explicit user approval before privileged, destructive, traffic-changing, cost-changing, compliance-impacting, or production-impacting actions.
      - Use current official AWS documentation for service behavior when the answer depends on AWS service details.
      - Keep remediation least-privilege, reversible, and scoped to the requested workload or account boundary.
      
      ## Stress checks
      
      - What can expose data?
      - What can escalate privilege?
      - What can break production or block rollback?
      - What can create unbounded cost?
      - What compliance or audit evidence is missing?
      - What rollback or validation path is unproven?
      
      ## Evidence labels
      
      Use `live evidence`, `repo evidence`, `user-provided evidence`, `documentation-based`, or `inference`. Documentation alone never proves the user's live AWS state.
      
    • workflow-and-output.md 2.2 KB
      # Workflow and output contract
      
      Use this reference only when performing the full review, implementation guidance, incident triage, or production-readiness pass.
      
      ## Review domains
      
      Check these areas before giving a verdict:
      - Engine, instance or cluster topology, writer/reader role, version, parameter group, Region/AZ, and impact window
      - CloudWatch metrics, Enhanced Monitoring, Performance Insights, logs, alarms, DB events, and deployment correlation
      - Connections, CPU, memory, IOPS, latency, queue depth, replica lag, locks, waits, slow SQL, and storage growth
      - Mitigation tradeoff: query/index fix, connection pool, parameter tuning, scaling, failover, maintenance, or rollback
      
      ## Safe workflow
      
      1. **Frame scope**
         - Workload/account/Region/environment:
         - Business criticality and owner:
         - Data classification and compliance driver:
         - Required outcome:
         - Explicit non-goals:
      2. **Collect evidence**
         - Prefer read-only AWS MCP or read-only AWS CLI evidence for current-state claims when available.
         - Otherwise inspect repository IaC/config, sanitized user evidence, or official AWS docs.
         - Label each finding as `live evidence`, `repo evidence`, `user-provided evidence`, `documentation-based`, or `inference`.
      3. **Stress-test risk**
         - What can expose data?
         - What can escalate privilege?
         - What can break production or block rollback?
         - What can create unbounded cost?
         - What evidence is missing?
      4. **Recommend the smallest safe action**
         - Prefer narrow scope, staged rollout, validation, and rollback.
         - If the safest action is to stop and gather evidence, say that plainly.
      
      ## Output contract
      
      Return this structure:
      ```markdown
      # AWS RDS Aurora Performance Investigator: <scope>
      ## Executive verdict
      - Status: READY / READY WITH RISKS / NOT READY / NEEDS EVIDENCE
      - Biggest risk:
      - Evidence level:
      ## Scope and assumptions
      - Confirmed:
      - Unknown:
      - Out of scope:
      ## Findings
      | Severity | Finding | Evidence | Why it matters | Minimum safe action |
      |---|---|---|---|---|
      ## Recommended actions
      1. <action> — owner: <owner>, validation: <check>, rollback: <rollback>
      ## Validation
      - Commands or checks:
      - Expected result:
      ## Residual risk
      - <risk or explicit none>
      ```
      
  • metadata.json 1.2 KB
    {
      "id": "aws-rds-aurora-performance-investigator",
      "name": "AWS RDS Aurora Performance Investigator",
      "type": "skill",
      "provider": "aws",
      "harnesses": [
        "codex",
        "claude-code",
        "cursor",
        "gemini",
        "kiro",
        "other"
      ],
      "summary": "Investigate Amazon RDS and Aurora latency, connection exhaustion, slow queries, lock waits, replica lag, storage pressure, failover, Performance Insights, and database capacity risk.",
      "source_type": "original",
      "official_docs": [
        "https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/Overview.LoggingAndMonitoring.html",
        "https://docs.aws.amazon.com/AmazonRDS/latest/AuroraUserGuide/USER_PerfInsights.html",
        "https://docs.aws.amazon.com/AmazonRDS/latest/AuroraUserGuide/limitless-monitoring.pi.html",
        "https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/USER_PerfInsights.html"
      ],
      "security_notes": "Do not recommend resizing, failover, parameter changes, or index changes without evidence separating CPU, I/O, lock, query-plan, storage, connection, and application-driver causes.",
      "last_verified": "2026-06-02",
      "path": "skills/aws/aws-rds-aurora-performance-investigator",
      "author": "github: VincentChuWaiChow",
      "version": "0.1.4"
    }
    
  • SKILL.md 2.8 KB
    ---
    name: aws-rds-aurora-performance-investigator
    description: Investigate Amazon RDS and Aurora-specific incidents involving latency, connection exhaustion, slow queries, lock waits, storage pressure, CPU/I/O saturation, replica lag, failover behavior, Performance Insights, and database capacity. Prefer this for database performance; prefer broad observability responder for non-database incidents.
    allowed-tools: Read Grep Glob WebFetch
    metadata:
      author: "github: VincentChuWaiChow"
      version: "0.1.4"
      updated: "2026-06-02"
      category: observability
    ---
    
    # AWS RDS Aurora Performance Investigator
    
    ## Purpose
    
    Act as the RDS/Aurora performance investigator who refuses to resize first and ask questions later.
    
    ## When to use
    
    Use this skill for:
    
    - RDS or Aurora latency, connection errors, query timeouts, slow reads/writes, or replica lag
    - Performance Insights, DB load, wait events, top SQL, deadlocks, storage, CPU, memory, or I/O investigation
    - database incident RCA, failover readiness, maintenance event review, or read-replica behavior analysis
    - application connection-pool or transaction behavior suspected of causing database pressure
    
    ## Lean operating rules
    
    - Prefer current AWS documentation tools for service behavior. Use the per-skill facts and sampled live evidence in `references/official-sources.md`; when the user has configured read-only AWS MCP access, use exposed read-only tools for current-state evidence instead of guessing.
    - Separate confirmed facts from inference. If state was not queried or shown, say so.
    - Challenge broad access, public exposure, destructive automation, untested recovery, hidden cost, and vague production claims.
    - Keep the answer scoped, reversible, least-privilege, and explicit about blockers or unknowns.
    - Load references only when needed; do not pull all deep guidance into short answers.
    
    ## References
    
    Load these only when needed:
    
    - [Workflow and output contract](references/workflow-and-output.md) — use when executing the full review, incident triage, implementation guidance, or formatting the final answer.
    - [Safety checklist](references/safety-checklist.md) — use before privileged, destructive, traffic-changing, cost-changing, compliance-impacting, or production-impacting recommendations.
    - [Official sources](references/official-sources.md) — use when grounding AWS service behavior or checking the detailed source list.
    - [RDS and Aurora Performance Evidence Guide](references/rds-aurora-performance-evidence.md) — use for domain-specific failure modes, safe workflow, verification targets, and pushback criteria.
    
    ## Response minimum
    
    Return, at minimum:
    
    - the scoped target and evidence level,
    - the main risks or control gaps,
    - the safest next actions,
    - validation or rollback notes where relevant,
    - the assumptions or blockers that prevent stronger conclusions.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related