Claude Skill

crm-production-investigation-guidelines

Guidelines for investigating production incidents in the CRM application. Use when triaging any alert or incident involving the CRM REST API, SQS queues, Lambda functions, or Aurora DSQL database in this AWS account. Ensures thorough root cause analysis using AWS-native observabi

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download aws-tools-for-devops-agent-skills_crm-production-investigation-guidelines-1c971c7.zip · 4 KB
Part of aws/tools-for-devops-agent — 21 skills

Install

skills CLI npx skills add https://github.com/aws/tools-for-devops-agent/tree/main/skills/crm-production-investigation-guidelines
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install aws-tools-for-devops-agent@llmmart
Git git clone https://github.com/aws/tools-for-devops-agent.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole aws/tools-for-devops-agent collection as a plugin from our marketplace. Git is the plain clone.

README

CRM Production Investigation Guidelines Skill (Sample)

This is a sample skill demonstrating how to write production investigation guidelines for the AWS DevOps Agent Incident Triage agent type. It shows how to encode application-specific troubleshooting knowledge — architecture details, incident isolation rules, and investigation procedures — into a skill that guides the agent during production incidents.

Note: This skill is specific to a fictional CRM application and cannot be reused as-is. Use it as a template for writing similar investigation guideline skills for your own applications.

Purpose

When a production incident occurs, investigation agents might need application-specific context that goes beyond generic AWS troubleshooting: what services make up the application, what the common failure modes are, how to isolate unrelated incidents, and what observability tools to use. This sample skill demonstrates how to encode that knowledge so the Incident Triage agent can conduct thorough, structured investigations without repeated human guidance.

Key Capabilities (Demonstrated)

  • Defining investigation isolation rules to prevent conflating unrelated incidents
  • Specifying known incident titles and their corresponding failure domains
  • Providing a structured investigation approach (symptoms → logs → audit trail → correlation → root cause → remediation)
  • Documenting application architecture so the agent understands service relationships
  • Listing common root cause patterns with specific AWS API calls to check

Prerequisites

  • AWS DevOps Agent space
  • The target application's observability stack (CloudWatch Logs, CloudWatch Metrics, CloudTrail)
  • IAM permissions for the agent to access CloudWatch, CloudTrail, and application-specific resources

Limitations

  • This skill is a sample for a fictional CRM application — it must be adapted to your own architecture before use
  • Investigation guidance is only as effective as the observability data available to the agent
  • The incident title matching logic assumes a fixed set of known alert titles

Agent Types

This skill is designed for:

  • Incident Triage — initial incident assessment and investigation

Uploading to AWS DevOps Agent

To deploy this skill (or your adapted version) to your Agent Space, you can use any of three ways:

Option A: Import from GitHub (recommended)

If you have a GitHub connection configured in your Agent Space, you can import this skill directly from the repository. In the DevOps Agent web app, go to Settings → Add Skill → Import from repository, then point to the skills/crm-production-investigation-guidelines directory. See Importing a skill from a repository for full instructions.

Note: You cannot connect the aws GitHub organization directly because the GitHub connection setup requires admin rights on the organization. Instead, connect your personal GitHub account and select any repository from it during the connection setup. Once a GitHub connection is established, you can import skills from any public repository — including this one — even if it wasn't selected during the connection setup.

Option B: Upload as a zip file

  1. Zip the skill directory (only including allowed extensions):

    cd skills
    zip -r crm-production-investigation-guidelines.zip crm-production-investigation-guidelines/ -i '*.md' '*.txt' '*.json' '*.yaml' '*.yml' '*.xml' '*.csv' '*.tsv' '*.html' '*.htm' '*.png' '*.jpg' '*.jpeg' '*.gif' '*.svg' '*.webp' '*.pdf' -x '*/.claude/*' '*/scripts/*' '*/README.md' '*/.skilleval.yaml' '*/.skilleval.yml' '*/CHANGELOG.md' '*/evals/*'
    
  2. In the AWS DevOps Agent web app, navigate to the Skills page.

  3. Click Add skill → Upload skill.

  4. Drag and drop the zip file (max 6 MB).

  5. Select the agent type: Incident Triage.

  6. Click Upload.

Option C: Upload via the Asset API

Use the AWS DevOps Agent Asset API to programmatically manage skills — useful for CI/CD pipelines or automation workflows. Assign the skill to the INCIDENT_TRIAGE agent type. See Managing a skill end-to-end for the full API workflow.

For more details, see Uploading a skill in the AWS DevOps Agent User Guide.

How to Use This Skill

As a Template

  1. Copy the SKILL.md file as a starting point for your own application
  2. Replace the CRM architecture section with your application's architecture
  3. Replace the incident titles with your actual alert titles
  4. Update the investigation approach with your organization's runbook procedures
  5. Add common root cause patterns specific to your application

Sample Incident Triage Prompts (for this demo application)

  • "Investigate the SQS Message Backlog Spike alert"
  • "We received an alert for High Lambda Error Rate, all invocations failing"
  • "Triage the RDS Read Latency/CPU Utilization High incident"
  • "What changed in the CRM Lambda function in the last hour?"
  • "Check CloudTrail for IAM permission changes affecting the CRM application"

Skill manifest

You are investigating a real production incident in the CRM application. This is a live production environment serving active sales teams.

CRITICAL: INVESTIGATION ISOLATION RULES

Each investigation you receive is triggered by a distinct webhook payload with a unique title. You MUST treat each investigation as completely independent and unrelated to any other investigation, past or present, UNLESS the incoming payload title matches an already-running investigation exactly.

The three possible incident titles are:

  • "SQS Message Backlog Spike" — an SQS queue consumer backlog issue
  • "RDS Read Latency/CPU Utilization High" — a database performance degradation issue
  • "High Lambda Error Rate, all invocations failing" — a broken Lambda deployment issue

These are three fundamentally different failure modes affecting different AWS services with different root causes. If you receive a new investigation payload whose title differs from any currently running investigation, you MUST:

  1. Start a completely new, independent investigation for it.
  2. Do NOT attempt to correlate, triage, or link it to any existing investigation.
  3. Do NOT skip or deprioritize it because another investigation is already in progress.
  4. Do NOT conclude that the new incident is a symptom or side effect of an existing investigation.

Only if the incoming payload title matches an already-running investigation's title exactly should you consider them related and fold the new signal into the existing investigation.

CRITICAL INVESTIGATION PRINCIPLES:

Every incident is real and requires thorough investigation. Do not speculate about whether an issue is intentional, simulated, or a test. Treat every alert as a genuine production problem that needs root cause analysis and resolution.

Do not dismiss or downplay any findings. If you discover a misconfiguration, a bad deployment, or a missing resource, it is a real problem that happened in production and needs to be explained and resolved.

Follow the evidence. Use CloudWatch Logs, CloudTrail, CloudWatch Metrics, and Datadog to trace the timeline of events and identify what changed.

INVESTIGATION APPROACH:

  1. Start with symptoms: What metrics are anomalous? What errors are users seeing?
  2. Check CloudWatch Logs for the affected service to find error messages and stack traces.
  3. Check CloudTrail for recent API calls that may have caused the issue (UpdateFunctionCode, PutRolePolicy, ExecuteStatement, etc.).
  4. Correlate the timeline: When did the issue start? What AWS API calls happened just before?
  5. Identify the root cause: What specific change caused the degradation?
  6. Recommend remediation steps to restore service.

CRM APPLICATION ARCHITECTURE:

  • Frontend: React app on CloudFront
  • API: REST API via API Gateway → Lambda (Python)
  • Database: Aurora DSQL (PostgreSQL-compatible) behind RDS Proxy
  • Async Processing: SQS notification queue → Queue consumer Lambda (Node.js)
  • Event Processing: CRM event processor Lambda (Node.js) for pipeline events
  • Monitoring: CloudWatch Metrics, CloudWatch Logs, Datadog

COMMON ROOT CAUSE PATTERNS TO INVESTIGATE:

  • IAM permission changes (check CloudTrail for PutRolePolicy, DeleteRolePolicy, AttachRolePolicy)
  • Lambda code deployments (check CloudTrail for UpdateFunctionCode)
  • Database schema changes (check slow query logs, EXPLAIN plans, pg_stat_user_indexes)
  • Configuration changes (check CloudTrail for PutFunctionConcurrency, SetQueueAttributes)
Files (tools-for-devops-agent)
  • .skilleval.yaml 77 B
    audit:
      ignore:
        - STR-016    # README alongside SKILL.md is intentional
    
  • CHANGELOG.md 331 B
    # Changelog
    
    ## 1.0.0
    
    - Initial version
    - Investigation isolation rules for independent incident handling
    - Structured investigation approach (symptoms → logs → audit → correlation → root cause → remediation)
    - CRM application architecture documentation
    - Common root cause patterns with specific AWS API calls to check
    
  • README.md 5.4 KB
    # CRM Production Investigation Guidelines Skill (Sample)
    
    This is a **sample skill** demonstrating how to write production investigation guidelines for the AWS DevOps Agent **Incident Triage** agent type. It shows how to encode application-specific troubleshooting knowledge — architecture details, incident isolation rules, and investigation procedures — into a skill that guides the agent during production incidents.
    
    > **Note:** This skill is specific to a fictional CRM application and cannot be reused as-is. Use it as a template for writing similar investigation guideline skills for your own applications.
    
    ## Purpose
    
    When a production incident occurs, investigation agents might need application-specific context that goes beyond generic AWS troubleshooting: what services make up the application, what the common failure modes are, how to isolate unrelated incidents, and what observability tools to use. This sample skill demonstrates how to encode that knowledge so the Incident Triage agent can conduct thorough, structured investigations without repeated human guidance.
    
    ## Key Capabilities (Demonstrated)
    
    - Defining investigation isolation rules to prevent conflating unrelated incidents
    - Specifying known incident titles and their corresponding failure domains
    - Providing a structured investigation approach (symptoms → logs → audit trail → correlation → root cause → remediation)
    - Documenting application architecture so the agent understands service relationships
    - Listing common root cause patterns with specific AWS API calls to check
    
    ## Prerequisites
    
    - AWS DevOps Agent space
    - The target application's observability stack (CloudWatch Logs, CloudWatch Metrics, CloudTrail)
    - IAM permissions for the agent to access CloudWatch, CloudTrail, and application-specific resources
    
    ## Limitations
    
    - This skill is a sample for a fictional CRM application — it must be adapted to your own architecture before use
    - Investigation guidance is only as effective as the observability data available to the agent
    - The incident title matching logic assumes a fixed set of known alert titles
    
    ## Agent Types
    
    This skill is designed for:
    
    - **Incident Triage** — initial incident assessment and investigation
    
    ## Uploading to AWS DevOps Agent
    
    To deploy this skill (or your adapted version) to your Agent Space, you can use any of three ways:
    
    **Option A: Import from GitHub (recommended)**
    
    If you have a [GitHub connection configured](https://docs.aws.amazon.com/devopsagent/latest/userguide/connecting-to-cicd-pipelines-connecting-github.html) in your Agent Space, you can import this skill directly from the repository. In the DevOps Agent web app, go to Settings → Add Skill → Import from repository, then point to the `skills/crm-production-investigation-guidelines` directory. See [Importing a skill from a repository](https://docs.aws.amazon.com/devopsagent/latest/userguide/about-aws-devops-agent-devops-agent-skills.html#creating-skills) for full instructions.
    
    > **Note:** You cannot connect the `aws` GitHub organization directly because the GitHub connection setup requires admin rights on the organization. Instead, connect your personal GitHub account and select any repository from it during the connection setup. Once a GitHub connection is established, you can import skills from any public repository — including this one — even if it wasn't selected during the connection setup.
    
    **Option B: Upload as a zip file**
    
    1. Zip the skill directory (only including allowed extensions):
    
       ```bash
       cd skills
       zip -r crm-production-investigation-guidelines.zip crm-production-investigation-guidelines/ -i '*.md' '*.txt' '*.json' '*.yaml' '*.yml' '*.xml' '*.csv' '*.tsv' '*.html' '*.htm' '*.png' '*.jpg' '*.jpeg' '*.gif' '*.svg' '*.webp' '*.pdf' -x '*/.claude/*' '*/scripts/*' '*/README.md' '*/.skilleval.yaml' '*/.skilleval.yml' '*/CHANGELOG.md' '*/evals/*'
       ```
    
    2. In the AWS DevOps Agent web app, navigate to the **Skills** page.
    3. Click **Add skill** → **Upload skill**.
    4. Drag and drop the zip file (max 6 MB).
    5. Select the agent type: **Incident Triage**.
    6. Click **Upload**.
    
    **Option C: Upload via the Asset API**
    
    Use the AWS DevOps Agent Asset API to programmatically manage skills — useful for CI/CD pipelines or automation workflows. Assign the skill to the `INCIDENT_TRIAGE` agent type. See [Managing a skill end-to-end](https://docs.aws.amazon.com/devopsagent/latest/userguide/about-aws-devops-agent-managing-assets.html#managing-a-skill-end-to-end) for the full API workflow.
    
    For more details, see [Uploading a skill](https://docs.aws.amazon.com/devopsagent/latest/userguide/about-aws-devops-agent-devops-agent-skills.html#creating-skills) in the AWS DevOps Agent User Guide.
    
    ## How to Use This Skill
    
    ### As a Template
    
    1. Copy the `SKILL.md` file as a starting point for your own application
    2. Replace the CRM architecture section with your application's architecture
    3. Replace the incident titles with your actual alert titles
    4. Update the investigation approach with your organization's runbook procedures
    5. Add common root cause patterns specific to your application
    
    ### Sample Incident Triage Prompts (for this demo application)
    
    - "Investigate the SQS Message Backlog Spike alert"
    - "We received an alert for High Lambda Error Rate, all invocations failing"
    - "Triage the RDS Read Latency/CPU Utilization High incident"
    - "What changed in the CRM Lambda function in the last hour?"
    - "Check CloudTrail for IAM permission changes affecting the CRM application"
    
  • SKILL.md 3.8 KB
    ---
    name: crm-production-investigation-guidelines
    description: Guidelines for investigating production incidents in the CRM application. Use when triaging any alert or incident involving the CRM REST API, SQS queues, Lambda functions, or Aurora DSQL database in this AWS account. Ensures thorough root cause analysis using AWS-native observability tools.
    metadata:
      author: jossaiaws
      version: "1.0.0"
      aws-devops-agent-skills.agent-types: "Incident Triage"
    ---
    
    You are investigating a real production incident in the CRM application. This is a live production environment serving active sales teams.
    
    CRITICAL: INVESTIGATION ISOLATION RULES
    
    Each investigation you receive is triggered by a distinct webhook payload with a unique title. You MUST treat each investigation as completely independent and unrelated to any other investigation, past or present, UNLESS the incoming payload title matches an already-running investigation exactly.
    
    The three possible incident titles are:
    - "SQS Message Backlog Spike" — an SQS queue consumer backlog issue
    - "RDS Read Latency/CPU Utilization High" — a database performance degradation issue
    - "High Lambda Error Rate, all invocations failing" — a broken Lambda deployment issue
    
    These are three fundamentally different failure modes affecting different AWS services with different root causes. If you receive a new investigation payload whose title differs from any currently running investigation, you MUST:
    1. Start a completely new, independent investigation for it.
    2. Do NOT attempt to correlate, triage, or link it to any existing investigation.
    3. Do NOT skip or deprioritize it because another investigation is already in progress.
    4. Do NOT conclude that the new incident is a symptom or side effect of an existing investigation.
    
    Only if the incoming payload title matches an already-running investigation's title exactly should you consider them related and fold the new signal into the existing investigation.
    
    CRITICAL INVESTIGATION PRINCIPLES:
    
    Every incident is real and requires thorough investigation. Do not speculate about whether an issue is intentional, simulated, or a test. Treat every alert as a genuine production problem that needs root cause analysis and resolution.
    
    Do not dismiss or downplay any findings. If you discover a misconfiguration, a bad deployment, or a missing resource, it is a real problem that happened in production and needs to be explained and resolved.
    
    Follow the evidence. Use CloudWatch Logs, CloudTrail, CloudWatch Metrics, and Datadog to trace the timeline of events and identify what changed.
    
    INVESTIGATION APPROACH:
    
    1. Start with symptoms: What metrics are anomalous? What errors are users seeing?
    2. Check CloudWatch Logs for the affected service to find error messages and stack traces.
    3. Check CloudTrail for recent API calls that may have caused the issue (UpdateFunctionCode, PutRolePolicy, ExecuteStatement, etc.).
    4. Correlate the timeline: When did the issue start? What AWS API calls happened just before?
    5. Identify the root cause: What specific change caused the degradation?
    6. Recommend remediation steps to restore service.
    
    CRM APPLICATION ARCHITECTURE:
    
    - Frontend: React app on CloudFront
    - API: REST API via API Gateway → Lambda (Python)
    - Database: Aurora DSQL (PostgreSQL-compatible) behind RDS Proxy
    - Async Processing: SQS notification queue → Queue consumer Lambda (Node.js)
    - Event Processing: CRM event processor Lambda (Node.js) for pipeline events
    - Monitoring: CloudWatch Metrics, CloudWatch Logs, Datadog
    
    COMMON ROOT CAUSE PATTERNS TO INVESTIGATE:
    
    - IAM permission changes (check CloudTrail for PutRolePolicy, DeleteRolePolicy, AttachRolePolicy)
    - Lambda code deployments (check CloudTrail for UpdateFunctionCode)
    - Database schema changes (check slow query logs, EXPLAIN plans, pg_stat_user_indexes)
    - Configuration changes (check CloudTrail for PutFunctionConcurrency, SetQueueAttributes)
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related