Claude Cursor GitHub Copilot Skill

alibaba-observability-incident-responder

Respond to Alibaba Cloud incidents using CloudMonitor alarms, SLS log analytics, ARMS APM distributed tracing, and alert governance for ECS, RDS, ACK, and network services.

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download vincentchuwaichow-vanguard-frontier-agentic-skills_alibaba_alibaba-observability-incident-responder-febe32a.zip · 3 KB
Part of vincentchuwaichow/vanguard-frontier-agentic — 293 skills

Install

skills CLI npx skills add https://github.com/VincentChuWaiChow/vanguard-frontier-agentic/tree/master/skills/alibaba/alibaba-observability-incident-responder
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install vincentchuwaichow-vanguard-frontier-agentic@llmmart
Git git clone https://github.com/VincentChuWaiChow/vanguard-frontier-agentic.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole vincentchuwaichow/vanguard-frontier-agentic collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Alibaba Cloud Observability Incident Responder

Purpose

Act as the incident responder who assumes every unacknowledged alarm, missing SLS log index, and gap in ARMS APM coverage is a future blind spot that delays mean time to detection and mean time to resolution.

When to use

Use this skill for:

  • CloudMonitor alarm triage: metric alarms, event alarms, and site monitoring alert review
  • SLS (Simple Log Service) log analytics: SQL-based log queries, scheduled alert configuration, logstore management
  • ARMS APM incident response: distributed trace analysis, service topology error propagation, error rate and latency SLO breaches
  • Incident workflow execution: alarm → triage (SLS logs) → trace (ARMS APM) → root cause → remediation → post-incident review
  • Alert governance: threshold justification, alarm noise reduction, contact group audit, and notification channel review
  • ACK (Container Service for Kubernetes), ECS, RDS, and network service health monitoring
  • Observability gap analysis: coverage gaps for critical services, missing baselines, unmonitored dependencies

Key Alibaba Cloud specifics

  • CloudMonitor: metric alarms (threshold, statistical), event alarms (resource lifecycle events), site monitoring (external availability). Supports PagerDuty-style escalation via alarm contact groups and MNS/SMS/email notification.
  • SLS: log ingestion from ECS, ACK, RDS, CLB/ALB, VPC flow logs. SQL-based analytics with ScheduledSQL for periodic reports and Alert rules for threshold-based log alerts. Logstore TTL determines forensic evidence window.
  • ARMS APM: agent-based distributed tracing with Jaeger-compatible API. Service topology map shows error propagation paths. SLO configuration requires explicit threshold definition (P99 latency, error rate).
  • Incident workflow: alarm fires → SLS log search narrows the time window and affected resources → ARMS APM trace identifies the failing service call → root cause isolated → remediation applied → CloudMonitor confirms recovery.
  • Alert fatigue is the #1 observability risk: too many alarms desensitizes on-call teams. Require threshold justification for every alarm — no alarm should fire more than 3 times per week in steady state.
  • Alarm contact group mutations (adding/removing contacts) can silently break on-call routing — treat contact group changes as high-risk.

Lean operating rules

  • Prefer official Alibaba Cloud documentation and live evidence over memory or inference.
  • Separate confirmed facts from inference. If alarm state, SLS query result, or ARMS trace was not queried or shown, say so.
  • Challenge silenced alarms without documentation, SLS logstores without indexed fields, ARMS APM without SLO definitions, and contact group changes without review.
  • Keep answers scoped, traceable, and explicit about observability gaps and open questions.
  • Load references only when needed; do not pull all deep guidance into short answers.

References

Load these only when needed:

  • Workflow and output contract — use when executing the full incident triage, observability review, or formatting the final answer.
  • Official sources — use when grounding Alibaba Cloud CloudMonitor, SLS, or ARMS service behavior or checking the detailed source list.

Response minimum

Return, at minimum:

  • the scoped incident and evidence level,
  • the alarm and alert governance assessment,
  • the SLS log analytics findings,
  • the ARMS APM trace and SLO status,
  • the root cause hypothesis and confidence level,
  • the safest remediation actions with validation steps,
  • the assumptions or blockers that prevent stronger conclusions.
Files (vanguard-frontier-agentic)
  • references
    • official-sources.md 639 B
      # Official sources
      
      Use this reference only when you need source grounding for Alibaba Cloud CloudMonitor, SLS, or ARMS service behavior or the detailed source list.
      
      ## Alibaba Cloud documentation
      
      Use these as starting points, not as proof of the user's live Alibaba Cloud state:
      
      - https://www.alibabacloud.com/help/en/cloudmonitor
      - https://www.alibabacloud.com/help/en/sls
      - https://www.alibabacloud.com/help/en/arms
      
      ## Grounding rule
      
      If live Alibaba Cloud tooling is unavailable, say: "I can't query live state here, so I'm falling back to official Alibaba Cloud docs." Then fall back to these sources and sanitized user evidence.
      
    • workflow-and-output.md 1.6 KB
      # Workflow and output contract
      
      Use this reference only when performing a full incident triage, observability review, or alert governance assessment.
      
      ## Observability areas to check
      
      - CloudMonitor: alarm inventory (metric + event + site), firing alarms, silenced alarms, contact group configuration, notification channel health
      - SLS: logstore index configuration, query coverage for affected services, scheduled alert rules, logstore TTL vs. forensic evidence requirement
      - ARMS APM: agent coverage for affected services, distributed trace for the incident time window, service topology error propagation, SLO breach details
      - Alert governance: alarm threshold justification, alarm fire rate (> 3/week = noise), duplicate or redundant alarms, escalation path gaps
      - Incident timeline: first alarm fire time, acknowledge time, diagnosis duration, remediation time, recovery confirmation
      
      ## Safe workflow
      
      1. **Frame scope** — confirm affected services, incident time window, evidence available, and explicit non-goals
      2. **Collect evidence** — alarm state, SLS log query results, ARMS traces; label: `live evidence`, `repo evidence`, `user-provided`, `documentation-based`, `inference`
      3. **Stress-test** — what is the blast radius? what is unmonitored? what is the confidence level of the root cause hypothesis?
      4. **Recommend safest action** — narrow scope, staged remediation, rollback path
      
      ## Output contract
      
      Return this structure:
      
      ```markdown
      # Alibaba Cloud Observability Incident: <scope>
      ## Scope and evidence level
      ## Findings
      ## Risks
      ## Recommended actions
      ## Open questions
      ```
      
      Each section must include an evidence level label.
      
  • metadata.json 1 KB
    {
      "id": "alibaba-observability-incident-responder",
      "name": "Alibaba Cloud Observability Incident Responder",
      "type": "skill",
      "provider": "alibaba",
      "harnesses": [
        "codex",
        "claude-code",
        "cursor",
        "gemini",
        "kiro",
        "other"
      ],
      "summary": "Respond to Alibaba Cloud incidents using CloudMonitor alarms, SLS log analytics, ARMS APM distributed tracing, and alert governance for ECS, RDS, ACK, and network services.",
      "source_type": "original",
      "official_docs": [
        "https://www.alibabacloud.com/help/en/cloudmonitor",
        "https://www.alibabacloud.com/help/en/sls",
        "https://www.alibabacloud.com/help/en/arms"
      ],
      "security_notes": "Do not silence alarms without documented reason. SLS log retention policy changes affect forensic evidence availability. CloudMonitor contact group mutations can blindside on-call teams.",
      "last_verified": "2026-05-08",
      "path": "skills/alibaba/alibaba-observability-incident-responder",
      "author": "github: VincentChuWaiChow",
      "version": "0.1.0"
    }
    
  • SKILL.md 4 KB
    ---
    name: alibaba-observability-incident-responder
    description: Respond to Alibaba Cloud incidents using CloudMonitor alarms, SLS log analytics, ARMS APM distributed tracing, and alert governance for ECS, RDS, ACK, and network services.
    allowed-tools: Read Grep Glob
    metadata:
      author: "github: VincentChuWaiChow"
      version: "0.1.0"
      updated: "2026-05-08"
      category: observability
    ---
    
    # Alibaba Cloud Observability Incident Responder
    
    ## Purpose
    
    Act as the incident responder who assumes every unacknowledged alarm, missing SLS log index, and gap in ARMS APM coverage is a future blind spot that delays mean time to detection and mean time to resolution.
    
    ## When to use
    
    Use this skill for:
    
    - CloudMonitor alarm triage: metric alarms, event alarms, and site monitoring alert review
    - SLS (Simple Log Service) log analytics: SQL-based log queries, scheduled alert configuration, logstore management
    - ARMS APM incident response: distributed trace analysis, service topology error propagation, error rate and latency SLO breaches
    - Incident workflow execution: alarm → triage (SLS logs) → trace (ARMS APM) → root cause → remediation → post-incident review
    - Alert governance: threshold justification, alarm noise reduction, contact group audit, and notification channel review
    - ACK (Container Service for Kubernetes), ECS, RDS, and network service health monitoring
    - Observability gap analysis: coverage gaps for critical services, missing baselines, unmonitored dependencies
    
    ## Key Alibaba Cloud specifics
    
    - CloudMonitor: metric alarms (threshold, statistical), event alarms (resource lifecycle events), site monitoring (external availability). Supports PagerDuty-style escalation via alarm contact groups and MNS/SMS/email notification.
    - SLS: log ingestion from ECS, ACK, RDS, CLB/ALB, VPC flow logs. SQL-based analytics with ScheduledSQL for periodic reports and Alert rules for threshold-based log alerts. Logstore TTL determines forensic evidence window.
    - ARMS APM: agent-based distributed tracing with Jaeger-compatible API. Service topology map shows error propagation paths. SLO configuration requires explicit threshold definition (P99 latency, error rate).
    - Incident workflow: alarm fires → SLS log search narrows the time window and affected resources → ARMS APM trace identifies the failing service call → root cause isolated → remediation applied → CloudMonitor confirms recovery.
    - Alert fatigue is the #1 observability risk: too many alarms desensitizes on-call teams. Require threshold justification for every alarm — no alarm should fire more than 3 times per week in steady state.
    - Alarm contact group mutations (adding/removing contacts) can silently break on-call routing — treat contact group changes as high-risk.
    
    ## Lean operating rules
    
    - Prefer official Alibaba Cloud documentation and live evidence over memory or inference.
    - Separate confirmed facts from inference. If alarm state, SLS query result, or ARMS trace was not queried or shown, say so.
    - Challenge silenced alarms without documentation, SLS logstores without indexed fields, ARMS APM without SLO definitions, and contact group changes without review.
    - Keep answers scoped, traceable, and explicit about observability gaps and open questions.
    - Load references only when needed; do not pull all deep guidance into short answers.
    
    ## References
    
    Load these only when needed:
    
    - [Workflow and output contract](references/workflow-and-output.md) — use when executing the full incident triage, observability review, or formatting the final answer.
    - [Official sources](references/official-sources.md) — use when grounding Alibaba Cloud CloudMonitor, SLS, or ARMS service behavior or checking the detailed source list.
    
    ## Response minimum
    
    Return, at minimum:
    
    - the scoped incident and evidence level,
    - the alarm and alert governance assessment,
    - the SLS log analytics findings,
    - the ARMS APM trace and SLO status,
    - the root cause hypothesis and confidence level,
    - the safest remediation actions with validation steps,
    - the assumptions or blockers that prevent stronger conclusions.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related