Claude Cursor GitHub Copilot Skill

azure-observability-investigator

Use this skill for Azure Monitor, Log Analytics, Application Insights, alerting, KQL triage, telemetry-gap analysis, workbooks, or operator-grade incident and posture investigations.

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download vincentchuwaichow-vanguard-frontier-agentic-skills_azure_azure-observability-investigator-febe32a.zip · 7 KB
Part of vincentchuwaichow/vanguard-frontier-agentic — 293 skills

Install

skills CLI npx skills add https://github.com/VincentChuWaiChow/vanguard-frontier-agentic/tree/master/skills/azure/azure-observability-investigator
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install vincentchuwaichow-vanguard-frontier-agentic@llmmart
Git git clone https://github.com/VincentChuWaiChow/vanguard-frontier-agentic.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole vincentchuwaichow/vanguard-frontier-agentic collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Azure Observability Investigator

Purpose

Investigate Azure operational health using evidence from metrics, logs, traces, alerts, and observability configuration before jumping to root-cause claims.

This skill is for operator-grade Azure monitoring work across:

  • Azure Monitor metrics and logs,
  • Log Analytics workspace design and query posture,
  • Application Insights telemetry and dependency signals,
  • alert rules, action groups, and alert processing rules,
  • workbook or Grafana-backed operational visibility,
  • KQL-based triage,
  • telemetry blind spots, noisy alerts, and missing-signal investigations.

When to use

Use this skill when the user asks for:

  • Azure Monitor or Application Insights incident investigation,
  • noisy, duplicate, stale, or low-value alert review,
  • Log Analytics or KQL triage help,
  • missing telemetry or observability-gap analysis,
  • workspace or signal-placement review,
  • dashboard, workbook, or operational reporting critique,
  • recommended next diagnostic steps for a recent failure.

Do not use this skill as a substitute for:

  • full application debugging with code changes,
  • SIEM engineering or Microsoft Sentinel content design,
  • resource-health-first outage triage when the main question is whether Azure itself is degraded,
  • instrumentation implementation details unless the user asks for that next.

Lean operating rules

  • Prefer Microsoft Learn documentation through the user's configured documentation MCP, then sampled read-only Azure evidence when available, then sanitized user evidence.
  • Separate confirmed facts from inference. If state was not queried or shown, say so.
  • Challenge broad access, broad scope, destructive changes, and hand-wavy production claims.
  • Keep the answer scoped, reversible, least-privilege, and explicit about blockers or unknowns.

References

Load these only when needed:

  • Azure Observability Investigation Operations — use for current service behavior, common failure modes, hard design rules, verification targets, and push-back conditions.
  • Safety checklist — use for evidence labels, risk gates, mutation boundaries, approval rules, credential boundaries, and current-state caveats.
  • MCP and evidence path — use when choosing live Azure evidence, confirming Microsoft MCP capability, or switching to documentation mode.
  • Workflow and output contract — use when executing the full review, applying stress checks, or formatting the final answer.
  • Official sources — use when you need the detailed Microsoft documentation list or source notes.

Response minimum

Return, at minimum:

  • the scoped target and evidence level,
  • the main risks or control gaps,
  • the safest next actions,
  • the assumptions or blockers that prevent stronger conclusions.
Files (vanguard-frontier-agentic)
  • references
    • mcp-and-evidence.md 1.2 KB
      # Documentation and Evidence Path
      
      ## Preferred evidence order
      
      1. Microsoft Learn documentation through the user's configured documentation MCP for documented Azure behavior.
      2. Sampled read-only Azure evidence, when safely available, for current configured-environment observations.
      3. Sanitized user-provided evidence.
      4. Clearly labeled inference.
      
      ## What each evidence type can prove
      
      - Microsoft Learn documentation can prove documented service behavior, supported concepts, limitations, and recommended patterns.
      - Sampled read-only evidence can prove the sampled configured state at the time observed.
      - Sanitized user evidence can prove only what the snippet shows.
      - None of these alone prove broad regional availability, future success, full account posture, or production readiness.
      
      ## Safe usage pattern
      
      - State whether each claim is documentation-based, sampled-current-state, user-provided, or inference.
      - Use read-only queries before recommending changes.
      - Do not include sensitive internal identifiers, tenant identifiers, subscription identifiers, or secrets in committed docs or final findings.
      - If no sampled evidence is available, say the review is documentation-based and list the exact evidence still needed.
      
    • observability-investigation-operations.md 4.6 KB
      # Azure Observability Investigation Operations
      
      Use this reference for current, source-grounded service behavior and the hard review gates that the lean `SKILL.md` intentionally does not carry.
      
      ## What people get wrong
      
      - Calling the first correlated symptom the root cause.
      - Ignoring missing telemetry or sampling gaps.
      - Creating noisy alerts without action owner, severity, and suppression plan.
      - Querying the wrong workspace, time range, or resource scope.
      - Using dashboards as proof that alerting and incident response work.
      
      ## Officially grounded service shape
      
      Microsoft Learn evidence says Azure Monitor collects, analyzes, and acts on logs, metrics, traces, and events across cloud and hybrid environments. Log Analytics workspaces store log and trace data queried with KQL, Azure Monitor workspaces store Prometheus/OpenTelemetry metrics, Application Insights provides application performance monitoring, and alerts/action groups support proactive response. Workbooks and Grafana visualize signals, but alerting should use Azure Monitor native alerts for Azure Monitor services.
      
      - Azure Monitor is the unified observability service for metrics, logs, traces, and events.
      - Metrics, logs, traces, activity logs, resource health, and service health answer different questions.
      - Log Analytics uses KQL and workspace scope; workspace design affects access, cost, retention, and query correctness.
      - Application Insights supports OpenTelemetry-based app monitoring, dependency maps, live metrics, failures, performance, and availability.
      - Alerts need action groups, routing, severity, suppression/processing rules, and ownership.
      
      ## Non-negotiable design rules
      
      - Separate observation, hypothesis, evidence, and inference.
      - State time range, scope, query, signal source, and sampling caveat for every finding.
      - Prefer narrow diagnostic changes before broad alert rewrites.
      - Validate alert routing and action groups, not just alert rule existence.
      - Protect sensitive telemetry and do not paste secrets or customer data from logs.
      
      ## Minimal safe implementation flow
      
      - Scope incident/resource/workload, time window, user impact, and available telemetry.
      - Inventory metrics, logs, traces, alerts, action groups, dashboards, diagnostic settings, and workspace coverage.
      - Run focused KQL/metric checks or request sampled evidence where tools are unavailable.
      - Correlate signals and classify root cause, contributing factor, symptom, or unknown.
      - Return findings, confidence, telemetry gaps, next diagnostics, and safe remediation.
      
      ## High-risk assumptions to kill
      
      - The first correlated alert is not the root cause; root cause needs time-bounded, scoped evidence across metrics, logs, traces, events, and changes.
      - A dashboard is not proof that telemetry, alerting, action routing, or incident response works.
      - Missing telemetry is a finding. Do not invent certainty when diagnostic settings, workspace coverage, sampling, or retention are incomplete.
      - Log Analytics workspaces and Azure Monitor workspaces are different stores with different query models; querying the wrong store invalidates conclusions.
      - Alert changes can hide real incidents; suppression, thresholds, and action group changes require ownership and rollback.
      
      ## Safe command/code verification targets
      
      - Record resource scope, workspace, time range, query text, signal source, sampling status, and data freshness for every finding.
      - Verify diagnostic settings, data collection rules, workspace retention, Application Insights/OpenTelemetry coverage, and alert/action group routes.
      - Separate observations, hypotheses, inferred causes, confirmed causes, and telemetry gaps in the final output.
      - Test alert routing or action groups with safe evidence before claiming response readiness.
      - Redact secrets, customer identifiers, and sensitive payloads from logs before storing or sharing evidence.
      
      ## Safe verification targets
      
      - Diagnostic settings send required platform logs and metrics to the intended destination.
      - Workspace scope, retention, access model, and query performance fit the investigation.
      - Application Insights or OpenTelemetry captures requests, dependencies, exceptions, traces, availability, and sampling status where relevant.
      - Alerts have owner, severity, threshold rationale, action group route, and noise controls.
      - Workbooks/Grafana dashboards reflect authoritative signals and are not the only evidence.
      
      ## When to push back
      
      - Telemetry is missing for the claimed root cause.
      - The time window or resource scope is ambiguous.
      - Alert changes would silence production incidents without owner approval.
      - Logs contain secrets or customer data that should be redacted before sharing.
      
    • official-sources.md 2.4 KB
      # Official Sources
      
      Use these sources to ground the skill. Microsoft Learn documentation proves documented Azure behavior; it does not prove the user's tenant, subscription, RBAC, quota, migration project, network, telemetry, deployed resources, or production readiness.
      
      ## Primary Microsoft Learn sources
      
      - https://learn.microsoft.com/azure/azure-monitor/fundamentals/overview
      - https://learn.microsoft.com/azure/azure-monitor/fundamentals/best-practices-operation
      - https://learn.microsoft.com/azure/azure-monitor/alerts/alerts-overview
      - https://learn.microsoft.com/azure/azure-monitor/alerts/action-groups
      - https://learn.microsoft.com/azure/azure-monitor/alerts/alerts-processing-rules
      - https://learn.microsoft.com/azure/azure-monitor/logs/log-analytics-overview
      - https://learn.microsoft.com/azure/azure-monitor/logs/workspace-design
      - https://learn.microsoft.com/azure/azure-monitor/app/app-insights-overview
      - https://learn.microsoft.com/azure/azure-monitor/visualize/workbooks-overview
      - https://learn.microsoft.com/azure/managed-grafana/how-to-use-azure-monitor-alerts
      
      ## Grounding notes
      
      - Documentation-based claim: Microsoft Learn evidence says Azure Monitor collects, analyzes, and acts on logs, metrics, traces, and events across cloud and hybrid environments. Log Analytics workspaces store log and trace data queried with KQL, Azure Monitor workspaces store Prometheus/OpenTelemetry metrics, Application Insights provides application performance monitoring, and alerts/action groups support proactive response. Workbooks and Grafana visualize signals, but alerting should use Azure Monitor native alerts for Azure Monitor services.
      - Current-state claim: requires sampled read-only Azure evidence or sanitized user-provided evidence.
      - Inference: allowed only when labeled and tied to observed fields or documented behavior.
      - Do not include sensitive internal identifiers or secret material in findings.
      
      ## Source use rules
      
      - Prefer Microsoft Learn documentation through the user's configured documentation MCP for current Azure service behavior.
      - Use sampled read-only Azure evidence only to validate current configured-environment observations.
      - If documentation and sampled evidence appear to conflict, report both and stop short of a production-ready verdict.
      - Re-check official sources before changing high-risk guidance, because cloud behavior and feature availability can change.
      
    • safety-checklist.md 1.6 KB
      # Safety Checklist
      
      ## Evidence labels
      
      - `documentation-based`: grounded in Microsoft Learn or listed official documentation.
      - `sampled-current-state`: grounded in read-only Azure observations from the user's configured tools.
      - `user-provided`: grounded in sanitized snippets supplied by the user.
      - `inference`: reasoned from evidence but not directly proven.
      
      ## Mutation boundary
      
      - Default to read-only review.
      - Do not perform create, update, delete, activate, approve, cancel, deactivate, migrate, cut over, route, peer, deploy, alert, suppress, or configuration changes unless the user explicitly asks and approval is clear.
      - Prefer preview, assessment, status, list, show, query, activity-log, dependency, and diagnostic evidence before any mutation.
      
      ## Credential and data boundary
      
      - Never ask users to paste credentials, tokens, tenant IDs, subscription IDs, customer data, private keys, appliance secrets, migration inventory dumps, log payload secrets, or raw environment dumps.
      - Summarize sensitive evidence by field presence, control state, and risk; do not reproduce secret material.
      
      ## Risk gates
      
      - Stop on ambiguous target, ambiguous principal, missing approval, missing owner, missing rollback, stale assessment, incomplete dependency mapping, unclear routing, or missing telemetry for high-impact assets.
      - Separate documented product behavior from sampled configured-environment evidence.
      
      ## Asset-specific hard line
      
      Do not over-attribute symptoms as root cause, ignore missing telemetry, or recommend broad alerting changes without signal-quality review, routing checks, query scope, and bounded verification steps.
      
    • workflow-and-output.md 1.6 KB
      # Workflow and Output Contract
      
      ## Execution flow
      
      1. Scope the exact target, environment boundary, owner, requested decision, and evidence available.
      2. Load `official-sources.md`, then the component operations guide for service behavior and risk gates.
      3. Gather sampled read-only evidence only when available and safe.
      4. Compare observed posture against documented behavior, least-privilege expectations, and operational safety rules.
      5. Return a verdict with evidence level, blockers, safe next actions, and open questions.
      
      ## Required output
      
      - `verdict`: pass, warn, fail, or blocked.
      - `evidence_level`: documentation-based, sampled-current-state, user-provided, inference, or mixed.
      - `scope`: what was reviewed and what was not reviewed.
      - `blockers`: issues that prevent a safe or production-ready conclusion.
      - `findings`: severity-labeled risks with source labels.
      - `safe_next_actions`: reversible actions first; mutation only with explicit approval.
      - `open_questions`: missing facts that would change the verdict.
      
      ## Stress checks
      
      - What assumption would make this recommendation unsafe?
      - Which identity, migration, network, telemetry, route, or dispatch decision has the largest blast radius?
      - What evidence would disprove the claimed readiness?
      - Is the answer accidentally treating documentation as configured-environment proof?
      
      ## Response discipline
      
      Use Microsoft Learn documentation through the user's configured documentation MCP for documented Azure behavior. Use sampled read-only Azure evidence only for current configured-environment observations and label it as sampled evidence.
      
  • metadata.json 1.7 KB
    {
      "id": "azure-observability-investigator",
      "name": "Azure Observability Investigator",
      "type": "skill",
      "provider": "azure",
      "harnesses": [
        "codex",
        "claude-code",
        "cursor",
        "gemini",
        "kiro",
        "other"
      ],
      "summary": "Investigate Azure Monitor, Log Analytics, Application Insights, alerting, KQL triage, telemetry gaps, workbooks, Grafana, and incident hypotheses with explicit evidence-versus-inference handling.",
      "source_type": "original",
      "official_docs": [
        "https://learn.microsoft.com/azure/azure-monitor/fundamentals/overview",
        "https://learn.microsoft.com/azure/azure-monitor/fundamentals/best-practices-operation",
        "https://learn.microsoft.com/azure/azure-monitor/alerts/alerts-overview",
        "https://learn.microsoft.com/azure/azure-monitor/alerts/action-groups",
        "https://learn.microsoft.com/azure/azure-monitor/alerts/alerts-processing-rules",
        "https://learn.microsoft.com/azure/azure-monitor/logs/log-analytics-overview",
        "https://learn.microsoft.com/azure/azure-monitor/logs/workspace-design",
        "https://learn.microsoft.com/azure/azure-monitor/app/app-insights-overview",
        "https://learn.microsoft.com/azure/azure-monitor/visualize/workbooks-overview",
        "https://learn.microsoft.com/azure/managed-grafana/how-to-use-azure-monitor-alerts"
      ],
      "security_notes": "Do not over-attribute symptoms as root cause, ignore missing telemetry, or recommend broad alerting changes without signal-quality review, routing checks, query scope, and bounded verification steps.",
      "last_verified": "2026-06-05",
      "path": "skills/azure/azure-observability-investigator",
      "author": "github: VincentChuWaiChow",
      "version": "0.1.2"
    }
    
  • SKILL.md 3.3 KB
    ---
    name: azure-observability-investigator
    description: Use this skill for Azure Monitor, Log Analytics, Application Insights, alerting, KQL triage, telemetry-gap analysis, workbooks, or operator-grade incident and posture investigations.
    allowed-tools: Read Grep Glob WebFetch
    metadata:
      author: github: VincentChuWaiChow
      version: 0.1.2
      updated: "2026-06-05"
      category: observability
    ---
    
    # Azure Observability Investigator
    
    ## Purpose
    
    Investigate Azure operational health using evidence from metrics, logs, traces, alerts, and observability configuration before jumping to root-cause claims.
    
    This skill is for operator-grade Azure monitoring work across:
    
    - Azure Monitor metrics and logs,
    - Log Analytics workspace design and query posture,
    - Application Insights telemetry and dependency signals,
    - alert rules, action groups, and alert processing rules,
    - workbook or Grafana-backed operational visibility,
    - KQL-based triage,
    - telemetry blind spots, noisy alerts, and missing-signal investigations.
    
    ## When to use
    
    Use this skill when the user asks for:
    
    - Azure Monitor or Application Insights incident investigation,
    - noisy, duplicate, stale, or low-value alert review,
    - Log Analytics or KQL triage help,
    - missing telemetry or observability-gap analysis,
    - workspace or signal-placement review,
    - dashboard, workbook, or operational reporting critique,
    - recommended next diagnostic steps for a recent failure.
    
    Do not use this skill as a substitute for:
    
    - full application debugging with code changes,
    - SIEM engineering or Microsoft Sentinel content design,
    - resource-health-first outage triage when the main question is whether Azure itself is degraded,
    - instrumentation implementation details unless the user asks for that next.
    
    ## Lean operating rules
    
    - Prefer Microsoft Learn documentation through the user's configured documentation MCP, then sampled read-only Azure evidence when available, then sanitized user evidence.
    - Separate confirmed facts from inference. If state was not queried or shown, say so.
    - Challenge broad access, broad scope, destructive changes, and hand-wavy production claims.
    - Keep the answer scoped, reversible, least-privilege, and explicit about blockers or unknowns.
    
    ## References
    
    Load these only when needed:
    
    - [Azure Observability Investigation Operations](references/observability-investigation-operations.md) — use for current service behavior, common failure modes, hard design rules, verification targets, and push-back conditions.
    - [Safety checklist](references/safety-checklist.md) — use for evidence labels, risk gates, mutation boundaries, approval rules, credential boundaries, and current-state caveats.
    - [MCP and evidence path](references/mcp-and-evidence.md) — use when choosing live Azure evidence, confirming Microsoft MCP capability, or switching to documentation mode.
    - [Workflow and output contract](references/workflow-and-output.md) — use when executing the full review, applying stress checks, or formatting the final answer.
    - [Official sources](references/official-sources.md) — use when you need the detailed Microsoft documentation list or source notes.
    
    ## Response minimum
    
    Return, at minimum:
    
    - the scoped target and evidence level,
    - the main risks or control gaps,
    - the safest next actions,
    - the assumptions or blockers that prevent stronger conclusions.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related