azure-cosmosdb-performance-investigator
Use this skill for Azure Cosmos DB performance investigation, especially RU spikes, query latency, throttling, hot partitions, indexing inefficiency, partition-skew analysis, request-charge profiling, diagnostic-log review, and evidence-driven remediation planning.
Install
npx skills add https://github.com/VincentChuWaiChow/vanguard-frontier-agentic/tree/master/skills/azure/azure-cosmosdb-performance-investigator
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install vincentchuwaichow-vanguard-frontier-agentic@llmmart
git clone https://github.com/VincentChuWaiChow/vanguard-frontier-agentic.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole vincentchuwaichow/vanguard-frontier-agentic collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Azure Cosmos DB Performance Investigator
Purpose
Investigate Azure Cosmos DB performance pathologies with evidence-first profiling instead of lazy “add more RUs” advice.
This skill is for deep performance work across:
- RU inefficiency and unexpected request-charge spikes,
- query latency and scan-heavy query behavior,
- hot partitions and partition-key skew,
- throttling, retry inflation, and client-perceived latency,
- indexing gaps and poor query/index alignment,
- container, partition, and workload-level profiling,
- diagnostic-log and metrics-backed remediation planning.
When to use
Use this skill when the user asks for:
- slow Azure Cosmos DB queries or workload latency,
- high RU cost or suspicious request-charge behavior,
- 429 throttling analysis,
- hot partition or partition-skew investigation,
- indexing or query-performance tuning,
- a step-by-step Cosmos DB profiling plan.
Do not use this skill as a substitute for:
- initial data-model design when the main problem is greenfield schema modeling,
- pure account/platform governance review when performance is incidental,
- generic application debugging unrelated to Cosmos DB workload behavior,
- vector-search-specific Mongo vCore tuning unless the user explicitly asks for that API surface.
Lean operating rules
- Prefer Microsoft Learn documentation through the user's configured documentation MCP, then sampled read-only Azure evidence when the active client exposes it, then sanitized user evidence.
- Separate confirmed facts from inference. If state was not queried or shown, say so.
- Challenge throughput-first fixes that ignore partition skew, query scans, indexing, or client retry inflation.
- Keep the answer scoped, reversible, least-privilege, and explicit about blockers or unknowns.
References
Load these only when needed:
- Operations guide — use for service-specific pitfalls, design rules, verification targets, and pushback criteria.
- MCP and evidence path — use when choosing documentation-based evidence, sampled read-only Azure evidence, or sanitized user evidence.
- Safety checklist — use for evidence labels, risk gates, mutation boundaries, approval rules, and credential boundaries.
- Workflow and output contract — use when executing the full investigation, applying stress checks, or formatting the final answer.
- Data profiling playbook — use when you need the detailed step-by-step profiling sequence.
- Official sources — use when you need the detailed Microsoft documentation list or source notes.
Response minimum
Return, at minimum:
- the scoped target and evidence level,
- the main performance pathologies observed or still unproven,
- the safest next profiling or remediation steps,
- the assumptions or blockers that prevent stronger conclusions.
Files (vanguard-frontier-agentic)
-
references
-
cosmosdb-performance-investigation.md 4.5 KB
# Cosmos DB performance investigation ## What people get wrong - They add RUs before proving whether the problem is global load, hot partitions, query scans, indexing, document size, consistency, or client retries. - They look only at average latency and miss p95/p99, 429s, partition-level normalized RU, and retry inflation. - They inspect query text without query metrics or request charge. - They treat physical partition pressure as proof of a specific logical key before checking diagnostic logs. - They change indexing or partition strategy without rollback and backfill planning. ## Officially grounded service shape Microsoft Learn states that RU consumption depends on document size, indexing, consistency, query shape, and operation type. Each request exposes request charge. Query metrics break down query work, and Azure Monitor/Insights expose Total Request Units, normalized RU by partition key range, throttling, storage, and latency signals. Hot partitions appear when one or a few partition key ranges consume disproportionate RU/s. ## Non-negotiable design rules 1. Establish a time window before collecting metrics. 2. Separate account/container-level symptoms from operation/query-level symptoms. 3. Capture request charge and query metrics for representative operations. 4. Check normalized RU by partition key range before assuming global under-provisioning. 5. Check status codes and retry behavior before blaming service latency. 6. Treat indexing changes, throughput changes, and data reshaping as mutations requiring approval. 7. Prefer reversible experiments before structural changes. ## Minimal safe implementation flow 1. Define symptom, workload path, API type, region, and time window. 2. Collect aggregate metrics: total RU, 429s, latency percentiles, operation type, status code, and region. 3. Collect partition metrics: normalized RU by partition key range and top logical partition evidence when diagnostics are available. 4. Profile representative operations: request charge, query metrics, index metrics, result count, page count, and continuation behavior. 5. Profile client: SDK version, connection mode, retries, timeouts, diagnostics, and deployment region path. 6. Rank root causes and propose one low-risk experiment per suspected cause. 7. Document approval and rollback before changing throughput, index policy, partitioning, consistency, or SDK retry behavior. ## High-risk assumptions to kill - More RU/s is not a root-cause fix when one partition key range is hot, a query scans, indexing is wrong, or clients are retrying badly. - Average latency hides tail latency, retries, page count, continuation behavior, regional routing, and partition-level pressure. - A physical partition hot spot does not identify the logical offender until diagnostic logs or equivalent partition-key evidence is available. - Index metrics are troubleshooting evidence, not a reason to enable expensive diagnostics or change indexing permanently without rollback. - Autoscale reaching 100% normalized RU briefly is not automatically unhealthy; rate of 429s and end-to-end latency decide urgency. ## Safe command/code verification targets - Inspect application code for request-charge logging, SDK diagnostics capture, retry configuration, connection mode, timeout settings, and preferred region behavior. - Check query code and captured metrics for retrieved-versus-output document counts, index hit ratio, page count, continuation-token handling, and cross-partition execution. - Review Azure Monitor/Kusto queries or dashboards for normalized RU by partition key range, 429 percentage, operation type, p95/p99 latency, and regional splits. - Verify index-policy changes are represented as reviewed diffs with before/after RU and latency measurements plus rollback steps. - Confirm throughput, consistency, partition-key, or SDK tuning recommendations include an experiment boundary and post-change measurement window. ## Safe verification targets - `RequestCharge` per representative read, write, and query. - Query metrics: index lookup, document load, output document count, retrieved document count, and execution time components. - Index metrics for filter and sort support. - Normalized RU consumption split by partition key range. - 429 count, retry count, p95/p99 latency, and region routing. - Before/after measurement for each remediation experiment. ## When to push back Push back when the proposed fix is "increase RUs" without partition/query proof, repartitioning without distribution evidence, index changes without query metrics, or SDK changes without diagnostics. -
data-profiling-playbook.md 6.2 KB
# Data profiling playbook Use this reference when the user needs a detailed, step-by-step Cosmos DB performance investigation path. ## Goal Determine whether the workload problem is caused primarily by: - bad query shape, - index mismatch, - partition-key skew, - insufficient or badly distributed throughput, - cross-region or client-side latency, - retry amplification, - poor workload access patterns, - or multiple issues at once. ## Step 1: Define the symptom precisely Capture exactly which of these is true: 1. RU charge is too high. 2. Query latency is too high. 3. 429 throttling is frequent. 4. One partition or one workload slice is much worse than others. 5. The issue is intermittent rather than constant. If the user only says “Cosmos is slow,” push back and force narrower symptom definition. ## Step 2: Fix the observation window Before analysis, anchor the time range. Recommended windows: - **1 hour** for a sharp incident or deploy regression - **24 hours** for day-pattern investigation - **7 days** for partition-skew or periodic workload behavior Do not compare unrelated windows. ## Step 3: Establish workload scope Identify: - account - database n- container - API type - top queries or endpoints - whether the issue is read-heavy, write-heavy, or mixed If the user cannot identify the hot path, say the conclusion will remain weaker. ## Step 4: Measure request charge first For code-facing workloads, collect request charge from the SDK or response headers. Why: - RU cost is the fastest way to distinguish inefficient operations from mere latency complaints. - A stable high request charge suggests query/data/index issues. - A low request charge with high latency suggests retries, distance, contention, or client behavior. Collect examples for: - representative point read - representative query - representative write or batch operation ## Step 5: Get query metrics For each expensive or slow query: 1. collect query metrics 2. compare **Retrieved Document Count** vs **Output Document Count** 3. flag scan-heavy behavior when retrieved greatly exceeds output Interpretation: - **Retrieved much higher than output** → likely index miss or scan-heavy predicate/function behavior - **Retrieved approximately equal to output** with high RU → maybe wide result sets, cross-partition cost, or costly ordering/filter combinations - **RU acceptable but latency high** → move to proximity, retries, concurrency, and buffering checks ## Step 6: Check index metrics only when troubleshooting Use index metrics when query behavior is suspicious and the indexing path is unclear. Look for: - utilized indexed paths - recommended indexed paths - evidence that the current indexing policy does not support the query efficiently Do not blindly expand indexing on every field without considering write cost. ## Step 7: Test for hot partitions Use portal insights or equivalent telemetry to inspect: - **Normalized RU Consumption (%) By PartitionKeyRangeID** Interpretation: - one or a few ranges consistently near saturation while others are low strongly suggests skew - even usage with high overall pressure suggests general underprovisioning or globally inefficient workload design ## Step 8: Move from physical skew to logical key offenders If hot partitions are suspected: 1. enable and inspect diagnostic logs if available 2. use partition-key RU consumption data to identify the top logical partition keys 3. review at least several days when possible, not just one bursty hour This is where “add more RUs” often fails. If a few keys dominate, throughput growth can mask the problem without fixing it. ## Step 9: Profile query shape versus access pattern Ask or verify: - should this be a point read instead of a query? - is the partition key included where it should be? - is the query filtering on properties that align with the partition strategy? - is ORDER BY combined with filters in a way the index supports? - is the application doing repeated small round trips that should be consolidated? Common anti-patterns: - using queries where a point read would do - cross-partition fan-out for tenant-local data - filtering on secondary properties without an index strategy - embedding data that grows without bound and inflates read/write cost ## Step 10: Separate server-side from client-side latency If RU looks acceptable but latency is still poor, check: - region proximity between app and Cosmos DB account - throttling retries inflating end-to-end time - MaxConcurrency for parallel queries - MaxBufferedItemCount / prefetch behavior - client connection reuse / SDK posture Do not blame the database alone without this separation. ## Step 11: Profile storage and index footprint Inspect storage insights when available: - data usage - index usage - document usage Use this to spot cases where index footprint is high relative to value or where document shape is inflating costs. ## Step 12: Classify the root cause before prescribing remediation Use one or more of these buckets: 1. **Query/index mismatch** 2. **Partition-key skew / hot partitions** 3. **General underprovisioning** 4. **Client-side latency / retries / distance** 5. **Data-model inefficiency** 6. **Mixed causes** If evidence is incomplete, say so. ## Step 13: Recommend bounded remediations Good remediation examples: - capture request charge for the top 3 slow queries before changing throughput - enable query metrics on one hot query and compare retrieved vs output counts - inspect normalized RU consumption by partition key range for 7 days - review whether a hot query should become a point read - test one indexing-policy adjustment against one proven expensive query Bad remediation examples: - “just increase RUs” - “just repartition everything” - “turn on every index” - “switch consistency” without naming the application tradeoff ## Step 14: Decide when to hand off to another role Hand off or recommend adjacent roles when: - the platform posture itself is the main issue → **Azure Cosmos DB Platform Operator** - the data model / query code is the main issue → **Azure Cosmos DB Application Developer** - alerting / visibility is the main issue → **Azure Observability Investigator** - ongoing RU waste and budget posture dominate → **Azure Cost Optimization Governor** -
mcp-and-evidence.md 1.8 KB
# MCP and evidence path for Cosmos DB performance investigation Use Microsoft Learn documentation through the user's configured documentation MCP as the first grounding path for Azure service behavior. This file defines evidence boundaries; it must not imply that documentation proves the user's tenant, subscription, RBAC, quotas, deployed resources, or production readiness. ## Evidence ladder 1. `docs_only`: Microsoft Learn documentation and official architecture guidance. Use for documented behavior, caveats, and safe review criteria. 2. `sampled_read_only`: configured-environment evidence from read-only tools, if available and explicitly scoped. Use only for the sampled resource/time window. 3. `user_supplied`: sanitized outputs, IaC, diagrams, or metrics provided by the user. Treat as unverified unless independently checked. 4. `mutation_ready`: documentation plus current-state evidence plus explicit approval, blast-radius statement, and rollback path. ## Rules - Do not expose environment-specific implementation details in committed docs or user-facing guidance. - Do not ask for credentials, tokens, tenant identifiers, subscription identifiers, connection strings, private keys, customer data, or raw secrets. - If current-state evidence was not sampled, say `not sampled`; do not imply it. - If evidence is representative or partial, say so. A sample does not prove broad regional availability or production readiness. - Prefer read-only evidence before mutation planning. Stop for approval before write operations. ## Final-answer evidence language Use phrases like: - "Based on Microsoft Learn documentation..." - "Configured-environment evidence was not sampled in this review." - "The following is an inference from the provided configuration, not proven live state." - "This recommendation is mutation-ready only after explicit approval and rollback review." -
official-sources.md 2.7 KB
# Official sources for Azure Cosmos DB Performance Investigator Use Microsoft Learn documentation through the user's configured documentation MCP before recommending performance fixes. Documentation proves service behavior; it does not prove this user's hot partitions, RU budget, indexes, SDK behavior, retry inflation, or latency root cause. ## Primary Microsoft Learn sources | Source | Review implication | | --- | --- | | [Troubleshoot query performance](https://learn.microsoft.com/en-us/azure/cosmos-db/troubleshoot-query-performance) | Use for slow query triage, query metrics, index utilization, and query-shape remediation. | | [Query metrics](https://learn.microsoft.com/en-us/azure/cosmos-db/query-metrics) | Require query metrics before claiming where query time or RU cost is going. | | [Index metrics](https://learn.microsoft.com/en-us/azure/cosmos-db/index-metrics) | Use to validate whether current indexes support filters and sorts. | | [Understand request units consumption](https://learn.microsoft.com/en-us/azure/cosmos-db/understand-request-unit-consumption) | Ground RU cost in document size, indexing, consistency, and query shape. | | [Monitor and debug with insights](https://learn.microsoft.com/en-us/azure/cosmos-db/use-metrics) | Use for throughput, storage, normalized RU, request charge, and partition-key-range views. | | [Monitor normalized request units](https://learn.microsoft.com/en-us/azure/cosmos-db/monitor-normalized-request-units) | Use to detect physical partition pressure and hot partition indicators. | | [Diagnose and troubleshoot request rate too large](https://learn.microsoft.com/en-us/azure/cosmos-db/troubleshoot-request-rate-too-large) | Use for 429 and throttling diagnostics, including hot partition analysis. | | [Redistribute throughput across partitions](https://learn.microsoft.com/en-us/azure/cosmos-db/how-to-redistribute-throughput-across-partitions) | Treat as preview/limited remediation for eligible accounts; verify limitations before suggesting. | | [Performance tips for .NET SDK v3](https://learn.microsoft.com/en-us/azure/cosmos-db/performance-tips-dotnet-sdk-v3) | Use for SDK-side diagnostics, retry, connection, and request-charge instrumentation patterns when .NET is in scope. | ## Source-grounding rules - Never recommend "add RUs" before separating hot partition, inefficient query, indexing, consistency, document size, and client retry causes. - Query metrics and request charge are stronger evidence than aggregate charts for a single query path. - Normalized RU by partition key range is a hot-partition signal, not proof of the exact logical key unless diagnostics identify it. - Throughput redistribution has eligibility and preview caveats; do not present it as a default fix. -
safety-checklist.md 2.2 KB
# Safety checklist for Azure Cosmos DB Performance Investigator ## Non-negotiable gates - Never ask for account keys, connection strings, full documents, customer data, tenant identifiers, subscription identifiers, or database dumps. - Do not recommend throughput changes, indexing changes, repartitioning, SDK retry changes, consistency changes, or data reshaping without evidence and rollback. - Treat diagnostic logs as potentially sensitive. Ask only for sanitized aggregates, query text with literals removed, or metrics screenshots summarized in text. - Do not claim root cause from a single metric. Correlate RU, latency, status codes, partition distribution, query metrics, and client retries. - Require explicit approval before any write or configuration mutation. ## High-risk assumptions to kill - "429 means under-provisioned globally." It may be a hot partition or retry pattern. - "High latency means service issue." It may be query fan-out, index miss, large documents, consistency, region path, or client connection mode. - "Normalized RU at 100% always equals failures." It can spike without broad 429s; correlate with status and workload. - "Indexing more properties improves performance." It can increase write cost and may not fix query shape. - "Throughput redistribution fixes a bad partition key." It may reduce symptoms but not remove skewed design risk. ## Evidence labels - `docs_only`: Microsoft Learn guidance only. - `metrics_sample`: sanitized Azure Monitor or Insights metrics reviewed. - `query_profile`: request charge, query metrics, index metrics, diagnostics, or SDK logs reviewed. - `remediation_ready`: proposed change has measured baseline, blast radius, rollback, and approval. ## Minimum safe evidence - Time window, region, API type, account/container, workload path, and symptom. - Total Request Units, 429 counts, normalized RU by partition key range, latency percentiles, and operation type. - Representative query text with sensitive literals removed, request charge, query metrics, index metrics, and result counts. - SDK version, connection mode, retry policy, timeout behavior, and diagnostics for client-perceived latency. - Recent changes in traffic, data volume, indexing, regions, consistency, or deployments. -
workflow-and-output.md 1.5 KB
# Workflow and output contract for Azure Cosmos DB Performance Investigator ## Minimal safe workflow 1. Classify symptom: high RU, latency, 429, hot partition, index miss, SDK/client issue, or mixed. 2. Ground the investigation with Microsoft Learn through the user's configured documentation MCP. 3. Establish time window and scope without exposing sensitive identifiers. 4. Build a four-lane evidence table: service metrics, query profile, partition profile, and client profile. 5. Rank likely causes only after correlating request charge, query metrics, normalized RU, status codes, latency, and retries. 6. Propose lowest-risk experiments before irreversible changes. 7. Require approval and rollback for any throughput, index, data model, partition, consistency, or SDK policy mutation. ## Output contract ```markdown ## Verdict <root cause likely | mixed causes | insufficient evidence | docs-only advisory> ## Evidence level - Documentation: <sources used> - Metrics/query evidence: <metrics_sample | query_profile | not provided> ## Findings 1. <finding> — Evidence: <docs_only|metrics_sample|query_profile|inference> ## Root-cause ranking 1. <cause> — why it is likely / what would disprove it ## Safe next profiling steps - <specific metric/query/diagnostic to collect> ## Remediation boundaries - <changes that require approval and rollback> ``` ## Pushback triggers Push back on blind RU increases, unmeasured index changes, repartitioning without hot-key proof, query rewrites without query metrics, or SDK tuning without client diagnostics.
-
-
metadata.json 1.7 KB
{ "id": "azure-cosmosdb-performance-investigator", "name": "Azure Cosmos DB Performance Investigator", "version": "0.1.3", "type": "skill", "provider": "azure", "harnesses": [ "codex", "claude-code", "cursor", "gemini", "kiro", "other" ], "summary": "Investigate Azure Cosmos DB query latency, RU inefficiency, throttling, hot partitions, indexing gaps, and workload-level performance pathologies using explicit evidence, metrics, and step-by-step profiling discipline.", "source_type": "original", "official_docs": [ "https://learn.microsoft.com/en-us/azure/cosmos-db/troubleshoot-query-performance", "https://learn.microsoft.com/en-us/azure/cosmos-db/query-metrics", "https://learn.microsoft.com/en-us/azure/cosmos-db/index-metrics", "https://learn.microsoft.com/en-us/azure/cosmos-db/use-metrics", "https://learn.microsoft.com/en-us/azure/cosmos-db/how-to-redistribute-throughput-across-partitions", "https://learn.microsoft.com/en-us/azure/cosmos-db/performance-tips-dotnet-sdk-v3", "https://learn.microsoft.com/en-us/azure/well-architected/service-guides/cosmos-db", "https://learn.microsoft.com/en-us/azure/cosmos-db/monitor-normalized-request-units", "https://learn.microsoft.com/en-us/azure/cosmos-db/autoscale-faq", "https://learn.microsoft.com/en-us/azure/cosmos-db/understand-request-unit-consumption" ], "security_notes": "Do not recommend throughput increases, repartitioning, indexing changes, or SDK tuning before separating RU cost, latency, partition skew, and query-shape evidence. Avoid speculative fixes that hide workload design defects.", "last_verified": "2026-06-05", "path": "skills/azure/azure-cosmosdb-performance-investigator", "author": "github: VincentChuWaiChow" } -
SKILL.md 3.4 KB
--- name: azure-cosmosdb-performance-investigator description: Use this skill for Azure Cosmos DB performance investigation, especially RU spikes, query latency, throttling, hot partitions, indexing inefficiency, partition-skew analysis, request-charge profiling, diagnostic-log review, and evidence-driven remediation planning. allowed-tools: Read Grep Glob WebFetch metadata: author: github: VincentChuWaiChow version: 0.1.3 updated: "2026-06-05" category: data --- # Azure Cosmos DB Performance Investigator ## Purpose Investigate Azure Cosmos DB performance pathologies with evidence-first profiling instead of lazy “add more RUs” advice. This skill is for deep performance work across: - RU inefficiency and unexpected request-charge spikes, - query latency and scan-heavy query behavior, - hot partitions and partition-key skew, - throttling, retry inflation, and client-perceived latency, - indexing gaps and poor query/index alignment, - container, partition, and workload-level profiling, - diagnostic-log and metrics-backed remediation planning. ## When to use Use this skill when the user asks for: - slow Azure Cosmos DB queries or workload latency, - high RU cost or suspicious request-charge behavior, - 429 throttling analysis, - hot partition or partition-skew investigation, - indexing or query-performance tuning, - a step-by-step Cosmos DB profiling plan. Do not use this skill as a substitute for: - initial data-model design when the main problem is greenfield schema modeling, - pure account/platform governance review when performance is incidental, - generic application debugging unrelated to Cosmos DB workload behavior, - vector-search-specific Mongo vCore tuning unless the user explicitly asks for that API surface. ## Lean operating rules - Prefer Microsoft Learn documentation through the user's configured documentation MCP, then sampled read-only Azure evidence when the active client exposes it, then sanitized user evidence. - Separate confirmed facts from inference. If state was not queried or shown, say so. - Challenge throughput-first fixes that ignore partition skew, query scans, indexing, or client retry inflation. - Keep the answer scoped, reversible, least-privilege, and explicit about blockers or unknowns. ## References Load these only when needed: - [Operations guide](references/cosmosdb-performance-investigation.md) — use for service-specific pitfalls, design rules, verification targets, and pushback criteria. - [MCP and evidence path](references/mcp-and-evidence.md) — use when choosing documentation-based evidence, sampled read-only Azure evidence, or sanitized user evidence. - [Safety checklist](references/safety-checklist.md) — use for evidence labels, risk gates, mutation boundaries, approval rules, and credential boundaries. - [Workflow and output contract](references/workflow-and-output.md) — use when executing the full investigation, applying stress checks, or formatting the final answer. - [Data profiling playbook](references/data-profiling-playbook.md) — use when you need the detailed step-by-step profiling sequence. - [Official sources](references/official-sources.md) — use when you need the detailed Microsoft documentation list or source notes. ## Response minimum Return, at minimum: - the scoped target and evidence level, - the main performance pathologies observed or still unproven, - the safest next profiling or remediation steps, - the assumptions or blockers that prevent stronger conclusions.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.