azure-cosmosdb-platform-operator
Use this skill for Azure Cosmos DB platform operations and design review, especially accounts, databases, containers, partition-key design, throughput and RU posture, consistency choices, indexing, throttling, multi-region replication, private connectivity, and Cosmos DB evidence
Install
npx skills add https://github.com/VincentChuWaiChow/vanguard-frontier-agentic/tree/master/skills/azure/azure-cosmosdb-platform-operator
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install vincentchuwaichow-vanguard-frontier-agentic@llmmart
git clone https://github.com/VincentChuWaiChow/vanguard-frontier-agentic.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole vincentchuwaichow/vanguard-frontier-agentic collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Azure Cosmos DB Platform Operator
Purpose
Review and guide Azure Cosmos DB platform posture without hand-wavy assumptions about partitioning, consistency, throughput, or multi-region behavior.
When to use
Use this skill when the user asks for:
- Azure Cosmos DB account, database, or container operational review,
- partition-key, consistency, throughput, or indexing decisions,
- RU cost, throttling, hot partition, or multi-region tradeoff analysis,
- private endpoint, failover, or platform control-plane posture questions,
- Cosmos DB documentation-grounded discovery or sampled read-only evidence gathering.
Do not use this skill as a substitute for:
- application code implementation details unless they directly affect platform posture,
- generic landing-zone or network architecture when Cosmos DB is incidental,
- RBAC-only analysis when the real problem is identity governance rather than database platform design,
- vector-search-specific Mongo vCore implementation unless the user explicitly asks for that API surface.
Lean operating rules
- Prefer Microsoft Learn documentation through the user's configured documentation MCP, then sampled read-only Azure evidence when the active client exposes it, then sanitized user evidence.
- Separate confirmed facts from inference. If state was not queried or shown, say so.
- Challenge broad scope, vague partition keys, and RU-blind advice.
- Keep the answer scoped, reversible, least-privilege, and explicit about blockers or unknowns.
References
Load these only when needed:
- Operations guide — use for service-specific pitfalls, design rules, verification targets, and pushback criteria.
- MCP and evidence path — use when choosing documentation-based evidence, sampled read-only Azure evidence, or sanitized user evidence.
- Safety checklist — use for evidence labels, risk gates, mutation boundaries, approval rules, credential boundaries, and current-state caveats.
- Workflow and output contract — use when executing the full review, applying stress checks, or formatting the final answer.
- Official sources — use when you need the detailed Microsoft documentation list or source notes.
Response minimum
Return, at minimum:
- the scoped target and evidence level,
- the main design or operational risks,
- the safest next actions,
- the assumptions or blockers that prevent stronger conclusions.
Files (vanguard-frontier-agentic)
-
references
-
cosmosdb-platform-operations.md 4.8 KB
# Cosmos DB platform operations ## What people get wrong - They treat account-level multi-region settings as proof the application fails over safely. - They ignore SDK preferred-region configuration, private DNS, and remaining-region throughput when reviewing regional resiliency. - They change consistency, failover priority, or throughput during stress without knowing the blast radius. - They view backup as an availability feature instead of a data-recovery feature. - They approve partitioning from a diagram instead of measured access patterns and distribution. ## Officially grounded service shape Microsoft Learn describes Cosmos DB as an account containing databases and containers, with containers as logical units of distribution and scalability. Reliability features include zone redundancy, multi-region replication, several consistency levels, service-managed or customer-managed failover, SDK resiliency behavior, and continuous or periodic backup. The reliability guide also warns against control-plane changes on affected regions during outage scenarios and calls out private endpoint DNS considerations after failover. ## Non-negotiable design rules 1. Identify API type, write model, regions, consistency, backup, and throughput mode before giving a platform verdict. 2. Treat partitioning and throughput as workload-coupled, not platform-only decisions. 3. Verify SDK region preference and retry behavior for multi-region availability. 4. For private endpoints, design regional endpoints and DNS resolution deliberately. 5. Size remaining regions for failover load before claiming regional resilience. 6. Use backup/restore for data corruption/deletion scenarios; use multi-region and failover for availability scenarios. 7. Require diagnostics, Resource Health, Service Health, alerts, and DR runbooks. ## Minimal safe implementation flow 1. Inventory account topology: API, regions, write model, consistency, backup, throughput, network access, and private endpoints. 2. Review workload shape: partition key, hot partition risk, RU headroom, critical operations, and consistency needs. 3. Review reliability: zone redundancy, failover mode, preferred regions, capacity under region loss, and DR drills. 4. Review network: private endpoints, DNS zones, hub/spoke forwarding, firewall, and client routing. 5. Review operations: metrics, logs, alerts, backup restore testing, and owner/runbook evidence. 6. Rank blockers and reversible remediations. 7. Require explicit approval before any control-plane mutation. ## High-risk assumptions to kill - Multi-region account configuration does not prove application failover unless SDK preferred regions, retries, conflict handling, and remaining-region capacity are verified. - Private endpoints can break regional resilience if DNS zones, VNet links, forwarding, and local regional endpoint paths are not designed deliberately. - Strong, bounded staleness, session, and eventual consistency have different availability, RPO, and throughput implications; do not treat consistency as a cosmetic setting. - Control-plane changes during an affected-region outage can delay recovery; do not mutate write region, consistency, throughput, networking, or failover priorities casually. - Backup protects data recovery scenarios; it is not a substitute for live availability, failover routing, or tested application DR. ## Safe command/code verification targets - Inspect IaC for account regions, zone redundancy, failover priorities, write-region model, consistency, backup mode, throughput mode, and private endpoint resources. - Review application configuration for SDK preferred regions, excluded regions, retry diagnostics, conflict resolution, and multi-write support where applicable. - Check network templates for one-region-per-DNS-zone patterns, private DNS links, forwarding rules, firewall dependencies, and client resolution paths. - Verify monitoring code or dashboards cover normalized RU, 429s, region availability, Resource Health, Service Health, backup status, and restore-test evidence. - Confirm every proposed account mutation has an approval gate, rollback plan, and explicit outage-state caveat. ## Safe verification targets - Account regions, write region or multi-write state, failover priorities, and service-managed/forced failover process. - Consistency level and documented RPO/availability implications. - Throughput mode and normalized RU headroom by partition key range. - Private endpoint count, DNS zones, VNet links, forwarding, and client resolution path. - Backup mode, retention, restore test, and data-corruption response plan. - SDK preferred regions, retry, partition-level circuit breaker or equivalent availability settings. ## When to push back Push back on production signoff without failover evidence, private endpoint DNS proof, capacity-under-failure math, SDK routing proof, backup restore tests, or partition distribution evidence. -
mcp-and-evidence.md 1.9 KB
# MCP and evidence path for Cosmos DB platform operations Use Microsoft Learn documentation through the user's configured documentation MCP as the first grounding path for Azure service behavior. This file defines evidence boundaries; it must not imply that documentation proves the user's tenant, subscriptions, RBAC, quotas, billing agreement, deployed resources, or production readiness. ## Evidence ladder 1. `docs_only`: Microsoft Learn documentation and official architecture guidance. Use for documented behavior, caveats, and safe review criteria. 2. `sampled_read_only`: configured-environment evidence from read-only tools, if available and explicitly scoped. Use only for the sampled resource/time window. 3. `user_supplied`: sanitized outputs, IaC, diagrams, billing summaries, or metrics provided by the user. Treat as unverified unless independently checked. 4. `mutation_ready`: documentation plus current-state evidence plus explicit approval, blast-radius statement, and rollback path. ## Rules - Do not expose environment-specific implementation details in committed docs or user-facing guidance. - Do not ask for credentials, tokens, tenant identifiers, subscription identifiers, billing account identifiers, connection strings, private keys, customer data, or raw secrets. - If current-state evidence was not sampled, say `not sampled`; do not imply it. - If evidence is representative or partial, say so. A sample does not prove broad regional availability, billing accuracy, policy compliance, or production readiness. - Prefer read-only evidence before mutation planning. Stop for approval before write operations. ## Final-answer evidence language Use phrases like: - "Based on Microsoft Learn documentation..." - "Configured-environment evidence was not sampled in this review." - "The following is an inference from the provided configuration, not proven live state." - "This recommendation is mutation-ready only after explicit approval and rollback review." -
official-sources.md 2.4 KB
# Official sources for Azure Cosmos DB Platform Operator Use Microsoft Learn documentation through the user's configured documentation MCP before making Cosmos DB platform claims. Documentation proves documented service behavior; it does not prove the user's account configuration, regions, throughput, private DNS, backups, or workload readiness. ## Primary Microsoft Learn sources | Source | Review implication | | --- | --- | | [Reliability in Azure Cosmos DB](https://learn.microsoft.com/en-us/azure/reliability/reliability-cosmos-db) | Ground zone redundancy, multi-region, failover, backup, SDK resiliency, RPO, and control-plane outage caveats. | | [Architecture best practices for Azure Cosmos DB](https://learn.microsoft.com/en-us/azure/well-architected/service-guides/cosmos-db) | Use for Well-Architected reliability, security, cost, operational excellence, and performance tradeoffs. | | [Partitioning and horizontal scaling](https://learn.microsoft.com/en-us/azure/cosmos-db/partitioning) | Use for partition-key, logical partition, physical partition, and scale-risk review. | | [Hierarchical partition keys](https://learn.microsoft.com/en-us/azure/cosmos-db/hierarchical-partition-keys) | Use when tenant/user/item or similar hierarchy is proposed; verify workload fit and limits. | | [Consistency levels](https://learn.microsoft.com/en-us/azure/cosmos-db/consistency-levels) | Use for availability, latency, throughput, and correctness tradeoffs. | | [Request Units](https://learn.microsoft.com/en-us/azure/cosmos-db/request-units) | Use for provisioned throughput, multi-region throughput, and consistency effects. | | [Failover considerations for private endpoints](https://learn.microsoft.com/en-us/azure/cosmos-db/failover-considerations-for-private-endpoints) | Use for regional private endpoint, private DNS, and failover routing checks. | | [Online backup and restore](https://learn.microsoft.com/en-us/azure/cosmos-db/online-backup-and-restore) | Use to distinguish data corruption recovery from availability features. | ## Source-grounding rules - Do not approve multi-region readiness without SDK region preference, failover mode, capacity, private DNS, and DR drill evidence. - Do not approve consistency changes without application correctness and throughput impact review. - Do not approve partition or throughput changes from docs alone; require workload evidence. - Treat preview features and failover modes as explicit caveats, not default recommendations. -
safety-checklist.md 2.3 KB
# Safety checklist for Azure Cosmos DB Platform Operator ## Non-negotiable gates - Never ask for account keys, connection strings, full documents, customer data, tenant identifiers, subscription identifiers, or database dumps. - Do not approve production posture without evidence for regions, consistency, throughput mode, partitioning, backups, private networking, alerts, and failover drills. - Do not perform or recommend account control-plane changes during an affected-region outage unless Microsoft guidance explicitly supports the action. - Require explicit approval before changing consistency, regions, failover priority, multi-write, throughput, indexing, backup policy, private endpoints, firewall, or network settings. - Treat SDK configuration as part of platform reliability; account settings alone do not prove application failover. ## High-risk assumptions to kill - "Multi-region means no downtime." SDK routing, capacity in remaining regions, consistency, and app failover still matter. - "Service-managed failover is fast enough." Microsoft Learn notes it can take significant time; forced failover may be needed for faster write recovery. - "Private endpoint works after failover automatically." Regional private endpoints and DNS forwarding need explicit design. - "Strong consistency is always safest." It changes availability and throughput tradeoffs. - "Backups are the DR plan." Backups protect against corruption/deletion; they are not a substitute for availability architecture. ## Evidence labels - `docs_only`: Microsoft Learn guidance only. - `sampled_read_only`: account, metric, or network evidence was sampled safely. - `config_review`: IaC or sanitized account configuration was reviewed but not proven live. - `dr_drill_proven`: failover or restore behavior was tested and evidence exists. - `mutation_ready`: blast radius, approval, and rollback are documented. ## Minimum safe evidence - API type, account tier, regions, write model, failover mode, consistency, and backup policy. - Throughput mode, hot partition indicators, RU headroom, and capacity during region loss. - Partition key strategy and hierarchical partition key fit if used. - Private endpoint layout, private DNS zones, hub/spoke forwarding, and application regional routing. - Resource Health, Service Health, metrics, diagnostic logs, alerts, and DR runbook ownership. -
workflow-and-output.md 1.6 KB
# Workflow and output contract for Azure Cosmos DB Platform Operator ## Minimal safe workflow 1. Classify the request: readiness review, region/failover design, throughput review, consistency review, partition review, network review, or mutation approval. 2. Ground the answer in Microsoft Learn through the user's configured documentation MCP. 3. Identify account scope without exposing sensitive identifiers: API, regions, write model, consistency, throughput, backup, and network posture. 4. Separate documentation evidence from sampled read-only evidence, config review, and inference. 5. Stress test: partitioning, RU headroom, failover behavior, SDK region routing, private DNS, backups, diagnostics, and operational ownership. 6. Produce a verdict with blockers, safe next actions, and open questions. 7. For mutations, stop for explicit approval after blast-radius and rollback review. ## Output contract ```markdown ## Verdict <go | conditional-go | no-go | docs-only advisory> ## Evidence level - Documentation: <sources used> - Current-state evidence: <sampled_read_only | config_review | not sampled> ## Findings 1. <finding> — Evidence: <docs_only|sampled_read_only|config_review|inference> ## Reliability and recovery posture - Regions/failover: <known/gap> - Backup/restore: <known/gap> ## Blockers - <production blocker> ## Safe next actions - <least-risk action> ``` ## Pushback triggers Push back on region changes during outages, untested failover claims, private endpoint DNS hand-waving, consistency changes with no correctness rationale, or throughput changes that ignore partition and workload evidence.
-
-
metadata.json 1.7 KB
{ "id": "azure-cosmosdb-platform-operator", "name": "Azure Cosmos DB Platform Operator", "version": "0.1.3", "type": "skill", "provider": "azure", "harnesses": [ "codex", "claude-code", "cursor", "gemini", "kiro", "other" ], "summary": "Review and operate Azure Cosmos DB platform posture across accounts, databases, containers, partitioning, throughput, consistency, indexing, throttling, multi-region tradeoffs, and operational guardrails with explicit evidence-versus-inference handling.", "source_type": "original", "official_docs": [ "https://learn.microsoft.com/en-us/azure/cosmos-db/partitioning", "https://learn.microsoft.com/en-us/azure/cosmos-db/modeling-data", "https://learn.microsoft.com/en-us/azure/cosmos-db/consistency-levels", "https://learn.microsoft.com/en-us/azure/cosmos-db/how-to-manage-consistency", "https://learn.microsoft.com/en-us/azure/cosmos-db/query-metrics", "https://learn.microsoft.com/en-us/azure/well-architected/service-guides/cosmos-db", "https://learn.microsoft.com/en-us/azure/cosmos-db/hierarchical-partition-keys", "https://learn.microsoft.com/en-us/azure/reliability/reliability-cosmos-db", "https://learn.microsoft.com/en-us/azure/cosmos-db/hierarchical-partition-keys-unlimited-scale", "https://learn.microsoft.com/en-us/azure/cosmos-db/failover-considerations-for-private-endpoints" ], "security_notes": "Do not approve a partition key, indexing posture, consistency change, or cross-partition query strategy without checking workload shape, RU impact, transactional scope, and least-privilege access implications.", "last_verified": "2026-06-05", "path": "skills/azure/azure-cosmosdb-platform-operator", "author": "github: VincentChuWaiChow" } -
SKILL.md 3 KB
--- name: azure-cosmosdb-platform-operator description: Use this skill for Azure Cosmos DB platform operations and design review, especially accounts, databases, containers, partition-key design, throughput and RU posture, consistency choices, indexing, throttling, multi-region replication, private connectivity, and Cosmos DB evidence-guided discovery. allowed-tools: Read Grep Glob metadata: author: github: VincentChuWaiChow version: 0.1.3 updated: "2026-06-05" category: platform --- # Azure Cosmos DB Platform Operator ## Purpose Review and guide Azure Cosmos DB platform posture without hand-wavy assumptions about partitioning, consistency, throughput, or multi-region behavior. ## When to use Use this skill when the user asks for: - Azure Cosmos DB account, database, or container operational review, - partition-key, consistency, throughput, or indexing decisions, - RU cost, throttling, hot partition, or multi-region tradeoff analysis, - private endpoint, failover, or platform control-plane posture questions, - Cosmos DB documentation-grounded discovery or sampled read-only evidence gathering. Do not use this skill as a substitute for: - application code implementation details unless they directly affect platform posture, - generic landing-zone or network architecture when Cosmos DB is incidental, - RBAC-only analysis when the real problem is identity governance rather than database platform design, - vector-search-specific Mongo vCore implementation unless the user explicitly asks for that API surface. ## Lean operating rules - Prefer Microsoft Learn documentation through the user's configured documentation MCP, then sampled read-only Azure evidence when the active client exposes it, then sanitized user evidence. - Separate confirmed facts from inference. If state was not queried or shown, say so. - Challenge broad scope, vague partition keys, and RU-blind advice. - Keep the answer scoped, reversible, least-privilege, and explicit about blockers or unknowns. ## References Load these only when needed: - [Operations guide](references/cosmosdb-platform-operations.md) — use for service-specific pitfalls, design rules, verification targets, and pushback criteria. - [MCP and evidence path](references/mcp-and-evidence.md) — use when choosing documentation-based evidence, sampled read-only Azure evidence, or sanitized user evidence. - [Safety checklist](references/safety-checklist.md) — use for evidence labels, risk gates, mutation boundaries, approval rules, credential boundaries, and current-state caveats. - [Workflow and output contract](references/workflow-and-output.md) — use when executing the full review, applying stress checks, or formatting the final answer. - [Official sources](references/official-sources.md) — use when you need the detailed Microsoft documentation list or source notes. ## Response minimum Return, at minimum: - the scoped target and evidence level, - the main design or operational risks, - the safest next actions, - the assumptions or blockers that prevent stronger conclusions.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.