Claude Cursor Skill

cargo-observability

Watch a Cargo workspace and get told when something breaks — scheduled threshold alerts over workflow telemetry (spans, runs, records), a storage model freshness or row count, or any SQL query, firing a connector, tool, or agent when a metric breaches. Triggers: "alert me when",

LLM Mart · 0 points · 18 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download getcargohq-cargo-skills-cargo-observability-addc862.zip · 14 KB
Part of getcargohq/cargo-skills — 22 skills

Install

skills CLI npx skills add https://github.com/getcargohq/cargo-skills/tree/main/cargo-observability
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install getcargohq-cargo-skills@llmmart
Git git clone https://github.com/getcargohq/cargo-skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole getcargohq/cargo-skills collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Cargo CLI — Observability

Alerts. An alert is a scheduled threshold check. On every cron tick it measures a scope (what to watch), compares the measured value against a threshold (the breach condition), and on breach fires actions — each as its own run — and records an event. This is the proactive counterpart to cargo-diagnostics: diagnostics explains a failure after you notice it; an alert tells you the moment a metric crosses a line.

Everything lives under one CLI domain:

cargo-ai observability alert   …   # the alert CRUD + preview surface
cargo-ai observability event   …   # an alert's firing history

Bootstrap

Already signed in (cargo-ai whoami returns a workspace)? Skip to the next section.

npm install -g @cargo-ai/cli            # no global install? prefix every command with `npx @cargo-ai/cli`
cargo-ai login --email you@company.com  # emailed code, no browser; creates the account on first use
                                        # alternatives: --oauth (browser) · --token <api-token> (CI)
cargo-ai whoami                         # confirm the active workspace before any write

Every command prints JSON to stdout; failures exit non-zero with {"errorMessage": "..."}. Anything that creates a run or a batch is async — pass --wait-until-finished or poll the matching get. Alerts are guarded by observability:read / observability:write permissions. If a create/update/remove returns a permission error, the token lacks observability:write — use an admin token or have one granted (../cargo-workspace-management/SKILL.md). When the full skill bundle is installed, ../cargo/references/prerequisites.md adds the CLI version pin, token scopes, and the admin-only surface.

The three moving parts of every alert

Part Flag What it is
Scope --scope <json> What to measure — one of six sources: spans, runs, records, orchestrationQuery, storageQuery, model.
Threshold --threshold <json> When it breaches — a metric + operator (gte/lte) + value. The metric menu depends on the scope.
Actions --actions <json> What happens on breach — an Action[] (connector / tool / agent / native nodes), each fired as its own run. Optional; omit for a silent alert whose breaches you read from its events.

The scope and threshold are a matched pair — a metric can only be computed over the scopes that produce it (e.g. errorRate needs telemetry, freshness needs a model). The full compatibility matrix, every metric's meaning and units, and every scope filter field are in references/scopes-and-thresholds.md — read it before writing a --scope/--threshold pair you haven't used before.

The golden rule: preview before you create

alert preview evaluates a scope + threshold right now, without firing actions or writing an event. It returns the value the alert would measure and whether that value breaches — so you calibrate the threshold against reality instead of guessing, and you confirm the scope/threshold pairing is even valid before committing it to a schedule.

cargo-ai observability alert preview \
  --scope '{"kind":"runs","workflowUuid":"<uuid>","statuses":["error"]}' \
  --threshold '{"metric":"errorRate","operator":"gte","value":10}' \
  --window-minutes 1440          # last 24h; default 60
  • outcome: "computed" → { value, total, failed, isBreached }. Set your threshold from value.
  • outcome: "empty" → the window had nothing to measure (see the empty-vs-zero rule in references/alert-lifecycle.md).
  • outcome: "notComputed" → { errorMessage }. A bad SQL query, a deleted model, or an invalid scope/threshold pairing all land here — fix it before creating.

--window-minutes only shapes the window for telemetry scopes (spans/runs/records). A model is measured as it stands right now; a query scope windows itself in its SQL.

Always preview first. It is free, it is the only way to size a threshold correctly, and it catches an invalid pairing before it becomes a schedule that writes an error event every tick.

Commands

All commands output JSON. Reads need a token with observability:read; create/update/remove need observability:write (an admin token has both; a plain member token may not — see Bootstrap above).

Create an alert

cargo-ai observability alert create \
  --name "CRM sync error rate" \
  --description "Page when the HubSpot sync starts failing" \
  --cron "*/30 * * * *" \
  --scope '{"kind":"runs","workflowUuid":"<workflow-uuid>","statuses":["error"]}' \
  --threshold '{"metric":"errorRate","operator":"gte","value":10}' \
  --actions '[{"kind":"agent","agentUuid":"<agent-uuid>","config":{"message":"{{alert.name}} breached: {{event.value}}% errors. {{alert.url}}"}}]'
  • --cron — 5-field cron or @every <interval> (e.g. @every 30m), always UTC, at most once a minute. The UI presets bottom out at 30 minutes; go tighter only with reason (every tick scans ClickHouse and can fire paid runs).
  • --disabled — create it paused (evaluate nothing until you update --enabled true).
  • --folder <uuid> — file it under a folder (from cargo-workspace-management).
  • --actions — optional. Omit for a silent alert. Each entry is a configured action: unlike orchestration action execute, which carries no config at all, an alert action requires one — that is where the templated message lives. The config is templated against the firing context ({{alert.*}}, {{event.*}}) — see references/alert-lifecycle.md for the full variable list. Each action's target (agentUuid/toolUuid/connectorUuid) is validated to exist in the workspace at create time.

List, get, update, remove

cargo-ai observability alert list                      # all alerts, each with its lastEvent
cargo-ai observability alert get <uuid>                # one alert + its lastEvent

cargo-ai observability alert update --uuid <uuid> \
  --enabled false                                      # pause it (true/false — must be literal)
cargo-ai observability alert update --uuid <uuid> \
  --threshold '{"metric":"errorRate","operator":"gte","value":20}'   # raise the bar
cargo-ai observability alert update --uuid <uuid> \
  --description none                                    # "none" clears; --folder none unfiles

cargo-ai observability alert remove <uuid>

--enabled is strict: only the literal true or false are accepted — --enabled yes is rejected rather than silently disabling the alert. On update, any flag you omit is left unchanged; --description none / --folder none are the explicit "clear it" spellings.

Inspect firing history

cargo-ai observability event list <alertUuid>          # latest evaluation events, newest first

Each event carries status (healthy / unhealthy / error), the measured value, a snapshot of the scope/threshold/actions as they were when it fired (the alert can change afterwards), runUuids (the runs the actions spawned — feed these to cargo-diagnostics or orchestration run get), the evaluation window, and errorMessage for error events. unhealthy = breached and fired; error = the metric could not be computed.

How evaluation actually works

The lifecycle — cron windows and the ClickHouse indexing lag, the at-most-once firing guarantee (an alert never re-fires on the same rows; a sustained breach is re-detected on the next tick), the empty-window-vs-real-zero rule that makes lte a dead-man's switch, and the full {{alert.*}}/{{event.*}} templating context — is documented in references/alert-lifecycle.md. Read it before you rely on an alert for anything time-sensitive.

Worked recipes

references/examples/recipes.md — copy-paste starting points: error-rate pager, credit-budget guard, p95-latency watch, a dead-man's switch (count lte 0 — alert when a workflow stops running), model freshness / empty-model alerts, and a custom SQL-query alert.

Declarative alternative: defineAlert (CDK)

This skill is the imperative surface — one-off cargo-ai observability alert … calls. To manage an alert as code (in git, reproducible, deployed alongside the workflow it watches), use CDK's defineAlert builder instead — see ../cargo-project/SKILL.md and "Declarative vs imperative" in the router. Same scope/threshold/action model; different authoring mode.

Cost discipline

An alert's actions fire as real runs — if an action calls a paid connector action or an agent, every breach re-bills. A poorly-sized threshold on a tight cron can breach (and bill) every tick. Two safeguards:

  • Preview to size the threshold so it fires on genuine anomalies, not normal variance.
  • If an action node calls a credits-based provider action, treat it like any scheduled paid workflow: read that provider's playbook (esp. its Recurring use section) in ../cargo-gtm/provider-playbooks/, and apply the spend rules in ../cargo-gtm/references/cost-discipline.md. Prefer cheap notification actions (an agent that posts to Slack, a connector notification) over anything that fans out.

When the CLI surprises you

If a documented flag, scope field, or response shape doesn't match what you observe (a fix may have shipped, or the docs may have drifted), re-refresh the CLI and skills; if it still doesn't add up, file a report — it's read by the team:

cargo-ai workspaceManagement report create \
  --title "<one-line summary>" \
  --description "<exact command(s), errorMessage verbatim, expected vs actual, UUIDs>"

Presenting results

Follow ../cargo/references/interaction.md: lead with the outcome ("alert created, will page the on-call agent when the CRM sync's error rate hits 10% over 30 min"), summarize an alert or its events as a compact table, never dump raw alert get / event list JSON into the conversation.

Files (cargo-skills)
  • references
    • examples
      • recipes.md 6.6 KB
        # Alert recipes
        
        Copy-paste starting points. Every one is **`preview` first** to size the threshold, then `create`. Replace `<…>` placeholders with real UUIDs (discover them via `cargo-orchestration` / `cargo-storage` / `cargo-ai`). See `../scopes-and-thresholds.md` for every field and `../alert-lifecycle.md` for firing semantics.
        
        ---
        
        ## 1. Error-rate pager for one workflow
        
        Page the on-call agent when a workflow's runs start failing.
        
        ```bash
        # Preview against the last 24h to see the current rate
        cargo-ai observability alert preview \
          --scope '{"kind":"runs","workflowUuid":"<workflow-uuid>"}' \
          --threshold '{"metric":"errorRate","operator":"gte","value":10}' \
          --window-minutes 1440
        
        # Create it on a 30-min cadence with an agent notification
        cargo-ai observability alert create \
          --name "CRM sync — error rate" \
          --description "Error rate ≥10% over 30 min" \
          --cron "*/30 * * * *" \
          --scope '{"kind":"runs","workflowUuid":"<workflow-uuid>"}' \
          --threshold '{"metric":"errorRate","operator":"gte","value":10}' \
          --actions '[{"kind":"agent","agentUuid":"<agent-uuid>","config":{"message":"🚨 {{alert.name}}: {{event.value}}% errors ({{event.windowStart}}–{{event.windowEnd}}). {{alert.url}}"}}]'
        ```
        
        `errorRate` is a percent of *finished* runs; a quiet window with nothing finished is `empty`, not a false `0%`.
        
        ---
        
        ## 2. Credit-budget guard
        
        Catch cost blowouts before the invoice does — total credits burned by a workflow over the window.
        
        ```bash
        cargo-ai observability alert create \
          --name "Enrichment play — hourly credit ceiling" \
          --cron "0 * * * *" \
          --scope '{"kind":"spans","workflowUuid":"<workflow-uuid>","nodeKind":"connector"}' \
          --threshold '{"metric":"credits","aggregation":"sum","operator":"gte","value":500}' \
          --actions '[{"kind":"agent","agentUuid":"<agent-uuid>","config":{"message":"💸 {{alert.name}} burned {{event.value}} credits in the last hour ({{event.threshold}} ceiling). {{alert.url}}"}}]'
        ```
        
        Scoping to `nodeKind: "connector"` measures only the paid provider calls. Preview at a couple of window sizes to learn the normal hourly spend before setting `value`.
        
        ---
        
        ## 3. p95 latency watch
        
        Fire when a workflow (or a specific node) gets slow.
        
        ```bash
        cargo-ai observability alert preview \
          --scope '{"kind":"spans","workflowUuid":"<workflow-uuid>","nodeActionSlug":"<action-slug>"}' \
          --threshold '{"metric":"duration","aggregation":"p95","operator":"gte","value":30}' \
          --window-minutes 180
        ```
        
        Use the previewed `value` to set a realistic `p95` threshold, then `create` with a `--cron`. Aggregations: `avg` | `p50` | `p95` | `p99`.
        
        ---
        
        ## 4. Dead-man's switch — alert when a workflow STOPS running
        
        The one pattern that needs `count lte 0`: `count` reports a real `0` on an empty window (not `empty`), so silence breaches.
        
        ```bash
        cargo-ai observability alert create \
          --name "Nightly sync — did it run?" \
          --description "No runs in the last 24h = something broke upstream" \
          --cron "0 8 * * *" \
          --scope '{"kind":"runs","workflowUuid":"<workflow-uuid>"}' \
          --threshold '{"metric":"count","operator":"lte","value":0}' \
          --actions '[{"kind":"agent","agentUuid":"<agent-uuid>","config":{"message":"⚠️ {{alert.name}}: the nightly sync produced no runs. {{alert.url}}"}}]'
        ```
        
        The cron interval **is** the window — run this once daily so "no runs" means "none in the last day".
        
        ---
        
        ## 5. Model freshness — stale sync
        
        Breach when a model hasn't emitted new data in too long. `freshness` is in **minutes** and ignores the scope filter.
        
        ```bash
        cargo-ai observability alert create \
          --name "Companies model — freshness" \
          --cron "*/30 * * * *" \
          --scope '{"kind":"model","modelUuid":"<model-uuid>"}' \
          --threshold '{"metric":"freshness","operator":"gte","value":120}' \
          --actions '[{"kind":"agent","agentUuid":"<agent-uuid>","config":{"message":"🕒 {{alert.name}}: {{event.value}} min since last emit. {{alert.url}}"}}]'
        ```
        
        Related model metrics: `syncDuration gte <seconds>` (sync got slow), `recordsShare gte <percent>` (a filtered slice grew — needs a scope `filter`).
        
        ---
        
        ## 6. Empty-model dead-man's switch
        
        `recordsCount` returns a real `0` on an empty model, so `lte 0` catches a model that emptied out or never populated.
        
        ```bash
        cargo-ai observability alert create \
          --name "Leads model — not empty" \
          --cron "0 */6 * * *" \
          --scope '{"kind":"model","modelUuid":"<model-uuid>"}' \
          --threshold '{"metric":"recordsCount","operator":"lte","value":0}' \
          --actions '[{"kind":"agent","agentUuid":"<agent-uuid>","config":{"message":"📉 {{alert.name}}: the model is empty. {{alert.url}}"}}]'
        ```
        
        Add a `filter` (segmentation shape, spelled `conjonction`) to count only a slice — e.g. records missing an enrichment column, then alert if that count climbs with `gte`.
        
        ---
        
        ## 7. Custom SQL-query alert
        
        When no built-in metric fits, compute the value yourself. The query must return a **single number** and window itself.
        
        Orchestration runtime (error rate over the last hour, self-windowed):
        
        ```bash
        cargo-ai observability alert create \
          --name "Workspace-wide error rate" \
          --cron "*/15 * * * *" \
          --scope '{"kind":"orchestrationQuery","query":"select countIf(status = '"'"'error'"'"') * 100 / count() from runs where created_at > now() - interval 1 hour"}' \
          --threshold '{"metric":"query","operator":"gte","value":5}'
        ```
        
        Storage warehouse (records missing enrichment):
        
        ```bash
        cargo-ai observability alert create \
          --name "Unenriched companies backlog" \
          --cron "0 */4 * * *" \
          --scope '{"kind":"storageQuery","query":"select count() from default.companies where enriched_at is null"}' \
          --threshold '{"metric":"query","operator":"gte","value":1000}'
        ```
        
        Notes:
        - Validate the SQL with `cargo-ai orchestration query execute` / `cargo-ai storage query execute` first, then `alert preview` the whole scope+threshold.
        - An aggregate over no rows is `NULL` → treated as `empty` (won't breach `lte`). For a "went silent" query alert, return a real `0` via `count()`.
        - Shell-quoting SQL with single quotes is fiddly — the `'"'"'` dance above escapes a literal `'`. Alternatively build the JSON in a file and pass `--scope "$(cat scope.json)"`.
        
        ---
        
        ## Reading the results
        
        ```bash
        cargo-ai observability alert list                 # every alert + its lastEvent status/value
        cargo-ai observability event list <alert-uuid>    # firing history, newest first
        ```
        
        An `unhealthy` event's `runUuids` are the runs its actions spawned — trace them with `cargo-ai orchestration run get <uuid>` or hand them to `cargo-diagnostics`. An `error` event means the metric couldn't be computed (bad query, deleted model) — read its `errorMessage` and fix the scope.
        
    • alert-lifecycle.md 6.9 KB
      # Alert lifecycle — how evaluation and firing actually work
      
      What happens on each cron tick, why an alert never double-fires, the empty-window rule, and the full templating context an action gets. Read this before you rely on an alert for anything time-sensitive.
      
      ## One evaluation tick
      
      On every scheduled tick, for an enabled alert:
      
      1. **Build the window.** `windowEndedAt = now − ClickHouse indexing lag` (spans land through Kinesis + a materialized view, so the tail of the window is left for the next tick to avoid missing late rows). `windowStartedAt = the alert's last `lastEvaluatedAt`, or its `updatedAt` on the first ever tick. Telemetry scopes measure over `[windowStartedAt, windowEndedAt]`; a `model` is measured point-in-time; a query scope windows itself.
      2. **Compute the value** for the scope + threshold (see `scopes-and-thresholds.md`).
      3. **Claim the window atomically** — advance the cursor (`lastEvaluatedAt → windowEndedAt`) only if the alert is still enabled, not deleted, and no other activity already advanced it. If the claim is lost, the tick records nothing and fires nothing.
      4. **Record an event** and, on breach, **fire the actions**.
      
      If `windowStartedAt >= windowEndedAt` (an empty or already-evaluated window — e.g. a retry, or overlapping ticks during a schedule change) the tick is a no-op: an alert never re-fires on spans it already saw.
      
      ## The three event outcomes
      
      Every tick that claims its window writes exactly one event:
      
      | Compute outcome | Event `status` | Fires actions? | `value` |
      | --- | --- | --- | --- |
      | breached | `unhealthy` | **yes** | the measured value |
      | not breached | `healthy` | no | the measured value |
      | `empty` (nothing to measure) | `healthy` | no | `null` |
      | `notComputed` (bad SQL, deleted model, corrupt pairing) | `error` | no | `null` (+ `errorMessage`) |
      
      `event list <alertUuid>` returns these newest-first. A run of `unhealthy` events is a sustained breach; an `error` event means the alert can't measure what it was told to — fix the scope/query/model.
      
      ## At-most-once firing (and what that means for you)
      
      Firing is deliberately **at-most-once**, not at-least-once. The cursor is claimed *before* actions fire, so a Temporal retry or an overlapping cron can't fire the same breach twice. Actions spawn runs — which cost credits and can include agents that open PRs or send messages — so a rare *miss* is preferred to a *duplicate*.
      
      Consequences:
      
      - **A sustained breach is re-detected, not re-fired on the same rows.** Each tick only sees rows since the last cursor advance. If the condition is still breaching on the *next* window's rows, you get another `unhealthy` event then. So on a 30-minute cron, an ongoing error spike pages roughly every 30 minutes — it does not spam.
      - **A one-tick blip fires once.** Good for "tell me the moment X happens".
      - **Disabling or deleting an alert mid-evaluation cancels the firing** for that tick.
      - **Design actions to be safe to receive repeatedly** (a notification, an idempotent ticket), since a long breach produces one firing per tick it's true for.
      
      ## The empty-vs-zero rule
      
      An idle/empty window is reported as **`empty`** (→ `healthy`, no fire) for almost every metric — you don't want a latency or error-rate alert firing "0" every quiet hour.
      
      The exceptions are **`count`** (telemetry scopes) and **`recordsCount`** (model scope): an empty window is a real **`0`**. Paired with **`lte 0`** they become **dead-man's switches** that breach *because* nothing happened — the only way to alert on *absence* (a workflow that stopped, a model that emptied). See `scopes-and-thresholds.md` for the per-metric table.
      
      The same principle protects query alerts: an aggregate over no rows is `NULL` (ClickHouse also renders `NaN`/`0÷0` as `NULL`), which the alert treats as `empty` rather than `0` — so a rate query on an idle window won't false-breach an `lte` threshold. If you *want* silence to breach, write a `count()` that returns a genuine `0`.
      
      ## Actions: what fires, and the templating context
      
      On breach, **each action in `--actions` is fired as its own run** through the orchestration action service (`skipConcurrencyCheck` is on — an alert must fire even when the workspace is at its run-concurrency limit). All the runs of one firing share a single trace; their uuids are stored on the event as `runUuids`. Firing is best-effort per action: one action failing to start is logged and skipped, never blocking the others or the event.
      
      `--actions` is the shared orchestration `Action[]` union — the same shape used everywhere in `cargo-orchestration`:
      
      ```json
      [
        {"kind":"agent","agentUuid":"…","config":{ "message":"…" }},
        {"kind":"connector","integrationSlug":"…","actionSlug":"…","config":{ … }},
        {"kind":"tool","toolUuid":"…","config":{ … }},
        {"kind":"native","actionSlug":"…","config":{ … }}
      ]
      ```
      
      Each action's target (`agentUuid` / `toolUuid` / `connectorUuid`) is **validated to exist in the workspace** at create/update time — a mistyped uuid is rejected up front, not silently at breach. (An action whose target is deleted *afterwards* fails at fire time and is recorded on the event.)
      
      ### Firing context (templating)
      
      Before each action runs, its `config` is interpolated against the firing context. The action's evaluated `config` becomes the run's input **data** (the action executes with an empty config), so put your bindings in `config`:
      
      | Variable | Value |
      | --- | --- |
      | `{{alert.uuid}}` | The alert's UUID. |
      | `{{alert.name}}` | The alert's name. |
      | `{{alert.url}}` | Deep link to the alert in the app. |
      | `{{event.value}}` | The measured value (rounded to 2 dp). |
      | `{{event.threshold}}` | The threshold `value` it crossed. |
      | `{{event.operator}}` | `gte` / `lte`. |
      | `{{event.windowStart}}` | Window start, ISO 8601. |
      | `{{event.windowEnd}}` | Window end, ISO 8601. |
      | `{{event.spansUrl}}` | Deep link to the workspace's spans view. |
      
      Example agent action that composes a human-readable page:
      
      ```json
      [{"kind":"agent","agentUuid":"<agent-uuid>","config":{
        "message":"🚨 {{alert.name}} breached: {{event.value}} {{event.operator}} {{event.threshold}} over {{event.windowStart}}–{{event.windowEnd}}. Alert: {{alert.url}} · Spans: {{event.spansUrl}}"
      }}]
      ```
      
      ## Schedules
      
      - `--cron` accepts a 5-field cron expression **or** `@every <interval>` (e.g. `@every 15m`), evaluated in **UTC**.
      - Minimum interval is **once a minute** — tighter `@every` values are rejected (every tick scans ClickHouse and may fire paid runs). Named-weekday cron expressions the interval parser can't measure are allowed through, since a 5-field cron can't fire more than once a minute anyway.
      - The UI presets bottom out at **30 minutes**; the CDK template defaults to **5 minutes**. Pick the loosest cadence that still catches the problem in time — it's cheaper and quieter.
      - Create with `--disabled` to stage an alert without evaluating it; `alert update --uuid <uuid> --enabled true` starts the schedule. `--enabled false` pauses it (the schedule stops; the alert and its history remain).
      
    • scopes-and-thresholds.md 8.1 KB
      # Scopes & thresholds — the compatibility matrix
      
      An alert measures a **scope** and breaches on a **threshold**. They are a matched pair: each metric can only be computed over the scopes that produce it. Get the pairing wrong and the alert (or a `preview`) returns `outcome: "notComputed"` with `The "<metric>" metric cannot be computed over a "<scope>" scope.`
      
      ## The matrix at a glance
      
      | Scope `kind` | Source | Allowed threshold metrics |
      | --- | --- | --- |
      | `spans` | Per-node executions of a workflow (ClickHouse) | `errorRate`, `duration`, `credits`, `count` |
      | `runs` | Whole runs (the eight-value run status) | `errorRate`, `duration`, `credits`, `count` |
      | `records` | One row per record, latest state | `errorRate`, `duration`, `credits`, `count` |
      | `orchestrationQuery` | Your SQL over `runs`/`batches`/`spans`/`records` | `query` |
      | `storageQuery` | Your SQL over the workspace data warehouse | `query` |
      | `model` | A storage model's records + sync state | `recordsCount`, `recordsShare`, `freshness`, `syncDuration` |
      
      `operator` is always `gte` or `lte`; `value` is always a number.
      
      ---
      
      ## Scopes (the `--scope` JSON)
      
      ### Telemetry scopes — `spans`, `runs`, `records`
      
      All three are windowed over the evaluation interval and share the same four metrics. They differ in what they count and how they filter.
      
      **`spans`** — one row per node execution. The richest filter set:
      
      ```json
      {
        "kind": "spans",
        "workflowUuid": "…",
        "parentAgentUuid": "…",
        "nodeKind": "native | connector | tool | agent",
        "nodeIntegrationSlug": "…",
        "nodeConnectorUuid": "…",
        "nodeActionSlug": "…",
        "nodeToolUuid": "…",
        "nodeAgentUuid": "…",
        "executionTitleOrErrorMessage": "substring match",
        "executionStatuses": ["pending", "success", "error"],
        "userUuid": "…"
      }
      ```
      
      Every field is optional; omit them all to watch every span in the workspace. Use `nodeActionSlug` / `nodeConnectorUuid` to pin the alert to a single provider action, `nodeAgentUuid` to watch one agent's calls.
      
      **`runs`** — one row per run. Filters on the **full** run status set (not the three execution statuses):
      
      ```json
      {
        "kind": "runs",
        "workflowUuid": "…",
        "statuses": ["error"],
        "releaseUuid": "…",
        "recordTitleOrErrorMessage": "substring match",
        "userUuid": "…"
      }
      ```
      
      **`records`** — one row per record holding its latest state (no idle/skipped; the same work as `runs`, keyed by record):
      
      ```json
      {
        "kind": "records",
        "workflowUuid": "…",
        "statuses": ["error"],
        "releaseUuid": "…",
        "titleOrErrorMessage": "substring match",
        "userUuid": "…"
      }
      ```
      
      > `statuses` on `runs` uses the full run-status enum; on `records` it uses the record-status enum. When in doubt, `preview` with the statuses you want and check the `total`/`failed` counts. (Discover valid status values from `cargo-orchestration`.)
      
      ### Query scopes — `orchestrationQuery`, `storageQuery`
      
      You supply the SQL; it must return a **single numeric value** (the first column of the first row). The query is expected to **window itself** — `--window-minutes` does not apply.
      
      ```json
      { "kind": "orchestrationQuery", "query": "select countIf(status='error')*100/count() from runs where created_at > now() - interval 1 hour" }
      ```
      
      ```json
      { "kind": "storageQuery", "query": "select count() from default.companies where enriched_at is null" }
      ```
      
      - `orchestrationQuery` runs against the orchestration runtime tables (`runs`, `batches`, `spans`, `records`; no schema prefix; workspace-scoped) — same engine as `cargo-ai orchestration query execute`.
      - `storageQuery` runs against the workspace data warehouse using `<datasetSlug>.<modelSlug>` table names — same engine as `cargo-ai storage query execute`. If no warehouse is connected, the alert errors with *"No data warehouse is connected to this workspace."*
      - An aggregate over no rows is `NULL` (and ClickHouse renders `NaN`/`0/0` as `NULL` too) → the evaluation is treated as **`empty`**, not `0`. So a rate query on an idle window won't false-breach an `lte` threshold. If you want silence to breach, use a `count()` that returns a real `0` (see the dead-man's switch recipe).
      
      Pair either query scope with the `query` threshold — the SQL computes the value, the threshold just carries the comparison:
      
      ```json
      { "metric": "query", "operator": "gte", "value": 10 }
      ```
      
      ### Model scope — `model`
      
      Watches a storage model's records and its sync health:
      
      ```json
      { "kind": "model", "modelUuid": "…", "filter": { "conjonction": "and", "groups": [ … ] } }
      ```
      
      - `filter` is optional and uses the **segmentation filter shape** — note the spelling **`conjonction`** (silently ignored if misspelled). It narrows the record metrics (`recordsCount`, `recordsShare`) and is **ignored** by the sync metrics (`freshness`, `syncDuration`). Discover the filter shape and model UUIDs from `cargo-storage` / `cargo-orchestration`.
      - A model is measured **point-in-time** — as it stands at the tick. The cron controls *how often* it's checked, not what's measured; `--window-minutes` doesn't apply.
      
      ---
      
      ## Thresholds (the `--threshold` JSON)
      
      ### Telemetry metrics (for `spans` / `runs` / `records`)
      
      | Metric | Extra field | Value means | Empty window |
      | --- | --- | --- | --- |
      | `errorRate` | — | `failed × 100 / finished` (**percent**). `total` on the event is the finished denominator. | Nothing *finished* → `empty` (healthy, no fire). |
      | `duration` | `aggregation`: `avg`\|`p50`\|`p95`\|`p99` | The chosen aggregate of duration **in seconds**, over finished rows only. Preview to see the current level before setting `value`. | No rows → `empty`. |
      | `credits` | `aggregation`: `sum`\|`avg`\|`p95` | Credit spend aggregated over the window. | No rows → `empty`. |
      | `count` | — | Number of rows **started** in the window (running ones included) — "did work happen". | A real **`0`**, *not* `empty` — so `count lte 0` is a dead-man's switch. |
      
      ```json
      {"metric":"errorRate","operator":"gte","value":10}
      {"metric":"duration","aggregation":"p95","operator":"gte","value":30}
      {"metric":"credits","aggregation":"sum","operator":"gte","value":500}
      {"metric":"count","operator":"lte","value":0}
      ```
      
      ### Query metric (for `orchestrationQuery` / `storageQuery`)
      
      ```json
      {"metric":"query","operator":"gte","value":10}
      ```
      
      ### Model metrics (for `model`)
      
      | Metric | Value means | Notes |
      | --- | --- | --- |
      | `recordsCount` | Live count of records matching the scope `filter`. | An empty model is a real **`0`** (not `empty`), so `recordsCount lte 0` is a dead-man's switch for an empty/broken model. |
      | `recordsShare` | `matching × 100 / total` (**percent**) — the filter's share of the model. | Needs a scope `filter` to be meaningful. A model with **0** total records is `empty` (no share to report), not `0%`. Costs two warehouse queries. |
      | `freshness` | **Minutes** since the model last emitted (falls back to the model's `createdAt` if it never has). | A model that never synced ages from creation, so `freshness gte <mins>` can breach a sync that never ran. |
      | `syncDuration` | **Seconds** the model's last completed sync took (`finishedAt − createdAt` of `lastRun`). | While a sync is still in flight there's no duration → `empty` (won't false-breach an `lte`). |
      
      ```json
      {"metric":"recordsCount","operator":"lte","value":0}
      {"metric":"recordsShare","operator":"gte","value":30}
      {"metric":"freshness","operator":"gte","value":60}
      {"metric":"syncDuration","operator":"gte","value":300}
      ```
      
      ---
      
      ## The empty-vs-zero rule (why it matters)
      
      Most metrics report an idle/empty window as **`empty`** → a `healthy` event, no fire. That's deliberate: you don't want an error-rate or latency alert screaming "0!" every quiet night.
      
      Two metrics are the exception and return a real **`0`** on an empty window — **`count`** (telemetry) and **`recordsCount`** (model). Paired with **`lte`**, they become **dead-man's switches**: they breach *because* nothing happened. This is the only way to alert on absence — a workflow that stopped running, a model that emptied out. Every other metric treats absence as "nothing to measure", not "a low value". See the lifecycle reference for the full statement, and the recipes for a ready-made dead-man's switch.
      
  • skill-metadata.json 872 B
    {
      "$comment": "Generated by .github/scripts/skills-metadata.mjs — do not hand-edit. Regenerate with: node .github/scripts/skills-metadata.mjs --write .",
      "name": "cargo-observability",
      "version": "1.0.3",
      "documents": [
        {
          "path": "SKILL.md",
          "kind": "entrypoint",
          "title": "Cargo CLI — Observability"
        },
        {
          "path": "references/alert-lifecycle.md",
          "kind": "reference",
          "title": "Alert lifecycle — how evaluation and firing actually work"
        },
        {
          "path": "references/examples/recipes.md",
          "kind": "example",
          "title": "Alert recipes"
        },
        {
          "path": "references/scopes-and-thresholds.md",
          "kind": "reference",
          "title": "Scopes & thresholds — the compatibility matrix"
        }
      ],
      "contentHash": "08683568fc29d3613e339821aeb169edfc42e02db816dcb52a5bd4d1430da8dc"
    }
    
  • SKILL.md 11.4 KB
    ---
    name: cargo-observability
    description: "Watch a Cargo workspace and get told when something breaks — scheduled threshold alerts over workflow telemetry (spans, runs, records), a storage model freshness or row count, or any SQL query, firing a connector, tool, or agent when a metric breaches. Triggers: \"alert me when\", \"notify me if\", \"let me know when the error rate\", \"monitor this workflow\", \"tell me if the sync stops\", \"warn me before I run out of credits\", \"dead man’s switch\", \"is this still running\", \"set up monitoring\", plus listing, previewing, editing, and reviewing an alert firing history. Skip when: diagnosing something that already went wrong — use cargo-diagnostics."
    version: "1.0.3"
    compatibility: Requires @cargo-ai/cli (npm). Sign in or create an account with `cargo-ai login --email` (emailed code, no browser), `--oauth`, or an API token
    homepage: https://github.com/getcargohq/cargo-skills
    metadata:
      author: getcargo
      openclaw:
        requires:
          bins:
            - cargo-ai
        install:
          - kind: node
            package: "@cargo-ai/cli@latest"
            bins:
              - cargo-ai
        homepage: https://github.com/getcargohq/cargo-skills
    ---
    
    # Cargo CLI — Observability
    
    **Alerts.** An alert is a scheduled threshold check. On every cron tick it measures a **scope** (what to watch), compares the measured value against a **threshold** (the breach condition), and on breach fires **actions** — each as its own run — and records an **event**. This is the proactive counterpart to `cargo-diagnostics`: diagnostics explains a failure *after* you notice it; an alert *tells you* the moment a metric crosses a line.
    
    Everything lives under one CLI domain:
    
    ```bash
    cargo-ai observability alert   …   # the alert CRUD + preview surface
    cargo-ai observability event   …   # an alert's firing history
    ```
    
    ## Bootstrap
    
    Already signed in (`cargo-ai whoami` returns a workspace)? Skip to the next section.
    
    ```bash
    npm install -g @cargo-ai/cli            # no global install? prefix every command with `npx @cargo-ai/cli`
    cargo-ai login --email you@company.com  # emailed code, no browser; creates the account on first use
                                            # alternatives: --oauth (browser) · --token <api-token> (CI)
    cargo-ai whoami                         # confirm the active workspace before any write
    ```
    
    Every command prints JSON to stdout; failures exit non-zero with `{"errorMessage": "..."}`. Anything that creates a run or a batch is async — pass `--wait-until-finished` or poll the matching `get`. Alerts are guarded by `observability:read` / `observability:write` permissions. If a create/update/remove returns a permission error, the token lacks `observability:write` — use an admin token or have one granted ([`../cargo-workspace-management/SKILL.md`](../cargo-workspace-management/SKILL.md)). When the full skill bundle is installed, [`../cargo/references/prerequisites.md`](../cargo/references/prerequisites.md) adds the CLI version pin, token scopes, and the admin-only surface.
    
    ## The three moving parts of every alert
    
    | Part | Flag | What it is |
    | --- | --- | --- |
    | **Scope** | `--scope <json>` | *What* to measure — one of six sources: `spans`, `runs`, `records`, `orchestrationQuery`, `storageQuery`, `model`. |
    | **Threshold** | `--threshold <json>` | *When it breaches* — a `metric` + `operator` (`gte`/`lte`) + `value`. The metric menu depends on the scope. |
    | **Actions** | `--actions <json>` | *What happens on breach* — an `Action[]` (connector / tool / agent / native nodes), each fired as its own run. Optional; omit for a silent alert whose breaches you read from its events. |
    
    The scope and threshold are a **matched pair** — a metric can only be computed over the scopes that produce it (e.g. `errorRate` needs telemetry, `freshness` needs a model). The full compatibility matrix, every metric's meaning and units, and every scope filter field are in **[`references/scopes-and-thresholds.md`](references/scopes-and-thresholds.md)** — read it before writing a `--scope`/`--threshold` pair you haven't used before.
    
    ## The golden rule: `preview` before you `create`
    
    `alert preview` evaluates a scope + threshold **right now, without firing actions or writing an event**. It returns the value the alert would measure and whether that value breaches — so you calibrate the threshold against reality instead of guessing, and you confirm the scope/threshold pairing is even valid before committing it to a schedule.
    
    ```bash
    cargo-ai observability alert preview \
      --scope '{"kind":"runs","workflowUuid":"<uuid>","statuses":["error"]}' \
      --threshold '{"metric":"errorRate","operator":"gte","value":10}' \
      --window-minutes 1440          # last 24h; default 60
    ```
    
    - `outcome: "computed"` → `{ value, total, failed, isBreached }`. Set your threshold from `value`.
    - `outcome: "empty"` → the window had nothing to measure (see the empty-vs-zero rule in `references/alert-lifecycle.md`).
    - `outcome: "notComputed"` → `{ errorMessage }`. A bad SQL query, a deleted model, or an **invalid scope/threshold pairing** all land here — fix it before creating.
    
    `--window-minutes` only shapes the window for telemetry scopes (`spans`/`runs`/`records`). A `model` is measured as it stands right now; a query scope windows itself in its SQL.
    
    **Always preview first.** It is free, it is the only way to size a threshold correctly, and it catches an invalid pairing before it becomes a schedule that writes an `error` event every tick.
    
    ## Commands
    
    All commands output JSON. Reads need a token with `observability:read`; create/update/remove need `observability:write` (an admin token has both; a plain member token may not — see Bootstrap above).
    
    ### Create an alert
    
    ```bash
    cargo-ai observability alert create \
      --name "CRM sync error rate" \
      --description "Page when the HubSpot sync starts failing" \
      --cron "*/30 * * * *" \
      --scope '{"kind":"runs","workflowUuid":"<workflow-uuid>","statuses":["error"]}' \
      --threshold '{"metric":"errorRate","operator":"gte","value":10}' \
      --actions '[{"kind":"agent","agentUuid":"<agent-uuid>","config":{"message":"{{alert.name}} breached: {{event.value}}% errors. {{alert.url}}"}}]'
    ```
    
    - `--cron` — 5-field cron **or** `@every <interval>` (e.g. `@every 30m`), always **UTC**, at most once a minute. The UI presets bottom out at 30 minutes; go tighter only with reason (every tick scans ClickHouse and can fire paid runs).
    - `--disabled` — create it paused (evaluate nothing until you `update --enabled true`).
    - `--folder <uuid>` — file it under a folder (from `cargo-workspace-management`).
    - `--actions` — optional. Omit for a silent alert. Each entry is a **configured** action: unlike `orchestration action execute`, which carries no `config` at all, an alert action **requires** one — that is where the templated message lives. The config is templated against the firing context (`{{alert.*}}`, `{{event.*}}`) — see [`references/alert-lifecycle.md`](references/alert-lifecycle.md) for the full variable list. Each action's target (`agentUuid`/`toolUuid`/`connectorUuid`) is validated to exist in the workspace at create time.
    
    ### List, get, update, remove
    
    ```bash
    cargo-ai observability alert list                      # all alerts, each with its lastEvent
    cargo-ai observability alert get <uuid>                # one alert + its lastEvent
    
    cargo-ai observability alert update --uuid <uuid> \
      --enabled false                                      # pause it (true/false — must be literal)
    cargo-ai observability alert update --uuid <uuid> \
      --threshold '{"metric":"errorRate","operator":"gte","value":20}'   # raise the bar
    cargo-ai observability alert update --uuid <uuid> \
      --description none                                    # "none" clears; --folder none unfiles
    
    cargo-ai observability alert remove <uuid>
    ```
    
    `--enabled` is strict: only the literal `true` or `false` are accepted — `--enabled yes` is rejected rather than silently disabling the alert. On `update`, any flag you omit is left unchanged; `--description none` / `--folder none` are the explicit "clear it" spellings.
    
    ### Inspect firing history
    
    ```bash
    cargo-ai observability event list <alertUuid>          # latest evaluation events, newest first
    ```
    
    Each event carries `status` (`healthy` / `unhealthy` / `error`), the measured `value`, a **snapshot** of the `scope`/`threshold`/`actions` as they were when it fired (the alert can change afterwards), `runUuids` (the runs the actions spawned — feed these to `cargo-diagnostics` or `orchestration run get`), the evaluation window, and `errorMessage` for `error` events. `unhealthy` = breached and fired; `error` = the metric could not be computed.
    
    ## How evaluation actually works
    
    The lifecycle — cron windows and the ClickHouse indexing lag, the **at-most-once** firing guarantee (an alert never re-fires on the same rows; a *sustained* breach is re-detected on the next tick), the empty-window-vs-real-zero rule that makes `lte` a dead-man's switch, and the full `{{alert.*}}`/`{{event.*}}` templating context — is documented in **[`references/alert-lifecycle.md`](references/alert-lifecycle.md)**. Read it before you rely on an alert for anything time-sensitive.
    
    ## Worked recipes
    
    **[`references/examples/recipes.md`](references/examples/recipes.md)** — copy-paste starting points: error-rate pager, credit-budget guard, p95-latency watch, a **dead-man's switch** (`count lte 0` — alert when a workflow *stops* running), model freshness / empty-model alerts, and a custom SQL-query alert.
    
    ## Declarative alternative: `defineAlert` (CDK)
    
    This skill is the **imperative** surface — one-off `cargo-ai observability alert …` calls. To manage an alert **as code** (in git, reproducible, deployed alongside the workflow it watches), use CDK's `defineAlert` builder instead — see [`../cargo-project/SKILL.md`](../cargo-project/SKILL.md) and "Declarative vs imperative" in the router. Same scope/threshold/action model; different authoring mode.
    
    ## Cost discipline
    
    An alert's **actions fire as real runs** — if an action calls a paid connector action or an agent, every breach re-bills. A poorly-sized threshold on a tight cron can breach (and bill) every tick. Two safeguards:
    
    - **Preview to size the threshold** so it fires on genuine anomalies, not normal variance.
    - If an action node calls a **credits-based provider action**, treat it like any scheduled paid workflow: read that provider's playbook (esp. its *Recurring use* section) in `../cargo-gtm/provider-playbooks/`, and apply the spend rules in [`../cargo-gtm/references/cost-discipline.md`](../cargo-gtm/references/cost-discipline.md). Prefer cheap notification actions (an agent that posts to Slack, a connector notification) over anything that fans out.
    
    ## When the CLI surprises you
    
    If a documented flag, scope field, or response shape doesn't match what you observe (a fix may have shipped, or the docs may have drifted), re-refresh the CLI and skills; if it still doesn't add up, file a report — it's read by the team:
    
    ```bash
    cargo-ai workspaceManagement report create \
      --title "<one-line summary>" \
      --description "<exact command(s), errorMessage verbatim, expected vs actual, UUIDs>"
    ```
    
    ## Presenting results
    
    Follow [`../cargo/references/interaction.md`](../cargo/references/interaction.md): lead with the outcome ("alert created, will page the on-call agent when the CRM sync's error rate hits 10% over 30 min"), summarize an alert or its events as a compact table, never dump raw `alert get` / `event list` JSON into the conversation.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related