cargo-observability
Watch a Cargo workspace and get told when something breaks — scheduled threshold alerts over workflow telemetry (spans, runs, records), a storage model freshness or row count, or any SQL query, firing a connector, tool, or agent when a metric breaches. Triggers: "alert me when",
Install
npx skills add https://github.com/getcargohq/cargo-skills/tree/main/cargo-observability
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install getcargohq-cargo-skills@llmmart
git clone https://github.com/getcargohq/cargo-skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole getcargohq/cargo-skills collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Cargo CLI — Observability
Alerts. An alert is a scheduled threshold check. On every cron tick it measures a scope (what to watch), compares the measured value against a threshold (the breach condition), and on breach fires actions — each as its own run — and records an event. This is the proactive counterpart to cargo-diagnostics: diagnostics explains a failure after you notice it; an alert tells you the moment a metric crosses a line.
Everything lives under one CLI domain:
cargo-ai observability alert … # the alert CRUD + preview surface
cargo-ai observability event … # an alert's firing history
Bootstrap
Already signed in (cargo-ai whoami returns a workspace)? Skip to the next section.
npm install -g @cargo-ai/cli # no global install? prefix every command with `npx @cargo-ai/cli`
cargo-ai login --email you@company.com # emailed code, no browser; creates the account on first use
# alternatives: --oauth (browser) · --token <api-token> (CI)
cargo-ai whoami # confirm the active workspace before any write
Every command prints JSON to stdout; failures exit non-zero with {"errorMessage": "..."}. Anything that creates a run or a batch is async — pass --wait-until-finished or poll the matching get. Alerts are guarded by observability:read / observability:write permissions. If a create/update/remove returns a permission error, the token lacks observability:write — use an admin token or have one granted (../cargo-workspace-management/SKILL.md). When the full skill bundle is installed, ../cargo/references/prerequisites.md adds the CLI version pin, token scopes, and the admin-only surface.
The three moving parts of every alert
| Part | Flag | What it is |
|---|---|---|
| Scope | --scope <json> |
What to measure — one of six sources: spans, runs, records, orchestrationQuery, storageQuery, model. |
| Threshold | --threshold <json> |
When it breaches — a metric + operator (gte/lte) + value. The metric menu depends on the scope. |
| Actions | --actions <json> |
What happens on breach — an Action[] (connector / tool / agent / native nodes), each fired as its own run. Optional; omit for a silent alert whose breaches you read from its events. |
The scope and threshold are a matched pair — a metric can only be computed over the scopes that produce it (e.g. errorRate needs telemetry, freshness needs a model). The full compatibility matrix, every metric's meaning and units, and every scope filter field are in references/scopes-and-thresholds.md — read it before writing a --scope/--threshold pair you haven't used before.
The golden rule: preview before you create
alert preview evaluates a scope + threshold right now, without firing actions or writing an event. It returns the value the alert would measure and whether that value breaches — so you calibrate the threshold against reality instead of guessing, and you confirm the scope/threshold pairing is even valid before committing it to a schedule.
cargo-ai observability alert preview \
--scope '{"kind":"runs","workflowUuid":"<uuid>","statuses":["error"]}' \
--threshold '{"metric":"errorRate","operator":"gte","value":10}' \
--window-minutes 1440 # last 24h; default 60
outcome: "computed"→{ value, total, failed, isBreached }. Set your threshold fromvalue.outcome: "empty"→ the window had nothing to measure (see the empty-vs-zero rule inreferences/alert-lifecycle.md).outcome: "notComputed"→{ errorMessage }. A bad SQL query, a deleted model, or an invalid scope/threshold pairing all land here — fix it before creating.
--window-minutes only shapes the window for telemetry scopes (spans/runs/records). A model is measured as it stands right now; a query scope windows itself in its SQL.
Always preview first. It is free, it is the only way to size a threshold correctly, and it catches an invalid pairing before it becomes a schedule that writes an error event every tick.
Commands
All commands output JSON. Reads need a token with observability:read; create/update/remove need observability:write (an admin token has both; a plain member token may not — see Bootstrap above).
Create an alert
cargo-ai observability alert create \
--name "CRM sync error rate" \
--description "Page when the HubSpot sync starts failing" \
--cron "*/30 * * * *" \
--scope '{"kind":"runs","workflowUuid":"<workflow-uuid>","statuses":["error"]}' \
--threshold '{"metric":"errorRate","operator":"gte","value":10}' \
--actions '[{"kind":"agent","agentUuid":"<agent-uuid>","config":{"message":"{{alert.name}} breached: {{event.value}}% errors. {{alert.url}}"}}]'
--cron— 5-field cron or@every <interval>(e.g.@every 30m), always UTC, at most once a minute. The UI presets bottom out at 30 minutes; go tighter only with reason (every tick scans ClickHouse and can fire paid runs).--disabled— create it paused (evaluate nothing until youupdate --enabled true).--folder <uuid>— file it under a folder (fromcargo-workspace-management).--actions— optional. Omit for a silent alert. Each entry is a configured action: unlikeorchestration action execute, which carries noconfigat all, an alert action requires one — that is where the templated message lives. The config is templated against the firing context ({{alert.*}},{{event.*}}) — seereferences/alert-lifecycle.mdfor the full variable list. Each action's target (agentUuid/toolUuid/connectorUuid) is validated to exist in the workspace at create time.
List, get, update, remove
cargo-ai observability alert list # all alerts, each with its lastEvent
cargo-ai observability alert get <uuid> # one alert + its lastEvent
cargo-ai observability alert update --uuid <uuid> \
--enabled false # pause it (true/false — must be literal)
cargo-ai observability alert update --uuid <uuid> \
--threshold '{"metric":"errorRate","operator":"gte","value":20}' # raise the bar
cargo-ai observability alert update --uuid <uuid> \
--description none # "none" clears; --folder none unfiles
cargo-ai observability alert remove <uuid>
--enabled is strict: only the literal true or false are accepted — --enabled yes is rejected rather than silently disabling the alert. On update, any flag you omit is left unchanged; --description none / --folder none are the explicit "clear it" spellings.
Inspect firing history
cargo-ai observability event list <alertUuid> # latest evaluation events, newest first
Each event carries status (healthy / unhealthy / error), the measured value, a snapshot of the scope/threshold/actions as they were when it fired (the alert can change afterwards), runUuids (the runs the actions spawned — feed these to cargo-diagnostics or orchestration run get), the evaluation window, and errorMessage for error events. unhealthy = breached and fired; error = the metric could not be computed.
How evaluation actually works
The lifecycle — cron windows and the ClickHouse indexing lag, the at-most-once firing guarantee (an alert never re-fires on the same rows; a sustained breach is re-detected on the next tick), the empty-window-vs-real-zero rule that makes lte a dead-man's switch, and the full {{alert.*}}/{{event.*}} templating context — is documented in references/alert-lifecycle.md. Read it before you rely on an alert for anything time-sensitive.
Worked recipes
references/examples/recipes.md — copy-paste starting points: error-rate pager, credit-budget guard, p95-latency watch, a dead-man's switch (count lte 0 — alert when a workflow stops running), model freshness / empty-model alerts, and a custom SQL-query alert.
Declarative alternative: defineAlert (CDK)
This skill is the imperative surface — one-off cargo-ai observability alert … calls. To manage an alert as code (in git, reproducible, deployed alongside the workflow it watches), use CDK's defineAlert builder instead — see ../cargo-project/SKILL.md and "Declarative vs imperative" in the router. Same scope/threshold/action model; different authoring mode.
Cost discipline
An alert's actions fire as real runs — if an action calls a paid connector action or an agent, every breach re-bills. A poorly-sized threshold on a tight cron can breach (and bill) every tick. Two safeguards:
- Preview to size the threshold so it fires on genuine anomalies, not normal variance.
- If an action node calls a credits-based provider action, treat it like any scheduled paid workflow: read that provider's playbook (esp. its Recurring use section) in
../cargo-gtm/provider-playbooks/, and apply the spend rules in../cargo-gtm/references/cost-discipline.md. Prefer cheap notification actions (an agent that posts to Slack, a connector notification) over anything that fans out.
When the CLI surprises you
If a documented flag, scope field, or response shape doesn't match what you observe (a fix may have shipped, or the docs may have drifted), re-refresh the CLI and skills; if it still doesn't add up, file a report — it's read by the team:
cargo-ai workspaceManagement report create \
--title "<one-line summary>" \
--description "<exact command(s), errorMessage verbatim, expected vs actual, UUIDs>"
Presenting results
Follow ../cargo/references/interaction.md: lead with the outcome ("alert created, will page the on-call agent when the CRM sync's error rate hits 10% over 30 min"), summarize an alert or its events as a compact table, never dump raw alert get / event list JSON into the conversation.
Files (cargo-skills)
-
references
-
examples
-
recipes.md 6.6 KB
# Alert recipes Copy-paste starting points. Every one is **`preview` first** to size the threshold, then `create`. Replace `<…>` placeholders with real UUIDs (discover them via `cargo-orchestration` / `cargo-storage` / `cargo-ai`). See `../scopes-and-thresholds.md` for every field and `../alert-lifecycle.md` for firing semantics. --- ## 1. Error-rate pager for one workflow Page the on-call agent when a workflow's runs start failing. ```bash # Preview against the last 24h to see the current rate cargo-ai observability alert preview \ --scope '{"kind":"runs","workflowUuid":"<workflow-uuid>"}' \ --threshold '{"metric":"errorRate","operator":"gte","value":10}' \ --window-minutes 1440 # Create it on a 30-min cadence with an agent notification cargo-ai observability alert create \ --name "CRM sync — error rate" \ --description "Error rate ≥10% over 30 min" \ --cron "*/30 * * * *" \ --scope '{"kind":"runs","workflowUuid":"<workflow-uuid>"}' \ --threshold '{"metric":"errorRate","operator":"gte","value":10}' \ --actions '[{"kind":"agent","agentUuid":"<agent-uuid>","config":{"message":"🚨 {{alert.name}}: {{event.value}}% errors ({{event.windowStart}}–{{event.windowEnd}}). {{alert.url}}"}}]' ``` `errorRate` is a percent of *finished* runs; a quiet window with nothing finished is `empty`, not a false `0%`. --- ## 2. Credit-budget guard Catch cost blowouts before the invoice does — total credits burned by a workflow over the window. ```bash cargo-ai observability alert create \ --name "Enrichment play — hourly credit ceiling" \ --cron "0 * * * *" \ --scope '{"kind":"spans","workflowUuid":"<workflow-uuid>","nodeKind":"connector"}' \ --threshold '{"metric":"credits","aggregation":"sum","operator":"gte","value":500}' \ --actions '[{"kind":"agent","agentUuid":"<agent-uuid>","config":{"message":"💸 {{alert.name}} burned {{event.value}} credits in the last hour ({{event.threshold}} ceiling). {{alert.url}}"}}]' ``` Scoping to `nodeKind: "connector"` measures only the paid provider calls. Preview at a couple of window sizes to learn the normal hourly spend before setting `value`. --- ## 3. p95 latency watch Fire when a workflow (or a specific node) gets slow. ```bash cargo-ai observability alert preview \ --scope '{"kind":"spans","workflowUuid":"<workflow-uuid>","nodeActionSlug":"<action-slug>"}' \ --threshold '{"metric":"duration","aggregation":"p95","operator":"gte","value":30}' \ --window-minutes 180 ``` Use the previewed `value` to set a realistic `p95` threshold, then `create` with a `--cron`. Aggregations: `avg` | `p50` | `p95` | `p99`. --- ## 4. Dead-man's switch — alert when a workflow STOPS running The one pattern that needs `count lte 0`: `count` reports a real `0` on an empty window (not `empty`), so silence breaches. ```bash cargo-ai observability alert create \ --name "Nightly sync — did it run?" \ --description "No runs in the last 24h = something broke upstream" \ --cron "0 8 * * *" \ --scope '{"kind":"runs","workflowUuid":"<workflow-uuid>"}' \ --threshold '{"metric":"count","operator":"lte","value":0}' \ --actions '[{"kind":"agent","agentUuid":"<agent-uuid>","config":{"message":"⚠️ {{alert.name}}: the nightly sync produced no runs. {{alert.url}}"}}]' ``` The cron interval **is** the window — run this once daily so "no runs" means "none in the last day". --- ## 5. Model freshness — stale sync Breach when a model hasn't emitted new data in too long. `freshness` is in **minutes** and ignores the scope filter. ```bash cargo-ai observability alert create \ --name "Companies model — freshness" \ --cron "*/30 * * * *" \ --scope '{"kind":"model","modelUuid":"<model-uuid>"}' \ --threshold '{"metric":"freshness","operator":"gte","value":120}' \ --actions '[{"kind":"agent","agentUuid":"<agent-uuid>","config":{"message":"🕒 {{alert.name}}: {{event.value}} min since last emit. {{alert.url}}"}}]' ``` Related model metrics: `syncDuration gte <seconds>` (sync got slow), `recordsShare gte <percent>` (a filtered slice grew — needs a scope `filter`). --- ## 6. Empty-model dead-man's switch `recordsCount` returns a real `0` on an empty model, so `lte 0` catches a model that emptied out or never populated. ```bash cargo-ai observability alert create \ --name "Leads model — not empty" \ --cron "0 */6 * * *" \ --scope '{"kind":"model","modelUuid":"<model-uuid>"}' \ --threshold '{"metric":"recordsCount","operator":"lte","value":0}' \ --actions '[{"kind":"agent","agentUuid":"<agent-uuid>","config":{"message":"📉 {{alert.name}}: the model is empty. {{alert.url}}"}}]' ``` Add a `filter` (segmentation shape, spelled `conjonction`) to count only a slice — e.g. records missing an enrichment column, then alert if that count climbs with `gte`. --- ## 7. Custom SQL-query alert When no built-in metric fits, compute the value yourself. The query must return a **single number** and window itself. Orchestration runtime (error rate over the last hour, self-windowed): ```bash cargo-ai observability alert create \ --name "Workspace-wide error rate" \ --cron "*/15 * * * *" \ --scope '{"kind":"orchestrationQuery","query":"select countIf(status = '"'"'error'"'"') * 100 / count() from runs where created_at > now() - interval 1 hour"}' \ --threshold '{"metric":"query","operator":"gte","value":5}' ``` Storage warehouse (records missing enrichment): ```bash cargo-ai observability alert create \ --name "Unenriched companies backlog" \ --cron "0 */4 * * *" \ --scope '{"kind":"storageQuery","query":"select count() from default.companies where enriched_at is null"}' \ --threshold '{"metric":"query","operator":"gte","value":1000}' ``` Notes: - Validate the SQL with `cargo-ai orchestration query execute` / `cargo-ai storage query execute` first, then `alert preview` the whole scope+threshold. - An aggregate over no rows is `NULL` → treated as `empty` (won't breach `lte`). For a "went silent" query alert, return a real `0` via `count()`. - Shell-quoting SQL with single quotes is fiddly — the `'"'"'` dance above escapes a literal `'`. Alternatively build the JSON in a file and pass `--scope "$(cat scope.json)"`. --- ## Reading the results ```bash cargo-ai observability alert list # every alert + its lastEvent status/value cargo-ai observability event list <alert-uuid> # firing history, newest first ``` An `unhealthy` event's `runUuids` are the runs its actions spawned — trace them with `cargo-ai orchestration run get <uuid>` or hand them to `cargo-diagnostics`. An `error` event means the metric couldn't be computed (bad query, deleted model) — read its `errorMessage` and fix the scope.
-
-
alert-lifecycle.md 6.9 KB
# Alert lifecycle — how evaluation and firing actually work What happens on each cron tick, why an alert never double-fires, the empty-window rule, and the full templating context an action gets. Read this before you rely on an alert for anything time-sensitive. ## One evaluation tick On every scheduled tick, for an enabled alert: 1. **Build the window.** `windowEndedAt = now − ClickHouse indexing lag` (spans land through Kinesis + a materialized view, so the tail of the window is left for the next tick to avoid missing late rows). `windowStartedAt = the alert's last `lastEvaluatedAt`, or its `updatedAt` on the first ever tick. Telemetry scopes measure over `[windowStartedAt, windowEndedAt]`; a `model` is measured point-in-time; a query scope windows itself. 2. **Compute the value** for the scope + threshold (see `scopes-and-thresholds.md`). 3. **Claim the window atomically** — advance the cursor (`lastEvaluatedAt → windowEndedAt`) only if the alert is still enabled, not deleted, and no other activity already advanced it. If the claim is lost, the tick records nothing and fires nothing. 4. **Record an event** and, on breach, **fire the actions**. If `windowStartedAt >= windowEndedAt` (an empty or already-evaluated window — e.g. a retry, or overlapping ticks during a schedule change) the tick is a no-op: an alert never re-fires on spans it already saw. ## The three event outcomes Every tick that claims its window writes exactly one event: | Compute outcome | Event `status` | Fires actions? | `value` | | --- | --- | --- | --- | | breached | `unhealthy` | **yes** | the measured value | | not breached | `healthy` | no | the measured value | | `empty` (nothing to measure) | `healthy` | no | `null` | | `notComputed` (bad SQL, deleted model, corrupt pairing) | `error` | no | `null` (+ `errorMessage`) | `event list <alertUuid>` returns these newest-first. A run of `unhealthy` events is a sustained breach; an `error` event means the alert can't measure what it was told to — fix the scope/query/model. ## At-most-once firing (and what that means for you) Firing is deliberately **at-most-once**, not at-least-once. The cursor is claimed *before* actions fire, so a Temporal retry or an overlapping cron can't fire the same breach twice. Actions spawn runs — which cost credits and can include agents that open PRs or send messages — so a rare *miss* is preferred to a *duplicate*. Consequences: - **A sustained breach is re-detected, not re-fired on the same rows.** Each tick only sees rows since the last cursor advance. If the condition is still breaching on the *next* window's rows, you get another `unhealthy` event then. So on a 30-minute cron, an ongoing error spike pages roughly every 30 minutes — it does not spam. - **A one-tick blip fires once.** Good for "tell me the moment X happens". - **Disabling or deleting an alert mid-evaluation cancels the firing** for that tick. - **Design actions to be safe to receive repeatedly** (a notification, an idempotent ticket), since a long breach produces one firing per tick it's true for. ## The empty-vs-zero rule An idle/empty window is reported as **`empty`** (→ `healthy`, no fire) for almost every metric — you don't want a latency or error-rate alert firing "0" every quiet hour. The exceptions are **`count`** (telemetry scopes) and **`recordsCount`** (model scope): an empty window is a real **`0`**. Paired with **`lte 0`** they become **dead-man's switches** that breach *because* nothing happened — the only way to alert on *absence* (a workflow that stopped, a model that emptied). See `scopes-and-thresholds.md` for the per-metric table. The same principle protects query alerts: an aggregate over no rows is `NULL` (ClickHouse also renders `NaN`/`0÷0` as `NULL`), which the alert treats as `empty` rather than `0` — so a rate query on an idle window won't false-breach an `lte` threshold. If you *want* silence to breach, write a `count()` that returns a genuine `0`. ## Actions: what fires, and the templating context On breach, **each action in `--actions` is fired as its own run** through the orchestration action service (`skipConcurrencyCheck` is on — an alert must fire even when the workspace is at its run-concurrency limit). All the runs of one firing share a single trace; their uuids are stored on the event as `runUuids`. Firing is best-effort per action: one action failing to start is logged and skipped, never blocking the others or the event. `--actions` is the shared orchestration `Action[]` union — the same shape used everywhere in `cargo-orchestration`: ```json [ {"kind":"agent","agentUuid":"…","config":{ "message":"…" }}, {"kind":"connector","integrationSlug":"…","actionSlug":"…","config":{ … }}, {"kind":"tool","toolUuid":"…","config":{ … }}, {"kind":"native","actionSlug":"…","config":{ … }} ] ``` Each action's target (`agentUuid` / `toolUuid` / `connectorUuid`) is **validated to exist in the workspace** at create/update time — a mistyped uuid is rejected up front, not silently at breach. (An action whose target is deleted *afterwards* fails at fire time and is recorded on the event.) ### Firing context (templating) Before each action runs, its `config` is interpolated against the firing context. The action's evaluated `config` becomes the run's input **data** (the action executes with an empty config), so put your bindings in `config`: | Variable | Value | | --- | --- | | `{{alert.uuid}}` | The alert's UUID. | | `{{alert.name}}` | The alert's name. | | `{{alert.url}}` | Deep link to the alert in the app. | | `{{event.value}}` | The measured value (rounded to 2 dp). | | `{{event.threshold}}` | The threshold `value` it crossed. | | `{{event.operator}}` | `gte` / `lte`. | | `{{event.windowStart}}` | Window start, ISO 8601. | | `{{event.windowEnd}}` | Window end, ISO 8601. | | `{{event.spansUrl}}` | Deep link to the workspace's spans view. | Example agent action that composes a human-readable page: ```json [{"kind":"agent","agentUuid":"<agent-uuid>","config":{ "message":"🚨 {{alert.name}} breached: {{event.value}} {{event.operator}} {{event.threshold}} over {{event.windowStart}}–{{event.windowEnd}}. Alert: {{alert.url}} · Spans: {{event.spansUrl}}" }}] ``` ## Schedules - `--cron` accepts a 5-field cron expression **or** `@every <interval>` (e.g. `@every 15m`), evaluated in **UTC**. - Minimum interval is **once a minute** — tighter `@every` values are rejected (every tick scans ClickHouse and may fire paid runs). Named-weekday cron expressions the interval parser can't measure are allowed through, since a 5-field cron can't fire more than once a minute anyway. - The UI presets bottom out at **30 minutes**; the CDK template defaults to **5 minutes**. Pick the loosest cadence that still catches the problem in time — it's cheaper and quieter. - Create with `--disabled` to stage an alert without evaluating it; `alert update --uuid <uuid> --enabled true` starts the schedule. `--enabled false` pauses it (the schedule stops; the alert and its history remain). -
scopes-and-thresholds.md 8.1 KB
# Scopes & thresholds — the compatibility matrix An alert measures a **scope** and breaches on a **threshold**. They are a matched pair: each metric can only be computed over the scopes that produce it. Get the pairing wrong and the alert (or a `preview`) returns `outcome: "notComputed"` with `The "<metric>" metric cannot be computed over a "<scope>" scope.` ## The matrix at a glance | Scope `kind` | Source | Allowed threshold metrics | | --- | --- | --- | | `spans` | Per-node executions of a workflow (ClickHouse) | `errorRate`, `duration`, `credits`, `count` | | `runs` | Whole runs (the eight-value run status) | `errorRate`, `duration`, `credits`, `count` | | `records` | One row per record, latest state | `errorRate`, `duration`, `credits`, `count` | | `orchestrationQuery` | Your SQL over `runs`/`batches`/`spans`/`records` | `query` | | `storageQuery` | Your SQL over the workspace data warehouse | `query` | | `model` | A storage model's records + sync state | `recordsCount`, `recordsShare`, `freshness`, `syncDuration` | `operator` is always `gte` or `lte`; `value` is always a number. --- ## Scopes (the `--scope` JSON) ### Telemetry scopes — `spans`, `runs`, `records` All three are windowed over the evaluation interval and share the same four metrics. They differ in what they count and how they filter. **`spans`** — one row per node execution. The richest filter set: ```json { "kind": "spans", "workflowUuid": "…", "parentAgentUuid": "…", "nodeKind": "native | connector | tool | agent", "nodeIntegrationSlug": "…", "nodeConnectorUuid": "…", "nodeActionSlug": "…", "nodeToolUuid": "…", "nodeAgentUuid": "…", "executionTitleOrErrorMessage": "substring match", "executionStatuses": ["pending", "success", "error"], "userUuid": "…" } ``` Every field is optional; omit them all to watch every span in the workspace. Use `nodeActionSlug` / `nodeConnectorUuid` to pin the alert to a single provider action, `nodeAgentUuid` to watch one agent's calls. **`runs`** — one row per run. Filters on the **full** run status set (not the three execution statuses): ```json { "kind": "runs", "workflowUuid": "…", "statuses": ["error"], "releaseUuid": "…", "recordTitleOrErrorMessage": "substring match", "userUuid": "…" } ``` **`records`** — one row per record holding its latest state (no idle/skipped; the same work as `runs`, keyed by record): ```json { "kind": "records", "workflowUuid": "…", "statuses": ["error"], "releaseUuid": "…", "titleOrErrorMessage": "substring match", "userUuid": "…" } ``` > `statuses` on `runs` uses the full run-status enum; on `records` it uses the record-status enum. When in doubt, `preview` with the statuses you want and check the `total`/`failed` counts. (Discover valid status values from `cargo-orchestration`.) ### Query scopes — `orchestrationQuery`, `storageQuery` You supply the SQL; it must return a **single numeric value** (the first column of the first row). The query is expected to **window itself** — `--window-minutes` does not apply. ```json { "kind": "orchestrationQuery", "query": "select countIf(status='error')*100/count() from runs where created_at > now() - interval 1 hour" } ``` ```json { "kind": "storageQuery", "query": "select count() from default.companies where enriched_at is null" } ``` - `orchestrationQuery` runs against the orchestration runtime tables (`runs`, `batches`, `spans`, `records`; no schema prefix; workspace-scoped) — same engine as `cargo-ai orchestration query execute`. - `storageQuery` runs against the workspace data warehouse using `<datasetSlug>.<modelSlug>` table names — same engine as `cargo-ai storage query execute`. If no warehouse is connected, the alert errors with *"No data warehouse is connected to this workspace."* - An aggregate over no rows is `NULL` (and ClickHouse renders `NaN`/`0/0` as `NULL` too) → the evaluation is treated as **`empty`**, not `0`. So a rate query on an idle window won't false-breach an `lte` threshold. If you want silence to breach, use a `count()` that returns a real `0` (see the dead-man's switch recipe). Pair either query scope with the `query` threshold — the SQL computes the value, the threshold just carries the comparison: ```json { "metric": "query", "operator": "gte", "value": 10 } ``` ### Model scope — `model` Watches a storage model's records and its sync health: ```json { "kind": "model", "modelUuid": "…", "filter": { "conjonction": "and", "groups": [ … ] } } ``` - `filter` is optional and uses the **segmentation filter shape** — note the spelling **`conjonction`** (silently ignored if misspelled). It narrows the record metrics (`recordsCount`, `recordsShare`) and is **ignored** by the sync metrics (`freshness`, `syncDuration`). Discover the filter shape and model UUIDs from `cargo-storage` / `cargo-orchestration`. - A model is measured **point-in-time** — as it stands at the tick. The cron controls *how often* it's checked, not what's measured; `--window-minutes` doesn't apply. --- ## Thresholds (the `--threshold` JSON) ### Telemetry metrics (for `spans` / `runs` / `records`) | Metric | Extra field | Value means | Empty window | | --- | --- | --- | --- | | `errorRate` | — | `failed × 100 / finished` (**percent**). `total` on the event is the finished denominator. | Nothing *finished* → `empty` (healthy, no fire). | | `duration` | `aggregation`: `avg`\|`p50`\|`p95`\|`p99` | The chosen aggregate of duration **in seconds**, over finished rows only. Preview to see the current level before setting `value`. | No rows → `empty`. | | `credits` | `aggregation`: `sum`\|`avg`\|`p95` | Credit spend aggregated over the window. | No rows → `empty`. | | `count` | — | Number of rows **started** in the window (running ones included) — "did work happen". | A real **`0`**, *not* `empty` — so `count lte 0` is a dead-man's switch. | ```json {"metric":"errorRate","operator":"gte","value":10} {"metric":"duration","aggregation":"p95","operator":"gte","value":30} {"metric":"credits","aggregation":"sum","operator":"gte","value":500} {"metric":"count","operator":"lte","value":0} ``` ### Query metric (for `orchestrationQuery` / `storageQuery`) ```json {"metric":"query","operator":"gte","value":10} ``` ### Model metrics (for `model`) | Metric | Value means | Notes | | --- | --- | --- | | `recordsCount` | Live count of records matching the scope `filter`. | An empty model is a real **`0`** (not `empty`), so `recordsCount lte 0` is a dead-man's switch for an empty/broken model. | | `recordsShare` | `matching × 100 / total` (**percent**) — the filter's share of the model. | Needs a scope `filter` to be meaningful. A model with **0** total records is `empty` (no share to report), not `0%`. Costs two warehouse queries. | | `freshness` | **Minutes** since the model last emitted (falls back to the model's `createdAt` if it never has). | A model that never synced ages from creation, so `freshness gte <mins>` can breach a sync that never ran. | | `syncDuration` | **Seconds** the model's last completed sync took (`finishedAt − createdAt` of `lastRun`). | While a sync is still in flight there's no duration → `empty` (won't false-breach an `lte`). | ```json {"metric":"recordsCount","operator":"lte","value":0} {"metric":"recordsShare","operator":"gte","value":30} {"metric":"freshness","operator":"gte","value":60} {"metric":"syncDuration","operator":"gte","value":300} ``` --- ## The empty-vs-zero rule (why it matters) Most metrics report an idle/empty window as **`empty`** → a `healthy` event, no fire. That's deliberate: you don't want an error-rate or latency alert screaming "0!" every quiet night. Two metrics are the exception and return a real **`0`** on an empty window — **`count`** (telemetry) and **`recordsCount`** (model). Paired with **`lte`**, they become **dead-man's switches**: they breach *because* nothing happened. This is the only way to alert on absence — a workflow that stopped running, a model that emptied out. Every other metric treats absence as "nothing to measure", not "a low value". See the lifecycle reference for the full statement, and the recipes for a ready-made dead-man's switch.
-
-
skill-metadata.json 872 B
{ "$comment": "Generated by .github/scripts/skills-metadata.mjs — do not hand-edit. Regenerate with: node .github/scripts/skills-metadata.mjs --write .", "name": "cargo-observability", "version": "1.0.3", "documents": [ { "path": "SKILL.md", "kind": "entrypoint", "title": "Cargo CLI — Observability" }, { "path": "references/alert-lifecycle.md", "kind": "reference", "title": "Alert lifecycle — how evaluation and firing actually work" }, { "path": "references/examples/recipes.md", "kind": "example", "title": "Alert recipes" }, { "path": "references/scopes-and-thresholds.md", "kind": "reference", "title": "Scopes & thresholds — the compatibility matrix" } ], "contentHash": "08683568fc29d3613e339821aeb169edfc42e02db816dcb52a5bd4d1430da8dc" } -
SKILL.md 11.4 KB
--- name: cargo-observability description: "Watch a Cargo workspace and get told when something breaks — scheduled threshold alerts over workflow telemetry (spans, runs, records), a storage model freshness or row count, or any SQL query, firing a connector, tool, or agent when a metric breaches. Triggers: \"alert me when\", \"notify me if\", \"let me know when the error rate\", \"monitor this workflow\", \"tell me if the sync stops\", \"warn me before I run out of credits\", \"dead man’s switch\", \"is this still running\", \"set up monitoring\", plus listing, previewing, editing, and reviewing an alert firing history. Skip when: diagnosing something that already went wrong — use cargo-diagnostics." version: "1.0.3" compatibility: Requires @cargo-ai/cli (npm). Sign in or create an account with `cargo-ai login --email` (emailed code, no browser), `--oauth`, or an API token homepage: https://github.com/getcargohq/cargo-skills metadata: author: getcargo openclaw: requires: bins: - cargo-ai install: - kind: node package: "@cargo-ai/cli@latest" bins: - cargo-ai homepage: https://github.com/getcargohq/cargo-skills --- # Cargo CLI — Observability **Alerts.** An alert is a scheduled threshold check. On every cron tick it measures a **scope** (what to watch), compares the measured value against a **threshold** (the breach condition), and on breach fires **actions** — each as its own run — and records an **event**. This is the proactive counterpart to `cargo-diagnostics`: diagnostics explains a failure *after* you notice it; an alert *tells you* the moment a metric crosses a line. Everything lives under one CLI domain: ```bash cargo-ai observability alert … # the alert CRUD + preview surface cargo-ai observability event … # an alert's firing history ``` ## Bootstrap Already signed in (`cargo-ai whoami` returns a workspace)? Skip to the next section. ```bash npm install -g @cargo-ai/cli # no global install? prefix every command with `npx @cargo-ai/cli` cargo-ai login --email you@company.com # emailed code, no browser; creates the account on first use # alternatives: --oauth (browser) · --token <api-token> (CI) cargo-ai whoami # confirm the active workspace before any write ``` Every command prints JSON to stdout; failures exit non-zero with `{"errorMessage": "..."}`. Anything that creates a run or a batch is async — pass `--wait-until-finished` or poll the matching `get`. Alerts are guarded by `observability:read` / `observability:write` permissions. If a create/update/remove returns a permission error, the token lacks `observability:write` — use an admin token or have one granted ([`../cargo-workspace-management/SKILL.md`](../cargo-workspace-management/SKILL.md)). When the full skill bundle is installed, [`../cargo/references/prerequisites.md`](../cargo/references/prerequisites.md) adds the CLI version pin, token scopes, and the admin-only surface. ## The three moving parts of every alert | Part | Flag | What it is | | --- | --- | --- | | **Scope** | `--scope <json>` | *What* to measure — one of six sources: `spans`, `runs`, `records`, `orchestrationQuery`, `storageQuery`, `model`. | | **Threshold** | `--threshold <json>` | *When it breaches* — a `metric` + `operator` (`gte`/`lte`) + `value`. The metric menu depends on the scope. | | **Actions** | `--actions <json>` | *What happens on breach* — an `Action[]` (connector / tool / agent / native nodes), each fired as its own run. Optional; omit for a silent alert whose breaches you read from its events. | The scope and threshold are a **matched pair** — a metric can only be computed over the scopes that produce it (e.g. `errorRate` needs telemetry, `freshness` needs a model). The full compatibility matrix, every metric's meaning and units, and every scope filter field are in **[`references/scopes-and-thresholds.md`](references/scopes-and-thresholds.md)** — read it before writing a `--scope`/`--threshold` pair you haven't used before. ## The golden rule: `preview` before you `create` `alert preview` evaluates a scope + threshold **right now, without firing actions or writing an event**. It returns the value the alert would measure and whether that value breaches — so you calibrate the threshold against reality instead of guessing, and you confirm the scope/threshold pairing is even valid before committing it to a schedule. ```bash cargo-ai observability alert preview \ --scope '{"kind":"runs","workflowUuid":"<uuid>","statuses":["error"]}' \ --threshold '{"metric":"errorRate","operator":"gte","value":10}' \ --window-minutes 1440 # last 24h; default 60 ``` - `outcome: "computed"` → `{ value, total, failed, isBreached }`. Set your threshold from `value`. - `outcome: "empty"` → the window had nothing to measure (see the empty-vs-zero rule in `references/alert-lifecycle.md`). - `outcome: "notComputed"` → `{ errorMessage }`. A bad SQL query, a deleted model, or an **invalid scope/threshold pairing** all land here — fix it before creating. `--window-minutes` only shapes the window for telemetry scopes (`spans`/`runs`/`records`). A `model` is measured as it stands right now; a query scope windows itself in its SQL. **Always preview first.** It is free, it is the only way to size a threshold correctly, and it catches an invalid pairing before it becomes a schedule that writes an `error` event every tick. ## Commands All commands output JSON. Reads need a token with `observability:read`; create/update/remove need `observability:write` (an admin token has both; a plain member token may not — see Bootstrap above). ### Create an alert ```bash cargo-ai observability alert create \ --name "CRM sync error rate" \ --description "Page when the HubSpot sync starts failing" \ --cron "*/30 * * * *" \ --scope '{"kind":"runs","workflowUuid":"<workflow-uuid>","statuses":["error"]}' \ --threshold '{"metric":"errorRate","operator":"gte","value":10}' \ --actions '[{"kind":"agent","agentUuid":"<agent-uuid>","config":{"message":"{{alert.name}} breached: {{event.value}}% errors. {{alert.url}}"}}]' ``` - `--cron` — 5-field cron **or** `@every <interval>` (e.g. `@every 30m`), always **UTC**, at most once a minute. The UI presets bottom out at 30 minutes; go tighter only with reason (every tick scans ClickHouse and can fire paid runs). - `--disabled` — create it paused (evaluate nothing until you `update --enabled true`). - `--folder <uuid>` — file it under a folder (from `cargo-workspace-management`). - `--actions` — optional. Omit for a silent alert. Each entry is a **configured** action: unlike `orchestration action execute`, which carries no `config` at all, an alert action **requires** one — that is where the templated message lives. The config is templated against the firing context (`{{alert.*}}`, `{{event.*}}`) — see [`references/alert-lifecycle.md`](references/alert-lifecycle.md) for the full variable list. Each action's target (`agentUuid`/`toolUuid`/`connectorUuid`) is validated to exist in the workspace at create time. ### List, get, update, remove ```bash cargo-ai observability alert list # all alerts, each with its lastEvent cargo-ai observability alert get <uuid> # one alert + its lastEvent cargo-ai observability alert update --uuid <uuid> \ --enabled false # pause it (true/false — must be literal) cargo-ai observability alert update --uuid <uuid> \ --threshold '{"metric":"errorRate","operator":"gte","value":20}' # raise the bar cargo-ai observability alert update --uuid <uuid> \ --description none # "none" clears; --folder none unfiles cargo-ai observability alert remove <uuid> ``` `--enabled` is strict: only the literal `true` or `false` are accepted — `--enabled yes` is rejected rather than silently disabling the alert. On `update`, any flag you omit is left unchanged; `--description none` / `--folder none` are the explicit "clear it" spellings. ### Inspect firing history ```bash cargo-ai observability event list <alertUuid> # latest evaluation events, newest first ``` Each event carries `status` (`healthy` / `unhealthy` / `error`), the measured `value`, a **snapshot** of the `scope`/`threshold`/`actions` as they were when it fired (the alert can change afterwards), `runUuids` (the runs the actions spawned — feed these to `cargo-diagnostics` or `orchestration run get`), the evaluation window, and `errorMessage` for `error` events. `unhealthy` = breached and fired; `error` = the metric could not be computed. ## How evaluation actually works The lifecycle — cron windows and the ClickHouse indexing lag, the **at-most-once** firing guarantee (an alert never re-fires on the same rows; a *sustained* breach is re-detected on the next tick), the empty-window-vs-real-zero rule that makes `lte` a dead-man's switch, and the full `{{alert.*}}`/`{{event.*}}` templating context — is documented in **[`references/alert-lifecycle.md`](references/alert-lifecycle.md)**. Read it before you rely on an alert for anything time-sensitive. ## Worked recipes **[`references/examples/recipes.md`](references/examples/recipes.md)** — copy-paste starting points: error-rate pager, credit-budget guard, p95-latency watch, a **dead-man's switch** (`count lte 0` — alert when a workflow *stops* running), model freshness / empty-model alerts, and a custom SQL-query alert. ## Declarative alternative: `defineAlert` (CDK) This skill is the **imperative** surface — one-off `cargo-ai observability alert …` calls. To manage an alert **as code** (in git, reproducible, deployed alongside the workflow it watches), use CDK's `defineAlert` builder instead — see [`../cargo-project/SKILL.md`](../cargo-project/SKILL.md) and "Declarative vs imperative" in the router. Same scope/threshold/action model; different authoring mode. ## Cost discipline An alert's **actions fire as real runs** — if an action calls a paid connector action or an agent, every breach re-bills. A poorly-sized threshold on a tight cron can breach (and bill) every tick. Two safeguards: - **Preview to size the threshold** so it fires on genuine anomalies, not normal variance. - If an action node calls a **credits-based provider action**, treat it like any scheduled paid workflow: read that provider's playbook (esp. its *Recurring use* section) in `../cargo-gtm/provider-playbooks/`, and apply the spend rules in [`../cargo-gtm/references/cost-discipline.md`](../cargo-gtm/references/cost-discipline.md). Prefer cheap notification actions (an agent that posts to Slack, a connector notification) over anything that fans out. ## When the CLI surprises you If a documented flag, scope field, or response shape doesn't match what you observe (a fix may have shipped, or the docs may have drifted), re-refresh the CLI and skills; if it still doesn't add up, file a report — it's read by the team: ```bash cargo-ai workspaceManagement report create \ --title "<one-line summary>" \ --description "<exact command(s), errorMessage verbatim, expected vs actual, UUIDs>" ``` ## Presenting results Follow [`../cargo/references/interaction.md`](../cargo/references/interaction.md): lead with the outcome ("alert created, will page the on-call agent when the CRM sync's error rate hits 10% over 30 min"), summarize an alert or its events as a compact table, never dump raw `alert get` / `event list` JSON into the conversation.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.