Claude Cursor Skill

monte-carlo-monitoring-advisor

Analyze data coverage, create monitors for warehouse tables and AI agents. Covers coverage gaps, use-case analysis, data monitor creation, and agent observability.

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download monte-carlo-data-mc-agent-toolkit-skills_monitoring-advisor-bcc7373.zip · 81 KB
Part of monte-carlo-data/mc-agent-toolkit — 20 skills

Install

skills CLI npx skills add https://github.com/monte-carlo-data/mc-agent-toolkit/tree/main/skills/monitoring-advisor
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install monte-carlo-data-mc-agent-toolkit@llmmart
Git git clone https://github.com/monte-carlo-data/mc-agent-toolkit.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole monte-carlo-data/mc-agent-toolkit collection as a plugin from our marketplace. Git is the plain clone.

README

Monte Carlo Monitoring Advisor Skill

Analyze data coverage, create monitors for warehouse tables and AI agents. Walks users through warehouse discovery, use-case exploration, coverage gap analysis, data monitor creation, and agent observability — all through natural conversation. This single skill handles all monitoring needs: coverage analysis, data quality monitors (metric, validation, custom SQL, comparison, table), and AI agent monitors (metric, evaluation, trajectory, validation).

Editor & Stack Compatibility

The skill works with any AI editor that supports MCP and the Agent Skills format — including Claude Code, Cursor, and VS Code.

All warehouses supported by Monte Carlo work with the monitoring advisor. The skill validates table and column references against your actual warehouse schema via the Monte Carlo API.

Prerequisites

  • Claude Code, Cursor, VS Code or any editor with MCP support
  • Monte Carlo account with Editor role or above
  • MC CLI installed for monitor deployment (pip install montecarlodata)
  • All monitor creation capabilities are built in — no additional skills needed

Setup

Via the mc-agent-toolkit plugin (recommended)

Install the plugin for your editor — it bundles the skill, hooks, MCP server, and permissions automatically. See the main README for editor-specific instructions.

Standalone

  1. Configure the Monte Carlo MCP server:

    claude mcp add --transport http monte-carlo-mcp https://mcp.getmontecarlo.com/mcp
    
  2. Install the skill:

    npx skills add monte-carlo-data/mc-agent-toolkit --skill monitoring-advisor
    
  3. Authenticate: run /mcp in your editor, select monte-carlo-mcp, and complete the OAuth flow.

  4. Verify: ask your editor "Test my Monte Carlo connection" — it should call test_connection and confirm.

Legacy: header-based auth (for MCP clients without HTTP transport)

If your MCP client doesn't support HTTP transport, use .mcp.json.example with npx mcp-remote and header-based authentication. See the MCP server docs for details.

How to use it

Ask your AI editor about your monitoring coverage — describe what you want to understand or protect. The skill guides the agent through warehouse discovery, use-case analysis, coverage gap identification, and monitor creation. No special commands needed.

Example prompts

  • "What are my coverage gaps?"
  • "Show me my use cases and what's monitored"
  • "Which tables should I monitor first?"
  • "Analyze monitoring coverage for my warehouse"
  • "Find unmonitored tables with recent anomalies"
  • "Help me set up monitoring for my critical use cases"
  • "Create a freshness monitor on the orders table"
  • "Set up a null check on the email column"
  • "Monitor my AI agent's latency and token usage"
  • "Track my agent's response quality"

What it does

  1. Discovers your warehouses, use cases, and AI agents
  2. Analyzes coverage — which tables are monitored, which aren't, and which have active anomalies
  3. Prioritizes gaps by criticality, importance score, and anomaly activity
  4. Creates data quality monitors (metric, validation, custom SQL, comparison, table) with full parameter validation
  5. Creates AI agent monitors (metric, evaluation, trajectory, validation) for agent observability
  6. Generates monitors-as-code YAML ready for deployment

Deploying generated monitors

When the advisor generates a monitor, it returns MaC YAML. Deploy with:

montecarlo monitors apply --dry-run    # preview
montecarlo monitors apply --auto-yes   # apply

Your project needs a montecarlo.yml config in the working directory:

version: 1
namespace: <your-namespace>
default_resource: <your-warehouse-name>

Skill manifest

Monte Carlo Monitoring Advisor Skill

This skill handles all monitoring requests -- coverage analysis, data monitor creation, and AI agent monitoring. It routes to the right reference file based on the user's intent.

Monte Carlo tool routing (required): Always call Monte Carlo MCP tools through this plugin's bundled server, whose fully-qualified tool names are mcp__plugin_mc-agent-toolkit_monte-carlo-mcp__<tool> (e.g. mcp__plugin_mc-agent-toolkit_monte-carlo-mcp__get_alerts). Bare tool names used in this skill (get_alerts, search, get_table, …) refer to that bundled server. If the session also has a separately-configured monte-carlo-mcp server, do not route to it — it may point at a different endpoint or credentials.

Reference files live next to this skill file. Use the Read tool (not MCP resources) to access them:

  • Data monitor creation procedure: references/data-monitor-creation.md (relative to this file)
  • Agent monitor creation procedure: references/agent-monitor-creation.md (relative to this file)
  • Per-type references: references/data-*.md and references/agent-*.md (relative to this file)

When to activate this skill

Activate when the user:

  • Asks about monitoring coverage, data coverage, or coverage gaps
  • Wants to understand what's monitored vs. not in their warehouse
  • Asks about use cases, use-case criticality, or use-case analysis
  • Wants to explore their data estate and find what needs monitoring
  • Says things like "what should I monitor?", "where are my coverage gaps?", "show me my use cases"
  • Asks about unmonitored tables with anomalies or importance-based prioritization
  • Asks to create, add, or set up a monitor (e.g. "add a monitor for...", "create a freshness check on...", "set up validation for...")
  • Mentions monitoring a specific table, field, or metric
  • Wants to check data quality rules or enforce data contracts
  • Asks about monitoring options for a table or dataset
  • Requests monitors-as-code YAML generation
  • Wants to add monitoring after new transformation logic (when the prevent skill is not active)
  • Asks about monitoring AI agents, agent latency, agent token usage, or agent quality
  • Wants to set up alerts on agent behavior or execution patterns
  • Says things like "monitor my agent", "track agent latency", "alert on agent errors", "set up performance monitoring for my agent", or asks for an agent latency SLO
  • Asks about agent evaluation monitors, trajectory monitors, or validation monitors
  • Mentions agent observability or agent monitoring

When NOT to activate this skill

Do not activate when the user is:

  • Just querying data or exploring table contents
  • Triaging or responding to active alerts (use the prevent skill's Workflow 3)
  • Running impact assessments before code changes (use the prevent skill's Workflow 4)
  • Asking about existing monitor configuration (use get_monitors directly)
  • Editing or deleting existing monitors
  • Investigating agent alerts or agent traces (this skill creates agent monitors; investigating what they catch uses the monte-carlo-troubleshoot-agent-traces skill)

Prerequisites

  • Required: Monte Carlo MCP server (monte-carlo-mcp) must be configured and authenticated
  • Optional: A database MCP server (Snowflake, BigQuery, Redshift, Databricks) for SQL profiling of table usage patterns

Available MCP tools

All tools are available via the monte-carlo-mcp MCP server.

Coverage and discovery tools

Tool Purpose
get_warehouses List accessible warehouses (needed first -- get_use_cases requires warehouse_id)
get_use_cases List use cases with criticality, descriptions, table counts, precomputed tag names
get_use_case_table_summary Criticality distribution (HIGH/MEDIUM/LOW table counts) for a use case
get_use_case_tables Paginated tables with criticality, golden-table status, MCONs
get_monitors Check monitoring status on specific tables via mcons filter
get_asset_lineage Upstream/downstream dependencies for tables (takes MCONs + direction)
get_audiences List notification audiences
get_unmonitored_tables_with_anomalies Tables with muted OOTB anomalies but no monitors (takes ISO 8601 time range)
search Find tables by name; supports is_monitored filter
get_table Table details, fields, stats, domain membership
get_queries_for_table Query logs for a table (source/destination)
get_field_metric_definitions Available metrics per field type for a warehouse
get_domains List Monte Carlo domains
get_validation_predicates Available validation rule types

Data monitor creation tools

All five tools follow a two-call preview-then-confirm pattern: the first call (with the default dry_run=True) returns rendered MaC YAML for review; the second call (dry_run=False) deploys the monitor live and returns a deep link to it. Pass monitor_uuid on either call to update an existing monitor in place instead of creating a new one. See references/data-monitor-creation.md for the full flow.

Tool Purpose
create_or_update_table_monitor Create or update a table monitor (preview YAML on dry_run=True, deploy on dry_run=False)
create_or_update_metric_monitor Create or update a metric monitor (preview YAML on dry_run=True, deploy on dry_run=False)
create_or_update_validation_monitor Create or update a validation monitor (preview YAML on dry_run=True, deploy on dry_run=False)
create_or_update_sql_monitor Create or update a custom SQL monitor (preview YAML on dry_run=True, deploy on dry_run=False)
create_or_update_comparison_monitor Create or update a comparison monitor (preview YAML on dry_run=True, deploy on dry_run=False)

Data product tool

Tool Purpose
create_or_update_data_product Create or update a data product — a named grouping of warehouse assets with reliability tracking (asset-footprint preview on dry_run=True, live create on dry_run=False). Used by the agent Context pillar to wrap an agent's upstream tables (see references/agent-monitor-creation.md)

Agent monitoring tools

Tool Purpose
get_agent_metadata List AI agents -- returns agent names, agentReference values (the agent arg for monitor creation), trace table MCONs, source types, backend classes (backend_class), and each agent's warehouse_uuid/warehouse_name (the warehouse arg -- show the name, pass the uuid)
get_agent_conversations List recent conversations for an agent (newest first; filter by errors/status/turns/tokens/duration; optional inline transcripts)
get_agent_conversation Retrieve one conversation's full prompt/completion thread by conversation_id
get_agent_traces List traces with per-trace workflows, tasks, models, LLM-call counts, tokens, duration, and error counts
get_agent_trace Inspect one execution trace's full span tree
get_agent_segments Enumerate the distinct workflow / task / model values to scope a monitor to a real segment
create_or_update_agent_metric_monitor Create or update monitors for quantitative span-level metrics (preview YAML on dry_run=True, deploy on dry_run=False)
create_or_update_agent_evaluation_monitor Create or update monitors for LLM-evaluated quality metrics (preview YAML on dry_run=True, deploy on dry_run=False)
create_or_update_agent_trajectory_monitor Create or update trajectory monitors for execution pattern alerts (preview YAML on dry_run=True, deploy on dry_run=False)
create_or_update_agent_validation_monitor Create or update validation monitors for logical assertions (preview YAML on dry_run=True, deploy on dry_run=False)

Routing

When the user's request comes in, determine which workflow to follow:

User intent Workflow
Coverage analysis, use-case exploration, "what should I monitor?" Coverage workflow (below)
Create a specific data monitor for a known table Read references/data-monitor-creation.md and follow its procedure
Monitor AI agents, agent latency, agent quality, agent traces Read references/agent-monitor-creation.md and follow its procedure — propose coverage across the four POBC pillars (Performance, Output, Behavior, Context)
Coverage analysis leads to monitor creation Complete coverage workflow, then read references/data-monitor-creation.md for creation

When reading reference files, always use the Read tool with the path relative to this skill file.


Coverage workflow

This is the primary flow when the user asks about monitoring coverage, coverage gaps, or what to monitor.

Step 1: Discover warehouses

Call get_warehouses to list all accessible warehouses.

  • If one warehouse: select it automatically, proceed to Step 2.
  • If multiple warehouses: present warehouse names (never UUIDs) and ask the user which one to explore.

Step 2: Discover use cases

Call get_use_cases(warehouse_id=<selected>) to discover use cases for the chosen warehouse.

  • If use cases exist --> proceed to the Use-case exploration (below).
  • If no use cases --> proceed to the Importance-based fallback (below).

Step 3: Check for database MCP (optional)

Check if the user has a database MCP server available by looking for tools containing snowflake, bigquery, redshift, or databricks in the tool list. If found, note it for the SQL profiling step later. If not found, skip SQL profiling gracefully.


Use-case exploration

This is the primary flow when use cases are defined.

Present use cases

  • Sort by criticality: HIGH before MEDIUM before LOW.
  • For each use case, show the description and explain the reasoning for its criticality level so the user understands why it matters.
  • Call get_use_case_tables with golden_tables_only=true and mention specific golden-table names as concrete examples. Golden tables are the last layer in the warehouse -- they feed ML models, dashboards, and reports. Explain this when relevant.
  • Use get_asset_lineage to explain how tables in a use case are connected and why certain tables are important (e.g. a golden table with many upstream dependencies).

"Create a use case" requests

You cannot create use cases -- they are generated automatically by Monte Carlo (along with their criticality), and there is no tool to author one. When the user asks to "create", "set up", or "define" a use case: briefly say so, and do NOT silently substitute monitor deployment. Then offer what you can do for the table(s) they named -- look up the existing use case / criticality, recommend field monitors, generate monitor previews, or analyze coverage gaps -- and act on the do-able part without expanding to sibling tables.

Analyze coverage

  1. Call get_use_case_table_summary to show how many tables exist at each criticality level (HIGH / MEDIUM / LOW) for the use case.
  2. Call get_use_case_tables to obtain table MCONs, then call get_monitors(mcons=[...]) to report how many are already monitored vs. not.
  3. Default to HIGH + MEDIUM criticality scope. This covers the most important tables without overwhelming the user. Do NOT ask the user which scope to use -- just proceed. If they want LOW-criticality tables included, they'll ask.
  4. You may suggest covering multiple use cases in one session.
  5. Bias toward action, not questions. When the scope is clear (HIGH + MEDIUM for the selected use case), proceed directly to generating monitor previews for all recommended monitors. Frame it as opt-out, not opt-in: "I'll generate previews for all N monitors -- tell me if you want to skip any." Do NOT ask "which would you like me to create?" one at a time -- batch them.

Identify coverage gaps with anomaly data

Use get_unmonitored_tables_with_anomalies to discover tables that are not monitored but already have muted out-of-the-box anomalies. This reveals real coverage gaps -- places where Monte Carlo detected data issues but no monitor was configured to alert anyone.

  • Call it with a recent time window (e.g. last 7-30 days) using ISO 8601 timestamps.
  • Results are ranked by importance score -- the most critical gaps appear first.
  • Each result includes a sample of anomaly events showing what types of issues were detected (freshness, volume, schema changes).
  • Use this to prioritize which unmonitored tables to cover first -- a table with recent anomalies is a stronger candidate than one with no activity.
  • Cross-reference with use-case data: if an unmonitored table with anomalies belongs to a critical use case, escalate its priority.

Importance-based fallback

When no use cases are defined, fall back to importance-based table discovery.

  1. Find unmonitored tables: Use search(query="", is_monitored=false) to find unmonitored tables sorted by importance.
  2. Find tables with anomalies: Use get_unmonitored_tables_with_anomalies with a recent time window (last 14-30 days) to find tables with recent anomalies but no monitors.
  3. Inspect top candidates: Use get_table to check table details, fields, and stats for the most important unmonitored tables.
  4. Understand criticality via lineage: Use get_asset_lineage with direction="DOWNSTREAM" to understand which tables are most connected -- a table with many downstream dependents is a stronger candidate for monitoring.
  5. Prioritize: Rank candidates by importance score and anomaly activity. Present the top candidates to the user with reasoning.

Important

  • Do NOT present importance scores as business criticality. Always explain that the importance score is a computed metric (query frequency, downstream dependencies, usage patterns), not business-defined criticality.
  • Tell the user their account doesn't have use-case data yet -- use cases are generated automatically by Monte Carlo from warehouse metadata and exposed as asset tags; they are not manually configured through a UI.
  • You can still create metric, validation, and custom SQL monitors for individual tables in this mode -- you just won't use tag-based table monitors, since there are no use-case tags.

SQL profiling (optional)

If a database MCP server was detected in Step 3 of the coverage workflow:

  1. Call get_queries_for_table to see recent query patterns on candidate tables.
  2. Use the database MCP tools (e.g. snowflake_query, bigquery_query) to profile table usage -- identify which tables are queried most frequently, which columns are used in JOINs and WHERE clauses.
  3. Use this information to refine monitor suggestions -- heavily-queried tables with no monitors are high-priority gaps.

If no database MCP is available, skip this step entirely. Do not ask the user to configure one.


Pre-creation context (coverage-driven)

When coverage analysis leads to monitor creation, gather this context before reading the creation reference file:

  1. Dedup first. Before generating a use-case tag monitor, call get_monitors with the same tag pair (and monitor_types=["TABLE"]) you'd put in the monitor's asset_selection.filters. If a monitor already covers that (tag, domain) scope, surface it (description, uuid) and ask whether to update it (pass its monitor_uuid), add one with a distinct scope, or skip -- do NOT silently re-create. The backend upserts a table monitor on its (description, domain), so a same-description definition silently overwrites the prior monitor's settings.
  2. Call get_audiences to list notification audiences. Suggest one or more relevant audiences (match by team or use-case context) and ask the user which they want -- they can pick one or several. This is the one question to ask before generating; do NOT also ask about draft/active or schedule. Default to draft (is_draft=True); the user can flip to active after seeing the preview.
  3. When passing audiences or failure_audiences, use the audience name/label (not UUID), as a list -- one entry per selected audience.
  4. Never fabricate credit costs. Do not give a generic per-monitor or per-field MC credit rate -- cost scales with the specific spec (segmentation, schedule, field count). If a preview response includes a backend estimate (e.g. estimated_credits.credits_per_day), report that; otherwise decline and offer to preview a specific monitor or use case to get the real estimate.

Use-case tag monitors

The most common output of coverage analysis is a table monitor scoped by use-case tags via create_or_update_table_monitor. The asset_selection parameter uses this structure:

{
  "databases": ["<database_name>"],
  "schemas": ["<schema_name>"],
  "filters": [
    {
      "type": "TABLE_TAG",
      "tableTags": ["<tag_key>:<criticality>"],
      "tableTagsOperator": "HAS_ANY"
    }
  ]
}

Rules:

  • Filter type is always TABLE_TAG for use-case monitors.
  • tableTagsOperator should be HAS_ANY.
  • Each entry in tableTags is "<tag_key>:<value>" where the tag key is the precomputed tag name from get_use_cases output and the value is the criticality level in lowercase (high, medium, low).
  • To monitor only HIGH-criticality tables: ["tag_name:high"]
  • To monitor MEDIUM + HIGH: ["tag_name:high", "tag_name:medium"]
  • To monitor ALL: ["tag_name:high", "tag_name:medium", "tag_name:low"]

Monitor title (description) and reasoning (notes)

Keep these distinct -- both are accepted by the creation tools. The backend auto-generates the monitor name slug; description is the title users see.

  • description -- the title. Short and scannable (≤ ~80 chars), plain English, naming the asset/use case and criticality scope. Do NOT cram reasoning here.
  • notes -- the reasoning. 1-3 sentences answering "why this monitor?", grounded in criticality, scope, and downstream impact.

Example for a use-case tag monitor:

  • Bad description (this is reasoning, not a title): "Monitor HIGH criticality tables in the Revenue Reporting use case to catch issues before they affect dashboards and financial reports."
  • Good description: "Revenue Reporting coverage -- HIGH + MEDIUM criticality tables"
  • Good notes (paired): "Covers HIGH/MEDIUM-criticality tables in the Revenue Reporting use case. Catches freshness, volume, and schema issues before they reach dashboards and financial reports."

Transient and truncate-and-reload tables

Some tables show 0 rows when queried directly but have recent write activity in Monte Carlo metadata. These are transient tables -- fully replaced on each pipeline run (truncate-and-reload pattern). Recognize this pattern early to avoid wasting time querying empty tables.

Signs of a transient table:

  • get_table shows a recent last_updated_on and high read/write activity
  • Direct SQL query returns 0 rows or all-NULL timestamp columns
  • Monte Carlo detected freshness anomalies (the table stayed empty longer than expected between loads)

Graceful degradation

Handle missing or unavailable tools gracefully:

Scenario Behavior
No use cases defined Fall back to importance-based discovery
No database MCP available Skip SQL profiling, rely on MC tools only
get_unmonitored_tables_with_anomalies returns empty Note that no recent anomalies were found; proceed with use-case or importance-based prioritization
get_use_case_tables returns no tables Note the use case has no tables; suggest exploring other use cases
get_audiences returns empty Inform user no audiences are configured; monitors can still be created without notification routing
User has no warehouses Inform user that no warehouses are accessible; they may need to check their Monte Carlo permissions

Never error out or stop the conversation because one tool returned empty results. Explain what happened and offer the next best path.


Rules

  • Never expose UUIDs, MCONs, or internal identifiers to the user -- always use human-readable names for warehouses, audiences, use cases, and tables. Keep internal identifiers for tool calls only.
  • When the user asks about relationships between tables, use get_asset_lineage to fetch upstream/downstream connections and explain the data flow.
  • Be concise but thorough. Use bullet points and tables for clarity.
  • Always use ISO 8601 format for datetime values in tool calls.
  • Never reformat YAML values returned by creation tools.
  • When passing audiences or failure_audiences to monitor creation tools, use the audience name/label (not UUID). The API accepts audience names.
Files (mc-agent-toolkit)
  • references
    • agent-evaluation-monitor.md 30.4 KB
      # Agent Evaluation Monitor
      
      ## When to use
      
      Run LLM-evaluated quality checks on agent outputs. Best for:
      
      - **Answer relevance scoring** — is the response relevant to the question?
      - **Helpfulness and clarity** — is the response useful and well-structured?
      - **Task completion** — did the agent complete what was asked?
      - **Banned-keyword check** — does the output avoid specific banned keywords (e.g. password, ssn, api secret)?
      - **Custom evaluation criteria** — a custom LLM check (`custom_prompt`) or SQL check (`custom_sql`) over span text
      
      Do NOT use this for a raw numeric metric like latency or token count (use
      `create_or_update_agent_metric_monitor`) — evaluation monitors add sampling and
      transforms, which those don't need.
      
      ## Constraints
      
      > **CRITICAL:** The monitor's source is the `agent` reference. Pass the
      > `agentReference` value from `get_agent_metadata` verbatim — a platform
      > `{database}:{schema}.{name}` reference or an OTel `service_name`. Never modify,
      > truncate, or reconstruct it, and never pass an MCON.
      
      > **CRITICAL:** `warehouse` is REQUIRED. Pass the agent's `warehouse_uuid` from
      > `get_agent_metadata`; omitting it fails with "Warehouse not found". Use
      > `get_warehouses` when `warehouse_uuid` is null or to resolve a warehouse by name.
      
      > **CRITICAL:** `sampling_config` is REQUIRED. Provide `percentage`, `count`, or
      > both. Per-span monitors cap `count` at 10,000; conversation-level monitors cap it
      > at 500 (a percentage-only conversation config is capped at 500 per run).
      
      > **IMPORTANT:** `transforms` is a TOP-LEVEL parameter. Each transform produces an
      > output field that `alert_conditions.fields` references.
      
      > **IMPORTANT:** `schedule_type` is `fixed` (default) or `manual` — never dynamic.
      > `interval_minutes` defaults to `60` and must be at least 60 **and** a multiple of 60.
      
      > **NEVER** set a `field` parameter on any transform. Predefined judges pull their
      > inputs automatically; `custom_prompt` reads from its prompt's template variables;
      > `custom_sql` reads from the columns in its expression. There is no `field` param,
      > and no `context` param either.
      
      > **NEVER** use `classification` or `sentiment` as a transform function. Their output
      > column is not added to the evaluated schema, so no `alert_conditions` field can
      > reference it — the monitor can't alert on the result (you get "Field `<alias>`
      > doesn't exist"). To bucket or pass/fail a response, use a `custom_prompt` with
      > `outputType: "boolean"` (a check) or `"string"`.
      
      ## Key characteristics
      
      - Requires `sampling_config` — controls how many spans/conversations are sampled
      - Supports a top-level `transforms` array — the evaluation logic
      - Transform output field names are what `alert_conditions.fields` reference
      - Optional `is_agent_conversation_aggregation=True` aggregates per conversation
        (see the conversation-grain section — OTel/ClickHouse, Snowflake Cortex, and
        Databricks Genie agents; Databricks MLflow agents are span-only)
      
      ## Parameters
      
      | Parameter | Type | Required | Description |
      |-----------|------|----------|-------------|
      | `description` | string | Yes | Human-readable monitor description (shown as display name) |
      | `agent` | string | Yes | Agent reference — `agentReference` from `get_agent_metadata` (`{db}:{schema}.{name}` or OTel `service_name`) |
      | `warehouse` | string | Yes | Warehouse name or UUID where the agent's traces live |
      | `alert_conditions` | array | Yes | Alert conditions using transform output field names |
      | `sampling_config` | object | Yes | `{"percentage": 10.0}`, `{"count": 100}`, or both |
      | `transforms` | array | No | Evaluation transforms (predefined or custom); top-level |
      | `is_agent_conversation_aggregation` | boolean | No | Aggregate evaluation per conversation (OTel/ClickHouse, Cortex, and Genie agents; MLflow agents are span-only) |
      | `trace_table` | string | No | Explicit trace table — only for non-ClickHouse OTel agents |
      | `agent_span_filters` | array | No | Optional span-scope refinement; at most ONE filter object. At conversation grain (`is_agent_conversation_aggregation=True`) only `agent`/`workflow` are allowed — not `task`/`spanName` |
      | `sensitivity` | string | No | Anomaly detection sensitivity for AUTO operators (`low`/`medium`/`high`) |
      | `aggregate_by` | string | No | Time-window bucketing (`hour`/`day`/`week`/`month`) |
      | `schedule_type` | string | No | `fixed` (default) or `manual` |
      | `interval_minutes` | int | No | Default `60`; at least 60 and a multiple of 60 |
      | `tags` | array | No | Key-value tags on the monitor. Each tag is `{"name": "<key>", "value": "<value>"}` — the key field is `name` (NOT `key`), and unknown fields are rejected. **Default: tag every agent monitor with its agent** — `[{"name": "agent", "value": "<AGENT_NAME>"}]` (the `agentName` from `get_agent_metadata`) — so one agent's monitors can be filtered as a group |
      | `domain_uuids` | array | No | Domain UUIDs to assign this monitor to — the agent-onboarding playbook passes the footprint's single resolved domain on every create (see agent-monitor-creation.md conventions) |
      | `monitor_uuid` | string | No | UUID of an existing monitor to update in place (PUT semantics) |
      | `dry_run` | boolean | No | Default `True` — preview YAML; set `False` to deploy |
      
      ## Predefined LLM transforms
      
      Pass only `function` (plus an optional `alias`, an optional `modelName` to pin the
      judge model — see **Judge model selection** below — and `modelConnectionId` on
      BigQuery only). Do NOT set `prompt`, `sqlExpression`, `outputType`, or `field` — the
      tool rejects them. Each writes a numeric score (1–5, except `semantic_similarity`
      which is 0–5) to its built-in output field:
      
      | Transform function | Output field | Output type | Description |
      |-------------------|-------------|-------------|-------------|
      | `answer_relevance` | `relevance_score` | number (1-5) | Is the response relevant to the question? |
      | `helpfulness` | `helpfulness_score` | number (1-5) | Is the response helpful? |
      | `task_completion` | `completion_score` | number (1-5) | Did the agent complete the task? |
      | `language_match` | `match_score` | number (1-5) | Does the response match the expected language? |
      | `clarity` | `clarity_score` | number (1-5) | Is the response clear and well-structured? |
      | `prompt_adherence` | `adherence_score` | number (1-5) | Does the response follow the prompt instructions? |
      | `semantic_similarity` | `similarity_score` | number (0-5) | How similar is the response to a reference? |
      
      ## Predefined SQL transforms (rule-based, no LLM needed)
      
      Same rule: pass only `function` (and an optional `alias`); do NOT set `prompt`,
      `sqlExpression`, `outputType`, `modelConnectionId`, `modelName`, or `field`.
      
      | Transform function | Output field | Output type | Description |
      |-------------------|-------------|-------------|-------------|
      | `output_length` | `word_count` | number | Non-whitespace word count of the first completion |
      | `json_validity` | `json_valid` | boolean | Is the first completion valid JSON? |
      | `keywords` | `content_safe` | boolean | TRUE if output does NOT contain banned keywords (password, ssn, api secret, credit card). Not a general PII/secrets detector. |
      
      ## Custom transforms
      
      Each writes an output column named by its `alias`, and that alias is what
      `alert_conditions.fields` references. `outputType` is **camelCase** and one of
      `"number"`, `"string"`, `"boolean"`.
      
      | Function | Set these | Do NOT set | Output type |
      |----------|-----------|------------|-------------|
      | `custom_prompt` | `prompt` (with a `{{variable}}`), `alias`, `outputType`, + optional `modelName` (see Judge model selection) | `field`, `sqlExpression` | number / string / boolean |
      | `custom_sql` | `sqlExpression`, `alias`, `outputType` | `field`, `prompt`, `modelConnectionId`, `modelName` | number / string / boolean |
      
      - **`custom_prompt` prompts MUST reference at least one template variable** —
        `{{prompts}}`, `{{completions}}`, or `{{expected_output}}` for a per-span monitor,
        or `{{conversation}}` for a conversation-level monitor. A prompt with no variable,
        an unknown variable, or the wrong variable for the grain is rejected. With
        `"includeToolCalls": true` on the transform, `{{conversation}}` also carries the
        agent's tool calls as clearly identifiable TOOL entries (see Conversation-grain
        judges below).
      - **`custom_sql` runs against warehouse columns.** On Snowflake Cortex agents,
        `prompts`/`completions` are arrays — reference the string columns
        `first_completion` / `full_completion` / `first_prompt` instead. The dry-run does
        not evaluate the SQL, so a bad column surfaces only at run time.
      ## Judge model selection (`modelName`)
      
      **`modelName`** pins the judge model for LLM-based transforms (predefined judges and
      `custom_prompt`). Optional — omit it to use the warehouse default. **Models are
      warehouse-specific** because the judge runs inside the warehouse hosting the agent's
      **trace table** — never offer models from the wrong pool:
      
      | Trace table's warehouse | Judge models (default first) |
      |-------------------|------------------------------|
      | Snowflake (Cortex) | `llama3.1-70b`, `llama3.1-8b`, `llama3.3-70b`, `llama4-maverick`, `mixtral-8x7b`, `mistral-large2`, `mistral-large3`, `claude-sonnet-5`, `claude-sonnet-4-6`, `claude-haiku-4-5`, `claude-opus-4-7`, `claude-opus-4-8`, `openai-gpt-5.1`, `openai-gpt-5`, `openai-gpt-5-mini`, `openai-gpt-5-nano`, `openai-gpt-4.1`, `gemini-3.1-pro` |
      | Databricks | `databricks-meta-llama-3-3-70b-instruct`, `databricks-meta-llama-3-1-8b-instruct`, `databricks-gpt-5`, `databricks-gpt-5-mini`, `databricks-gpt-5-nano`, `databricks-gpt-oss-20b`, `databricks-gpt-oss-120b`, `databricks-gemma-3-12b`, `databricks-llama-4-maverick` |
      | BigQuery | `gemini-2.5-flash`, `gemini-2.5-pro` |
      | Athena trace tables (any cloud), or ClickHouse trace tables on an AWS-hosted deployment (most accounts) | `us.anthropic.claude-sonnet-5`, `us.anthropic.claude-haiku-4-5-20251001-v1:0`, `us.anthropic.claude-sonnet-4-5-20250929-v1:0`, `us.anthropic.claude-opus-4-8`, `us.anthropic.claude-opus-4-1-20250805-v1:0` |
      | ClickHouse trace tables on a GCP/Azure-hosted deployment | short names — `claude-sonnet-5`, `claude-haiku-4-5` (GCP: `claude-haiku-4-5@20251001`), `claude-opus-4-8`. Athena trace tables always use the AWS `us.anthropic.*` ids (not cloud-resolved). |
      
      The pool follows the warehouse the agent's traces live in (`warehouse_uuid`/
      `warehouse_name` from `get_agent_metadata`), not the agent's `backend_class` — a
      customer OTel trace table on Snowflake uses the Snowflake pool.
      
      Snapshot as of 2026-07-29 (source: monolith `validations/llm_models.yaml`) —
      re-verify before quoting as exhaustive.
      
      On Snowflake, everything except `llama3.1-*`, `llama3.3-70b`, `mixtral-8x7b` and
      `mistral-large2` requires Cortex cross-region inference enabled
      (`CORTEX_ENABLED_CROSS_REGION`); the dry-run does NOT check this — flag it to the
      user before pinning.
      
      A known catalog model on the wrong warehouse is rejected at dry-run. An unrecognized
      name is NOT validated — it is accepted as a custom model and fails at evaluation
      time if the warehouse doesn't host it; a passing dry-run is not evidence the model
      exists. Prefer a listed model; if the user insists on an unlisted one, pin it but
      tell them it could not be verified and to check that the monitor's first run
      produced scores.
      
      **`modelConnectionId`** is **BigQuery-only** (and required there) — the BigQuery
      Cloud resource connection in the customer's own GCP project that runs the judge. It
      is NOT a Monte Carlo setting or integration and Monte Carlo cannot list it; on
      BigQuery, ask the user for their BigQuery connection ID. **On every other warehouse,
      omit it entirely and never ask the user for it.** To choose the judge model, use
      `modelName` — there is no "model connection" to configure outside BigQuery.
      
      ## Conversation-grain judges
      
      Each LLM judge has a `*_conversation` variant that evaluates a whole conversation
      instead of a single span. These require `is_agent_conversation_aggregation: true`
      **AND a conversation-capable agent** — OpenTelemetry/ClickHouse, Snowflake Cortex,
      or Databricks Genie (Databricks MLflow agents reject conversation aggregation).
      The output/score column is not always the span judge's name:
      
      | Conversation function | Output/score field |
      |-----------------------|--------------------|
      | `answer_relevance_conversation` | `relevance_score` |
      | `task_completion_conversation` | `task_completion_score` |
      | `helpfulness_conversation` | `helpfulness_score` |
      | `clarity_conversation` | `clarity_score` |
      | `prompt_adherence_conversation` | `adherence_score` |
      | `language_match_conversation` | `match_score` |
      | `satisfaction_conversation` | `satisfaction_score` (no per-span counterpart) |
      
      **`includeToolCalls` — default ON at conversation grain.** Every conversation-grain
      transform (judge or custom) accepts an optional `includeToolCalls` boolean. When
      `true`, the agent's tool calls (name, inputs, outputs, and errors) are included in
      the judged conversation as clearly identifiable TOOL entries between the messages
      that triggered them — in call order, autonomous agent steps included — so the judge
      scores what the agent did, not just what it said. **Set `"includeToolCalls": true`
      on every conversation-grain transform by default.** Omit it only for pure style/tone
      judges — `clarity_conversation`, `language_match_conversation`, or a wording-only
      custom prompt — where the transcript's wording alone is judged and tool noise
      dilutes the judge. The field is invalid at span grain: the tool rejects
      `includeToolCalls: true` on a monitor without
      `is_agent_conversation_aggregation=True`.
      
      At conversation grain, a `custom_prompt` may only reference `{{conversation}}`, and
      the predefined SQL checks and `custom_sql` are not supported. At span grain,
      `{{conversation}}` is not available. With `"includeToolCalls": true` on the
      transform, `{{conversation}}` also carries the agent's tool calls (name, inputs,
      outputs, errors) as clearly identifiable TOOL entries between the messages that
      triggered them, in call order, autonomous steps included — write conversation
      prompts that judge actions with that visibility in mind. `agent_span_filters` at
      conversation grain may scope only by `agent`/`workflow` — `task`/`spanName` are
      span-level and are rejected.
      
      ## Alert conditions
      
      Use `thresholdValue` (camelCase) for threshold operators — NOT `threshold_value`
      (snake_case). Each condition names one or more transform output fields in `fields`.
      
      ```json
      {
          "metric": "NUMERIC_MEAN",
          "operator": "LT",
          "fields": ["relevance_score"],
          "thresholdValue": 2
      }
      ```
      
      **Match the metric to the transform's output type** (a mismatch is rejected at
      dry-run):
      
      - **number** (the 1–5 judges, `output_length`, numeric `custom_prompt`/`custom_sql`)
        → `NUMERIC_MEAN` and other numeric metrics.
      - **boolean** (`json_validity`, `keywords`, boolean `custom_prompt`/`custom_sql`) →
        `TRUE_RATE` / `FALSE_RATE`. NEVER a numeric metric — a boolean field has no mean.
      - `NULL_RATE` works on any output type.
      
      `fields` must name a column in the evaluated schema — a predefined judge's built-in
      field, a custom transform's `alias`, or a raw source column (span grain:
      `duration_sec`, `total_tokens`, …; conversation grain: `turn_count`,
      `duration_seconds`, `status`). Duplicate `(metric, field)` pairs across conditions
      are rejected.
      
      ## Examples
      
      The `agent` value below comes from `get_agent_metadata`'s `agentReference` field.
      The first example uses a platform `{database}:{schema}.{name}` reference; the others
      use an OTel `service_name`. Use whichever form your agent returns.
      
      ### Answer relevance evaluation (platform agent reference)
      
      ```
      create_or_update_agent_evaluation_monitor(
          description="Chat Agent relevance evaluation",
          agent="analytics:agents.support_bot",
          warehouse="Prod Warehouse",
          transforms=[
              {"function": "answer_relevance"}
          ],
          alert_conditions=[
              {"metric": "NUMERIC_MEAN", "operator": "LT", "fields": ["relevance_score"],
               "thresholdValue": 2}
          ],
          sampling_config={"count": 100},
          dry_run=True
      )
      ```
      
      ### Banned-keyword check as a boolean rate (OTel service_name)
      
      `keywords` outputs the boolean `content_safe`, so alert on its false rate — a
      numeric metric on a boolean field is rejected.
      
      ```
      create_or_update_agent_evaluation_monitor(
          description="Chat Agent banned-keyword check",
          agent="checkout-agent",
          warehouse="Agent Observability",
          transforms=[
              {"function": "keywords"}
          ],
          alert_conditions=[
              {"metric": "FALSE_RATE", "operator": "GT", "fields": ["content_safe"],
               "thresholdValue": 0.05}
          ],
          sampling_config={"count": 100},
          dry_run=True
      )
      ```
      
      ### Custom prompt as a pass/fail (boolean) check
      
      The prompt references `{{completions}}`; there is no `field`; `outputType` is
      `boolean` so the alert watches the true/false rate. This is the right shape for
      "how often did the agent do X".
      
      ```
      create_or_update_agent_evaluation_monitor(
          description="Did the agent disambiguate the product before answering?",
          agent="checkout-agent",
          warehouse="Agent Observability",
          transforms=[
              {
                  "function": "custom_prompt",
                  "alias": "disambiguated_product",
                  "prompt": "Did this response either ask which product the user meant, or state which product it assumed, before answering? Response: {{completions}}. Answer true or false.",
                  "outputType": "boolean"
              }
          ],
          alert_conditions=[
              {"metric": "FALSE_RATE", "operator": "GT", "fields": ["disambiguated_product"],
               "thresholdValue": 0.2}
          ],
          sampling_config={"percentage": 20.0},
          dry_run=True
      )
      ```
      
      ### Custom SQL numeric check
      
      `custom_sql` needs `sqlExpression` + `alias` + `outputType`; alert with a numeric
      metric on the alias.
      
      ```
      create_or_update_agent_evaluation_monitor(
          description="Completion length floor",
          agent="checkout-agent",
          warehouse="Agent Observability",
          transforms=[
              {
                  "function": "custom_sql",
                  "alias": "answer_chars",
                  "sqlExpression": "LENGTH(first_completion)",
                  "outputType": "number"
              }
          ],
          alert_conditions=[
              {"metric": "NUMERIC_MEAN", "operator": "LT", "fields": ["answer_chars"],
               "thresholdValue": 40}
          ],
          sampling_config={"count": 100},
          dry_run=True
      )
      ```
      
      ### Conversation-level evaluation (OTel agent only)
      
      Set `is_agent_conversation_aggregation=True`, use a `*_conversation` judge, and alert
      on its score field (`task_completion_conversation` → `task_completion_score`).
      Sampling `count` ≤ 500. The transform carries `"includeToolCalls": true` (the
      conversation-grain default) so completion is judged against what the agent actually
      did, not just what it said.
      
      ```
      create_or_update_agent_evaluation_monitor(
          description="Task completion across full conversations",
          agent="checkout-agent",
          warehouse="Agent Observability",
          is_agent_conversation_aggregation=True,
          transforms=[
              {"function": "task_completion_conversation", "includeToolCalls": true}
          ],
          alert_conditions=[
              {"metric": "NUMERIC_MEAN", "operator": "AUTO", "fields": ["task_completion_score"]}
          ],
          sampling_config={"count": 100},
          dry_run=True
      )
      ```
      
      ## Custom-prompt template library
      
      Named, reusable `custom_prompt` templates for the most common Output-pillar checks. All three are
      **conversation-level**: set `is_agent_conversation_aggregation=True` and reference
      `{{conversation}}` — the only template variable a conversation-grain prompt may use. All three
      carry `"includeToolCalls": true` — the conversation-grain default; none is a pure style/tone
      judge (a wording-only check like a clarity- or language-style judge would omit it) — so the judge
      reads the agent's TOOL entries alongside the messages. On a span-only agent (Databricks MLflow),
      adapt the prompt to `{{completions}}` at span grain instead and drop `includeToolCalls` too — it
      is rejected at span grain.
      
      **Render, don't recite.** Before proposing a template, replace every `<AGENT_NAME>` placeholder
      with the agent's actual name (from `get_agent_metadata`) and tailor the wording to the intents and
      failure modes you observed in its real conversations — a template proposed as generic boilerplate
      judges generically. `<AGENT_NAME>` is an authoring placeholder, **not** a template variable: never
      send it verbatim or turn it into a curly-brace variable — the agent name goes into the prompt as
      plain literal text. Always show the user the full rendered prompt text for approval before
      creating the monitor (dry-run first, as usual).
      
      ### `frustration_free_score` — was the experience frustration-free?
      
      1–5 score for user-visible friction: rephrasing, repeated corrections, complaints, giving up.
      Complements `helpfulness` — helpfulness judges the answers, this judges the user's experience
      across the whole conversation. Part of the **baseline pack** — propose it for every agent.
      
      ```json
      {
          "function": "custom_prompt",
          "alias": "frustration_free_score",
          "prompt": "Read this conversation between a user and <AGENT_NAME>: {{conversation}}. Rate from 1 to 5 how frustration-free the user's experience was. 5 = no sign of frustration; the user got what they needed without friction. 4 = minor friction (one clarification or retry) but the user stayed satisfied. 3 = noticeable friction; the user had to rephrase or repeat themselves to get a useful answer. 2 = clear frustration; the user complained, corrected the agent repeatedly, or expressed annoyance. 1 = severe frustration; the user gave up, abandoned the task, or ended the conversation visibly dissatisfied. Answer with only the number.",
          "outputType": "number",
          "includeToolCalls": true
      }
      ```
      
      Recommended alert: `{"metric": "NUMERIC_MEAN", "operator": "LT", "fields": ["frustration_free_score"], "thresholdValue": 4}`
      
      ### `answer_attempt_score` — did the agent attempt a real answer?
      
      1–5 score for whether the agent actually attempted to answer the user's data questions, versus
      deflecting, refusing, asking clarifying questions without ever answering, or erroring out.
      Part of the **analytics pack** (Cortex/Genie — see below), where deflection is the dominant
      failure mode of NL2SQL/analytics agents. This judge benefits directly from `includeToolCalls`:
      the TOOL entries show whether a query actually ran, separating a real answer attempt from a
      confident deflection — so the prompt tells the judge to use them.
      
      ```json
      {
          "function": "custom_prompt",
          "alias": "answer_attempt_score",
          "prompt": "Read this conversation between a user and <AGENT_NAME>, an analytics agent that answers data questions: {{conversation}}. Rate from 1 to 5 how fully the agent attempted to answer the user's data questions. TOOL entries in the conversation show what the agent actually ran; a substantive answer attempt is normally backed by one. 5 = every question got a direct, substantive answer attempt (a query, a result, or a concrete data answer). 4 = answered with minor gaps or hedging. 3 = partial; some questions were deflected or met only with clarifying questions. 2 = mostly deflected, refused, or answered a different question than asked. 1 = no real answer attempt at all. Answer with only the number.",
          "outputType": "number",
          "includeToolCalls": true
      }
      ```
      
      Recommended alert: `{"metric": "NUMERIC_MEAN", "operator": "LT", "fields": ["answer_attempt_score"], "thresholdValue": 4}`
      
      ### `user_correction` — did the user have to correct the agent?
      
      Boolean detector for follow-up-turn corrections — the user saying an answer was wrong, restating
      what they actually meant, or re-asking the same question. A correction is the strongest observable
      ground-truth signal that an earlier answer missed. Part of the **analytics pack** (Cortex/Genie —
      see below). Framed so `true` = a correction occurred, making `TRUE_RATE` the correction rate.
      
      ```json
      {
          "function": "custom_prompt",
          "alias": "user_correction",
          "prompt": "Read this conversation between a user and <AGENT_NAME>: {{conversation}}. Did the user correct the agent in a follow-up turn - for example saying a previous answer was wrong, restating what they actually meant, or re-asking the same question because the answer missed it? A clarifying question from the agent does not count as a correction. Answer true if at least one correction occurred, false otherwise.",
          "outputType": "boolean",
          "includeToolCalls": true
      }
      ```
      
      Recommended alert: `{"metric": "TRUE_RATE", "operator": "GT", "fields": ["user_correction"], "thresholdValue": 0.2}` —
      0.2 is a conservative starting point, not a calibrated one. Tune it to the agent's observed
      correction rate after the first week of results, or switch the operator to `AUTO_HIGH` once
      enough history has accumulated for anomaly detection.
      
      ### Action-aware checks — judging what the agent did
      
      With tool calls included (`includeToolCalls: true`), a `custom_prompt` can judge **action
      correctness**, not just answer text — the TOOL entries are the evidence the judge reads. Propose
      one when the agent's job is to DO something and you observed the corresponding failure mode:
      
      - "Was the create-monitor tool called with the configuration the user asked for, and did it
        error?" — catches an agent that confirms an action it never (or incorrectly) performed.
      - "Did the agent run a SQL query before answering a data question?" — catches an analytics
        agent that answers a data question without ever touching the data.
      
      Frame these as booleans so `FALSE_RATE` is the failure rate (see the pass/fail boolean example
      above).
      
      ## Output-pillar eval packs
      
      When setting up evaluation coverage for an agent (the Output pillar of agent observability),
      propose these packs rather than inventing a one-off list. Shared defaults for every pack monitor:
      
      - **Schedule:** daily — `interval_minutes=1440`
      - **Sampling:** `{"count": 100}` (100 conversations per run; the conversation-grain cap is 500)
      - **Tags:** `[{"name": "agent", "value": "<AGENT_NAME>"}]` — tag every monitor with its agent so
        one agent's monitors can be filtered as a group
      - **Grain:** conversation (`is_agent_conversation_aggregation=True`) where the agent supports it;
        span grain with `{{completions}}` / span judges otherwise
      - **Tool calls:** `includeToolCalls: true` on every conversation-grain transform (the default —
        omit only for pure style/tone judges); drop it at span grain, where it is rejected
      
      ### Baseline pack — every agent
      
      | Monitor | Transform | Alert |
      |---------|-----------|-------|
      | Helpfulness | predefined `helpfulness_conversation` (plain `helpfulness` on span-only agents) | `NUMERIC_MEAN` `LT` 4 on `helpfulness_score` |
      | Frustration | `frustration_free_score` template | `NUMERIC_MEAN` `LT` 4 on `frustration_free_score` |
      
      Start with the fixed `LT 4` floor. `AUTO_LOW` (drift detection) is the alternative once the
      monitor has accumulated a baseline — but not alongside it in the same monitor: duplicate
      `(metric, field)` pairs across conditions are rejected, so moving to drift means changing the
      condition's operator, not adding a second condition.
      
      Full example — the frustration baseline monitor with all pack defaults applied:
      
      ```
      create_or_update_agent_evaluation_monitor(
          description="Support Bot - frustration-free conversations (baseline)",
          agent="analytics:agents.support_bot",
          warehouse="Prod Warehouse",
          is_agent_conversation_aggregation=True,
          transforms=[
              {
                  "function": "custom_prompt",
                  "alias": "frustration_free_score",
                  "prompt": "Read this conversation between a user and Support Bot: {{conversation}}. Rate from 1 to 5 how frustration-free the user's experience was. 5 = no sign of frustration; the user got what they needed without friction. 4 = minor friction (one clarification or retry) but the user stayed satisfied. 3 = noticeable friction; the user had to rephrase or repeat themselves to get a useful answer. 2 = clear frustration; the user complained, corrected the agent repeatedly, or expressed annoyance. 1 = severe frustration; the user gave up, abandoned the task, or ended the conversation visibly dissatisfied. Answer with only the number.",
                  "outputType": "number",
                  "includeToolCalls": true
              }
          ],
          alert_conditions=[
              {"metric": "NUMERIC_MEAN", "operator": "LT", "fields": ["frustration_free_score"],
               "thresholdValue": 4}
          ],
          interval_minutes=1440,
          sampling_config={"count": 100},
          tags=[{"name": "agent", "value": "Support Bot"}],
          dry_run=True
      )
      ```
      
      ### Analytics pack — Cortex and Genie agents only
      
      Propose **only when the agent's `backend_class` is `platform_agent` (Snowflake Cortex) or
      `databricks_genie`** — these are the NL2SQL/analytics agents where deflected answers and
      user-corrected answers are the dominant failure modes. Do not propose this pack for other agents.
      
      | Monitor | Transform | Alert |
      |---------|-----------|-------|
      | Answer attempts | `answer_attempt_score` template | `NUMERIC_MEAN` `LT` 4 on `answer_attempt_score` |
      | User corrections | `user_correction` template | `TRUE_RATE` `GT` 0.2 on `user_correction` (starting point — tune, or move to `AUTO_HIGH` with history) |
      
      Same defaults as the baseline pack: daily, `{"count": 100}` sampling, the `agent` tag,
      `includeToolCalls: true` on each transform, and full rendered prompt text shown for approval
      before creating.
      
      ## Common errors
      
      | Error message | Cause | Fix |
      |--------------|-------|-----|
      | Warehouse not found | `warehouse` omitted or wrong | Pass the agent's `warehouse_uuid` from `get_agent_metadata`; if null, list warehouses via `get_warehouses` |
      | invalid / unresolvable `agent` reference | The `agent` value wasn't taken from `get_agent_metadata` | Use the exact `agentReference` value — do not construct it by hand, and never pass an MCON |
      | "Field X doesn't exist" | Wrong transform output field name, or a `classification`/`sentiment` output that isn't in the schema | Use the documented output field (e.g. `relevance_score`) or a custom transform's `alias`; replace `classification`/`sentiment` with a `custom_prompt` (`outputType` `boolean`/`string`) |
      | metric/output-type mismatch | Numeric metric on a boolean field (or vice versa) | `NUMERIC_MEAN` for numbers, `TRUE_RATE`/`FALSE_RATE` for booleans, `NULL_RATE` for any type |
      | `task`/`spanName` rejected in `agent_span_filters` | Used a span-level filter dimension at conversation grain | At conversation grain (`is_agent_conversation_aggregation=True`), scope only by `agent`/`workflow` — `task`/`spanName` are span-level |
      | `includeToolCalls` rejected | The field was set on a per-span monitor | `includeToolCalls` is conversation-grain only — keep it (default ON) with `is_agent_conversation_aggregation=True`; drop it at span grain |
      
    • agent-metric-monitor.md 16 KB
      # Agent Metric Monitor
      
      ## When to use
      
      Track quantitative span-level metrics over time. Best for:
      
      - **Latency monitoring** — `duration_sec` trending up
      - **Token usage tracking** — `total_tokens`, `prompt_tokens`, `completion_tokens` per call
      - **Volume monitoring** — number of spans per time window (`ROW_COUNT_CHANGE`)
      - **Boolean rates** — e.g. tool-call rate via `is_tool_call`
      - **Anomaly detection** on any of the above with automatic thresholds
      
      Do NOT use this for LLM-evaluated quality (relevance, correctness, tone) — that's
      `create_or_update_agent_evaluation_monitor`, which adds sampling + transforms.
      
      ## Constraints
      
      > **CRITICAL:** The monitor's source is the `agent` reference. Pass the
      > `agentReference` value from `get_agent_metadata` verbatim — a platform
      > `{database}:{schema}.{name}` reference or an OTel `service_name`. Never modify,
      > truncate, or reconstruct it, and never pass an MCON.
      
      > **CRITICAL:** `warehouse` is REQUIRED. Pass the agent's `warehouse_uuid` from
      > `get_agent_metadata`; use `get_warehouses` when it is null or to resolve by name.
      
      > **IMPORTANT:** `schedule_type` is `fixed` (default) or `manual` — never dynamic.
      > `interval_minutes` defaults to `60` and must be at least 60 **and** a multiple of 60.
      
      > **IMPORTANT:** Use `duration_sec` (not `duration_ms`) for latency. The field is
      > named `duration_sec` in the PARSED_SPANS layer — `duration_ms` does not exist.
      
      > **IMPORTANT:** `ROW_COUNT_CHANGE` is table-level — do NOT include a `fields` array,
      > and use only an anomaly operator (`AUTO`/`AUTO_HIGH`/`AUTO_LOW`) or `NOOP`. A manual
      > comparison operator has no field to bind to and is rejected.
      
      > **IMPORTANT:** Match the metric to the field type — numeric metrics need numeric
      > fields, boolean metrics need boolean fields. A numeric metric on a non-numeric
      > field is rejected at dry-run.
      
      ## Key characteristics
      
      - Uses `alert_conditions` with metric + operator (no transforms, no sampling)
      - Supports anomaly (`AUTO`) and threshold operators
      - Optional `is_agent_trace_aggregation=True` aggregates per trace instead of per span
        (OTel agents only — see Trace aggregation)
      - Optional `sensitivity` (`low`/`medium`/`high`) tunes AUTO operators
      
      ## Parameters
      
      | Parameter | Type | Required | Description |
      |-----------|------|----------|-------------|
      | `description` | string | Yes | Human-readable monitor description (shown as display name) |
      | `agent` | string | Yes | Agent reference — `agentReference` from `get_agent_metadata` (`{db}:{schema}.{name}` or OTel `service_name`) |
      | `warehouse` | string | Yes | Warehouse name or UUID holding the agent's traces |
      | `alert_conditions` | array | Yes | List of alert condition objects (see below) |
      | `trace_table` | string | No | Explicit trace table — only for non-ClickHouse OTel agents |
      | `agent_span_filters` | array | No | Optional span-scope refinement; at most ONE filter object |
      | `is_agent_trace_aggregation` | boolean | No | Aggregate per trace instead of per span (OTel only) |
      | `aggregate_by` | string | No | Time-window bucketing (`hour`/`day`/`week`/`month`) |
      | `sensitivity` | string | No | Anomaly detection sensitivity for AUTO operators |
      | `schedule_type` | string | No | `fixed` (default) or `manual` |
      | `interval_minutes` | int | No | Default `60`; at least 60 and a multiple of 60 |
      | `tags` | array | No | Key-value tags, e.g. `[{"name": "agent", "value": "<AGENT_NAME>"}]`. Tag every monitor you create for an agent with its name so they're groupable (see Performance pillar) |
      | `domain_uuids` | array | No | Domain UUIDs to assign this monitor to — the agent-onboarding playbook passes the footprint's single resolved domain on every create (see agent-monitor-creation.md conventions) |
      | `monitor_uuid` | string | No | UUID of an existing monitor to update in place (PUT semantics) |
      | `is_draft` | boolean | No | Save as a draft (not active). **On edit, omitting this un-drafts an existing draft** — re-pass `is_draft=True` when updating a draft that should stay a draft |
      | `dry_run` | boolean | No | Default `True` — preview YAML; set `False` to deploy |
      
      ## Alert conditions
      
      Each condition has:
      
      | Field | Required | Description |
      |-------|----------|-------------|
      | `metric` | Yes | The metric to compute (see Metrics below). |
      | `operator` | Yes | See Operators below. |
      | `fields` | Depends | PARSED_SPANS field name(s). Required for every manual operator and range. Omit for `ROW_COUNT_CHANGE`. |
      | `thresholdValue` | Depends | Required for single-value operators (`GT`/`GTE`/`LT`/`LTE`/`EQ`/`NEQ`). camelCase — NOT `threshold_value`. |
      | `lowerThreshold` / `upperThreshold` | Depends | Both required for `INSIDE_RANGE` / `OUTSIDE_RANGE`; `lowerThreshold` ≤ `upperThreshold`. |
      | `type` | No | `threshold` (default) or `noop` (collect without alerting; pair with `operator: "NOOP"`). |
      
      ### Operators
      
      - **Anomaly detection:** `AUTO`, `AUTO_HIGH`, `AUTO_LOW` — learn thresholds
        automatically. Do NOT pass any threshold.
      - **Single threshold:** `GT`, `GTE`, `LT`, `LTE`, `EQ`, `NEQ` — require
        `thresholdValue` **and** `fields`. (The not-equal operator is `NEQ`, not `NE`.)
      - **Range:** `INSIDE_RANGE`, `OUTSIDE_RANGE` — require both `lowerThreshold` and
        `upperThreshold` (and `fields`).
      - **Collect-only:** `NOOP` — record the metric without alerting; pair with
        `type: "noop"`.
      
      ### Metrics
      
      **Table-level metric (no `fields`; anomaly / NOOP operators only):**
      
      | Metric | Notes |
      |--------|-------|
      | `ROW_COUNT_CHANGE` | Anomalous span volume. `AUTO` / `AUTO_HIGH` / `AUTO_LOW` (or `NOOP`) only; no `fields`. |
      
      **Field-level metrics (must specify `fields`), by field type:**
      
      | Metric | Field type |
      |--------|-----------|
      | `NUMERIC_MEAN`, `NUMERIC_MEDIAN`, `NUMERIC_MIN`, `NUMERIC_MAX`, `NUMERIC_STDDEV`, `SUM` | numeric |
      | `PERCENTILE_20`, `PERCENTILE_40`, `PERCENTILE_60`, `PERCENTILE_80`, `PERCENTILE_95`, `PERCENTILE_99` | numeric |
      | `ZERO_RATE`, `ZERO_COUNT`, `NEGATIVE_RATE`, `NEGATIVE_COUNT` | numeric |
      | `TRUE_RATE`, `TRUE_COUNT`, `FALSE_RATE`, `FALSE_COUNT` | boolean |
      | `NULL_RATE`, `NULL_COUNT`, `NON_NULL_COUNT` | any |
      | `UNIQUE_COUNT` | numeric / text / date |
      
      Numeric metrics apply to `duration_sec` / `*_tokens` / `status_code`; boolean metrics
      apply to `is_tool_call` / `is_llm_call` / `has_prompts` / `has_completions` (the last
      three are OTel/ClickHouse only — platform/Cortex agents lack them); `NULL_RATE`
      applies to any field. Duplicate `(metric, field)` pairs across conditions
      are rejected.
      
      Common numeric fields: `duration_sec`, `total_tokens`, `prompt_tokens`,
      `completion_tokens`, `status_code` (span grain); `span_count`, `llm_call_count`,
      token totals, `duration_sec` (trace grain). Do NOT use `duration_ms` or raw table
      column names.
      
      **Databricks Genie agents (`backend_class: databricks_genie`) emit no token or
      model data** — token-usage metrics on them are permanently silent. Propose
      latency (`duration_sec`), volume, and error/outcome metrics instead.
      
      ## Trace aggregation
      
      `is_agent_trace_aggregation=True` rolls spans up per trace. Constraints:
      
      - **OTel agents only.** Platform agent references (`{database}:{schema}.{name}`) are
        rejected — the per-trace query isn't built for them. Target an OTel `service_name`,
        or pass an explicit `trace_table`.
      - **No span filters are supported** — remove `agent_span_filters` entirely (including
        `agent`). The monitor's agent is identified by the top-level `agent` parameter, not a
        span filter; the per-trace result has no columns a span filter could match against.
      - Use the trace-aggregation field names (`span_count`, `llm_call_count`, trace-summed
        tokens, total `duration_sec`). See `agent-span-fields.md`.
      
      ## Performance pillar (baseline monitor set)
      
      When the user asks for performance monitoring on a named agent — latency, token cost,
      errors, or a latency SLO ("set up performance monitoring for X", "alert me when X gets
      slow or expensive") — propose this exact monitor set rather than inventing one-offs.
      Discover the agent first (`get_agent_metadata` → `agentReference`, `backend_class`,
      `warehouse_uuid`), apply the backend gating below, then present the whole set with
      `dry_run=True` previews.
      
      Shared defaults for every monitor in the set:
      
      - **Daily schedule** — `interval_minutes=1440`.
      - **Tag the agent** — `tags=[{"name": "agent", "value": "<AGENT_NAME>"}]` on every
        monitor, so the pillar's monitors are groupable per agent.
      - **Draft-first when the user wants review** — pass `is_draft=True` to stage the set
        without activating it (and remember the un-draft-on-edit footgun in Parameters).
      - **Grain** — on OTel agents create monitors 1, 2, and 5 with
        `is_agent_trace_aggregation=True` so they track end-to-end interactions; other
        backends use the span-grain default. Monitors 3 and 4 stay span-grain everywhere
        (`SUM` of tokens is the same total either way, and `status_code` is not a
        trace-aggregation field).
      - **Cap-constrained surfaces** — if the consuming surface limits how many monitors may be
        proposed, keep priority order 1 → 4 → 3 → 2 → 5 and say which monitors were cut for the cap.
      
      The set:
      
      | # | Monitor | `alert_conditions` | Notes |
      |---|---------|--------------------|-------|
      | 1 | Latency anomaly | `NUMERIC_MEDIAN` + `PERCENTILE_95` on `duration_sec`, both `AUTO` | ONE monitor, TWO conditions (do not split) — catches drift in the typical and the worst experience. Distinct metrics on the same field are fine; only duplicate metric+field pairs are rejected. p50 = `NUMERIC_MEDIAN`: there is no `PERCENTILE_50` metric — never substitute `PERCENTILE_40`. |
      | 2 | Token anomaly | `NUMERIC_MEDIAN` + `PERCENTILE_95` on `total_tokens`, both `AUTO` | Per-interaction cost drift; same one-monitor-two-conditions shape. |
      | 3 | Daily token spend | `SUM` on `total_tokens`, `AUTO`, with `aggregate_by="day"` | Aggregate cost creep. `aggregate_by` buckets the datapoints; `interval_minutes` only sets the run cadence — set both. |
      | 4 | Error-level anomaly | `NUMERIC_MEAN` on `status_code`, `AUTO_HIGH` | Error-rate proxy — see rationale below. |
      | 5 | Latency SLO | `PERCENTILE_95` on `duration_sec`, `GT`, `thresholdValue` = measured p95 × 1.2 | Separate monitor, measure-then-propose — see below. |
      
      **Why monitor 4 is `NUMERIC_MEAN` on `status_code`:** `status_code` is the one error
      signal available on every backend (OTel semantics: 0 = unset, 1 = ok, 2 = error).
      Healthy spans sit at 0/1 and errored spans at 2, so the mean rises with the
      errored-span share — `AUTO_HIGH` on the mean fires when errors spike. Say so in the
      proposal: this is a *proxy* for error rate built from built-in metrics, not a true
      rate. Do NOT use `TRUE_RATE` (no backend has a boolean error field) or
      `exception_type` (not available on all backends). Monte Carlo may already have
      auto-created a dedicated error-rate monitor for the agent — check `get_monitors`
      first, and if one exists present it as the error coverage instead of duplicating it.
      
      **Monitor 5 is measure-then-propose — never invent an SLO threshold:**
      
      1. Sample recent traces: `get_agent_traces` for the agent (default 14-day lookback),
         `first=50`, paging with `after`/`end_cursor` up to ~3 pages (≤150 traces).
      2. Compute the 95th percentile of the sampled `duration_seconds`.
      3. Propose `thresholdValue` = p95 × 1.2 (20% headroom), rounded to a clean number.
      4. Show the evidence: measured p95, sample size and window, and the proposed
         threshold — stating explicitly that the headroom and threshold are the user's to
         adjust.
      
      Keep the SLO in its own monitor — adding a second `PERCENTILE_95` + `duration_sec`
      condition to monitor 1 is rejected as a duplicate metric+field pair. Propose it for
      OTel agents only (trace grain, so the trace-level measurement matches the monitored
      metric); on other backends skip it — a span-grain p95 tracks individual steps, not
      end-to-end latency, so a trace-based threshold would never fire — and note that
      monitor 1's anomaly conditions cover latency drift.
      
      **Backend gating** (from `backend_class`; explain each skip in one line in the
      proposal):
      
      | `backend_class` | Monitors | Adjustments |
      |---|---|---|
      | `ao_clickhouse_otel` / `customer_otel_trace_table` | 1–5 | 1, 2, 5 at trace grain |
      | `databricks_mlflow_sdk` | 1–4 | span grain; no trace aggregation, so no SLO monitor |
      | `platform_agent` (Cortex) | 1–4 | span grain; no trace aggregation, so no SLO monitor |
      | `databricks_mlflow_ka` | 1, 4 | token fields are NULL — skip 2–3 |
      | `databricks_genie` | 1 (median only), 4 | no token data — skip 2–3; latency percentiles are not meaningful on Genie's fabricated span tree, so drop the `PERCENTILE_95` condition from 1 and skip 5 |
      
      ## Examples
      
      The `agent` value below comes from `get_agent_metadata`'s `agentReference` field —
      a platform `{database}:{schema}.{name}` reference or an OTel `service_name`.
      
      ### Latency anomaly detection (platform agent reference)
      
      ```
      create_or_update_agent_metric_monitor(
          description="Chat Agent latency monitor",
          agent="analytics:agents.support_bot",
          warehouse="Prod Warehouse",
          alert_conditions=[
              {"metric": "NUMERIC_MEAN", "operator": "AUTO", "fields": ["duration_sec"]}
          ],
          dry_run=True
      )
      ```
      
      ### Span volume anomaly detection (OTel service_name, ROW_COUNT_CHANGE — no fields)
      
      ```
      create_or_update_agent_metric_monitor(
          description="Chat Agent span volume monitor",
          agent="checkout-agent",
          warehouse="Prod Warehouse",
          alert_conditions=[
              {"metric": "ROW_COUNT_CHANGE", "operator": "AUTO"}
          ],
          agent_span_filters=[
              {"workflow": {"value": "Chat Agent"}}
          ],
          dry_run=True
      )
      ```
      
      ### Token usage with threshold (span grain)
      
      ```
      create_or_update_agent_metric_monitor(
          description="Alert when mean token usage exceeds 5000",
          agent="analytics:agents.support_bot",
          warehouse="Prod Warehouse",
          alert_conditions=[
              {"metric": "NUMERIC_MEAN", "operator": "GT", "fields": ["total_tokens"],
               "thresholdValue": 5000}
          ],
          dry_run=True
      )
      ```
      
      ### Trace-level token rollup (OTel agent, trace aggregation)
      
      ```
      create_or_update_agent_metric_monitor(
          description="Alert on anomalous per-trace token totals",
          agent="checkout-agent",
          warehouse="Prod Warehouse",
          alert_conditions=[
              {"metric": "NUMERIC_MEAN", "operator": "AUTO", "fields": ["total_tokens"]}
          ],
          is_agent_trace_aggregation=True,
          dry_run=True
      )
      ```
      
      ### Tool-call rate (OTel agent, boolean field)
      
      ```
      create_or_update_agent_metric_monitor(
          description="Alert when tool-call rate drops",
          agent="checkout-agent",
          warehouse="Prod Warehouse",
          alert_conditions=[
              {"metric": "TRUE_RATE", "operator": "LT", "fields": ["is_tool_call"],
               "thresholdValue": 0.1}
          ],
          dry_run=True
      )
      ```
      
      ### Latency SLO threshold (Performance pillar monitor 5 — OTel agent, trace grain, draft)
      
      ```
      create_or_update_agent_metric_monitor(
          description="checkout-agent latency SLO — trace p95 under 42s (measured p95 35s + 20% headroom)",
          agent="checkout-agent",
          warehouse="Prod Warehouse",
          alert_conditions=[
              {"metric": "PERCENTILE_95", "operator": "GT", "fields": ["duration_sec"],
               "thresholdValue": 42}
          ],
          is_agent_trace_aggregation=True,
          interval_minutes=1440,
          tags=[{"name": "agent", "value": "checkout-agent"}],
          is_draft=True,
          dry_run=True
      )
      ```
      
      ## Common errors
      
      | Error message | Cause | Fix |
      |--------------|-------|-----|
      | Warehouse not found | `warehouse` omitted or wrong | Pass the agent's `warehouse_uuid` from `get_agent_metadata`; if null, list warehouses via `get_warehouses` |
      | invalid / unresolvable `agent` reference | The `agent` value wasn't taken from `get_agent_metadata` | Use the exact `agentReference` value — do not construct it by hand, and never pass an MCON |
      | "Field X doesn't exist" | Field name not in the PARSED_SPANS schema | Check `agent-span-fields.md`; use `duration_sec` not `duration_ms` |
      | metric/field-type mismatch | Numeric metric on a boolean field (or vice versa) | Numeric metrics on numeric fields, boolean metrics on boolean fields, `NULL_RATE` on any |
      | `ROW_COUNT_CHANGE` rejected | Included `fields`, or used a manual operator | Drop `fields`; use `AUTO`/`AUTO_HIGH`/`AUTO_LOW` or `NOOP` |
      
    • agent-monitor-creation.md 25.6 KB
      # Agent Monitor Creation Procedure
      
      This is the agent monitor creation procedure for AI agent observability. Follow
      these steps in order when a user asks to monitor their AI agents — setting up
      alerts on agent behavior or creating agent monitors.
      
      All tools are available via the `monte-carlo-mcp` MCP server.
      
      ---
      
      ## Step 1: Discover agents
      
      Call `get_agent_metadata` to list all AI agents in the account. Present agent
      names to the user (never expose MCONs or internal IDs). Ask which agent(s)
      they want to monitor.
      
      Key fields in the response:
      
      | Field | Description |
      |-------|-------------|
      | `agentName` | Human-readable agent name |
      | `agentReference` | The value to pass as the `agent` arg when creating monitors — a platform `{database}:{schema}.{name}` reference (Snowflake Cortex / Databricks) or an OpenTelemetry `service_name`. May be null for agents that cannot be referenced. |
      | `traceTableMcon` | Trace table MCON — used as the `trace_table_mcon` input for the read tools (`get_agent_conversations`, `get_agent_conversation`, `get_agent_traces`, `get_agent_segments`; the parameter is named `mcon` on `get_agent_trace`) |
      | `sourceType` | `TRACE_TABLE` (custom) or `PLATFORM_AGENT` (Monte Carlo native) |
      | `backend_class` | Which backend the agent's traces live in — `ao_clickhouse_otel`, `platform_agent`, `customer_otel_trace_table`, `databricks_genie`, `databricks_mlflow_sdk`, or `databricks_mlflow_ka`. Null when the server could not classify the agent (or predates the field). |
      | `warehouse_uuid` | Warehouse holding the agent's trace data — the value to pass as the `warehouse` arg when creating monitors. Null when the warehouse was deleted or cannot be resolved; fall back to `get_warehouses` (see Warehouse below). |
      | `warehouse_name` | Display name of that warehouse — what you show the user. Null alongside `warehouse_uuid`; fall back to `get_warehouses`. |
      
      **What `backend_class` tells you about capabilities:** conversation-grain
      evaluation monitors (`is_agent_conversation_aggregation=True`) are supported for
      `ao_clickhouse_otel`, `platform_agent` (Snowflake Cortex), and `databricks_genie`;
      the MLflow classes are span-only (the backend rejects conversation aggregation for
      them). At conversation grain, set `includeToolCalls: true` on every eval transform
      by default — it adds the agent's tool calls (name, inputs, outputs, errors) to the
      judged conversation as clearly identifiable TOOL entries, in call order, so evals
      score what the agent did, not just what it said. Omit it only for pure style/tone
      judges; the field is rejected at span grain. `databricks_genie` agents emit no
      token or model data — skip token-usage
      metrics for them (see `agent-metric-monitor.md`). A null `backend_class` means the
      server couldn't classify the agent — default to span-grain proposals.
      
      **Duplicate agent names:** The same agent name may appear more than once (e.g.,
      deployed in both prod and staging). Each entry is distinguished by its own
      `agentReference`, `traceTableMcon`, and warehouse — ask the user which one they
      want to monitor and pass that entry's `agentReference` verbatim. When you ask the
      user to choose, present each entry's `warehouse_name`, never a UUID.
      
      ---
      
      ## Step 2: Investigate agent behavior
      
      Use the read tools to understand the agent's behavior before suggesting monitors.
      These read tools are your sampling surface — do **not** query the trace store with
      SQL. Agents are tracked as agents, not tables: platform agents (Snowflake Cortex /
      Databricks) have no queryable trace table, and for custom agents the raw table's
      columns are not the fields monitors use.
      
      ### 2a. Review recent conversations
      
      Call `get_agent_conversations` with the agent's `agent_name` and `trace_table_mcon`
      to list recent conversations (newest first). Filter to surface interesting ones —
      `has_errors`, `status`, or turn/token/duration bounds — and set
      `include_transcript=True` to read the prompt/completion transcripts inline. Drill
      into one conversation you already have the id for with `get_agent_conversation`.
      Look for:
      
      - **Error patterns** — spans with error status or failure indicators
      - **Latency outliers** — unusually long durations
      - **Token usage** — high token counts that may indicate inefficiency
      - **Conversation quality** — check prompt/completion text for relevance
      
      ### 2b. Inspect execution shape and traces
      
      Call `get_agent_traces` to list traces with per-trace `workflows`, `tasks`,
      `models`, `count_llm_calls`, `total_tokens`, `duration_seconds`, and error counts —
      sort by the field you plan to monitor to see typical values and outliers. Call
      `get_agent_segments` to enumerate the distinct `workflow` / `task` / `model` values
      so you can scope a monitor to a real segment. Then pick a trace id and call
      `get_agent_trace` to see the full span tree. Look for:
      
      - **Excessive tool calls** — an agent calling the same tool many times
      - **Missing steps** — expected spans that don't appear
      - **Error cascades** — a failed span causing downstream failures
      - **Unusual paths** — the agent taking an unexpected execution route
      
      You are identifying **what** to monitor — you don't need exact percentiles up
      front; anomaly-detection operators (`AUTO`) learn the baseline themselves.
      
      ### 2c. Summarize the agent before proposing
      
      Condense the investigation into a short agent understanding and show it to the
      user — every monitor you propose should trace back to an item in it:
      
      - **Purpose** — 1–2 sentences on what this agent does, grounded in the sampled
        transcripts (e.g. "a revenue-analytics assistant that answers questions about
        bookings").
      - **Conversational?** — multi-turn user conversations (eval-worthy for
        satisfaction / task completion) vs. a batch / single-shot pipeline where
        structural and span checks fit better.
      - **Tools and the dominant span** — which tool spans the agent runs, and which
        one does its core work (the SQL execution tool for an analytics agent,
        retrieval for a RAG agent).
      - **Healthy trajectory shape** — how many times the dominant span runs per answer
        in healthy traces (the per-trace distribution and its max), typical turn counts,
        and latency/token magnitudes. This is the basis for every derived threshold.
      - **Recurring intents** — what users repeatedly ask (from the transcripts) —
        seeds for custom conversation evals.
      - **Observed failure modes** — what actually went wrong in the sample — seeds for
        evals and structural monitors.
      - **Existing monitors** — from `get_monitors`, so proposals don't duplicate
        coverage.
      
      ---
      
      ## Propose with the POBC framing (walk the user through all four pillars)
      
      When the user is setting up monitoring for an agent (rather than asking for one
      specific monitor), structure the proposal around the four pillars of agent
      observability — **Performance, Output, Behavior, Context (POBC)** — and walk
      through them one at a time. Do NOT dump every proposed monitor in one
      monolithic list.
      
      ### 1. Open with the framing
      
      Before presenting any monitors, briefly explain the framework. Use this copy,
      adapting the agent's actual name into the prose where it reads naturally:
      
      > A quick word on how we think about agent observability. Agents fail in four
      > distinct ways, so we monitor four distinct things — **Performance, Output,
      > Behavior, and Context (POBC)**: how efficiently the agent answers, what it
      > says, how it gets there, and the data it stands on.
      >
      > **Performance** — is it fast and affordable? Latency, token cost, and error
      > monitoring catch drift in both the typical experience and the worst one.
      >
      > **Output** — is the agent giving good answers? Evals score response quality
      > (helpfulness, non-answers, user corrections) so quality regressions surface
      > as alerts, not user complaints.
      >
      > **Behavior** — is it working sensibly under the hood? Trajectory monitoring
      > flags runs that loop or take paths a healthy run never takes — including
      > failures the agent recovers from and hides.
      >
      > **Context** — is the data it relies on healthy? The agent's answers are only
      > as good as its upstream tables; we monitor those for freshness, schema
      > changes, and anomalies.
      >
      > Everything below maps to one of these four. Here's the plan:
      
      ### Monitor conventions — every monitor in this playbook
      
      Four conventions apply to EVERY monitor created in this playbook — the agent
      monitors (metric, evaluation, trajectory, validation) AND any warehouse
      data-quality monitor created for the Context pillar (table or field monitors
      on the agent's upstream tables; see `data-monitor-creation.md`):
      
      - **Agent tag on every create.** Pass
        `tags=[{"name": "agent", "value": "<AGENT_NAME>"}]`, where `<AGENT_NAME>` is
        the agent's display name exactly as returned by `get_agent_metadata`
        (trimmed, case preserved — do not slugify or rename). This single tag is the
        footprint contract: one filter retrieves everything the playbook created for
        this agent.
      - **Audit / teardown contract.**
        `get_monitors(monitor_tags=["agent:<AGENT_NAME>"])` (or the UI monitors tag
        filter) returns the agent's full monitoring footprint — use it to audit
        what exists, tune, or tear down everything for an agent. This is why a
        create call without the tag is a defect: it silently drops the monitor out
        of the footprint.
      - **Audience — ask once, apply everywhere.** Before the FIRST create call of
        the playbook (not per monitor), ask the user once which audiences should be
        notified when these monitors fire — call `get_audiences` to list the
        options; the user can pick one, several, or none. Pass the chosen audience
        **names** (labels, never UUIDs) as `audiences` on EVERY monitor created in
        the playbook, and set `failure_audiences` to the same selection unless the
        user asks for a different failure-notification audience. Do not re-ask per
        pillar or per monitor; if the user declines, omit `audiences`.
      - **Domain — same for every monitor.** If the account uses domains, resolve
        one domain for the agent's footprint and pass the same `domain_uuids` on
        every monitor — including the Context-pillar DQ monitors (resolve via the
        domain-assignment steps in `data-monitor-creation.md`) — so the whole
        footprint lives in one domain.
      
      ### 2. Walk through the plan pillar by pillar
      
      Present the pillars in order — Performance, Output, Behavior, Context — one
      short block each:
      
      1. **Evidence** — one or two sentences of what you observed in Step 2 that
         motivates this pillar's monitors ("p95 latency is 40s with outliers over
         three minutes", "several conversations show repeated user corrections").
         If you found nothing notable for a pillar, say so and propose baseline
         coverage anyway — monitoring exists to catch what hasn't happened yet.
      2. **Proposed monitors** — the specific monitors for this pillar, each with
         its monitor type, field or judge, and alert condition.
      3. **Confirm** — ask whether to keep, adjust, or drop this pillar's monitors,
         and fold the answer in before moving to the next pillar.
      
      A healthy agent usually warrants coverage in every pillar you can serve —
      keep proposals broad across pillars, not deep in one.
      
      **Global defaults for proposed monitors** (apply unless the user asks
      otherwise):
      
      - **Daily schedule** — pass `interval_minutes=1440` explicitly; the tools'
        built-in default is hourly (see Schedule configuration below).
      - **Eval sampling** — `sampling_config={"count": 100}` (a fixed 100-sample
        budget per run), not a percentage, so evaluation cost stays predictable as
        traffic grows.
      - **Audience on every create** — apply the playbook-level audience selection
        (see "Monitor conventions — every monitor in this playbook" above); never
        create a monitor without that once-asked selection applied (Step 4).
      
      What each pillar maps to:
      
      | Pillar | Monitor types | Reference |
      |---|---|---|
      | **Performance** | Metric (validation for hard limits) — latency (`duration_sec`), token cost (`total_tokens`), error rate (`status_code`), volume (`ROW_COUNT_CHANGE`) | `agent-metric-monitor.md` |
      | **Output** | Evaluation — lead with the Output-pillar starting packs (see Step 3): baseline pack for every agent, analytics pack for Cortex/Genie. Add predefined judges (`answer_relevance`, `task_completion`, `clarity`, `prompt_adherence`), rule checks (`output_length`, `json_validity`), and one custom eval per recurring user intent or failure mode you observed | `agent-evaluation-monitor.md` |
      | **Behavior** | Trajectory (validation for aggregate assertions) — runaway loops (`SPAN_OCCURRENCE`), missing or mis-ordered steps (`SPAN_RELATION`), token budgets | `agent-trajectory-monitor.md`, `agent-validation-monitor.md` |
      | **Context** | Table monitors (freshness / schema changes / volume) on the agent's upstream tables, plus an optional `Context for {AGENT_NAME}` data product wrapper — see below | `data-table-monitor.md` |
      
      Mind the backend caveats from Step 1 (`backend_class`): no token or model
      metrics for Genie / Knowledge Assistant agents, and conversation-grain evals
      only on OTel/ClickHouse, Snowflake Cortex, and Genie — with
      `includeToolCalls: true` on every conversation-grain eval transform by default
      (rejected at span grain; see Step 1). Aggregate (per-trace)
      validation assertions require `is_agent_trace_aggregation=True`, supported
      only on `ao_clickhouse_otel` / `customer_otel_trace_table` agents — on other
      backends, use per-span assertions instead.
      
      **Context — monitor the tables the agent reads.** Unlike the other pillars,
      Context coverage lives on warehouse tables, not spans. Automatic
      lineage-derived table discovery is not available on this surface, so ask the
      user which upstream tables the agent depends on (the tables its SQL tools
      query, its knowledge bases are built from, or its features are loaded from) —
      never guess table names, and never drop the pillar silently. On the named
      tables:
      
      1. **Create the table monitors** with `create_or_update_table_monitor`,
         following `data-table-monitor.md` for warehouse resolution and asset
         selection (scope as narrowly as that reference allows). The tool's
         default alert conditions are exactly the Context coverage — freshness,
         schema changes, and volume — so omit `alert_conditions` unless the user
         asks for more. Apply the Monitor conventions above on every create: the
         `agent` tag, the playbook `audiences`, and `domain_uuids`. Dry-run
         preview first, deploy on explicit confirmation, like every other create
         in this playbook.
      2. **Optionally wrap the tables in a data product** named
         `Context for {AGENT_NAME}` via `create_or_update_data_product`: pass the
         tables' `mcons` (from `search` / `get_table` — do not guess them) and a
         `description` naming the agent. Keep the default `dry_run=True` to show
         the asset-footprint preview; set `dry_run=False` only on explicit
         confirmation. The preview's asset count can exceed the tables you
         named — the tool's backend automatically expands the footprint to
         include their upstream dependencies. When the count is notably larger
         than the named set, explain to the user what is being added before
         asking them to confirm the live create. Two caveats: data products
         take `audience_ids` — UUIDs from `get_audiences` — unlike monitors,
         which take audience names; and the data product itself carries no
         `agent` tag (the footprint contract rides on the monitors). If the
         live create is rejected because the account lacks the Data Mesh
         module, that is terminal for the wrapper: say so in one line (their
         Monte Carlo representative can enable it) and keep the table monitors
         — the data product is packaging; the monitors are the pillar.
      3. **Add field-level depth** where the user wants specific field checks
         (null rates, distributions, custom rules) on an upstream table: use the
         data-monitor workflow's field-level references (`data-metric-monitor.md`,
         `data-validation-monitor.md`, `data-custom-sql-monitor.md`), carrying
         the same tag, audiences, and domain on every create.
      
      If the user cannot name any upstream tables, present the pillar as a
      recommendation — name what you would monitor and why — rather than failing
      or silently dropping it.
      
      ### 3. Create the confirmed monitors
      
      Once the user has confirmed the pillars, continue to Step 3 (pick each
      monitor's reference doc) and Step 4 (dry-run preview, applying the
      already-collected audience selection, creation on explicit confirmation) for
      each approved monitor.
      
      ---
      
      ## The `agent` reference
      
      All four `create_or_update_agent_*_monitor` tools author the monitor's source from
      a single top-level **`agent`** argument (there is no `dw_id` and no `data_source`
      argument — the `agent` reference is the whole source). Two accepted forms:
      
      - **Platform agent reference** — `{database}:{schema}.{name}` (Snowflake Cortex /
        Databricks agents), e.g. `analytics:agents.support_bot`.
      - **OpenTelemetry `service_name`** — for OTel-instrumented agents, e.g. `checkout-agent`.
      
      Get the exact value from `get_agent_metadata`'s **`agentReference`** field and pass
      it verbatim — never construct, modify, or truncate it, and never pass an MCON. Two
      optional companions:
      
      - `trace_table` — only for non-ClickHouse OTel agents whose trace storage cannot be
        inferred from the agent reference.
      - The per-type reference tells you whether `warehouse` is required (see below).
      
      ---
      
      ## Warehouse
      
      `warehouse` names the warehouse the agent's trace data lives in — pass it as a name
      or UUID. Use the agent entry's `warehouse_uuid` from `get_agent_metadata` (and its
      `warehouse_name` when talking to the user); when both are null, use `get_warehouses`.
      Whether `warehouse` is required or optional depends on the monitor type — see the
      per-type reference.
      
      ---
      
      ## Agent span filters
      
      The optional `agent_span_filters` parameter refines which spans are monitored. It
      accepts **at most one** filter object. Each field of the object holds a
      `{"value": "..."}` sub-object.
      
      | Filter field | Description | Example |
      |-------------|-------------|---------|
      | `agent` | Filter by agent name | `{"agent": {"value": "My Agent"}}` |
      | `workflow` | Filter by workflow name | `{"workflow": {"value": "Chat Agent"}}` |
      | `task` | Filter by task name | `{"task": {"value": "call_model"}}` |
      | `spanName` | Filter by span name | `{"spanName": {"value": "ChatBedrockConverse.chat"}}` |
      
      Multiple fields can be combined in the single filter object:
      
      ```json
      [{"workflow": {"value": "Chat Agent"}, "task": {"value": "call_model"}}]
      ```
      
      `agent_span_filters` is a refinement and is optional — the `agent` reference already
      scopes the monitor. Some monitor types restrict which fields are allowed here (e.g.
      trajectory monitors, and trace-aggregated metric / validation monitors) — see the
      per-type reference for the exact rule.
      
      ---
      
      ## Schedule configuration
      
      **Propose daily schedules by default** — pass `interval_minutes=1440` explicitly.
      Schedule is set via two top-level args, not a nested object:
      
      - `schedule_type` — defaults to `fixed`. Valid values: `fixed`, `manual`.
      - `interval_minutes` — defaults to `60` (hourly), so omitting it creates an hourly
        monitor. The floor and alignment differ per monitor type — see each reference.
      
      No agent monitor accepts a dynamic schedule — use `fixed` or `manual`. Daily is
      the right cadence for most agents. If you judge an agent critical enough that a
      same-hour alert would matter, suggest hourly to the user and let them decide —
      the default stays daily unless they opt in.
      
      ---
      
      ## Time filter configuration
      
      Used by trajectory and validation monitors. The `timeField` is an object with
      a `field` property — always use `ingest_ts`:
      
      ```json
      {"timeField": {"field": "ingest_ts"}, "lookbackInHrs": 24}
      ```
      
      ---
      
      ## Step 3: Choose the right monitor type
      
      Based on your investigation, recommend one or more monitor types. Read the
      corresponding reference doc for the detailed creation guide.
      
      | I want to... | Monitor type | Reference file |
      |-------------|-------------|----------------|
      | Track a numeric metric trend (latency, tokens) | Agent Metric | `agent-metric-monitor.md` |
      | Set up performance coverage (latency, token cost, errors, SLO) | Agent Metric — Performance pillar | `agent-metric-monitor.md` |
      | Score output quality with LLM evaluation | Agent Evaluation | `agent-evaluation-monitor.md` |
      | Alert on execution patterns or span sequences | Agent Trajectory | `agent-trajectory-monitor.md` |
      | Assert a logical rule on span data | Agent Validation | `agent-validation-monitor.md` |
      | Monitor span volume over time | Agent Metric | `agent-metric-monitor.md` |
      | Detect answer relevance drops | Agent Evaluation | `agent-evaluation-monitor.md` |
      | Catch runaway tool call loops | Agent Trajectory | `agent-trajectory-monitor.md` |
      | Ensure token count stays below threshold | Agent Validation | `agent-validation-monitor.md` |
      
      **Output-pillar starting packs** — for a newly onboarded agent (or one with no eval coverage
      yet), lead with the named packs from `agent-evaluation-monitor.md` rather than inventing a
      one-off list:
      
      - **Baseline pack — every agent:** the predefined `helpfulness_conversation` judge (plain
        `helpfulness` on span-grain-only backends) plus the `frustration_free_score` template.
        Defaults: daily schedule (`interval_minutes=1440`), `{"count": 100}` sampling, an `agent`
        tag (`{"name": "agent", "value": "<AGENT_NAME>"}`) on every monitor, and
        `includeToolCalls: true` on every conversation-grain transform (see the `backend_class`
        capabilities in Step 1).
      - **Analytics pack — only when `backend_class` is `platform_agent` (Snowflake Cortex) or
        `databricks_genie`:** the `answer_attempt_score` and `user_correction` templates — the
        dominant NL2SQL/analytics failure modes are deflected answers and user-corrected answers.
        Do not propose this pack for other agents.
      
      Render each template with the agent's actual name and observed intents (never boilerplate) and
      show the full prompt text for approval — see "Custom-prompt template library" and
      "Output-pillar eval packs" in `agent-evaluation-monitor.md`.
      
      After selecting the monitor type, **read the reference doc** for that type to
      get the detailed parameter guide, examples, constraints, and creation workflow.
      For a blanket "performance monitoring" ask, follow the **Performance pillar**
      baseline set in `agent-metric-monitor.md` rather than assembling one-offs.
      
      ### Behavior monitors — two trajectory proposals for (almost) every agent
      
      Grounded in the Step 2c summary, propose these two patterns whenever they apply
      (`agent-trajectory-monitor.md` has the full playbooks and payload shapes):
      
      1. **Runaway loop — create live.** SPAN_OCCURRENCE on the agent's dominant tool
         span, threshold derived from the observed per-trace occurrence distribution:
         max observed + headroom, never a stock number. The proposal's evidence must
         show the dominant span, the distribution, and the derived threshold with its
         headroom rationale — plus a pre-create breach `preview` (dry run) proving zero
         historical matches; if the preview breaches, the sample missed the heavy tail
         (e.g. multi-turn accumulation in one trace) — re-derive from a wider window.
         Zero historical matches is the point — it is a regression guardrail that stays
         silent until the agent's behavior regresses.
      2. **Ungrounded-in-data — create as a DRAFT** (only for agents that answer
         questions from data). Negated `occurs_with` SPAN_RELATION: an answer was
         produced without the agent's data-access tool span. Show a breach `preview`
         (dry run) as evidence, then create with `is_draft=True` — generic questions
         legitimately skip the data tool, and the LLM-judge filter needed to separate
         them from real data questions cannot be combined with a trajectory condition
         yet.
      
      Tag both with the agent's name (`tags=[{"name": "agent", "value": "<AGENT_NAME>"}]`)
      and schedule them daily (`interval_minutes=1440`).
      
      Beyond these two, propose **agent-tailored behavioral custom prompts** for
      behaviors a span pattern can't see — e.g. "did the agent claim it ran a query it
      never executed?", "did the agent re-ask for information the user already gave?".
      One boolean `custom_prompt` per behavior, alerting on `TRUE_RATE` / `FALSE_RATE`,
      at conversation grain where the backend supports it (see `backend_class` in
      Step 1 and `agent-evaluation-monitor.md`) — keep `includeToolCalls: true` on
      these so the judge can see the tool calls it is judging.
      
      ---
      
      ## Step 4: Create the monitor
      
      All four tools follow the same **two-call preview-then-confirm pattern** as the data
      monitor tools: the first call (`dry_run=True`, the default) returns the rendered MaC
      YAML for review; the second call (`dry_run=False`) deploys the monitor live and
      returns its UUID. Pass `monitor_uuid` on either call to update an existing agent
      monitor in place instead of creating a new one (PUT semantics — re-pass every field
      you want to keep, since omitted fields revert to defaults).
      
      1. **Always start with `dry_run=True`** (the default). Show the user the
         configuration preview (the rendered YAML).
      2. **Apply the playbook-level audience selection** (see "Monitor conventions —
         every monitor in this playbook"): the audience question was already asked
         once before the playbook's first create — pass that same selection of
         audience **names** (not UUIDs) as the `audiences` list, and default
         `failure_audiences` to the same selection. Fall back to asking here (one
         question, `get_audiences` for options) only when the playbook-level ask has
         not happened (e.g. the user jumped straight to a single monitor outside the
         walkthrough).
      3. After showing the preview, offer to create or adjust settings.
      4. Only set `dry_run=False` when the user explicitly confirms creation.
      
      ---
      
      ## Field name reference
      
      See `agent-span-fields.md` for the complete list of known span field names
      available in agent monitors. Do not guess field names — use only the ones
      documented there.
      
    • agent-span-fields.md 5.1 KB
      # Known Span Field Names (PARSED_SPANS Layer)
      
      > **CRITICAL:** Do NOT run `SHOW COLUMNS` or `SELECT *` on the trace table to discover
      > field names — the raw table has different columns. Use the field names listed below.
      
      Monte Carlo maintains a **PARSED_SPANS** transformation view on top of each
      agent's raw trace table. This view extracts structured fields from the raw
      OTLP JSON. **You cannot query PARSED_SPANS directly via SQL** — it is not
      a real table. However, all agent monitors operate on these parsed fields,
      so you MUST use these field names when creating monitors.
      
      Do NOT run `SHOW COLUMNS` or `SELECT *` on the trace table to discover
      field names for monitors — the raw table has different columns (e.g.,
      `VALUE`, `FILENAME`, `INGEST_TS`, `DATE_PART`) that are NOT the fields
      monitors use. Instead, use the field names listed below.
      
      > **The read tools use different names and units for some fields.** The names below
      > are the **monitor** field names. The tools you sample with do not all match them:
      > `get_agent_traces` returns `duration_seconds` and `count_llm_calls` (the monitor
      > fields are `duration_sec` and `llm_call_count`), and `get_agent_trace` reports
      > `duration` in **milliseconds** (the monitor field `duration_sec` is in seconds).
      > Always use the monitor field names below in monitor payloads — never copy a read
      > tool's column name or unit into a monitor.
      
      ## Span-level fields (default, per-span rows)
      
      | Field | Type | Description |
      |-------|------|-------------|
      | `agent` | STRING | Agent name |
      | `trace_id` | STRING | Trace identifier |
      | `span_id` | STRING | Span identifier |
      | `parent_span_id` | STRING | Parent span identifier |
      | `workflow` | STRING | Workflow name |
      | `task` | STRING | Task name |
      | `span_name` | STRING | Span operation name |
      | `model_name` | STRING | Model used |
      | `prompts` | ARRAY | LLM prompt messages (evaluation transforms) |
      | `completions` | ARRAY | LLM completion messages (evaluation transforms) |
      | `total_tokens` | INTEGER | Total token count (prompt + completion) |
      | `prompt_tokens` | INTEGER | Input/prompt token count |
      | `completion_tokens` | INTEGER | Output/completion token count |
      | `duration_sec` | FLOAT | Span duration in seconds |
      | `status_code` | INTEGER | Span status code (`2` = error) — numeric; compare with `"2"`, not `"ERROR"` |
      | `is_tool_call` | BOOLEAN | Whether the span is a tool call |
      | `is_llm_call` | BOOLEAN | Whether the span is an LLM call *(OTel/ClickHouse only)* |
      | `has_prompts` | BOOLEAN | Whether the span has prompt messages *(OTel/ClickHouse only)* |
      | `has_completions` | BOOLEAN | Whether the span has completion messages *(OTel/ClickHouse only)* |
      | `start_time` | TIMESTAMP | Span start timestamp |
      | `end_time` | TIMESTAMP | Span end timestamp |
      | `ingest_ts` | TIMESTAMP | Ingestion timestamp (use for time filters) |
      
      ### Platform vs OpenTelemetry availability
      
      Platform (Snowflake Cortex / Databricks native) agents expose the core numeric and
      text fields — `duration_sec`, `total_tokens`, `prompt_tokens`, `completion_tokens`,
      `status_code`, `is_tool_call` — but **not** `is_llm_call`, `has_prompts`, or
      `has_completions`. Those presence/kind flags exist only for OTel/ClickHouse agents.
      Stick to the core fields unless you know the agent is OTel-instrumented.
      
      **Databricks Genie agents (`backend_class: databricks_genie`) emit NO token or
      model data** — `total_tokens` / `prompt_tokens` / `completion_tokens` are always
      empty, so don't build token-usage metrics for them; prefer latency
      (`duration_sec`), volume, and error/outcome signals.
      
      ## Trace-aggregation fields (is_agent_trace_aggregation=True)
      
      When `is_agent_trace_aggregation=True`, rows are aggregated per trace (OpenTelemetry
      agents only):
      
      | Field | Type | Description |
      |-------|------|-------------|
      | `agent` | STRING | Agent name |
      | `trace_id` | STRING | Trace identifier |
      | `span_count` | INT | Number of spans in the trace |
      | `llm_call_count` | INT | Number of LLM calls in the trace |
      | `prompt_tokens` | INTEGER | Total prompt tokens across the trace |
      | `completion_tokens` | INTEGER | Total completion tokens across the trace |
      | `total_tokens` | INTEGER | Total tokens across the trace |
      | `duration_sec` | FLOAT | Total trace duration in seconds |
      | `start_time` | TIMESTAMP | Earliest span start in the trace |
      | `end_time` | TIMESTAMP | Latest span end in the trace |
      | `ingest_ts` | TIMESTAMP | Ingestion timestamp |
      
      ## Conversation-aggregation fields (is_agent_conversation_aggregation=True)
      
      Agent evaluation monitors can aggregate per conversation
      (`is_agent_conversation_aggregation=True` — supported for OpenTelemetry/ClickHouse,
      Snowflake Cortex, and Databricks Genie agents; Databricks MLflow agents are
      span-only). At this grain,
      `alert_conditions.fields` may reference these raw conversation columns alongside any
      judge output field:
      
      | Field | Type | Description |
      |-------|------|-------------|
      | `turn_count` | INTEGER | Number of turns in the conversation |
      | `duration_seconds` | FLOAT | Total conversation duration in seconds |
      | `status` | STRING | Conversation status |
      
      Plus the conversation-grain judge output fields (e.g. `relevance_score`,
      `task_completion_score`) — see `agent-evaluation-monitor.md`.
      
    • agent-trajectory-monitor.md 16.1 KB
      # Agent Trajectory Monitor
      
      ## When to use
      
      Flag a whole trace by its execution pattern. Best for:
      
      - **Detecting excessive tool/step calls** — e.g. "search called more than 5 times"
      - **Detecting too-few calls** — e.g. "the retry step ran fewer than 2 times"
      - **Catching runaway loops**
      - **Broken execution order or a missing follow-up** — e.g. "generation runs before
        retrieval", "planning runs without validation"
      
      Do NOT use it to assert a rule on a single span's field values (token ceilings,
      non-null) — that's `create_or_update_agent_validation_monitor`. Do NOT use it to
      trend a numeric metric (mean latency, token counts) — that's
      `create_or_update_agent_metric_monitor`.
      
      ## Constraints
      
      > **CRITICAL:** The monitor's source is the `agent` reference. Pass the
      > `agentReference` value from `get_agent_metadata` verbatim — a platform
      > `{database}:{schema}.{name}` reference or an OTel `service_name`. Never modify,
      > truncate, or reconstruct it, and never pass an MCON.
      
      > **CRITICAL:** Conditions are OR-combined — a trace is flagged if ANY condition
      > matches. `operator` defaults to `OR`; **`operator: "AND"` is rejected.** To require
      > several patterns to all hold, use separate monitors.
      
      > **CRITICAL:** Trajectory `agent_span_filters` allow only the `agent` field — at
      > most one filter, e.g. `agent_span_filters=[{"agent": {"value": "My Agent"}}]`.
      > Setting `workflow`, `task`, or `spanName` there causes a validation error — those go
      > in the condition's `spanField` instead. For OpenTelemetry agents the filter's
      > `agent` value must equal the top-level `agent` reference.
      
      > **CRITICAL:** A span that never appears produces no rows to count, so "occurs 0
      > times" (a missing span) cannot be expressed with SPAN_OCCURRENCE — `EXACTLY 0` and
      > `LESS_THAN 1` are rejected. Use a SPAN_RELATION with a negated predicate to check a
      > span is absent relative to another span (see the missing-step example).
      
      > **IMPORTANT:** `time_filter` is REQUIRED; `timeField` is always
      > `{"field": "ingest_ts"}`.
      
      > **IMPORTANT:** `schedule_type` is `fixed` (default) or `manual` — never dynamic.
      > `interval_minutes` defaults to `60` and must be at least 5 (sub-hourly allowed).
      
      > **IMPORTANT:** `warehouse` is OPTIONAL here — when omitted, the backend falls back
      > to the account's default warehouse, which may not be where the agent's traces live.
      > Prefer passing the agent's `warehouse_uuid` from `get_agent_metadata` explicitly.
      
      ## Key characteristics
      
      - Uses `agent_span_alert_condition` (not `alert_conditions`) — `{"operator": "OR", "conditions": [...]}`
      - Two condition types: `SPAN_OCCURRENCE` (count) and `SPAN_RELATION` (relate two spans)
      - Requires `time_filter` with `timeField` (object `{"field": "ingest_ts"}`) and `lookbackInHrs`
      
      ## Parameters
      
      | Parameter | Type | Required | Description |
      |-----------|------|----------|-------------|
      | `description` | string | Yes | Human-readable monitor description (shown as display name) |
      | `agent` | string | Yes | Agent reference — `agentReference` from `get_agent_metadata` (`{db}:{schema}.{name}` or OTel `service_name`) |
      | `agent_span_alert_condition` | object | Yes | The span pattern that flags a trace (see below) |
      | `time_filter` | object | Yes | `{"timeField": {"field": "ingest_ts"}, "lookbackInHrs": 24}` |
      | `warehouse` | string | No | Warehouse name or UUID; defaults to the account default |
      | `trace_table` | string | No | Explicit trace table — only for non-ClickHouse OTel agents |
      | `agent_span_filters` | array | No | Optional — **only the `agent` field allowed** (no `workflow`/`task`/`spanName`); at most one |
      | `schedule_type` | string | No | `fixed` (default) or `manual` |
      | `interval_minutes` | int | No | Default `60`; at least 5 |
      | `monitor_uuid` | string | No | UUID of an existing monitor to update in place (PUT semantics) |
      | `dry_run` | boolean | No | Default `True` — preview YAML; set `False` to deploy |
      | `preview` | boolean | No | Only together with `dry_run=True`: also runs the monitor's query and returns a pre-create breach preview — whether the conditions would fire right now, a small sample of the underlying data, and the sample row count. Ignored on a real create (`dry_run=False`) |
      | `tags` | array | No | Key-value tags, e.g. `[{"name": "agent", "value": "Support Bot"}]`. Tag every agent monitor with its agent's name — `{"name": "agent", "value": "<AGENT_NAME>"}` — so all of one agent's monitors are filterable as a group |
      | `domain_uuids` | array | No | Domain UUIDs to assign this monitor to — the agent-onboarding playbook passes the footprint's single resolved domain on every create (see agent-monitor-creation.md conventions) |
      | `is_draft` | boolean | No | Default `False`. Save the monitor as a draft — visible in the UI but not running. **On edit, omitting this un-drafts an existing draft** — pass `is_draft=True` explicitly to keep a draft a draft |
      
      Because `preview` only works on a dry run, evidence and creation are **two separate
      calls**: first `dry_run=True, preview=True` to show the user what would fire, then
      (after they confirm) `dry_run=False` — with `is_draft=True` if the monitor should
      land as a draft.
      
      ## agent_span_alert_condition structure
      
      `{"operator": "OR", "conditions": [...]}`. Supply at least one condition. Each is one
      of two types. To OR several patterns together, list them as multiple entries in
      `conditions` — a trace is flagged if any one matches.
      
      ### SPAN_OCCURRENCE — count how many times a span occurs
      
      ```json
      {
        "type": "SPAN_OCCURRENCE",
        "predicate": {"name": "occurs"},
        "spanField": {
          "spanName": {"literal": "ChatBedrockConverse.chat"},
          "task": {"literal": "call_model"},
          "workflow": {"literal": "Chat Agent"},
          "type": "SPAN_FIELD"
        },
        "count": 5,
        "comparisonOperator": "MORE_THAN"
      }
      ```
      
      | Field | Required | Description |
      |-------|----------|-------------|
      | `type` | Yes | `"SPAN_OCCURRENCE"`. |
      | `predicate` | Yes | Always `{"name": "occurs"}`. `occurs` **cannot** be negated. |
      | `spanField` | Yes | The span to count (see spanField below). |
      | `comparisonOperator` | Yes | `"MORE_THAN"`, `"LESS_THAN"`, or `"EXACTLY"`. |
      | `count` | Yes | Occurrences to compare against. `EXACTLY` requires count ≥ 1; `LESS_THAN` requires count ≥ 2; `MORE_THAN` requires count ≥ 0. |
      
      Use `LESS_THAN 2` for "occurred exactly once when it should occur more". Occurrences
      are counted per `(trace, parent span, span name)` group — pin the `task`/`workflow`
      in `spanField` to the step you mean.
      
      ### SPAN_RELATION — relate two spans in a trace
      
      ```json
      {
        "type": "SPAN_RELATION",
        "predicate": {"name": "occurs_before"},
        "spanField": {
          "spanName": {"literal": "generate"},
          "task": {"literal": "generate_answer"},
          "workflow": {"literal": "RAG Agent"},
          "type": "SPAN_FIELD"
        },
        "relatedSpanFields": [
          {
            "spanName": {"literal": "retrieve"},
            "task": {"literal": "generate_answer"},
            "workflow": {"literal": "RAG Agent"},
            "type": "SPAN_FIELD"
          }
        ]
      }
      ```
      
      | Field | Required | Description |
      |-------|----------|-------------|
      | `type` | Yes | `"SPAN_RELATION"`. |
      | `predicate` | Yes | `{"name": "occurs_with" \| "occurs_before" \| "occurs_after"}`. Add `"negated": true` for the inverse (e.g. `occurs_with` + negated = "occurs without"). |
      | `spanField` | Yes | The primary span (see spanField below). |
      | `relatedSpanFields` | Yes | One or more related spans. Each must share the same coarser `workflow` / `task` as `spanField` — a span-name comparison needs matching `task` and `workflow`; a task-level comparison needs matching `workflow`. |
      
      ### spanField structure
      
      Identifies a span by workflow / task / span name. Fill in from coarse to fine.
      
      | Field | Required | Format |
      |-------|----------|--------|
      | `type` | No (defaults `"SPAN_FIELD"`) | `"SPAN_FIELD"` |
      | `workflow` | **Yes — always** | `{"literal": "workflow name"}` |
      | `task` | Yes when `spanName` is set | `{"literal": "task name"}` — requires `workflow` |
      | `spanName` | No | `{"literal": "span operation name"}` — requires `task` AND `workflow` |
      
      **`workflow` is the minimum.** If you set `spanName`, you MUST also set `task` and
      `workflow`. If you set `task`, you MUST also set `workflow`. All values use the
      `{"literal": "..."}` format. Discover the real names with `get_agent_segments` and
      `get_agent_trace`.
      
      ## Behavior playbooks
      
      Two trajectory-monitor patterns that apply to almost every agent. Both must be
      grounded in what THIS agent actually does — never propose them with stock values.
      
      ### Runaway loop — threshold derived from trace history
      
      Flags traces where the agent's dominant tool span repeats more times than any
      healthy run ever needed. Derive the threshold; never hardcode one:
      
      1. **Find the dominant tool span** — the span that does the agent's core work (the
         SQL execution tool for an analytics agent, retrieval for a RAG agent). Use
         `get_agent_traces` for per-trace shape and `get_agent_trace` on a few trace ids
         to see the span tree and which tool span dominates.
      2. **Build its per-trace occurrence distribution** from the sampled traces — e.g.
         "in 20 recent traces the SQL tool ran 1–3 times per trace; max observed: 3".
      3. **Set the threshold to max observed + headroom** — e.g. max 3 → `MORE_THAN`
         with `count: 5`. The headroom (roughly max + 2, or ~2× max for very tight
         distributions) keeps ordinary variance from alerting while still catching a loop.
      4. **Show the evidence** when proposing: the dominant span, the occurrence
         distribution, and the derived threshold with its headroom rationale.
      5. **Prove zero matches with a preview before creating.** Run `dry_run=True,
         preview=True` with the derived condition: the preview must report NOT breaching.
         If it reports breaching traces, your sample missed the heavy tail (long agentic
         sessions, multi-turn conversations accumulating in one trace) — re-derive from a
         wider window. Preview probes at increasing counts (e.g. more than 20/30/40 on a
         7-day `lookbackInHrs`) find the true historical max cheaply without pulling
         traces.
      
      A well-derived runaway-loop monitor matches **zero historical traces** — that is
      the point, not a defect. It is a regression guardrail: it stays silent until the
      agent's behavior actually regresses. At a design partner, exactly this monitor
      caught a silent-retry regression — a run that looped its dominant tool for ~3
      minutes with zero logged errors — within a week of being created, invisible to
      every error-based monitor.
      
      Create it live (`dry_run=False` after user confirmation), tagged
      `{"name": "agent", "value": "<AGENT_NAME>"}`, on a daily schedule
      (`interval_minutes=1440`, `lookbackInHrs: 24`).
      
      ### Ungrounded-in-data — create as a DRAFT with an evidence preview
      
      For agents that answer questions from data (analytics, RAG): flag traces where the
      agent produced an answer **without** executing its data-access tool — it likely
      answered from priors instead of the data. The shape is SPAN_RELATION `occurs_with`
      + `"negated": true` (the answer/LLM span occurs WITHOUT the data-tool span) —
      SPAN_OCCURRENCE cannot express "occurs 0 times".
      
      This naive pattern **also flags legitimate traffic**: generic questions ("what can
      you do?", "help") don't need a data query, so an active version alerts on healthy
      runs. Telling those apart needs an LLM-as-a-judge signal ("was this a data
      question?") combined with the trajectory condition, and that composition is not
      available yet. Therefore:
      
      - Run the evidence call first (`dry_run=True, preview=True`) and show the user
        what would currently fire and at what rate.
      - Create it as a **draft** (`dry_run=False, is_draft=True`) — same `agent` tag,
        same daily schedule — so the pattern is captured and reviewable without alerting
        on healthy runs.
      - Note the upgrade path: when trajectory monitors can be combined with an
        LLM-as-a-judge filter, add the "user asked a data question" judge and enable the
        monitor.
      - When later editing the monitor, keep passing `is_draft=True` — omitting it
        un-drafts.
      
      ## Examples
      
      The `agent` value below comes from `get_agent_metadata`'s `agentReference` field —
      a platform `{database}:{schema}.{name}` reference or an OTel `service_name`.
      
      ### Alert when a specific span occurs more than 5 times
      
      ```
      create_or_update_agent_trajectory_monitor(
          description="Alert when ChatBedrockConverse.chat exceeds 5 calls in a trace",
          agent="analytics:agents.support_bot",
          agent_span_alert_condition={
              "operator": "OR",
              "conditions": [
                  {
                      "type": "SPAN_OCCURRENCE",
                      "predicate": {"name": "occurs"},
                      "spanField": {
                          "spanName": {"literal": "ChatBedrockConverse.chat"},
                          "task": {"literal": "call_model"},
                          "workflow": {"literal": "Chat Agent"},
                          "type": "SPAN_FIELD"
                      },
                      "count": 5,
                      "comparisonOperator": "MORE_THAN"
                  }
              ]
          },
          time_filter={"timeField": {"field": "ingest_ts"}, "lookbackInHrs": 24},
          dry_run=True
      )
      ```
      
      ### Alert when a required step is missing (negated SPAN_RELATION)
      
      "occurs 0 times" can't be expressed with SPAN_OCCURRENCE. To catch a missing step,
      relate it to a step that always runs and negate the co-occurrence: alert when
      `generate` occurs **without** the `validate_output` step that should follow it.
      
      ```
      create_or_update_agent_trajectory_monitor(
          description="Alert when generation runs without a validation step",
          agent="checkout-agent",
          agent_span_alert_condition={
              "operator": "OR",
              "conditions": [
                  {
                      "type": "SPAN_RELATION",
                      "predicate": {"name": "occurs_with", "negated": true},
                      "spanField": {
                          "spanName": {"literal": "generate"},
                          "task": {"literal": "generate_answer"},
                          "workflow": {"literal": "Chat Agent"},
                          "type": "SPAN_FIELD"
                      },
                      "relatedSpanFields": [
                          {
                              "spanName": {"literal": "validate_output"},
                              "task": {"literal": "generate_answer"},
                              "workflow": {"literal": "Chat Agent"},
                              "type": "SPAN_FIELD"
                          }
                      ]
                  }
              ]
          },
          time_filter={"timeField": {"field": "ingest_ts"}, "lookbackInHrs": 24},
          dry_run=True
      )
      ```
      
      ### Alert when a step runs fewer than 2 times
      
      ```
      create_or_update_agent_trajectory_monitor(
          description="Alert when the safety_check step runs fewer than 2 times",
          agent="analytics:agents.support_bot",
          agent_span_alert_condition={
              "operator": "OR",
              "conditions": [
                  {
                      "type": "SPAN_OCCURRENCE",
                      "predicate": {"name": "occurs"},
                      "spanField": {
                          "spanName": {"literal": "safety_check"},
                          "task": {"literal": "validate_output"},
                          "workflow": {"literal": "Chat Agent"},
                          "type": "SPAN_FIELD"
                      },
                      "count": 2,
                      "comparisonOperator": "LESS_THAN"
                  }
              ]
          },
          time_filter={"timeField": {"field": "ingest_ts"}, "lookbackInHrs": 24},
          dry_run=True
      )
      ```
      
      ## Common errors
      
      | Error message | Cause | Fix |
      |--------------|-------|-----|
      | invalid / unresolvable `agent` reference | The `agent` value wasn't taken from `get_agent_metadata` | Use the exact `agentReference` value — do not construct it by hand, and never pass an MCON |
      | "workflow should not be set" in agentSpanFilters | Trajectory monitors only allow the `agent` field in `agent_span_filters` | Remove `workflow`/`task`/`spanName`; put them in the condition's `spanField` |
      | `AND` operator rejected | `agent_span_alert_condition.operator` set to `"AND"` | Use `"OR"` (or omit it); split an all-must-hold rule into separate monitors |
      | `EXACTLY 0` / `LESS_THAN 1` rejected | Tried to express "occurs 0 times" | Use a negated SPAN_RELATION for a missing span, or `LESS_THAN 2` for "occurred only once" |
      | spanField rejected | Set `spanName` without `task`+`workflow`, or `task` without `workflow` | Fill coarse-to-fine: `workflow` always; `task` needs `workflow`; `spanName` needs `task`+`workflow` |
      
    • agent-validation-monitor.md 11 KB
      # Agent Validation Monitor
      
      ## When to use
      
      Assert a logical condition on agent span data and alert on violations. Best for:
      
      - **Business rule assertions** — "total_tokens must be below 10000"
      - **Data quality checks on span attributes** — "completions must never be null"
      - **Compliance checks** — "PII detection must run on every trace"
      
      Do NOT use it to trend a numeric metric (use `create_or_update_agent_metric_monitor`)
      or to alert on span sequences / call counts (use
      `create_or_update_agent_trajectory_monitor`).
      
      ## Constraints
      
      > **CRITICAL:** The monitor's source is the `agent` reference. Pass the
      > `agentReference` value from `get_agent_metadata` verbatim — a platform
      > `{database}:{schema}.{name}` reference or an OTel `service_name`. Never modify,
      > truncate, or reconstruct it, and never pass an MCON.
      
      > **CRITICAL:** `warehouse` is REQUIRED. Pass the agent's `warehouse_uuid` from
      > `get_agent_metadata`; use `get_warehouses` when it is null or to resolve by name.
      
      > **CRITICAL:** `alert_condition` matches the rows to ALERT on. A `null` predicate
      > alerts on spans where the field IS null. Express negation with the `negated` flag —
      > there is NO `not_equal` and NO `not_null` predicate.
      
      > **IMPORTANT:** BINARY conditions use `left`/`right`; UNARY conditions use `value`
      > (NOT `left`). Getting this wrong is the most common failure.
      
      > **IMPORTANT:** `time_filter` is REQUIRED and `timeField` is always
      > `{"field": "ingest_ts"}`. `time_filter` is `{"timeField": {"field": "ingest_ts"}, "lookbackInHrs": <hours>}`.
      
      > **IMPORTANT:** `schedule_type` is `fixed` (default) or `manual` — never dynamic.
      > `interval_minutes` defaults to `60` and must be at least 5 (sub-hourly is allowed;
      > no 60-minute alignment).
      
      ## Key characteristics
      
      - Uses `alert_condition` as a `FilterGroup` — an `operator` (`AND`/`OR`) plus a
        `conditions` array of `BINARY` / `UNARY` / `SQL` / `GROUP` entries
      - Requires `time_filter` (time field is always `ingest_ts`)
      - Optional `is_agent_trace_aggregation` for trace-level assertions (OTel agents only)
      
      ## Parameters
      
      | Parameter | Type | Required | Description |
      |-----------|------|----------|-------------|
      | `description` | string | Yes | Human-readable monitor description (shown as display name) |
      | `agent` | string | Yes | Agent reference — `agentReference` from `get_agent_metadata` (`{db}:{schema}.{name}` or OTel `service_name`) |
      | `alert_condition` | object | Yes | FilterGroup — the condition that marks INVALID rows (see below) |
      | `time_filter` | object | Yes | `{"timeField": {"field": "ingest_ts"}, "lookbackInHrs": 24}` |
      | `warehouse` | string | Yes | Warehouse name or UUID where the agent's traces live |
      | `trace_table` | string | No | Explicit trace table — only for non-ClickHouse OTel agents |
      | `agent_span_filters` | array | No | Optional span-scope refinement; at most ONE filter object |
      | `is_agent_trace_aggregation` | boolean | No | Aggregate per trace for trace-level assertions (OTel only) |
      | `schedule_type` | string | No | `fixed` (default) or `manual` |
      | `interval_minutes` | int | No | Default `60`; at least 5 |
      | `tags` | array | No | Key-value tags, e.g. `[{"name": "agent", "value": "Support Bot"}]`. Tag every agent monitor with its agent's name — `{"name": "agent", "value": "<AGENT_NAME>"}` — so all of one agent's monitors are filterable as a group |
      | `domain_uuids` | array | No | Domain UUIDs to assign this monitor to — the agent-onboarding playbook passes the footprint's single resolved domain on every create (see agent-monitor-creation.md conventions) |
      | `monitor_uuid` | string | No | UUID of an existing monitor to update in place (PUT semantics) |
      | `dry_run` | boolean | No | Default `True` — preview YAML; set `False` to deploy |
      
      ## alert_condition structure (FilterGroup)
      
      The top level is a group: an `operator` (`AND`/`OR`) and a `conditions` array. Each
      condition is one of `BINARY`, `UNARY`, `SQL`, or `GROUP`.
      
      ### BINARY (compare two values)
      
      ```json
      {
        "type": "BINARY",
        "predicate": {"name": "greater_than"},
        "left": [{"type": "FIELD", "field": "total_tokens"}],
        "right": [{"type": "LITERAL", "literal": "10000"}]
      }
      ```
      
      - `left`: exactly one `FIELD` value (the column being validated).
      - `right`: exactly one value — usually a `LITERAL` (a string, even for numbers). The
        `in_set` predicate is the exception: it takes several `LITERAL`s in `right`.
      
      ### UNARY (single-value check)
      
      ```json
      {
        "type": "UNARY",
        "predicate": {"name": "null"},
        "value": [{"type": "FIELD", "field": "completions"}]
      }
      ```
      
      - The field list is named **`value`** (NOT `left`), and holds exactly one `FIELD`.
      - The example above matches spans where `completions` **is null**. To alert on
        non-null instead, add `"negated": true` (→ IS NOT NULL).
      
      ### SQL (custom boolean expression)
      
      ```json
      {"type": "SQL", "sql": "total_tokens > 1000 AND duration_sec < 60"}
      ```
      
      Use only when the condition can't be expressed with a predicate.
      
      ### GROUP (nested conditions)
      
      ```json
      {
        "type": "GROUP",
        "operator": "OR",
        "conditions": [
          {"type": "BINARY", "...": "..."},
          {"type": "UNARY", "...": "..."}
        ]
      }
      ```
      
      ### Predicates
      
      Predicate names are matched by exact name — call `get_validation_predicates` to list
      the full set. Key rules:
      
      - **Express negation with the `negated` flag** — e.g. `{"name": "equal", "negated": true}`
        or `{"name": "null", "negated": true}`. Do NOT prefix names with `not_`. There is
        **no `not_equal` and no `not_null` predicate**.
      - **BINARY predicates** include `equal`, `in_set`, `greater_than`,
        `greater_than_or_equal`, `less_than`, `less_than_or_equal`, `contains`,
        `starts_with`, `ends_with`, `matches_regex`. The four comparators (`greater_than` /
        `less_than` / `*_or_equal`) cannot be negated — use the inverse comparator instead.
      - **UNARY predicates** include `null`, `empty_string`, `is_zero`, `is_negative`,
        `is_nan`, `is_between_0_and_1`, `is_between_0_and_100`, `is_uuid`, plus many locale /
        PII / timestamp checks. Fetch the full list with `get_validation_predicates`.
      
      ### Value types
      
      | Type | Format |
      |------|--------|
      | `FIELD` | `{"type": "FIELD", "field": "column_name"}` — references a span field |
      | `LITERAL` | `{"type": "LITERAL", "literal": "value_string"}` — a static value, always a string even for numbers |
      | `SQL` | `{"type": "SQL", "sql": "..."}` — a SQL expression (used inside a BINARY `right`) |
      
      Condition/value/operator keywords are uppercase: `BINARY`, `UNARY`, `SQL`, `GROUP`;
      `FIELD`, `LITERAL`; `AND`, `OR`.
      
      ## Field notes
      
      Use the span field names from `agent-span-fields.md` in `FIELD` values — not raw
      table columns. Note in particular:
      
      - `status_code` is **numeric** — the OTel span status code (`2` = error). Compare
        with a numeric literal (e.g. `equal` / `"2"`), not `"ERROR"`.
      - The time field is always `ingest_ts`.
      
      ## Trace aggregation
      
      `is_agent_trace_aggregation=True` is **OpenTelemetry-only.** A platform agent reference
      (`{database}:{schema}.{name}`) is rejected — target an OTel `service_name`, or pass an
      explicit `trace_table` to force the agent to be read as OpenTelemetry. At trace grain,
      use trace-level fields (`span_count`, `llm_call_count`, `total_tokens`, …) and filter
      only by `agent`.
      
      ## Examples
      
      The `agent` value below comes from `get_agent_metadata`'s `agentReference` field —
      a platform `{database}:{schema}.{name}` reference or an OTel `service_name`.
      
      ### Assert total_tokens stays below a threshold (platform agent reference)
      
      ```
      create_or_update_agent_validation_monitor(
          description="Alert when total_tokens exceeds 10000",
          agent="analytics:agents.support_bot",
          warehouse="Analytics WH",
          tags=[{"name": "agent", "value": "Support Bot"}],
          alert_condition={
              "operator": "AND",
              "conditions": [
                  {
                      "type": "BINARY",
                      "predicate": {"name": "greater_than"},
                      "left": [{"type": "FIELD", "field": "total_tokens"}],
                      "right": [{"type": "LITERAL", "literal": "10000"}]
                  }
              ]
          },
          time_filter={"timeField": {"field": "ingest_ts"}, "lookbackInHrs": 24},
          dry_run=True
      )
      ```
      
      ### Alert when a required field is null (UNARY uses `value`)
      
      Business rule "completions must always be populated" → alert on the violating rows,
      i.e. spans where `completions` **is null**, so the condition is a plain `null`
      predicate.
      
      ```
      create_or_update_agent_validation_monitor(
          description="Alert when the completions field is null",
          agent="analytics:agents.support_bot",
          warehouse="Analytics WH",
          alert_condition={
              "operator": "AND",
              "conditions": [
                  {
                      "type": "UNARY",
                      "predicate": {"name": "null"},
                      "value": [{"type": "FIELD", "field": "completions"}]
                  }
              ]
          },
          time_filter={"timeField": {"field": "ingest_ts"}, "lookbackInHrs": 24},
          dry_run=True
      )
      ```
      
      ### Span-level assertion scoped to a workflow
      
      ```
      create_or_update_agent_validation_monitor(
          description="Alert when Chat Agent spans exceed 120s",
          agent="analytics:agents.support_bot",
          warehouse="Analytics WH",
          alert_condition={
              "operator": "AND",
              "conditions": [
                  {
                      "type": "BINARY",
                      "predicate": {"name": "greater_than"},
                      "left": [{"type": "FIELD", "field": "duration_sec"}],
                      "right": [{"type": "LITERAL", "literal": "120"}]
                  }
              ]
          },
          time_filter={"timeField": {"field": "ingest_ts"}, "lookbackInHrs": 24},
          agent_span_filters=[{"workflow": {"value": "Chat Agent"}}],
          dry_run=True
      )
      ```
      
      ### Trace-level assertion (OTel agent)
      
      ```
      create_or_update_agent_validation_monitor(
          description="Alert when a trace has more than 50 spans",
          agent="checkout-agent",
          warehouse="OTel WH",
          alert_condition={
              "operator": "AND",
              "conditions": [
                  {
                      "type": "BINARY",
                      "predicate": {"name": "greater_than"},
                      "left": [{"type": "FIELD", "field": "span_count"}],
                      "right": [{"type": "LITERAL", "literal": "50"}]
                  }
              ]
          },
          time_filter={"timeField": {"field": "ingest_ts"}, "lookbackInHrs": 24},
          is_agent_trace_aggregation=True,
          dry_run=True
      )
      ```
      
      ## Common errors
      
      | Error message | Cause | Fix |
      |--------------|-------|-----|
      | Warehouse not found | `warehouse` omitted or wrong | Pass the agent's `warehouse_uuid` from `get_agent_metadata`; if null, list warehouses via `get_warehouses` |
      | invalid / unresolvable `agent` reference | The `agent` value wasn't taken from `get_agent_metadata` | Use the exact `agentReference` value — do not construct it by hand, and never pass an MCON |
      | unknown predicate `not_equal` / `not_null` | Used a `not_`-prefixed predicate | Use the base predicate (`equal` / `null`) with `"negated": true` |
      | UNARY condition rejected | Used `left` instead of `value` | UNARY conditions put the field in `value`; only BINARY uses `left`/`right` |
      | "Field X doesn't exist" | Field name not in the PARSED_SPANS schema | Check `agent-span-fields.md`; `status_code` is numeric (compare with `"2"`, not `"ERROR"`) |
      
    • data-comparison-monitor.md 19.1 KB
      # Comparison Monitor Reference
      
      Detailed reference for building `create_or_update_comparison_monitor` tool calls. The tool follows the **two-call preview-then-confirm pattern** — see `data-monitor-creation.md` for the full flow.
      
      ## Critical Constraints
      
      - **NEVER guess column names.** Always get them from `get_table` for both the source and target tables. Verify that `sourceField` exists in the source table and `targetField` exists in the target table before building alert conditions.
      
      ---
      
      ## When to Use
      
      Use a comparison monitor when the user wants to:
      
      - Compare data between two tables (e.g., source vs target, dev vs prod)
      - Validate data consistency after migration or replication
      - Check row count parity across environments
      - Compare field-level metrics between tables (null counts, sums, distributions)
      
      ---
      
      ## Pre-Step: Verify Both Tables and Fields
      
      Before constructing alert conditions, you MUST verify that both tables exist and that any referenced fields are real columns. This is the most common source of comparison monitor failures.
      
      1. **Resolve both MCONs.** Use `search` to find the source and target tables. If the user provided `database:schema.table` format, search for each to get the MCON.
      2. **Get full schemas.** Call `get_table` with `include_fields: true` on BOTH the source table and the target table. You need the column lists from both.
      3. **For field-level metrics, verify fields exist on both sides.** Confirm that `sourceField` exists in the source table's column list AND `targetField` exists in the target table's column list. Field names are case-sensitive on most warehouses.
      4. **Check field type compatibility.** The metric must be compatible with the column types on both sides. For example, `NUMERIC_MEAN` requires numeric columns in both the source and target tables. If the source column is numeric but the target is a string, the comparison will fail.
      5. If any field does not exist or types are incompatible, stop and ask the user to clarify. Do not guess.
      
      ---
      
      ## Required Parameters
      
      | Parameter | Type | Description |
      |-----------|------|-------------|
      | `name` | string | Unique identifier for the monitor. Use a descriptive slug (e.g., `orders_dev_prod_compare`). |
      | `description` | string | Human-readable description of what the monitor checks. |
      | `source_table` | string | Source table MCON (preferred) or `database:schema.table` format. If not MCON, also pass `source_warehouse`. |
      | `target_table` | string | Target table MCON (preferred) or `database:schema.table` format. If not MCON, also pass `target_warehouse`. |
      | `alert_conditions` | array | List of comparison conditions (see Alert Conditions below). |
      
      ## Optional Parameters
      
      | Parameter | Type | Description |
      |-----------|------|-------------|
      | `source_warehouse` | string | Warehouse name or UUID for the source table. Required if `source_table` is not an MCON. |
      | `target_warehouse` | string | Warehouse name or UUID for the target table. Required if `target_table` is not an MCON. |
      | `segment_fields` | array of string | Fields to segment the comparison by. Must exist in BOTH tables with the same name. |
      | `domain_uuids` | array of string (uuid) | Domain UUIDs (use `get_domains` to list). Data monitors accept exactly one UUID in the list. |
      | `schedule_type` | string | Schedule type: `"fixed"` (default), `"dynamic"`, `"manual"`. |
      | `interval_minutes` | int | Schedule interval in minutes (only for `schedule_type="fixed"`). |
      | `audiences` | array of string | Notification audience **names** (not UUIDs) to alert when the monitor triggers. |
      | `failure_audiences` | array of string | Notification audience names to alert on query execution failures. |
      | `notes` | string | Free-text notes shown in the UI (separate from `description`). |
      | `priority` | string | Monitor priority (e.g. `"P1"`, `"P2"`). |
      | `tags` | array of `{name, value}` | Key-value tags to attach. |
      | `is_draft` | bool | When `True`, saves the monitor as a draft (not active). Default `False`. |
      | `monitor_uuid` | string (uuid) | UUID of an existing monitor to update in place. Omit to create a new monitor. **PUT semantics:** the call fully replaces the monitor's configuration — fields you omit revert to tool defaults, they are NOT left untouched. Before editing, read the current config with `get_monitors(monitor_ids=[<uuid>], include_fields=["config"])` and re-pass every field you want to keep. See `data-monitor-creation.md` (Step 7) for the safe-edit workflow. |
      | `dry_run` | bool | Default `True`. Preview mode. When omitted or `True`, returns YAML preview in `result.yaml`. When `False`, actually creates/updates the monitor and returns `result.monitor_uuid` + a deep link in `result.instructions`. See `data-monitor-creation.md`. |
      
      ---
      
      ## Cross-Warehouse Comparisons
      
      When the source and target tables live in different warehouses (e.g., comparing a Snowflake staging table against a BigQuery production table), you MUST provide both `source_warehouse` and `target_warehouse` explicitly. The tool cannot auto-resolve warehouses when tables are in different environments.
      
      Even when both tables are MCONs, if they belong to different warehouses, pass both warehouse parameters to be safe. Omitting them in cross-warehouse scenarios causes silent failures or incorrect results.
      
      Common cross-warehouse patterns:
      - **Dev vs prod:** same warehouse type, different databases or schemas
      - **Migration validation:** source in old warehouse, target in new warehouse
      - **Replication checks:** primary warehouse vs replica or downstream warehouse
      
      ---
      
      ## Alert Conditions
      
      Each condition compares a metric between the source and target tables.
      
      | Field | Type | Required | Description |
      |-------|------|----------|-------------|
      | `metric` | string | Yes | The metric to compare (see Metrics Reference below). |
      | `type` | string | Yes (for non-AUTO thresholds) | Threshold type — one of `comparison_delta` (static threshold on source↔target diff) or `AUTO` (anomaly detection). Omitting `type` defaults to AUTO-style behavior. |
      | `sourceField` | string | For field-level metrics | Column in the source table. Required for ALL metrics except `ROW_COUNT`. |
      | `targetField` | string | For field-level metrics | Column in the target table. Required for ALL metrics except `ROW_COUNT`. |
      | `thresholdValue` | number | **Required for `comparison_delta`-type conditions**; optional for `AUTO`-type anomaly detection. | Threshold for acceptable difference between source and target. Omitting it on a delta-type condition is rejected with `threshold_value is required for comparison_delta type`. |
      | `isThresholdRelative` | boolean | No | `false` = absolute difference (default), `true` = percentage difference. |
      | `customMetric` | object | No | Custom SQL expressions for source and target (see Custom Metrics below). |
      
      ### Threshold types
      
      Two threshold types are supported on comparison alert conditions:
      
      | `type` | Behavior | Required fields |
      |---|---|---|
      | `AUTO` (default when `type` is omitted) | Monte Carlo learns normal variance and alerts on anomalies. | `metric`, `sourceField` / `targetField` (unless `ROW_COUNT`) |
      | `comparison_delta` | Static threshold on the source↔target difference. | `metric`, `sourceField` / `targetField` (unless `ROW_COUNT`), `thresholdValue` |
      
      If the user wants a specific numeric tolerance (e.g. "alert if source and target row counts differ by more than 100"), use `comparison_delta` and set `thresholdValue`. If they want "alert when the difference looks unusual," use `AUTO` — no `thresholdValue` needed.
      
      ---
      
      ## ROW_COUNT and Fields: A Critical Rule
      
      > **NEVER pass `sourceField` or `targetField` when using the `ROW_COUNT` metric.**
      
      `ROW_COUNT` is a table-level metric -- it counts all rows in the table, not values in a column. Passing field names with `ROW_COUNT` causes the API call to fail or produce unexpected behavior.
      
      This is the single most common mistake with comparison monitors. Before submitting any alert condition with `ROW_COUNT`, verify that `sourceField` and `targetField` are both absent from the condition object.
      
      | Metric | Fields needed? | What happens if you pass fields? |
      |--------|---------------|----------------------------------|
      | `ROW_COUNT` | **No -- NEVER pass fields** | API error or undefined behavior |
      | All other metrics | **Yes -- always pass both fields** | Required for the comparison to work |
      
      ---
      
      ## Metrics Reference
      
      ### Table-level metric (no fields needed)
      
      | Metric | Description |
      |--------|-------------|
      | `ROW_COUNT` | Compare total row counts between source and target. |
      
      ### Field-level metrics (require `sourceField` and `targetField`)
      
      #### Uniqueness and duplicates
      
      | Metric | Description |
      |--------|-------------|
      | `UNIQUE_COUNT` | Count of distinct values. |
      | `DUPLICATE_COUNT` | Count of duplicate (non-unique) values. |
      | `APPROX_DISTINCT_COUNT` | Approximate distinct count (faster on large tables). |
      
      #### Null and empty checks
      
      | Metric | Description |
      |--------|-------------|
      | `NULL_COUNT` | Count of null values. |
      | `NON_NULL_COUNT` | Count of non-null values. |
      | `EMPTY_STRING_COUNT` | Count of empty string values. |
      | `TEXT_ALL_SPACES_COUNT` | Count of values that are all whitespace. |
      | `NAN_COUNT` | Count of NaN values. |
      | `TEXT_NULL_KEYWORD_COUNT` | Count of values containing null-like keywords (e.g., "NULL", "None"). |
      
      #### Numeric statistics
      
      | Metric | Description |
      |--------|-------------|
      | `NUMERIC_MEAN` | Mean of numeric field. |
      | `NUMERIC_MEDIAN` | Median of numeric field. |
      | `NUMERIC_MIN` | Minimum value. |
      | `NUMERIC_MAX` | Maximum value. |
      | `NUMERIC_STDDEV` | Standard deviation. |
      | `SUM` | Sum of numeric field. |
      | `ZERO_COUNT` | Count of zero values. |
      | `NEGATIVE_COUNT` | Count of negative values. |
      
      #### Percentiles
      
      | Metric | Description |
      |--------|-------------|
      | `PERCENTILE_20` | 20th percentile value. |
      | `PERCENTILE_40` | 40th percentile value. |
      | `PERCENTILE_60` | 60th percentile value. |
      | `PERCENTILE_80` | 80th percentile value. |
      
      #### Text statistics
      
      | Metric | Description |
      |--------|-------------|
      | `TEXT_MAX_LENGTH` | Maximum string length. |
      | `TEXT_MIN_LENGTH` | Minimum string length. |
      | `TEXT_MEAN_LENGTH` | Mean string length. |
      | `TEXT_STD_LENGTH` | Standard deviation of string length. |
      
      #### Text format checks
      
      | Metric | Description |
      |--------|-------------|
      | `TEXT_NOT_INT_COUNT` | Count of values not parseable as integers. |
      | `TEXT_NOT_NUMBER_COUNT` | Count of values not parseable as numbers. |
      | `TEXT_NOT_UUID_COUNT` | Count of values not matching UUID format. |
      | `TEXT_NOT_SSN_COUNT` | Count of values not matching SSN format. |
      | `TEXT_NOT_US_PHONE_COUNT` | Count of values not matching US phone format. |
      | `TEXT_NOT_US_STATE_CODE_COUNT` | Count of values not matching US state codes. |
      | `TEXT_NOT_US_ZIP_CODE_COUNT` | Count of values not matching US zip codes. |
      | `TEXT_NOT_EMAIL_ADDRESS_COUNT` | Count of values not matching email format. |
      | `TEXT_NOT_TIMESTAMP_COUNT` | Count of values not parseable as timestamps. |
      
      #### Boolean
      
      | Metric | Description |
      |--------|-------------|
      | `TRUE_COUNT` | Count of true values. |
      | `FALSE_COUNT` | Count of false values. |
      
      #### Timestamp
      
      | Metric | Description |
      |--------|-------------|
      | `FUTURE_TIMESTAMP_COUNT` | Count of timestamps in the future. |
      | `PAST_TIMESTAMP_COUNT` | Count of timestamps unreasonably far in the past. |
      | `UNIX_ZERO_COUNT` | Count of timestamps equal to Unix epoch zero (1970-01-01). |
      
      ---
      
      ## Choosing the Right Metric
      
      | User intent | Correct metric | Fields needed? |
      |-------------|---------------|----------------|
      | Row count parity | `ROW_COUNT` | **No** -- never pass fields |
      | Distinct values in a column | `UNIQUE_COUNT` | Yes |
      | Null values in a column | `NULL_COUNT` | Yes |
      | Sum, average, min, max | `SUM`, `NUMERIC_MEAN`, `NUMERIC_MIN`, `NUMERIC_MAX` | Yes |
      | Data completeness | `NON_NULL_COUNT` | Yes |
      | String format validation | `TEXT_NOT_EMAIL_ADDRESS_COUNT`, `TEXT_NOT_UUID_COUNT`, etc. | Yes |
      | Custom computed expressions | Use `customMetric` instead of `metric` | No (SQL handles it) |
      
      ---
      
      ## Custom Metrics
      
      Use custom metrics when:
      
      - **Column names differ** between source and target and you need a computed expression (not just a direct field comparison).
      - **You need a derived calculation** like `SUM(quantity * unit_price)` rather than a simple column metric.
      - **Standard metrics do not cover the comparison** (e.g., comparing a ratio, a conditional aggregate, or a windowed calculation).
      
      If the columns simply have different names but you want a standard metric (e.g., compare `SUM` of `revenue` in source vs `total_revenue` in target), you do NOT need a custom metric -- just use the standard metric with different `sourceField` and `targetField` values.
      
      Custom metric structure:
      
      ```json
      {
        "customMetric": {
          "displayName": "Revenue Sum",
          "sourceSqlExpression": "SUM(revenue)",
          "targetSqlExpression": "SUM(total_revenue)"
        }
      }
      ```
      
      | Field | Type | Required | Description |
      |-------|------|----------|-------------|
      | `displayName` | string | Yes | Human-readable name for the metric in alerts and dashboards. |
      | `sourceSqlExpression` | string | Yes | SQL expression evaluated against the source table. |
      | `targetSqlExpression` | string | Yes | SQL expression evaluated against the target table. |
      
      When using `customMetric`, do NOT also pass `metric`, `sourceField`, or `targetField` in the same alert condition. The custom metric replaces all of those.
      
      ---
      
      ## Threshold Guidance
      
      ### Absolute thresholds (`isThresholdRelative: false` or omitted)
      
      The `thresholdValue` is the maximum acceptable absolute difference between the source and target metric values.
      
      - `thresholdValue: 0` -- source and target must match exactly.
      - `thresholdValue: 100` -- up to 100 units of difference is acceptable.
      
      ### Relative (percentage) thresholds (`isThresholdRelative: true`)
      
      The `thresholdValue` is the maximum acceptable percentage difference.
      
      - `thresholdValue: 5` -- up to 5% difference is acceptable.
      - `thresholdValue: 0.1` -- up to 0.1% difference is acceptable.
      
      ### When to use each
      
      | Scenario | Recommended threshold type |
      |----------|---------------------------|
      | Exact replication (row counts must match) | Absolute, `thresholdValue: 0` |
      | Near-real-time sync with small lag | Absolute, small value (e.g., 10-100) |
      | Tables at different scales | Relative, percentage-based |
      | Aggregated metrics (sums, means) | Relative, to handle floating-point differences |
      
      ---
      
      ## Examples
      
      ### Row count parity with absolute threshold
      
      Compare row counts between dev and prod, alerting if they differ by more than 100 rows.
      
      ```json
      {
        "name": "orders_dev_prod_row_count",
        "description": "Verify dev and prod orders tables have similar row counts",
        "source_table": "MCON++a1b2c3d4-e5f6-7890-abcd-ef1234567890++1++1++dev_warehouse:core.orders",
        "target_table": "MCON++b2c3d4e5-f6a7-8901-bcde-f12345678901++1++1++prod_warehouse:core.orders",
        "alert_conditions": [
          {
            "metric": "ROW_COUNT",
            "thresholdValue": 100,
            "isThresholdRelative": false
          }
        ]
      }
      ```
      
      Note: no `sourceField` or `targetField` -- `ROW_COUNT` is table-level.
      
      ### Row count parity with percentage threshold
      
      Alert if row counts differ by more than 5%.
      
      ```json
      {
        "name": "orders_replication_check",
        "description": "Verify replicated orders table is within 5% of source row count",
        "source_table": "MCON++a1b2c3d4-e5f6-7890-abcd-ef1234567890++1++1++primary:sales.orders",
        "target_table": "MCON++b2c3d4e5-f6a7-8901-bcde-f12345678901++1++1++replica:sales.orders",
        "alert_conditions": [
          {
            "metric": "ROW_COUNT",
            "thresholdValue": 5,
            "isThresholdRelative": true
          }
        ]
      }
      ```
      
      ### Field-level comparison (different column names)
      
      Compare the sum of `revenue` in the source table against `total_revenue` in the target table.
      
      ```json
      {
        "name": "revenue_source_target_sum",
        "description": "Verify revenue sums match between staging and production",
        "source_table": "MCON++a1b2c3d4-e5f6-7890-abcd-ef1234567890++1++1++staging:finance.transactions",
        "target_table": "MCON++b2c3d4e5-f6a7-8901-bcde-f12345678901++1++1++production:finance.transactions",
        "alert_conditions": [
          {
            "metric": "SUM",
            "sourceField": "revenue",
            "targetField": "total_revenue",
            "thresholdValue": 1,
            "isThresholdRelative": true
          }
        ]
      }
      ```
      
      ### Segmented comparison
      
      Compare null counts on `email` between source and target, segmented by `country`. The `country` field must exist in both tables.
      
      ```json
      {
        "name": "email_nulls_by_country",
        "description": "Compare email null counts by country between ETL source and target",
        "source_table": "MCON++a1b2c3d4-e5f6-7890-abcd-ef1234567890++1++1++raw:crm.contacts",
        "target_table": "MCON++b2c3d4e5-f6a7-8901-bcde-f12345678901++1++1++analytics:crm.contacts",
        "segment_fields": ["country"],
        "alert_conditions": [
          {
            "metric": "NULL_COUNT",
            "sourceField": "email",
            "targetField": "email",
            "thresholdValue": 0,
            "isThresholdRelative": false
          }
        ]
      }
      ```
      
      ### Cross-warehouse comparison with explicit warehouses
      
      When source and target are in different warehouses, both warehouse parameters must be provided.
      
      ```json
      {
        "name": "migration_users_row_count",
        "description": "Validate user row counts match after Snowflake to BigQuery migration",
        "source_table": "snowflake_db:public.users",
        "source_warehouse": "snowflake-prod",
        "target_table": "bigquery_project:public.users",
        "target_warehouse": "bigquery-prod",
        "alert_conditions": [
          {
            "metric": "ROW_COUNT",
            "thresholdValue": 0,
            "isThresholdRelative": false
          }
        ]
      }
      ```
      
      ### Custom metric comparison
      
      Compare a computed revenue expression when the SQL differs between source and target.
      
      ```json
      {
        "name": "computed_revenue_compare",
        "description": "Compare total revenue computation between legacy and new schema",
        "source_table": "MCON++a1b2c3d4-e5f6-7890-abcd-ef1234567890++1++1++warehouse:legacy.orders",
        "target_table": "MCON++b2c3d4e5-f6a7-8901-bcde-f12345678901++1++1++warehouse:v2.orders",
        "alert_conditions": [
          {
            "customMetric": {
              "displayName": "Total Revenue",
              "sourceSqlExpression": "SUM(quantity * unit_price)",
              "targetSqlExpression": "SUM(total_amount)"
            },
            "thresholdValue": 0.01,
            "isThresholdRelative": true
          }
        ]
      }
      ```
      
      ### Multiple alert conditions
      
      Compare both row counts and field-level metrics in a single monitor.
      
      ```json
      {
        "name": "orders_full_comparison",
        "description": "Full comparison of orders between staging and production",
        "source_table": "MCON++a1b2c3d4-e5f6-7890-abcd-ef1234567890++1++1++staging:core.orders",
        "target_table": "MCON++b2c3d4e5-f6a7-8901-bcde-f12345678901++1++1++production:core.orders",
        "domain_uuids": ["f47ac10b-58cc-4372-a567-0e02b2c3d479"],
        "alert_conditions": [
          {
            "metric": "ROW_COUNT",
            "thresholdValue": 0,
            "isThresholdRelative": false
          },
          {
            "metric": "NULL_COUNT",
            "sourceField": "customer_id",
            "targetField": "customer_id",
            "thresholdValue": 0,
            "isThresholdRelative": false
          },
          {
            "metric": "SUM",
            "sourceField": "amount",
            "targetField": "amount",
            "thresholdValue": 0.1,
            "isThresholdRelative": true
          }
        ]
      }
      ```
      
      Note: the `ROW_COUNT` condition has no fields, while the field-level conditions each specify both `sourceField` and `targetField`.
      
    • data-custom-sql-monitor.md 12.2 KB
      # Custom SQL Monitor Reference
      
      Detailed reference for building `create_or_update_sql_monitor` tool calls. The tool follows the **two-call preview-then-confirm pattern** — see `data-monitor-creation.md` for the full flow.
      
      ## Critical Constraints
      
      - **NEVER guess column names.** Always verify column names from `get_table` before referencing them in SQL queries. A typo or assumed column name causes the monitor to fail on every scheduled run.
      
      ---
      
      ## When to Use
      
      Use a custom SQL monitor when the user wants to:
      
      - Run a specific SQL query and alert on its result
      - Implement cross-table logic (joins, subqueries, CTEs)
      - Apply business-specific aggregations or calculations that don't map to a single metric
      - Monitor a condition that spans multiple columns or tables
      - Use a SQL query they already have in mind
      
      ---
      
      ## The Universal Fallback
      
      Custom SQL is the fallback monitor type. Reach for it whenever another monitor type cannot express what the user needs:
      
      - **Validation monitor won't work** because the column doesn't exist yet, or the logic requires joins across tables.
      - **Metric monitor can't express the business logic** -- for example, a ratio between two columns, a conditional aggregation, or a calculation that spans multiple tables.
      - **Cross-table joins are needed** -- metric and validation monitors operate on a single table. If the check requires data from two or more tables, custom SQL is the only option.
      - **The user already has a SQL query** -- don't force it into another monitor type. Wrap it in a custom SQL monitor.
      
      If you find yourself contorting another monitor type to fit the user's intent, stop and use custom SQL instead.
      
      ---
      
      ## Required Parameters
      
      | Parameter | Type | Description |
      |-----------|------|-------------|
      | `name` | string | Unique identifier for the monitor. Use a descriptive slug (e.g., `orphan_orders_check`). |
      | `description` | string | Human-readable description of what the monitor checks. |
      | `warehouse` | string | Warehouse name or UUID where the SQL query will be executed. |
      | `sql` | string | SQL query that returns a **single numeric value** (one row, one column). |
      | `alert_condition` | object | When the monitor should fire (see Alert Conditions below). Singular — the tool takes exactly one condition object, not an array. |
      
      ## Optional Parameters
      
      | Parameter | Type | Description |
      |-----------|------|-------------|
      | `domain_uuids` | array of string (uuid) | Domain UUIDs (use `get_domains` to list). Data monitors accept exactly one UUID in the list. |
      | `query_result_type` | string | What the SQL query returns — see the tool's enum for accepted values (single numeric is the most common). |
      | `custom_sampling_sql` | string | Optional SQL used to sample rows that contributed to the result (shown in alert detail). |
      | `variable_definitions` | object | Named variables that can be referenced in `sql` / `custom_sampling_sql`. |
      | `schedule_type` | string | Schedule type: `"fixed"` (default), `"dynamic"`, `"manual"`. |
      | `interval_minutes` | int | Schedule interval in minutes (only for `schedule_type="fixed"`). |
      | `audiences` | array of string | Notification audience **names** (not UUIDs) to alert when the monitor triggers. |
      | `failure_audiences` | array of string | Notification audience names to alert on query execution failures. |
      | `notes` | string | Free-text notes shown in the UI (separate from `description`). |
      | `priority` | string | Monitor priority (e.g. `"P1"`, `"P2"`). |
      | `tags` | array of `{name, value}` | Key-value tags to attach. |
      | `is_draft` | bool | When `True`, saves the monitor as a draft (not active). Default `False`. |
      | `monitor_uuid` | string (uuid) | UUID of an existing monitor to update in place. Omit to create a new monitor. **PUT semantics:** the call fully replaces the monitor's configuration — fields you omit revert to tool defaults, they are NOT left untouched. Before editing, read the current config with `get_monitors(monitor_ids=[<uuid>], include_fields=["config"])` and re-pass every field you want to keep. See `data-monitor-creation.md` (Step 7) for the safe-edit workflow. |
      | `dry_run` | bool | Default `True`. Preview mode. When omitted or `True`, returns YAML preview in `result.yaml`. When `False`, actually creates/updates the monitor and returns `result.monitor_uuid` + a deep link in `result.instructions`. See `data-monitor-creation.md`. |
      
      ---
      
      ## Alert Conditions
      
      Field names inside `alert_condition` are camelCase (`thresholdValue`, `thresholdSensitivity`, `baselineAggFunction`, ...) — NOT snake_case. Snake_case keys like `threshold_value` are rejected with an `extra_forbidden` validation error.
      
      | Field | Type | Required | Description |
      |-------|------|----------|-------------|
      | `operator` | string | Yes | One of: `EQ`, `NEQ`, `LT`, `LTE`, `GT`, `GTE`, `OUTSIDE_RANGE`, `INSIDE_RANGE`, `AUTO`, `AUTO_HIGH`, `AUTO_LOW`, `NOOP`. Note: the inequality operator is `NEQ` (not `NE`). |
      | `thresholdValue` | number | For explicit operators | Numeric threshold to compare the query result against. Pair with `GT`, `GTE`, `LT`, `LTE`, `EQ`, `NEQ`. |
      | `type` | string | No | Comparison semantics — see Threshold types below. Default: `threshold`. |
      
      ### Threshold types
      
      | `type` | Behavior | Fields (in addition to `operator`) |
      |---|---|---|
      | `threshold` (default) | Compare the query result directly against a fixed value. | `thresholdValue` |
      | `dynamic_threshold` | ML anomaly detection on the query result. | `operator` = `AUTO` / `AUTO_HIGH` / `AUTO_LOW`; optional `thresholdSensitivity` (`low` / `medium` / `high`, default `medium`) |
      | `change` | Compare the query result against a recent baseline (e.g. 2-day rolling MAX). | `thresholdValue`, `baselineAggFunction`, `baselineIntervalMinutes`, `isThresholdRelative` |
      | `noop` | Collect data without alerting. | `operator` = `NOOP`, no threshold |
      
      **`change`-type fields:**
      
      - `baselineAggFunction` — how to aggregate baseline samples. One of: `AVG`, `MIN`, `MAX`. Backend rejects anything else: `Must be one of: AVG, MIN, MAX.`
      - `baselineIntervalMinutes` — lookback window for the baseline, in minutes (e.g. `1440` = last 24h). Required with `type="change"`; the backend accepts up to `129600` (90 days).
      - `isThresholdRelative` — `true` if `thresholdValue` is a percentage (relative to baseline), `false` if it is an absolute delta. Defaults to `false`.
      
      Omitting these on a `change`-type condition produces stacked `required` / `Aggregate function is required` / `Lookback Interval in minutes should be between 0 and 129600` backend errors.
      
      ### Operator and type pairing
      
      Not every operator is accepted for every threshold type. The default `threshold` type only supports: `EQ`, `NEQ`, `LT`, `LTE`, `GT`, `GTE`, `OUTSIDE_RANGE`, `INSIDE_RANGE`. Pair `INSIDE_RANGE` / `OUTSIDE_RANGE` with `lowerThreshold` + `upperThreshold` instead of `thresholdValue`.
      
      If the user is unsure what threshold to set, help them reason about it: "What value would indicate a problem? If the query returns X, should that fire an alert?"
      
      ---
      
      ## SQL Query Requirements
      
      The SQL query MUST return exactly **one row with one numeric column**. This is non-negotiable -- the monitor compares that single value against the alert conditions.
      
      ### Rules
      
      - Use aggregate functions: `COUNT(*)`, `SUM()`, `AVG()`, `MAX()`, `MIN()`, or similar.
      - Can reference any table, view, or materialized view accessible in the warehouse.
      - Can use joins, subqueries, CTEs, window functions -- any valid SQL.
      - Do **NOT** include trailing semicolons.
      - Do **NOT** include comments (`--` or `/* */`) -- some warehouses strip them inconsistently.
      
      ### SQL Validation Tips
      
      These are the most common mistakes that cause custom SQL monitors to fail or produce misleading results:
      
      1. **Handle NULLs with COALESCE.** If your aggregate could return NULL (e.g., `SUM(amount)` on an empty result set), wrap it: `SELECT COALESCE(SUM(amount), 0) FROM ...`. A NULL result cannot be compared against a threshold and will not trigger alerts.
      
      2. **Ensure exactly one row, one column.** If your query could return zero rows (e.g., a filtered `SELECT` with no `GROUP BY`), wrap it in an outer aggregate: `SELECT COUNT(*) FROM (SELECT ...) sub`. If it returns multiple columns, select only the one you need.
      
      3. **Test the query mentally.** Before finalizing, ask: "If this query returns 5, will the alert condition fire correctly?" Walk through the logic with a concrete number.
      
      4. **For time-windowed checks, use appropriate date functions.** SQL syntax for date arithmetic varies by warehouse (see Warehouse-Specific SQL Notes below). Always scope time windows to avoid scanning the entire table history.
      
      5. **Avoid non-deterministic results.** Queries using `LIMIT` without `ORDER BY`, or `RANDOM()`, produce unpredictable results that make alerting unreliable.
      
      ---
      
      ## Warehouse-Specific SQL Notes
      
      SQL syntax for date arithmetic and functions varies across warehouses. When writing time-windowed queries, use the correct syntax for the user's warehouse:
      
      | Operation | Snowflake | BigQuery | Redshift |
      |-----------|-----------|----------|----------|
      | Subtract 1 day from now | `DATEADD(day, -1, CURRENT_TIMESTAMP())` | `DATE_SUB(CURRENT_TIMESTAMP(), INTERVAL 1 DAY)` | `DATEADD(day, -1, GETDATE())` |
      | Subtract 1 hour from now | `DATEADD(hour, -1, CURRENT_TIMESTAMP())` | `TIMESTAMP_SUB(CURRENT_TIMESTAMP(), INTERVAL 1 HOUR)` | `DATEADD(hour, -1, GETDATE())` |
      | Current timestamp | `CURRENT_TIMESTAMP()` | `CURRENT_TIMESTAMP()` | `GETDATE()` |
      | Date truncation | `DATE_TRUNC('day', col)` | `DATE_TRUNC(col, DAY)` | `DATE_TRUNC('day', col)` |
      
      When unsure which warehouse the user is on, ask. Getting the syntax wrong causes the monitor to fail on every scheduled run.
      
      ---
      
      ## Examples
      
      ### Orphan records (GT 0)
      
      Alert when orders reference customers that don't exist.
      
      ```json
      {
        "name": "orphan_orders_check",
        "description": "Detect orders referencing non-existent customers",
        "warehouse": "production_snowflake",
        "sql": "SELECT COUNT(*) FROM analytics.core.orders o LEFT JOIN analytics.core.customers c ON o.customer_id = c.id WHERE c.id IS NULL",
        "alert_condition": {
          "operator": "GT",
          "thresholdValue": 0
        }
      }
      ```
      
      ### Daily revenue floor (LT threshold)
      
      Alert when total revenue for the past 24 hours drops below a minimum.
      
      ```json
      {
        "name": "daily_revenue_floor",
        "description": "Alert when daily revenue falls below $10,000",
        "warehouse": "production_snowflake",
        "sql": "SELECT COALESCE(SUM(amount), 0) FROM analytics.billing.transactions WHERE created_at >= DATEADD(day, -1, CURRENT_TIMESTAMP())",
        "alert_condition": {
          "operator": "LT",
          "thresholdValue": 10000
        }
      }
      ```
      
      ### Duplicate rate exceeds threshold
      
      Alert when the duplicate rate on a key field exceeds 1%.
      
      ```json
      {
        "name": "order_id_duplicate_rate",
        "description": "Alert when order_id duplicate rate exceeds 1%",
        "warehouse": "production_snowflake",
        "sql": "SELECT COALESCE(1.0 - (COUNT(DISTINCT order_id) * 1.0 / NULLIF(COUNT(*), 0)), 0) FROM analytics.core.orders WHERE created_at >= DATEADD(day, -1, CURRENT_TIMESTAMP())",
        "alert_condition": {
          "operator": "GT",
          "thresholdValue": 0.01
        }
      }
      ```
      
      ### Range check (OUTSIDE_RANGE)
      
      Alert when a value falls outside an acceptable range. A two-sided range is a single condition with `lowerThreshold` + `upperThreshold`, not two separate conditions.
      
      ```json
      {
        "name": "avg_order_amount_range",
        "description": "Alert when average order amount is outside the $20-$500 range",
        "warehouse": "production_snowflake",
        "sql": "SELECT COALESCE(AVG(amount), 0) FROM analytics.core.orders WHERE created_at >= DATEADD(day, -1, CURRENT_TIMESTAMP()) AND status = 'completed'",
        "alert_condition": {
          "operator": "OUTSIDE_RANGE",
          "lowerThreshold": 20,
          "upperThreshold": 500
        }
      }
      ```
      
      ### Cross-table freshness check (BigQuery syntax)
      
      Alert when the latest row in a downstream table is more than 2 hours behind the source.
      
      ```json
      {
        "name": "pipeline_lag_check",
        "description": "Alert when downstream table lags source by more than 2 hours",
        "warehouse": "production_bigquery",
        "sql": "SELECT COALESCE(TIMESTAMP_DIFF(s.max_ts, t.max_ts, MINUTE), 9999) FROM (SELECT MAX(event_timestamp) AS max_ts FROM project.raw.events) s CROSS JOIN (SELECT MAX(processed_at) AS max_ts FROM project.analytics.events_processed) t",
        "alert_condition": {
          "operator": "GT",
          "thresholdValue": 120
        }
      }
      ```
      
    • data-metric-monitor.md 17.6 KB
      # Metric Monitor Reference
      
      Detailed reference for building `create_or_update_metric_monitor` tool calls. The tool follows the **two-call preview-then-confirm pattern** — see `data-monitor-creation.md` for the full flow.
      
      ## Critical Constraints
      
      - **NEVER guess column names.** Always get them from `get_table`. This is the most common source of monitor creation failures.
      - **`aggregate_time_field` MUST be a real timestamp column** from the table schema. Never assume or guess this value -- verify it exists in the `get_table` output.
      
      ---
      
      ## When to Use
      
      Use a metric monitor when the user wants to:
      
      - Track row count changes over time
      - Monitor null rates, unique counts, or other statistical metrics on specific fields
      - Detect anomalies in numeric distributions (mean, max, min, percentiles)
      - Monitor data freshness (time since last row count change)
      - Segment metrics by dimensions (e.g., by country, status)
      
      ---
      
      ## Required Parameters
      
      | Parameter | Type | Description |
      |-----------|------|-------------|
      | `name` | string | Unique identifier for the monitor. Use a descriptive slug (e.g., `orders_null_check`). |
      | `description` | string | Human-readable description of what the monitor checks. |
      | `table` | string | Table MCON (preferred) or `database:schema.table` format. If not MCON, also pass `warehouse`. |
      | `alert_conditions` | array | List of alert condition objects (see Alert Conditions below). |
      
      ## Optional Parameters
      
      | Parameter | Type | Default | Description |
      |-----------|------|---------|-------------|
      | `aggregate_time_field` | string | none | Timestamp/datetime column for time-windowed aggregation. **When provided, MUST be a real column from the table — NEVER guess this value.** When omitted, the monitor queries all rows on each run (whole-table scan). Omit for tables without a suitable timestamp column. |
      | `aggregate_time_sql` | string | none | SQL expression that produces the timestamp for bucketing (e.g. `CAST(payload:event_time AS TIMESTAMP)`). Use when the timestamp is embedded in a variant/object column or needs transformation. Mutually exclusive with `aggregate_time_field`. |
      | `warehouse` | string | auto-resolved | Warehouse name or UUID. Required if `table` is not an MCON. |
      | `segment_fields` | array of string | none | Fields to group/segment metrics by (e.g., `["country", "status"]`). |
      | `segment_sql` | array of string | none | SQL expressions to segment by (e.g. `["CASE WHEN amount > 100 THEN 'high' ELSE 'low' END"]`). |
      | `aggregate_by` | string | `"day"` | Time interval: `"hour"`, `"day"`, `"week"`, `"month"`. |
      | `where_condition` | string | none | SQL WHERE clause (without `WHERE` keyword) to filter rows before computing metrics. |
      | `sensitivity` | string | none | Anomaly detection sensitivity for AUTO operators: `"low"`, `"medium"`, `"high"`. |
      | `collection_lag_hours` | int | none | Hours to wait after expected data arrival before running the monitor. |
      | `schedule_type` | string | `"fixed"` | Schedule type: `"fixed"`, `"dynamic"`, `"manual"`. |
      | `interval_minutes` | int | auto | Schedule interval in minutes. Must be compatible with `aggregate_by` (see note below). If not specified, the tool defaults to the minimum valid interval for the chosen `aggregate_by`. |
      | `domain_uuids` | array of string (uuid) | none | Domain UUIDs (use `get_domains` to list). Data monitors accept exactly one UUID in the list. |
      | `audiences` | array of string | none | Notification audience **names** (not UUIDs) to alert when the monitor triggers. |
      | `failure_audiences` | array of string | none | Notification audience names to alert on query execution failures. |
      | `notes` | string | none | Free-text notes shown in the UI (separate from `description`). |
      | `priority` | string | none | Monitor priority (e.g. `"P1"`, `"P2"`). |
      | `tags` | array of `{name, value}` | none | Key-value tags to attach. |
      | `is_draft` | bool | `False` | When `True`, saves the monitor as a draft (not active). |
      | `monitor_uuid` | string (uuid) | none | UUID of an existing monitor to update in place. Omit to create a new monitor. **PUT semantics:** the call fully replaces the monitor's configuration — fields you omit revert to tool defaults, they are NOT left untouched. Before editing, read the current config with `get_monitors(monitor_ids=[<uuid>], include_fields=["config"])` and re-pass every field you want to keep. See `data-monitor-creation.md` (Step 7) for the safe-edit workflow. |
      | `dry_run` | bool | `True` | Preview mode. When omitted or `True`, returns YAML preview in `result.yaml`. When `False`, actually creates/updates the monitor and returns `result.monitor_uuid` + a deep link in `result.instructions`. See `data-monitor-creation.md`. |
      
      ---
      
      ## Schedule and Aggregation Compatibility
      
      The schedule interval must be compatible with `aggregate_by`. Daily aggregation requires an interval that is a multiple of 1440 minutes (24 hours), weekly requires a multiple of 10080, etc. If you pass `interval_minutes`, make sure it satisfies this constraint. If you omit it, the tool picks a sensible default.
      
      | `aggregate_by` | Minimum `interval_minutes` | Default if omitted |
      |---|---|---|
      | `hour` | 60 | 60 |
      | `day` | 1440 | 1440 |
      | `week` | 10080 | 10080 |
      | `month` | 43200 | 43200 |
      
      For example, to run a daily-aggregated monitor every other day, pass `aggregate_by: "day"` and `interval_minutes: 2880`.
      
      ---
      
      ## Choosing the Timestamp Field
      
      The `aggregate_time_field` controls whether the monitor uses time-windowed aggregation or whole-table scans. When provided, it MUST be a real column from the table — this is the number one source of monitor creation failures.
      
      ### When to omit it
      
      Omit `aggregate_time_field` when:
      - The table has **no timestamp or datetime columns** at all.
      - The table uses a **truncate-and-reload** pattern (fully replaced on each pipeline run) — time-windowed aggregation is meaningless since all rows share the same load time.
      - The user wants to monitor the **entire table state** on each run (e.g., `RELATIVE_ROW_COUNT` segmented by a dimension).
      
      When omitted, the monitor queries all rows on each run. This works well for small-to-medium tables but can be expensive for very large tables.
      
      ### How to pick it
      
      1. You should already have the column names **and their data types** from `get_table` with `include_fields: true` (done in Step 2 of the main skill).
      2. Look for columns whose names suggest a timestamp: `created_at`, `updated_at`, `modified_at`, `timestamp`, `event_timestamp`, or columns with `_ts`, `_dt`, `_time` suffixes, or `date`, `datetime`.
      3. **Verify the column's data type is an actual datetime/timestamp/date type** — not a string, number, or other type that happens to have a timestampy name. The backend rejects non-datetime types with `Field <name> is not a valid type to group the metrics by; it cannot be interpreted as a datetime.`
      4. If the user specified one, verify it exists in the column list AND has a datetime type.
      5. If exactly one obvious candidate exists (correct type), suggest it.
      6. If multiple candidates exist, present them and ask the user.
      7. If NO datetime-typed columns exist, omit the field — the monitor will do a whole-table scan. For very large tables, consider whether a custom SQL monitor would be more efficient.
      
      **NEVER** guess a timestamp field name, and never pick a column based on its name alone — always confirm the datatype from `get_table`, or omit the field.
      
      ### Common timestamp field mistakes
      
      - **Using a DATE column (not TIMESTAMP):** This may work, but aggregation granularity is limited. For example, `aggregate_by: "hour"` is meaningless on a DATE column because the time component is always midnight. Warn the user and default to `aggregate_by: "day"` or coarser.
      - **Using a field that contains many nulls:** If the timestamp column has significant null values, rows with null timestamps are excluded from aggregation windows, producing unreliable or misleading results. Check the column's null rate from `get_table` field stats if available, and warn the user if it is high.
      - **Guessing a field name that does not exist:** Always verify the column name against the `get_table` output. A typo or assumed name (e.g., `created_date` when the actual column is `created_at`) causes the monitor creation to fail silently or error.
      
      ---
      
      ## Field-Type-to-Metric Compatibility Matrix
      
      **Before selecting a metric, check the column's data type from `get_table` results.** Passing a metric incompatible with the column type is the most common source of creation failures after timestamp issues.
      
      | Column Type | Compatible Metrics |
      |-------------|-------------------|
      | **Numeric** (int, float, decimal, bigint) | `NUMERIC_MEAN`, `NUMERIC_MEDIAN`, `NUMERIC_MIN`, `NUMERIC_MAX`, `NUMERIC_STDDEV`, `SUM`, `ZERO_COUNT`, `ZERO_RATE`, `NEGATIVE_COUNT`, `NEGATIVE_RATE`, `NULL_COUNT`, `NULL_RATE`, `UNIQUE_COUNT`, `UNIQUE_RATE`, `DUPLICATE_COUNT` |
      | **String / Text** (varchar, char, text) | `TEXT_MAX_LENGTH`, `TEXT_MIN_LENGTH`, `TEXT_MEAN_LENGTH`, `TEXT_INT_RATE`, `TEXT_NUMBER_RATE`, `TEXT_UUID_RATE`, `TEXT_EMAIL_ADDRESS_RATE`, `EMPTY_STRING_COUNT`, `EMPTY_STRING_RATE`, `NULL_COUNT`, `NULL_RATE`, `UNIQUE_COUNT`, `UNIQUE_RATE`, `DUPLICATE_COUNT` |
      | **Boolean** | `TRUE_COUNT`, `FALSE_COUNT`, `NULL_COUNT`, `NULL_RATE` |
      | **Timestamp / Date** | `FUTURE_TIMESTAMP_COUNT`, `PAST_TIMESTAMP_COUNT`, `UNIX_ZERO_TIMESTAMP_COUNT`, `NULL_COUNT`, `NULL_RATE`, `UNIQUE_COUNT`, `UNIQUE_RATE` |
      | **Any type** | `NULL_COUNT`, `NULL_RATE`, `UNIQUE_COUNT`, `UNIQUE_RATE`, `DUPLICATE_COUNT` |
      
      ### Rules
      
      - **NEVER** apply `NUMERIC_*`, `SUM`, `ZERO_*`, or `NEGATIVE_*` metrics to string, boolean, or timestamp columns.
      - **NEVER** apply `TEXT_*` or `EMPTY_STRING_*` metrics to numeric, boolean, or timestamp columns.
      - **NEVER** apply `TRUE_COUNT` or `FALSE_COUNT` to non-boolean columns.
      - **NEVER** apply `FUTURE_TIMESTAMP_COUNT`, `PAST_TIMESTAMP_COUNT`, or `UNIX_ZERO_TIMESTAMP_COUNT` to non-timestamp columns.
      - When in doubt, `NULL_COUNT`, `NULL_RATE`, `UNIQUE_COUNT`, and `UNIQUE_RATE` are safe for any column type.
      
      ### Common metric-name mistakes
      
      The `NUMERIC_*` prefix pattern covers mean/median/min/max/stddev but **not** sum: the metric is `SUM`, not `NUMERIC_SUM`. Backend rejects with `Invalid metric: NUMERIC_SUM`.
      
      Other names agents guess-and-get-wrong:
      
      | Guessed (wrong) | Use instead |
      |---|---|
      | `NUMERIC_SUM` | `SUM` |
      | `APPROX_DISTINCT_COUNT`, `COUNT_DISTINCT` | `UNIQUE_COUNT` |
      | `COUNT_NULL`, `NULLS` | `NULL_COUNT` |
      | `ROW_COUNT` (as a column metric) | `ROW_COUNT_CHANGE` (table-level only) |
      
      If the metric you want isn't in the compatibility matrix above, it doesn't exist — use the closest alternative or fall back to a custom SQL monitor.
      
      ---
      
      ## Alert Conditions
      
      Alert-condition field names are camelCase (`thresholdValue`, not `threshold_value` or `threshold`) — snake_case keys are rejected with an `extra_forbidden` validation error.
      
      Each alert condition has:
      
      | Field | Type | Required | Description |
      |-------|------|----------|-------------|
      | `metric` | string | Yes | The metric to monitor (see Metrics Reference below). |
      | `operator` | string | Yes | `"AUTO"` (anomaly detection), `"GT"`, `"LT"`, `"EQ"`, `"GTE"`, `"LTE"`, `"NEQ"`. Note: the inequality operator is `NEQ`, not `NE`. |
      | `thresholdValue` | number | For explicit operators | The threshold value. Required when using `GT`, `LT`, `EQ`, `GTE`, `LTE`, or `NEQ`. Not used with `AUTO`. |
      | `fields` | array of string | Depends | Column names to apply the metric to. Required for field-level metrics. Not needed for table-level metrics. |
      
      ---
      
      ## Operator Guidance
      
      ### When to use `AUTO` (anomaly detection)
      
      - Best when you do not know the expected range of values and want Monte Carlo's ML to learn normal patterns and alert on deviations.
      - Works well for organic metrics that vary day-to-day (row counts, null rates on evolving data, numeric distributions).
      - Some metrics **require** `AUTO` -- see the table below.
      
      ### When to use explicit thresholds (`GT`, `LT`, `EQ`, `GTE`, `LTE`, `NEQ`)
      
      - Use when there is a known business rule or data contract (e.g., "null rate on `email` should never exceed 5%", "order amount must always be greater than 0").
      - Provides deterministic alerting -- no training period needed, alerts fire immediately when the condition is met.
      - Requires a `thresholdValue` in the alert condition.
      
      ### Operator restrictions by metric
      
      | Metric | Allowed Operators | Notes |
      |--------|-------------------|-------|
      | `ROW_COUNT_CHANGE` | `AUTO` only | Anomaly detection on row count delta. |
      | `TIME_SINCE_LAST_ROW_COUNT_CHANGE` | `AUTO` only | Anomaly detection on staleness duration. |
      | `RELATIVE_ROW_COUNT` | `AUTO` only | Anomaly detection on segment distribution. Requires `segment_fields`. |
      | All other metrics | `AUTO`, `GT`, `LT`, `EQ`, `GTE`, `LTE`, `NEQ` | Any operator is valid. |
      
      ---
      
      ## Metrics Reference
      
      ### Table-level metrics (no `fields` needed)
      
      | Metric | Operator | Description |
      |--------|----------|-------------|
      | `ROW_COUNT_CHANGE` | Must use `AUTO` | Alert on anomalous changes in total row count. |
      | `TIME_SINCE_LAST_ROW_COUNT_CHANGE` | Must use `AUTO` | Alert when the table has not been updated for an unusual duration. |
      
      ### Field-level metrics (must specify `fields`)
      
      | Metric | Column Types | Description |
      |--------|-------------|-------------|
      | `NULL_COUNT` | Any | Count of null values. |
      | `NULL_RATE` | Any | Rate of null values (0.0 to 1.0). |
      | `UNIQUE_COUNT` | Any | Count of distinct values. |
      | `UNIQUE_RATE` | Any | Rate of distinct values (0.0 to 1.0). |
      | `DUPLICATE_COUNT` | Any | Count of duplicate (non-unique) values. |
      | `EMPTY_STRING_COUNT` | String/Text | Count of empty string values. |
      | `EMPTY_STRING_RATE` | String/Text | Rate of empty string values. |
      | `NUMERIC_MEAN` | Numeric | Mean of numeric field. |
      | `NUMERIC_MEDIAN` | Numeric | Median of numeric field. |
      | `NUMERIC_MIN` | Numeric | Minimum value of numeric field. |
      | `NUMERIC_MAX` | Numeric | Maximum value of numeric field. |
      | `NUMERIC_STDDEV` | Numeric | Standard deviation of numeric field. |
      | `SUM` | Numeric | Sum of numeric field. |
      | `ZERO_COUNT` | Numeric | Count of zero values. |
      | `ZERO_RATE` | Numeric | Rate of zero values. |
      | `NEGATIVE_COUNT` | Numeric | Count of negative values. |
      | `NEGATIVE_RATE` | Numeric | Rate of negative values. |
      | `TRUE_COUNT` | Boolean | Count of true values. |
      | `FALSE_COUNT` | Boolean | Count of false values. |
      | `TEXT_MAX_LENGTH` | String/Text | Maximum string length. |
      | `TEXT_MIN_LENGTH` | String/Text | Minimum string length. |
      | `TEXT_MEAN_LENGTH` | String/Text | Mean string length. |
      | `TEXT_INT_RATE` | String/Text | Rate of values parseable as integers. |
      | `TEXT_NUMBER_RATE` | String/Text | Rate of values parseable as numbers. |
      | `TEXT_UUID_RATE` | String/Text | Rate of values matching UUID format. |
      | `TEXT_EMAIL_ADDRESS_RATE` | String/Text | Rate of values matching email format. |
      | `FUTURE_TIMESTAMP_COUNT` | Timestamp/Date | Count of timestamps in the future. |
      | `PAST_TIMESTAMP_COUNT` | Timestamp/Date | Count of timestamps unreasonably far in the past. |
      | `UNIX_ZERO_TIMESTAMP_COUNT` | Timestamp/Date | Count of timestamps equal to Unix epoch zero (1970-01-01). |
      
      ### Segmentation metric
      
      | Metric | Operator | Description |
      |--------|----------|-------------|
      | `RELATIVE_ROW_COUNT` | Must use `AUTO` | Alert on anomalous changes in distribution across segments. MUST use `segment_fields`. |
      
      ---
      
      ## Examples
      
      ### Row count anomaly detection
      
      ```json
      {
        "name": "orders_row_count",
        "description": "Detect anomalous changes in daily order volume",
        "table": "MCON++a1b2c3d4-e5f6-7890-abcd-ef1234567890++1++1++analytics:core.orders",
        "aggregate_time_field": "created_at",
        "aggregate_by": "day",
        "alert_conditions": [
          {
            "metric": "ROW_COUNT_CHANGE",
            "operator": "AUTO"
          }
        ]
      }
      ```
      
      ### Null monitoring on specific fields
      
      ```json
      {
        "name": "orders_null_check",
        "description": "Alert when email or user_id nulls exceed 50 per day",
        "table": "MCON++a1b2c3d4-e5f6-7890-abcd-ef1234567890++1++1++analytics:core.orders",
        "aggregate_time_field": "created_at",
        "aggregate_by": "day",
        "alert_conditions": [
          {
            "metric": "NULL_COUNT",
            "operator": "GT",
            "thresholdValue": 50,
            "fields": ["email", "user_id"]
          }
        ]
      }
      ```
      
      ### Segmented monitoring
      
      ```json
      {
        "name": "orders_by_country_distribution",
        "description": "Detect anomalous shifts in order distribution across countries",
        "table": "MCON++a1b2c3d4-e5f6-7890-abcd-ef1234567890++1++1++analytics:core.orders",
        "aggregate_time_field": "created_at",
        "aggregate_by": "day",
        "segment_fields": ["country"],
        "alert_conditions": [
          {
            "metric": "RELATIVE_ROW_COUNT",
            "operator": "AUTO"
          }
        ]
      }
      ```
      
      ### Numeric range monitoring with filter
      
      ```json
      {
        "name": "completed_orders_amount_check",
        "description": "Detect anomalous max order amounts for completed orders",
        "table": "MCON++a1b2c3d4-e5f6-7890-abcd-ef1234567890++1++1++analytics:core.orders",
        "aggregate_time_field": "created_at",
        "aggregate_by": "day",
        "where_condition": "status = 'completed'",
        "alert_conditions": [
          {
            "metric": "NUMERIC_MAX",
            "operator": "AUTO",
            "fields": ["amount"]
          }
        ]
      }
      ```
      
      ### Multiple alert conditions in one monitor
      
      ```json
      {
        "name": "payments_quality_check",
        "description": "Monitor payment amount stats and null rate on transaction_id",
        "table": "MCON++a1b2c3d4-e5f6-7890-abcd-ef1234567890++1++1++warehouse:billing.payments",
        "aggregate_time_field": "processed_at",
        "aggregate_by": "day",
        "domain_uuids": ["f47ac10b-58cc-4372-a567-0e02b2c3d479"],
        "alert_conditions": [
          {
            "metric": "NUMERIC_MEAN",
            "operator": "AUTO",
            "fields": ["amount"]
          },
          {
            "metric": "NULL_RATE",
            "operator": "GT",
            "thresholdValue": 0.01,
            "fields": ["transaction_id"]
          }
        ]
      }
      ```
      
    • data-monitor-creation.md 22.7 KB
      # Data Monitor Creation Procedure
      
      This is the data monitor creation procedure for Monte Carlo warehouse tables. Use this reference when a user wants to create monitors for their data warehouse tables -- it walks through the full workflow from understanding the request through generating monitors-as-code (MaC) YAML and (optionally) deploying the monitor.
      
      All five `create_or_update_*_monitor` tools follow a **two-call preview-then-confirm pattern**:
      
      1. **First call -- preview.** Invoke with `dry_run=True` (this is the default -- you can omit the argument). The tool returns rendered MaC YAML in `result.yaml` and a DRY RUN notice in `result.instructions`. Show the YAML to the user and confirm.
      2. **Second call -- live create/update.** After the user confirms, invoke the same tool again with `dry_run=False` and the same other parameters. The tool actually creates or updates the monitor and returns `result.monitor_uuid` plus a `result.instructions` string containing a deep link `<webapp_url>/monitors/<monitor_uuid>` to the live monitor. `result.yaml` is intentionally `None` on this call -- the monitor is already deployed.
      
      To **update an existing monitor** instead of creating a new one, pass its `monitor_uuid`. This works on both the preview and live calls. **Important:** `create_or_update_*_monitor` with `monitor_uuid` has **PUT semantics** -- the call fully replaces the monitor's configuration. Fields you omit revert to the tool's defaults; they are NOT left untouched. See Step 7 ("Updating an existing monitor") for the safe-edit workflow. To save the monitor as a draft (not active), pass `is_draft=True`.
      
      The user may also choose to skip the live call and take the preview YAML themselves and apply it via the Monte Carlo CLI or CI/CD. Always present the YAML on the preview call regardless.
      
      ---
      
      ## Validation Phase (Steps 1-3)
      
      **CRITICAL: Do not call creation tools before the validation phase is complete.** The number one error pattern is agents skipping validation and calling a creation tool with guessed or incomplete parameters. Every field in the creation call must be grounded in data retrieved during this phase.
      
      ### Step 1: Understand the request
      
      Ask yourself:
      - What does the user want to monitor? (a specific table, a metric, a data quality rule, cross-table consistency, freshness/volume at schema level)
      - Which monitor type fits? Use the monitor type selection table below.
      - Does the user have all the details, or do they need guidance?
      
      If the user's intent is unclear, ask a focused question before proceeding.
      
      ### Step 2: Identify the table(s) and columns
      
      If you don't have the table MCON:
      1. Use `search` with the table name and `include_fields: ["field_names"]` to find the MCON and get column names.
      2. If the user provided a full table ID like `database:schema.table`, search for it.
      3. Once you have the MCON, call `get_table` with `include_fields: true` and `include_table_capabilities: true` to verify capabilities and get domain info.
      
      If you already have the MCON:
      1. Call `get_table` with the MCON, `include_fields: true`, and `include_table_capabilities: true`.
      
      **If `search` returns zero results, or `get_table` shows the table is not ingested:** STOP. The table must already exist in Monte Carlo before a monitor can be created against it. Ask the user to confirm the correct table name, or to ingest the table first — do not call the creation tool with an unverified table.
      
      **CRITICAL: You need the actual column names from `get_table` results. NEVER guess or hallucinate column names.** This is the most common source of monitor creation failures.
      
      **Pre-call column-verification gate (run this immediately before calling any creation tool):**
      
      1. List every column name you plan to put in the tool arguments, in every slot the per-type reference describes (see the Tier-3 file for the authoritative list of column-bearing parameters).
      2. For each name, confirm it appears verbatim in the `get_table.fields` list you fetched in this step. Names are case-sensitive on most warehouses (Snowflake often returns uppercase column names — match exactly).
      3. If any name is missing, STOP. Do not call the creation tool. Ask the user to confirm the correct column name, or suggest the closest matches from the actual column list — do NOT substitute a similar-sounding name on your own.
      
      If you reached this step without calling `get_table` (or equivalent) for the target table, go back — you cannot skip the fetch.
      
      For monitor types that require a timestamp column (metric monitors), review the column names and identify likely timestamp candidates. Present them to the user if ambiguous.
      
      **CRITICAL: The `warehouse` parameter on creation tools is a UUID, not a name.** Extract it from the `get_table` response (the resource / warehouse UUID). If you only have a warehouse name and no MCON, call `get_warehouses` to resolve it -- NEVER pass a warehouse name string like `databricks-aws-agent` or `snowflake-prod`, the backend will reject with `Warehouse not found`.
      
      ### Step 3: Handle domain assignment
      
      **ALWAYS resolve a `domain_uuids` value BEFORE calling any creation tool.** Missing or empty domain assignment is one of the top failure modes — the backend will reject the monitor with `Domain assignment is required for this monitor. Please provide one and only one valid domain UUID.`
      
      The tool field is `domain_uuids` (a list). For data monitors, provide exactly one UUID.
      
      Use the `domains` list on the `get_table` response (each entry has `uuid` and `name`):
      
      1. If the table's `domains` has exactly one entry: default `domain_uuids` to `[<that uuid>]`.
      2. If the table's `domains` has multiple entries: present only those domains and ask the user to pick.
      3. If the table's `domains` is empty: call `get_domains` to see the account's domains. If the account has one or more, ask the user to pick one (do not invent a selection) -- note that domains that don't contain the table may still be rejected on apply. If `get_domains` returns zero domains, only then may `domain_uuids` be omitted.
      
      Do NOT present all account domains as options when the table already has domains listed -- prefer domains that contain the table.
      
      **Agent-onboarding context.** If this create is part of an agent-onboarding flow (the table is a customer agent's upstream/golden table and an agent is in scope — e.g. you arrived here from the Context pillar of `agent-monitor-creation.md`), every monitor you create must carry that agent's footprint tag `tags=[{"name": "agent", "value": "<AGENT_NAME>"}]` (display name from the onboarding flow, verbatim) and reuse the same `audiences` and `domain_uuids` the agent's monitors use — this keeps the agent's whole footprint retrievable with a single tag filter (`get_monitors(monitor_tags=["agent:<AGENT_NAME>"])`). Outside an agent-onboarding flow, do NOT add an `agent` tag.
      
      ### Step 3b: Ground thresholds and predicates in real data (profiling)
      
      Verified column names are not enough. Schema alone does not tell you whether a column is mostly-null, whether a status code is truly a closed set, or whether a numeric range is stable enough to alert on. Profiling once up front is the difference between a useful monitor and a noisy one the user mutes the next day.
      
      | Monitor / config | Profiling | How to ground it |
      | --- | --- | --- |
      | Metric monitor, **auto / ML threshold** | **Not required** | The backend learns the baseline from history -- you don't pick a number. |
      | Metric monitor, **manual `min`/`max`** | **Required** (unless the user pre-opts-out) | Sample the metric over a recent window and pick a threshold outside normal variation but tight enough to catch regressions. |
      | **Validation** monitor (any predicate) | **Required** | Membership predicates (`in_set`/`not_in_set`/equality): sample distinct values. Range/regex/cross-field: sample the actual distribution. |
      | **Custom SQL** monitor | **Required** | Run the proposed SQL once first -- confirm it parses, returns the expected shape, and produces values consistent with the threshold. |
      
      How to profile depends on the environment: use `get_table` field stats where they suffice, or an optional database MCP (`snowflake_query`, `bigquery_query`, etc.) for distributions and distinct values. **If you cannot profile** (no query access, permission denied, user declines): do NOT invent values. Fall back to **auto / ML thresholds** where the monitor supports them, propose a metadata-only sketch for the user to refine, or pause -- and say which. A user instruction like "use my number directly", "skip profiling", or "just give me the dry-run" is a valid pre-opt-out; honor it without re-prompting.
      
      ### Step 3c: Field monitors require a live table monitor (prerequisite)
      
      A metric or validation monitor on a table only runs if that table has an active **user-deployed table monitor** -- not the auto-applied out-of-the-box freshness/volume. A field monitor on an OOTB-only table is silently inert. Before proposing a metric/validation monitor, check the table's monitors (`get_monitors` for the MCON); if there is no user table monitor, tell the user the field monitor would not actually run, and offer to create a table monitor first. (Custom SQL monitors are exempt.)
      
      ---
      
      ## Creation Phase (Steps 4-8)
      
      Only enter this phase after the validation phase is complete with real data from MCP tools.
      
      ### Step 4: Load the per-type reference
      
      Based on the monitor type, read the detailed reference for parameter guidance:
      
      | Type           | Reference file               |
      | -------------- | ---------------------------- |
      | **Metric**     | `data-metric-monitor.md`     |
      | **Validation** | `data-validation-monitor.md` |
      | **Custom SQL** | `data-custom-sql-monitor.md` |
      | **Comparison** | `data-comparison-monitor.md` |
      | **Table**      | `data-table-monitor.md`      |
      
      All reference files are in the same directory as this file.
      
      **CRITICAL: Every enum value comes from the per-type reference.** `metric`, `operator`, predicate `name`, `schedule.type`, `aggregate_by`, and any other enum-shaped parameter must match the exact strings documented in the Tier-3 file for this monitor type. Never invent values by analogy or adjust casing — the backend rejects anything outside the documented set. Subsets apply per threshold type (e.g. custom_sql Absolute Threshold allows fewer operators than the full list); the per-type file spells those out too. If you're unsure, ask the user rather than guessing.
      
      ### Step 5: Ask about scheduling
      
      **Skip this step for table monitors.** Table monitors do not support the `schedule` field in MaC YAML -- adding it will cause a validation error on `montecarlo monitors apply`. Table monitor scheduling is managed automatically by Monte Carlo.
      
      For all other monitor types, the creation tools default to a fixed schedule running every 60 minutes. Present these options:
      
      1. **Fixed interval** -- any integer for `interval_minutes` (30, 60, 90, 120, 360, 720, 1440, etc.)
      2. **Dynamic** -- MC auto-determines when to run based on table update patterns.
      3. **Manual** -- runs only on demand.
      
      Pass the user's choice to the creation tool as `schedule_type` and (for fixed schedules) `interval_minutes`. **Both the preview (`dry_run=True`) and the live (`dry_run=False`) call must use the same schedule arguments** -- the tool re-renders the schedule from these parameters when it deploys, so editing the `schedule` section of the preview YAML by hand does NOT change what the live call creates. Without explicit arguments the backend falls back to fixed/60 regardless of what the YAML displayed to the user.
      
      Valid arguments:
      
      - Fixed: `schedule_type="fixed"`, `interval_minutes=<N>` (any integer, e.g. 30, 60, 90, 360, 720, 1440)
      - Dynamic: `schedule_type="dynamic"` (omit `interval_minutes`)
      - Manual: `schedule_type="manual"` (omit `interval_minutes`)
      
      **Views require a fixed schedule.** A view has no independent "last update" timestamp, so dynamic scheduling never triggers correctly. When the target is a view, always use a fixed schedule and surface that in the preview so the user sees the correct config.
      
      ### Step 6: Confirm with the user
      
      **NEVER skip the confirmation step.**
      
      Before calling the creation tool, present the monitor configuration in plain language:
      - Monitor type
      - Target table (and columns if applicable)
      - What it checks / what triggers an alert
      - Domain assignment
      - Schedule
      - Whether this is a new monitor or an in-place update (i.e. is `monitor_uuid` set?)
      - Whether to save as draft (`is_draft=True`) or active
      
      Ask: "Does this look correct? I'll generate the monitor configuration."
      
      Also ask how the user wants to deploy it:
      
      > **Deployment preference:** Deploy live now (via MCP), or save as a Monitors-as-Code YAML file to apply through your repo?
      >
      > - **Live (MCP):** I'll call the creation tool and the monitor will be active immediately.
      > - **MaC YAML:** I'll generate the YAML definition so you can commit it to your repo and apply it with `montecarlo monitors apply`. Use `/monte-carlo-manage-mac` if you want to validate or edit the file first.
      
      If the user chooses MaC YAML: generate the preview YAML (dry_run=True) as usual, present it wrapped in the standard MaC structure (see MaC YAML Format), and stop -- do not call with `dry_run=False`. The user takes the YAML from there.
      
      If the user chooses live or does not express a preference, proceed with the standard two-call sequence in Step 7.
      
      ### Step 7: Create the monitor
      
      This step is a **two-call sequence**. Do NOT skip the preview call.
      
      1. **Preview call.** Call the appropriate creation tool with the parameters built in previous steps. Omit `dry_run` (it defaults to `True`) or pass `dry_run=True` explicitly. Always pass an MCON when possible. If only a table name is available, also pass `warehouse`. The tool returns rendered YAML in `result.yaml` and a DRY RUN notice in `result.instructions`. Present the YAML per Step 8 and ask the user to confirm before proceeding.
      2. **Live call.** After the user confirms and explicitly opts in to deploying directly, call the same tool again with **the same parameters** plus `dry_run=False`. The tool actually creates (or updates) the monitor; the response carries the new `monitor_uuid` and a deep link in `result.instructions`. On this call `result.yaml` is `None` by design -- the monitor is already deployed.
      
      **Updating an existing monitor.** If the user wants to edit a monitor they (or a previous call) already created, pass `monitor_uuid=<uuid>` on both the preview and live calls. The tool will update that monitor in place rather than creating a new one. Use a previously returned `monitor_uuid`, or look one up via `get_monitors`. If the underlying monitor was deleted between read and write, the tool will raise a clear error instructing you to retry without `monitor_uuid` (turning the intent from "update" into "create").
      
      **PUT semantics -- do not skip this step.** `create_or_update_*_monitor` with `monitor_uuid` replaces the monitor configuration in full. Every parameter you omit reverts to the tool's default (e.g. schedule resets to fixed/60 minutes); it is NOT left untouched. To edit safely:
      
      1. **Read the current config first.** Call `get_monitors(monitor_ids=[<uuid>], include_fields=["config"])` to get the full monitor configuration. `config` is excluded by default for performance -- you must request it explicitly.
      2. **Carry over every value you want to keep**, in addition to the ones you're changing. Do not pass only the changed fields -- anything you leave out is overwritten with the tool default.
      3. **Preview with `dry_run=True` and diff** the rendered YAML against the original config. If anything you meant to preserve is missing or changed, fix the call before running `dry_run=False`.
      
      **Drafts.** Pass `is_draft=True` to save the monitor in draft state (not active). Omit it to create the monitor as active.
      
      ### Step 8: Present results
      
      Handle both response shapes.
      
      **Preview response (`dry_run=True`)** -- `result.yaml` is set; `result.monitor_uuid` is `None`; `result.instructions` includes a DRY RUN notice. You MUST include the YAML in your reply -- the user needs copy-pasteable YAML in the **same** message where you ask for confirmation. Do NOT refer back to "the YAML I showed you" or give deployment instructions without the actual YAML.
      
      1. The YAML comes verbatim from `result.yaml` -- the tool has already rendered the schedule from the `schedule_type` / `interval_minutes` you passed in. Do NOT post-edit the `schedule` section to change values; if the schedule is wrong, re-call the preview with corrected arguments.
      2. ALWAYS present the full YAML in a ```yaml code block. Present ALL YAML values exactly as returned by the tool. Do NOT reformat, convert, or "humanize" any values -- especially dates, timestamps, UUIDs, and identifiers.
      3. Wrap the YAML in the standard MaC structure before presenting it (see MaC YAML Format below).
      4. ALWAYS use ISO 8601 format for any datetime values you author (e.g. `start_time: '2026-03-25T09:00:00+00:00'`).
      5. **NEVER reformat YAML values returned by creation tools.**
      6. Explain the user's two options once they confirm: (a) let you re-call the tool with `dry_run=False` to deploy it directly in Monte Carlo, or (b) take the YAML and apply it themselves via Monte Carlo CLI or CI/CD.
      
      **Live response (`dry_run=False`)** -- `result.yaml` is `None`; `result.monitor_uuid` is the new (or updated) monitor's UUID; `result.instructions` contains a deep link of the form `<webapp_url>/monitors/<monitor_uuid>`.
      
      1. Confirm to the user that the monitor was created (or updated) and surface the deep link from `result.instructions` so they can click through to it in the Monte Carlo web app.
      2. Do NOT try to re-render or invent YAML -- it is intentionally not returned for live calls.
      
      ---
      
      ## Monitor Type Selection
      
      | Type           | Creation tool                         | Use when                                                                                                                               |
      | -------------- | ------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------- |
      | **Metric**     | `create_or_update_metric_monitor`     | Track statistical metrics on fields (null rates, unique counts, numeric stats) or row count changes over time. Requires a timestamp field for aggregation. |
      | **Validation** | `create_or_update_validation_monitor` | Row-level data quality checks with conditions (e.g. "field X is never null", "status is in allowed set"). Alerts on INVALID data.      |
      | **Custom SQL** | `create_or_update_sql_monitor`        | Run arbitrary SQL returning a single number and alert on thresholds. Most flexible; use when other types don't fit.                    |
      | **Comparison** | `create_or_update_comparison_monitor` | Compare metrics between two tables (e.g. dev vs prod, source vs target).                                                              |
      | **Table**      | `create_or_update_table_monitor`      | Monitor groups of tables for freshness, schema changes, and volume. Uses asset selection at database/schema level.                     |
      
      Per-type reference files with detailed parameter guidance, constraints, and examples:
      - `data-metric-monitor.md`
      - `data-validation-monitor.md`
      - `data-custom-sql-monitor.md`
      - `data-comparison-monitor.md`
      - `data-table-monitor.md`
      
      ---
      
      ## MaC YAML Format
      
      The YAML returned on the preview call (`dry_run=True`) is the monitor definition. It must be wrapped in the standard MaC structure to be applied:
      
      ```yaml
      montecarlo:
        <monitor_type>:
          - <returned yaml>
      ```
      
      For example, a metric monitor would look like:
      
      ```yaml
      montecarlo:
        metric:
          - <yaml returned by create_or_update_metric_monitor>
      ```
      
      **Important:** `montecarlo.yml` (without a directory path) is a separate Monte Carlo project configuration file -- it is NOT the same as a monitor definition file. Monitor definitions go in their own `.yml` files, typically in a `monitors/` directory or alongside dbt model schema files.
      
      If the user prefers to deploy via CLI/CI rather than the live tool call:
      - Save the YAML to a `.yml` file (e.g. `monitors/<table_name>.yml` or in their dbt schema)
      - Apply via the Monte Carlo CLI: `montecarlo monitors apply --namespace <namespace>`
      - Or integrate into CI/CD for automatic deployment on merge
      
      ---
      
      ## Schema Validation
      
      Always add the following comment as the **first line** of any MaC YAML file you create or edit:
      
      ```yaml
      # yaml-language-server: $schema=https://clidocs.getmontecarlo.com/mac/schema.json
      ```
      
      The published schema is available at `https://clidocs.getmontecarlo.com/mac/schema.json`. Use WebFetch to inspect it if you're uncertain whether a field name or value is valid for a given monitor type.
      
      Generated YAML must not include fields that don't appear in the schema for that monitor type. Unknown fields are silently ignored by the CLI but indicate a misconfiguration and may break future validation.
      
      **Schema scope:** The schema validates field names, types, and enum values only. Cross-field semantic constraints (e.g. required field combinations, mutually exclusive options, conditional required fields) are NOT checked by the schema — they are enforced by the Monte Carlo backend at apply time. A file that passes schema validation may still fail on `montecarlo monitors apply`.
      
      ---
      
      ## Available MCP Tools
      
      All tools are available via the `monte-carlo-mcp` MCP server.
      
      | Tool                            | Purpose                                                      |
      | ------------------------------- | ------------------------------------------------------------ |
      | `test_connection`               | Verify auth and connectivity before starting                 |
      | `search`                        | Find tables/assets by name; use `include_fields` for columns |
      | `get_table`                     | Schema, stats, metadata, domain membership, capabilities     |
      | `get_validation_predicates`     | List available validation rule types for a warehouse         |
      | `get_domains`                   | List MC domains (only needed if table has no domain info)    |
      | `get_warehouses`                          | Resolve warehouse UUIDs from names; needed when a name is the only identifier |
      | `get_monitors`                            | Look up an existing monitor's UUID for in-place updates via `monitor_uuid` |
      | `create_or_update_metric_monitor`         | Create or update a metric monitor (preview on `dry_run=True`, deploy on `dry_run=False`) |
      | `create_or_update_validation_monitor`     | Create or update a validation monitor (preview on `dry_run=True`, deploy on `dry_run=False`) |
      | `create_or_update_comparison_monitor`     | Create or update a comparison monitor (preview on `dry_run=True`, deploy on `dry_run=False`) |
      | `create_or_update_sql_monitor`            | Create or update a custom SQL monitor (preview on `dry_run=True`, deploy on `dry_run=False`) |
      | `create_or_update_table_monitor`          | Create or update a table monitor (preview on `dry_run=True`, deploy on `dry_run=False`) |
      
    • data-table-monitor.md 10.9 KB
      # Table Monitor Reference
      
      Detailed reference for building `create_or_update_table_monitor` tool calls. The tool follows the **two-call preview-then-confirm pattern** — see `data-monitor-creation.md` for the full flow.
      
      ## Critical Constraints
      
      - **NEVER guess column names.** Always verify table and schema names from `get_table` or `search` before building the asset selection.
      - **`alert_conditions` is a flat list of strings** — metric names like `"last_updated_on"`, `"schema"`, `"total_row_count"`. NEVER pass dicts like `{"metric": "last_updated_on", "operator": "AUTO"}`. That shape is rejected with `Input should be a valid string [type=string_type, input_value={'metric': '...', 'operator': 'AUTO'}]`. Table monitors do not take per-condition operators — they use anomaly detection on the named metrics by default.
      
      ---
      
      ## When to Use
      
      Use a table monitor when the user wants to:
      
      - Monitor many tables at once across an entire database or schema
      - Track freshness (when was each table last updated?)
      - Detect schema changes (columns added, removed, or type-changed)
      - Monitor volume changes (row count anomalies) across a broad set of tables
      - Apply broad coverage with anomaly detection (no custom thresholds needed)
      
      **Do NOT use a table monitor when the user wants to:**
      
      - Track field-level metrics on a single table (use a metric monitor)
      - Apply custom thresholds or explicit operators like GT/LT (use a metric monitor)
      - Validate row-level business rules or referential integrity (use a validation monitor)
      
      ---
      
      ## Required Parameters
      
      | Parameter | Type | Description |
      |-----------|------|-------------|
      | `name` | string | Unique identifier for the table monitor. Must be unique across all table monitors in the same namespace. |
      | `description` | string | Human-readable description of what the monitor checks (max 512 characters). |
      | `warehouse` | string | Warehouse name or UUID. Use `get_table` or `search` to find it. |
      | `asset_selection` | object | Asset selection config defining which tables to monitor (see Asset Selection below). |
      
      ## Optional Parameters
      
      | Parameter | Type | Default | Description |
      |-----------|------|---------|-------------|
      | `alert_conditions` | array of strings | `["last_updated_on", "schema", "total_row_count", "total_row_count_last_changed_on"]` | Metric names to monitor (see Alert Conditions below). |
      | `domain_uuids` | array of string (uuid) | none | Domain UUIDs (use `get_domains` to list). Data monitors accept exactly one UUID in the list. |
      | `audiences` | array of string | none | Notification audience **names** (not UUIDs) to alert when the monitor triggers. |
      | `failure_audiences` | array of string | none | Notification audience names to alert on query execution failures. |
      | `notes` | string | none | Free-text notes shown in the UI (separate from `description`). |
      | `priority` | string | none | Monitor priority (e.g. `"P1"`, `"P2"`). |
      | `tags` | array of `{name, value}` | none | Key-value tags to attach. |
      | `is_draft` | bool | `False` | When `True`, saves the monitor as a draft (not active). |
      | `monitor_uuid` | string (uuid) | none | UUID of an existing monitor to update in place. Omit to create a new monitor. **PUT semantics:** the call fully replaces the monitor's configuration — fields you omit revert to tool defaults, they are NOT left untouched. Before editing, read the current config with `get_monitors(monitor_ids=[<uuid>], include_fields=["config"])` and re-pass every field you want to keep. See `data-monitor-creation.md` (Step 7) for the safe-edit workflow. |
      | `dry_run` | bool | `True` | Preview mode. When omitted or `True`, returns YAML preview in `result.yaml`. When `False`, actually creates/updates the monitor and returns `result.monitor_uuid` + a deep link in `result.instructions`. See `data-monitor-creation.md`. |
      
      ---
      
      ## Pre-Step: Verify Warehouse
      
      Before creating a table monitor, resolve the warehouse name or UUID. The `warehouse` parameter is required and must match an existing warehouse in the Monte Carlo account.
      
      1. If the user provides a table name, call `get_table` to retrieve the table details -- the response includes the warehouse name and UUID.
      2. If the user provides a database or schema name without a specific table, call `search` with the database or schema name to find assets and identify the warehouse.
      3. Use either the warehouse name or UUID in the `warehouse` parameter.
      
      **NEVER guess the warehouse value.** If you cannot resolve it, ask the user.
      
      ---
      
      ## Asset Selection
      
      The `asset_selection` object defines which tables the monitor covers. It must include a `databases` list.
      
      **Use database and schema scoping to select which tables to monitor.** This is the reliable approach and covers most use cases.
      
      > **Known limitation:** The MCP tool supports `filters` and `exclusions` parameters, but the tool's schema describes the wrong format for them. Until this is fixed ([K2-269](https://linear.app/montecarlodata/issue/K2-269)), **do not pass `filters` or `exclusions`** — they will cause errors. Use database/schema scoping instead to narrow the set of monitored tables. If the user needs regex or pattern-based filtering, explain this limitation and suggest either (a) using schema-level scoping to get close, or (b) creating individual metric monitors for specific tables.
      
      ### Database-Level Selection
      
      To monitor all tables in an entire database, specify only the database name with no `schemas` list:
      
      ```json
      {
        "databases": [
          {"name": "analytics"}
        ]
      }
      ```
      
      This monitors every table in every schema within the `analytics` database.
      
      ### Schema-Level Selection
      
      To monitor all tables in specific schemas, include the `schemas` list:
      
      ```json
      {
        "databases": [
          {
            "name": "analytics",
            "schemas": ["core", "staging"]
          }
        ]
      }
      ```
      
      This monitors every table in the `core` and `staging` schemas within `analytics`, but not tables in other schemas.
      
      ### Multiple Databases
      
      You can monitor tables across multiple databases in a single monitor:
      
      ```json
      {
        "databases": [
          {"name": "analytics", "schemas": ["core"]},
          {"name": "raw_data"},
          {"name": "reporting", "schemas": ["public", "internal"]}
        ]
      }
      ```
      
      ---
      
      ## Alert Conditions
      
      Alert conditions define which metrics the table monitor tracks. The operator is always AUTO (anomaly detection) -- custom thresholds are not available for table monitors.
      
      | Metric | Description |
      |--------|-------------|
      | `last_updated_on` | Freshness monitoring. Alerts when a table has not been updated within its normal cadence. |
      | `schema` | Any schema change. Alerts when columns are added, removed, or their types change. |
      | `schema_fields_added` | New columns detected. Alerts only when new columns appear in the table. |
      | `schema_fields_removed` | Columns removed. Alerts only when existing columns are dropped from the table. |
      | `schema_fields_type_change` | Column type changes. Alerts only when a column's data type changes. |
      | `total_row_count` | Row count changes. Alerts on anomalous changes in total row count. |
      | `total_row_count_last_changed_on` | Time since last volume change. Alerts when the row count has not changed for an unusual duration. |
      
      ### Notes
      
      - **All operators are AUTO (anomaly detection).** Table monitors do not support custom thresholds like GT, LT, or explicit operators. If the user needs custom thresholds, use a metric monitor instead.
      - **No `schedule` field.** Table monitors do not support the `schedule` field in MaC YAML. Adding it will cause a validation error on `montecarlo monitors apply`. Table monitor scheduling is managed automatically by Monte Carlo. Do NOT add a schedule block to the generated YAML.
      - The default set (`last_updated_on`, `schema`, `total_row_count`, `total_row_count_last_changed_on`) provides broad coverage and is appropriate for most use cases. Only override the defaults when the user specifically requests a subset.
      - `schema` is a superset of `schema_fields_added`, `schema_fields_removed`, and `schema_fields_type_change`. If using `schema`, there is no need to also include the granular schema metrics.
      
      ---
      
      ## Examples
      
      ### Monitor all tables in a database (minimal config)
      
      ```json
      {
        "name": "analytics_db_monitor",
        "description": "Monitor all tables in the analytics database for freshness, schema changes, and volume",
        "warehouse": "production_warehouse",
        "asset_selection": {
          "databases": [
            {"name": "analytics"}
          ]
        }
      }
      ```
      
      Uses the default alert conditions (`last_updated_on`, `schema`, `total_row_count`, `total_row_count_last_changed_on`).
      
      ### Monitor specific schemas with default alerts
      
      ```json
      {
        "name": "core_schemas_monitor",
        "description": "Monitor all tables in core and reporting schemas",
        "warehouse": "production_warehouse",
        "asset_selection": {
          "databases": [
            {
              "name": "analytics",
              "schemas": ["core", "reporting"]
            }
          ]
        }
      }
      ```
      
      Monitors every table in the `core` and `reporting` schemas, leaving other schemas unmonitored.
      
      ### Monitor multiple schemas across databases
      
      ```json
      {
        "name": "prod_tables_monitor",
        "description": "Monitor production tables across analytics and raw_data databases",
        "warehouse": "production_warehouse",
        "asset_selection": {
          "databases": [
            {
              "name": "analytics",
              "schemas": ["core", "reporting"]
            },
            {
              "name": "raw_data",
              "schemas": ["ingestion"]
            }
          ]
        }
      }
      ```
      
      Monitors tables in specific production schemas, leaving development and staging schemas unmonitored.
      
      ### Schema change monitoring only
      
      ```json
      {
        "name": "warehouse_schema_watch",
        "description": "Track schema changes across the entire data warehouse",
        "warehouse": "production_warehouse",
        "asset_selection": {
          "databases": [
            {"name": "analytics"},
            {"name": "raw_data"}
          ]
        },
        "alert_conditions": [
          "schema_fields_added",
          "schema_fields_removed",
          "schema_fields_type_change"
        ]
      }
      ```
      
      Monitors only schema changes (not freshness or volume) across multiple databases. Uses the granular schema metrics instead of `schema` to allow selectively enabling/disabling each type.
      
      ---
      
      ## Table Monitor vs Metric Monitor
      
      | Aspect | Table Monitor | Metric Monitor |
      |--------|---------------|----------------|
      | **Scope** | Multiple tables (database/schema level) | Single table |
      | **Metrics** | Freshness, schema changes, row count | Field-level metrics (null rate, mean, sum, etc.) |
      | **Operator** | AUTO only (anomaly detection) | AUTO or explicit thresholds (GT, LT, EQ, etc.) |
      | **Asset selection** | Database/schema with filters and exclusions | Single table specified by MCON or name |
      | **Timestamp field** | Not required | Required (`aggregate_time_field`) |
      | **Segmentation** | Not available | Available via `segment_fields` |
      | **Best for** | Broad coverage, freshness, schema drift | Targeted field-level data quality checks |
      
      **Rule of thumb:** If the user wants to monitor a specific field on a specific table with specific thresholds, use a metric monitor. If the user wants broad monitoring across many tables with automatic anomaly detection, use a table monitor.
      
    • data-validation-monitor.md 19.3 KB
      # Validation Monitor Reference
      
      Detailed reference for building `create_or_update_validation_monitor` tool calls. The tool follows the **two-call preview-then-confirm pattern** — see `data-monitor-creation.md` for the full flow.
      
      ## Critical Constraints
      
      - **NEVER guess column names.** Always get them from `get_table`. Every field referenced in a validation condition must exist in the table schema exactly as spelled.
      - **IMPORTANT: Conditions match INVALID data, not valid data.** The monitor alerts when it finds rows matching the condition, so the condition must describe the BAD rows. Getting this backwards is the number one mistake with validation monitors.
      - **NEVER put a SELECT statement in a condition-level `SQL` node.** `{"type": "SQL", "sql": "..."}` as a top-level condition must be a boolean predicate expression (e.g. `amount < 0 OR amount > 1e9`), not a full query. Backend error: `Invalid SQL expression. Please provide a direct expression; it shouldn't begin with SELECT.`
      - **NEVER use an aggregate or SQL expression in a `FIELD` value.** `{"type": "FIELD", "field": "COUNT(*)"}` is rejected as `Field "COUNT(*)" doesn't exist`. Fields are column names only — use `get_table` to list valid ones. For counts/aggregates, fall back to a custom SQL monitor.
      - **NEVER put a `SQL` value on the LEFT side of a BINARY condition.** Only `FIELD` references are allowed on the left. A `SQL` value is valid only on the right side (typically as a scalar subquery). Backend error: `Filter left side value must be a field or map key: FilterValueSql(...)`.
      - **`alert_condition` is a dict (JSON object), NEVER a JSON-encoded string.** Pass the condition tree as a structured object — `{"type": "GROUP", "operator": "AND", "conditions": [...]}`. Serializing it to a string first is rejected with `Input should be a valid dictionary [type=dict_type, input_value='{"conditions":[...]'...]`.
      
      ---
      
      ## When to Use
      
      Use a validation monitor when the user wants to:
      
      - Check that specific fields are never null
      - Validate that values are within an allowed set (e.g., status in 'active', 'pending', 'inactive')
      - Enforce referential integrity (field values exist in another table)
      - Apply row-level business rules (e.g., "amount must be positive")
      - Combine multiple conditions with AND/OR logic
      
      ---
      
      ## Getting the Logic Right: Conditions Match INVALID Data
      
      This is the single most confusing aspect of validation monitors and the number one source of mistakes. **Conditions describe what INVALID data looks like -- the data you want to be alerted about.** They do NOT describe what valid data looks like.
      
      Think of it this way: the monitor scans rows and fires an alert when it finds rows matching the condition. So the condition must match the BAD rows.
      
      | User wants | Condition should match | Common mistake |
      |------------|----------------------|----------------|
      | "id should never be null" | id IS NULL (alert when null found) | id IS NOT NULL (would alert on every valid row) |
      | "status must be in [active, pending]" | status NOT IN [active, pending] (alert on unexpected values) | status IN [active, pending] (would alert on valid rows) |
      | "amount must be positive" | amount IS NEGATIVE (alert on bad values) | amount > 0 (would alert on valid rows) |
      | "email must not be empty" | email IS NULL **OR** email = '' (alert on missing) | email IS NOT NULL (would alert on valid rows) |
      
      **Before building any condition, ask yourself: "If a row matches this condition, is the row INVALID?" If the answer is no, the logic is backwards.**
      
      ---
      
      ## Pre-Step: Verify Field Existence
      
      Before constructing the `alert_condition`, verify that every field name you plan to reference exists in the table's column list. This is the number two source of validation monitor failures -- referencing columns that do not exist or are misspelled.
      
      1. You should already have the column list from `get_table` with `include_fields: true` (done in Step 2 of the main skill).
      2. For every field name in your planned conditions, confirm it appears in the column list exactly as spelled (field names are case-sensitive on most warehouses).
      3. If a field does not exist, stop and ask the user to clarify the correct column name. Do not guess.
      
      ---
      
      ## Required Parameters
      
      | Parameter | Type | Description |
      |-----------|------|-------------|
      | `name` | string | Unique identifier for the monitor. Use a descriptive slug (e.g., `orders_not_null_check`). |
      | `description` | string | Human-readable description of what the monitor checks. |
      | `table` | string | Table MCON (preferred) or `database:schema.table` format. If not MCON, also pass `warehouse`. |
      | `alert_condition` | object | Condition tree defining when to alert (see Alert Condition Structure below). |
      
      ## Optional Parameters
      
      | Parameter | Type | Description |
      |-----------|------|-------------|
      | `warehouse` | string | Warehouse name or UUID. Required if `table` is not an MCON. |
      | `domain_uuids` | array of string (uuid) | Domain UUIDs (use `get_domains` to list). Data monitors accept exactly one UUID in the list. |
      | `schedule_type` | string | Schedule type: `"fixed"` (default), `"dynamic"`, `"manual"`. |
      | `interval_minutes` | int | Schedule interval in minutes (only for `schedule_type="fixed"`). |
      | `audiences` | array of string | Notification audience **names** (not UUIDs) to alert when the monitor triggers. |
      | `failure_audiences` | array of string | Notification audience names to alert on query execution failures. |
      | `notes` | string | Free-text notes shown in the UI (separate from `description`). |
      | `priority` | string | Monitor priority (e.g. `"P1"`, `"P2"`). |
      | `tags` | array of `{name, value}` | Key-value tags to attach. |
      | `is_draft` | bool | When `True`, saves the monitor as a draft (not active). Default `False`. |
      | `monitor_uuid` | string (uuid) | UUID of an existing monitor to update in place. Omit to create a new monitor. **PUT semantics:** the call fully replaces the monitor's configuration — fields you omit revert to tool defaults, they are NOT left untouched. Before editing, read the current config with `get_monitors(monitor_ids=[<uuid>], include_fields=["config"])` and re-pass every field you want to keep. See `data-monitor-creation.md` (Step 7) for the safe-edit workflow. |
      | `dry_run` | bool | Default `True`. Preview mode. When omitted or `True`, returns YAML preview in `result.yaml`. When `False`, actually creates/updates the monitor and returns `result.monitor_uuid` + a deep link in `result.instructions`. See `data-monitor-creation.md`. |
      
      ---
      
      ## Alert Condition Structure
      
      The top level of `alert_condition` must always be a GROUP node. This GROUP contains one or more conditions combined with AND or OR logic.
      
      ```json
      {
        "type": "GROUP",
        "operator": "AND",
        "conditions": [...]
      }
      ```
      
      ### Condition Types
      
      There are four condition types: UNARY, BINARY, SQL, and GROUP.
      
      #### UNARY (single-value checks)
      
      Used for predicates that operate on a single field with no comparison value.
      
      ```json
      {
        "type": "UNARY",
        "predicate": {"name": "null", "negated": false},
        "value": [{"type": "FIELD", "field": "column_name"}]
      }
      ```
      
      - `predicate.name` -- the predicate to apply (see Predicates Reference below).
      - `predicate.negated` -- set to `true` to invert the predicate (e.g., `null` with `negated: true` means "is NOT null").
      - `value` -- an array with a single value descriptor (usually a FIELD reference).
      
      #### BINARY (comparison checks)
      
      Used for predicates that compare a field against a value.
      
      ```json
      {
        "type": "BINARY",
        "predicate": {"name": "greater_than", "negated": false},
        "left": [{"type": "FIELD", "field": "column_name"}],
        "right": [{"type": "LITERAL", "literal": "0"}]
      }
      ```
      
      - `left` -- the left-hand side of the comparison (typically a FIELD reference).
      - `right` -- the right-hand side (typically a LITERAL value, SQL expression, or FIELD reference).
      - Both `left` and `right` are arrays of value descriptors.
      
      #### SQL (custom SQL expression)
      
      Used for complex conditions that are difficult to express with UNARY/BINARY nodes. The SQL expression should evaluate to true for INVALID rows.
      
      ```json
      {
        "type": "SQL",
        "sql": "amount > 0 AND amount < 1000000"
      }
      ```
      
      #### GROUP (nested conditions)
      
      Used to combine multiple conditions with AND or OR logic. Groups can be nested.
      
      ```json
      {
        "type": "GROUP",
        "operator": "OR",
        "conditions": [
          {"type": "UNARY", "...": "..."},
          {"type": "BINARY", "...": "..."}
        ]
      }
      ```
      
      ---
      
      ## Value Types
      
      Value descriptors appear in the `value`, `left`, and `right` arrays of UNARY and BINARY conditions.
      
      | Type | Field | Description | Example |
      |------|-------|-------------|---------|
      | `FIELD` | `"field": "column_name"` | References a column in the table. Must be a plain column name — never an aggregate like `COUNT(*)` or a SQL snippet. | `{"type": "FIELD", "field": "user_id"}` |
      | `LITERAL` | `"literal": "value"` | A static value (always a string, even for numbers). | `{"type": "LITERAL", "literal": "100"}` |
      | `SQL` | `"sql": "..."` | A scalar SQL expression or subquery. **Right-side only** — cannot appear on the `left` of a BINARY. Valid forms: a scalar subquery (`SELECT MAX(id) FROM ref_table`) or a scalar expression. | `{"type": "SQL", "sql": "SELECT MAX(id) FROM ref_table"}` |
      
      ---
      
      ## Predicates Reference
      
      Before building conditions, call `get_validation_predicates` to get the full list of supported predicates for the connected warehouse. The list below covers common predicates but may not be exhaustive.
      
      ### Unary Predicates
      
      These predicates take no comparison value -- they check a property of the field itself.
      
      | Predicate | Description | Example use |
      |-----------|-------------|-------------|
      | `null` | Field value is null. | Alert on null ids. |
      | `is_negative` | Field value is negative. | Alert on negative amounts. |
      | `is_between_0_and_1` | Field value is between 0 and 1 (inclusive). | Alert on rates that should be percentages (0-100). |
      | `is_future_date` | Field value is a date/timestamp in the future. | Alert on future-dated records. |
      | `is_uuid` | Field value matches UUID format. | Alert on non-UUID values in a UUID field (use with `negated: true`). |
      
      ### Binary Predicates
      
      These predicates compare a field against a value.
      
      | Predicate | Right-hand side | Description | Example use |
      |-----------|----------------|-------------|-------------|
      | `equal` | Single LITERAL | Field equals the given value. | Alert when `status` equals `'deleted'`. |
      | `greater_than` | Single LITERAL | Field is greater than the given value. | Alert when `discount_pct` exceeds 100. |
      | `less_than` | Single LITERAL | Field is less than the given value. | Alert when `quantity` is below 0. |
      | `in_set` | Multiple LITERALs | Field value is in the given set. | Alert when `status` is in an invalid set (see example below). |
      | `contains` | Single LITERAL | Field value contains the given substring. | Alert when `email` contains `'test@'`. |
      | `starts_with` | Single LITERAL | Field value starts with the given prefix. | Alert when `phone` starts with `'000'`. |
      | `between` | Two LITERALs | Field value is between the two given values (inclusive). | Alert when `score` is between 0 and 10 (if that range is invalid). |
      
      ### Using `negated` to Invert Predicates
      
      Any predicate can be inverted by setting `"negated": true` in the predicate object. This is essential for "must be in set" validations:
      
      - **"status must be in [active, pending]"** becomes `in_set` with values `["active", "pending"]` and `negated: true` -- meaning "alert when status is NOT in [active, pending]".
      - **"id must not be null"** becomes `null` with `negated: false` -- meaning "alert when id IS null" (no inversion needed since the condition already matches invalid data).
      
      ### Semantic gotchas
      
      Two constraints the backend enforces that don't show up in the predicate list:
      
      - **Predicate/field-type compatibility.** Predicates have a target data type. `is_not_a_number` (NaN detection), `is_negative`, `is_between_0_and_1`, `numeric_*` — these only work on numeric columns; the backend rejects them on string/text fields with messages like `'not a number (NaN)' does not support fields of type 'string'`. Same pattern for `is_future_date` on non-timestamp columns. Use `get_validation_predicates` to confirm a predicate is supported for the target column type, and check the column's type from `get_table` before picking one.
      - **Field references on the RIGHT side of a BINARY.** The right-hand side is usually a `LITERAL` or `SQL` (subquery). You can put a `FIELD` reference on the right **only** when the left field is a date/timestamp column, OR when the operator is `in_set` (`in`). Other combinations are rejected with `Fields are only allowed on the right side for date and timestamp fields, or when using the 'in' operator`.
      
      ---
      
      ## Examples
      
      ### Alert when id is null
      
      Verify that `id` exists in the table schema from `get_table` before proceeding.
      
      ```json
      {
        "name": "orders_id_not_null",
        "description": "Alert when order id is null",
        "table": "MCON++a1b2c3d4-e5f6-7890-abcd-ef1234567890++1++1++analytics:core.orders",
        "alert_condition": {
          "type": "GROUP",
          "operator": "AND",
          "conditions": [
            {
              "type": "UNARY",
              "predicate": {"name": "null", "negated": false},
              "value": [{"type": "FIELD", "field": "id"}]
            }
          ]
        }
      }
      ```
      
      The condition matches rows where `id` IS NULL -- these are the invalid rows we want to be alerted about.
      
      ### Alert when status is not in allowed set
      
      Verify that `status` exists in the table schema from `get_table` before proceeding.
      
      ```json
      {
        "name": "orders_status_allowed_values",
        "description": "Alert when order status is outside the allowed set",
        "table": "MCON++a1b2c3d4-e5f6-7890-abcd-ef1234567890++1++1++analytics:core.orders",
        "alert_condition": {
          "type": "GROUP",
          "operator": "AND",
          "conditions": [
            {
              "type": "BINARY",
              "predicate": {"name": "in_set", "negated": true},
              "left": [{"type": "FIELD", "field": "status"}],
              "right": [
                {"type": "LITERAL", "literal": "active"},
                {"type": "LITERAL", "literal": "pending"},
                {"type": "LITERAL", "literal": "inactive"}
              ]
            }
          ]
        }
      }
      ```
      
      Note `negated: true` -- the predicate is `in_set`, but we want to alert when the value is NOT in the set. This catches any unexpected status values.
      
      ### Alert when amount is negative
      
      Verify that `amount` exists in the table schema from `get_table` before proceeding.
      
      ```json
      {
        "name": "orders_positive_amount",
        "description": "Alert when order amount is negative",
        "table": "MCON++a1b2c3d4-e5f6-7890-abcd-ef1234567890++1++1++analytics:core.orders",
        "alert_condition": {
          "type": "GROUP",
          "operator": "AND",
          "conditions": [
            {
              "type": "UNARY",
              "predicate": {"name": "is_negative", "negated": false},
              "value": [{"type": "FIELD", "field": "amount"}]
            }
          ]
        }
      }
      ```
      
      The condition matches rows where `amount` is negative -- these are the invalid rows.
      
      ### Combined conditions: null OR negative
      
      Verify that both `amount` and `quantity` exist in the table schema from `get_table` before proceeding.
      
      ```json
      {
        "name": "orders_amount_quality",
        "description": "Alert when amount is null or quantity is negative",
        "table": "MCON++a1b2c3d4-e5f6-7890-abcd-ef1234567890++1++1++analytics:core.orders",
        "alert_condition": {
          "type": "GROUP",
          "operator": "OR",
          "conditions": [
            {
              "type": "UNARY",
              "predicate": {"name": "null", "negated": false},
              "value": [{"type": "FIELD", "field": "amount"}]
            },
            {
              "type": "UNARY",
              "predicate": {"name": "is_negative", "negated": false},
              "value": [{"type": "FIELD", "field": "quantity"}]
            }
          ]
        }
      }
      ```
      
      The OR operator means an alert fires if either condition matches -- the row has a null amount OR a negative quantity.
      
      ### Between check with nested AND/OR
      
      Verify that `score` and `status` exist in the table schema from `get_table` before proceeding.
      
      ```json
      {
        "name": "records_score_validation",
        "description": "Alert when score is outside 0-100 range for active records",
        "table": "MCON++a1b2c3d4-e5f6-7890-abcd-ef1234567890++1++1++warehouse:metrics.records",
        "alert_condition": {
          "type": "GROUP",
          "operator": "AND",
          "conditions": [
            {
              "type": "BINARY",
              "predicate": {"name": "equal", "negated": false},
              "left": [{"type": "FIELD", "field": "status"}],
              "right": [{"type": "LITERAL", "literal": "active"}]
            },
            {
              "type": "BINARY",
              "predicate": {"name": "between", "negated": true},
              "left": [{"type": "FIELD", "field": "score"}],
              "right": [
                {"type": "LITERAL", "literal": "0"},
                {"type": "LITERAL", "literal": "100"}
              ]
            }
          ]
        }
      }
      ```
      
      This uses `between` with `negated: true` to alert when score is outside the 0-100 range, but only for active records (the AND operator requires both conditions to match).
      
      ### Referential integrity with SQL subquery
      
      Verify that `customer_id` exists in the table schema from `get_table` before proceeding.
      
      ```json
      {
        "name": "orders_valid_customer",
        "description": "Alert when customer_id does not exist in customers table",
        "table": "MCON++a1b2c3d4-e5f6-7890-abcd-ef1234567890++1++1++analytics:core.orders",
        "alert_condition": {
          "type": "GROUP",
          "operator": "AND",
          "conditions": [
            {
              "type": "SQL",
              "sql": "customer_id IS NOT NULL AND customer_id NOT IN (SELECT id FROM analytics.core.customers)"
            }
          ]
        }
      }
      ```
      
      The SQL condition type is useful for referential integrity checks that require subqueries. The `customer_id IS NOT NULL` guard avoids alerting on null values (which should be caught by a separate null check if needed).
      
      ### Contains and starts_with checks
      
      Verify that `email` and `phone` exist in the table schema from `get_table` before proceeding.
      
      ```json
      {
        "name": "contacts_format_validation",
        "description": "Alert when email contains test data or phone has invalid prefix",
        "table": "MCON++a1b2c3d4-e5f6-7890-abcd-ef1234567890++1++1++warehouse:crm.contacts",
        "alert_condition": {
          "type": "GROUP",
          "operator": "OR",
          "conditions": [
            {
              "type": "BINARY",
              "predicate": {"name": "contains", "negated": false},
              "left": [{"type": "FIELD", "field": "email"}],
              "right": [{"type": "LITERAL", "literal": "@test.example.com"}]
            },
            {
              "type": "BINARY",
              "predicate": {"name": "starts_with", "negated": false},
              "left": [{"type": "FIELD", "field": "phone"}],
              "right": [{"type": "LITERAL", "literal": "000"}]
            }
          ]
        }
      }
      ```
      
      ---
      
      ## Fallback: Custom SQL Monitor
      
      If `create_or_update_validation_monitor` fails -- for example because a referenced column does not exist yet in the live table, or the warehouse does not support a particular predicate -- fall back to `create_or_update_sql_monitor` with an explicit SQL query instead.
      
      A custom SQL monitor lets you express any validation logic as a SQL query that returns rows or a count. This is always available as a backup when the structured validation condition tree cannot express what you need or encounters an API error.
      
      When falling back:
      
      1. Translate the intended validation logic into a SQL query.
      2. The SQL should select rows that violate the rule (matching the same "conditions match INVALID data" principle).
      3. Use `create_or_update_sql_monitor` with the translated query.
      4. Inform the user that you used a custom SQL monitor as a fallback and explain why.
      
  • README.md 3.9 KB
    # Monte Carlo Monitoring Advisor Skill
    
    Analyze data coverage, create monitors for warehouse tables and AI agents. Walks users through warehouse discovery, use-case exploration, coverage gap analysis, data monitor creation, and agent observability — all through natural conversation. This single skill handles all monitoring needs: coverage analysis, data quality monitors (metric, validation, custom SQL, comparison, table), and AI agent monitors (metric, evaluation, trajectory, validation).
    
    ## Editor & Stack Compatibility
    
    The skill works with any AI editor that supports MCP and the Agent Skills format — including Claude Code, Cursor, and VS Code.
    
    All warehouses supported by Monte Carlo work with the monitoring advisor. The skill validates table and column references against your actual warehouse schema via the Monte Carlo API.
    
    ## Prerequisites
    
    - Claude Code, Cursor, VS Code or any editor with MCP support
    - Monte Carlo account with Editor role or above
    - [MC CLI](https://docs.getmontecarlo.com/docs/using-the-cli) installed for monitor deployment (`pip install montecarlodata`)
    - All monitor creation capabilities are built in — no additional skills needed
    
    ## Setup
    
    ### Via the mc-agent-toolkit plugin (recommended)
    
    Install the plugin for your editor — it bundles the skill, hooks, MCP server, and permissions automatically. See the [main README](../../README.md#installing-the-plugin-recommended) for editor-specific instructions.
    
    ### Standalone
    
    1. Configure the Monte Carlo MCP server:
       ```
       claude mcp add --transport http monte-carlo-mcp https://mcp.getmontecarlo.com/mcp
       ```
    
    2. Install the skill:
       ```bash
       npx skills add monte-carlo-data/mc-agent-toolkit --skill monitoring-advisor
       ```
    
    3. Authenticate: run `/mcp` in your editor, select `monte-carlo-mcp`, and complete the OAuth flow.
    
    4. Verify: ask your editor "Test my Monte Carlo connection" — it should call `test_connection` and confirm.
    
    <details>
    <summary>Legacy: header-based auth (for MCP clients without HTTP transport)</summary>
    
    If your MCP client doesn't support HTTP transport, use `.mcp.json.example` with `npx mcp-remote` and header-based authentication. See the [MCP server docs](https://docs.getmontecarlo.com/docs/mcp-server) for details.
    
    </details>
    
    ## How to use it
    
    Ask your AI editor about your monitoring coverage — describe what you want to understand or protect. The skill guides the agent through warehouse discovery, use-case analysis, coverage gap identification, and monitor creation. No special commands needed.
    
    ### Example prompts
    
    - "What are my coverage gaps?"
    - "Show me my use cases and what's monitored"
    - "Which tables should I monitor first?"
    - "Analyze monitoring coverage for my warehouse"
    - "Find unmonitored tables with recent anomalies"
    - "Help me set up monitoring for my critical use cases"
    - "Create a freshness monitor on the orders table"
    - "Set up a null check on the email column"
    - "Monitor my AI agent's latency and token usage"
    - "Track my agent's response quality"
    
    ### What it does
    
    1. **Discovers** your warehouses, use cases, and AI agents
    2. **Analyzes** coverage — which tables are monitored, which aren't, and which have active anomalies
    3. **Prioritizes** gaps by criticality, importance score, and anomaly activity
    4. **Creates** data quality monitors (metric, validation, custom SQL, comparison, table) with full parameter validation
    5. **Creates** AI agent monitors (metric, evaluation, trajectory, validation) for agent observability
    6. **Generates** monitors-as-code YAML ready for deployment
    
    ### Deploying generated monitors
    
    When the advisor generates a monitor, it returns MaC YAML. Deploy with:
    
    ```bash
    montecarlo monitors apply --dry-run    # preview
    montecarlo monitors apply --auto-yes   # apply
    ```
    
    Your project needs a `montecarlo.yml` config in the working directory:
    
    ```yaml
    version: 1
    namespace: <your-namespace>
    default_resource: <your-warehouse-name>
    ```
    
  • SKILL.md 20.7 KB
    ---
    name: monte-carlo-monitoring-advisor
    description: Analyze data coverage, create monitors for warehouse tables and AI agents. Covers coverage gaps, use-case analysis, data monitor creation, and agent observability.
    bucket: Monitoring
    version: 2.1.1
    ---
    
    # Monte Carlo Monitoring Advisor Skill
    
    This skill handles all monitoring requests -- coverage analysis, data monitor creation, and AI agent monitoring. It routes to the right reference file based on the user's intent.
    
    > **Monte Carlo tool routing (required):** Always call Monte Carlo MCP tools through this plugin's
    > bundled server, whose fully-qualified tool names are
    > `mcp__plugin_mc-agent-toolkit_monte-carlo-mcp__<tool>` (e.g.
    > `mcp__plugin_mc-agent-toolkit_monte-carlo-mcp__get_alerts`). Bare tool names used in this skill
    > (`get_alerts`, `search`, `get_table`, …) refer to that bundled server. If the session also has a
    > separately-configured `monte-carlo-mcp` server, do **not** route to it — it may point at a
    > different endpoint or credentials.
    
    Reference files live next to this skill file. **Use the Read tool** (not MCP resources) to access them:
    
    - Data monitor creation procedure: `references/data-monitor-creation.md` (relative to this file)
    - Agent monitor creation procedure: `references/agent-monitor-creation.md` (relative to this file)
    - Per-type references: `references/data-*.md` and `references/agent-*.md` (relative to this file)
    
    ## When to activate this skill
    
    Activate when the user:
    
    - Asks about monitoring coverage, data coverage, or coverage gaps
    - Wants to understand what's monitored vs. not in their warehouse
    - Asks about use cases, use-case criticality, or use-case analysis
    - Wants to explore their data estate and find what needs monitoring
    - Says things like "what should I monitor?", "where are my coverage gaps?", "show me my use cases"
    - Asks about unmonitored tables with anomalies or importance-based prioritization
    - Asks to create, add, or set up a monitor (e.g. "add a monitor for...", "create a freshness check on...", "set up validation for...")
    - Mentions monitoring a specific table, field, or metric
    - Wants to check data quality rules or enforce data contracts
    - Asks about monitoring options for a table or dataset
    - Requests monitors-as-code YAML generation
    - Wants to add monitoring after new transformation logic (when the prevent skill is not active)
    - Asks about monitoring AI agents, agent latency, agent token usage, or agent quality
    - Wants to set up alerts on agent behavior or execution patterns
    - Says things like "monitor my agent", "track agent latency", "alert on agent errors",
      "set up performance monitoring for my agent", or asks for an agent latency SLO
    - Asks about agent evaluation monitors, trajectory monitors, or validation monitors
    - Mentions agent observability or agent monitoring
    
    ## When NOT to activate this skill
    
    Do not activate when the user is:
    
    - Just querying data or exploring table contents
    - Triaging or responding to active alerts (use the prevent skill's Workflow 3)
    - Running impact assessments before code changes (use the prevent skill's Workflow 4)
    - Asking about existing monitor configuration (use `get_monitors` directly)
    - Editing or deleting existing monitors
    - Investigating agent alerts or agent traces (this skill creates agent monitors; investigating what they catch uses the `monte-carlo-troubleshoot-agent-traces` skill)
    
    ---
    
    ## Prerequisites
    
    - **Required:** Monte Carlo MCP server (`monte-carlo-mcp`) must be configured and authenticated
    - **Optional:** A database MCP server (Snowflake, BigQuery, Redshift, Databricks) for SQL profiling of table usage patterns
    
    ---
    
    ## Available MCP tools
    
    All tools are available via the `monte-carlo-mcp` MCP server.
    
    ### Coverage and discovery tools
    
    | Tool | Purpose |
    | --- | --- |
    | `get_warehouses` | List accessible warehouses (needed first -- `get_use_cases` requires `warehouse_id`) |
    | `get_use_cases` | List use cases with criticality, descriptions, table counts, precomputed tag names |
    | `get_use_case_table_summary` | Criticality distribution (HIGH/MEDIUM/LOW table counts) for a use case |
    | `get_use_case_tables` | Paginated tables with criticality, golden-table status, MCONs |
    | `get_monitors` | Check monitoring status on specific tables via `mcons` filter |
    | `get_asset_lineage` | Upstream/downstream dependencies for tables (takes MCONs + direction) |
    | `get_audiences` | List notification audiences |
    | `get_unmonitored_tables_with_anomalies` | Tables with muted OOTB anomalies but no monitors (takes ISO 8601 time range) |
    | `search` | Find tables by name; supports `is_monitored` filter |
    | `get_table` | Table details, fields, stats, domain membership |
    | `get_queries_for_table` | Query logs for a table (source/destination) |
    | `get_field_metric_definitions` | Available metrics per field type for a warehouse |
    | `get_domains` | List Monte Carlo domains |
    | `get_validation_predicates` | Available validation rule types |
    
    ### Data monitor creation tools
    
    All five tools follow a **two-call preview-then-confirm pattern**: the first call (with the default `dry_run=True`) returns rendered MaC YAML for review; the second call (`dry_run=False`) deploys the monitor live and returns a deep link to it. Pass `monitor_uuid` on either call to update an existing monitor in place instead of creating a new one. See `references/data-monitor-creation.md` for the full flow.
    
    | Tool | Purpose |
    | --- | --- |
    | `create_or_update_table_monitor` | Create or update a table monitor (preview YAML on `dry_run=True`, deploy on `dry_run=False`) |
    | `create_or_update_metric_monitor` | Create or update a metric monitor (preview YAML on `dry_run=True`, deploy on `dry_run=False`) |
    | `create_or_update_validation_monitor` | Create or update a validation monitor (preview YAML on `dry_run=True`, deploy on `dry_run=False`) |
    | `create_or_update_sql_monitor` | Create or update a custom SQL monitor (preview YAML on `dry_run=True`, deploy on `dry_run=False`) |
    | `create_or_update_comparison_monitor` | Create or update a comparison monitor (preview YAML on `dry_run=True`, deploy on `dry_run=False`) |
    
    ### Data product tool
    
    | Tool | Purpose |
    | --- | --- |
    | `create_or_update_data_product` | Create or update a data product — a named grouping of warehouse assets with reliability tracking (asset-footprint preview on `dry_run=True`, live create on `dry_run=False`). Used by the agent Context pillar to wrap an agent's upstream tables (see `references/agent-monitor-creation.md`) |
    
    ### Agent monitoring tools
    
    | Tool | Purpose |
    | --- | --- |
    | `get_agent_metadata` | List AI agents -- returns agent names, `agentReference` values (the `agent` arg for monitor creation), trace table MCONs, source types, backend classes (`backend_class`), and each agent's `warehouse_uuid`/`warehouse_name` (the `warehouse` arg -- show the name, pass the uuid) |
    | `get_agent_conversations` | List recent conversations for an agent (newest first; filter by errors/status/turns/tokens/duration; optional inline transcripts) |
    | `get_agent_conversation` | Retrieve one conversation's full prompt/completion thread by `conversation_id` |
    | `get_agent_traces` | List traces with per-trace workflows, tasks, models, LLM-call counts, tokens, duration, and error counts |
    | `get_agent_trace` | Inspect one execution trace's full span tree |
    | `get_agent_segments` | Enumerate the distinct `workflow` / `task` / `model` values to scope a monitor to a real segment |
    | `create_or_update_agent_metric_monitor` | Create or update monitors for quantitative span-level metrics (preview YAML on `dry_run=True`, deploy on `dry_run=False`) |
    | `create_or_update_agent_evaluation_monitor` | Create or update monitors for LLM-evaluated quality metrics (preview YAML on `dry_run=True`, deploy on `dry_run=False`) |
    | `create_or_update_agent_trajectory_monitor` | Create or update trajectory monitors for execution pattern alerts (preview YAML on `dry_run=True`, deploy on `dry_run=False`) |
    | `create_or_update_agent_validation_monitor` | Create or update validation monitors for logical assertions (preview YAML on `dry_run=True`, deploy on `dry_run=False`) |
    
    ---
    
    ## Routing
    
    When the user's request comes in, determine which workflow to follow:
    
    | User intent | Workflow |
    | --- | --- |
    | Coverage analysis, use-case exploration, "what should I monitor?" | **Coverage workflow** (below) |
    | Create a specific data monitor for a known table | **Read `references/data-monitor-creation.md`** and follow its procedure |
    | Monitor AI agents, agent latency, agent quality, agent traces | **Read `references/agent-monitor-creation.md`** and follow its procedure — propose coverage across the four POBC pillars (Performance, Output, Behavior, Context) |
    | Coverage analysis leads to monitor creation | Complete coverage workflow, then **read `references/data-monitor-creation.md`** for creation |
    
    When reading reference files, always use the **Read tool** with the path relative to this skill file.
    
    ---
    
    ## Coverage workflow
    
    This is the primary flow when the user asks about monitoring coverage, coverage gaps, or what to monitor.
    
    ### Step 1: Discover warehouses
    
    Call `get_warehouses` to list all accessible warehouses.
    
    - If **one** warehouse: select it automatically, proceed to Step 2.
    - If **multiple** warehouses: present warehouse **names** (never UUIDs) and ask the user which one to explore.
    
    ### Step 2: Discover use cases
    
    Call `get_use_cases(warehouse_id=<selected>)` to discover use cases for the chosen warehouse.
    
    - If **use cases exist** --> proceed to the **Use-case exploration** (below).
    - If **no use cases** --> proceed to the **Importance-based fallback** (below).
    
    ### Step 3: Check for database MCP (optional)
    
    Check if the user has a database MCP server available by looking for tools containing `snowflake`, `bigquery`, `redshift`, or `databricks` in the tool list. If found, note it for the SQL profiling step later. If not found, skip SQL profiling gracefully.
    
    ---
    
    ## Use-case exploration
    
    This is the primary flow when use cases are defined.
    
    ### Present use cases
    
    - Sort by criticality: **HIGH** before **MEDIUM** before **LOW**.
    - For each use case, show the **description** and explain the **reasoning for its criticality level** so the user understands why it matters.
    - Call `get_use_case_tables` with `golden_tables_only=true` and mention specific golden-table names as concrete examples. Golden tables are the last layer in the warehouse -- they feed ML models, dashboards, and reports. Explain this when relevant.
    - Use `get_asset_lineage` to explain how tables in a use case are connected and why certain tables are important (e.g. a golden table with many upstream dependencies).
    
    ### "Create a use case" requests
    
    You **cannot** create use cases -- they are generated automatically by Monte Carlo (along with their criticality), and there is no tool to author one. When the user asks to "create", "set up", or "define" a use case: briefly say so, and do NOT silently substitute monitor deployment. Then offer what you *can* do for the table(s) they named -- look up the existing use case / criticality, recommend field monitors, generate monitor previews, or analyze coverage gaps -- and act on the do-able part without expanding to sibling tables.
    
    ### Analyze coverage
    
    1. Call `get_use_case_table_summary` to show how many tables exist at each criticality level (HIGH / MEDIUM / LOW) for the use case.
    2. Call `get_use_case_tables` to obtain table MCONs, then call `get_monitors(mcons=[...])` to report how many are already monitored vs. not.
    3. **Default to HIGH + MEDIUM criticality scope.** This covers the most important tables without overwhelming the user. Do NOT ask the user which scope to use -- just proceed. If they want LOW-criticality tables included, they'll ask.
    4. You may suggest covering **multiple** use cases in one session.
    5. **Bias toward action, not questions.** When the scope is clear (HIGH + MEDIUM for the selected use case), proceed directly to generating monitor previews for all recommended monitors. Frame it as opt-out, not opt-in: "I'll generate previews for all N monitors -- tell me if you want to skip any." Do NOT ask "which would you like me to create?" one at a time -- batch them.
    
    ### Identify coverage gaps with anomaly data
    
    Use `get_unmonitored_tables_with_anomalies` to discover tables that are **not monitored** but already have muted out-of-the-box anomalies. This reveals real coverage gaps -- places where Monte Carlo detected data issues but no monitor was configured to alert anyone.
    
    - Call it with a recent time window (e.g. last 7-30 days) using ISO 8601 timestamps.
    - Results are ranked by **importance score** -- the most critical gaps appear first.
    - Each result includes a sample of anomaly events showing what types of issues were detected (freshness, volume, schema changes).
    - Use this to **prioritize** which unmonitored tables to cover first -- a table with recent anomalies is a stronger candidate than one with no activity.
    - Cross-reference with use-case data: if an unmonitored table with anomalies belongs to a critical use case, escalate its priority.
    
    ---
    
    ## Importance-based fallback
    
    When no use cases are defined, fall back to importance-based table discovery.
    
    1. **Find unmonitored tables:** Use `search(query="", is_monitored=false)` to find unmonitored tables sorted by importance.
    2. **Find tables with anomalies:** Use `get_unmonitored_tables_with_anomalies` with a recent time window (last 14-30 days) to find tables with recent anomalies but no monitors.
    3. **Inspect top candidates:** Use `get_table` to check table details, fields, and stats for the most important unmonitored tables.
    4. **Understand criticality via lineage:** Use `get_asset_lineage` with `direction="DOWNSTREAM"` to understand which tables are most connected -- a table with many downstream dependents is a stronger candidate for monitoring.
    5. **Prioritize:** Rank candidates by importance score and anomaly activity. Present the top candidates to the user with reasoning.
    
    ### Important
    
    - **Do NOT present importance scores as business criticality.** Always explain that the importance score is a *computed* metric (query frequency, downstream dependencies, usage patterns), not business-defined criticality.
    - Tell the user their account doesn't have use-case data **yet** -- use cases are generated automatically by Monte Carlo from warehouse metadata and exposed as asset tags; they are not manually configured through a UI.
    - You can still create metric, validation, and custom SQL monitors for individual tables in this mode -- you just won't use tag-based table monitors, since there are no use-case tags.
    
    ---
    
    ## SQL profiling (optional)
    
    If a database MCP server was detected in Step 3 of the coverage workflow:
    
    1. Call `get_queries_for_table` to see recent query patterns on candidate tables.
    2. Use the database MCP tools (e.g. `snowflake_query`, `bigquery_query`) to profile table usage -- identify which tables are queried most frequently, which columns are used in JOINs and WHERE clauses.
    3. Use this information to refine monitor suggestions -- heavily-queried tables with no monitors are high-priority gaps.
    
    If no database MCP is available, skip this step entirely. Do not ask the user to configure one.
    
    ---
    
    ## Pre-creation context (coverage-driven)
    
    When coverage analysis leads to monitor creation, gather this context before reading the creation reference file:
    
    1. **Dedup first.** Before generating a use-case tag monitor, call `get_monitors` with the same tag pair (and `monitor_types=["TABLE"]`) you'd put in the monitor's `asset_selection.filters`. If a monitor already covers that `(tag, domain)` scope, surface it (description, uuid) and ask whether to update it (pass its `monitor_uuid`), add one with a distinct scope, or skip -- do NOT silently re-create. The backend upserts a table monitor on its `(description, domain)`, so a same-description definition silently overwrites the prior monitor's settings.
    2. Call `get_audiences` to list notification audiences. Suggest one or more relevant audiences (match by team or use-case context) and ask the user which they want -- they can pick **one or several**. This is the **one** question to ask before generating; do NOT also ask about draft/active or schedule. Default to **draft** (`is_draft=True`); the user can flip to active after seeing the preview.
    3. When passing `audiences` or `failure_audiences`, use the audience **name/label** (not UUID), as a list -- one entry per selected audience.
    4. **Never fabricate credit costs.** Do not give a generic per-monitor or per-field MC credit rate -- cost scales with the specific spec (segmentation, schedule, field count). If a preview response includes a backend estimate (e.g. `estimated_credits.credits_per_day`), report that; otherwise decline and offer to preview a specific monitor or use case to get the real estimate.
    
    ### Use-case tag monitors
    
    The most common output of coverage analysis is a **table monitor scoped by use-case tags** via `create_or_update_table_monitor`. The `asset_selection` parameter uses this structure:
    
    ```json
    {
      "databases": ["<database_name>"],
      "schemas": ["<schema_name>"],
      "filters": [
        {
          "type": "TABLE_TAG",
          "tableTags": ["<tag_key>:<criticality>"],
          "tableTagsOperator": "HAS_ANY"
        }
      ]
    }
    ```
    
    Rules:
    - Filter `type` is **always** `TABLE_TAG` for use-case monitors.
    - `tableTagsOperator` should be `HAS_ANY`.
    - Each entry in `tableTags` is `"<tag_key>:<value>"` where the tag key is the precomputed tag name from `get_use_cases` output and the value is the criticality level in lowercase (`high`, `medium`, `low`).
    - To monitor only HIGH-criticality tables: `["tag_name:high"]`
    - To monitor MEDIUM + HIGH: `["tag_name:high", "tag_name:medium"]`
    - To monitor ALL: `["tag_name:high", "tag_name:medium", "tag_name:low"]`
    
    ### Monitor title (`description`) and reasoning (`notes`)
    
    Keep these distinct -- both are accepted by the creation tools. The backend auto-generates the monitor `name` slug; `description` is the title users see.
    
    - **`description` -- the title.** Short and scannable (≤ ~80 chars), plain English, naming the asset/use case and criticality scope. Do NOT cram reasoning here.
    - **`notes` -- the reasoning.** 1-3 sentences answering "why this monitor?", grounded in criticality, scope, and downstream impact.
    
    Example for a use-case tag monitor:
    
    - **Bad description** (this is reasoning, not a title): `"Monitor HIGH criticality tables in the Revenue Reporting use case to catch issues before they affect dashboards and financial reports."`
    - **Good description:** `"Revenue Reporting coverage -- HIGH + MEDIUM criticality tables"`
    - **Good notes** (paired): `"Covers HIGH/MEDIUM-criticality tables in the Revenue Reporting use case. Catches freshness, volume, and schema issues before they reach dashboards and financial reports."`
    
    ---
    
    ## Transient and truncate-and-reload tables
    
    Some tables show 0 rows when queried directly but have recent write activity in Monte Carlo metadata. These are **transient tables** -- fully replaced on each pipeline run (truncate-and-reload pattern). Recognize this pattern early to avoid wasting time querying empty tables.
    
    Signs of a transient table:
    - `get_table` shows a recent `last_updated_on` and high read/write activity
    - Direct SQL query returns 0 rows or all-NULL timestamp columns
    - Monte Carlo detected freshness anomalies (the table stayed empty longer than expected between loads)
    
    ---
    
    ## Graceful degradation
    
    Handle missing or unavailable tools gracefully:
    
    | Scenario | Behavior |
    | --- | --- |
    | No use cases defined | Fall back to importance-based discovery |
    | No database MCP available | Skip SQL profiling, rely on MC tools only |
    | `get_unmonitored_tables_with_anomalies` returns empty | Note that no recent anomalies were found; proceed with use-case or importance-based prioritization |
    | `get_use_case_tables` returns no tables | Note the use case has no tables; suggest exploring other use cases |
    | `get_audiences` returns empty | Inform user no audiences are configured; monitors can still be created without notification routing |
    | User has no warehouses | Inform user that no warehouses are accessible; they may need to check their Monte Carlo permissions |
    
    Never error out or stop the conversation because one tool returned empty results. Explain what happened and offer the next best path.
    
    ---
    
    ## Rules
    
    - **Never expose UUIDs, MCONs, or internal identifiers** to the user -- always use human-readable names for warehouses, audiences, use cases, and tables. Keep internal identifiers for tool calls only.
    - When the user asks about relationships between tables, use `get_asset_lineage` to fetch upstream/downstream connections and explain the data flow.
    - Be concise but thorough. Use bullet points and tables for clarity.
    - Always use **ISO 8601** format for datetime values in tool calls.
    - Never reformat YAML values returned by creation tools.
    - When passing `audiences` or `failure_audiences` to monitor creation tools, use the audience **name/label** (not UUID). The API accepts audience names.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related