Claude Skill

benchmark

Benchmark one session (or a small recent set) against the rolling average using Agent Monitor data — cost, total tokens, tool count, and workflow complexity score — and report where each metric lands as a percentile of the population. Tells you whether a session was normal, cheap

LLM Mart · 0 points · 9 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download hoangsonww-claude-code-agent-monitor-plugins_ccam-insights_skills_benchmark-83d4df5.zip · 1 KB
Part of hoangsonww/claude-code-agent-monitor — 86 skills

Install

skills CLI npx skills add https://github.com/hoangsonww/Claude-Code-Agent-Monitor/tree/master/plugins/ccam-insights/skills/benchmark
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install hoangsonww-claude-code-agent-monitor@llmmart
Git git clone https://github.com/hoangsonww/Claude-Code-Agent-Monitor.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole hoangsonww/claude-code-agent-monitor collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Benchmark

Score a session against the rolling population average and report its percentile on cost, tokens, tool count, and complexity using Agent Monitor data.

Input

The user provides: $ARGUMENTS

This may be:

  • A single session ID — benchmark that session
  • "latest" — benchmark the most recent session
  • "latest N" — benchmark the N most recent sessions, each vs the average
  • empty — benchmark the most recent session (default)

Data Sources

Endpoint Returns
GET /api/sessions?limit=N Population of sessions with cost, model, started_at, metadata (turn_count, total_turn_duration_ms) — builds the rolling baseline
GET /api/pricing/cost/{sessionId} { total_cost, breakdown:[{ input_tokens, output_tokens, cache_read_tokens, cache_write_tokens, cost }] } — the target session's cost and tokens
GET /api/workflows/{sessionId} complexity (score), stats (tool/event counts), toolFlow (distinct tools used) — the target session's tool count and complexity
GET /api/analytics avg_events_per_session, tool_usage, daily_sessions — corroborates population-level averages

Report Sections

1. Build the Baseline

Fetch the population with GET /api/sessions?limit=200 (the rolling set). For each session gather cost (GET /api/pricing/cost/{id} or the list cost field), total tokens (sum of the 4 token types from the pricing breakdown), tool count and complexity (GET /api/workflows/{id}). Compute mean, median, and standard deviation for each metric across the population.

2. Measure the Target

For the requested session, pull the same four metrics:

  • Cost — total_cost from GET /api/pricing/cost/{id}.
  • Total tokens — input + output + cache_read + cache_write summed from the breakdown.
  • Tool count — distinct/total tools from GET /api/workflows/{id} stats/toolFlow.
  • Complexity score — complexity.score from GET /api/workflows/{id}.

3. Percentile and Deviation

For each metric report the target's percentile within the population (share of sessions at or below it) and its z-score (value − mean) / stddev. Label each: below average / typical / above average / outlier (|z| > 2).

4. Verdict

State whether the session was normal overall. If it is an outlier, name which metric drove it (e.g., complexity p96, cost p91 → an unusually heavy session).

Output

  • A Markdown table: metric | session value | population mean | percentile | z-score | label.
  • Currency in USD to 4 decimals; tokens and tool counts as integers; complexity to 2 decimals.
  • Use ▲ for above-average and ▼ for below-average vs the mean.
  • One-line verdict: "Normal session" or "Outlier — driven by
  • When benchmarking multiple sessions, one row block per session plus a summary line.
  • Read-only: percentiles come only from the fetched population; never fabricate the baseline.
Files (claude-code-agent-monitor)
  • agents
    • openai.yaml 256 B
      interface:
        display_name: "Benchmark"
        short_description: "Benchmark one session (or a small recent set) against the..."
        default_prompt: "Use $benchmark to inspect CCAM data and complete this workflow safely."
      policy:
        allow_implicit_invocation: true
      
  • SKILL.md 3.3 KB
    ---
    name: benchmark
    description: >
      Benchmark one session (or a small recent set) against the rolling average using
      Agent Monitor data — cost, total tokens, tool count, and workflow complexity
      score — and report where each metric lands as a percentile of the population.
      Tells you whether a session was normal, cheap, or an outlier. Use when judging
      whether a session was typical or out of band.
    ---
    
    # Benchmark
    
    Score a session against the rolling population average and report its percentile on
    cost, tokens, tool count, and complexity using Agent Monitor data.
    
    ## Input
    
    The user provides: **$ARGUMENTS**
    
    This may be:
    - A single session ID — benchmark that session
    - "latest" — benchmark the most recent session
    - "latest N" — benchmark the N most recent sessions, each vs the average
    - empty — benchmark the most recent session (default)
    
    ## Data Sources
    
    | Endpoint | Returns |
    |----------|---------|
    | `GET /api/sessions?limit=N` | Population of sessions with `cost`, `model`, `started_at`, `metadata` (turn_count, total_turn_duration_ms) — builds the rolling baseline |
    | `GET /api/pricing/cost/{sessionId}` | `{ total_cost, breakdown:[{ input_tokens, output_tokens, cache_read_tokens, cache_write_tokens, cost }] }` — the target session's cost and tokens |
    | `GET /api/workflows/{sessionId}` | `complexity` (score), `stats` (tool/event counts), `toolFlow` (distinct tools used) — the target session's tool count and complexity |
    | `GET /api/analytics` | `avg_events_per_session`, `tool_usage`, `daily_sessions` — corroborates population-level averages |
    
    ## Report Sections
    
    ### 1. Build the Baseline
    Fetch the population with `GET /api/sessions?limit=200` (the rolling set). For each
    session gather cost (`GET /api/pricing/cost/{id}` or the list `cost` field), total
    tokens (sum of the 4 token types from the pricing breakdown), tool count and
    complexity (`GET /api/workflows/{id}`). Compute mean, median, and standard
    deviation for each metric across the population.
    
    ### 2. Measure the Target
    For the requested session, pull the same four metrics:
    - **Cost** — `total_cost` from `GET /api/pricing/cost/{id}`.
    - **Total tokens** — `input + output + cache_read + cache_write` summed from the breakdown.
    - **Tool count** — distinct/total tools from `GET /api/workflows/{id}` `stats`/`toolFlow`.
    - **Complexity score** — `complexity.score` from `GET /api/workflows/{id}`.
    
    ### 3. Percentile and Deviation
    For each metric report the target's percentile within the population (share of
    sessions at or below it) and its z-score `(value − mean) / stddev`. Label each:
    below average / typical / above average / outlier (|z| > 2).
    
    ### 4. Verdict
    State whether the session was normal overall. If it is an outlier, name which
    metric drove it (e.g., complexity p96, cost p91 → an unusually heavy session).
    
    ## Output
    
    - A Markdown table: metric | session value | population mean | percentile | z-score | label.
    - Currency in USD to 4 decimals; tokens and tool counts as integers; complexity to 2 decimals.
    - Use ▲ for above-average and ▼ for below-average vs the mean.
    - One-line verdict: "Normal session" or "Outlier — driven by <metric> (pNN)".
    - When benchmarking multiple sessions, one row block per session plus a summary line.
    - Read-only: percentiles come only from the fetched population; never fabricate the baseline.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related