Claude Skill

token-doctor

Personal diagnosis of where your Claude Code + Cowork spend goes. Reads local transcripts, prints your conversation length distribution, marathon share, cache rebuild costs, and per-project diagnosis (good projects and problem projects) right in the terminal. Then offers a deeper

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download techwolf-ai-ai-first-toolkit-plugins_ai-adoption_skills_token-doctor-2ee7841.zip · 34 KB
Part of techwolf-ai/ai-first-toolkit — 20 skills

Install

skills CLI npx skills add https://github.com/techwolf-ai/ai-first-toolkit/tree/main/plugins/ai-adoption/skills/token-doctor
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install techwolf-ai-ai-first-toolkit@llmmart
Git git clone https://github.com/techwolf-ai/ai-first-toolkit.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole techwolf-ai/ai-first-toolkit collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Token Doctor

Platforms: Claude Code / Cowork and Codex. scripts/inventory.py detects the host (via the platform stamp install.sh writes, or AI_FIRST_PLATFORM) and routes: Claude Code (~/.claude/projects) + Cowork transcripts, or Codex rollouts (~/.codex/sessions). Codex token usage comes from Codex's own per-response token_count events; cost uses OpenAI list rates in pricing.py (gpt-5.4 family; unknown models show token counts with no fabricated cost). Antigravity is unsupported: its IDE store is AEAD-encrypted at rest and its CLI store has no parseable turn/token content, so the skill prints a clear "not available" message and exits.

Two-stage diagnostic. Stage 1 is fast and lands directly in the terminal so the user always walks away with their numbers. Stage 2 is opt-in, fans out subagents over hotspots, and writes a tight Markdown report.

Read this whole file before running.

When to run

Trigger phrases: "diagnose my Claude habits", "why am I spending so much", "where are my tokens going", "audit my spend", "token doctor", "what's driving my Claude bill".

Do NOT trigger for:

  • "what tasks do I do with Claude" → that's task-profile.
  • "where did I work on X" → that's session-search.

The line is: token-doctor is about cost shape, not task inventory or recall.

Prerequisites

  • Claude Code transcripts: ~/.claude/projects/*/*.jsonl (CLI and desktop app)
  • Sub-agent transcripts: ~/.claude/projects/*/<sid>/subagents/**/*.jsonl, including workflow agents under subagents/workflows/<wf>/
  • Claude Cowork transcripts (optional): ~/Library/Application Support/Claude/local-agent-mode-sessions/*/*/local_*/audit.jsonl
  • Python 3, stdlib only. No external services.

If neither path exists, stop and say so.

What counts as one session

Sub-agent transcripts are separate files but the same piece of work, so their cost rolls into the parent session rather than appearing as sessions of their own. This matters when you read the report:

  • cost_usd is main conversation plus fan-out. main_cost_usd and subagent_cost_usd split it.
  • Turn counts, the timeline, cache rebuilds and the re-read ratio are main-session figures. They describe how that one conversation's context grew; sub-agents have their own context.
  • So a session can show 100 turns and $641. That is not a contradiction, it is fan-out. Say so rather than letting the reader trip over it.
  • One turn = one assistant message, not one transcript line. Claude Code writes a line per content block (thinking, text, each tool_use) and repeats the same usage object on every one, so counting lines inflates turns and cost by roughly 2-3x. The inventory dedupes by message.id. Turn counts from an older run of this skill are not comparable with these.

Automation excluded by default: sdk-cli background dispatch, paperclip, ditto-routines, scheduled tasks, and the automation slash commands. Desktop-app sessions are interactive work and are counted.


STAGE 1 — Fast diagnosis (always runs, you write the report)

The goal is: the user invokes the skill, sees a clean doctor's report within 10 seconds, knows which projects are healthy and which are bleeding, and can decide whether to go deeper. You write the report directly in your message based on the JSON the scripts produce. The scripts compute, you communicate.

Step 1.1 — Inventory (deterministic)

~/.claude/skills/token-doctor/scripts/inventory.py --since YYYY-MM-DD --out out/sessions.jsonl

Default window: last 90 days. Flags: --since, --until, --all, --include-automation, --no-cowork. Automation runs (sdk-cli background dispatch, paperclip, /loop, /schedule, ditto-routines, scheduled-tasks) excluded by default.

--until YYYY-MM-DD means midnight at the start of that day, so it excludes that day's sessions. To include today, leave --until off.

The scan prints how many sub-agent transcripts it rolled into parent sessions. If it also reports sub-agents with no parent session on disk, that cost is not in the totals; mention it only if it is material.

Step 1.2 — Aggregate (deterministic)

~/.claude/skills/token-doctor/scripts/personal_stats.py --in out/sessions.jsonl --out out/user-stats.json

Prints a single confirmation line. The full data is in out/user-stats.json.

Step 1.3 — Read the JSON and write the doctor's report

Read out/user-stats.json. Then write the report directly in your message as terminal-style ASCII with emojis. The user reads your message; no intermediate file.

Report structure (mandatory sections, in order)

🩺 ┌────────────────────────────────────────────────────────────────────┐
   │            TOKEN DOCTOR · personal diagnosis                       │
   └────────────────────────────────────────────────────────────────────┘

  Patient: <user's first name or "you">
  Window:  <window dates from inventory>
  Spend:   $<total> list-price equivalent · <conv count> conversations

  ── Vital signs ─────────────────────────────────────────────────────────

  🔴/🟡/🟢 Marathon (≥300 turns)    <N> conv  ·  $<X>  ·  <Y>% of spend
  🔴/🟡/🟢 Fan-out (≥5 sub-agents)  <N> conv  ·  $<X>  ·  <Y>% of spend
  🔴/🟡/🟢 Zombie (≥4h wall clock)  <N> conv  ·  $<X>  ·  <Y>% of spend
  🔴/🟡/🟢 Cache rebuilds            <N> events · <Z>M tokens · ~$<X>
  🔴/🟡/🟢 Re-read ratio             <X>×   (healthy ≤15×, org avg 30×)
  📈 Peak context observed       <X>k tokens
  🌳 Sub-agent cost               $<X> of $<total>  (<Y>%) across <N> transcripts

  ── Spend by conversation length ────────────────────────────────────────

       1 to 5        <bar>   <%>   (<N> conv)
       6 to 20       <bar>   <%>   (<N> conv)
       …
       1,000+        <bar>   <%>   (<N> conv)

  ── Model mix ───────────────────────────────────────────────────────────

    <model>        $<X>   <%>   <N> turns    main $<X>/<N>t · sub $<X>/<N>t
    <model>        $<X>   <%>   <N> turns    main $<X>/<N>t · sub $<X>/<N>t
    …

  ── Diagnosis ───────────────────────────────────────────────────────────

  <2-4 sentences synthesizing the vitals into one clear picture. Lead with
  the dominant antipattern in this user's data, then the corollary cost. End
  with one line about the strongest positive signal you see.>

  ── Project chart ───────────────────────────────────────────────────────

  ✅ clean · 🏃 marathon · 🌳 fanout · 🔄 rebuilds · 🧟 zombie · ⚠️ multiple

  <emoji>  $<X>  <truncated cwd>                              <meta line>
  <emoji>  $<X>  <truncated cwd>                              <meta line>
  … up to 10-12 rows from by_cwd_top

  ── Treatment plan ──────────────────────────────────────────────────────

  💚 Keep doing:
     <bullet referencing a clean project from by_cwd_top or by_cwd_clean_deeper,
      OR a positive structural signal like a high short_share if no clean cwd is in top 12>
     <2-3 bullets total>

  🎯 Change first:
     <one concrete action tied to the biggest lever, with cited project>
     <2-3 bullets total, ordered by expected impact>

  ── Want a deeper look? ─────────────────────────────────────────────────

  <one-line question asking if they want the deep dive>

Rules when writing the report

  • Use the emojis above consistently. Box-drawing characters (─ ┌ └ │) are fine and make the report look like a medical printout.
  • Traffic-light dots: 🔴 = bad, 🟡 = watch, 🟢 = healthy. Apply the bands in the rubric below.
  • Per-project emoji must come from by_cwd_top[i].emoji in the JSON. Do not re-classify.
  • Model mix comes from model_mix, already sorted by cost. Show every model down to 1% of spend, then stop. Use readable names (claude-opus-5 → Opus 5, claude-fable-5-1 → Fable 5.1). The main / sub split is the point of the section: a model that is cheap in the main conversation and expensive across sub-agents is the clearest lever in the whole report, because sub-agent model tier is a one-line change in an Agent(...) call. Call that out when you see it.
  • Fan-out uses fanout_conv / fanout_cost / fanout_share (sessions with ≥ 5 sub-agents) and subagent_cost / subagent_share (fan-out's share of total spend). If subagent_files is 0, drop both the fan-out vital and the sub-agent line rather than printing zeroes.
  • Bars for the length distribution: build them with █ characters proportional to the share. Use a fixed width like 36 chars.
  • The diagnosis paragraph is yours to write — it is the doctor's read on the data. Be specific. Don't restate the numbers; conclude from them. Aim for 3-5 sentences max. Examples of good diagnostic sentences:
    • "Your spend is concentrated in a small number of very long sessions: 22 conversations carry 70% of your bill."
    • "Cache rebuilds are minor at $388, but the re-read ratio of 23× tells me your context grows fast inside those long sessions."
    • "Three of your top five projects are evaluation runs from last month. Each is one long session; splitting them would compound."
  • Treatment plan must include both "keep doing" AND "change first" sections. Skipping the positive section is forbidden. Pick from:
    • by_cwd_top entries with emoji == "✅" for clean
    • by_cwd_clean_deeper for clean projects below the top 12
    • short_share if neither is available — frame as "X% of your spend is in short focused sessions, so the habit is there, you just don't use it everywhere"
  • Cite specific projects. Truncate cwds to the last 36-44 chars with a leading … if they're long. Drop the /Users/<name>/ prefix when it makes the line cleaner.
  • No em-dashes. Use commas, semicolons, or periods.
  • No "waste", "burning", "bad habit". Use "cost", "spend", "context", "rebuild".

Traffic-light thresholds

Metric 🟢 🟡 🔴
Marathon share < 20% 20-50% ≥ 50%
Fan-out share < 20% 20-50% ≥ 50%
Zombie share < 20% 20-50% ≥ 50%
Cache rebuild $ < $50 $50-$200 ≥ $200
Re-read ratio ≤ 15× 15-35× ≥ 35×

Asking about the deep dive

End with one short line, not a paragraph. Example:

Want me to pull apart your top sessions one by one — what specifically drove each marathon, plus a couple of your most efficient runs to learn from? Takes about a minute, runs ~15 Haiku subagents in parallel.

If they say no, stop. The report is the deliverable.


STAGE 2 — Deep dive (opt-in)

Step 2.1 — Pick hotspots

~/.claude/skills/token-doctor/scripts/pick_hotspots.py --in out/sessions.jsonl --out out/hotspots.json

Selects ~14 sessions:

  • 8 by absolute cost
  • 3 by cache rebuild tokens
  • 3 by raw turn count
  • 3 by read:create ratio at ≥$5 cost
  • 3 positive examples (lowest cost-per-turn at 20-100 turns)

Briefly tell the user the list before fan-out so they can drop sensitive sids. Keep it to one line per session: $X · N turns · short title.

Step 2.2 — Build payloads

~/.claude/skills/token-doctor/scripts/build_payloads.py --sessions out/sessions.jsonl --hotspots out/hotspots.json --outdir out/payloads

Writes one redacted payload per hotspot. Payloads include token shape, tool-call counts, timeline samples, and a 120-char title. They do NOT include user prompt bodies or model output text.

Step 2.3 — Parallel subagent fan-out

For each payload in out/payloads/, dispatch one subagent. Send all calls in one message with multiple tool blocks so they run in parallel.

Agent(
  description="Diagnose one session",
  subagent_type="general-purpose",
  model="haiku",
  run_in_background=true,
  prompt="""
Diagnose the session at out/payloads/<sid>.json.

Read first, in order:
  1. ~/.claude/skills/token-doctor/references/antipattern-taxonomy.md
  2. ~/.claude/skills/token-doctor/references/diagnosis-rubric.md
  3. out/payloads/<sid>.json

Apply the rubric. Emit strictly the JSON schema (see rubric §Output) to out/analyses/<sid>.json. Hard length limits: what_happened max 2 sentences (~25 words), trigger_moment.what max 14 words, would_have_helped max 18 words. Lead with structural facts (turn count, context size, key signal). No restating the schema.

Tone: second-person, neutral, no "waste" / "burning".
"""
)

Wait for all to complete.

Step 2.4 — Synthesize the report (main agent)

Read every out/analyses/*.json. Then:

  1. Group by cwd. For each cwd that has ≥2 analyzed sessions, decide if it shows a dominant pattern.
  2. Find the signature: the one antipattern that recurs most across the user's data, and the one positive habit they have consistently.
  3. Write out/recommendations.md — tight, scannable, no padding. Structure:
# Token Doctor — your diagnosis

**Signature.** <one sentence: dominant antipattern + dominant strength>

**Bottom-line lever.** <one sentence: the single habit change with the biggest expected impact>

## What you're doing well

- **<positive pattern>** in `<cwd>`. <one sentence with one cited sid>
- **<positive pattern>** in `<cwd>`. <one sentence with one cited sid>

(2-3 bullets. At least one is mandatory; do not skip this section.)

## What's driving your bill

- **<antipattern>** in `<cwd>`. <one sentence. Cite the worst sid and one specific turn or signal>
- ...

(3-5 bullets ordered by estimated savings)

## Per-session diagnoses

| Cost | Turns | Verdict | What happened |
|---:|---:|---|---|
| $XXX | NNN | <label> | <1-2 line what_happened from the analysis> |
| ...

(Only the analyzed sessions. Use the `what_happened` field verbatim from each analysis JSON.)

Also write out/recommendations.json with structured form for re-use:

{
  "signature": {
    "primary_antipattern": "marathon | drip-feed | zombie | bloat | grind | drift | fanout | none",
    "primary_strength": "focused | front-loaded | time-bounded | lean | directed | none",
    "one_line": "<one sentence about how this user's spend is shaped>"
  },
  "bottom_line_lever": "<one sentence>",
  "positives": [{"pattern": "...", "cwd": "...", "evidence_sid": "...", "note": "..."}],
  "antipatterns": [{"pattern": "...", "cwd": "...", "evidence_sids": ["..."], "note": "...", "expected_impact": "small|medium|large"}],
  "session_table": [{"sid": "...", "cost": 0, "turns": 0, "verdict": "...", "what_happened": "..."}]
}

Keep the prose terse. The user already has the terminal numbers; the report's job is to point at specific projects and habits, not to recite stats.


Privacy contract

  • Everything runs locally. No transcript text leaves the machine.
  • Subagents receive token counts, tool-call names, turn indices, and a 120-char title. They do not receive user prompt text or model output text.
  • Output files in out/ contain session ids and short titles. They do not contain conversation content.
  • The user can rm -rf out/ to wipe everything.

Tone

  • Descriptive, not punitive. The user is reading their own data.
  • Both antipatterns and positive habits get airtime. Skipping the "what you're doing well" section is forbidden.
  • Specific numbers, specific turn indices, specific sids. Avoid hedging.
  • No em-dashes. No "waste", "burning", "bad habit".
  • Emojis are allowed in the terminal output (the personal_stats.py script uses them). Keep them out of the Markdown report — there it should look like an engineering doc.

Failure modes

  • No transcripts found. Stop with a clear message.
  • Subagent emitted invalid JSON. Skip that sid, log a warning once, continue.
  • All sessions are automation. Tell the user to re-run with --include-automation if they want those analyzed; otherwise note the interactive count.
  • Single-cwd user. Skip the per-cwd grouping in the report; recommendations still work as a flat list.
Files (ai-first-toolkit)
  • references
    • antipattern-taxonomy.md 6.8 KB
      # Token-doctor taxonomy
      
      Seven cost-shaping patterns. Each one has a **token signature** (something we can detect from the inventory) and a **positive counterpart** (the pattern that the same person could be running instead). Diagnosis output should always cite both sides where the data supports it.
      
      The point of this file is consistency: every subagent and every recommendation uses the same names, the same definitions, and the same neutral tone.
      
      ---
      
      ## 1. Marathon
      
      **Signature**
      - `turn_count >= 300` (p95 of the org distribution; 4.4% of sessions cross this).
      - Cumulative `cache_read` per turn keeps growing because context is never reset.
      
      **Why it's expensive**
      Every turn re-pays ~10% of the entire context. At 500k context × 300 turns, that is 15B re-billed tokens. The model is doing real work on each turn but the cost grows quadratically with session length.
      
      **Positive counterpart: Focused session**
      - 20-100 turns, single task, ends cleanly with `/clear` or a session exit.
      - `cache_read` per turn is bounded because the context never balloons past what the task actually needs.
      
      **Diagnosis cue**
      A marathon is bad when the turn-by-turn `cache_read` keeps climbing AND the work covers multiple unrelated topics. A 600-turn session that stays on one cohesive refactor is appropriate, not antipattern.
      
      ---
      
      ## 2. Drip-feed (late context)
      
      **Signature**
      - One or more `cache_create > 20k tokens` events at `turn_index >= 5`.
      - Often paired with a sudden jump in `cache_read` from that turn forward.
      
      **Why it's expensive**
      Each late-arriving file forces the prefix cache to rebuild and every subsequent turn pays the bigger re-read bill. Late context tokens are billed at full input rate ($3/M for Sonnet), not the 10% cache-read rate.
      
      **Positive counterpart: Front-loaded context**
      - All relevant files dropped into turn 1 or 2 via `@file` references.
      - After turn 5, `cache_create` events are small or zero. The cache warms once and stays warm.
      
      **Diagnosis cue**
      Count the number of `>20k cache_create at turn >= 5` events. Zero is excellent. 1-2 might be unavoidable (mid-session pivot). 3+ is a habit.
      
      ---
      
      ## 3. Zombie session
      
      **Signature**
      - `duration_min >= 240` (wall-clock 4h+) AND the session resumed multiple times AND turn count is non-trivial (>30).
      - Often coupled with stale context: cache_read of last turn is comparable to earliest turns even if topic has drifted.
      
      **Why it's expensive**
      The user steps away, comes back, asks a new question, but the session is still carrying yesterday's 400k context as overhead. Every new turn pays the full context tax for context that is no longer load-bearing.
      
      **Positive counterpart: Time-bounded session**
      - Most conversations end within a single working block (<2 hours).
      - New tasks the next day start with a new session.
      
      **Diagnosis cue**
      Wall-clock alone is not the signal. A 6-hour focused refactor with related turns is fine. A 6-hour session that touches three different files for three different reasons is a zombie.
      
      ---
      
      ## 4. Context bloat
      
      **Signature**
      - `peak_cache_read >= 400_000` tokens (peak single-turn context).
      - Especially when this occurs in 50-200 turn sessions, well below the marathon threshold.
      
      **Why it's expensive**
      The model is operating near the 1M context ceiling, paying the cache-read tax on hundreds of thousands of tokens each turn. The session feels normal but every turn costs $0.30-0.60 just to re-read the context.
      
      **Positive counterpart: Lean context**
      - Peak context stays under 100k for typical sessions.
      - The user is asking targeted questions about a small slice of code, not letting the model `Read` the whole tree.
      
      **Diagnosis cue**
      Look at the trajectory of `cache_read` over turns. Steady climb past 200k without a `/clear` is bloat. Spike followed by a reset is healthy editing pattern.
      
      ---
      
      ## 5. Repeated re-reads (cache-grind)
      
      **Signature**
      - `cache_read / cache_create` ratio >= 50x at the session level, or >= 45x averaged across the user's sessions.
      - Org healthy baseline is around 10-15x. Org average is 30x.
      
      **Why it's expensive**
      A high ratio means the user is making many small turns over a large unchanging context. Each turn the model writes ~1k of output but pays for re-reading 500k of context. Tokens per useful action are skewed.
      
      **Positive counterpart: Context churn that matches activity**
      - Read 50k of new code → produce 5k of edits → ratio of around 10x.
      - Or: tight back-and-forth with small context (chat-mode) → ratio of 5-15x.
      
      **Diagnosis cue**
      This pattern hides inside long sessions and looks like "thinking out loud" or "the agent keeps asking clarifying questions." Reduce by tightening the initial prompt so the model doesn't need to re-confirm.
      
      ---
      
      ## 6. Off-task drift
      
      **Signature**
      - One session spans 3+ unrelated topics in its user-prompts.
      - Often shows as `cache_create > 20k` mid-session (a new file pulled in for the new topic) coupled with the same `session.id`.
      
      **Why it's expensive**
      Each topic-switch keeps the cumulative context of the prior topic alive. By topic 3, the session is paying for context from topics 1 and 2 with no benefit.
      
      **Positive counterpart: Clean task boundaries**
      - A new task = a `/clear` or a fresh session.
      - Each session has a coherent, single topic in its first and last user prompt.
      
      **Diagnosis cue**
      Compare the first 3 user prompts to the last 3. If they look like different conversations, that is drift. This is a behavioural signal; subagents can flag it.
      
      ---
      
      ## 7. Heavy tool fan-out
      
      **Signature**
      - Many `tool_use` calls per turn (especially `Read`) without any `Edit` or output.
      - Cache_create explodes because each `Read` pulls a new file into context.
      
      **Why it's expensive**
      The agent is reading widely without committing to a plan. Tokens go in, nothing comes out. Often the user could have pointed at the right file directly.
      
      **Positive counterpart: Directed reads**
      - The user names the relevant file or function in the first prompt.
      - The agent reads 1-3 files, makes the edit, done.
      
      **Diagnosis cue**
      Ratio of `Read` calls to `Edit/Write` calls in a session. If > 10:1 with no clear research intent, it is fan-out.
      
      ---
      
      ## Pattern colour-coding for the explorer
      
      | Pattern | Token signature in one phrase | Output colour |
      |---|---|---|
      | Marathon | Conversations beyond turn 300 | warm orange |
      | Drip-feed | Late `cache_create > 20k` events | yellow |
      | Zombie | 4h+ wall clock, multi-resume | red-pink |
      | Context bloat | Peak cache_read > 400k | purple |
      | Cache-grind | r/c ratio > 45x | warm |
      | Off-task drift | Multi-topic single session | purple |
      | Heavy fan-out | Read:edit > 10:1 | aqua-muted |
      | **Focused session** | 20-100 turns, single task | aqua |
      | **Front-loaded** | No late-create events | aqua |
      | **Time-bounded** | < 2h wall clock | aqua |
      | **Lean context** | Peak < 100k | aqua |
      
      Aqua is reserved for positive patterns. The skill should never call out a positive pattern in a punitive colour.
      
    • diagnosis-rubric.md 6 KB
      # Diagnosis rubric
      
      Decision rules a Haiku subagent applies to a single session payload. Read this **before** the payload.
      
      The output JSON must follow the schema in §Output. Be concise, neutral, and tie every claim to specific turn indices in the payload.
      
      ---
      
      ## Step 1: Classify the session shape
      
      Read the payload header. Use the numeric signals first, before the transcript.
      
      | Signal | Threshold | Label |
      |---|---|---|
      | `turn_count >= 300` AND `cost >= 50` | true | **marathon-candidate** |
      | `late_create_count >= 3` | true | **drip-feed-candidate** |
      | `duration_min >= 240` AND `resume_count >= 2` | true | **zombie-candidate** |
      | `peak_cache_read >= 400_000` | true | **bloat-candidate** |
      | `rc_ratio >= 45` | true | **grind-candidate** |
      | `tool_read_count / max(tool_edit_count, 1) >= 10` AND turns > 10 | true | **fanout-candidate** |
      | None of the above + `turn_count <= 100` + `cost <= 5` | true | **focused-candidate** |
      
      A session can carry multiple labels.
      
      ## Step 2: Confirm with transcript reading
      
      For each candidate label, look at the transcript sample to confirm or downgrade:
      
      - **marathon**: Are the user prompts about one cohesive task across the whole arc, or do they jump topics? If single task → downgrade to "long-focused-task" (positive). If multi-topic → confirm.
      - **drip-feed**: Look at the turns flagged in `late_create_events`. Were they file `@`-references the user typed at turn 50, or were they `Read` calls the model made because the user asked something new mid-flight? Either way, they were avoidable.
      - **zombie**: Compare first 3 user prompts to last 3. If topics differ → confirm. If same task throughout → downgrade to "extended-focus" (positive).
      - **bloat**: Was the peak context warranted (large codebase the user genuinely needed) or accidental (model `Read` ten irrelevant files)? Inspect the tool_calls list at the peak turn.
      - **grind**: Look at output volume across turns. Many turns producing little output = grind. Few turns producing lots of output = healthy.
      - **fanout**: Are the reads concentrated in early turns (research phase) or sprinkled throughout (lost agent)?
      
      ## Step 3: Identify the trigger moment
      
      Pick **one** specific turn where the session went wrong, if applicable. This becomes the user's "next time, don't do this" anchor.
      
      Examples of good trigger moments:
      - Turn 47: user pasted a 50k file body inline instead of using `@path/to/file`.
      - Turn 95: user said "actually let's also look at X" — should have been a new session.
      - Turn 12: model read 8 files via `Read` when the user had named just one.
      
      If the session is genuinely positive (no trigger), say so.
      
      ## Step 4: Estimate savings
      
      Conservative rule of thumb:
      - If marathon + drift confirmed: 30-50% of the session cost could have been avoided.
      - If drip-feed confirmed: 10-25% (the late-cache rebuild plus its tail of inflated re-reads).
      - If bloat confirmed: 20-40% (reading fewer files at the start).
      - If grind confirmed: 15-30% (tighter prompts mean fewer back-and-forth turns).
      - If focused/positive: 0% savings, name it as a model session to repeat.
      
      Round to nearest 5%. State your confidence.
      
      ## Step 5: Suggest one habit
      
      One sentence, second person, action-oriented. Tie it back to the trigger moment.
      
      Bad: "Try to use shorter sessions in the future."
      Good: "When you switched topics at turn 47, that was the moment to `/clear` and start fresh."
      
      Bad: "Drop fewer files."
      Good: "The 95k file pasted at turn 23 could have been an `@` reference, which keeps the prefix cache stable."
      
      For positive sessions, the "habit" becomes "what to keep doing":
      - "This is a model of a focused session: one task, front-loaded context, ended at 67 turns."
      
      ---
      
      ## Output schema
      
      Write strictly this JSON to `out/analyses/<sid>.json`:
      
      ```json
      {
        "sid": "<session id>",
        "verdict": "antipattern | positive | mixed",
        "primary_label": "marathon | drip-feed | zombie | bloat | grind | drift | fanout | focused | front-loaded | time-bounded | lean",
        "secondary_labels": ["<additional labels, ordered by severity or strength>"],
        "what_happened": "<at most 2 short sentences, ~25 words total>",
        "trigger_moment": {
          "turn": 47,
          "what": "<one short clause describing the specific user or agent action>"
        },
        "would_have_helped": "<one short sentence concrete habit, second person>",
        "estimated_savings_pct": 30,
        "confidence": "high | medium | low",
        "is_model_session": false
      }
      ```
      
      **Hard length limits.** Reports get scanned, not read.
      
      - `what_happened`: max 2 sentences. Lead with the structural fact (turn count + context size + one key signal). No restating the schema.
      - `trigger_moment.what`: max 14 words.
      - `would_have_helped`: max 18 words.
      
      Good `what_happened` examples:
      - "944 turns, 650k context never reset. 10 late-create events at turns 214-557 locked in a 65× re-read tax."
      - "Clean 67-turn session, peak 80k context, no late context. Worth repeating."
      - "Two unrelated tasks glued together starting turn 95. Topic 2 paid for topic 1's 400k context."
      
      Bad `what_happened` examples (too long):
      - "You ran a marathon session that spanned nearly 1,000 turns. During this time, the context kept growing and you didn't reset it. Each late injection at the various turns I listed forced the prefix cache to rebuild..." → over 40 words, restates obvious things.
      
      Rules:
      - `verdict = "positive"` means do NOT include `estimated_savings_pct` over 0. Use 0.
      - `is_model_session = true` only if the session is a particularly strong positive example worth pointing other team members to.
      - For mixed sessions, primary_label is the dominant pattern; put the others in secondary_labels.
      - If `trigger_moment` does not apply (positive sessions), set it to `{"turn": null, "what": "none, this session was clean"}`.
      
      ## Tone
      
      - Second person ("you", not "the user").
      - Descriptive, not punitive. The user is looking at their own data and trying to learn.
      - No "waste", "wasted", "burning", "guilty". Use "cost", "spend", "context", "rebuild".
      - Cite specific turn numbers when possible.
      - One sentence per field where the schema allows it; do not over-explain.
      
    • pricing.md 1.4 KB
      # Pricing table
      
      Per-million-token list rates from Anthropic, last verified 2026-05-12. These are the rates the deterministic cost calculator uses. Token-doctor reports list-price equivalent (your actual billing may differ if you're on a fixed-cost plan; the report makes the list-price equivalent visible).
      
      | Model family | Input ($/M) | Output ($/M) | Cache read ($/M) | Cache write 5m ($/M) | Cache write 1h ($/M) |
      |---|---:|---:|---:|---:|---:|
      | Opus 4.x (`claude-opus-4-*`) | 15.00 | 75.00 | 1.50 | 18.75 | 30.00 |
      | Sonnet 4.x (`claude-sonnet-4-*`) | 3.00 | 15.00 | 0.30 | 3.75 | 6.00 |
      | Haiku 4.x (`claude-haiku-4-*`) | 1.00 | 5.00 | 0.10 | 1.25 | 2.00 |
      
      `scripts/pricing.py` exposes a single helper:
      
      ```python
      from pricing import cost_for_turn
      
      cost = cost_for_turn(
          model="claude-sonnet-4-6",
          input_tokens=120,
          output_tokens=850,
          cache_read=42_000,
          cache_create=8_500,
      )
      ```
      
      The helper maps any model id to a family by prefix match. Unknown models fall back to Sonnet rates with a logged warning.
      
      The 1M-context tier suffix (`[1m]`) does not change the per-token rate; only the maximum window. Token-doctor treats `claude-opus-4-7[1m]` as `claude-opus-4-7`.
      
      ## Cache_create rate
      
      The default is the 5-minute ephemeral cache rate (1.25× input). If a session uses 1h caching (rare; opt-in via beta header), the cost calculator will undercount. We accept this for now; the deviation is small (under 5% for typical sessions).
      
    • recommendation-templates.md 3 KB
      # Recommendation sentence frames
      
      Templates the main agent uses when synthesizing per-project recommendations. Keep the voice consistent across the report.
      
      ## When the report describes a pattern
      
      | Pattern | Frame |
      |---|---|
      | Marathon | "Your sessions in `<cwd>` regularly cross 300 turns. Around half that time is on one task, the other half drifts into adjacent work that could have started fresh." |
      | Drip-feed | "About `<N>` of your sessions in `<cwd>` had files arrive after turn 5. Each one forced the prefix cache to rebuild and made every later turn more expensive." |
      | Zombie | "You kept `<N>` sessions in `<cwd>` alive across multiple days. The longest one ran `<X>` hours of wall clock across `<Y>` resume events." |
      | Context bloat | "Your peak context in `<cwd>` sits near `<Xk>` tokens for routine work. That is `<X×>` higher than the same task takes in other projects." |
      | Cache-grind | "In `<cwd>` you read `<X>×` more cached context than you wrote new context. That ratio means many small turns over a large frozen context." |
      | Drift | "`<N>` of your `<cwd>` sessions touched three or more unrelated topics in a single conversation." |
      | Fan-out | "When you open a session in `<cwd>`, the agent typically reads `<X>` files before making the first edit." |
      
      ## When the report celebrates a pattern
      
      | Pattern | Frame |
      |---|---|
      | Focused | "Your `<cwd>` work runs in tight 30-80 turn sessions. Median cost per session is `<$X>`, well below your average." |
      | Front-loaded | "You drop files into the first prompt in nearly all `<cwd>` sessions. Late-context events: `<N>`." |
      | Time-bounded | "Your `<cwd>` sessions almost always end within a working block. No session crossed `<X>` hours in the window." |
      | Lean | "Peak context in `<cwd>` stays around `<Xk>`. You point the agent at the exact file rather than letting it read the tree." |
      
      ## Habit suggestions (action lines)
      
      Action lines should be specific to a trigger moment if there is one. Templates:
      
      - "Next time you switch topics like at turn `<N>` of `<sid>`, run `/clear` first. New task, new conversation."
      - "The `<Xk>` file you pasted at turn `<N>` could have been `@<path>`. The prefix cache would stay warm for the rest of the session."
      - "When the agent starts reading files you did not name (like the `<N>` reads in `<sid>`), interrupt and point at the one file you actually need."
      - "If a `<cwd>` task starts feeling like yesterday's session, open a fresh one. Resume across days is the most expensive pattern in your data."
      
      ## What to avoid
      
      - Do not use words like "waste", "wasted", "burning tokens", "abuse", "guilty", "bad habit".
      - Do not compare the user to colleagues by name unless the user is in the top-5 and the comparison helps. Even then, frame as "people in your cohort" rather than calling out individuals.
      - Do not estimate dollar savings to more than two significant figures. `~$120` is fine; `$117.45` reads as fake precision.
      - Do not recommend a habit the user already practices well in another project. Cite the good example and say "you already do this in `<good cwd>`, the same approach works here".
      
  • scripts
    • build_payloads.py 6.1 KB
      #!/usr/bin/env python3
      """Generate one payload JSON per hotspot for the Haiku subagent fan-out.
      
      Each payload has:
        - header: numeric signals the subagent uses for step-1 classification
        - timeline: per-turn token + tool summary across the whole session
        - sample: a redacted sample of user prompts (first 5, every late-create context, last 3)
      
      Subagents read references/diagnosis-rubric.md and references/antipattern-taxonomy.md,
      then this payload, and emit out/analyses/<sid>.json.
      """
      from __future__ import annotations
      
      import argparse
      import json
      import re
      from pathlib import Path
      
      REDACTIONS = [
          (re.compile(r"-----BEGIN [A-Z ]+ PRIVATE KEY-----[\s\S]*?-----END [A-Z ]+ PRIVATE KEY-----"), "[REDACTED:pk]"),
          (re.compile(r"eyJ[A-Za-z0-9_\-]+\.eyJ[A-Za-z0-9_\-]+\.[A-Za-z0-9_\-]+"), "[REDACTED:jwt]"),
          (re.compile(r"sk-[A-Za-z0-9\-_]{20,}"), "[REDACTED:apikey]"),
          (re.compile(r"ghp_[A-Za-z0-9]{30,}"), "[REDACTED:gh]"),
          (re.compile(r"github_pat_[A-Za-z0-9_]{20,}"), "[REDACTED:gh]"),
          (re.compile(r"xox[baprs]-[A-Za-z0-9-]{10,}"), "[REDACTED:slack]"),
          (re.compile(r"AIza[A-Za-z0-9\-_]{30,}"), "[REDACTED:apikey]"),
          (re.compile(r"AKIA[A-Z0-9]{16}"), "[REDACTED:aws]"),
      ]
      
      def redact(s: str) -> str:
          if not s: return s
          for rx, repl in REDACTIONS:
              s = rx.sub(repl, s)
          return s
      
      def load_sessions(path: Path) -> dict[str, dict]:
          out: dict[str, dict] = {}
          with path.open() as f:
              for line in f:
                  line = line.strip()
                  if not line: continue
                  s = json.loads(line)
                  out[s["sid"]] = s
          return out
      
      def pick_sample_turns(timeline: list[dict], late_events: list[dict], max_total: int = 40) -> list[int]:
          """Return ordered turn indices to include in the sample."""
          n = len(timeline)
          if n <= max_total:
              return list(range(n))
          keep = set()
          # First 5 turns
          keep.update(range(min(5, n)))
          # Late-create events + neighbours
          for e in late_events:
              idx = e["turn"]
              keep.update([idx-1, idx, idx+1])
          # Last 3 turns
          keep.update(range(max(0, n-3), n))
          # Fill with evenly-spaced ones until we have max_total
          if len(keep) < max_total:
              remaining = [i for i in range(n) if i not in keep]
              step = max(1, len(remaining) // (max_total - len(keep)))
              keep.update(remaining[::step][:max_total - len(keep)])
          return sorted(i for i in keep if 0 <= i < n)
      
      def main():
          ap = argparse.ArgumentParser()
          ap.add_argument("--sessions", default="out/sessions.jsonl")
          ap.add_argument("--hotspots", default="out/hotspots.json")
          ap.add_argument("--outdir", default="out/payloads")
          args = ap.parse_args()
      
          sessions = load_sessions(Path(args.sessions))
          hotspots = json.loads(Path(args.hotspots).read_text())["hotspots"]
          outdir = Path(args.outdir); outdir.mkdir(parents=True, exist_ok=True)
      
          written = 0
          for h in hotspots:
              sid = h["sid"]
              s = sessions.get(sid)
              if not s: continue
              timeline = s.get("timeline", [])
              late = s.get("late_create_events", [])
      
              # Tool-call counters
              tool_counts = s.get("tool_counts", {})
              tool_read = tool_counts.get("Read", 0)
              tool_edit = tool_counts.get("Edit", 0) + tool_counts.get("Write", 0) + tool_counts.get("NotebookEdit", 0)
      
              # Resume count proxy: count timeline gaps > 60 minutes between consecutive turns
              resume_count = 0
              from datetime import datetime
              last_ts = None
              for t in timeline:
                  ts_raw = t.get("ts")
                  if not ts_raw: continue
                  try:
                      ts = datetime.fromisoformat(ts_raw.replace("Z","+00:00"))
                  except Exception:
                      continue
                  if last_ts is not None:
                      gap_s = (ts - last_ts).total_seconds()
                      if gap_s > 3600: resume_count += 1
                  last_ts = ts
      
              keep = set(pick_sample_turns(timeline, late))
              sample_timeline = [t for t in timeline if t["i"] in keep]
      
              # Sample user prompts at the same indices, if we have them
              # We didn't store user text in inventory by design (privacy). Just expose tool calls + token shape.
              # Cap title at 120 chars for the payload. Subagents only need a short tag of what the session was.
              short_title = redact(s.get("title") or "")
              if len(short_title) > 120:
                  short_title = short_title[:117] + "..."
              payload = {
                  "sid": sid,
                  "cwd": s.get("cwd"),
                  "title": short_title,
                  "surface": s.get("surface"),
                  "reasons_picked": h["reasons"],
                  "header": {
                      "turn_count": s["turn_count"],
                      # cost_usd is main + fan-out, while turn_count, by_model and the
                      # timeline below are main-session only. Without the split the
                      # subagent attributes fan-out cost to the main conversation's
                      # habits and recommends the wrong fix.
                      "cost_usd": s["cost_usd"],
                      "main_cost_usd": s.get("main_cost_usd"),
                      "subagent_cost_usd": s.get("subagent_cost_usd"),
                      "subagent_files": s.get("subagent_files"),
                      "subagent_turns": s.get("subagent_turns"),
                      "duration_min": s.get("duration_min"),
                      "resume_count": resume_count,
                      "peak_cache_read": s.get("peak_cache_read"),
                      "rc_ratio": s.get("rc_ratio"),
                      "late_create_count": s.get("late_create_count"),
                      "late_create_tokens": s.get("late_create_tokens"),
                      "tool_read_count": tool_read,
                      "tool_edit_count": tool_edit,
                      "tool_counts": tool_counts,
                  },
                  "late_create_events": late,
                  "by_model": s.get("by_model"),
                  "subagent_by_model": s.get("subagent_by_model"),
                  "sample_timeline": sample_timeline,
              }
              out_path = outdir / f"{sid}.json"
              out_path.write_text(json.dumps(payload, indent=1, default=str))
              written += 1
      
          print(f"Wrote {written} payloads to {outdir}/")
      
      if __name__ == "__main__":
          main()
      
    • codex_sessions.py 8.8 KB
      """Codex session adapter.
      
      Codex (OpenAI) stores each session as a JSONL "rollout" at
      `~/.codex/sessions/<YYYY>/<MM>/<DD>/rollout-*.jsonl`. Lines are wrapper objects
      `{type, timestamp, payload}`. The types we use:
      
        session_meta              -> payload.cwd, payload.timestamp, payload.model_provider
        turn_context              -> payload.model (the concrete model id)
        event_msg/user_message    -> payload.message|text (first one = session title)
        event_msg/agent_message   -> payload.message|text
        response_item/message     -> payload.role + payload.content ([{type,text}] or str)
      
      This adapter exposes the same shape session-search uses for Claude transcripts
      (cwd, title, plus a (role, text) iterator for grep), so routing is a drop-in.
      """
      from __future__ import annotations
      
      import json
      from pathlib import Path
      from typing import Iterator
      
      CODEX_ROOT = Path.home() / ".codex" / "sessions"
      
      
      def _content_text(content) -> str:
          if isinstance(content, str):
              return content
          if isinstance(content, list):
              parts = []
              for c in content:
                  if isinstance(c, dict):
                      t = c.get("text") or c.get("content")
                      if isinstance(t, str) and t:
                          parts.append(t)
              return "\n".join(parts)
          return ""
      
      
      def iter_turns(path: Path) -> Iterator[tuple[str, str]]:
          """Yield (role, text) for the conversational turns in a rollout.
      
          Codex records each turn twice: a clean `event_msg` (user_message /
          agent_message) and a lower-level `response_item/message` that also carries
          system/developer preamble. We use the clean channel and fall back to
          response_item only if a rollout has no event_msg turns at all.
          """
          event_turns: list[tuple[str, str]] = []
          item_turns: list[tuple[str, str]] = []
          try:
              with path.open(encoding="utf-8", errors="replace") as f:
                  for line in f:
                      try:
                          e = json.loads(line)
                      except json.JSONDecodeError:
                          continue
                      p = e.get("payload") or {}
                      if not isinstance(p, dict):
                          continue
                      typ, ptyp = e.get("type"), p.get("type")
                      if typ == "event_msg" and ptyp == "user_message":
                          txt = p.get("message") or p.get("text") or ""
                          if isinstance(txt, str) and txt:
                              event_turns.append(("user", txt))
                      elif typ == "event_msg" and ptyp == "agent_message":
                          txt = p.get("message") or p.get("text") or ""
                          if isinstance(txt, str) and txt:
                              event_turns.append(("assistant", txt))
                      elif typ == "response_item" and ptyp == "message":
                          txt = _content_text(p.get("content"))
                          if txt and p.get("role") in ("user", "assistant"):
                              item_turns.append((p.get("role"), txt))
          except OSError:
              return
          yield from (event_turns or item_turns)
      
      
      def peek(path: Path) -> tuple[str, str]:
          """Return (cwd, first_user_line) for a rollout, reading only the head."""
          cwd = ""
          first_user = ""
          try:
              with path.open(encoding="utf-8", errors="replace") as f:
                  for i, line in enumerate(f):
                      if i > 120 and cwd and first_user:
                          break
                      try:
                          e = json.loads(line)
                      except json.JSONDecodeError:
                          continue
                      p = e.get("payload") or {}
                      if not isinstance(p, dict):
                          continue
                      if not cwd and e.get("type") == "session_meta":
                          cwd = p.get("cwd", "") or ""
                      if (not first_user and e.get("type") == "event_msg"
                              and p.get("type") == "user_message"):
                          txt = p.get("message") or p.get("text") or ""
                          if isinstance(txt, str) and txt.strip():
                              first_user = txt.strip().splitlines()[0]
          except OSError:
              pass
          return cwd, first_user
      
      
      def iter_sessions() -> Iterator[dict]:
          """Yield session records shaped like session-search's Claude records."""
          if not CODEX_ROOT.is_dir():
              return
          for jsonl in CODEX_ROOT.glob("*/*/*/rollout-*.jsonl"):
              try:
                  st = jsonl.stat()
              except OSError:
                  continue
              cwd, title = peek(jsonl)
              yield {
                  "kind": "codex",
                  "mtime": st.st_mtime,
                  "size": st.st_size,
                  "title": title,
                  "cwd": cwd,
                  "path": str(jsonl),
              }
      
      
      def _usage_from_last(lt: dict) -> dict:
          """Map a Codex last_token_usage record to the Anthropic-style usage dict the
          token-doctor pipeline expects (uncached input + cache_read split; OpenAI has
          no separate cache-write metric).
          """
          inp = int(lt.get("input_tokens") or 0)
          cached = int(lt.get("cached_input_tokens") or 0)
          out = int(lt.get("output_tokens") or 0) + int(lt.get("reasoning_output_tokens") or 0)
          return {
              "input": max(inp - cached, 0),
              "output": out,
              "cache_read": cached,
              "cache_creation": 0,
          }
      
      
      def session_meta(path: Path) -> dict:
          """Return {cwd, model, start_ts} for a rollout."""
          cwd, model, start_ts = "", "", ""
          try:
              with path.open(encoding="utf-8", errors="replace") as f:
                  for i, line in enumerate(f):
                      if i > 200 and cwd and model:
                          break
                      try:
                          e = json.loads(line)
                      except json.JSONDecodeError:
                          continue
                      p = e.get("payload") or {}
                      if not isinstance(p, dict):
                          continue
                      if e.get("type") == "session_meta":
                          cwd = cwd or p.get("cwd", "") or ""
                          start_ts = start_ts or p.get("timestamp", "") or e.get("timestamp", "")
                      elif e.get("type") == "turn_context" and p.get("model"):
                          model = model or p["model"]
          except OSError:
              pass
          return {"cwd": cwd, "model": model, "start_ts": start_ts}
      
      
      def codex_turns(path: Path, model_hint: str = "") -> list[dict]:
          """Parse a rollout into turn dicts matching the token-doctor/task-profile
          shape: {role, ts, text, tool_calls, [model, usage]}.
      
          Usage is attached from each `token_count` event's per-response
          `last_token_usage` to the assistant turn it belongs to.
          """
          model = model_hint
          turns: list[dict] = []
          pending_tools: list[str] = []
          try:
              with path.open(encoding="utf-8", errors="replace") as f:
                  for line in f:
                      try:
                          e = json.loads(line)
                      except json.JSONDecodeError:
                          continue
                      p = e.get("payload") or {}
                      if not isinstance(p, dict):
                          continue
                      typ, ptyp = e.get("type"), p.get("type")
                      ts = e.get("timestamp")
                      if typ == "turn_context" and p.get("model"):
                          model = p["model"]
                      elif typ == "event_msg" and ptyp == "user_message":
                          txt = p.get("message") or p.get("text") or ""
                          if isinstance(txt, str) and txt:
                              turns.append({"role": "user", "ts": ts, "text": txt, "tool_calls": []})
                      elif typ == "event_msg" and ptyp == "agent_message":
                          txt = p.get("message") or p.get("text") or ""
                          turns.append({"role": "assistant", "ts": ts,
                                        "text": txt if isinstance(txt, str) else "",
                                        "tool_calls": pending_tools})
                          pending_tools = []
                      elif typ == "response_item" and ptyp == "function_call":
                          name = p.get("name") or p.get("tool_name")
                          if name:
                              pending_tools.append(name)
                      elif typ == "event_msg" and ptyp == "token_count":
                          lt = (p.get("info") or {}).get("last_token_usage") or {}
                          if not lt:
                              continue
                          target = next((t for t in reversed(turns)
                                         if t["role"] == "assistant" and "usage" not in t), None)
                          if target is None:
                              target = {"role": "assistant", "ts": ts, "text": "", "tool_calls": []}
                              turns.append(target)
                          if pending_tools:
                              target["tool_calls"] = target.get("tool_calls", []) + pending_tools
                              pending_tools = []
                          target["model"] = model or "gpt-5"
                          target["usage"] = _usage_from_last(lt)
          except OSError:
              return []
          return turns
      
    • host_platform.py 3.8 KB
      """Host-platform detection for the ai-adoption skills.
      
      One mechanism, shared by every script that reads agent session history. An
      identical copy ships in each skill's scripts/ dir so it travels with the script
      under all install shapes (Claude Code native, Codex flat, Antigravity nested).
      
      Resolution order (first hit wins):
        1. AI_FIRST_PLATFORM env var, if set to a known platform (explicit override).
        2. The "platform" field stamped into .techwolf-plugin.json by install.sh.
           The installer knows the target IDE (--ide codex|antigravity), so this is
           deterministic for Codex/Antigravity installs.
        3. Fallback: if ~/.claude exists, assume Claude Code. Claude Code uses the
           native plugin system and never runs install.sh, so it is never stamped.
        4. Default: "claude".
      
      Per-platform session-data reality (see each skill's SKILL.md):
        - claude       Claude Code (~/.claude/projects) + Cowork transcripts. Full.
        - codex        ~/.codex/sessions/**/rollout-*.jsonl, plaintext JSONL with
                       cwd, model, token usage, and turns. Parseable (session-search
                       routes here).
        - antigravity  IDE conversations are AEAD-encrypted at rest
                       (~/.gemini/antigravity/conversations/*.pb); the unencrypted CLI
                       store (~/.gemini/antigravity-cli/conversations/*.db) carries no
                       parseable turn/token content. No honest analysis path -> degrade.
      """
      from __future__ import annotations
      
      import json
      import os
      from pathlib import Path
      
      CLAUDE = "claude"
      CODEX = "codex"
      ANTIGRAVITY = "antigravity"
      _VALID = {CLAUDE, CODEX, ANTIGRAVITY}
      
      
      def _from_stamp() -> str | None:
          # scripts/ -> skill root holds .techwolf-plugin.json (written by install.sh).
          here = Path(__file__).resolve()
          for d in (here.parent, here.parent.parent, here.parent.parent.parent):
              stamp = d / ".techwolf-plugin.json"
              if not stamp.is_file():
                  continue
              try:
                  platform = json.loads(stamp.read_text(encoding="utf-8")).get("platform")
              except (json.JSONDecodeError, OSError):
                  return None
              if isinstance(platform, str) and platform.lower() in _VALID:
                  return platform.lower()
              return None
          return None
      
      
      def detect_platform() -> str:
          env = os.environ.get("AI_FIRST_PLATFORM", "").strip().lower()
          if env in _VALID:
              return env
          stamped = _from_stamp()
          if stamped:
              return stamped
          return CLAUDE
      
      
      _DEGRADE = {
          CODEX: (
              "this analysis is not available on Codex yet.\n"
              "  Codex sessions (~/.codex/sessions) are parseable, but this skill does not\n"
              "  read them yet. Run it under Claude Code. (session-search already supports\n"
              "  Codex.)"
          ),
          ANTIGRAVITY: (
              "session analysis is not available on Antigravity.\n"
              "  Antigravity stores IDE conversations encrypted at rest\n"
              "  (~/.gemini/antigravity/conversations/*.pb, AEAD), and its unencrypted CLI\n"
              "  store carries no parseable turn/token content. There is no honest local\n"
              "  data path to analyse. Run this skill under Claude Code."
          ),
      }
      
      
      def degrade(skill: str, platform: str | None = None) -> None:
          """Print a clear, platform-specific 'not available' message and exit 0.
      
          Degrading is not an error: the skill simply isn't available on this host.
          """
          platform = platform or detect_platform()
          print(f"{skill}: {_DEGRADE.get(platform, f'not available on {platform}.')}")
          raise SystemExit(0)
      
      
      def require_claude(skill: str) -> str:
          """Return the platform if Claude; otherwise degrade. For skills that only
          support Claude transcripts today (token-doctor, task-profile)."""
          platform = detect_platform()
          if platform == CLAUDE:
              return platform
          degrade(skill, platform)
      
    • inventory.py 20.9 KB
      #!/usr/bin/env python3
      """Walk local Claude Code + Cowork transcripts into out/sessions.jsonl.
      
      Per-session row contains everything later stages need: tokens by model, per-turn
      timeline, late-create events, peak cache_read, duration, automation flag, cwd,
      title, and computed cost.
      
      Default window: last 90 days. Override with --since YYYY-MM-DD / --until / --all.
      """
      from __future__ import annotations
      
      import argparse
      import json
      import re
      import sys
      from datetime import datetime, timedelta, timezone
      from pathlib import Path
      
      sys.path.insert(0, str(Path(__file__).parent))
      from pricing import cost_for_usage  # noqa: E402
      import codex_sessions  # noqa: E402
      from host_platform import ANTIGRAVITY, CLAUDE, CODEX, degrade, detect_platform  # noqa: E402
      
      HOME = Path.home()
      CODE_ROOT = HOME / ".claude" / "projects"
      
      def _cowork_root() -> Path:
          if sys.platform == "darwin":
              return HOME / "Library" / "Application Support" / "Claude" / "local-agent-mode-sessions"
          if sys.platform.startswith("win"):
              import os
              return Path(os.environ.get("APPDATA", str(HOME))) / "Claude" / "local-agent-mode-sessions"
          return HOME / ".config" / "Claude" / "local-agent-mode-sessions"
      
      COWORK_ROOT = _cowork_root()
      
      # Automation markers we exclude by default (long-running automated processes inflate cost)
      # Entrypoints that are background dispatch rather than a person at a keyboard.
      # Deny-list on purpose: `cli`, `claude-desktop` and any future interactive surface
      # count as spend. An unknown entrypoint is treated as interactive, because silently
      # dropping real work is the worse failure for a cost tool.
      AUTOMATION_ENTRYPOINTS = {"sdk-cli", "sdk"}
      AUTOMATION_SLASH_CMDS = {
          "/loop", "/schedule", "/babysit-prs", "/ultrareview", "/autonomous-loop",
          "/productivity:update", "/productivity:start",
      }
      TASK_NOTIF_RE = re.compile(r"<task-notification>.*?</task-notification>", re.DOTALL)
      COMMAND_WRAPPER_RE = re.compile(r"<command-(?:name|message|args)>.*?</command-\w+>", re.DOTALL)
      
      def parse_ts(s: str | None) -> datetime | None:
          if not s: return None
          try:
              if s.endswith("Z"): s = s[:-1] + "+00:00"
              return datetime.fromisoformat(s)
          except Exception:
              return None
      
      def _text_of(entry: dict) -> str:
          msg = entry.get("message") or {}
          c = msg.get("content")
          if isinstance(c, str): return c
          if isinstance(c, list):
              out = []
              for p in c:
                  if isinstance(p, dict) and p.get("type") == "text":
                      t = p.get("text") or ""
                      if t: out.append(t)
              return "\n".join(out)
          return ""
      
      def _tool_calls(entry: dict) -> list[str]:
          msg = entry.get("message") or {}
          c = msg.get("content")
          if not isinstance(c, list): return []
          return [p.get("name") for p in c if isinstance(p, dict) and p.get("type") == "tool_use" and p.get("name")]
      
      def _usage(entry: dict) -> tuple[str, dict] | None:
          msg = entry.get("message") or {}
          u = msg.get("usage")
          if not isinstance(u, dict): return None
          return msg.get("model") or "unknown", {
              "input":          int(u.get("input_tokens") or 0),
              "output":         int(u.get("output_tokens") or 0),
              "cache_read":     int(u.get("cache_read_input_tokens") or 0),
              "cache_creation": int(u.get("cache_creation_input_tokens") or 0),
          }
      
      def _strip_noise(text: str) -> str:
          return TASK_NOTIF_RE.sub("", text).strip()
      
      def _first_user_text(turns: list[dict]) -> str:
          for t in turns:
              if t["role"] == "user" and t.get("text"):
                  txt = COMMAND_WRAPPER_RE.sub("", t["text"]).strip()
                  if txt:
                      return txt[:400]
          return ""
      
      def _is_automation(path: Path, turns: list[dict], first_meta: dict) -> str | None:
          p = str(path)
          ep = first_meta.get("entrypoint") or ""
          if ep in AUTOMATION_ENTRYPOINTS: return f"entrypoint:{ep}"
          if "/agent/local_ditto_" in p or "/agent/local_routine_" in p: return "ditto-routine"
          if "--paperclip-instances-" in p: return "paperclip"
          if turns:
              first_user = next((t for t in turns if t["role"] == "user"), None)
              if first_user:
                  text = first_user.get("text") or ""
                  if "<scheduled-task" in text: return "scheduled-task"
                  m = re.search(r"<command-name>\s*(/\S+)", text)
                  if m and m.group(1) in AUTOMATION_SLASH_CMDS:
                      return f"slash-command:{m.group(1)}"
          return None
      
      def _subagent_usage(path: Path) -> dict:
          """Cost and turns for one sub-agent transcript. Usage only: no text, no timeline.
      
          One turn = one assistant message, deduped by message id, matching how the main
          transcript is counted.
          """
          agg = {"cost": 0.0, "turns": 0, "by_model": {}}
          try:
              lines = path.read_text(encoding="utf-8", errors="replace").splitlines()
          except OSError:
              return agg
          seen_msgs: set[str] = set()
          for line in lines:
              try:
                  e = json.loads(line)
              except json.JSONDecodeError:
                  continue
              if e.get("type") != "assistant":
                  continue
              # Same per-content-block duplication as the main transcript: count each
              # assistant message once, so a sub-agent turn is the same unit as a main turn.
              mid = (e.get("message") or {}).get("id")
              if mid:
                  if mid in seen_msgs:
                      continue
                  seen_msgs.add(mid)
              u = _usage(e)
              if not u:
                  continue
              model, usage = u
              cost = cost_for_usage(model, usage)
              agg["cost"] += cost
              agg["turns"] += 1
              bm = agg["by_model"].setdefault(model, {"turns": 0, "cost": 0.0})
              bm["turns"] += 1
              bm["cost"] += cost
          return agg
      
      def _rollup_subagents(paths: list[Path]) -> dict:
          """Merge every sub-agent transcript belonging to one parent session."""
          agg = {"cost": 0.0, "turns": 0, "files": 0, "by_model": {}}
          for p in paths:
              u = _subagent_usage(p)
              if not u["turns"]:
                  continue
              agg["files"] += 1
              agg["cost"] += u["cost"]
              agg["turns"] += u["turns"]
              for model, bm in u["by_model"].items():
                  tgt = agg["by_model"].setdefault(model, {"turns": 0, "cost": 0.0})
                  tgt["turns"] += bm["turns"]
                  tgt["cost"] += bm["cost"]
          return agg
      
      def _split_project(project_dir: Path) -> tuple[list[Path], dict[str, list[Path]]]:
          """Split a project dir into main session files and sub-agent files by parent sid.
      
          Main sessions are <project>/<sid>.jsonl. Sub-agents live at
          <project>/<sid>/subagents/agent-*.jsonl, and workflow agents one level deeper at
          <project>/<sid>/subagents/workflows/<wf>/agent-*.jsonl. Both shapes carry the
          parent session id as the first path component, so one rule covers them.
          """
          mains: list[Path] = []
          subs: dict[str, list[Path]] = {}
          for p in sorted(project_dir.rglob("*.jsonl")):
              parts = p.relative_to(project_dir).parts
              if len(parts) == 1:
                  mains.append(p)
              elif "subagents" in parts:
                  subs.setdefault(parts[0], []).append(p)
          return mains, subs
      
      def _process_code_transcript(path: Path, subagents: dict | None = None) -> dict | None:
          """Parse a Claude Code .jsonl transcript."""
          try:
              lines = path.read_text(encoding="utf-8", errors="replace").splitlines()
          except OSError:
              return None
          if not lines: return None
      
          sid = None
          cwd = None
          first_meta: dict = {}
          turns: list[dict] = []
          # Claude Code writes ONE JSONL LINE PER CONTENT BLOCK of an assistant message
          # (thinking, text, and each tool_use), and every one of those lines repeats the
          # SAME message.usage object. Treating each line as a turn therefore counts the
          # same tokens once per block: measured ~2.4x on line counts and up to ~3x on
          # token totals. Lines sharing a message id are merged into one turn here, so
          # cost, token totals, the timeline, the cache metrics and turn_count are all
          # per-message. Keyed by message id rather than requestId: a retried request
          # reuses the id, and counting a retry twice would be the same bug again.
          by_msg_id: dict[str, dict] = {}
      
          for line in lines:
              try:
                  e = json.loads(line)
              except json.JSONDecodeError:
                  continue
              t = e.get("type")
              if not sid:
                  sid = e.get("sessionId") or e.get("session_id")
              if not cwd:
                  cwd = e.get("cwd") or (e.get("metadata") or {}).get("cwd")
                  if cwd:
                      first_meta["cwd"] = cwd
              if "entrypoint" not in first_meta:
                  # Independent of cwd: an early entry can carry one without the other.
                  ep = e.get("entryPointType") or e.get("entrypoint")
                  if ep:
                      first_meta["entrypoint"] = ep
              if t not in ("user", "assistant"):
                  continue
              text = _strip_noise(_text_of(e))
              tools = _tool_calls(e)
              u = _usage(e)
              mid = (e.get("message") or {}).get("id") if t == "assistant" else None
              prev = by_msg_id.get(mid) if mid else None
              if prev is not None:
                  # Another content block of a message already seen: fold it in, do not
                  # re-count its usage.
                  if text:
                      prev["text"] = f"{prev['text']}\n{text}".strip() if prev["text"] else text
                  if tools:
                      prev["tool_calls"].extend(tools)
                  if u and "usage" not in prev:
                      prev["model"], prev["usage"] = u
                  continue
              ent: dict = {
                  "role": t,
                  "ts": e.get("timestamp"),
                  "text": text,
                  "tool_calls": tools,
              }
              if u:
                  ent["model"], ent["usage"] = u
              turns.append(ent)
              if mid:
                  by_msg_id[mid] = ent
      
          if not turns:
              return None
      
          # The filename is the canonical session id. A resumed or forked transcript can
          # still carry the originating sessionId on an early line, and two files sharing
          # a sid would collide on out/payloads/<sid>.json in stage 2.
          sid = path.stem or sid
          autom = _is_automation(path, turns, first_meta)
          return _summarize_session(sid=sid, surface="code", path=path, cwd=cwd, turns=turns,
                                    automation=autom, subagents=subagents)
      
      def _process_codex_transcript(path: Path) -> dict | None:
          """Parse a Codex rollout (~/.codex/sessions/.../rollout-*.jsonl)."""
          meta = codex_sessions.session_meta(path)
          turns = codex_sessions.codex_turns(path, model_hint=meta.get("model", ""))
          if not turns:
              return None
          return _summarize_session(sid=path.stem, surface="codex", path=path,
                                    cwd=meta.get("cwd") or None, turns=turns, automation=None)
      
      def _process_cowork_transcript(audit_path: Path) -> dict | None:
          """Parse a Cowork audit.jsonl + companion json sidecar."""
          try:
              lines = audit_path.read_text(encoding="utf-8", errors="replace").splitlines()
          except OSError:
              return None
          if not lines: return None
      
          # Sidecar metadata
          sidecar = audit_path.parent.parent / (audit_path.parent.name.replace("local_", "") + ".json")
          title = None
          if sidecar.exists():
              try:
                  sd = json.loads(sidecar.read_text(encoding="utf-8", errors="replace"))
                  title = sd.get("title") or (sd.get("metadata") or {}).get("title")
              except Exception:
                  pass
      
          sid = audit_path.parent.name
          turns: list[dict] = []
          for line in lines:
              try:
                  e = json.loads(line)
              except json.JSONDecodeError:
                  continue
              t = e.get("type")
              if t not in ("user", "assistant"):
                  continue
              text = _strip_noise(_text_of(e))
              ent: dict = {
                  "role": t,
                  "ts": e.get("timestamp"),
                  "text": text,
                  "tool_calls": _tool_calls(e),
              }
              u = _usage(e)
              if u:
                  ent["model"], ent["usage"] = u
              turns.append(ent)
      
          if not turns: return None
          autom = _is_automation(audit_path, turns, {})
          return _summarize_session(sid=sid, surface="cowork", path=audit_path, cwd=None,
                                    turns=turns, automation=autom, title_override=title)
      
      def _summarize_session(sid: str, surface: str, path: Path, cwd: str | None,
                             turns: list[dict], automation: str | None,
                             title_override: str | None = None,
                             subagents: dict | None = None) -> dict:
          # One model turn = one assistant MESSAGE carrying usage. The caller has already
          # merged the per-content-block lines of a message into a single turn, so this is
          # a message count, not a JSONL line count.
          model_turns = [t for t in turns if t.get("usage")]
          turn_count = len(model_turns)
          # Per-turn timeline (compact)
          timeline = []
          cre_tot = 0; read_tot = 0; inp_tot = 0; out_tot = 0
          by_model: dict[str, dict] = {}
          peak_read = 0
          late_events = []  # (turn_index, tokens)
          for i, t in enumerate(model_turns):
              u = t["usage"]
              c_read = u["cache_read"]; c_cre = u["cache_creation"]
              inp = u["input"]; out = u["output"]
              model = t.get("model") or "unknown"
              cost = cost_for_usage(model, u)
              timeline.append({
                  "i": i, "ts": t.get("ts"),
                  "model": model,
                  "inp": inp, "out": out, "read": c_read, "cre": c_cre,
                  "cost": round(cost, 6),
                  "tools": t.get("tool_calls") or [],
              })
              cre_tot += c_cre; read_tot += c_read; inp_tot += inp; out_tot += out
              peak_read = max(peak_read, c_read)
              if i >= 5 and c_cre >= 20_000:
                  late_events.append({"turn": i, "tokens": c_cre})
              bm = by_model.setdefault(model, {"input":0,"output":0,"cache_read":0,"cache_creation":0,"cost":0.0,"turns":0})
              bm["input"]          += inp
              bm["output"]         += out
              bm["cache_read"]     += c_read
              bm["cache_creation"] += c_cre
              bm["cost"]           += cost
              bm["turns"]          += 1
      
          # Total cost. Sub-agent transcripts are separate files but the same piece of work,
          # so their cost rolls into the parent session. Turn counts, the timeline and the
          # cache metrics stay main-session-only: they describe this conversation's own
          # context growth, and the length buckets downstream are main-turn buckets.
          total_cost = sum(b["cost"] for b in by_model.values())
          for bm in by_model.values():
              bm["cost"] = round(bm["cost"], 6)
          sub = subagents or {}
          sub_cost = sub.get("cost", 0.0)
          sub_by_model = {m: {"turns": v["turns"], "cost": round(v["cost"], 6)}
                          for m, v in (sub.get("by_model") or {}).items()}
      
          # Duration
          first_ts = parse_ts(turns[0].get("ts"))
          last_ts  = parse_ts(turns[-1].get("ts"))
          dur_min = None
          if first_ts and last_ts:
              try:
                  dur_min = round((last_ts - first_ts).total_seconds() / 60, 1)
              except Exception:
                  pass
      
          # Tool-call counters
          tool_counts: dict[str, int] = {}
          for t in turns:
              for n in t.get("tool_calls") or []:
                  tool_counts[n] = tool_counts.get(n, 0) + 1
      
          title = title_override or _first_user_text(turns)
          return {
              "sid": sid,
              "surface": surface,
              "path": str(path),
              "cwd": cwd,
              "title": title,
              "automation": automation,
              "turn_count": turn_count,
              "start_ts": turns[0].get("ts"),
              "end_ts": turns[-1].get("ts"),
              "duration_min": dur_min,
              "input": inp_tot, "output": out_tot,
              "cache_read": read_tot, "cache_creation": cre_tot,
              "peak_cache_read": peak_read,
              "late_create_events": late_events,
              "late_create_count": len(late_events),
              "late_create_tokens": sum(e["tokens"] for e in late_events),
              "rc_ratio": round(read_tot / max(cre_tot, 1), 2),
              "cost_usd": round(total_cost + sub_cost, 4),
              "main_cost_usd": round(total_cost, 4),
              "subagent_cost_usd": round(sub_cost, 4),
              "subagent_files": sub.get("files", 0),
              "subagent_turns": sub.get("turns", 0),
              "subagent_by_model": sub_by_model,
              "by_model": by_model,
              "tool_counts": tool_counts,
              "timeline": timeline,
          }
      
      def _within_window(sess: dict, since: datetime | None, until: datetime | None) -> bool:
          ts = parse_ts(sess.get("end_ts") or sess.get("start_ts"))
          if not ts: return True  # keep unknown timestamps
          if since and ts < since: return False
          if until and ts > until: return False
          return True
      
      def main():
          ap = argparse.ArgumentParser()
          ap.add_argument("--since", type=str, default=None, help="YYYY-MM-DD")
          ap.add_argument("--until", type=str, default=None, help="YYYY-MM-DD")
          ap.add_argument("--all", action="store_true", help="No window filter (overrides --since/--until)")
          ap.add_argument("--include-automation", action="store_true",
                          help="Include sessions flagged as automation (paperclip, schedule, ditto, /loop, etc.)")
          ap.add_argument("--out", type=str, default="out/sessions.jsonl")
          ap.add_argument("--include-cowork", action="store_true", default=True)
          ap.add_argument("--no-cowork", dest="include_cowork", action="store_false")
          args = ap.parse_args()
      
          platform = detect_platform()
          if platform == ANTIGRAVITY:
              degrade("token-doctor", platform)
      
          if args.all:
              since = until = None
          else:
              until = datetime.now(timezone.utc) if not args.until else datetime.fromisoformat(args.until).replace(tzinfo=timezone.utc)
              if args.since:
                  since = datetime.fromisoformat(args.since).replace(tzinfo=timezone.utc)
              else:
                  since = until - timedelta(days=90)
      
          out_path = Path(args.out)
          out_path.parent.mkdir(parents=True, exist_ok=True)
      
          n_in = n_out = n_autom = n_sub = n_orphan = 0
          orphan_cost = 0.0
          with out_path.open("w", encoding="utf-8") as f_out:
              if platform == CODEX:
                  # Codex rollouts (~/.codex/sessions). Token usage comes from Codex's
                  # own per-response token_count events; cost via OpenAI rates.
                  for p in sorted(codex_sessions.CODEX_ROOT.glob("*/*/*/rollout-*.jsonl")):
                      n_in += 1
                      sess = _process_codex_transcript(p)
                      if not sess: continue
                      if not _within_window(sess, since, until): continue
                      f_out.write(json.dumps(sess, default=str) + "\n")
                      n_out += 1
              else:
                  # Code transcripts. Sub-agent transcripts are rolled into their parent
                  # session rather than emitted as sessions of their own: they are the same
                  # unit of work, and counting them separately would double the session count
                  # and corrupt the turn-count length buckets.
                  if CODE_ROOT.exists():
                      for project_dir in sorted(CODE_ROOT.iterdir()):
                          if not project_dir.is_dir() or project_dir.name.startswith("_archive"):
                              continue
                          mains, subs = _split_project(project_dir)
                          rollups = {sid: _rollup_subagents(paths) for sid, paths in subs.items()}
                          for p in mains:
                              n_in += 1
                              roll = rollups.pop(p.stem, None)
                              if roll:
                                  n_sub += roll["files"]
                              sess = _process_code_transcript(p, roll)
                              if not sess: continue
                              if not _within_window(sess, since, until): continue
                              if sess.get("automation") and not args.include_automation:
                                  n_autom += 1
                                  continue
                              f_out.write(json.dumps(sess, default=str) + "\n")
                              n_out += 1
                          # Sub-agents whose parent transcript is gone have nothing to roll into.
                          for roll in rollups.values():
                              n_orphan += roll["files"]
                              orphan_cost += roll["cost"]
                  # Cowork transcripts
                  if args.include_cowork and COWORK_ROOT.exists():
                      for p in sorted(COWORK_ROOT.glob("*/*/local_*/audit.jsonl")):
                          n_in += 1
                          sess = _process_cowork_transcript(p)
                          if not sess: continue
                          if not _within_window(sess, since, until): continue
                          if sess.get("automation") and not args.include_automation:
                              n_autom += 1
                              continue
                          f_out.write(json.dumps(sess, default=str) + "\n")
                          n_out += 1
      
          print(f"Scanned {n_in} transcripts, wrote {n_out} sessions to {out_path}, excluded {n_autom} automation runs.")
          if n_sub:
              print(f"Rolled {n_sub} sub-agent transcripts into their parent sessions.")
          if n_orphan:
              print(f"Skipped {n_orphan} sub-agent transcripts (~${orphan_cost:,.0f}) with no parent session on disk.")
          print(f"Window: {since.date() if since else 'all'} to {until.date() if until else 'now'}")
      
      if __name__ == "__main__":
          main()
      
    • personal_stats.py 9.7 KB
      #!/usr/bin/env python3
      """Aggregate out/sessions.jsonl into out/user-stats.json and print a terminal summary.
      
      Reproduces the team-dashboard metrics for one person: length buckets, marathon share,
      cache rebuild $, read:create ratio, per-cwd breakdown with health classification.
      """
      from __future__ import annotations
      
      import argparse
      import json
      from collections import defaultdict
      from pathlib import Path
      
      MARATHON_THRESHOLD = 300  # p95 of org turn distribution
      FANOUT_THRESHOLD = 5      # sub-agent transcripts before a session counts as fan-out
      
      def length_bucket(n: int) -> str:
          if n <= 5:    return "b1_1_5"
          if n <= 20:   return "b2_6_20"
          if n <= 50:   return "b3_21_50"
          if n <= 100:  return "b4_51_100"
          if n <= 300:  return "b5_101_300"
          if n <= 1000: return "b6_301_1000"
          return "b7_1001_plus"
      
      def classify_cwd(c: dict) -> tuple[str, str]:
          """Return (emoji, one-word health) for a cwd given its rollup stats."""
          mara_share = c["mara_cost"] / c["cost"] if c["cost"] else 0
          rebuild_share = (c["late_tok"]/1e6 * 3) / c["cost"] if c["cost"] else 0
          fanout_share = c["sub_cost"] / c["cost"] if c["cost"] else 0
          issues = []
          if mara_share >= 0.5: issues.append("marathon")
          if fanout_share >= 0.5: issues.append("fanout")
          if rebuild_share >= 0.05: issues.append("rebuilds")
          if c["zomb"] >= max(2, c["conv"] // 3): issues.append("zombie")
          if not issues:
              return ("✅", "clean")
          if len(issues) >= 2:
              return ("⚠️ ", f"{','.join(issues)}")
          if "marathon" in issues:  return ("🏃", "marathon")
          if "fanout" in issues:    return ("🌳", "fanout")
          if "rebuilds" in issues:  return ("🔄", "rebuilds")
          return ("🧟", "zombie")
      
      def main():
          ap = argparse.ArgumentParser()
          ap.add_argument("--in", dest="inp", default="out/sessions.jsonl")
          ap.add_argument("--out", default="out/user-stats.json")
          ap.add_argument("--marathon-threshold", type=int, default=MARATHON_THRESHOLD)
          ap.add_argument("--fanout-threshold", type=int, default=FANOUT_THRESHOLD)
          args = ap.parse_args()
      
          sessions = []
          with open(args.inp) as f:
              for line in f:
                  line = line.strip()
                  if not line: continue
                  sessions.append(json.loads(line))
      
          if not sessions:
              print("🩺 Nothing to diagnose. No sessions in inventory.")
              Path(args.out).parent.mkdir(parents=True, exist_ok=True)
              Path(args.out).write_text(json.dumps({"empty": True}))
              return
      
          total_cost = sum(s["cost_usd"] for s in sessions)
          total_conv = len(sessions)
          bucket_cost = defaultdict(float)
          bucket_count = defaultdict(int)
          mara_cost = 0.0; mara_conv = 0
          zomb_cost = 0.0; zomb_conv = 0
          late_tokens = 0; late_evt = 0
          read_total = 0; cre_total = 0
          peak_max = 0
          by_cwd = defaultdict(lambda: {"cost":0.0,"conv":0,"mara_cost":0.0,"mara":0,"zomb":0,
                                        "late_tok":0,"sub_cost":0.0,"fan":0})
          by_surface = defaultdict(lambda: {"cost":0.0,"conv":0})
          sub_cost = 0.0; sub_files = 0; sub_turns = 0
          fan_cost = 0.0; fan_conv = 0
          model_mix = defaultdict(lambda: {"cost":0.0,"turns":0,"main_cost":0.0,"main_turns":0,
                                           "sub_cost":0.0,"sub_turns":0})
      
          for s in sessions:
              n = s["turn_count"]
              b = length_bucket(n)
              bucket_cost[b] += s["cost_usd"]
              bucket_count[b] += 1
              is_mara = n >= args.marathon_threshold
              is_zomb = (s.get("duration_min") or 0) >= 240
              if is_mara: mara_cost += s["cost_usd"]; mara_conv += 1
              if is_zomb: zomb_cost += s["cost_usd"]; zomb_conv += 1
              late_tokens += s.get("late_create_tokens", 0)
              late_evt    += s.get("late_create_count", 0)
              read_total  += s.get("cache_read", 0)
              cre_total   += s.get("cache_creation", 0)
              peak_max     = max(peak_max, s.get("peak_cache_read", 0))
              # Sub-agent cost is already inside cost_usd; these break out how much of it
              # came from fan-out rather than the main conversation.
              s_sub = s.get("subagent_cost_usd", 0.0)
              s_files = s.get("subagent_files", 0)
              sub_cost += s_sub; sub_files += s_files; sub_turns += s.get("subagent_turns", 0)
              is_fan = s_files >= args.fanout_threshold
              if is_fan: fan_cost += s["cost_usd"]; fan_conv += 1
              for m, v in (s.get("by_model") or {}).items():
                  e = model_mix[m]
                  e["main_cost"] += v["cost"]; e["main_turns"] += v["turns"]
              for m, v in (s.get("subagent_by_model") or {}).items():
                  e = model_mix[m]
                  e["sub_cost"] += v["cost"]; e["sub_turns"] += v["turns"]
              cwd = s.get("cwd") or "(no cwd / Cowork)"
              c = by_cwd[cwd]
              c["cost"] += s["cost_usd"]; c["conv"] += 1
              if is_mara: c["mara"] += 1; c["mara_cost"] += s["cost_usd"]
              if is_zomb: c["zomb"] += 1
              if is_fan: c["fan"] += 1
              c["sub_cost"] += s_sub
              c["late_tok"] += s.get("late_create_tokens", 0)
              bs = by_surface[s.get("surface","unknown")]
              bs["cost"] += s["cost_usd"]; bs["conv"] += 1
      
          # Top 12 cwds by spend.
          # Skip the synthetic "(no cwd / Cowork)" bucket — it's not a project, just an
          # aggregate of all cowork sessions. The surface line already shows cowork totals.
          real_cwds = [(k, v) for k, v in by_cwd.items() if k != "(no cwd / Cowork)"]
          top_cwds = sorted(real_cwds, key=lambda kv: -kv[1]["cost"])[:12]
          cwd_breakdown = []
          for k, v in top_cwds:
              emoji, health = classify_cwd(v)
              cwd_breakdown.append({
                  "cwd": k,
                  "cost": round(v["cost"], 2),
                  "conv": v["conv"],
                  "subagent_cost": round(v["sub_cost"], 2),
                  "fanout_conv": v["fan"],
                  "marathon_conv": v["mara"],
                  "marathon_share": round(v["mara_cost"]/v["cost"], 3) if v["cost"] else 0,
                  "zombie_conv": v["zomb"],
                  "late_tokens_M": round(v["late_tok"]/1e6, 2),
                  "health": health,
                  "emoji": emoji,
              })
      
          # Also expose deeper clean projects (beyond top 12) so the main agent can cite them
          # when no top project is clean.
          smallest_top_cost = top_cwds[-1][1]["cost"] if top_cwds else 0
          deeper_clean = []
          for k, v in real_cwds:
              if v["cost"] < 20: continue
              if v["cost"] >= smallest_top_cost: continue
              emoji, health = classify_cwd(v)
              if health == "clean":
                  deeper_clean.append({
                      "cwd": k,
                      "cost": round(v["cost"], 2),
                      "conv": v["conv"],
                  })
          deeper_clean.sort(key=lambda x: -x["cost"])
          deeper_clean = deeper_clean[:5]
      
          for e in model_mix.values():
              e["cost"] = round(e["main_cost"] + e["sub_cost"], 2)
              e["turns"] = e["main_turns"] + e["sub_turns"]
              e["main_cost"] = round(e["main_cost"], 2)
              e["sub_cost"] = round(e["sub_cost"], 2)
              e["share"] = round(e["cost"] / total_cost, 4) if total_cost else 0
          model_mix = dict(sorted(model_mix.items(), key=lambda kv: -kv[1]["cost"]))
      
          stats = {
              "marathon_threshold": args.marathon_threshold,
              "fanout_threshold": args.fanout_threshold,
              "total_cost": round(total_cost, 2),
              "total_conv": total_conv,
              "by_surface": {k: {"cost": round(v["cost"], 2), "conv": v["conv"]} for k, v in by_surface.items()},
              "buckets": {
                  b: {
                      "cost": round(bucket_cost[b], 2),
                      "count": bucket_count[b],
                      "share": round(bucket_cost[b] / total_cost, 4) if total_cost else 0,
                  } for b in [
                      "b1_1_5","b2_6_20","b3_21_50","b4_51_100",
                      "b5_101_300","b6_301_1000","b7_1001_plus",
                  ]
              },
              "marathon_cost": round(mara_cost, 2),
              "marathon_share": round(mara_cost / total_cost, 4) if total_cost else 0,
              "marathon_conv": mara_conv,
              "fanout_cost": round(fan_cost, 2),
              "fanout_share": round(fan_cost / total_cost, 4) if total_cost else 0,
              "fanout_conv": fan_conv,
              "subagent_cost": round(sub_cost, 2),
              "subagent_share": round(sub_cost / total_cost, 4) if total_cost else 0,
              "subagent_files": sub_files,
              "subagent_turns": sub_turns,
              "model_mix": model_mix,
              "zombie_cost": round(zomb_cost, 2),
              "zombie_share": round(zomb_cost / total_cost, 4) if total_cost else 0,
              "zombie_conv": zomb_conv,
              "cache_rebuild_tokens_M": round(late_tokens / 1e6, 2),
              "cache_rebuild_events": late_evt,
              "cache_rebuild_cost_est": round(late_tokens / 1e6 * 3, 2),
              "read_total_B": round(read_total / 1e9, 3),
              "create_total_B": round(cre_total / 1e9, 3),
              "rc_ratio": round(read_total / max(cre_total, 1), 1),
              "peak_cache_read_max_k": round(peak_max / 1000, 0),
              "by_cwd_top": cwd_breakdown,
              "by_cwd_clean_deeper": deeper_clean,
              "short_share": round(
                  (bucket_cost["b1_1_5"] + bucket_cost["b2_6_20"] + bucket_cost["b3_21_50"]) / total_cost,
                  4) if total_cost else 0,
          }
      
          Path(args.out).parent.mkdir(parents=True, exist_ok=True)
          Path(args.out).write_text(json.dumps(stats, indent=2))
      
          # Minimal confirmation only. The main agent reads user-stats.json and writes
          # the doctor's report directly to the terminal with framing and conclusions.
          def money(v): return f"${v:,.0f}" if v >= 100 else f"${v:.2f}"
          print(f"🩺  Stats ready: {money(stats['total_cost'])} across {stats['total_conv']:,} conversations, "
                f"{len(stats['by_cwd_top'])} top projects · saved to {args.out}")
          if stats["subagent_files"]:
              print(f"    including {money(stats['subagent_cost'])} of sub-agent fan-out "
                    f"({stats['subagent_share']*100:.0f}%) across {stats['subagent_files']:,} sub-agent transcripts")
      
      if __name__ == "__main__":
          main()
      
    • pick_hotspots.py 3.9 KB
      #!/usr/bin/env python3
      """Select 12-20 sessions for deep analysis by subagents.
      
      Balanced selection so the report covers expensive sessions AND positive examples:
        - Top 8 by absolute cost (where money is going)
        - Top 3 by cache_rebuild tokens (late context drip-feed)
        - Top 3 by turn count (raw marathon length)
        - Top 3 by read:create ratio at >= $5 cost (cache grind)
        - Top 3 positive examples: highest cost-per-turn efficiency among 20-100 turn sessions
      
      Dedupes by sid. Writes out/hotspots.json with the selection + reason for each.
      """
      from __future__ import annotations
      
      import argparse
      import json
      from pathlib import Path
      
      def load_sessions(path: str) -> list[dict]:
          rows = []
          with open(path) as f:
              for line in f:
                  line = line.strip()
                  if line:
                      rows.append(json.loads(line))
          return rows
      
      def main():
          ap = argparse.ArgumentParser()
          ap.add_argument("--in", dest="inp", default="out/sessions.jsonl")
          ap.add_argument("--out", default="out/hotspots.json")
          ap.add_argument("--top-cost",     type=int, default=8)
          ap.add_argument("--top-rebuild",  type=int, default=3)
          ap.add_argument("--top-turns",    type=int, default=3)
          ap.add_argument("--top-grind",    type=int, default=3)
          ap.add_argument("--top-positive", type=int, default=3)
          args = ap.parse_args()
      
          sessions = load_sessions(args.inp)
      
          def pick(rows, key_fn, n, reason, min_cost=0.0):
              sel = [s for s in rows if s["cost_usd"] >= min_cost]
              sel.sort(key=key_fn, reverse=True)
              return [(s, reason) for s in sel[:n]]
      
          bucket_cost    = pick(sessions, lambda s: s["cost_usd"], args.top_cost, "top-cost")
          bucket_rebuild = pick(sessions, lambda s: s.get("late_create_tokens", 0), args.top_rebuild, "drip-feed", min_cost=1)
          bucket_turns   = pick(sessions, lambda s: s["turn_count"], args.top_turns, "marathon", min_cost=1)
          bucket_grind   = pick(sessions, lambda s: s.get("rc_ratio", 0), args.top_grind, "cache-grind", min_cost=5)
      
          # Positive: efficient short-to-mid focused sessions (20-100 turns, lowest cost per turn)
          short_focused = [s for s in sessions if 20 <= s["turn_count"] <= 100 and s["cost_usd"] >= 0.5]
          short_focused.sort(key=lambda s: s["cost_usd"] / max(s["turn_count"], 1))
          bucket_positive = [(s, "positive-focused") for s in short_focused[:args.top_positive]]
      
          # Merge with dedup, preserving reason for each pick
          seen: dict[str, dict] = {}
          for s, reason in (bucket_cost + bucket_rebuild + bucket_turns + bucket_grind + bucket_positive):
              sid = s["sid"]
              if sid not in seen:
                  seen[sid] = {"session": s, "reasons": [reason]}
              else:
                  if reason not in seen[sid]["reasons"]:
                      seen[sid]["reasons"].append(reason)
      
          # Sort final list by cost desc for output
          hotspots = sorted(seen.values(), key=lambda x: -x["session"]["cost_usd"])
      
          summary = [{
              "sid": h["session"]["sid"],
              "cost": round(h["session"]["cost_usd"], 2),
              "turns": h["session"]["turn_count"],
              "duration_min": h["session"].get("duration_min"),
              "rc_ratio": h["session"].get("rc_ratio"),
              "late_create_count": h["session"].get("late_create_count"),
              "peak_cache_read_k": round((h["session"].get("peak_cache_read") or 0)/1000),
              "cwd": h["session"].get("cwd"),
              "title": (h["session"].get("title") or "")[:120],
              "reasons": h["reasons"],
          } for h in hotspots]
      
          Path(args.out).parent.mkdir(parents=True, exist_ok=True)
          Path(args.out).write_text(json.dumps({"count": len(hotspots), "hotspots": summary}, indent=2))
          print(f"Picked {len(hotspots)} hotspots to {args.out}")
          for h in summary:
              rs = ",".join(h["reasons"])
              title = h["title"] or "(no title)"
              print(f"  ${h['cost']:>7.2f}  {h['turns']:>4} turns  [{rs}]  {title[:60]}")
      
      if __name__ == "__main__":
          main()
      
    • pricing.py 2.5 KB
      """Per-token cost calculation.
      
      Per-million-token list rates, USD: (input, output, cache_read, cache_write_5m).
      
      Anthropic rates: see references/pricing.md.
      OpenAI rates (for Codex sessions): public list prices, verified 2026-06.
        gpt-5.4       $2.50 in / $15.00 out / $0.25 cached-in   (cache_write n/a)
        gpt-5.4-mini  $0.75 in / $4.50  out / $0.075 cached-in
        gpt-5.4-nano  $0.20 in / $1.20  out / $0.02 cached-in
        Source: openai.com/api/pricing (gpt-5.4 family). Update when rates change.
      
      Unknown models return no rate, and cost_for_usage() yields 0.0 so the caller can
      show token counts without a fabricated dollar figure.
      """
      from __future__ import annotations
      
      # (input, output, cache_read, cache_write_5m) per million tokens, USD.
      _RATES = {
          # Anthropic
          "opus":   (15.00, 75.00, 1.50, 18.75),
          "sonnet": ( 3.00, 15.00, 0.30,  3.75),
          "haiku":  ( 1.00,  5.00, 0.10,  1.25),
          # OpenAI (Codex). cache_write n/a -> 0.0.
          "gpt-5-nano": (0.20,  1.20, 0.02, 0.0),
          "gpt-5-mini": (0.75,  4.50, 0.075, 0.0),
          "gpt-5":      (2.50, 15.00, 0.25, 0.0),
      }
      
      
      def _family(model: str) -> str | None:
          m = (model or "").lower()
          # OpenAI / Codex
          if "gpt" in m or m.startswith("o1") or m.startswith("o3") or m.startswith("o4"):
              if "nano" in m:
                  return "gpt-5-nano"
              if "mini" in m:
                  return "gpt-5-mini"
              return "gpt-5"
          # Anthropic. Empty/"unknown" keeps the historical sonnet default (Claude
          # transcripts occasionally omit the model id).
          if "opus" in m:
              return "opus"
          if "haiku" in m:
              return "haiku"
          if "claude" in m or m in ("", "unknown"):
              return "sonnet"
          return None
      
      
      def rates_for(model: str) -> tuple[float, float, float, float] | None:
          fam = _family(model)
          return _RATES[fam] if fam else None
      
      
      def cost_for_turn(
          model: str,
          input_tokens: int = 0,
          output_tokens: int = 0,
          cache_read: int = 0,
          cache_create: int = 0,
      ) -> float:
          rates = rates_for(model)
          if rates is None:
              return 0.0
          inp, out, cr, cw = rates
          return (
              input_tokens   / 1e6 * inp +
              output_tokens  / 1e6 * out +
              cache_read     / 1e6 * cr  +
              cache_create   / 1e6 * cw
          )
      
      
      def cost_for_usage(model: str, usage: dict) -> float:
          return cost_for_turn(
              model,
              input_tokens=int(usage.get("input") or 0),
              output_tokens=int(usage.get("output") or 0),
              cache_read=int(usage.get("cache_read") or 0),
              cache_create=int(usage.get("cache_creation") or usage.get("cache_create") or 0),
          )
      
  • SKILL.md 17.3 KB
    ---
    name: token-doctor
    description: Personal diagnosis of where your Claude Code + Cowork spend goes. Reads local transcripts, prints your conversation length distribution, marathon share, cache rebuild costs, and per-project diagnosis (good projects and problem projects) right in the terminal. Then offers a deeper dive that fans out parallel Haiku subagents over your most expensive (and most efficient) sessions and writes a tight Markdown report. Use when the user asks "why is my Claude spend so high", "where am I burning tokens", "diagnose my Claude habits", "audit my Claude usage", or asks for a personal token-cost diagnosis.
    ---
    
    # Token Doctor
    
    > **Platforms: Claude Code / Cowork and Codex.** `scripts/inventory.py` detects the host (via the `platform` stamp `install.sh` writes, or `AI_FIRST_PLATFORM`) and routes: Claude Code (`~/.claude/projects`) + Cowork transcripts, or Codex rollouts (`~/.codex/sessions`). Codex token usage comes from Codex's own per-response `token_count` events; cost uses OpenAI list rates in `pricing.py` (gpt-5.4 family; unknown models show token counts with no fabricated cost). **Antigravity** is unsupported: its IDE store is AEAD-encrypted at rest and its CLI store has no parseable turn/token content, so the skill prints a clear "not available" message and exits.
    
    Two-stage diagnostic. Stage 1 is fast and lands directly in the terminal so the user always walks away with their numbers. Stage 2 is opt-in, fans out subagents over hotspots, and writes a tight Markdown report.
    
    **Read this whole file before running.**
    
    ## When to run
    
    Trigger phrases: "diagnose my Claude habits", "why am I spending so much", "where are my tokens going", "audit my spend", "token doctor", "what's driving my Claude bill".
    
    Do NOT trigger for:
    - "what tasks do I do with Claude" → that's `task-profile`.
    - "where did I work on X" → that's `session-search`.
    
    The line is: token-doctor is about cost **shape**, not task inventory or recall.
    
    ## Prerequisites
    
    - Claude Code transcripts: `~/.claude/projects/*/*.jsonl` (CLI **and** desktop app)
    - Sub-agent transcripts: `~/.claude/projects/*/<sid>/subagents/**/*.jsonl`, including workflow agents under `subagents/workflows/<wf>/`
    - Claude Cowork transcripts (optional): `~/Library/Application Support/Claude/local-agent-mode-sessions/*/*/local_*/audit.jsonl`
    - Python 3, stdlib only. No external services.
    
    If neither path exists, stop and say so.
    
    ### What counts as one session
    
    Sub-agent transcripts are separate files but the same piece of work, so their cost
    rolls into the parent session rather than appearing as sessions of their own. This
    matters when you read the report:
    
    - `cost_usd` is main conversation **plus** fan-out. `main_cost_usd` and `subagent_cost_usd` split it.
    - Turn counts, the timeline, cache rebuilds and the re-read ratio are **main-session figures**. They describe how that one conversation's context grew; sub-agents have their own context.
    - So a session can show 100 turns and $641. That is not a contradiction, it is fan-out. Say so rather than letting the reader trip over it.
    - One turn = one assistant **message**, not one transcript line. Claude Code writes a line per content block (thinking, text, each tool_use) and repeats the same `usage` object on every one, so counting lines inflates turns and cost by roughly 2-3x. The inventory dedupes by `message.id`. Turn counts from an older run of this skill are not comparable with these.
    
    Automation excluded by default: `sdk-cli` background dispatch, paperclip, ditto-routines,
    scheduled tasks, and the automation slash commands. Desktop-app sessions are interactive
    work and **are** counted.
    
    ---
    
    ## STAGE 1 — Fast diagnosis (always runs, you write the report)
    
    The goal is: the user invokes the skill, sees a clean doctor's report within 10 seconds, knows which projects are healthy and which are bleeding, and can decide whether to go deeper. **You write the report directly in your message** based on the JSON the scripts produce. The scripts compute, you communicate.
    
    ### Step 1.1 — Inventory (deterministic)
    
    ```bash
    ~/.claude/skills/token-doctor/scripts/inventory.py --since YYYY-MM-DD --out out/sessions.jsonl
    ```
    
    Default window: last 90 days. Flags: `--since`, `--until`, `--all`, `--include-automation`, `--no-cowork`. Automation runs (`sdk-cli` background dispatch, paperclip, `/loop`, `/schedule`, ditto-routines, scheduled-tasks) excluded by default.
    
    `--until YYYY-MM-DD` means midnight at the **start** of that day, so it excludes that day's sessions. To include today, leave `--until` off.
    
    The scan prints how many sub-agent transcripts it rolled into parent sessions. If it also reports sub-agents with no parent session on disk, that cost is not in the totals; mention it only if it is material.
    
    ### Step 1.2 — Aggregate (deterministic)
    
    ```bash
    ~/.claude/skills/token-doctor/scripts/personal_stats.py --in out/sessions.jsonl --out out/user-stats.json
    ```
    
    Prints a single confirmation line. The full data is in `out/user-stats.json`.
    
    ### Step 1.3 — Read the JSON and write the doctor's report
    
    Read `out/user-stats.json`. Then **write the report directly in your message** as terminal-style ASCII with emojis. The user reads your message; no intermediate file.
    
    #### Report structure (mandatory sections, in order)
    
    ```
    🩺 ┌────────────────────────────────────────────────────────────────────┐
       │            TOKEN DOCTOR · personal diagnosis                       │
       └────────────────────────────────────────────────────────────────────┘
    
      Patient: <user's first name or "you">
      Window:  <window dates from inventory>
      Spend:   $<total> list-price equivalent · <conv count> conversations
    
      ── Vital signs ─────────────────────────────────────────────────────────
    
      🔴/🟡/🟢 Marathon (≥300 turns)    <N> conv  ·  $<X>  ·  <Y>% of spend
      🔴/🟡/🟢 Fan-out (≥5 sub-agents)  <N> conv  ·  $<X>  ·  <Y>% of spend
      🔴/🟡/🟢 Zombie (≥4h wall clock)  <N> conv  ·  $<X>  ·  <Y>% of spend
      🔴/🟡/🟢 Cache rebuilds            <N> events · <Z>M tokens · ~$<X>
      🔴/🟡/🟢 Re-read ratio             <X>×   (healthy ≤15×, org avg 30×)
      📈 Peak context observed       <X>k tokens
      🌳 Sub-agent cost               $<X> of $<total>  (<Y>%) across <N> transcripts
    
      ── Spend by conversation length ────────────────────────────────────────
    
           1 to 5        <bar>   <%>   (<N> conv)
           6 to 20       <bar>   <%>   (<N> conv)
           …
           1,000+        <bar>   <%>   (<N> conv)
    
      ── Model mix ───────────────────────────────────────────────────────────
    
        <model>        $<X>   <%>   <N> turns    main $<X>/<N>t · sub $<X>/<N>t
        <model>        $<X>   <%>   <N> turns    main $<X>/<N>t · sub $<X>/<N>t
        …
    
      ── Diagnosis ───────────────────────────────────────────────────────────
    
      <2-4 sentences synthesizing the vitals into one clear picture. Lead with
      the dominant antipattern in this user's data, then the corollary cost. End
      with one line about the strongest positive signal you see.>
    
      ── Project chart ───────────────────────────────────────────────────────
    
      ✅ clean · 🏃 marathon · 🌳 fanout · 🔄 rebuilds · 🧟 zombie · ⚠️ multiple
    
      <emoji>  $<X>  <truncated cwd>                              <meta line>
      <emoji>  $<X>  <truncated cwd>                              <meta line>
      … up to 10-12 rows from by_cwd_top
    
      ── Treatment plan ──────────────────────────────────────────────────────
    
      💚 Keep doing:
         <bullet referencing a clean project from by_cwd_top or by_cwd_clean_deeper,
          OR a positive structural signal like a high short_share if no clean cwd is in top 12>
         <2-3 bullets total>
    
      🎯 Change first:
         <one concrete action tied to the biggest lever, with cited project>
         <2-3 bullets total, ordered by expected impact>
    
      ── Want a deeper look? ─────────────────────────────────────────────────
    
      <one-line question asking if they want the deep dive>
    ```
    
    #### Rules when writing the report
    
    - **Use the emojis above consistently.** Box-drawing characters (─ ┌ └ │) are fine and make the report look like a medical printout.
    - **Traffic-light dots**: 🔴 = bad, 🟡 = watch, 🟢 = healthy. Apply the bands in the rubric below.
    - **Per-project emoji** must come from `by_cwd_top[i].emoji` in the JSON. Do not re-classify.
    - **Model mix** comes from `model_mix`, already sorted by cost. Show every model down to 1% of spend, then stop. Use readable names (`claude-opus-5` → Opus 5, `claude-fable-5-1` → Fable 5.1). The `main` / `sub` split is the point of the section: a model that is cheap in the main conversation and expensive across sub-agents is the clearest lever in the whole report, because sub-agent model tier is a one-line change in an `Agent(...)` call. Call that out when you see it.
    - **Fan-out** uses `fanout_conv` / `fanout_cost` / `fanout_share` (sessions with ≥ 5 sub-agents) and `subagent_cost` / `subagent_share` (fan-out's share of total spend). If `subagent_files` is 0, drop both the fan-out vital and the sub-agent line rather than printing zeroes.
    - **Bars** for the length distribution: build them with `█` characters proportional to the share. Use a fixed width like 36 chars.
    - **The diagnosis paragraph is yours to write** — it is the doctor's read on the data. Be specific. Don't restate the numbers; conclude from them. Aim for 3-5 sentences max. Examples of good diagnostic sentences:
      - "Your spend is concentrated in a small number of very long sessions: 22 conversations carry 70% of your bill."
      - "Cache rebuilds are minor at $388, but the re-read ratio of 23× tells me your context grows fast inside those long sessions."
      - "Three of your top five projects are evaluation runs from last month. Each is one long session; splitting them would compound."
    - **Treatment plan must include both "keep doing" AND "change first" sections.** Skipping the positive section is forbidden. Pick from:
      - `by_cwd_top` entries with `emoji == "✅"` for clean
      - `by_cwd_clean_deeper` for clean projects below the top 12
      - `short_share` if neither is available — frame as "X% of your spend is in short focused sessions, so the habit is there, you just don't use it everywhere"
    - **Cite specific projects.** Truncate cwds to the last 36-44 chars with a leading `…` if they're long. Drop the `/Users/<name>/` prefix when it makes the line cleaner.
    - **No em-dashes.** Use commas, semicolons, or periods.
    - **No "waste", "burning", "bad habit".** Use "cost", "spend", "context", "rebuild".
    
    #### Traffic-light thresholds
    
    | Metric | 🟢 | 🟡 | 🔴 |
    |---|---|---|---|
    | Marathon share | < 20% | 20-50% | ≥ 50% |
    | Fan-out share | < 20% | 20-50% | ≥ 50% |
    | Zombie share | < 20% | 20-50% | ≥ 50% |
    | Cache rebuild $ | < $50 | $50-$200 | ≥ $200 |
    | Re-read ratio | ≤ 15× | 15-35× | ≥ 35× |
    
    #### Asking about the deep dive
    
    End with one short line, not a paragraph. Example:
    
    > Want me to pull apart your top sessions one by one — what specifically drove each marathon, plus a couple of your most efficient runs to learn from? Takes about a minute, runs ~15 Haiku subagents in parallel.
    
    If they say no, stop. The report is the deliverable.
    
    ---
    
    ## STAGE 2 — Deep dive (opt-in)
    
    ### Step 2.1 — Pick hotspots
    
    ```bash
    ~/.claude/skills/token-doctor/scripts/pick_hotspots.py --in out/sessions.jsonl --out out/hotspots.json
    ```
    
    Selects ~14 sessions:
    - 8 by absolute cost
    - 3 by cache rebuild tokens
    - 3 by raw turn count
    - 3 by read:create ratio at ≥$5 cost
    - **3 positive examples** (lowest cost-per-turn at 20-100 turns)
    
    Briefly tell the user the list before fan-out so they can drop sensitive sids. Keep it to one line per session: `$X · N turns · short title`.
    
    ### Step 2.2 — Build payloads
    
    ```bash
    ~/.claude/skills/token-doctor/scripts/build_payloads.py --sessions out/sessions.jsonl --hotspots out/hotspots.json --outdir out/payloads
    ```
    
    Writes one redacted payload per hotspot. Payloads include token shape, tool-call counts, timeline samples, and a 120-char title. They do NOT include user prompt bodies or model output text.
    
    ### Step 2.3 — Parallel subagent fan-out
    
    For each payload in `out/payloads/`, dispatch one subagent. **Send all calls in one message with multiple tool blocks so they run in parallel.**
    
    ```
    Agent(
      description="Diagnose one session",
      subagent_type="general-purpose",
      model="haiku",
      run_in_background=true,
      prompt="""
    Diagnose the session at out/payloads/<sid>.json.
    
    Read first, in order:
      1. ~/.claude/skills/token-doctor/references/antipattern-taxonomy.md
      2. ~/.claude/skills/token-doctor/references/diagnosis-rubric.md
      3. out/payloads/<sid>.json
    
    Apply the rubric. Emit strictly the JSON schema (see rubric §Output) to out/analyses/<sid>.json. Hard length limits: what_happened max 2 sentences (~25 words), trigger_moment.what max 14 words, would_have_helped max 18 words. Lead with structural facts (turn count, context size, key signal). No restating the schema.
    
    Tone: second-person, neutral, no "waste" / "burning".
    """
    )
    ```
    
    Wait for all to complete.
    
    ### Step 2.4 — Synthesize the report (main agent)
    
    Read every `out/analyses/*.json`. Then:
    
    1. **Group by cwd**. For each cwd that has ≥2 analyzed sessions, decide if it shows a dominant pattern.
    2. **Find the signature**: the one antipattern that recurs most across the user's data, and the one positive habit they have consistently.
    3. **Write `out/recommendations.md`** — tight, scannable, no padding. Structure:
    
    ```markdown
    # Token Doctor — your diagnosis
    
    **Signature.** <one sentence: dominant antipattern + dominant strength>
    
    **Bottom-line lever.** <one sentence: the single habit change with the biggest expected impact>
    
    ## What you're doing well
    
    - **<positive pattern>** in `<cwd>`. <one sentence with one cited sid>
    - **<positive pattern>** in `<cwd>`. <one sentence with one cited sid>
    
    (2-3 bullets. At least one is mandatory; do not skip this section.)
    
    ## What's driving your bill
    
    - **<antipattern>** in `<cwd>`. <one sentence. Cite the worst sid and one specific turn or signal>
    - ...
    
    (3-5 bullets ordered by estimated savings)
    
    ## Per-session diagnoses
    
    | Cost | Turns | Verdict | What happened |
    |---:|---:|---|---|
    | $XXX | NNN | <label> | <1-2 line what_happened from the analysis> |
    | ...
    
    (Only the analyzed sessions. Use the `what_happened` field verbatim from each analysis JSON.)
    ```
    
    Also write `out/recommendations.json` with structured form for re-use:
    
    ```json
    {
      "signature": {
        "primary_antipattern": "marathon | drip-feed | zombie | bloat | grind | drift | fanout | none",
        "primary_strength": "focused | front-loaded | time-bounded | lean | directed | none",
        "one_line": "<one sentence about how this user's spend is shaped>"
      },
      "bottom_line_lever": "<one sentence>",
      "positives": [{"pattern": "...", "cwd": "...", "evidence_sid": "...", "note": "..."}],
      "antipatterns": [{"pattern": "...", "cwd": "...", "evidence_sids": ["..."], "note": "...", "expected_impact": "small|medium|large"}],
      "session_table": [{"sid": "...", "cost": 0, "turns": 0, "verdict": "...", "what_happened": "..."}]
    }
    ```
    
    Keep the prose terse. The user already has the terminal numbers; the report's job is to point at specific projects and habits, not to recite stats.
    
    ---
    
    ## Privacy contract
    
    - Everything runs locally. No transcript text leaves the machine.
    - Subagents receive token counts, tool-call names, turn indices, and a 120-char title. They do not receive user prompt text or model output text.
    - Output files in `out/` contain session ids and short titles. They do not contain conversation content.
    - The user can `rm -rf out/` to wipe everything.
    
    ## Tone
    
    - Descriptive, not punitive. The user is reading their own data.
    - Both antipatterns and positive habits get airtime. Skipping the "what you're doing well" section is forbidden.
    - Specific numbers, specific turn indices, specific sids. Avoid hedging.
    - No em-dashes. No "waste", "burning", "bad habit".
    - Emojis are allowed in the terminal output (the `personal_stats.py` script uses them). Keep them out of the Markdown report — there it should look like an engineering doc.
    
    ## Failure modes
    
    - **No transcripts found.** Stop with a clear message.
    - **Subagent emitted invalid JSON.** Skip that sid, log a warning once, continue.
    - **All sessions are automation.** Tell the user to re-run with `--include-automation` if they want those analyzed; otherwise note the interactive count.
    - **Single-cwd user.** Skip the per-cwd grouping in the report; recommendations still work as a flat list.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related