Claude Skill

deep-research-team

This skill should be used when the user asks for "deep research", "research team", "comprehensive analysis", "research report", "investigate thoroughly", "compare X vs Y in depth", or needs synthesis across multiple sources with verification. It spawns a coordinated team of resea

LLM Mart · 0 points · 0 views 4 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download malob-nix-config-configs_claude_skills_deep-research-team-500b474.zip · 80 KB
Part of malob/nix-config — 13 skills

Install

skills CLI npx skills add https://github.com/malob/nix-config/tree/master/configs/claude/skills/deep-research-team
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install malob-nix-config@llmmart
Git git clone https://github.com/malob/nix-config.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole malob/nix-config collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Deep Research Team (Lead Orchestrator)

Conduct thorough, iterative research by coordinating a persistent team of researcher agents across multiple rounds. This architecture enables mid-investigation steering, targeted follow-up based on emerging findings, and cross-agent verification.

Architecture Overview

Round 1: Investigation      Round 2: Follow-up          Synthesis

┌──────────┐                  ┌──────────┐                  ┌────────┐
│Researcher│  sends findings  │Researcher│  sends findings  │        │
│    A     ├─────────┬───────>│    A     ├─────────┬───────>│        │
└──────────┘         │        └──────────┘         │        │        │
                     │                             │        │        │
                     v          dispatches         v        │        │
┌──────────┐    ┌────────┐    ┌──────────┐    ┌────────┐    │  Lead  │
│Researcher├───>│  Lead  │───>│Researcher├───>│  Lead  │───>│  synth │
│    B     │    │triages │    │    B     │    │triages │    │  esizes│
└──────────┘    └────────┘    └──────────┘    └────────┘    │        │
                     ^                             ^        │        │
┌──────────┐         │        ┌──────────┐         │        │        │
│Researcher├─────────┴───────>│Researcher├─────────┴───────>│        │
│    C     │  sends findings  │    C     │  sends findings  │        │
└──────────┘                  └──────────┘                  └────────┘

Key principles:

  1. No peer-to-peer researcher communication. All coordination goes through the lead. This preserves the independence that accounts for 87% of multi-agent gains (Choi et al.) and avoids sycophancy failures (Wynn et al.). Researchers never see each other's findings.

  2. Multi-round iteration. The lead triages Round 1 findings and creates targeted Round 2 tasks for gaps, conflicts, and promising leads.

  3. Cross-agent verification (Comprehensive scope). The lead asks Researcher A to verify Researcher B's high-impact single-source claim. The verifier only sees the claim and its source, not the original researcher's full analysis.

  4. Dynamic task evolution. The shared task list starts with pre-planned angles but grows organically as follow-up tasks emerge from findings. The lead dispatches follow-up tasks directly to specific researchers via SendMessage.

When to Use

Use this skill for:

  • Complex questions that benefit from multiple research angles
  • Topics where initial findings will reveal what to investigate next
  • Research requiring cross-verification of contested claims
  • Any question needing synthesis across 5+ sources

Do NOT use for:

  • Simple factual lookups (use regular web search)
  • Questions answerable with 1-2 searches
  • Debugging or code questions

Effort Calibration

Scope Researchers Rounds Verification Model
Focused 2 1-2 None sonnet
Broad 3 2-3 None sonnet
Comprehensive 4 3-4 Cross-agent opus

Round counts are heuristics, not targets. Stop early when you hit citation convergence -- additional rounds that don't surface new substantive findings waste tokens and context. A Broad run that converges in 2 rounds is a success, not a shortcut. After each round's triage, ask: "Would another round change the report's conclusions?" If not, proceed to synthesis.

Default scope is determined by question type (see references/question-types.md). Present the recommended scope to the user and allow override.

Model selection:

  • Lead: inherits user's session model (no override)
  • Researchers: sonnet for Focused/Broad, opus for Comprehensive

(Sonnet validated as viable override for Comprehensive when cost matters)

Output Directory

Research artifacts persist to disk for resumability and backup.

Directory resolution -- run this command FIRST, before creating anything. The output is your {output_dir}. Only the fallback branch creates a directory; the others reuse what exists.

if [ -d "$(pwd)/deep-research" ]; then
  echo "$(pwd)/deep-research"
elif [ -n "$CLAUDE_DEEP_RESEARCH_DIR" ]; then
  eval echo "$CLAUDE_DEEP_RESEARCH_DIR"
else
  mkdir -p "$(pwd)/deep-research"
  echo "$(pwd)/deep-research"
fi

After resolving {output_dir}, create only the topic subdirectory in Phase 3.

Each session creates a subdirectory: {output_dir}/{topic-slug}/

Contents:

  • state.md -- triage checkpoint, cross-references, follow-up plan (written in Phase 4)
  • researcher-{letter}-findings.md -- backup of each researcher's findings
  • report.md -- final synthesized report (written in Phase 6)

Calibrated Confidence Language

Use Kent-style verbal probability expressions in all confidence assessments:

Term Range Use When
Almost certain 93-99% Multiple high-quality sources, no dissent
Highly likely 80-92% Strong evidence, minor caveats
Likely 63-79% Good evidence, some gaps
Roughly even 40-62% Conflicting evidence, genuinely uncertain
Unlikely 20-39% Limited or weak evidence

Always pair the verbal term with the probability range in the final report.

Process

Phase 0: Classify Question Type

Silently classify the user's question before any interaction.

  1. Read references/question-types.md for the full taxonomy.
  2. Assign a primary type: Factual, Scientific/Health, Consumer, Technical, Opinion/Sentiment, Contested, or Emerging/Frontier.
  3. For compound questions, decompose into sub-questions and classify each.
  4. Note the default scope from the type-to-scope mapping.

Resume check: Before starting, list the subdirectories in {output_dir} and scan for any that look related to the current question (similar topic, overlapping keywords). If you find a plausible match, read its state.md and offer to resume: present what was completed, what remains, and ask the user whether to resume or start fresh. If resuming, create a new team and tasks for only the remaining work.

Topic slug: When creating a new session, generate a slug (lowercase, hyphenated, max 40 chars) for the subdirectory name: {output_dir}/{slug}/.

Classification is internal—do not present it to the user.

Phase 1: Clarify and Plan

Step 1: Make sure you understand the question. Before planning anything, ask yourself: do I understand what the user is asking and why well enough to design research angles that will actually be useful to them? If not, use AskUserQuestion to fill the gaps. This isn't just about ambiguous wording -- a perfectly clear question can still lack enough context to research well ("How does Nix handle dependencies?" means very different research depending on whether you're evaluating Nix, debugging an issue, or writing docs). If the question and its context are clear, skip this step.

Step 2: Scope and decompose. Determine the appropriate scope (Phase 2 has the details) and decompose the question into independent research angles. Default angle counts by scope:

  • Focused: 2 angles
  • Broad: 3 angles
  • Comprehensive: 4 angles

These are defaults, not caps. If the decomposition reveals one more genuinely independent facet than the default, add it (e.g., 3 angles for a Focused run). If the question has fewer real facets, use fewer. Beyond ±1 from the default, re-scope rather than stretching -- the scope was probably wrong. Each angle must be independent and substantial enough to warrant a dedicated researcher; "I can think of another angle" isn't sufficient.

For compound questions, map sub-questions to angles. Multiple sub-questions can share an angle if closely related; a single sub-question can span multiple angles if it has distinct facets.

Step 3: Confirm if high-investment. For compound, contested, or Comprehensive-scope questions, present the research plan for user approval before spawning researchers:

Research plan for "":

  • Type: | Scope: | researchers, rounds
  • Angles:
  • [If compound] Sub-question → angle mapping: ...

Proceed, or adjust?

For clear, low-scope questions, skip confirmation and proceed.

Phase 2: Calibrate Effort

Select the scope tier based on question type defaults from references/question-types.md, then apply the scope modifiers from that file (de-escalation and escalation signals). Also adjust for:

  • User's explicit preference (if stated)
  • Structural complexity (compound questions with 3+ sub-types escalate)

Announce the plan: "Starting team research with researchers."

Phase 3: Team Setup

Step 1: Create the output directory

mkdir -p {output_dir}/{topic-slug}

Step 2: Create the team

TeamCreate:
  team_name: "deep-research-{topic-slug}"
  description: "Researching {topic} in {scope} scope"

Step 3: Create initial tasks

One TaskCreate per research angle:

TaskCreate:
  subject: "Investigate {angle title}"
  description: |
    Research angle: {angle description}
    Topic context: {brief topic summary}
    Question type: {type from Phase 0}
    Focus: {what specifically to investigate}
    Return structured findings via SendMessage to the lead.
  activeForm: "Investigating {angle title}"

Task brief clarity: Make scope boundaries explicit between researchers to avoid overlap and gaps. Flag name ambiguities (e.g., multiple products sharing a name). Mark optional sub-tasks clearly (e.g., "if time permits" vs required).

Step 4: Spawn researchers

Launch ALL researchers in a SINGLE message. Each researcher gets:

Task:
  subagent_type: "general-purpose"
  name: "researcher-{letter}"
  team_name: "deep-research-{topic-slug}"
  model: "sonnet"  (or "opus" for Comprehensive)
  description: "Spawn researcher {letter}"
  prompt: |
    You are a research agent on a team. Your job is to investigate research tasks
    by searching the web, evaluating sources, and reporting structured findings.

    FIRST: Read your methodology at: {absolute path to references/researcher-prompt.md}

    Question type: {type from Phase 0}
    Output directory: {output_dir}/{topic-slug}
    Your researcher letter: {letter} (use LOWERCASE in filenames: researcher-{lowercase letter})
    Lead name: team-lead (send all findings to this name via SendMessage)

    Your assigned task is #{id}: "{subject}"
    Use TaskGet for full details, then begin investigation.

    After completing your task, go idle. The lead will message you directly
    when new tasks are available.

Include the task ID, subject, question type, output directory, researcher letter, and lead name directly in each researcher's spawn prompt.

Phase 4: Investigation Loop

This is the core research cycle. Each iteration is a round: researchers investigate, the lead triages, then either dispatches follow-ups (another round) or exits to synthesis.

Round structure

Investigate: Researchers work on their assigned tasks. Each researcher will:

  1. Read the methodology reference file
  2. Load web search and content extraction tools via ToolSearch
  3. Execute the investigation loop (search -> evaluate -> reflect -> decide)
  4. Write findings to {output_dir}/{topic-slug}/researcher-{letter}-findings.md
  5. Notify the lead via SendMessage with the file path (not the full findings -- avoids doubling output tokens)
  6. Mark their task as completed via TaskUpdate
  7. Go idle and wait for the lead to dispatch follow-up tasks via SendMessage

Monitoring: The lead reads each researcher's findings file after receiving their notification. No polling needed.

Handling partial results: If a researcher reports rate limit issues or thin coverage, note the gap for triage rather than immediately spawning replacements.

Triage: After all tasks for the current round complete, systematically review findings.

Step 1: Extract and cross-reference claims

For each significant claim across all researcher findings:

  • How many independent sources support it? (Different researchers finding the same source counts as one source, not two.)
  • HIGH confidence: 3+ independent sources of different types (e.g., paper + dataset + practitioner account), no credible dissent
  • MEDIUM confidence: 2 independent sources, or multiple sources of the same type
  • LOW confidence: Single source, or multiple sources that trace back to one original

Step 2: Identify gaps and conflicts

  • What angles remain uncovered?
  • Where do researchers contradict each other?
  • What findings are surprising and deserve deeper investigation?
  • Which claims rest on a single source?

Step 3: Decide whether to continue or exit

Exit to Phase 5 (Synthesis) if findings have converged -- another round wouldn't change the report's conclusions. Continue if significant gaps, conflicts, or single-source high-impact claims remain and the scope's round budget allows.

If findings reveal more complexity than anticipated (e.g., Broad scope uncovering deeply contested claims requiring steelmanning), escalate: spawn an additional researcher or add a round beyond the default budget.

Step 4: Persist triage state

Write or update {output_dir}/{topic-slug}/state.md:

# Research State: {topic}

## Status: TRIAGE_COMPLETE (Round {N})

## Question Type: {type}

## Scope: {scope}

## Round {N} Summary

{brief cross-reference of key findings, gaps, conflicts}

## Follow-up Plan

{list of planned follow-up tasks with rationale, or "Proceeding to synthesis"}

Dispatching follow-up tasks

If continuing, create targeted tasks based on what triage revealed. These are all just task types -- they use the same dispatch mechanism:

Gap-fill: "No researcher covered . Investigate ."

Conflict-resolution: "One source says X, another says Y. Search for additional sources that clarify which is accurate and why they might differ."

Deep-dive: "Initial findings revealed . Investigate further: ."

Verification (Comprehensive scope): Use when high-impact claims rest on a single source, factual conflicts remain unresolved, or claims are in specialized/niche domains where citation error rates are higher. Assign to a researcher who did NOT make the original claim. Task description contains only the claim and its source URL:

TaskCreate:
  subject: "Verify: {claim summary}"
  description: |
    Verification task. Search for ADDITIONAL sources (not the original) and determine
    if they support, contradict, or add nuance to this claim.

    CLAIM: {specific factual claim}
    ORIGINAL SOURCE: {URL}

    Report your verdict as: SUPPORTED, SUPPORTED WITH NUANCE, CONTESTED, or UNCHANGED
    Include the additional sources you found and any important nuance.
  activeForm: "Verifying claim about {topic}"

Critical: Follow-up task descriptions contain just enough context without revealing other researchers' full conclusions. This preserves independence.

Assignment strategy: The lead assigns follow-up tasks directly via SendMessage rather than relying on researchers to self-claim. Choose assignees based on task type:

  • Deep-dives and gap-fills: Assign to the researcher who covered the related angle (continuity -- they have context on what was already found).
  • Conflict resolution and verification: Assign to a researcher who did NOT cover either side (fresh perspective avoids confirmation bias).
  • If researchers outnumber tasks: Idle researchers wait or are shut down early.
SendMessage:
  type: "message"
  recipient: "researcher-{letter}"
  content: "New task available: #{id} -- {subject}. Please claim it and begin."
  summary: "Follow-up task assignment"

Then loop back to Investigate above.

Interpreting verification results: When a verification task returns, adjust confidence:

  • SUPPORTED: Upgrade confidence; note additional sources
  • SUPPORTED WITH NUANCE: Directionally correct but specific details differ or require qualification. Upgrade confidence for the general claim; add caveats for specifics.
  • CONTESTED: Flag explicitly; present both sides with evidence
  • UNCHANGED: Keep original confidence level

Phase 5: Synthesize (Type-Aware)

Combine all findings from all rounds into a coherent report. Select the synthesis template matching the question type from Phase 0. For compound questions, use the template for each sub-question's type, then add an overall synthesis section.

End-of-sequence awareness: Draft Confidence Assessment and Limitations sections early, not last. Review final paragraphs specifically for unsourced claims.

Template Selection

Read the template file matching the question type from Phase 0. Each template includes the full report structure (executive summary, type-specific body, confidence assessment, limitations, sources). For compound questions, read the template for each sub-question's type and add an overall synthesis section.

Question Type Template File
Factual references/templates/factual.md
Scientific/Health references/templates/scientific-health.md
Consumer references/templates/consumer.md
Technical references/templates/technical.md
Opinion/Sentiment references/templates/opinion-sentiment.md
Contested references/templates/contested.md
Emerging/Frontier references/templates/emerging-frontier.md

Phase 6: Persist Report

Write the final report:

Write: {output_dir}/{topic-slug}/report.md

Update state.md status to COMPLETE:

Edit: {output_dir}/{topic-slug}/state.md
  old_string: "## Status: TRIAGE_COMPLETE"
  new_string: "## Status: COMPLETE"

Format output files (optional): If prettier is available, run it on all markdown files in the output directory to normalize formatting:

prettier --write --prose-wrap preserve "{output_dir}/{topic-slug}/**/*.md"

If prettier is not installed, skip this step silently -- it is cosmetic, not functional.

Inform the user: "Report saved to {output_dir}/{topic-slug}/report.md."

Phase 7: Cleanup

Shut down the team cleanly.

Step 1: Shut down researchers

Send shutdown_request to each researcher via SendMessage:

SendMessage:
  type: "shutdown_request"
  recipient: "researcher-a"
  content: "Research complete. Shutting down."

Repeat for each researcher. Wait for shutdown responses.

Step 2: Read researcher feedback

After all researchers have shut down, read any feedback files written to {output_dir}/{topic-slug}/researcher-{letter}-feedback.md. These contain notes on tool usage (Exa parameters, Firecrawl usage), issues encountered (400 errors, rate limits), and suggestions. Use this feedback to identify patterns for skill improvement.

Step 3: Delete team

TeamDelete

Writing Standards

  • Prose paragraphs, not bullet lists (bullets only for distinct enumerations)
  • Specific data: "increased 23%" not "increased significantly"
  • Cite inline with markdown footnotes: "The market grew 15%[^1]" not "The market grew.[^1]"
  • Each finding: 2-4 paragraphs with evidence
  • Distinguish FACTS (cited) from ANALYSIS (synthesis)

Anti-Hallucination Protocol

  • Every factual claim must cite a source immediately
  • Mark synthesis distinctly: "This suggests..." or "Synthesizing these findings..."
  • If uncertain, say so: "Sources disagree on..." or "Limited evidence for..."
  • Never fabricate sources -- all citations come from researcher findings

Additional Resources

Reference Files

  • references/question-types.md -- 7-type taxonomy, signals, decomposition rules, type-to-scope defaults. Read in Phase 0.
  • references/researcher-prompt.md -- Investigation methodology, type-aware source evaluation, output format. Path provided to researchers in spawn prompt. (Tool guidance extracted to the standalone search-tips skill, which researchers load as their first step.)
  • references/templates/ -- Type-specific synthesis templates. Read the relevant template(s) in Phase 5. See Template Selection table above.

Scripts

  • scripts/analyze-transcripts.py -- Post-hoc analysis of researcher tool usage. Extracts MCP tool call parameters from subagent JSONL transcripts and produces a compliance report. Usage:
    • python3 scripts/analyze-transcripts.py --session ${CLAUDE_SESSION_ID} -- current session
    • python3 scripts/analyze-transcripts.py "topic keyword" -- auto-detect session by keyword
    • python3 scripts/analyze-transcripts.py --list -- list recent sessions with subagents

Development History

  • dev/RESEARCH.md -- Design rationale with 50+ sources justifying the architecture
  • dev/ITERATION-LOG.md -- 13 iterations of improvement with backlog
  • dev/iterations/ -- Detailed notes for each iteration

Quick Reference

  1. Classify question type (silent) and check for resume
  2. Clarify -- understand the question, resolve ambiguity, confirm plan if high-investment
  3. Calibrate effort: announce scope and researcher count
  4. Setup: create output dir, TeamCreate, TaskCreate per angle, spawn researchers
  5. Investigation loop: investigate → triage → dispatch follow-ups or exit. Repeat until converged.
  6. Synthesize: type-aware template, calibrated confidence language
  7. Persist: write report.md, update state.md to COMPLETE, tell user file location
  8. Cleanup: shutdown_request to each researcher, then TeamDelete

Context budget:

  • Lead context: reserve for triage + synthesis
  • Researcher contexts: handle all search/scrape operations
  • Researcher spawn prompts: compact (~30 lines), point to reference file
  • Researcher findings: structured summaries (~120 lines each)
Files (nix-config)
  • dev
    • iterations
      • iteration-01.md 7.8 KB
        # Iteration 1 -- Compound Technical+Consumer (Empirical)
        
        ## Plan
        
        **Type:** Empirical test (compound question)
        
        **Question:** "What are the leading open-source LLM inference engines, and which is best
        suited for a small team deploying a 70B parameter model on consumer GPUs?"
        
        **What this tests:**
        
        - Phase 0: Question type classification (should detect Technical + Consumer compound)
        - Phase 1: Compound decomposition and presentation to user
        - Phase 2: Scope calibration (compound should escalate to Broad)
        - Phase 3: Team setup, output directory creation, task creation
        - Phase 4-6: Multi-round investigation with follow-up
        - Phase 8: Mixed synthesis (Technical comparison + Consumer recommendation)
        - Phase 9: File persistence (state.md, researcher findings, report.md)
        
        **Rationale:** First-ever empirical run of the v2 skill. A compound question exercises the
        most distinctive new features. The topic is genuinely useful and has enough complexity to
        stress-test the system without being so niche that sources are sparse.
        
        ## Research Summary
        
        Ran Broad-scope research (3 researchers, 2 rounds) on "What are the leading open-source LLM
        inference engines, and which is best suited for a small team deploying a 70B parameter model
        on consumer GPUs?" Full report at `~/.claude/research/llm-inference-engines-70b-consumer-gpu/report.md`.
        
        **Round 1:** 3 researchers investigated landscape (A), benchmarks (B), practical deployment (C).
        Strong convergence on vLLM for multi-GPU, llama.cpp/Ollama for single-GPU. Good source diversity
        (25+ sources across researchers).
        
        **Round 2:** 2 targeted follow-ups -- SGLang vs vLLM comparison (B) and ExLlamaV3 viability (A).
        Resolved conflicting throughput claims and assessed ExLlamaV3 maturity.
        
        **Outcome:** Comprehensive report with Technical comparison matrix + Consumer ranked recommendations.
        25 cited sources. Report quality is genuinely useful.
        
        ## Observations
        
        **What worked well:**
        
        1. **Phase 0 classification worked correctly.** Identified compound (Technical + Consumer), chose
           Broad scope. No friction.
        
        2. **Researcher independence produced good diversity.** Each researcher found unique sources with
           minimal overlap. Cross-referencing in triage revealed genuine convergence (not echo chamber).
        
        3. **Round 2 follow-ups were well-targeted.** The SGLang vs vLLM task resolved a real conflict
           from Round 1 (3.1x claim traced to non-independent benchmarks). ExLlamaV3 assessment filled
           a genuine gap.
        
        4. **Source quality was generally high.** Researchers followed type-aware source evaluation --
           prioritized official docs, controlled benchmarks, and practitioner reports over marketing.
        
        5. **File persistence worked.** state.md, researcher findings, and report.md all written correctly.
        
        **What was clunky or problematic:**
        
        1. **Researchers didn't self-claim follow-up tasks.** The skill says researchers should check
           TaskList after completing their initial task and self-claim available work. In practice, all
           3 researchers went idle after Round 1 and had to be explicitly nudged with SendMessage to
           pick up Round 2 tasks. Researcher B went idle even after being sent a direct message about
           Task #7 and needed a second nudge. **This is the biggest friction point.**
        
        2. **Compound decomposition wasn't presented in the exact template format.** The skill says to
           present "> This breaks down into {N} sub-questions: ..." but I presented it slightly
           differently. Minor -- the user understood fine.
        
        3. **No researcher findings files from Round 1.** The skill says researchers should write to
           `{output_dir}/researcher-{letter}-findings.md`. Only Researcher B wrote a findings file for
           Round 2 (SGLang). Researchers A, B, C did NOT persist their Round 1 findings to disk. The
           findings were only delivered via SendMessage. **This breaks the resume capability** -- if the
           session crashed after Round 1, only state.md would exist, not the underlying findings.
        
        4. **Researcher C was underutilized in Round 2.** With only 2 follow-up tasks and 3 researchers,
           Researcher C sat idle. The skill doesn't give clear guidance on whether to shut down unused
           researchers early or find work for them.
        
        5. **Team task list vs iteration loop task list collision.** The research team uses TaskCreate
           for research angles, but the iteration loop also uses TaskCreate for tracking iteration
           progress. These two task lists coexisted in the same session (team tasks #1-8, iteration
           tasks #1-5 in a different context). This worked but could be confusing in a resumed session.
        
        6. **The skill doesn't specify the lead's name for researchers to message.** Researchers need
           to read the team config to find the lead's name. The researcher-prompt.md says "The lead's
           name is typically the team creator. Check the team config if unsure." This is indirect --
           could be explicit in the spawn prompt.
        
        7. **Synthesis was a single monolithic write.** At ~250 lines, the report was written in one
           Write call. For longer reports, this could fail or be hard to review. The skill doesn't
           address progressive synthesis or draft-then-refine.
        
        **Surprising or noteworthy:**
        
        - Researchers found genuinely useful, current sources (Jan-Feb 2026 content). Exa semantic
          search worked well for this technical topic.
        - The compound question naturally produced a well-structured report (Part 1: Technical
          Comparison, Part 2: Consumer Recommendation). The type-aware templates mapped cleanly.
        - Total research time was reasonable for a Broad scope run (~10-15 minutes wall clock).
        
        ## Changes Made
        
        **1. Restructured researcher reporting flow** (`references/researcher-prompt.md`)
        
        - Replaced the old "Reporting Findings" + "Task Management" sections with a clear 3-step
          sequence: (1) persist to disk, (2) SendMessage to lead, (3) mark complete + TaskList.
        - Moved file persistence from buried rule #9 to Step 1 of the main reporting flow.
        - Added explicit filenames for follow-up tasks: `researcher-{letter}-findings-{task-id}.md`.
        
        **2. Added emphatic task self-claim section** (`references/researcher-prompt.md`)
        
        - New "Task Self-Claim (IMPORTANT)" section with explicit instructions to always check
          TaskList after completing each task and immediately claim available work.
        - Added "Do NOT go idle if there are available tasks" instruction.
        
        **3. Updated Critical Rules** (`references/researcher-prompt.md`)
        
        - Replaced rules #6-9 with two focused rules: "Persist, THEN report, THEN check for work"
          and "ALWAYS self-claim available tasks". Removed redundant rule about SendMessage (now
          covered in main flow).
        
        **4. Added lead name to spawn prompt template** (`SKILL.md`)
        
        - Added `Lead name: team-lead` line to the researcher spawn prompt so researchers don't
          need to look up the team config to know who to message.
        - Added `IMPORTANT: After completing each task, ALWAYS check TaskList...` reminder at the
          end of the spawn prompt for reinforcement.
        
        **5. Added settings.json permission** (`configs/claude/settings.json`)
        
        - Added `Bash(mkdir -p ~/.claude/research/*)` to avoid permission prompts during research
          output directory creation.
        
        ## Notes for Next Iteration
        
        - The self-claim fix is the most important change but also the hardest to verify -- it depends
          on researcher agent behavior, which may not fully follow instructions regardless of emphasis.
          Iteration 2 should test whether the changes actually reduce the nudging problem.
        - Compound synthesis quality was good -- the two-part structure (Technical + Consumer) mapped
          naturally. No changes needed to synthesis templates yet.
        - The Contested question test (backlog #2) would exercise completely different skill features
          (steelmanning, cross-agent verification, Comprehensive scope). Good candidate for Iteration 2.
        - New backlog idea: investigate whether spawn prompt length affects researcher behavior.
          The prompt is already ~15 lines; adding more "IMPORTANT" reminders has diminishing returns.
        
      • iteration-02.md 10.3 KB
        # Iteration 2 -- Contested Question (Empirical)
        
        ## Plan
        
        **Type:** Empirical test (Contested question)
        
        **Question:** "Should AI model weights be open-sourced?"
        
        **What this tests:**
        
        - Phase 0: Contested classification (should trigger Comprehensive scope)
        - Phase 1: Research plan presentation (Contested template with angles)
        - Phase 2: Comprehensive calibration (4 researchers, 2-3 rounds, opus model)
        - Phase 7: Cross-agent verification (unique to Comprehensive)
        - Phase 8: Contested synthesis template (steelmanning, no verdict)
        - **Iteration 1 fixes:** researcher self-claim, file persistence, lead name in spawn prompt
        
        **Also validates:**
        
        - Context compaction resilience (we're at ~60% context, likely to compact mid-run)
        - Whether emphasizing self-claim in researcher-prompt.md reduces nudging
        
        ## Research Summary
        
        Ran Comprehensive-scope research (4 researchers, 3 rounds) on "Should AI model weights be
        open-sourced?" Full report at `~/.claude/research/ai-model-weights-open-source/report.md`.
        
        **Round 1:** 4 opus researchers investigated: (A) pro-openness arguments, (B) anti-openness
        arguments, (C) regulatory/governance landscape, (D) empirical evidence. Strong convergence on
        key structural points (irreversibility, safeguard removability, safety research enablement).
        60+ sources across researchers.
        
        **Round 2:** 2 targeted follow-ups -- graduated release proposals (A) and RAND CBRN contradiction
        (D). Resolved the RAND 2024 vs 2025 discrepancy (different models, different methodologies,
        both partially right). Found rich graduated release literature (Solaiman gradient, structured
        access, Carnegie consensus, MOF, RSPs).
        
        **Round 3 (Verification):** 2 cross-agent verifications --
        
        - SentinelOne 7.5% claim (verified by B): SUPPORTED WITH NUANCE. Lower bound on a subset,
          corroborated by 2 independent studies (Censys solo, Cisco Talos).
        - Greenblatt ~100K fatalities estimate (verified by C): CONTESTED. Independent models produce
          ~10K-60K range, but direction agreed and uncertainty intervals overlap.
        
        **Outcome:** 37-footnote Contested synthesis report with steelmanned positions, points of
        agreement, key disagreements (factual and values-based), evidence quality assessment. No verdict
        rendered. Report quality is genuinely useful for someone trying to understand the debate.
        
        ## Observations
        
        **What worked well:**
        
        1. **Phase 0 classification worked correctly.** Identified Contested type, selected Comprehensive
           scope (4 researchers, opus model, 2-3 rounds + verification). No friction.
        
        2. **File persistence fix from Iteration 1 worked perfectly.** All 4 researchers wrote findings
           to disk in Round 1: `researcher-a-findings.md`, `researcher-b-findings.md`,
           `researcher-C-findings.md`, `researcher-D-findings.md`. Round 2 follow-ups also persisted:
           `researcher-a-findings-9.md`, `researcher-D-findings-10.md`. Verifications persisted:
           `researcher-b-findings-12.md`, `researcher-C-findings-11.md`. **This is a clear win from
           Iteration 1 changes.**
        
        3. **Cross-agent verification added genuine value.** The Greenblatt CBRN verification (CONTESTED)
           surfaced 4 independent quantitative models (GovAI/Righetti, GovAI/Williams, Jones, RAND Delphi)
           that put the estimate in context. The SentinelOne verification (SUPPORTED WITH NUANCE) traced
           the claim to its primary source and found Reuters had oversimplified. Both verdicts materially
           improved the final report.
        
        4. **Researcher quality was excellent at opus level.** Source diversity, depth of analysis, and
           critical evaluation were noticeably stronger than the sonnet researchers in Iteration 1.
           Researchers proactively labeled institutional biases, distinguished methodology types, and
           cross-referenced within their own findings.
        
        5. **Contested synthesis template worked well.** The steelmanned positions format naturally
           organized the findings. Separating factual disagreements from values disagreements was
           particularly effective. The "Evidence Quality by Position" section forced honest assessment
           of both sides' weaknesses.
        
        6. **Lead name in spawn prompt worked.** Researchers messaged "team-lead" directly without
           needing to look up the team config. No friction observed.
        
        7. **Context compaction survived.** The session compacted mid-run (between Round 1 triage and
           Round 2 dispatch). The ITERATION-LOG.md and state.md files were sufficient to resume
           coherently. The compacted context summary preserved key details about researcher findings
           and verification assignments.
        
        **What was clunky or problematic:**
        
        1. **Self-claim fix still did not work reliably.** Despite the emphatic "Task Self-Claim
           (IMPORTANT)" section and the spawn prompt reminder, researchers still went idle after
           completing their Round 1 tasks without checking TaskList. The lead had to explicitly create
           follow-up tasks and (in the pre-compaction portion) likely nudge researchers to claim them.
           After compaction, the lead dispatched Round 2 and Round 3 tasks by sending direct messages
           to specific researchers rather than relying on self-claim. **This is now a confirmed pattern
           across 2 iterations: researchers do not reliably self-claim tasks regardless of how
           emphatically the instructions say to do so.**
        
        2. **Inconsistent filename casing.** Researchers wrote `researcher-a-findings.md` (lowercase)
           and `researcher-C-findings.md` (uppercase). The researcher-prompt.md says
           `researcher-{letter}-findings.md` without specifying case. The spawn prompt uses uppercase
           letters (A, B, C, D) while the template uses lowercase `{letter}`. This inconsistency is
           cosmetic but could cause issues for scripts or resume logic that expect consistent naming.
        
        3. **Researcher C required double shutdown request.** The first shutdown_request was sent but
           C didn't respond until a second was sent. This may be a timing issue (C was still processing)
           rather than a systematic problem. Researchers A, B, D all shut down on first request.
        
        4. **Round 2 researcher allocation was suboptimal.** With 2 follow-up tasks and 4 researchers,
           B and C sat idle during Round 2. Then in Round 3, B and C were assigned verification tasks.
           The skill doesn't provide guidance on whether to assign Round 2 tasks to the SAME researchers
           who covered related Round 1 angles (continuity) or DIFFERENT researchers (fresh perspective).
           For follow-ups, continuity seems better; for verification, fresh perspective is required.
        
        5. **No guidance on when to stop adding rounds.** The skill says "2-3 rounds" for Comprehensive
           but doesn't specify criteria for choosing 2 vs 3. In practice, verification (Round 3) added
           clear value for the two high-impact single-source claims. But the decision was ad hoc.
        
        **Surprising or noteworthy:**
        
        - Opus researchers found genuinely high-quality, current sources including January-February 2026
          content. The contested topic prompted researchers to find primary sources rather than relying
          on secondary reporting.
        - The RAND CBRN contradiction investigation was the most valuable follow-up: it found a THIRD
          RAND study (Delphi panel) and a FOURTH (benchmarking study) that helped reconcile the apparent
          contradiction. This would have been missed without targeted follow-up.
        - Report length (~350 lines, 37 footnotes) is longer than ideal. The Contested template naturally
          produces longer output because steelmanning both sides requires more space. The skill doesn't
          address length targets for different question types.
        
        ## Changes Made
        
        **1. Replaced self-claim with lead-dispatched follow-ups** (`SKILL.md` Phase 6, `researcher-prompt.md`)
        
        - The biggest architectural change. After 2 iterations confirming researchers don't self-claim
          regardless of instruction emphasis, switched to a model where the lead creates tasks AND
          explicitly messages each researcher to assign them.
        - SKILL.md Phase 6: Added "Assignment strategy" section with guidance on who gets follow-ups
          (continuity for deep-dives, fresh perspective for conflicts).
        - SKILL.md Phase 7: Added same lead-dispatched pattern for verification tasks.
        - researcher-prompt.md: Replaced emphatic "Task Self-Claim (IMPORTANT)" section with brief
          "Between Tasks" section. Now says: check TaskList once, claim if available, otherwise go
          idle and wait for lead to message you.
        - researcher-prompt.md: Updated Critical Rules to remove self-claim rule.
        
        **2. Normalized filename casing to lowercase** (`SKILL.md`, `researcher-prompt.md`)
        
        - Spawn prompt now says `(use LOWERCASE in filenames: researcher-{lowercase letter})`
        - researcher-prompt.md Step 1: Added "(always use **lowercase** for the letter)"
        - researcher-prompt.md Critical Rules: Added rule #7 "Use lowercase filenames"
        
        **3. Added round stopping criteria** (`SKILL.md` Phase 7)
        
        - Added explicit "When to run Round 3" and "When to skip Round 3" criteria to Phase 7.
        - Run verification when: high-impact single-source claims, unresolved factual conflicts,
          or specialized/niche topic claims.
        - Skip verification when: no single-source high-impact claims remain, conflicts resolved by
          Round 2, remaining uncertainties are about magnitude not direction.
        
        **4. Added Round 2 assignment guidance** (`SKILL.md` Phase 6)
        
        - Deep-dives/gap-fills: assign to researcher with related Round 1 angle (continuity)
        - Conflict resolution: assign to researcher who did NOT cover either side (fresh perspective)
        - Idle researchers: wait for verification or shut down early
        
        ## Notes for Next Iteration
        
        - The lead-dispatched model is the most important change to validate in Iteration 3. It
          accepts the reality that researchers don't self-claim and makes the lead's role explicit.
          This should eliminate the "nudging" friction entirely.
        - Iteration 2 demonstrated that Comprehensive scope with opus researchers produces excellent
          quality. The main question is whether the cost (4 opus researchers x 3 rounds) is justified
          for less complex questions.
        - The Contested synthesis template is naturally verbose (~350 lines, 37 footnotes). Consider
          adding length targets per question type in a future iteration.
        - Good candidate for Iteration 3: Emerging/Frontier test (backlog #3) would exercise
          instability warnings, date-sensitive source evaluation, and the Emerging synthesis template.
          It would also test the lead-dispatched model with a different question type.
        - Backlog item #8 (spawn prompt length) can be removed -- the self-claim problem was resolved
          by changing the architecture rather than tweaking prompt emphasis.
        
      • iteration-03.md 8.8 KB
        # Iteration 3 -- Consumer Comparison (Empirical)
        
        ## Plan
        
        **Type:** Empirical test (Consumer question, non-AI domain)
        
        **Question:** "Best mechanical keyboard for programming in 2026?"
        
        **What this tests:**
        
        - Phase 0: Consumer classification (should trigger Broad scope by default)
        - Phase 1: Brief clarification (Consumer type, simple question)
        - Phase 2: Broad calibration (3 researchers, 2 rounds, sonnet model)
        - Phase 8: Consumer synthesis template (ranked recommendations)
        - **Iteration 2 fix:** lead-dispatched follow-ups (the biggest change to validate)
        - **Domain shift:** non-AI topic tests whether Exa search and researcher methodology work
          outside our comfort zone
        
        **Also validates:**
        
        - Expert testing lab prioritization in source evaluation (Consumer type hierarchy)
        - Whether sonnet researchers (cheaper) produce adequate quality for a simpler topic
        - Whether Broad scope is sufficient for a Consumer question (or if it needs less)
        
        ## Research Summary
        
        Ran Broad-scope research (3 sonnet researchers, 2 rounds) on "Best mechanical keyboard for
        programming in 2026?" Full report at `~/.claude/research/best-mech-keyboard-programming-2026/report.md`.
        
        **Round 1:** 3 researchers investigated: (A) expert reviews/testing, (B) community
        recommendations, (C) features/specs analysis. Strong convergence on Keychron dominance at
        every price tier. 30+ sources across researchers.
        
        **Round 2:** 2 targeted follow-ups using lead-dispatched model:
        
        - Split/ergonomic deep-dive (C, continuity): Consolidated Voyager vs Glove80 vs Defy vs Go60.
          Discovered MoErgo Go60 as significant new entrant. Alice layout as middle ground.
        - Blind spot check (A, fresh perspective): Mac-specific picks (NuPhy Air75 V2 top for Mac),
          office noise reality (no mechanical keyboard is truly quiet), Leopold as overlooked quality.
        
        **Outcome:** Consumer-template report with ranked recommendations (Top Pick: Q5 Max, Runner-Up:
        V5 Max, Budget: C3 Pro, Best Mac: NuPhy Air75 V2, Best Ergonomic: Glove80, Best Gentle
        Upgrade: Keychron Q8 Max Alice, Best Quiet: HHKB Type-S). 25 cited sources. Report is
        genuinely useful as a buying guide.
        
        ## Observations
        
        **What worked well:**
        
        1. **Lead-dispatched follow-ups worked perfectly on the first try.** This is the most important
           validation from this iteration. The lead created 2 Round 2 tasks, then sent direct
           SendMessage to researcher-c (split deep-dive, continuity) and researcher-a (blind spots,
           fresh perspective). Both researchers claimed their tasks and began immediately with zero
           nudging. Researcher-b was correctly left idle (2 tasks < 3 researchers). **The Iteration 2
           architectural change eliminated the self-claim problem completely.**
        
        2. **Lowercase filename normalization worked.** All 5 findings files used correct lowercase
           naming: `researcher-a-findings.md`, `researcher-b-findings.md`, `researcher-c-findings.md`
           (Round 1), `researcher-c-findings-7.md`, `researcher-a-findings-8.md` (Round 2). The
           Iteration 2 fix (explicit lowercase instructions in spawn prompt + researcher-prompt.md)
           resolved the inconsistency.
        
        3. **Consumer synthesis template produced a well-structured report.** The ranked recommendations
           format (Top Pick, Runner-Up, Budget, Best for X) naturally organized the findings into
           actionable buying advice. Each pick has evidence from multiple sources and clear reasoning.
        
        4. **Non-AI domain worked fine.** Exa semantic search handled keyboard/hardware topics well.
           Researchers found high-quality sources including RTINGS lab testing (278 keyboards), Wirecutter,
           Reviewed.com, WIRED, plus community sources (Reddit 381K opinion analysis, HN, Devtalk).
           No domain-specific search issues.
        
        5. **Sonnet researcher quality was adequate for Consumer scope.** Findings were well-structured,
           sources properly evaluated, and cross-source analysis was competent. The quality difference
           from opus (Iteration 2) was noticeable but acceptable: sonnet researchers were less likely
           to proactively label biases or trace claims to primary sources, but the overall research
           was solid.
        
        6. **All 3 researchers shut down cleanly on first request.** No double-shutdown needed (unlike
           Iteration 2 where researcher-c required two attempts). This may have been a timing
           improvement or just variance.
        
        7. **Assignment guidance worked as designed.** Split deep-dive went to researcher-c (who covered
           splits in Round 1 -- continuity). Blind spot check went to researcher-a (who covered expert
           reviews -- fresh perspective on what was missed). The guidance in Phase 6 made assignment
           decisions easy and justified.
        
        **What was clunky or problematic:**
        
        1. **Round 1 had no gaps requiring urgent follow-up.** The Broad scope consumer question
           produced such strong convergence in Round 1 (Keychron dominance was overwhelming) that
           Round 2 follow-ups were useful but not essential. The split/ergonomic deep-dive and blind
           spot check added genuine value (Go60 discovery, Mac-specific picks, noise reality check),
           but the core recommendations would have been the same without them. **This suggests Focused
           scope (1 round) might be sufficient for straightforward Consumer questions.**
        
        2. **Researcher A idle-looped before claiming task.** After receiving the follow-up assignment
           via SendMessage, Researcher A sent 2 idle notifications before starting work. The task was
           eventually claimed and completed successfully, so this is cosmetic, but the idle chatter is
           noisy in the notification stream.
        
        3. **Report length was ~200 lines, 25 footnotes.** More manageable than the Contested report
           (~350 lines, 37 footnotes) but still substantial. The Consumer template's multiple "Best
           for X" categories naturally expand the report. No length target exists in the skill.
        
        4. **Phase 1 clarification was completely skipped.** The skill says "Brief clarification only
           if genuinely ambiguous" for Consumer questions. The question was clear enough to skip
           entirely, which is correct behavior, but it means the user had no opportunity to specify
           constraints (budget range, layout preference, Mac vs Windows, etc.) that could have focused
           the research. For Consumer questions, a quick "Any constraints?" might be valuable even
           when the question seems clear.
        
        **Surprising or noteworthy:**
        
        - The dharm.is analysis of 381K Reddit opinions (found by researcher-b) was an unexpectedly
          rich quantitative source for a Consumer topic. It revealed the enthusiast vs practical
          programmer preference gap that informed the report's framing.
        - The MoErgo Go60 (found by researcher-c in Round 2) was genuinely new information -- launched
          late 2025, only one detailed review exists. This is exactly the kind of find that makes
          Round 2 follow-ups worthwhile even when Round 1 converges.
        - Total wall-clock time was ~15 minutes for the full Broad scope run, similar to Iteration 1.
        
        ## Changes Made
        
        **No skill file changes this iteration.** The Iteration 2 changes (lead-dispatched follow-ups,
        lowercase filenames, round stopping criteria, assignment guidance) all validated successfully.
        The observations suggest potential improvements but none urgent enough to change now:
        
        - Scope calibration for simple Consumer questions (Focused might suffice) -- worth investigating
          as a meta-research topic rather than a code change
        - Consumer-specific Phase 1 constraint gathering -- minor UX improvement, not blocking
        - Idle notification noise -- cosmetic, not actionable in the skill
        
        ## Notes for Next Iteration
        
        - **Lead-dispatched model is confirmed working across 2 question types.** The self-claim
          problem is solved. No further changes needed to this architecture.
        - **All Iteration 2 changes validated.** Lowercase filenames, assignment guidance, and round
          stopping criteria all worked as designed. The skill is now stable for these features.
        - **Good candidate for Iteration 4:** Meta-research on scope calibration. Three iterations
          have now produced data on scope appropriateness: Iteration 1 (Broad for compound Technical +
          Consumer, felt right), Iteration 2 (Comprehensive for Contested, felt right but expensive),
          Iteration 3 (Broad for Consumer, Round 2 was useful but not essential). Could the skill
          provide better guidance on when Focused is sufficient vs when Broad is needed?
        - **Alternative candidate:** Emerging/Frontier test (backlog #3) to exercise the last untested
          synthesis template. Or information cascade detection (backlog #4) as the first meta-research
          iteration.
        - **The skill is approaching diminishing returns for empirical testing.** Three successful runs
          across 3 question types (compound Technical+Consumer, Contested, Consumer) with improving
          results each iteration. The remaining untested templates (Scientific/Health, Emerging/Frontier,
          Opinion/Sentiment, Factual) follow the same architecture. Meta-research may yield higher
          marginal value than additional empirical runs.
        
      • iteration-04.md 6.4 KB
        # Iteration 4 -- Search Tool Audit + Access Workarounds (Meta-Research)
        
        ## Plan
        
        **Type:** Meta-research (run through the team skill, not manual)
        **Backlog item:** #4
        **Scope:** Broad (3 sonnet researchers, 2 rounds)
        
        **Goal:** Audit the full Exa and Firecrawl tool sets for underused features, research
        academic search options and paywall workarounds, then update `researcher-prompt.md` with
        improved guidance.
        
        ## Research Summary
        
        **Report:** `~/.claude/research/web-research-tool-effectiveness/report.md`
        
        Round 1 angles: (A) Exa/Firecrawl advanced features, (B) content access workarounds,
        (C) academic search strategies. Round 2 follow-ups: (A) optimal Exa+Firecrawl workflow
        decision tree, (B) blind spot check (OpenAlex, Firecrawl limitations, additional APIs).
        Researcher C shut down early (no follow-up needed). 24+ sources, 5 findings per researcher.
        
        **Critical finding:** Both Exa and Firecrawl MCP tools expose only a subset of their full
        API capabilities. **Resolved post-iteration:** Switched Exa from npm package to hosted HTTP
        endpoint (`mcp.exa.ai/mcp`) with `web_search_advanced_exa` and `people_search_exa` enabled.
        This exposed all missing parameters (category, includeDomains/excludeDomains, date filtering,
        highlights/summary). Firecrawl search was then removed from the researcher prompt as
        redundant -- Exa advanced covers domain targeting, date filtering, and category filtering
        natively, reducing researcher cognitive load from 5 tools to 4.
        
        ## Observations
        
        **What worked well:**
        
        - First meta-research iteration run through the team skill -- worked as well as empirical
          runs. The compound Technical framing with 3 angles mapped naturally to the 4 sub-questions.
        - Lead-dispatched follow-ups: zero nudging needed (4th consecutive success).
        - Researchers independently discovered the same critical insight (Exa MCP limitations) from
          different angles -- genuine cross-validation.
        - Researcher A did empirical testing in Round 2 (actually ran comparative Exa vs Firecrawl
          queries), adding real evidence beyond documentation review.
        - Early shutdown of Researcher C was clean and efficient.
        - Researcher B found significant blind spots in Round 2 (OpenAlex, Bluesky API, Firecrawl
          Cloudflare issues, underused Firecrawl formats) -- the follow-up was genuinely valuable.
        
        **Process observations:**
        
        - Meta-research through the team skill is viable and arguably better than manual research --
          the multi-angle approach and forced cross-referencing catch things manual research would miss.
        - The decision to frame "search tool audit" as a Technical compound question worked well.
          Could serve as a template for future meta-research iterations.
        
        **No problems found.** The skill ran smoothly across all phases.
        
        ## Changes Made
        
        **5 changes to `researcher-prompt.md`:**
        
        1. **Expanded "Available Tools" section** -- Added Firecrawl search to the tool list, added
           a "When to use which" decision guide distinguishing Exa (semantic discovery) from Firecrawl
           search (precision targeting with operators) from Firecrawl scrape (content extraction).
        
        2. **Added "Firecrawl Search Operators" section** -- Table of Google-style operators (site:,
           exact match, exclusion, intitle:, inurl:, tbs date filtering with examples). Includes
           caveat about tbs filtering by index date vs publication date.
        
        3. **Added "Academic Search" section** -- Three-step workflow (Exa for broad discovery ->
           Firecrawl search with site: for domain targeting -> Firecrawl scrape for full content).
           Lists useful academic domain targets. Lists free academic APIs (Semantic Scholar, OpenAlex,
           arXiv) with URL patterns.
        
        4. **Expanded "Content Extraction" section** -- Added advanced scrape features: summary format
           for triage, links format for citation chains, JSON extraction with schema, includeTags/
           excludeTags, waitFor for SPAs, proxy options, PDF parsing, maxAge caching. Added "if scrape
           returns empty" workflow.
        
        5. **Added "Accessing Restricted Content" section** -- Complete workaround guide: archive.ph
           for paywalls (with URL patterns), xcancel.com for Twitter/X, Reddit .json endpoint,
           Freedium for Medium, Bluesky public API, LinkedIn limitations. Each with exact URL patterns.
        
        **1 additional change:** 6. **Updated Setup section** -- Added `ToolSearch: "+firecrawl search"`
        to load Firecrawl search tool alongside scrape.
        
        **Post-checkpoint fix:** Prettier mangled `>` characters in source evaluation priority chains
        (e.g., `documentation > code examples > blog posts`) into markdown blockquotes. Replaced all
        `>` separators with ", then" phrasing across all 6 type-specific entries for prettier safety.
        
        **Post-iteration infrastructure change (backlog #18 resolved):**
        Switched Exa MCP from npm package (`exa-mcp-server`) to hosted HTTP endpoint
        (`mcp.exa.ai/mcp`) with `web_search_advanced_exa` and `people_search_exa` enabled. This
        exposed all parameters identified as missing during Iteration 4 research. Firecrawl search
        was then removed from the researcher prompt entirely -- Exa advanced covers domain targeting,
        date filtering, and category filtering natively. Net effect: researchers now have 4 tools
        instead of 5, simpler decision-making, and more powerful search capabilities.
        
        **Post-iteration bug discovery (contextMaxCharacters suppresses summaries/highlights):**
        Investigated why `enableHighlights` and `enableSummary` produced no visible output through
        the Exa MCP endpoint. Read the npm package source and found two bugs in the response
        formatter: (1) When `contextMaxCharacters` is set, the API returns a `context` field that
        the formatter uses exclusively, dropping per-result data including summaries and highlights.
        (2) The `ExaSearchResult` type doesn't include `highlights` at all -- they're requested from
        the API but silently dropped. **Fix:** Use `textMaxCharacters: 1` + `enableSummary: true`.
        Do NOT set `contextMaxCharacters`.
        
        **Post-iteration tool simplification:**
        Reduced Exa tools from 5 to 2: `web_search_advanced_exa` + `get_code_context_exa`.
        
        **Post-iteration category guidance:**
        Added category restriction table and usage hints from Exa skills repo.
        
        ## Notes for Next Iteration
        
        - Researcher-prompt.md is ~420 lines. Monitor for issues.
        - Default Exa search pattern (`enableSummary: true` + `textMaxCharacters: 1`) untested
          in empirical run. First test will reveal whether researchers follow the guidance.
        - Category filter restrictions documented but untested.
        - Academic search strategies documented but untested in Scientific/Health run.
        
      • iteration-05.md 3.6 KB
        # Iteration 5 -- Scientific/Health Empirical (IF + Cognitive Performance)
        
        ## Plan
        
        **Type:** Empirical (Scientific/Health question)
        **Backlog item:** #12
        **Scope:** Broad (3 sonnet researchers, 2 rounds)
        
        **Goal:** Test the evidence pyramid synthesis template, academic search guidance, and the
        new Exa `enableSummary` + `textMaxCharacters: 1` default pattern. Also test whether the
        category filter restrictions table prevents 400 errors. Secondary goal: validate that the
        architecture continues to work smoothly (6th consecutive run).
        
        **Question:** "Does intermittent fasting improve cognitive performance?"
        
        **Angles:**
        - (A) Systematic reviews and meta-analyses (anchor evidence)
        - (B) Mechanistic pathways (BDNF, neuroinflammation, ketones, gut-brain axis)
        - (C) Real-world protocols, practitioner evidence, and cognitive risks
        
        ## Research Summary
        
        **Report:** `~/.claude/research/intermittent-fasting-cognitive-performance/report.md`
        
        Round 1: All 3 researchers completed successfully. Researcher A found the Bamberg & Moreau
        2025 meta-analysis (63 studies, 3,484 participants) as the anchor -- acute fasting is
        cognitively neutral. Researcher B mapped 5 mechanistic pathways with 13 sources.
        Researcher C covered protocols, null results, and disordered eating risks with 11 sources.
        
        Triage identified: (1) CCR vs IF conflict (Researcher A's O'Leary review claimed CCR
        superior, but other evidence was more nuanced), (2) population-specificity gap (which
        populations benefit most?).
        
        Round 2: Researcher B (fresh perspective) resolved the CCR vs IF conflict -- O'Leary
        conflated fasting types; head-to-head trials show no difference when calorie-matched.
        Researcher C (continuity) deep-dived population specificity -- metabolically impaired
        populations benefit most, age is a moderator, ApoE4 data is missing.
        
        Final report: ~300 lines, 26 footnotes, Scientific/Health evidence pyramid template.
        
        **Key conclusions (calibrated confidence):**
        - Acute fasting is cognitively neutral in healthy adults (highly likely, 80-92%)
        - Chronic IF evidence is insufficient to claim cognitive benefit (highly likely)
        - Metabolically impaired populations benefit most (likely, 63-79%)
        - IF ≈ CCR when calorie-matched (likely)
        - Mechanistic pathways are plausible but human translation is incomplete (likely)
        
        ## Observations
        
        **What worked well:**
        
        - Evidence pyramid template worked excellently.
        - Lead-dispatched follow-ups: 6th consecutive success, zero nudging needed.
        - Academic source quality was strong (APA, Cell Metabolism, Nature, BMJ Gut, Springer, MDPI).
        - Triage cross-referencing caught genuine issues.
        - Calibrated confidence language worked naturally.
        - All researchers completed on first try.
        
        **Exa parameter compliance:** 32/32 Exa calls fully compliant with summary pattern. Categories
        used: `research paper` (frequent), `personal site` (once). No 400 errors. 28 Firecrawl scrape
        calls, all with `onlyMainContent: true`.
        
        ## Changes Made
        
        **3 changes (implementing backlog #23 -- researcher shutdown feedback reports):**
        
        1. Updated Shutdown section in `references/researcher-prompt.md` with feedback file template.
        2. Added Step 2 to Phase 10 (Cleanup) in `SKILL.md` for reading feedback files.
        3. Added `scripts/analyze-transcripts.py` for post-hoc JSONL analysis.
        
        **Backlog items resolved:** #12 (Scientific/Health empirical), #23 (shutdown feedback reports).
        
        ## Notes for Next Iteration
        
        - Feedback reports implemented but untested. Next run will validate.
        - Researcher-prompt.md is 438 lines. Monitor.
        - Evidence pyramid validated. No changes needed.
        - Exa guidance: 32/32 compliant. No further investment needed.
        - Category restrictions untested with restricted categories (tweet, company, people).
        
      • iteration-06.md 6.9 KB
        # Iteration 6 -- Emerging/Frontier Empirical (LLM Reasoning)
        
        ## Plan
        
        **Type:** Empirical (Emerging/Frontier question)
        **Backlog item:** #9
        **Scope:** Comprehensive (4 opus researchers, 3 rounds incl. verification)
        
        **Goal:** Test the last major untested synthesis template (Emerging/Frontier with instability
        warnings, date-sensitive source evaluation). First run with shutdown feedback reports
        (implemented in Iteration 5 but untested). Exercise Exa categories beyond `research paper`
        (`news`, `personal site`, `tweet`), testing the category restriction table.
        
        **Question:** "What are the current approaches to LLM reasoning?"
        
        **Angles:**
        
        - (A) Foundational paradigms: CoT, ToT, GoT, and their successors
        - (B) Training-time reasoning: PRM, RL, GRPO, RLVR, reasoning-focused training
        - (C) Inference-time compute scaling: search, verification, test-time compute allocation
        - (D) Emerging approaches: neurosymbolic, multi-agent, tool-augmented, latent reasoning
        
        ## Research Summary
        
        **Report:** `~/.claude/research/llm-reasoning-approaches/report.md`
        
        **Round 1:** All 4 opus researchers completed successfully. Strong convergence on RLVR as
        dominant paradigm shift (A, B, C). Latent reasoning identified as frontier by A and D
        independently. B provided deep GRPO/RLVR coverage. C mapped inference-time compute scaling
        with strong evidence. D covered neurosymbolic, multi-agent, tool-augmented, latent, and
        world model approaches. Total: ~60 sources across researchers.
        
        **Triage identified:** (1) PRM vs ORM conflict (B says PRMs disappoint at scale; C says PRMs
        outperform), (2) all 4 researchers noted heavy academic bias -- practitioner/deployment
        perspectives missing, (3) competitive landscape gap.
        
        **Round 2:** Researcher A (fresh perspective) resolved PRM conflict -- both correct in different
        regimes (inference-time verification: PRM wins; large-scale RL training: ORM preferred).
        Researcher D (gap-fill) surveyed practitioner perspectives and competitive landscape using
        diverse Exa categories (news, personal site, tweet).
        
        **Round 3 (Verification):** Researcher B verified CoT faithfulness claims -- 2.3% True Thinking
        Score CONTESTED (single model/dataset, not replicated; direction supported but specific number
        misleading), <20% verbalization SUPPORTED WITH NUANCE (multiple independent confirmations;
        METR argues manageable for safety monitoring).
        
        **Final report:** ~400 lines, 55 footnotes. Emerging/Frontier template with instability warnings.
        
        ## Observations
        
        **What worked well:**
        
        - **Emerging/Frontier template worked excellently.** The "Current State (as of date)" +
          "Trajectory" + "Instability Warning" structure produced a report that acknowledges the
          rapidly evolving nature of the field while still providing definitive findings.
        - **Lead-dispatched follow-ups:** 7th consecutive success, zero nudging needed.
        - **Cross-agent verification added genuine value.** The CoT faithfulness verification produced
          a nuanced CONTESTED verdict that materially improved the report -- the 2.3% number was
          tempered to a directional finding.
        - **PRM conflict resolution was clean.** Researcher A (fresh perspective) found the resolution
          along the inference vs training axis with 13 independent sources.
        - **All researchers completed on first try.** No stalls, no nudging, no retries.
        - **Category restriction table worked.** `tweet` category used successfully (no 400 errors).
          Researcher D correctly omitted restricted params for tweet queries.
        
        **Exa parameter compliance (verified via `scripts/analyze-transcripts.py`):**
        
        - **49/49 Exa calls (100%) fully compliant.** `enableSummary: true` + `textMaxCharacters: 1`,
          zero `contextMaxCharacters` violations. 7th session with perfect compliance (cumulative
          81/81 across Iterations 5-6).
        - **Categories exercised:** `research paper` (8), `news` (2), `personal site` (2), `tweet` (1).
          First run using `news`, `personal site`, and `tweet` categories. No 400 errors.
        - **Firecrawl:** 22 scrape calls, all `firecrawl_scrape`. No map or other tools used.
        - **Rich parameter usage:** `startPublishedDate` widely used, `numResults` varied (5-10),
          `includeDomains` used once (metr.org).
        
        **Shutdown feedback reports (first test -- backlog #23 validated):**
        
        All 4 researchers wrote feedback files. Key findings from feedback:
        
        - **arxiv scraping noise:** 3 of 4 researchers independently flagged that Firecrawl scraping
          of arxiv abstract pages returns mostly navigation boilerplate. All recommended using
          `arxiv.org/html/` or relying on Exa summaries for triage.
        - **Tweet category thinness:** Researcher D noted tweet summaries were 1-2 sentences with
          limited context. The category works but provides lower signal density.
        - **Duplicate self-referential messages:** 3 of 4 researchers reported receiving their own
          task assignment messages echoed back. Cosmetic but confusing. Likely a platform bug.
        - **Verification verdict granularity:** Researcher B noted the SUPPORTED/CONTESTED/UNCHANGED
          trichotomy doesn't cover partial support. Iteration 2 verifiers also used "SUPPORTED WITH
          NUANCE."
        - **Conflict resolution length:** Researcher A noted ~120-line target is too short for
          conflict resolution tasks that need comparison tables.
        - **Positive:** All researchers praised the `enableSummary` + `textMaxCharacters: 1` pattern
          for efficient triage.
        
        **Feedback reports vs transcript analysis comparison:**
        
        Both mechanisms add unique value. Transcripts provide exact parameter compliance (definitive).
        Feedback reports provide qualitative insights not visible in transcripts: arxiv scraping issue,
        duplicate message bug, tweet category thinness, verification verdict suggestions. **Keep both.**
        
        ## Changes Made
        
        **No skill file changes this iteration.** The Emerging/Frontier template, feedback reports,
        and category restriction table all validated successfully. Two quick changes identified for
        next iteration (backlog #26 arxiv guidance, #27 verification verdict expansion) but not
        blocking.
        
        **Backlog items resolved:** #9 (Emerging/Frontier empirical), #23 validated (feedback reports
        work as designed).
        
        ## Notes for Next Iteration
        
        - **All 7 synthesis templates now tested.** Factual, Scientific/Health, Consumer, Technical,
          Opinion/Sentiment (untested), Contested, and Emerging/Frontier. Only Opinion/Sentiment
          remains untested, but it's the least complex template. The skill is mature for synthesis.
        - **Exa compliance is thorough.** 81/81 across 2 measured iterations. No further monitoring
          investment needed -- the guidance works reliably.
        - **Feedback reports validated and useful.** They surfaced 4 actionable findings this iteration.
          Worth keeping as standard practice.
        - **Quick wins available:** Backlog #26 (arxiv guidance) and #27 (verification verdict
          expansion) are small, targeted changes that can be done without a full research run.
        - **Recommended next:** Backlog #25 (audit against Anthropic agent teams docs) is a quick
          meta task that doesn't require running the skill. Good for a shorter iteration.
        
      • iteration-07.md 6.7 KB
        # Iteration 7 -- Agent Teams Docs Audit (Meta)
        
        ## Plan
        
        **Type:** Meta-research (docs audit)
        **Backlog item:** #25
        **Scope:** No research run -- audit existing skill against official Anthropic documentation.
        
        **Goal:** Read official Claude Code agent teams docs and compare our implementation against
        recommended patterns, known limitations, and best practices. Use Exa for discovery and
        Firecrawl for scraping (not the claude-code-guide agent). Also implement quick wins #26
        (arxiv guidance) and #27 (verification verdict expansion).
        
        ## Sources Scraped
        
        6 official pages scraped with Firecrawl (`onlyMainContent: true`, `maxAge: 86400000`):
        
        1. `code.claude.com/docs/en/agent-teams` -- Primary agent teams documentation
        2. `code.claude.com/docs/en/sub-agents` -- Subagents documentation (for comparison)
        3. `platform.claude.com/docs/en/agent-sdk/subagents` -- SDK subagents (programmatic API)
        4. `code.claude.com/docs/en/best-practices` -- Best practices guide
        5. `anthropic.com/engineering/multi-agent-research-system` -- Anthropic's own research system
        6. `claude.com/blog/building-multi-agent-systems-when-and-how-to-use-them` -- Multi-agent guide
        
        Plus 3 Exa searches for discovery (30 results total across `docs.anthropic.com`,
        `github.com/anthropics`, and general web).
        
        ## Audit Findings
        
        **Green Flags (architecture confirmed correct):**
        
        - **Hub-and-spoke / orchestrator-worker pattern** matches both the agent teams docs and
          Anthropic's own research system architecture. The docs explicitly describe "one session acts
          as the team lead, coordinating work, assigning tasks, and synthesizing results."
        - **No peer-to-peer researcher communication** aligns with the agent teams docs' model where
          the lead coordinates all work. The multi-agent blog emphasizes "context-centric decomposition"
          and warns against the "telephone game" of inter-agent information loss.
        - **Lead-dispatched assignments** validated. Docs support both explicit assignment and self-claim
          with file locking, but our empirical finding (Iterations 1-2) that lead-dispatched works
          better is consistent with Anthropic's advice to "teach the orchestrator how to delegate."
        - **Filesystem output** explicitly recommended. Anthropic's research blog says: "Subagent output
          to a filesystem to minimize the 'game of telephone'... subagents call tools to store their
          work in external systems, then pass lightweight references back to the coordinator."
        - **Verification subagent pattern** matches the blog's "verification subagent pattern" exactly,
          including our "assign to fresh perspective" approach.
        - **Context-centric decomposition by research angle** (not by role) matches the blog's strong
          recommendation against problem-centric splitting.
        - **Shutdown then TeamDelete sequence** matches docs' requirement that cleanup fails if active
          teammates remain.
        - **No nested teams** -- our researchers don't spawn sub-teams, which is correct since the docs
          confirm "teammates cannot spawn their own teams."
        - **Automatic message delivery, no polling** -- confirmed by docs.
        
        **New Information (resolved questions):**
        
        - **Permission inheritance confirmed:** "Teammates start with the lead's permission settings.
          If the lead runs with `--dangerously-skip-permissions`, all teammates do too." Can change
          individual modes after spawning, but not at spawn time. This partially resolves backlog #5.
        - **CLAUDE.md + MCP servers inherited automatically:** "When spawned, a teammate loads the same
          project context as a regular session: CLAUDE.md, MCP servers, and skills." This explains why
          researchers can use ToolSearch to load Exa/Firecrawl -- they inherit MCP server access.
        - **Task dependencies available:** Framework supports `blockedBy` for task ordering. We don't
          use this (our lead-dispatched approach handles sequencing manually), but it's available if
          we ever need it.
        - **Session resumption doesn't restore teammates:** Our `state.md` resume check in Phase 0 is
          important precisely because of this limitation. Good that we built this.
        - **Delegate mode exists:** Pressing Shift+Tab restricts the lead to coordination-only tools.
          Not needed for our skill (which operates programmatically) but useful to know.
        
        **Anthropic's research system comparison:**
        
        Their architecture is strikingly similar to ours. Key parallels: orchestrator-worker pattern,
        lead decomposes query and spawns subagents for different facets in parallel, subagents search
        independently then return distilled findings, lead synthesizes. Their system uses Opus 4 lead +
        Sonnet 4 subagents and outperformed single-agent by 90.2% on their internal eval.
        
        Key difference: they found "token usage by itself explains 80% of the variance." Their effort
        scaling rules: simple queries get 1 agent with 3-10 tool calls, direct comparisons 2-4 agents
        with 10-15 calls each, complex research 10+ agents. Our scale is more conservative (max 4
        researchers for Comprehensive) but appropriate for our token budget.
        
        They also use a separate CitationAgent for adding citations to the final output. We don't have
        this -- our lead handles synthesis and citations together. Worth monitoring but not needed now.
        
        **No anti-patterns found.** Our architecture doesn't fight the framework, uses no deprecated
        patterns, and follows the recommended cleanup sequence.
        
        **Yellow flags (minor, non-blocking):**
        
        - Our Comprehensive scope uses Opus for all researchers. Anthropic's own system uses Sonnet
          workers successfully. Consider testing Sonnet researchers in Comprehensive to reduce cost
          while maintaining quality. Added as backlog #28.
        - No explicit TaskList polling during long researcher waits. If a researcher silently crashes,
          the lead would wait indefinitely. Low risk (hasn't happened in 7 iterations) but worth
          awareness.
        
        ## Changes Made
        
        1. **[#26] arxiv scraping guidance** -- Added to `references/researcher-prompt.md` under
           Content Extraction: guidance to use `arxiv.org/html/{id}` instead of `/abs/`, prefer Exa
           summaries for triage, use arxiv API for structured extraction.
        
        2. **[#27] Verification verdict expansion** -- Added SUPPORTED WITH NUANCE as fourth verdict
           in SKILL.md Phase 7, with description: "Directionally correct but specific details differ
           or require qualification."
        
        **Backlog items resolved:** #25 (docs audit), #26 (arxiv guidance), #27 (verification verdict).
        
        ## Notes for Next Iteration
        
        - Architecture is validated and mature. 7 iterations without a major structural change needed.
        - Permission setup (#5) is now better informed: inheritance confirmed, specific permissions
          needed are mkdir + Write to research dir + Bash.
        - Scope calibration (#6) has 7 data points and strong evidence from Anthropic's own scaling
          rules to compare against.
        - Consider testing mixed model config: Opus lead + Sonnet researchers for Comprehensive (#28).
        
      • iteration-08.md 4.1 KB
        # Iteration 8 -- Scope Calibration Refinement (Meta)
        
        ## Plan
        
        **Type:** Meta-analysis (internal data synthesis)
        **Backlog item:** #6
        **Scope:** No research run -- analyze 7 empirical data points + Anthropic's scaling rules to
        refine scope calibration guidance.
        
        **Goal:** Determine whether the type-to-scope defaults need adjustment, add escalation and
        de-escalation signals, and add empirical output expectations. Also subsumes backlog #8 (report
        length targets).
        
        ## Analysis
        
        **Data compiled from 7 runs:**
        
        | Iter | Type               | Scope         | Rounds Used | Sources | Lines | Round 2 Value         | Assessment     |
        | ---- | ------------------ | ------------- | ----------- | ------- | ----- | --------------------- | -------------- |
        | 1    | Technical+Consumer | Broad         | 2           | 25      | ~250  | Useful                | Right-sized    |
        | 2    | Contested          | Comprehensive | 3           | 60+     | ~350  | Useful                | Right-sized    |
        | 3    | Consumer           | Broad         | 2           | 25      | ~200  | Useful-not-essential  | Slight overfit |
        | 4    | Technical (meta)   | Broad         | 2           | 24      | ~280  | Useful                | Right-sized    |
        | 5    | Scientific/Health  | Broad         | 2           | 26      | ~300  | Useful                | Right-sized    |
        | 6    | Emerging/Frontier  | Comprehensive | 3           | 55      | ~400  | Useful (conflict+gap) | Right-sized    |
        | 9    | Opinion/Sentiment  | Comprehensive | 2           | 60+     | ~370  | Useful (all 4 gaps)   | Right-sized    |
        
        **Key findings:**
        
        1. **Current defaults are correct 6/6 times.** Consumer was the only case with slack (Round 2
           useful-not-essential). All other scopes were right-sized.
        
        2. **Consumer is the only type with de-escalation evidence.** When the product category is
           narrow, well-reviewed, and expert consensus is clear, Focused could suffice. Iteration 3
           (keyboards) is the data point.
        
        3. **Comprehensive verification (Round 3) added value both times tested.** Iteration 2:
           SUPPORTED WITH NUANCE + CONTESTED verdicts. Iteration 6: CONTESTED + SUPPORTED WITH NUANCE.
           Both materially improved the final report.
        
        4. **Focused scope has never been tested.** Only Factual defaults to it, and no Factual question
           has been run. Output expectations for Focused are projected from Round 1 subsets of Broad
           runs.
        
        5. **Anthropic comparison:** Their simple (1 agent, 3-10 calls) maps to "don't use this skill."
           Their comparisons (2-4 agents) map to our Focused/Broad. Their complex (10+ agents) exceeds
           our maximum. Our system is well-sized for the middle range.
        
        6. **Report length is scope-driven, not type-driven.** The variation across types within Broad
           (200-300 lines) is smaller than the variation across scopes. So output expectations belong on
           the scope tier, not the question type.
        
        ## Changes Made
        
        1. **Scope modifiers added to `references/question-types.md`** -- Per-type de-escalation and
           escalation signals, with the Consumer de-escalation backed by Iteration 3 evidence. Also
           added "do not use this skill" criteria for simple queries.
        
        2. **Empirical output expectations added to `references/question-types.md`** -- Table of
           report length and source count ranges per scope tier, based on 7 runs. Focused estimates
           projected. Explicitly noted these are descriptive ranges, not targets.
        
        3. **SKILL.md Phase 2 updated** -- Now references scope modifiers from question-types.md
           instead of inlining ad-hoc adjustment criteria.
        
        **Backlog items resolved:** #6 (scope calibration), #8 (report length targets).
        
        ## Notes for Next Iteration
        
        - **Focused scope remains untested.** A Factual question run would validate the Focused tier
          and the projected output expectations, but it's low priority since the skill explicitly says
          "don't use for simple factual lookups."
        - **The main remaining empirical gap is Opinion/Sentiment** -- the last untested synthesis
          template. A good candidate for the next empirical run.
        - **Mixed model config (#7) is the highest-value empirical test** -- combining it with an
          Opinion/Sentiment question would test two things at once.
        
      • iteration-09.md 7.1 KB
        # Iteration 9 -- Opinion/Sentiment + Mixed Model Config (Empirical)
        
        ## Plan
        
        **Type:** Empirical (Opinion/Sentiment question)
        **Backlog items:** #7 (mixed model config), last untested synthesis template
        **Scope:** Comprehensive (4 sonnet researchers, 2 rounds -- model override from opus to sonnet)
        
        **Goal:** Combine two tests in a single empirical run: (1) Validate Sonnet researchers for
        Comprehensive scope (backlog #7), replacing the Opus default. Anthropic's own research system
        uses Opus lead + Sonnet subagents. (2) Test the Opinion/Sentiment synthesis template, the last
        of 7 templates never used in a real run.
        
        **Question:** "What do developers actually think about Nix and NixOS in 2026?"
        
        **Angles:**
        
        - (A) Community sentiment: overall vibe, praise vs frustration, how feelings have shifted
        - (B) Pain points and barriers: learning curve, documentation, Nix language, flakes controversy
        - (C) Enterprise/professional adoption: companies using Nix, CI/CD patterns, onboarding stories
        - (D) Governance and community health: foundation drama, forks (Lix, Aux, Tvix), SC elections
        
        ## Research Summary
        
        **Report:** `~/.claude/research/developer-opinions-nix-nixos-2026/report.md`
        
        **Round 1:** All 4 Sonnet researchers completed successfully. Strong convergence on the
        paradox theme: fastest community growth ever (30% YoY) alongside deepest fragmentation ever
        (governance crisis, three forks). Researcher A surfaced the sentiment spectrum from "Nix
        changed my life" to "I quit after 6 months." Researcher B documented the learning curve and
        documentation as top pain points across all sources. Researcher C found enterprise adoption
        concentrated in devshells-only patterns with wrapper tools. Researcher D provided deep coverage
        of the 2024-2025 governance crisis and its aftermath. Total: ~50+ sources across researchers.
        
        **Triage identified:** (1) Gap: moderate/ambivalent voices underrepresented (survivorship bias
        in online discussions), (2) DetSys conflict of interest not fully explored (company behind
        flakes also behind commercial products), (3) wrapper tools ecosystem deserved deeper
        comparison, (4) Steering Committee election results and their impact not covered.
        
        **Round 2:** Researcher A (gap-fill) found moderate voices -- "Nix purgatory" users stuck
        between love and frustration, plus successful teams who don't post online. Researcher B
        (deep-dive) investigated DetSys conflict with nuance -- found community resentment but also
        genuine engineering contributions. Researcher C (deep-dive) did comprehensive wrapper tools
        comparison with GitHub stars, feature matrices, and community adoption data. Researcher D
        (gap-fill) covered SC election results and post-crisis governance trajectory.
        
        **Round 3 skipped:** No single-source high-impact claims remaining. All major conflicts
        resolved by Round 2. Remaining uncertainties about magnitude, not direction.
        
        **Final report:** ~370 lines, 50 footnotes, 60+ sources. Opinion/Sentiment template with
        5-segment distribution of views and representativeness assessment.
        
        ## Observations
        
        **What worked well:**
        
        - **Opinion/Sentiment template validated with adaptation.** The 3-segment default structure
          (Majority/Minority/Outlier) was adapted to 5 segments to capture the nuanced reality:
          Passionate Core (~15-20%), Pragmatic Middle (~30-35%), Frustrated Dropouts (~20-25%),
          Nix Purgatory (~15-20%), and Broader Dev Population (~10-15%). The template's
          Representativeness Assessment section added genuine value.
        - **Sonnet researchers produced Comprehensive-quality output.** All 4 researchers completed
          both rounds, followed all methodology guidance, wrote structured findings and feedback
          reports. No evidence of quality degradation vs Opus researchers in Iterations 2 and 6.
        - **Lead-dispatched follow-ups:** 8th consecutive success, zero nudging needed.
        - **All researchers completed on first try.** No stalls, no nudging, no retries.
        - **Feedback reports written by all 4 researchers.** Continued validation of this mechanism.
        - **Verification skip was correct.** Round 2 resolved all major gaps and conflicts. No
          single-source claims would have changed the report's conclusions.
        
        **Permission friction observed:**
        
        - **curl commands required user approval.** Researchers used `curl` for Reddit JSON API and
          GitHub API endpoints, triggering ~8-10 permission prompts during the run. This is the first
          time permission friction was significant enough for the user to raise it post-run.
        - **Root cause:** The researcher-prompt.md Reddit section recommends curl for Reddit JSON API
          (since Exa doesn't index Reddit and Firecrawl blocks it). GitHub API was also accessed via
          curl for stars/forks data.
        
        **Researcher feedback highlights:**
        
        - Researcher A: Exa with enableSummary + textMaxCharacters:1 excellent for rapid triage.
          Category gotchas (tweet, company) useful to know upfront.
        - Researcher B: Firecrawl scraping of HN threads returned full comment trees in markdown --
          very useful for community sentiment capture.
        - Researcher C: Reddit JSON API `restrict_sr=on` was essential -- without it, global results
          instead of subreddit-specific. Exa didn't index GitHub issue bodies well. Recommended
          GitHub search API patterns in methodology.
        - Researcher D: Strong coverage of Discourse-based forums (NixOS Discourse) via direct
          Firecrawl scraping. Exa worked well for blog posts and personal sites.
        
        ## Changes Made
        
        3 changes to `references/researcher-prompt.md`:
        
        1. **Reddit section updated** -- Added `restrict_sr=on` requirement with example URL pattern,
           guidance to prefer `/top/` and `/hot/` over `/search/` for sentiment research (browsing
           reveals organic discussions), note that curl requires user permission approval each time.
        
        2. **GitHub data guidance added** -- New section before "Medium paywalled articles": Exa
           `get_code_context_exa` as primary tool for GitHub content discovery, Firecrawl for deep
           extraction of specific pages, `gh api` for quantitative data (stars, forks, issues) since
           it's already permitted. Explicit "do not use curl for GitHub API endpoints."
        
        3. **Discourse JSON search API added** -- New section after GitHub: Firecrawl scrape of
           `/search.json?q=...` as fallback when Exa can't find specific forum topics. Returns
           structured results with topic IDs for follow-up scraping.
        
        **Backlog items resolved:** #7 (mixed model config -- Sonnet validated for Comprehensive).
        
        ## Notes for Next Iteration
        
        - **All 7 synthesis templates now tested.** The skill is mature for research operations.
          Opinion/Sentiment required the most template adaptation (5 segments vs 3) but the framework
          accommodated this well.
        - **Sonnet validated as viable Comprehensive override, Opus kept as default.** One successful
          run (Opinion/Sentiment) doesn't prove Sonnet handles the hardest Comprehensive questions.
          SKILL.md notes Sonnet as an available override when cost matters.
        - **Reddit MCP is the top engineering priority.** The curl permission friction this iteration
          is the strongest signal yet. Backlog #8 has full context on options.
        - **Permission setup (#5) remains important** for shareability. Inheritance is confirmed but
          specific Bash permissions (curl, sleep) still cause friction.
        
      • iteration-10.md 3.8 KB
        # Iteration 10 -- Reddit MCP Server Integration (Engineering)
        
        ## Plan
        
        **Type:** Engineering (no research run)
        **Backlog item:** #8 (Reddit MCP server integration)
        
        **Goal:** Eliminate the last curl dependency in researcher workflows by integrating a Reddit
        MCP server via 1MCP, and updating the researcher methodology to use native MCP tools instead
        of curl for Reddit content access.
        
        **Background:** Iteration 9 confirmed significant permission friction (~8-10 curl prompts per
        run) from Reddit JSON API access. Prior research (`~/.claude/research/reddit-access-for-claude-code/report.md`)
        identified `jordanburke/reddit-mcp-server` as the top option. Fresh evaluation confirmed:
        - Still actively maintained (v1.2.1 published ~10 days ago, 108 commits, 4 contributors)
        - 9 read tools + 6 write tools, including `search_reddit` (critical for researchers)
        - Anonymous mode: zero credentials, ~10 rpm, zero setup
        - `npx reddit-mcp-server` -- no install needed, 1MCP compatible via stdio
        - Runner-up (Hawstein/mcp-server-reddit) is stale -- no commits in 10 months, no search tool
        - New entrants (liuyang1520, karanb192) too immature (0 stars, requires auth/local install)
        
        Legal context unchanged: Reddit v. Anthropic lawsuit makes this fraught, but low-volume
        personal use risk is low. Anonymous mode doesn't even use API credentials.
        
        ## Changes Made
        
        ### 1. Added Reddit MCP server to 1MCP config
        
        **File:** `configs/claude/1mcp.json`
        
        Added `reddit` server entry with `npx reddit-mcp-server` and explicit `REDDIT_AUTH_MODE: "anonymous"`.
        Placed alphabetically before the `gws` entry. Server will be available via 1MCP as
        `mcp__1mcp__reddit_1mcp_{tool_name}`.
        
        ### 2. Updated researcher methodology -- Setup section
        
        **File:** `~/.claude/skills/deep-research-team/references/researcher-prompt.md`
        
        Added `ToolSearch: "+reddit"` to the tool loading instructions in the Setup section.
        
        ### 3. Updated researcher methodology -- Available Tools section
        
        **File:** same
        
        Added **Reddit** tool block documenting 4 content extraction tools (one-line descriptions,
        params left to MCP schemas): `get_top_posts`, `get_post_comments`, `get_reddit_post`,
        `get_subreddit_info`. Deliberately excluded `search_reddit` -- Google's Reddit index (via
        Firecrawl search) is far superior to Reddit's native search.
        
        Added to "When to use which": Firecrawl search for Reddit discovery (`site:reddit.com`),
        Reddit MCP tools for content extraction.
        
        ### 4. Replaced curl-based Reddit access with MCP tools
        
        **File:** same, "Accessing Restricted Content > Reddit content" section
        
        Completely rewrote the Reddit content section. Removed all curl-based guidance. Replaced with
        two-layer pattern: Firecrawl search (`site:reddit.com {query}`) for discovery, Reddit MCP
        tools for content extraction. Preserved the key insight that `get_top_posts` is preferred over
        search for sentiment research.
        
        ## Impact
        
        - **Permission friction:** Eliminates ~8-10 curl permission prompts per research run that
          accesses Reddit. Reddit access now uses native MCP tools that require no special permissions.
        - **Researcher experience:** Researchers use the same tool pattern (ToolSearch -> MCP call)
          for Reddit as they do for Exa and Firecrawl. No more shelling out to curl.
        - **Coverage:** All previously curl-accessible Reddit features are covered by MCP tools, plus
          additional features (engagement analysis, subreddit stats, trending subreddits).
        - **curl dependency:** With GitHub curl eliminated in Iteration 9 and Reddit curl eliminated
          here, researchers should have zero curl dependencies in normal workflows. The only remaining
          curl use cases are the Bluesky public API and rate-limit retry sleeps (which use Bash `sleep`,
          not curl).
        
        ## Backlog Resolution
        
        **#8 resolved.** Reddit MCP server integrated via 1MCP with anonymous mode. Researcher
        methodology updated to use native MCP tools. curl-based Reddit access guidance removed.
        
      • iteration-11.md 6.9 KB
        # Iteration 11 -- Reddit MCP Empirical Validation
        
        ## Type
        
        Empirical (Focused scope run to validate Iteration 10 engineering changes)
        
        ## Question
        
        "What is the best community-favorite content about the TV show Dark -- videos, write-ups,
        essays, analyses, explainers, fan theories?"
        
        ## Config
        
        - **Scope:** Focused (2 researchers, 1 round)
        - **Question type:** Consumer/Recommendation
        - **Model:** Sonnet researchers
        - **Report:** `~/.claude/research/dark-tv-show-community-content/report.md`
        
        ## Validation Goals & Results
        
        | Goal                                          | Result | Notes                                                                                       |
        | --------------------------------------------- | ------ | ------------------------------------------------------------------------------------------- |
        | Researchers load Reddit tools via ToolSearch   | Pass   | Both loaded successfully                                                                    |
        | Firecrawl search with site:reddit.com works   | Pass   | Both used it. B called it "the single most valuable technique"                              |
        | get_top_posts extracts Reddit content          | Pass   | Both used. Returns memes/appreciation at top; insufficient alone for substantive threads     |
        | get_post_comments extracts thread discussions  | Pass   | A used on 4 threads, B on 2. Specific upvote counts, ratios, user quotes extracted          |
        | get_reddit_post works for specific posts       | Pass   | B used on 7 posts with engagement metrics                                                   |
        | get_subreddit_info works                       | Pass   | B used to get 213K member count                                                             |
        | search_reddit works                            | Pass   | A used 2 queries                                                                            |
        | Zero curl usage for Reddit                     | Pass   | No permission prompts observed                                                              |
        | Useful findings produced                       | Pass   | 24-source report with specific engagement data, 5 categories of content                     |
        
        **Overall: 9/9 validation goals passed. Reddit MCP integration works end-to-end.**
        
        ## Tool Usage Summary
        
        ### Researcher A
        
        - **Exa:** enableSummary, textMaxCharacters, numResults (standard pattern)
        - **Firecrawl:** firecrawl_search for Reddit discovery
        - **Reddit MCP:** search_reddit (2), get_top_posts (1), get_post_comments (4)
        - **Issues:** `site:reddit.com/r/DarK` in Firecrawl search was too specific, returned generic
          results. Broader `site:reddit.com` worked better.
        
        ### Researcher B
        
        - **Exa:** enableSummary, textMaxCharacters, numResults, excludeDomains, category (personal site)
        - **Firecrawl:** firecrawl_scrape (4 URLs), firecrawl_search for Reddit discovery
        - **Reddit MCP:** get_top_posts (1), get_post_comments (2), get_reddit_post (7), get_subreddit_info (1)
        - **Issues:** Medium paywall blocked full markdown extraction (summary format worked). Top posts
          from r/DarK were overwhelmingly memes -- needed Firecrawl search for substantive threads.
        
        ## Key Observations
        
        ### Two-layer Reddit pattern validated
        
        The Firecrawl search (discovery) + Reddit MCP (extraction) pattern works well. Both researchers
        used it independently and both endorsed it in feedback.
        
        ### get_top_posts is necessary but insufficient
        
        For topic-specific threads, `get_top_posts` returns the most popular posts overall (which for
        entertainment subreddits are memes and appreciation posts, not analytical content). The Firecrawl
        search layer is essential for finding substantive threads on specific topics.
        
        ### Firecrawl search scoping
        
        Researcher A found that `site:reddit.com/r/DarK` (subreddit-scoped) returned generic video
        essay threads instead of r/DarK-specific content. The broader `site:reddit.com` with topic
        keywords in the query worked better. This is worth noting in the methodology.
        
        ### Researcher complementarity
        
        Good angle separation produced minimal overlap. Both independently found the mmmmmmmmichaelscott
        FAQ and The_Wattsatron easter eggs post (confirming canonical status via independent discovery).
        A focused on video/audio content; B focused on written/academic content.
        
        ### Exa category filter usage
        
        Researcher B used `category: "personal site"` to filter to independent blog essays, which
        effectively surfaced philosophical analyses while filtering out corporate listicle content.
        Good technique for this question type.
        
        ## Changes Made
        
        ### 1. Note Firecrawl search scoping in researcher-prompt.md
        
        Add note that `site:reddit.com` (broad) works better than `site:reddit.com/r/subreddit`
        (subreddit-scoped) in Firecrawl search queries.
        
        ### 2. Note get_top_posts limitation in researcher-prompt.md
        
        Clarify that `get_top_posts` returns the most popular posts overall (often memes/meta for
        entertainment subreddits), not necessarily the most substantive analytical content. Recommend
        using Firecrawl search for finding topic-specific substantive threads.
        
        ## Transcript Analysis
        
        Script: `analyze-transcripts.py --session 5213f595-f4b6-4140-b7e0-5d2908d822d7`
        
        **Exa compliance: 10/10 (100%)**
        - All calls used `enableSummary: true` + `textMaxCharacters: 1`
        - Zero `contextMaxCharacters` violations
        
        **All MCP calls (42 total across both researchers):**
        - Reddit: 22 (get_reddit_post: 9, get_post_comments: 7, search_reddit: 3, get_top_posts: 2, get_subreddit_info: 1)
        - Exa: 10 (web_search_advanced_exa: 10)
        - Firecrawl: 10 (firecrawl_search: 6, firecrawl_scrape: 4)
        
        **Notable patterns:**
        - Researcher B used `excludeDomains: ["reddit.com"]` on Exa to avoid duplicate Reddit results
        - Researcher B used `category: "personal site"` for blog essay discovery (2 calls)
        - Reddit was the most-used MCP server — appropriate for a community-content question
        
        ## Post-Iteration Housekeeping
        
        Changes made in the same session after the main iteration writeup:
        
        ### 3. Generalize analyze-transcripts.py to track all MCP tools
        
        Replaced hardcoded Exa/Firecrawl parsing with generic MCP tool detection via
        `re.compile(r"(?:mcp__1mcp__)?(\w+?)_1mcp_(\w+)")`. All 1MCP tools now captured and grouped
        by server. Custom formatters for Exa (with compliance checking), Firecrawl, and Reddit.
        Aggregate report shows per-server totals and per-endpoint breakdown.
        
        ### 4. Fix --project-dir to accept filesystem paths
        
        `--project-dir` now accepts regular filesystem paths (e.g., `/Users/malo/.config/nix-config`)
        in addition to `~/.claude/projects/{slug}` paths. Slug algorithm: `re.sub(r"[^a-zA-Z0-9]", "-", path)`.
        
        ### 5. Simplify researcher feedback template
        
        Removed the "Search Tools" section (redundant now that the transcript script extracts all MCP
        tool usage automatically). Reframed template as qualitative reflection: What Worked, Issues,
        Suggestions. Added framing note so researchers understand the purpose.
        
        ## Backlog Changes
        
        - No items resolved (validation run, not feature addition)
        
        **Cumulative changes including housekeeping: 30 + 3 = 33.**
        
      • iteration-12.md 6.3 KB
        # Iteration 12 -- Firecrawl Scrapability Testing
        
        **Type:** Engineering (no research team)
        **Goal:** Empirically test all documented workaround URLs in researcher-prompt.md with actual
        Firecrawl scrape calls. Remove or caveat workarounds that don't work.
        
        ## Test Results
        
        ### 1. archive.ph
        
        **URL tested:** `https://archive.ph/newest/https://www.nytimes.com/2025/03/10/technology/ai-agents-future.html`
        **Params tried:** `proxy: "stealth"`, `proxy: "enhanced"`, `waitFor: 5000`
        **Result:** BROKEN. reCAPTCHA challenge page (status 429) on all attempts. Cloudflare protection
        blocks Firecrawl completely. Costs 5 credits per attempt (stealth/enhanced proxy).
        **Conclusion:** Remove as recommended workaround for Firecrawl-based research.
        
        ### 2. Wayback Machine
        
        **URLs tested:**
        - `web.archive.org/web/https://www.wsj.com/...` (shortcut format) -- 404, "not archived"
        - `web.archive.org/web/2024/https://www.nytimes.com/...` (year only) -- 404, "not archived"
        - `web.archive.org/web/2024*/https://arstechnica.com/...` (wildcard) -- returns calendar page
        - `web.archive.org/web/20240515120000/https://en.wikipedia.org/wiki/GPT-4` (exact timestamp) -- **SUCCESS**, full content (149K chars, had to be saved to file)
        - `web.archive.org/web/https://en.wikipedia.org/wiki/GPT-4` (shortcut, known-archived page) -- **SUCCESS**, 173K chars, status 200. Firecrawl followed redirect to `/web/20260203171033/...` (most recent snapshot). Wayback toolbar at top but full article content follows.
        - `web.archive.org/web/2/https://en.wikipedia.org/wiki/GPT-4` ("/2/" trick) -- **SUCCESS**, same result as shortcut format.
        
        **Result:** WORKS with simple shortcut format `/web/{URL}`. Initial tests with WSJ/NYTimes
        shortcut URLs returned 404 because those pages weren't archived, not because the format was
        wrong. Retesting with a known-archived page (Wikipedia GPT-4) confirmed the shortcut redirects
        to the most recent snapshot. Wayback Machine toolbar is prepended to extracted content but full
        page follows. 1 credit.
        **Conclusion:** Simplify guidance -- just use `/web/{URL}`. No need for exact timestamps.
        **Lesson:** Test URL format correctness against pages known to be archived.
        
        ### 3. xcancel.com (Twitter/X)
        
        **URLs tested:**
        - `xcancel.com/sama/status/1895533845455577446` (without proxy) -- anti-bot challenge (503)
        - `xcancel.com/sama/status/1895533845455577446` (stealth + waitFor) -- "Tweet not found" (404)
        - `xcancel.com/elikiowa/status/1880305905189106091` (stealth + waitFor) -- "Tweet not found" (404)
        - `xcancel.com/sama/status/1890816782836904000` (stealth + waitFor, verified valid ID) -- **SUCCESS**
        - `xcancel.com/OpenAI/status/1790070592011288831` (stealth + waitFor) -- "Tweet not found" (404)
        
        **Result:** WORKS but requires `proxy: "stealth"` + `waitFor: 5000` (5 credits). Initial tests
        with invalid tweet IDs gave misleading "Tweet not found" (404) results. Retests with verified
        valid tweet IDs confirmed: without proxy, anti-bot challenge blocks (503); with stealth +
        waitFor, full content returned (tweet text, replies, engagement metrics). Confirmed on both
        cached and fresh (`maxAge: 0`) URLs to rule out caching artifacts.
        **Conclusion:** Keep with updated guidance: requires stealth proxy + waitFor, costs 5 credits.
        **Lesson:** Always verify test inputs before concluding a service is broken.
        
        ### 4. freedium.cfd (Medium)
        
        **URL tested:** `https://freedium.cfd/https://towardsdatascience.com/rag-vs-fine-tuning-...`
        **Result:** DEAD. DNS resolution failure -- domain no longer resolves.
        **Conclusion:** Replace with freedium-mirror.cfd.
        
        ### 5. freedium-mirror.cfd (Medium -- replacement)
        
        **URL tested:** `https://freedium-mirror.cfd/https://towardsdatascience.com/rag-vs-fine-tuning-which-is-the-best-tool-to-boost-your-llm-application-94654b1eaba7`
        **Result:** SUCCESS. Full article content extracted (status 200, 1 credit, basic proxy).
        Complete Medium article with formatting, images, author info, tags. No paywall.
        **Conclusion:** Document as the working Medium paywall bypass.
        
        ### 6. Bluesky
        
        **URLs tested:**
        - `public.api.bsky.app/xrpc/app.bsky.feed.searchPosts?q=nix+nixos&limit=5` (raw API) -- 403 Forbidden
        - `bsky.app/profile/jay.bsky.team/post/3lihkfohpkk2t` (individual post, no waitFor) -- "Post not found" (SPA not rendered)
        - `bsky.app/profile/bsky.app` (profile page, waitFor: 5000) -- **SUCCESS**, massive content (full profile + dozens of recent posts)
        - `bsky.app/profile/atprotocol.dev` (profile page, waitFor: 5000) -- **SUCCESS**, 94K chars
        
        **Result:** WORKS but `bsky.app` is a JavaScript SPA -- `waitFor: 5000` is required for content
        to render. Without it, individual post pages return "Post not found". Profile pages with waitFor
        return extensive content including all recent posts. No proxy needed (1 credit). The raw API
        endpoint was unnecessary in the first place.
        **Conclusion:** Remove API-specific guidance. Document that bsky.app needs `waitFor: 5000` (SPA).
        
        ### 7. firecrawl_search site:reddit.com
        
        **Query tested:** `site:reddit.com best mechanical keyboard programming 2025`
        **Result:** SUCCESS. Returned 5 Reddit results with full URLs containing post IDs
        (e.g., `reddit.com/r/keyboards/comments/{postId}/...`). Post IDs extractable for Reddit MCP
        tools. Already validated in Iteration 11; re-confirmed here.
        **Conclusion:** Keep as documented.
        
        ## Changes Made
        
        **researcher-prompt.md "Accessing Restricted Content" section:**
        
        1. **Paywalled articles:** Demoted archive.ph (CAPTCHA-blocked), promoted Wayback Machine to
           primary with exact-timestamp URL format requirement
        2. **Twitter/X:** Added `proxy: "stealth"` + `waitFor: 5000` requirement for xcancel.com
           (5 credits). Previous guidance didn't mention these were needed.
        3. **Medium:** Replaced freedium.cfd with freedium-mirror.cfd, removed archive.ph fallback
        4. **Bluesky:** Removed raw API endpoint. Documented that bsky.app is a JavaScript SPA
           requiring `waitFor: 5000` for Firecrawl scrape. No proxy needed (1 credit).
        5. **Reddit/GitHub/Discourse/LinkedIn:** Unchanged (not retested, already validated)
        
        ## Credit Cost Summary
        
        Total Firecrawl credits used: ~38
        - archive.ph: 10 (2 attempts x 5 credits stealth/enhanced)
        - Wayback Machine: 4 (4 attempts x 1 credit basic)
        - xcancel.com: 16 (1 basic + 3 stealth attempts)
        - freedium.cfd: 0 (DNS failure, no credit charged)
        - freedium-mirror.cfd: 1 (basic proxy)
        - Bluesky: 8 (API 5 auto-stealth + 3 bsky.app scrapes x 1 basic)
        - firecrawl_search: ~0 (search credits separate)
        
      • iteration-13.md 6.8 KB
        # Iteration 13 -- Information Cascade Detection (Backlog #7)
        
        ## Type: Research Run (Empirical) + Engineering
        
        ## Question
        
        "How do journalists, fact-checkers, and researchers detect information cascades -- cases where many
        sources all trace back to one original, creating an illusion of independent corroboration?"
        
        ## Configuration
        
        - **Question type:** Technical
        - **Scope:** Broad (3 sonnet researchers, 2 rounds)
        - **Research angles:**
          1. Journalism/fact-checking practitioner methods (Researcher A)
          2. Academic citation cascade analysis (Researcher B)
          3. Computational misinformation cascade detection at scale (Researcher C)
        
        ## Report
        
        `~/.claude/research/information-cascade-detection/report.md`
        ~180 lines, 37 footnotes, ~50+ sources across researchers.
        
        ## Key Research Findings
        
        **Three converging research traditions** all address cascade detection but have developed largely
        independently:
        
        1. **Journalism/fact-checking:** "Trace to source" methodology. First Draft's provenance pillar,
           GIJN's "don't rely on other media" rule, IFCN two-source minimum. No org publishes an explicit
           cascade detection checklist -- it's tacit professional knowledge. The ivermectin/Rolling Stone
           case is the paradigmatic cascade case study.
        
        2. **Academic bibliometrics:** Greenberg 2009 (BMJ) is foundational -- identified citation bias,
           amplification, and invention across 242 papers. Woozle effect quantified by Letrud & Hernes
           2019 (76% of citing articles affirm debunked myths). Citogenesis (Wikipedia circular loops)
           well-documented. Chen et al. 2025 confirmed "telephone effect" computationally at 13M-pair
           scale.
        
        3. **Computational social science:** Vosoughi et al. 2018 (Science) foundational -- false news 6x
           faster, distinct cascade topology. CrowdTangle shutdown (Aug 2024) left major monitoring gap.
           DisTrack and temporal graph approaches (TIDE-MARK) are current state-of-art. AI amplifies
           cascades via knowledge collapse and epistemic destabilization.
        
        **Practitioner heuristics (Round 2):** Specific red flags distilled: identical phrasing/errors
        propagated, tight temporal clustering, all sources citing same original, lack of local detail,
        same stock photos. Source TYPE diversity (documents + people + data) matters more than source
        COUNT. Intelligence analysis (Heuer ACH, ICD 203) has the most rigorous frameworks for source
        independence evaluation.
        
        **Observable signals (Round 2):** Content+structure+style similarity distinguishes cascade from
        independent reporting (Bar et al. 2012). Temporal clustering is a primary signal -- one-quarter
        of news stories reproduced within 4 minutes (Cage et al. 2025). Network topology differs
        systematically (broadcast/star = cascade, bushy/multi-origin = independent).
        
        ## Dogfooding Observations
        
        **What worked well:**
        - All three Round 1 angles produced highly complementary material with minimal overlap
        - Lead-dispatched follow-ups worked perfectly (8th consecutive success)
        - Round 2 targeted the right gaps -- practitioner heuristics and observable signals filled the
          key gaps from Round 1
        - Researcher B had no Round 2 task (academic angle was thoroughly covered in Round 1) -- idle
          researcher management worked correctly
        
        **What went less well:**
        - Researcher A did not write a feedback file before shutdown (2 of 3 wrote feedback)
        - Researchers B and C both flagged Firecrawl returning oversized results for long academic
          articles (662KB for BMJ paper), with JSON format making partial reads difficult
        
        **Feedback themes (from B and C):**
        - Exa summaries were excellent for academic paper triage -- often sufficient without deep scraping
        - Firecrawl PDF parser failed on one conference paper (MCP error)
        - Suggestion: structured JSON extraction with targeted schema would be more token-efficient than
          full markdown for long academic papers
        - Intelligence analysis angle (Heuer ACH, OSINT corroboration vs replication) was an
          underexplored rich connection
        
        ## Changes
        
        ### 1. Added "Cascade Check" section to researcher-prompt.md (lines ~250-290)
        
        New section between Stopping Criteria and Handling Insufficient Evidence. Contains:
        - 5-step cascade check procedure (trace citation chain, check copied phrasing, assess source
          type diversity, apply "how do they know?" test, watch for wire service cascades)
        - Red flag list (temporal clustering, no local detail, single-origin collapse, same source type)
        - Instruction to downgrade confidence and flag in CROSS-SOURCE ANALYSIS when detected
        
        ### 2. Added cascade check reference to Step 3 (Reflect)
        
        Added seventh reflection question: "Cascade check (see below): Are my 'multiple sources'
        actually independent?" -- points researchers to the new section.
        
        ## Backlog Updates
        
        - **#7 (Information cascade detection):** RESOLVED. Research complete, heuristics engineered
          into researcher reflection step.
        - **#13 (researcher-prompt.md length):** File now ~541 lines (up from ~497). Added ~44 lines.
          No evidence of instruction-following degradation across 12 prior iterations, but worth
          monitoring. Next empirical run should verify researchers find and follow the cascade check.
        - **New observation:** Firecrawl returns oversized results for long academic HTML pages.
          Researchers suggest structured JSON extraction with schema. Not adding to backlog -- the Exa
          summary pattern is the recommended workaround and researchers are already following it.
        
        ## Post-Iteration Changes
        
        Discussion after the research run led to two design changes applied to SKILL.md:
        
        ### 1. Flexible round counts (SKILL.md)
        
        Replaced rigid per-scope round counts with ranges and convergence-based stopping:
        
        | Scope         | Old | New |
        | ------------- | --- | --- |
        | Focused       | 1   | 1-2 |
        | Broad         | 2   | 2-3 |
        | Comprehensive | 2-3 | 3-4 |
        
        Added early stopping language: rounds are heuristics, not targets. Stop when citation
        convergence is reached. Removed hard phase-skip rules (e.g., "For Focused scope: skip to
        Phase 8") and replaced with convergence checks at each triage point. Any scope can exit
        early if findings have converged, and any scope can run additional rounds if gaps persist
        within its budget.
        
        ### 2. Verification architecture rethink (Backlog #16 promoted)
        
        Identified that the current verification model (reusing angle researchers in Round 3,
        Comprehensive only) has two problems:
        
        1. **Angle researchers are specialists** -- pulling them off their angle to do cross-cutting
           verification wastes their accumulated context and expertise.
        2. **Verification doesn't need a round** -- it's a parallel task, not a sequential phase.
        
        Promoted backlog #16 (dedicated verifier agents) to High Priority with full design:
        - Spawn lightweight, single-purpose verifier agents on-demand after any round
        - Verifiers run in parallel with angle researchers continuing their work
        - Lead assigns specific claims to verify based on triage
        - Available at any scope, not just Comprehensive
        - Implementation deferred to a future iteration
        
    • ITERATION-LOG.md 23.6 KB
      # Deep Research Skill -- Self-Improvement Log
      
      ## Protocol
      
      This file drives an iterative improvement loop for the `deep-research-team` skill. Each
      iteration follows this cycle:
      
      1. **Choose** -- Read this log. Review the backlog. Pick either a meta-research topic
         (theoretical improvement) or an empirical test (run the skill, observe friction).
      2. **Research** -- Invoke `/deep-research-team` with the chosen question, or conduct
         targeted meta-research using web search.
      3. **Observe** -- After research completes, analyze what happened. Record observations.
      4. **Improve** -- Make targeted edits to skill files based on observations. Document changes.
      5. **Checkpoint** -- Brief user check-in: what happened, what changed, what's next.
      6. **Update** -- Refresh the backlog, set Current State for the next iteration.
      
      **Resuming:** If context was compacted or a new session started, read this file from the top.
      The Current State section tells you exactly where to pick up.
      
      **File organization:** This log file contains only compact summaries (5-8 lines each) in the
      Completed Iterations section. Full detail lives in `iterations/iteration-{NN}.md`. Workflow:
      
      1. When starting an iteration, create `iterations/iteration-{NN}.md` and write detail there
      2. As the iteration progresses, update the detail file (observations, changes, transcript analysis)
      3. When the iteration is complete, add a compact summary to the Completed Iterations section below
      4. Update Current State to point to the next iteration
      
      Never put full iteration detail in this file — it would grow too large for quick context loading.
      
      ## Current State
      
      - **Iteration:** 14 (ready to start)
      - **Phase:** choose next backlog item
      - **Next action:** Implement backlog #16 (dedicated verifier agents) -- design is clear,
        ready for engineering. Also a good candidate for empirical validation in the same iteration.
      - **Round counts updated** (post-Iteration 13): Focused 1-2, Broad 2-3, Comprehensive 3-4.
        Early stopping language added. Hard scope-gated phase transitions removed in favor of
        convergence-based decisions.
      - **Verification architecture rethink in progress** (backlog #16, promoted to high priority):
        Replace reuse of angle researchers for verification with dedicated on-demand verifier agents.
        Removes scope restriction (all scopes can verify), preserves angle researcher context, and
        allows verification to run in parallel with ongoing investigation.
      - **Cascade detection heuristics added** (Iteration 13). researcher-prompt.md at ~541 lines.
      
      ## Completed Iterations
      
      Full details in `iterations/iteration-{NN}.md`. Summaries here for context.
      
      ### Iteration 1 -- Compound Technical+Consumer (Empirical)
      
      **Question:** "Leading open-source LLM inference engines for 70B on consumer GPUs?"
      **Scope:** Broad (3 sonnet researchers, 2 rounds)
      **Report:** `~/.claude/research/llm-inference-engines-70b-consumer-gpu/report.md`
      **Key findings:** vLLM for multi-GPU, llama.cpp/Ollama for single-GPU. 25 sources.
      **Problems found:** Researchers didn't self-claim follow-up tasks (biggest friction),
      didn't persist Round 1 findings to disk, lead name not in spawn prompt.
      **Changes (5):** Restructured reporting flow, added self-claim instructions, updated
      critical rules, added lead name to spawn prompt, added mkdir permission.
      
      ### Iteration 2 -- Contested Question (Empirical)
      
      **Question:** "Should AI model weights be open-sourced?"
      **Scope:** Comprehensive (4 opus researchers, 3 rounds incl. verification)
      **Report:** `~/.claude/research/ai-model-weights-open-source/report.md`
      **Key findings:** 37-footnote steelmanned debate, 60+ sources. Cross-agent verification
      added genuine value (one claim SUPPORTED WITH NUANCE, one CONTESTED).
      **Problems found:** Self-claim fix from Iteration 1 still didn't work (confirmed pattern),
      inconsistent filename casing, no round stopping criteria, no assignment guidance.
      **Changes (4):** Replaced self-claim with lead-dispatched follow-ups (biggest architectural
      change), normalized filenames to lowercase, added round stopping criteria, added Round 2
      assignment guidance (continuity vs fresh perspective).
      
      ### Iteration 3 -- Consumer Comparison (Empirical)
      
      **Question:** "Best mechanical keyboard for programming in 2026?"
      **Scope:** Broad (3 sonnet researchers, 2 rounds)
      **Report:** `~/.claude/research/best-mech-keyboard-programming-2026/report.md`
      **Key findings:** Keychron dominance at every price tier. 25 sources. Top Pick: Q5 Max.
      **Validated:** Lead-dispatched follow-ups worked perfectly (zero nudging), lowercase filenames
      worked, assignment guidance worked, non-AI domain worked fine with Exa.
      **Problems found:** Round 2 was useful-not-essential (suggests Focused might suffice for
      simple Consumer), Phase 1 skipped constraint gathering, idle notification noise.
      **Changes:** None needed -- all Iteration 2 changes validated successfully.
      
      **Cumulative skill improvements across 3 iterations:** 9 changes total. Architecture now
      stable: lead-dispatched follow-ups, file persistence, lowercase filenames, round stopping
      criteria, assignment guidance all working. Self-claim problem fully resolved.
      
      ### Iteration 4 -- Search Tool Audit + Access Workarounds (Meta-Research)
      
      **Question:** "What are the most effective web research tool combinations for deep research?"
      **Scope:** Broad (3 sonnet researchers, 2 rounds)
      **Report:** `~/.claude/research/web-research-tool-effectiveness/report.md`
      **Key findings:** Exa and Firecrawl MCP tools expose only a subset of their full API
      capabilities. Switched Exa to hosted endpoint, added academic search strategies, content
      access workarounds, and Firecrawl advanced features to researcher prompt.
      **Validated:** Lead-dispatched follow-ups (5th consecutive success), meta-research through
      the team skill is viable and effective.
      **Problems found:** None during run. Post-iteration discovered `contextMaxCharacters` bug
      suppressing Exa summaries/highlights.
      **Changes (6 to researcher-prompt.md + major post-iteration infrastructure):** Expanded
      Available Tools, added Firecrawl Search Operators, Academic Search, Content Extraction
      advanced features, Accessing Restricted Content sections. Post-iteration: switched Exa MCP
      to hosted endpoint, reduced tools from 5 to 2, added `enableSummary` + `textMaxCharacters: 1`
      default pattern, added category restriction table. Resolved backlog #18.
      
      **Cumulative skill improvements across 4 iterations:** 15 changes. Major post-iteration
      infrastructure overhaul to Exa tooling.
      
      ### Iteration 5 -- Scientific/Health Empirical (IF + Cognitive Performance)
      
      **Question:** "Does intermittent fasting improve cognitive performance?"
      **Scope:** Broad (3 sonnet researchers, 2 rounds)
      **Report:** `~/.claude/research/intermittent-fasting-cognitive-performance/report.md`
      **Key findings:** Acute fasting is cognitively neutral (meta-analysis of 63 studies). 26
      footnotes. Evidence pyramid template validated. Calibrated confidence language worked well.
      **Validated:** Exa summary pattern (32/32 compliant), lead-dispatched follow-ups (6th success),
      academic source quality strong.
      **Changes (3):** Implemented shutdown feedback reports (researcher-prompt.md + SKILL.md Phase 10),
      added `scripts/analyze-transcripts.py` for JSONL analysis. Resolved backlog #12, #23.
      
      **Cumulative skill improvements across 5 iterations:** 18 changes. All synthesis templates
      validated except Emerging/Frontier and Opinion/Sentiment.
      
      ### Iteration 6 -- Emerging/Frontier Empirical (LLM Reasoning)
      
      **Question:** "What are the current approaches to LLM reasoning?"
      **Scope:** Comprehensive (4 opus researchers, 3 rounds incl. verification)
      **Report:** `~/.claude/research/llm-reasoning-approaches/report.md`
      **Key findings:** ~400 lines, 55 footnotes. RLVR as dominant paradigm shift. PRM vs ORM conflict
      resolved (both correct in different regimes). CoT faithfulness 2.3% CONTESTED (single study),
      <20% verbalization SUPPORTED WITH NUANCE.
      **Validated:** Emerging/Frontier template, shutdown feedback reports (all 4 wrote files), category
      restriction table (tweet, news, personal site all worked), Exa compliance 49/49 (100%).
      **Changes (0):** No skill file changes. Identified backlog #26 (arxiv guidance) and #27
      (verification verdict expansion) for quick implementation.
      
      **Cumulative skill improvements across 6 iterations:** 18 changes. Only Opinion/Sentiment
      template remains untested.
      
      ### Iteration 7 -- Agent Teams Docs Audit (Meta)
      
      **Scope:** No research run -- audit skill against official Anthropic documentation.
      **Key findings:** Architecture fully validated against Anthropic's docs and their own research
      system. Hub-and-spoke, lead-dispatched, filesystem output, verification pattern all confirmed.
      Permission inheritance confirmed (teammates inherit lead's settings). Anthropic uses Opus lead +
      Sonnet subagents with 90.2% improvement over single-agent.
      **Changes (2):** arxiv scraping guidance (#26), SUPPORTED WITH NUANCE verdict (#27).
      
      ### Iteration 8 -- Scope Calibration Refinement (Meta)
      
      **Scope:** No research run -- analyze 7 empirical data points.
      **Key findings:** Current defaults correct 6/6 times. Consumer is only type with de-escalation
      evidence. Report length is scope-driven, not type-driven.
      **Changes (3):** Scope modifiers and empirical output expectations added to question-types.md,
      SKILL.md Phase 2 updated.
      
      **Cumulative skill improvements across 8 iterations:** 21 changes (counting Iteration 7's 2).
      
      ### Iteration 9 -- Opinion/Sentiment + Mixed Model Config (Empirical)
      
      **Question:** "What do developers actually think about Nix and NixOS in 2026?"
      **Scope:** Comprehensive (4 sonnet researchers, 2 rounds -- no Round 3 verification needed)
      **Report:** `~/.claude/research/developer-opinions-nix-nixos-2026/report.md`
      **Key findings:** ~370 lines, 50 footnotes, 60+ sources. Developer opinion defined by paradox:
      fastest community growth ever (30% YoY) alongside deepest fragmentation ever (governance crisis,
      three forks). Pragmatic "devshells only" is the dominant successful adoption pattern. Wrapper
      tools (Devbox 11K stars, devenv 6K, Flox 3.7K) work for 80% case, break at edges.
      **Tested:** Opinion/Sentiment template (last untested), Sonnet researchers in Comprehensive scope
      (backlog #7). Both validated successfully.
      **Changes (3):** Updated researcher-prompt.md Reddit section (restrict_sr=on, prefer /top/ over
      /search/, note curl permission cost), added GitHub data guidance (Exa code context + gh api,
      no curl), added Discourse JSON search API fallback.
      **Backlog items resolved:** #7 (mixed model config -- Sonnet validated for Comprehensive).
      
      **Cumulative skill improvements across 9 iterations:** 24 changes. All 7 synthesis templates
      tested. Sonnet validated as viable Comprehensive override (Opus kept as default).
      
      ### Iteration 10 -- Reddit MCP Server Integration (Engineering)
      
      **Scope:** No research run -- engineering integration.
      **Key findings:** Fresh evaluation confirmed jordanburke/reddit-mcp-server as best option
      (v1.2.1, actively maintained, 9 read tools + search, anonymous mode). Hawstein/mcp-server-reddit
      stale (no commits in 10 months, no search). New entrants too immature.
      **Changes (4):** Added reddit-mcp-server to 1MCP config (anonymous mode), added Reddit
      ToolSearch to researcher setup, added Reddit tool block to Available Tools section,
      replaced curl-based Reddit access with MCP tool guidance in researcher-prompt.md.
      **Backlog items resolved:** #8 (Reddit MCP server integration).
      
      **Cumulative skill improvements across 10 iterations:** 28 changes. Zero curl dependencies
      in researcher workflows (Reddit MCP + GitHub gh api).
      
      ### Iteration 11 -- Reddit MCP Empirical Validation
      
      **Question:** "What is the best community-favorite content about the TV show Dark?"
      **Scope:** Focused (2 sonnet researchers, 1 round)
      **Report:** `~/.claude/research/dark-tv-show-community-content/report.md`
      **Key findings:** 24-source report across 5 content categories (video essays, podcasts, Reddit
      mega-threads, visual companion tools, philosophical/academic writing). OneTake and Think Story
      top YouTube; r/DarK mega-threads are canonical community references; academic philosophy papers
      exist from Springer and Blackwell series.
      **Validated:** Reddit MCP integration end-to-end (9/9 goals passed). Both researchers loaded
      Reddit tools via ToolSearch, used Firecrawl search for discovery + Reddit MCP for extraction,
      zero curl usage, zero permission prompts from Reddit.
      **Problems found:** `firecrawl_search` with subreddit-scoped `site:reddit.com/r/DarK` returned
      generic results (broad `site:reddit.com` worked better). `get_top_posts` returns memes/meta for
      entertainment subreddits, insufficient alone for substantive threads.
      **Changes (2):** Updated researcher-prompt.md Reddit section: added Firecrawl search scoping
      note (broad > subreddit-specific), clarified get_top_posts limitation for entertainment subs.
      
      **Cumulative skill improvements across 11 iterations:** 33 changes (30 from iteration + 3
      post-iteration housekeeping: generalized transcript script, filesystem path support, simplified
      feedback template). Reddit MCP fully validated.
      
      ### Iteration 12 -- Firecrawl Scrapability Testing (Engineering)
      
      **Scope:** No research run -- empirical testing of all documented content workaround URLs.
      **Key findings:** archive.ph CAPTCHA-blocked (429) even with stealth/enhanced proxy.
      freedium.cfd DNS dead (replaced by freedium-mirror.cfd, works great). Bluesky bsky.app is a
      JavaScript SPA requiring `waitFor: 5000` (no proxy, 1 credit); raw API unnecessary.
      xcancel.com works but needs `proxy: "stealth"` + `waitFor: 5000` (5 credits). Wayback Machine
      shortcut `/web/{URL}` works -- Firecrawl follows redirect to latest snapshot (initial 404s were
      unarchived pages, not format errors). Lessons: validate test inputs (xcancel), test against
      known-good data (Wayback Machine).
      **Changes (5 to researcher-prompt.md):** Demoted archive.ph, simplified Wayback Machine to
      shortcut format, added xcancel proxy requirements, replaced freedium.cfd with
      freedium-mirror.cfd, documented Bluesky SPA waitFor requirement.
      **Backlog items resolved:** #12 (Firecrawl scrapability testing).
      
      **Cumulative skill improvements across 12 iterations:** 38 changes. All content access
      workarounds empirically validated. Researchers will no longer waste calls on dead endpoints.
      
      ### Iteration 13 -- Information Cascade Detection (Empirical + Engineering)
      
      **Question:** "How do journalists, fact-checkers, and researchers detect information cascades?"
      **Scope:** Broad (3 sonnet researchers, 2 rounds)
      **Report:** `~/.claude/research/information-cascade-detection/report.md`
      **Key findings:** Three converging research traditions (journalism, bibliometrics, computational
      social science) all address cascade detection independently. Greenberg 2009 foundational for
      citation distortion taxonomy. Woozle effect quantified (76% affirm debunked myths). Practitioner
      heuristics are tacit knowledge -- no org publishes explicit cascade checklist. Intelligence
      analysis (Heuer ACH, ICD 203) has most rigorous source independence frameworks. Composite
      signals (temporal clustering, textual overlap, network topology, metadata) most reliable.
      **Validated:** Lead-dispatched follow-ups (8th success), idle researcher management (B had no
      Round 2 task), all operational patterns stable.
      **Changes (2):** Added 5-step "Cascade Check" section to researcher-prompt.md with red flags
      and procedure. Added cascade check reference to Step 3 (Reflect) questions.
      **Backlog items resolved:** #7 (information cascade detection).
      
      **Cumulative skill improvements across 13 iterations:** 40 changes. Researchers now have
      explicit cascade detection heuristics in the reflection step. researcher-prompt.md at ~541 lines.
      
      ## Backlog
      
      Prioritized list of things to investigate or test. Re-prioritized each iteration.
      
      ### High Priority
      
      16. **[Engineering] Dedicated verifier agents** -- Restructure verification from reusing angle
          researchers to spawning lightweight, single-purpose verifier agents on-demand. Current
          approach wastes angle researchers' accumulated context and blocks investigation flow.
      
          **New design:** Angle researchers (A, B, C) own their angles for the full session and
          maintain context continuity. When the lead's triage identifies a high-impact single-source
          claim at any point, it spawns a dedicated verifier (sonnet, small context, focused task).
          The verifier receives only the claim + source URL, searches for independent evidence,
          reports a verdict (SUPPORTED / SUPPORTED WITH NUANCE / CONTESTED / UNCHANGED), and shuts
          down. Verifiers run in parallel with ongoing angle work -- no separate phase or round needed.
      
          **What this fixes:** (1) Removes the Comprehensive-only scope restriction -- any scope can
          verify because it's just a lightweight parallel agent. (2) Preserves angle researcher
          expertise -- they stay on their domain instead of context-switching to verify unrelated
          claims. (3) Eliminates the Phase 7 bottleneck -- verification no longer requires a
          dedicated round. (4) Cleaner separation of concerns -- investigation vs verification are
          different tasks that benefit from different agent configurations.
      
          **Implementation:** Refactor Phase 7 out of the sequential phase structure. Add verification
          dispatch logic to Phase 5 triage (and any subsequent triage). Update Effort Calibration
          table to show verification as available to all scopes. Update researcher spawn prompts
          (verifiers need a different, simpler prompt than angle researchers). Test empirically.
      
      17. **[Research] Confidence assessment methodology** -- The current triage confidence levels
          (HIGH/MEDIUM/LOW) are based primarily on independent source count. Source count is a crude
          proxy -- a single well-designed RCT can warrant more confidence than five blog posts citing
          each other. Source *quality* matters but is hard to assess systematically. Worth a meta-research
          run to explore: How do intelligence analysts, systematic reviewers, and epistemologists think
          about evidence quality? What frameworks exist beyond source counting? (Heuer's ACH and ICD 203
          from iteration 13 are starting points.) Goal: replace or augment the current source-count
          heuristic with something that accounts for source quality, methodology strength, and
          independence more holistically.
      
      5. **[Engineering] Permission setup / onboarding flow** -- Audit all permissions needed for
         the skill to run without prompts (mkdir, Write to research dir, researcher Bash calls for
         sleep/retry/curl). Build a setup check at skill invocation that verifies permissions exist
         and offers to add them. Permission inheritance now confirmed (Iteration 7 audit): teammates
         inherit the lead's permission settings. Iteration 9 confirmed curl is the biggest permission
         friction source. Key for making the skill shareable beyond Malo. Not a research task -- pure
         engineering.
      
      6. ~~**[Meta] Scope calibration guidance**~~ -- **Resolved (Iteration 8).** Added scope modifiers
         (de-escalation/escalation signals per type) and empirical output expectations to
         `references/question-types.md`. All 7 defaults validated. Consumer is the only type with
         observed de-escalation opportunity. Phase 2 updated to reference modifiers.
      
      7. ~~**[Empirical] Mixed model config for Comprehensive**~~ -- **Resolved (Iteration 9).** Tested
         Opus lead + 4 Sonnet researchers on Comprehensive scope (Opinion/Sentiment question). Sonnet
         validated as viable -- researchers completed both rounds, followed methodology, wrote structured
         findings and feedback. However, Opus kept as Comprehensive default: one successful run on a
         less demanding question type doesn't prove Sonnet handles the hardest Comprehensive questions
         (e.g., contested scientific topics with complex verification). SKILL.md notes Sonnet as an
         available override when cost matters.
      
      8. ~~**[Engineering] Reddit MCP server integration**~~ -- **Resolved (Iteration 10).** Integrated
         `jordanburke/reddit-mcp-server` via 1MCP with anonymous mode (~10 rpm). Researcher methodology
         updated to use native MCP tools (`search_reddit`, `get_top_posts`, `get_post_comments`, etc.)
         instead of curl. Eliminates ~8-10 permission prompts per research run. Zero curl dependencies
         remain in researcher workflows. Can upgrade to app-only auth (60 rpm) later if rate limits
         become an issue -- requires Reddit API application approval.
      
      ### Medium Priority
      
      7. ~~**[Meta] Information cascade detection**~~ -- **Resolved (Iteration 13).** Researched across
         three traditions (journalism, bibliometrics, computational social science). Distilled findings
         into a 5-step "Cascade Check" section in researcher-prompt.md with red flags and procedure.
         Researchers now check source independence during Step 3 (Reflect).
      
      8. ~~**[Meta] Report length targets per question type**~~ -- **Resolved (Iteration 8).** Added
         empirical output expectations table to `references/question-types.md` (Focused ~150-200
         lines/10-15 sources, Broad ~200-300/20-30, Comprehensive ~350-400+/40-60). Presented as
         descriptive ranges, not targets.
      
      9. **[Meta] Optimal angle decomposition** -- Research strategies for generating maximally
         complementary research angles. Could improve Phase 1 planning.
      
      10. **[Meta] Source diversity vs quality tradeoffs** -- When should researchers prefer
          breadth over authority? Could refine stopping criteria.
      
      11. **[Engineering] Exa `deep` search type via MCP** -- Exa's API supports `type: "deep"`
          which does automatic query expansion + parallel search + re-ranking ($0.015/search). This
          would replace the manual "generate 3-5 query variations" pattern with a single call +
          `additionalQueries`. Currently blocked: the hosted MCP endpoint (`mcp.exa.ai`) only
          exposes auto/fast/neural in the type enum. Monitor for when Exa adds deep to MCP.
      
      12. ~~**[Engineering] Firecrawl scrapability testing**~~ -- **Resolved (Iteration 12).** Tested
          all 7 documented workarounds with actual Firecrawl calls. 2 broken (archive.ph CAPTCHA,
          freedium.cfd DNS dead), 1 replaced (freedium-mirror.cfd works), 1 simplified (Wayback
          Machine shortcut `/web/{URL}` works -- initial 404s were unarchived pages), 2 updated
          with required params (xcancel.com needs stealth proxy, Bluesky bsky.app needs waitFor
          for SPA). Docs updated to prevent researchers from wasting calls on dead endpoints.
      
      13. **[Meta] Researcher-prompt.md length monitoring** -- File is ~438 lines. Researchers
          successfully followed all guidance in Iteration 6 (100% Exa compliance, feedback reports
          written). No evidence of instruction-following degradation. Continue monitoring but
          deprioritized -- no splitting needed yet.
      
      14. **[Meta/Engineering] Exa vs dedicated academic search** -- Iteration 6 confirmed strong
          academic search via `category: "research paper"` for CS/ML (arxiv, NeurIPS, ICLR, ACL,
          ICML sources found). Iterations 5+6 together show the current setup works well for both
          biomedical and CS domains. Deprioritized unless a specific academic search failure occurs.
      
      15. **[Engineering] Skill setup / environment bootstrapping** -- Investigate using skill
          `!` bang patterns or a setup phase to automatically verify/dump environment state at skill
          invocation. Related to backlog #5 (permission setup) but broader.
      
      ### Lower Priority
      
      13. **[Meta] Self-reflection patterns in multi-agent research** -- How does extended
          thinking interact with the investigation loop?
      
      14. **[Meta] Synthesis quality patterns** -- What makes a good research synthesis? Compare
          our output against professional research reports.
      
      15. **[Meta] Consumer Phase 1 constraint gathering** -- Should Consumer questions always get
          a quick "Any constraints?" even when the question seems clear? Minor UX improvement.
      
      16. **[Speculative] Researcher compaction robustness** -- Whether researchers need explicit
          compaction handling. Assessment: low risk. 6 iterations without compaction issues.
      
    • RESEARCH.md 13.3 KB
      # Deep Research Team v2 — Research Findings
      
      Research conducted 2026-02-06 using deep mode (4 Opus researchers, 2 rounds).
      50+ sources across academic papers, production systems, and practitioner reports.
      
      ## Core Question
      
      What is the optimal architecture for an LLM-agent-team-based research capability that handles
      the full spectrum of research questions — from factual lookups to subjective opinion synthesis
      to contested/controversial topics?
      
      ---
      
      ## Finding 1: Architecture-Task Alignment Matters More Than Agent Count
      
      Kim et al. (Google/MIT/DeepMind, Dec 2025) tested 180 configurations across 5 architectures,
      3 LLM families, and 4 benchmarks. Matching coordination topology to task structure is the
      critical variable, not scaling agent count[^1]. Centralized (hub-and-spoke) coordination
      improved parallelizable tasks by 80.8%, while every multi-agent variant degraded sequential
      reasoning by 39-70%. Communication overhead grows super-linearly (exponent 1.724). Effective
      team sizes: ~3-4 agents before costs outweigh benefits.
      
      Anthropic's production research system confirms: orchestrator-worker (Opus 4 lead + Sonnet 4
      subagents) outperformed single-agent Opus 4 by 90.2% on research tasks[^2]. Key insight:
      **token usage alone explains 80% of performance variance** on research benchmarks. Multi-agent
      architectures work primarily by scaling token usage across parallel context windows.
      
      Context dilution research (7+ studies: Stanford, Google, Meta, Microsoft, NVIDIA, Adobe,
      Chroma) shows more tokens in a single context _degrades_ performance (13.9-85% accuracy drop
      as length increases)[^3][^4][^5]. Multi-agent solves this by distributing tokens across separate
      windows — the fundamental architectural justification for agent teams.
      
      **Confidence: HIGH** (Kim et al. + Anthropic + 7+ context dilution studies)
      
      ## Finding 2: Classify-Then-Route Pattern Is Proven But Underdeployed
      
      Multiple working implementations demonstrate that classifying question type before selecting
      strategy improves accuracy and efficiency: n8n Adaptive RAG (4-type classification), LangGraph
      Adaptive RAG, Meilisearch (4 types with fallback escalation), Kore.ai Adaptive-RAG[^6][^7][^8][^9].
      Select-then-Route (EMNLP 2025) provides rigorous validation[^10].
      
      Yet no major commercial tool implements this. OpenAI Deep Research uses a single pipeline[^11].
      Perplexity requires manual mode control[^12]. Google has implicit sensitivity through E-E-A-T
      but no explicit classification[^13].
      
      All implementations converge on 4-5 categories. For compound questions, query decomposition
      handles this (Haystack, NVIDIA, AAAI 2023)[^14][^15][^16].
      
      **Confidence: HIGH** (4+ implementations + domain evidence; no major tool does this = differentiator)
      
      ## Finding 3: Evidence Hierarchies Must Vary by Question Type
      
      Four independent domains confirm:
      
      - **Evidence-based medicine**: Different "levels of evidence" for different question types (RCTs
        for therapy, cohort studies for prognosis, blind comparisons for diagnostics)[^17][^18]
      - **Library science**: Four question levels with escalating research requirements[^19]
      - **Consumer evaluation**: Expert testing + user satisfaction data as primary evidence (opposite
        of medical hierarchy)[^20][^21]
      - **Fact-checking**: Core methodology focuses exclusively on verifiable claims[^22]
      
      ACRL Framework: "Authority is Constructed and Contextual" — source credibility depends on the
      information need[^23]. Reddit is noise for health questions but signal for community opinion[^24].
      
      **Confidence: HIGH** (4 independent professional domains)
      
      ### Proposed Question-Type Taxonomy
      
      | Question Type               | Evidence Priority                        | Source Hierarchy                                                 | Synthesis Approach                                    |
      | --------------------------- | ---------------------------------------- | ---------------------------------------------------------------- | ----------------------------------------------------- |
      | **Factual/verifiable**      | Government databases, official records   | Authoritative ref > primary sources > news                       | Aggregate reasoning quality (AoR)                     |
      | **Scientific/health**       | Systematic reviews, RCTs, cohort studies | Peer-reviewed > gov health > expert > anecdotal                  | Evidence pyramid with study quality flags             |
      | **Consumer/recommendation** | Expert testing + user satisfaction data  | Expert reviewers > enthusiast forums > specs > ads               | Ranked recommendations (top/runner-up/budget)         |
      | **Technical/comparative**   | Documentation, benchmarks, adoption data | Official docs > code examples > expert blogs > SO/Reddit         | Working implementations + benchmarks                  |
      | **Opinion/sentiment**       | Community forums, surveys, social media  | Reddit/forums > blogs > news > academic surveys                  | Map distribution of views, assess representativeness  |
      | **Contested/controversial** | Multiple perspectives required           | Academic > think tanks > news analysis > advocacy (label biases) | Structure debate (MODS pattern), never single-verdict |
      | **Emerging/frontier**       | Recent sources only                      | Conference papers > expert blogs > vendor announcements          | Findings with recency/instability caveats             |
      
      ## Finding 4: Dynamic Roles and Complementarity Trump Fixed Assignments
      
      Multiple 2025 papers converge on dynamic role assignment > fixed roles: MetaGen (arXiv Jan 2026),
      MLC (ACL 2025), OSC (EMNLP 2025), AMAS (EMNLP Industry 2025), Heterogeneous Swarms[^25][^26][^27][^28][^29].
      
      Production systems use **topic-angle decomposition** (diversify by sub-question) not
      **functional-role decomposition** (searcher/critic/synthesizer)[^2][^30]. Practical pattern:
      topic decomposition during exploration, functional stages in the pipeline.
      
      Complementarity > diversity (NeurIPS Workshop 2025)[^31]. Implicit consensus with partial
      diversity outperforms forced agreement (EMNLP 2025)[^32]. DMoA (ICLR 2025): optimal
      diversity-consistency balance varies by task — factual favors consistency, exploratory
      favors diversity[^33].
      
      **Confidence: HIGH** (architecture pattern); **MEDIUM** (complementarity measurement in practice)
      
      ## Finding 5: Three-Stage Effort Calibration
      
      **Pre-research classification**: Question type, structural complexity, user intent. Anthropic's
      production rules: 1 agent/3-10 calls (simple), 2-4 subagents/10-15 calls (comparison),
      10+ subagents (complex research)[^2].
      
      **Mid-research adaptation**: Monitor source novelty, contradiction density, coverage adequacy.
      Dynamic subagent spawning only if warranted.
      
      **Stopping criteria**: Coverage-based (2+ sources per claim, citation convergence) + budget caps
      (OpenAI: 30-60 searches, 120-150 fetches, 20-30 min)[^34].
      
      Task complexity — not instruction specificity — drives failure (Tech Policy Institute, N=575):
      complexity reduces success by 24-37pp, instruction specificity only 8.6pp (not significant)[^35].
      
      Library science "reference interview" has AI analogs: STaR-GATE (72% preference), Ask-before-Plan
      (97.3% planning accuracy)[^36][^37]. Safest design: auto-classify, present plan, allow override.
      
      **Confidence: HIGH** (multi-source convergence)
      
      ## Finding 6: Synthesis Strategy Should Vary Along the Fact-Opinion Spectrum
      
      No current tool adapts synthesis by question type[^38][^39].
      
      - **Factual**: Aggregation of Reasoning > majority voting (1-12% improvement)[^40]
      - **Contested/debatable**: MODS framework (38-59% better coverage/balance) structures debate[^41];
        Plurals (CHI 2025) creates deliberative panels[^42]
      - **Confidence language**: Medium verbalized uncertainty optimizes trust (Xu et al., N=156)[^43];
        Kent's WEPs partially usable by LLMs but need calibration (Tang et al., 2026)[^44][^45]
      - **Recommendations**: Biased/opinionated AI improves user decision-making (Lai et al., N=2500)[^46];
        default to recommendation with dissent, not false balance
      
      **Confidence: HIGH** (peer-reviewed experimental studies)
      
      ## Finding 7: IC Tradecraft Provides Battle-Tested Analytical Frameworks
      
      CIA Structured Analytic Techniques[^47][^48]:
      
      - **ACH** (Analysis of Competing Hypotheses) — array evidence against hypotheses, focus on disproving
      - **Key Assumptions Check** — surface and challenge unstated assumptions
      - **Quality of Information Check** — systematic source reliability assessment
      
      Toulmin argumentation model (Data → Warrant → Claim + Qualifier + Rebuttal) forces explicit
      identification of inferential steps[^49][^50]. Useful analytically for lead's synthesis reasoning.
      
      **Confidence: HIGH** (established government/academic frameworks)
      
      ---
      
      ## Key Conflicts Resolved
      
      1. **"More tokens help" vs "more tokens hurt"**: Both true. Multi-agent distributes tokens across
         separate context windows (helps) rather than stuffing one window (hurts).
      
      2. **Kim et al. diminishing returns vs Anthropic 90.2% gains**: Research is parallelizable
         (where Kim et al. found the largest gains). The 45% saturation threshold applies to sequential tasks.
      
      3. **Centralized vs decentralized**: Anthropic's "centralized" is actually centralized-independent
         hybrid — orchestrator assigns/synthesizes, subagents work independently during exploration.
         Captures decentralized benefits with centralized quality control.
      
      ---
      
      ## Key Gaps (Unresolved)
      
      - No head-to-head comparison of question-type-adaptive vs one-size-fits-all research agents
      - No automated method for detecting information cascades (many sources citing one original)
      - No formal "fact-opinion spectrum" framework for routing synthesis strategies
      - Source diversity vs quality tradeoffs remain under-theorized
      - Cost-effectiveness curves for agent scaling are unpublished
      - Self-reflection + extended thinking interaction with multi-agent coordination understudied
      
      ---
      
      ## Design Implications for Skill Redesign
      
      1. **Add question-type classification as Phase 0** — lightweight LLM classification into 4-5 types,
         with decomposition for compound questions
      2. **Replace static quick/standard/deep with dynamic effort calibration** — auto-classify complexity,
         suggest tier, allow override, adapt mid-research
      3. **Vary source evaluation by question type** — different source hierarchies per type
      4. **Vary synthesis strategy by question type** — AoR for factual, MODS for contested,
         recommendations for consumer, opinion mapping for sentiment
      5. **Use calibrated confidence language** — Kent-style WEPs with explicit probability ranges
      6. **Keep centralized-independent hybrid** — validated by production systems and academic research
      7. **Keep topic-angle decomposition** — validated over functional-role decomposition for research
      8. **Add file-based persistence** — lead writes report + state files, researchers write findings to disk,
         enables cross-session resumability
      9. **Smarter clarification** — skip for simple factual, decompose for compound, plan-present for complex
      
      ---
      
      ## Sources
      
      [^1]: Kim et al., "Towards a Science of Scaling Agent Systems," arXiv:2512.08296, Dec 2025
      [^2]: Anthropic Engineering, "How we built our multi-agent research system," Jun 2025
      [^3]: Liu et al., "Lost in the Middle," Stanford/Meta, TACL 2024
      [^4]: "Context Length Alone Hurts," arXiv, Oct 2025
      [^5]: Chroma, "Context Rot," Jul 2025
      [^6]: n8n, "Adaptive RAG with Query Classification"
      [^7]: LangChain, "LangGraph Adaptive RAG"
      [^8]: Meilisearch, "Adaptive RAG Explained," Sep 2025
      [^9]: Kore.ai, "Adaptive-RAG," 2024
      [^10]: "Select-then-Route," EMNLP 2025 Industry Track
      [^11]: OpenAI, "Deep Research API Introduction," Jun 2025
      [^12]: Data Studios, "Perplexity AI Prompting Techniques," Dec 2025
      [^13]: Agenxus, "Google AI Overviews Source Prioritization," Nov 2025
      [^14]: Haystack, "Advanced RAG: Query Decomposition," Sep 2024
      [^15]: NVIDIA, "Query Decomposition for RAG Blueprint"
      [^16]: "Question Decomposition Tree," AAAI 2023
      [^17]: NHMRC Evidence Hierarchy
      [^18]: Icahn School of Medicine, "Evidence Based Medicine"
      [^19]: Hepler & Horalek, "Introduction to Library and Information Science"
      [^20]: Wirecutter, "Anatomy of a Guide," May 2024
      [^21]: Consumer Reports, "Rating Methods"
      [^22]: Ballotpedia, "Methodologies of Fact-Checking"
      [^23]: ACRL Framework for Information Literacy
      [^24]: White, "Reddit as Analogy for Scholarly Publishing," 2019
      [^25]: MetaGen, arXiv, Jan 2026
      [^26]: MLC, ACL 2025
      [^27]: OSC, EMNLP 2025
      [^28]: AMAS, EMNLP Industry 2025
      [^29]: Heterogeneous Swarms, OpenReview 2025
      [^30]: O-Researcher, arXiv:2601.03743, Jan 2026
      [^31]: Zhang, "Mixture of Complementary Agents," NeurIPS Workshop 2025
      [^32]: Wu & Ito, "The Hidden Strength of Disagreement," EMNLP 2025
      [^33]: "Balancing Act: DMoA," ICLR 2025
      [^34]: PromptLayer, "How OpenAI's Deep Research Works," Oct 2025
      [^35]: Lovin & Wallsten, Tech Policy Institute, Nov 2025
      [^36]: STaR-GATE, arXiv, 2024
      [^37]: Ask-before-Plan, 2024
      [^38]: 7minute.ai, "OpenAI vs Google Deep Research," 2025
      [^39]: Venkit et al., "Search Engines in an AI Era," 2024
      [^40]: Sun et al., "AoR," LREC-COLING 2024
      [^41]: Xiao et al., "MODS," NAACL 2025
      [^42]: Bakker et al., "Plurals," CHI 2025
      [^43]: Xu et al., "Confronting Verbalized Uncertainty," IJHCS 2025
      [^44]: Kent, "Words of Estimative Probability," CIA/CSI
      [^45]: Tang et al., npj Complexity 2026
      [^46]: Lai et al., arXiv:2508.09297, 2025
      [^47]: CIA Tradecraft Primer, 2009
      [^48]: RAND, RR1408, 2016
      [^49]: Verheij, "The Toulmin Argument Model in AI," 2009
      [^50]: Freedman et al., "Argumentative LLMs," AAAI 2025
      
  • references
    • templates
      • consumer.md 916 B
        # Synthesis Template: Consumer
        
        Ranked recommendations with methodology.
        
        ```markdown
        ## Executive Summary
        
        [2-3 sentences: key finding, confidence level, scope]
        
        ## Recommendations
        
        ### Top Pick: [Product]
        
        [Why, with evidence from expert testing]
        
        ### Runner-Up: [Product]
        
        [Why, differentiation from top pick]
        
        ### Budget Pick: [Product]
        
        [Trade-offs at this price point]
        
        ### Best for {Use Case}: [Product]
        
        [Specific use-case fit]
        
        ## Methodology
        
        [Criteria used, sources consulted, testing methodology of source reviewers]
        
        ## Confidence Assessment
        
        [Use calibrated language from the Confidence table in SKILL.md]
        
        - Almost certain (93-99%): [findings]
        - Highly likely (80-92%): [findings]
        - Likely (63-79%): [findings]
        - Roughly even / Unlikely: [findings, if any]
        
        ## Limitations & Gaps
        
        [What remains uncertain, missing, or contested]
        
        ## Sources
        
        [^1]: Author/Org. "Title". Publication. URL
        
        [^2]: ...
        ```
        
      • contested.md 994 B
        # Synthesis Template: Contested
        
        Structured debate -- NEVER resolve with a verdict.
        
        ```markdown
        ## Executive Summary
        
        [2-3 sentences: key finding, confidence level, scope]
        
        ## Perspectives
        
        ### Position A: [Label]
        
        [Strongest case for this position, steelmanned]
        
        ### Position B: [Label]
        
        [Strongest case for this position, steelmanned]
        
        ## Points of Agreement
        
        [Where the sides converge]
        
        ## Key Disagreements
        
        [Where they diverge and why -- identify factual vs. values-based disagreements]
        
        ## Evidence Quality by Position
        
        [Assess the strength of evidence supporting each side -- this is NOT a verdict]
        
        ## Confidence Assessment
        
        [Use calibrated language from the Confidence table in SKILL.md]
        
        - Almost certain (93-99%): [findings]
        - Highly likely (80-92%): [findings]
        - Likely (63-79%): [findings]
        - Roughly even / Unlikely: [findings, if any]
        
        ## Limitations & Gaps
        
        [What remains uncertain, missing, or contested]
        
        ## Sources
        
        [^1]: Author/Org. "Title". Publication. URL
        
        [^2]: ...
        ```
        
      • emerging-frontier.md 899 B
        # Synthesis Template: Emerging/Frontier
        
        State of play with instability warning.
        
        ```markdown
        ## Executive Summary
        
        [2-3 sentences: key finding, confidence level, scope]
        
        ## Current State (as of {date})
        
        [What is known now, with explicit date markers on key facts]
        
        ## Trajectory
        
        _Note: The following is projection based on current trends, not established fact._
        
        [Where things appear to be heading]
        
        ## Instability Warning
        
        [What is most likely to change and why -- flag specific claims with short shelf life]
        
        ## Confidence Assessment
        
        [Use calibrated language from the Confidence table in SKILL.md]
        
        - Almost certain (93-99%): [findings]
        - Highly likely (80-92%): [findings]
        - Likely (63-79%): [findings]
        - Roughly even / Unlikely: [findings, if any]
        
        ## Limitations & Gaps
        
        [What remains uncertain, missing, or contested]
        
        ## Sources
        
        [^1]: Author/Org. "Title". Publication. URL
        
        [^2]: ...
        ```
        
      • factual.md 752 B
        # Synthesis Template: Factual
        
        Clear answer with evidence weight.
        
        ```markdown
        ## Executive Summary
        
        [2-3 sentences: key finding, confidence level, scope]
        
        ## Verdict
        
        [Direct answer to the question with confidence level]
        
        ## Supporting Evidence
        
        ### [Claim 1]
        
        [Evidence with inline citations, weighted by source authority]
        
        ### [Claim 2]
        
        [Continue same pattern]
        
        ## Confidence Assessment
        
        [Use calibrated language from the Confidence table in SKILL.md]
        
        - Almost certain (93-99%): [findings]
        - Highly likely (80-92%): [findings]
        - Likely (63-79%): [findings]
        - Roughly even / Unlikely: [findings, if any]
        
        ## Limitations & Gaps
        
        [What remains uncertain, missing, or contested]
        
        ## Sources
        
        [^1]: Author/Org. "Title". Publication. URL
        
        [^2]: ...
        ```
        
      • opinion-sentiment.md 893 B
        # Synthesis Template: Opinion/Sentiment
        
        Distribution mapping with representativeness assessment.
        
        ```markdown
        ## Executive Summary
        
        [2-3 sentences: key finding, confidence level, scope]
        
        ## Distribution of Views
        
        ### Majority View (~X%)
        
        [What most people think and why]
        
        ### Minority View (~X%)
        
        [Dissenting perspective and reasoning]
        
        ### Outlier Positions
        
        [Fringe views worth noting]
        
        ## Representativeness Assessment
        
        [How representative is the source sample? What communities were/weren't covered?]
        
        ## Confidence Assessment
        
        [Use calibrated language from the Confidence table in SKILL.md]
        
        - Almost certain (93-99%): [findings]
        - Highly likely (80-92%): [findings]
        - Likely (63-79%): [findings]
        - Roughly even / Unlikely: [findings, if any]
        
        ## Limitations & Gaps
        
        [What remains uncertain, missing, or contested]
        
        ## Sources
        
        [^1]: Author/Org. "Title". Publication. URL
        
        [^2]: ...
        ```
        
      • scientific-health.md 1012 B
        # Synthesis Template: Scientific/Health
        
        Evidence pyramid with study quality assessment.
        
        ```markdown
        ## Executive Summary
        
        [2-3 sentences: key finding, confidence level, scope]
        
        ## Evidence Summary
        
        ### Systematic Reviews & Meta-Analyses
        
        [Findings from highest-quality evidence]
        
        ### Randomized Controlled Trials
        
        [Key RCT findings]
        
        ### Observational Studies
        
        [Cohort/case-control findings]
        
        ## Study Quality Notes
        
        [Sample sizes, methodology concerns, funding disclosures]
        
        ## Practical Implications
        
        _Note: The following reflects interpretation of the evidence, not sourced fact._
        
        [What the evidence means in practice]
        
        ## Confidence Assessment
        
        [Use calibrated language from the Confidence table in SKILL.md]
        
        - Almost certain (93-99%): [findings]
        - Highly likely (80-92%): [findings]
        - Likely (63-79%): [findings]
        - Roughly even / Unlikely: [findings, if any]
        
        ## Limitations & Gaps
        
        [What remains uncertain, missing, or contested]
        
        ## Sources
        
        [^1]: Author/Org. "Title". Publication. URL
        
        [^2]: ...
        ```
        
      • technical.md 942 B
        # Synthesis Template: Technical
        
        Comparison with context and benchmarks.
        
        ```markdown
        ## Executive Summary
        
        [2-3 sentences: key finding, confidence level, scope]
        
        ## Comparison Matrix
        
        | Criterion   | Option A | Option B | Option C |
        | ----------- | -------- | -------- | -------- |
        | [criterion] | ...      | ...      | ...      |
        
        ## Detailed Analysis
        
        ### [Criterion 1]
        
        [Evidence with benchmarks and citations]
        
        ### [Criterion 2]
        
        [Continue same pattern]
        
        ## Recommendation
        
        [Context-dependent recommendation with caveats about when each option wins]
        
        ## Confidence Assessment
        
        [Use calibrated language from the Confidence table in SKILL.md]
        
        - Almost certain (93-99%): [findings]
        - Highly likely (80-92%): [findings]
        - Likely (63-79%): [findings]
        - Roughly even / Unlikely: [findings, if any]
        
        ## Limitations & Gaps
        
        [What remains uncertain, missing, or contested]
        
        ## Sources
        
        [^1]: Author/Org. "Title". Publication. URL
        
        [^2]: ...
        ```
        
    • question-types.md 5.3 KB
      # Question Type Classification Reference
      
      Classify the user's question into one or more types before planning research. For compound
      questions, decompose into sub-questions and classify each independently.
      
      ## Taxonomy
      
      | Type                  | Signals                                                    | Example Questions                                            |
      | --------------------- | ---------------------------------------------------------- | ------------------------------------------------------------ |
      | **Factual**           | Verifiable claims, dates, numbers, "what is", "how many"   | "What year was GDPR enacted?" "How many GPUs did GPT-4 use?" |
      | **Scientific/Health** | Medical, biological, clinical, "is X safe", evidence-based | "Does intermittent fasting reduce inflammation?"             |
      | **Consumer**          | "Best X for Y", recommendations, buying decisions          | "Best noise-cancelling headphones under $300 for travel?"    |
      | **Technical**         | APIs, tools, benchmarks, "how to", architecture decisions  | "Redis vs Memcached for session caching at 10k RPS?"         |
      | **Opinion/Sentiment** | "What do people think", community views, satisfaction      | "How do developers feel about Tailwind CSS in 2026?"         |
      | **Contested**         | Political, ethical, actively debated, no consensus         | "Should AI models be open-sourced?"                          |
      | **Emerging/Frontier** | Very recent, rapidly evolving, limited sources             | "What are the leading approaches to LLM reasoning in 2026?"  |
      
      ## Compound Question Decomposition
      
      Many real questions span multiple types. Decompose and classify each part:
      
      > "What's the best RAG framework, and is RAG even the right approach for my use case?"
      
      - Sub-question 1: "What's the best RAG framework?" -> **Technical** (comparison)
      - Sub-question 2: "Is RAG the right approach for [use case]?" -> **Technical** (architecture decision)
      
      > "Is creatine safe, and which brand should I buy?"
      
      - Sub-question 1: "Is creatine safe?" -> **Scientific/Health**
      - Sub-question 2: "Which brand should I buy?" -> **Consumer**
      
      **Ordering rule for compound questions:** Investigate higher-complexity sub-questions first
      (Contested > Emerging > Scientific > Technical > Consumer > Opinion > Factual). Later
      sub-questions often depend on earlier answers.
      
      ## Type-to-Scope Defaults
      
      Each type has a natural default scope. Present the recommended scope with brief rationale
      and let the user confirm. Then apply the scope modifiers below to adjust.
      
      | Type                  | Default Scope | Rationale                                        |
      | --------------------- | ------------- | ------------------------------------------------ |
      | **Factual**           | Focused       | Few authoritative sources suffice                |
      | **Scientific/Health** | Broad         | Need evidence hierarchy across study types       |
      | **Consumer**          | Broad         | Need expert reviews + user satisfaction + specs  |
      | **Technical**         | Broad         | Need docs + benchmarks + practitioner experience |
      | **Opinion/Sentiment** | Broad         | Need representative sample across communities    |
      | **Contested**         | Comprehensive | Must steelman all major positions                |
      | **Emerging/Frontier** | Comprehensive | Sources are sparse; need breadth to find them    |
      
      ### Scope Modifiers
      
      After identifying the default scope, check for these signals to adjust up or down one tier.
      
      **De-escalate one tier when:**
      
      - **Consumer**: Narrow category with fewer than 5 viable options and established expert
        reviews (e.g., "best USB-C hub for MacBook"). When expert consensus is clear, Round 2
        adds only marginal value.
      - **Scientific/Health**: Single well-established finding with existing meta-analyses and no
        active controversy over mechanisms.
      - **Technical**: Well-documented comparison with published benchmarks and clear community
        consensus (e.g., mature tools with head-to-head comparisons already available).
      - **Opinion/Sentiment**: Question targets a single community with low controversy.
      
      **Escalate one tier when:**
      
      - **Factual**: Preliminary search reveals conflicting sources or the fact is embedded in a
        contested narrative.
      - **Consumer**: Emerging product category with sparse or unreliable reviews.
      - **Technical**: Novel or niche tools with sparse documentation, or the question involves
        architectural tradeoffs without clear benchmarks.
      - **Opinion/Sentiment**: Cross-community disagreement, politically charged, or culture-war
        adjacent topics.
      - **Any type**: Compound question with 3+ sub-questions of different types.
      
      **Do not use this skill when:**
      
      - The question is answerable with 1-2 web searches.
      - The question is a simple factual lookup with a single authoritative source.
      - The question is about debugging, code, or tasks better handled inline.
      
      ## Classification Procedure
      
      1. Read the question. Identify surface signals (keywords, phrasing, domain).
      2. Check for compound structure. If present, decompose into sub-questions.
      3. Assign a primary type to each sub-question (or the whole question if simple).
      4. Note the default scope from the table above.
      5. Proceed to Phase 1 (Clarify) with the classification ready.
      
      Classification is silent -- do not present it to the user. It informs scope defaults,
      source evaluation priorities, and synthesis template selection.
      
    • researcher-prompt.md 13.1 KB
      # Deep Researcher Methodology
      
      You are a researcher on a coordinated team. Your job is to investigate assigned tasks by
      searching the web, evaluating sources, and reporting structured findings back to the team lead.
      
      ## Setup (Do This First)
      
      Before starting any research:
      
      1. **Load your search toolkit** by invoking the `search-tips` skill using the Skill tool.
         This provides search strategy, tool guidance, and site workarounds.
      
      2. **Read your assigned task** using `TaskGet` with the task ID from your spawn prompt.
      
      3. **Begin investigation** following the loop below.
      
      ## Investigation Loop
      
      For each assigned task:
      
      1. **Search**: Formulate 2-4 queries using Exa (with summaries enabled) to discover and
         triage sources. Then scrape the promising ones with Firecrawl to read the actual content.
         **Summaries are for triage only** — any claim you include in your findings must be based
         on text you read from the source itself, not from any summary (AI-generated or otherwise)
      2. **Evaluate**: Assess findings for credibility, recency, depth, and bias
      3. **Reflect** (mandatory before synthesizing): Ask yourself:
         - Do my key claims have 2+ independent sources?
         - Are there contradictions between sources I haven't resolved?
         - Did a source reference a primary source I should find directly?
         - Am I missing a perspective (e.g., only found supporters, no critics)?
         - Did I find something surprising that deserves deeper investigation?
         - Am I relying on search snippets, or did I read the full content?
         - **Cascade check** (see below): Are my "multiple sources" actually independent?
      4. **Decide** based on reflection:
         - Gaps, contradictions, or thin coverage found -> formulate follow-up queries -> step 1
         - Citation convergence + good multi-source coverage -> step 5
      5. **Synthesize**: Structure findings into ~120 lines following the output format
      6. **Report**: Write findings to disk, notify the lead with the file path, mark task complete
      
      **Search budget:** Aim for 2-3 search rounds. Do not stop after round 1 unless you
      genuinely hit citation convergence. Do not exceed 4 rounds.
      
      ## Source Evaluation
      
      Evaluate each source for credibility, recency, depth, and bias. Then apply the type-specific
      hierarchy below. The question type is provided in your spawn prompt.
      
      ### Type-Specific Source Priorities
      
      **Factual/Verifiable:**
      Prioritize government databases, official records, authoritative references.
      Deprioritize blog posts, opinion pieces, secondary news coverage.
      A single authoritative primary source can be sufficient.
      
      **Scientific/Health:**
      Prioritize systematic reviews, then RCTs, then cohort studies, then case reports, then expert opinion.
      Deprioritize anecdotal evidence, supplement vendor sites, non-peer-reviewed claims.
      Note sample sizes, funding sources, and methodology quality.
      
      **Consumer/Recommendation:**
      Prioritize expert testing labs (Wirecutter, RTINGS, Consumer Reports), then enthusiast forums
      with hands-on experience, then manufacturer specs. Deprioritize paid reviews, affiliate-heavy
      listicles, AI-generated roundups with no testing methodology.
      
      For consumer research specifically: start Reddit early — practitioner threads surface real
      failure modes and durability data that reviews miss. Try cross-domain search terms when
      obvious ones return poor results (e.g., "studio rack" instead of just "server rack").
      
      **Technical/Comparative:**
      Prioritize official documentation, then working code examples, then expert blog posts with
      benchmarks, then Stack Overflow accepted answers. Deprioritize outdated tutorials, vendor
      marketing, theoretical comparisons without benchmarks.
      
      **Opinion/Sentiment:**
      Prioritize community forums (Reddit, HN, specialized forums), then surveys with methodology,
      then blog posts from practitioners. Deprioritize cherry-picked quotes, astroturfing indicators,
      single-voice opinion pieces presented as consensus.
      
      **Contested/Controversial:**
      Prioritize peer-reviewed research, then think tank analyses (label institutional bias), then
      long-form journalism, then advocacy organizations (label explicitly). Deprioritize social media hot takes,
      anonymous claims, sources that only present one side.
      
      **Emerging/Frontier:**
      Prioritize conference papers and preprints (with recency weight), then expert blogs from
      practitioners, then vendor announcements (label as interested party). Deprioritize older sources
      (>6 months may be outdated), speculative commentary without evidence.
      
      ## Cascade Check (Source Independence)
      
      Before claiming "2+ sources agree," verify they are genuinely independent. Information cascades
      occur when many sources trace back to one original, creating an illusion of corroboration.
      
      **Run this check during Step 3 (Reflect) for every key claim:**
      
      1. **Trace the citation chain.** For each source supporting a claim, ask: where did *they* get
         this? If two articles both cite the same original study, wire report, or press release, you
         have one source, not two. Follow citations backward until you hit primary evidence.
      
      2. **Check for copied phrasing.** Identical wording, the same errors/typos, or verbatim quotes
         across sources that claim independent reporting is the strongest cascade signal. If three
         articles contain the same unusual phrasing, they likely share a common upstream source.
      
      3. **Assess source type diversity.** Two blog posts citing the same press release is replication,
         not corroboration. True independence requires different *types* of evidence converging:
         e.g., an empirical study + practitioner experience + official data. Multiple sources of the
         same type (all news articles, all blog posts) with the same factual claims are suspect.
      
      4. **Apply the "how do they know?" test.** For each source, classify their knowledge as:
         - **First-hand**: direct observation, original research, primary data
         - **Second-hand**: told by a first-hand source, reporting on primary evidence
         - **Third-hand**: citing other secondary sources, "reports say"
         Only first-hand sources with independent access to the underlying reality count as genuine
         corroboration. Multiple second-hand sources tracing to the same first-hand source are a
         cascade.
      
      5. **Watch for wire service cascades.** When the same claim appears across many news outlets
         simultaneously with identical framing, check whether they all republished AP/Reuters wire
         copy. This is legitimate distribution but counts as one source, not many.
      
      **Red flags that suggest a cascade rather than independent corroboration:**
      
      - All sources appeared within a short time window (hours) with no independent reporting
      - No source adds local detail, original quotes, or independent verification
      - Removing one original source would collapse the entire evidence base
      - Sources are all the same type (all news, all blog posts, all citing one study)
      - A claim is widely repeated but every version traces to one interview, press release, or study
      
      **When you detect a cascade:** Downgrade confidence to single-source, note it in your
      CROSS-SOURCE ANALYSIS as "apparent multi-source but single origin," and search for genuinely
      independent evidence. If none exists, flag it in GAPS.
      
      **SEO content farms and information cascades:** AI-generated listicles, vendor marketing
      disguised as independent reviews, and content farms create an illusion of source breadth. When
      multiple sources make the same claim, verify they aren't all citing the same upstream source.
      Five blog posts citing one Bloomberg article is one source, not five. Deprioritize "Complete
      Guide 2026" articles with no original testing or methodology. For vendor-vs-vendor comparisons,
      independent benchmarks with published methodology outweigh vendor blog posts.
      
      ## Stopping Criteria
      
      **Continue searching if:**
      
      - Important claims rest on a single source
      - Sources contradict each other without resolution
      - You found only secondary sources -- hunt for the primary
      - Critical subtopics or perspectives remain uncovered
      - New searches still yield novel sources
      
      **Stop searching if:**
      
      - **Citation convergence**: same authoritative sources across multiple queries
      - Same URLs keep appearing in results
      - 3+ rounds without new substantive sources
      - 2+ credible sources for each major claim
      
      **Be explicit in reasoning:**
      
      - "Round 1 found 3 sources but all vendor blogs -> targeting independent reviews"
      - "Sources A and B contradict on pricing -> searching specifically for current pricing"
      - "Last two searches returned known URLs -> convergence reached, synthesizing"
      
      ## Handling Insufficient Evidence
      
      - **Explicitly state** "Insufficient evidence found for [X]" rather than inferring
      - **Do NOT synthesize** plausible-sounding content to fill gaps
      - **Flag gaps prominently** in your GAPS section
      - **Return fewer, well-sourced findings** rather than padding with uncertain claims
      - When the absence of evidence *is* the finding (e.g., "no major vendor supports this
        feature"), report it with the same rigor as positive findings: what you searched for, how
        many sources you checked, and why the absence is meaningful
      - When researching non-English-speaking subjects or regions, flag the language limitation in
        GAPS — key sources may exist in languages the current toolset cannot search
      
      ## Team Coordination
      
      ### Reporting Findings
      
      After completing investigation, do ALL THREE of these steps in order:
      
      **Step 1: Persist findings to disk.** Write your structured findings to
      `{output_dir}/researcher-{letter}-findings.md` using the Write tool (always use **lowercase**
      for the letter, e.g., `researcher-a-findings.md`). The output directory and your researcher
      letter are provided in your spawn prompt. For follow-up tasks, append with a suffix:
      `researcher-{letter}-findings-{task-id}.md`. This backup enables cross-session resume and
      survives if team coordination fails.
      
      **Step 2: Notify the lead.** Send the file path, not the full findings (avoids doubling
      output tokens). The lead's name is provided in your spawn prompt.
      
      ```
      SendMessage:
        type: "message"
        recipient: "{lead-name}"
        content: "Findings written to {output_dir}/researcher-{letter}-findings.md"
        summary: "Findings ready: {angle topic}"
      ```
      
      **Step 3: Mark task complete.**
      
      ```
      TaskUpdate: { taskId: "{id}", status: "completed" }
      ```
      
      Then go idle. The lead will message you with follow-up tasks if needed.
      
      ### Between Tasks
      
      After completing a task, go idle. The lead dispatches follow-up tasks directly via
      `SendMessage` -- watch for incoming messages. When you receive a task assignment, claim
      it with `TaskUpdate` and begin immediately. Note: you may receive echo notifications of
      your own task completions — ignore these.
      
      ### Shutdown
      
      When you receive a `shutdown_request`, write a brief feedback file BEFORE approving shutdown:
      
      **Step 1: Write feedback file** to `{output_dir}/researcher-{letter}-feedback.md`:
      
      Reflect on your research experience — what would help the next researcher on a similar task?
      Focus on qualitative observations, not tool call logs (those are extracted automatically from
      transcripts). Keep it brief (5-15 lines).
      
      ```markdown
      ## What Worked
      
      - [What strategies or tools were most effective for this task?]
      
      ## Issues
      
      - [What was frustrating, broken, or confusing? Tool failures, unclear instructions, dead ends?]
      
      ## Suggestions
      
      - [What would have made this easier? Missing tools, better guidance, different approach?]
      ```
      
      **Step 2: Approve shutdown:**
      
      ```
      SendMessage:
        type: "shutdown_response"
        request_id: "{from the request}"
        approve: true
      ```
      
      ## Output Format (STRICT)
      
      Your findings MUST follow this structure. Aim for ~120 lines.
      
      ```markdown
      ## ANGLE: [The specific question/angle you investigated]
      
      ## FINDINGS
      
      ### 1. [Claim title]
      
      [2-3 sentences with evidence] -- Source: [Author/Org], [Date]. [URL]
      Credibility: [High/Medium/Low]
      
      ### 2. [Claim title]
      
      [Continue same pattern -- aim for 3-5 findings total]
      
      ## CROSS-SOURCE ANALYSIS
      
      - Consensus (2+ sources agree): [list claims]
      - Conflicts: [list with both sides]
      - Single-source (lower confidence): [list claims]
      
      ## CONFIDENCE
      
      - High: [findings]
      - Medium: [findings]
      - Low/Uncertain: [findings]
      
      ## GAPS
      
      [What couldn't you find? What remains unanswered?]
      
      ## SOURCES
      
      1. [Title](URL) -- [credibility note]
      2. [Continue...]
      ```
      
      ## Critical Rules
      
      1. **NEVER return raw scraped content** -- only structured findings
      2. **ALWAYS cite sources inline** -- every factual claim needs attribution
      3. **ALWAYS assess confidence** -- distinguish multi-source from single-source claims
      4. **Aim for ~120 lines** -- compress aggressively but preserve important nuance
      5. **Flag contradictions explicitly** -- never silently pick one side
      6. **Persist, THEN notify, THEN go idle** -- follow the 3-step sequence in
         "Reporting Findings" every time: write to disk, notify lead with file path, mark task complete
      7. **Use lowercase filenames** -- always `researcher-a-findings.md`, never uppercase letters
      8. **Stay on-angle** -- investigate the specific topic in your assigned task, not adjacent topics
      9. **Verify claims against source text** -- summaries (AI-generated or human-written) are for
         triage and discovery only. Every factual claim in your findings must be grounded in text you
         read directly from the source. If you can't scrape the source, say so in GAPS rather than
         reporting the summary as fact
      
  • scripts
    • analyze-transcripts.py 14.6 KB
      #!/usr/bin/env python3
      """Analyze subagent transcripts from a deep-research-team session.
      
      Extracts MCP tool call parameters from researcher subagent JSONL transcripts
      to verify guidance compliance and understand search patterns.
      
      Usage:
          # Auto-detect: finds sessions with subagents matching a keyword
          python3 analyze-transcripts.py "intermittent fasting"
      
          # Explicit session ID
          python3 analyze-transcripts.py --session <session-id>
      
          # List recent sessions that have subagents
          python3 analyze-transcripts.py --list
      
      Output: Per-researcher summary of all MCP tool calls grouped by server, plus
      an aggregate compliance report for the Exa summary pattern.
      """
      
      import argparse
      import json
      import os
      import re
      import sys
      from pathlib import Path
      
      
      # MCP tool name pattern: mcp__1mcp__{server}_1mcp_{tool}
      # After ToolSearch loading, tools appear as: mcp__1mcp__{server}_1mcp_{tool}
      MCP_PATTERN = re.compile(r"(?:mcp__1mcp__)?(\w+?)_1mcp_(\w+)")
      
      
      def parse_mcp_tool_name(name):
          """Extract server and tool from an MCP tool name.
      
          Returns (server, tool) or None if not an MCP tool.
          """
          m = MCP_PATTERN.match(name)
          if m:
              return m.group(1), m.group(2)
          return None
      
      
      def find_project_dir(path=None):
          """Find the Claude project directory.
      
          Args:
              path: Optional path override. Accepts either:
                  - A filesystem path (e.g., /Users/malo/.config/nix-config) which
                    gets converted to the ~/.claude/projects/{slug} format
                  - A direct ~/.claude/projects/{slug} path (used as-is if it exists)
                  - None to auto-detect from cwd
          """
          base = Path.home() / ".claude" / "projects"
      
          if path:
              p = Path(path).expanduser().resolve()
              # If it's already a valid project dir, use it directly
              if p.exists() and p.parent == base:
                  return p
              # Otherwise treat it as a filesystem path and convert to slug
              source = str(p)
          else:
              source = os.getcwd()
      
          # Claude Code slugs: replace all non-alphanumeric characters with dashes
          slug = re.sub(r"[^a-zA-Z0-9]", "-", source)
          candidates = list(base.glob(f"{slug}*"))
          if len(candidates) == 1:
              return candidates[0]
          elif len(candidates) > 1:
              # Pick the one matching most closely
              for c in candidates:
                  if c.name == slug:
                      return c
              return candidates[0]
          return None
      
      
      def list_sessions(project_dir):
          """List sessions that have subagent directories, sorted by recency."""
          sessions = []
          for d in project_dir.iterdir():
              if d.is_dir() and (d / "subagents").is_dir():
                  subagent_count = len(list((d / "subagents").glob("agent-*.jsonl")))
                  if subagent_count > 0:
                      mtime = max(f.stat().st_mtime for f in (d / "subagents").glob("*.jsonl"))
                      sessions.append((d.name, subagent_count, mtime))
          sessions.sort(key=lambda x: x[2], reverse=True)
          return sessions
      
      
      def find_session_by_keyword(project_dir, keyword):
          """Find sessions whose subagents mention a keyword in their spawn message."""
          matches = []
          for d in project_dir.iterdir():
              subdir = d / "subagents"
              if not d.is_dir() or not subdir.is_dir():
                  continue
              for f in subdir.glob("agent-*.jsonl"):
                  if "compact" in f.name:
                      continue
                  try:
                      with open(f) as fh:
                          first = json.loads(fh.readline())
                          content = str(first.get("message", {}).get("content", ""))
                          if keyword.lower() in content.lower():
                              matches.append(d.name)
                              break
                  except (json.JSONDecodeError, OSError):
                      continue
          return list(set(matches))
      
      
      def parse_subagent(path):
          """Parse a subagent JSONL file, extracting identity and tool calls."""
          researcher = None
          task_subject = None
          question_type = None
          # mcp_calls: {server: [{tool, params}]}
          mcp_calls = {}
      
          with open(path) as f:
              for line in f:
                  try:
                      d = json.loads(line)
                  except json.JSONDecodeError:
                      continue
      
                  msg = d.get("message", {})
                  content = msg.get("content", "")
      
                  # Extract identity from spawn message or follow-up context
                  if isinstance(content, str):
                      m = re.search(r"researcher letter: (\w)", content, re.IGNORECASE)
                      if m:
                          researcher = m.group(1).upper()
                      m2 = re.search(r'task is #\d+: "([^"]+)"', content)
                      if m2:
                          task_subject = m2.group(1)
                      # Follow-up tasks have the subject in a different format
                      if not task_subject:
                          m2b = re.search(r'#\d+ -- "([^"]+)"', content)
                          if m2b:
                              task_subject = m2b.group(1)
                      m3 = re.search(r"Question type: (\w+)", content)
                      if m3:
                          question_type = m3.group(1)
      
                  # Extract identity from Write calls (follow-up agents write
                  # to researcher-{letter}-findings-{id}.md)
                  if isinstance(content, list):
                      for block in content:
                          if (
                              isinstance(block, dict)
                              and block.get("type") == "tool_use"
                              and block.get("name") == "Write"
                              and not researcher
                          ):
                              fp = block.get("input", {}).get("file_path", "")
                              m = re.search(r"researcher-([a-z])-findings", fp)
                              if m:
                                  researcher = m.group(1).upper()
      
                  # Extract MCP tool calls
                  if isinstance(content, list):
                      for block in content:
                          if not isinstance(block, dict) or block.get("type") != "tool_use":
                              continue
                          name = block.get("name", "")
                          inp = block.get("input", {})
      
                          parsed = parse_mcp_tool_name(name)
                          if not parsed:
                              continue
                          server, tool = parsed
                          mcp_calls.setdefault(server, []).append(
                              {"tool": tool, "params": dict(inp)}
                          )
      
          return {
              "researcher": researcher,
              "task": task_subject,
              "question_type": question_type,
              "mcp_calls": mcp_calls,
          }
      
      
      def format_exa_call(call):
          """Format an Exa call with compliance checking."""
          p = call["params"]
          q = p.get("query", "")[:80]
      
          has_summary = p.get("enableSummary") is True
          has_text_max = p.get("textMaxCharacters") == 1
          has_ctx_max = "contextMaxCharacters" in p
      
          compliance = ""
          if has_summary and has_text_max and not has_ctx_max:
              compliance = " [COMPLIANT]"
          elif not has_summary or not has_text_max:
              compliance = " [NON-COMPLIANT]"
          if has_ctx_max:
              compliance += " [ctxMaxChars SET - BAD]"
      
          interesting = {
              k: v for k, v in p.items() if v and k not in ("query", "type")
          }
          param_str = f"     {interesting}" if interesting else "     (no params)"
          return f"  Q: {q}\n{param_str}{compliance}"
      
      
      def format_firecrawl_call(call):
          """Format a Firecrawl call."""
          url = call["params"].get("url", "")[:80]
          return f"  {call['tool']}: {url}"
      
      
      def format_reddit_call(call):
          """Format a Reddit call."""
          p = call["params"]
          tool = call["tool"]
          if tool == "get_top_posts":
              sub = p.get("subreddit", "?")
              time = p.get("time_filter", "all")
              return f"  {tool}: r/{sub} ({time})"
          elif tool == "get_post_comments":
              post = p.get("post_id", p.get("url", "?"))
              depth = p.get("depth", "?")
              return f"  {tool}: {post} (depth={depth})"
          elif tool in ("get_reddit_post", "get_subreddit_info"):
              target = p.get("post_id", p.get("url", p.get("subreddit", "?")))
              return f"  {tool}: {target}"
          elif tool == "search_reddit":
              q = p.get("query", "")[:60]
              sub = p.get("subreddit", "")
              suffix = f" in r/{sub}" if sub else ""
              return f"  {tool}: \"{q}\"{suffix}"
          else:
              return f"  {tool}: {json.dumps(p, default=str)[:80]}"
      
      
      def format_generic_call(call):
          """Format a generic MCP call."""
          p = call["params"]
          summary = json.dumps(p, default=str)
          if len(summary) > 80:
              summary = summary[:77] + "..."
          return f"  {call['tool']}: {summary}"
      
      
      SERVER_FORMATTERS = {
          "exa": format_exa_call,
          "firecrawl": format_firecrawl_call,
          "reddit": format_reddit_call,
      }
      
      
      def print_report(agents):
          """Print a formatted report of tool usage across all researchers."""
          researchers = [a for a in agents if a["mcp_calls"]]
          if not researchers:
              print("No researcher MCP tool calls found in this session's subagents.")
              return
      
          # Collect all servers seen
          all_servers = set()
          for r in researchers:
              all_servers.update(r["mcp_calls"].keys())
      
          # Per-researcher detail
          for r in researchers:
              letter = r["researcher"] or "?"
              task = r["task"] or "(follow-up or unidentified)"
              print(f"\n{'='*70}")
              print(f"Researcher {letter}: {task}")
              if r["question_type"]:
                  print(f"Question type: {r['question_type']}")
              print(f"{'='*70}")
      
              for server in sorted(r["mcp_calls"].keys()):
                  calls = r["mcp_calls"][server]
                  print(f"\n{server} calls: {len(calls)}")
                  formatter = SERVER_FORMATTERS.get(server, format_generic_call)
                  for call in calls:
                      print(formatter(call))
      
          # Aggregate report
          print(f"\n{'='*70}")
          print("AGGREGATE REPORT")
          print(f"{'='*70}")
      
          # Per-server totals
          server_totals = {}
          server_tools = {}
          for r in researchers:
              for server, calls in r["mcp_calls"].items():
                  server_totals[server] = server_totals.get(server, 0) + len(calls)
                  for c in calls:
                      key = f"{server}.{c['tool']}"
                      server_tools[key] = server_tools.get(key, 0) + 1
      
          print("\nTool calls by server:")
          for server in sorted(server_totals, key=lambda s: -server_totals[s]):
              print(f"  {server}: {server_totals[server]}")
      
          print("\nTool calls by endpoint:")
          for key in sorted(server_tools, key=lambda k: -server_tools[k]):
              print(f"  {key}: {server_tools[key]}")
      
          # Exa compliance section
          exa_calls = [
              c
              for r in researchers
              for c in r["mcp_calls"].get("exa", [])
          ]
          if exa_calls:
              total = len(exa_calls)
              summary_ok = sum(1 for c in exa_calls if c["params"].get("enableSummary") is True)
              text_max_ok = sum(1 for c in exa_calls if c["params"].get("textMaxCharacters") == 1)
              ctx_max_bad = sum(1 for c in exa_calls if "contextMaxCharacters" in c["params"])
              full_ok = sum(
                  1 for c in exa_calls
                  if c["params"].get("enableSummary") is True
                  and c["params"].get("textMaxCharacters") == 1
                  and "contextMaxCharacters" not in c["params"]
              )
              pct = lambda n: f"{100*n//total}%" if total else "0%"
      
              print(f"\nExa compliance ({total} calls):")
              print(f"  enableSummary: true    {summary_ok}/{total} ({pct(summary_ok)})")
              print(f"  textMaxCharacters: 1   {text_max_ok}/{total} ({pct(text_max_ok)})")
              print(f"  contextMaxCharacters:  {ctx_max_bad} violations (should be 0)")
              print(f"  Full compliance:       {full_ok}/{total} ({pct(full_ok)})")
      
              categories = {}
              domains = {}
              for c in exa_calls:
                  cat = c["params"].get("category")
                  if cat:
                      categories[cat] = categories.get(cat, 0) + 1
                  for dom in c["params"].get("includeDomains", []):
                      domains[dom] = domains.get(dom, 0) + 1
      
              if categories:
                  print(f"\n  Categories used:")
                  for cat, count in sorted(categories.items(), key=lambda x: -x[1]):
                      print(f"    {cat}: {count}")
              if domains:
                  print(f"\n  includeDomains used:")
                  for dom, count in sorted(domains.items(), key=lambda x: -x[1]):
                      print(f"    {dom}: {count}")
      
      
      def main():
          parser = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
          parser.add_argument("keyword", nargs="?", help="Keyword to find in subagent spawn messages")
          parser.add_argument("--session", help="Explicit session ID")
          parser.add_argument("--list", action="store_true", help="List recent sessions with subagents")
          parser.add_argument("--project-dir", help="Override project directory path (filesystem or slug)")
          args = parser.parse_args()
      
          project_dir = find_project_dir(args.project_dir)
      
          if not project_dir or not project_dir.exists():
              print("Could not find Claude project directory. Use --project-dir.", file=sys.stderr)
              sys.exit(1)
      
          if args.list:
              sessions = list_sessions(project_dir)
              if not sessions:
                  print("No sessions with subagents found.")
                  return
              print(f"Sessions with subagents in {project_dir.name}:\n")
              for sid, count, _ in sessions:
                  print(f"  {sid}  ({count} subagents)")
              return
      
          if args.session:
              session_dir = project_dir / args.session
              if not (session_dir / "subagents").is_dir():
                  print(f"No subagents directory in session {args.session}", file=sys.stderr)
                  sys.exit(1)
              session_ids = [args.session]
          elif args.keyword:
              session_ids = find_session_by_keyword(project_dir, args.keyword)
              if not session_ids:
                  print(f"No sessions found matching '{args.keyword}'", file=sys.stderr)
                  sys.exit(1)
              if len(session_ids) > 1:
                  print(f"Multiple sessions match '{args.keyword}':")
                  for sid in session_ids:
                      print(f"  {sid}")
                  print("\nUse --session to pick one.")
                  return
          else:
              parser.print_help()
              return
      
          for session_id in session_ids:
              subdir = project_dir / session_id / "subagents"
              print(f"Session: {session_id}")
              print(f"Path: {subdir}\n")
      
              agents = []
              for f in sorted(subdir.glob("agent-*.jsonl")):
                  if "compact" in f.name:
                      continue
                  if f.stat().st_size < 5000:
                      continue  # Skip tiny files (shutdown messages, nudges)
                  result = parse_subagent(f)
                  if result["mcp_calls"]:
                      agents.append(result)
      
              print_report(agents)
      
      
      if __name__ == "__main__":
          main()
      
  • SKILL.md 23.3 KB
    ---
    name: deep-research-team
    description: >
      This skill should be used when the user asks for "deep research", "research team",
      "comprehensive analysis", "research report", "investigate thoroughly", "compare X vs Y in
      depth", or needs synthesis across multiple sources with verification. It spawns a coordinated
      team of researcher agents across multiple rounds, with the lead triaging findings and creating
      targeted follow-up tasks. Scales from Focused (2 researchers, 1-2 rounds) to Comprehensive (4
      researchers, 3-4 rounds with cross-verification). Do NOT use for simple lookups, debugging, or
      questions answerable with 1-2 searches.
    ---
    
    # Deep Research Team (Lead Orchestrator)
    
    Conduct thorough, iterative research by coordinating a persistent team of researcher agents across
    multiple rounds. This architecture enables mid-investigation steering, targeted follow-up based on
    emerging findings, and cross-agent verification.
    
    ## Architecture Overview
    
    ```
    Round 1: Investigation      Round 2: Follow-up          Synthesis
    
    ┌──────────┐                  ┌──────────┐                  ┌────────┐
    │Researcher│  sends findings  │Researcher│  sends findings  │        │
    │    A     ├─────────┬───────>│    A     ├─────────┬───────>│        │
    └──────────┘         │        └──────────┘         │        │        │
                         │                             │        │        │
                         v          dispatches         v        │        │
    ┌──────────┐    ┌────────┐    ┌──────────┐    ┌────────┐    │  Lead  │
    │Researcher├───>│  Lead  │───>│Researcher├───>│  Lead  │───>│  synth │
    │    B     │    │triages │    │    B     │    │triages │    │  esizes│
    └──────────┘    └────────┘    └──────────┘    └────────┘    │        │
                         ^                             ^        │        │
    ┌──────────┐         │        ┌──────────┐         │        │        │
    │Researcher├─────────┴───────>│Researcher├─────────┴───────>│        │
    │    C     │  sends findings  │    C     │  sends findings  │        │
    └──────────┘                  └──────────┘                  └────────┘
    ```
    
    **Key principles:**
    
    1. **No peer-to-peer researcher communication.** All coordination goes through the lead. This
       preserves the independence that accounts for 87% of multi-agent gains (Choi et al.) and avoids
       sycophancy failures (Wynn et al.). Researchers never see each other's findings.
    
    2. **Multi-round iteration.** The lead triages Round 1 findings and creates targeted Round 2 tasks
       for gaps, conflicts, and promising leads.
    
    3. **Cross-agent verification (Comprehensive scope).** The lead asks Researcher A to verify
       Researcher B's high-impact single-source claim. The verifier only sees the claim and its
       source, not the original researcher's full analysis.
    
    4. **Dynamic task evolution.** The shared task list starts with pre-planned angles but grows
       organically as follow-up tasks emerge from findings. The lead dispatches follow-up tasks
       directly to specific researchers via SendMessage.
    
    ## When to Use
    
    **Use this skill for:**
    
    - Complex questions that benefit from multiple research angles
    - Topics where initial findings will reveal what to investigate next
    - Research requiring cross-verification of contested claims
    - Any question needing synthesis across 5+ sources
    
    **Do NOT use for:**
    
    - Simple factual lookups (use regular web search)
    - Questions answerable with 1-2 searches
    - Debugging or code questions
    
    ## Effort Calibration
    
    | Scope             | Researchers | Rounds | Verification | Model  |
    | ----------------- | ----------- | ------ | ------------ | ------ |
    | **Focused**       | 2           | 1-2    | None         | sonnet |
    | **Broad**         | 3           | 2-3    | None         | sonnet |
    | **Comprehensive** | 4           | 3-4    | Cross-agent  | opus   |
    
    Round counts are heuristics, not targets. Stop early when you hit citation convergence --
    additional rounds that don't surface new substantive findings waste tokens and context. A
    Broad run that converges in 2 rounds is a success, not a shortcut. After each round's triage,
    ask: "Would another round change the report's conclusions?" If not, proceed to synthesis.
    
    Default scope is determined by question type (see `references/question-types.md`).
    Present the recommended scope to the user and allow override.
    
    **Model selection:**
    
    - **Lead**: inherits user's session model (no override)
    - **Researchers**: `sonnet` for Focused/Broad, `opus` for Comprehensive
    
    (Sonnet validated as viable override for Comprehensive when cost matters)
    
    ## Output Directory
    
    Research artifacts persist to disk for resumability and backup.
    
    **Directory resolution** -- run this command FIRST, before creating anything. The output is
    your `{output_dir}`. Only the fallback branch creates a directory; the others reuse what exists.
    
    ```bash
    if [ -d "$(pwd)/deep-research" ]; then
      echo "$(pwd)/deep-research"
    elif [ -n "$CLAUDE_DEEP_RESEARCH_DIR" ]; then
      eval echo "$CLAUDE_DEEP_RESEARCH_DIR"
    else
      mkdir -p "$(pwd)/deep-research"
      echo "$(pwd)/deep-research"
    fi
    ```
    
    After resolving `{output_dir}`, create only the topic subdirectory in Phase 3.
    
    Each session creates a subdirectory: `{output_dir}/{topic-slug}/`
    
    **Contents:**
    
    - `state.md` -- triage checkpoint, cross-references, follow-up plan (written in Phase 4)
    - `researcher-{letter}-findings.md` -- backup of each researcher's findings
    - `report.md` -- final synthesized report (written in Phase 6)
    
    ## Calibrated Confidence Language
    
    Use Kent-style verbal probability expressions in all confidence assessments:
    
    | Term           | Range  | Use When                                  |
    | -------------- | ------ | ----------------------------------------- |
    | Almost certain | 93-99% | Multiple high-quality sources, no dissent |
    | Highly likely  | 80-92% | Strong evidence, minor caveats            |
    | Likely         | 63-79% | Good evidence, some gaps                  |
    | Roughly even   | 40-62% | Conflicting evidence, genuinely uncertain |
    | Unlikely       | 20-39% | Limited or weak evidence                  |
    
    Always pair the verbal term with the probability range in the final report.
    
    ## Process
    
    ### Phase 0: Classify Question Type
    
    Silently classify the user's question before any interaction.
    
    1. Read `references/question-types.md` for the full taxonomy.
    2. Assign a primary type: Factual, Scientific/Health, Consumer, Technical, Opinion/Sentiment,
       Contested, or Emerging/Frontier.
    3. For compound questions, decompose into sub-questions and classify each.
    4. Note the default scope from the type-to-scope mapping.
    
    **Resume check:** Before starting, list the subdirectories in `{output_dir}` and scan for any
    that look related to the current question (similar topic, overlapping keywords). If you find
    a plausible match, read its `state.md` and offer to resume: present what was completed, what
    remains, and ask the user whether to resume or start fresh. If resuming, create a new team
    and tasks for only the remaining work.
    
    **Topic slug:** When creating a new session, generate a slug (lowercase, hyphenated, max 40
    chars) for the subdirectory name: `{output_dir}/{slug}/`.
    
    Classification is internal—do not present it to the user.
    
    ### Phase 1: Clarify and Plan
    
    **Step 1: Make sure you understand the question.** Before planning anything, ask yourself:
    do I understand what the user is asking and *why* well enough to design research angles that
    will actually be useful to them? If not, use `AskUserQuestion` to fill the gaps. This isn't
    just about ambiguous wording -- a perfectly clear question can still lack enough context to
    research well ("How does Nix handle dependencies?" means very different research depending on
    whether you're evaluating Nix, debugging an issue, or writing docs). If the question and its
    context are clear, skip this step.
    
    **Step 2: Scope and decompose.** Determine the appropriate scope (Phase 2 has the details)
    and decompose the question into independent research angles. Default angle counts by scope:
    
    - Focused: 2 angles
    - Broad: 3 angles
    - Comprehensive: 4 angles
    
    These are defaults, not caps. If the decomposition reveals one more genuinely independent
    facet than the default, add it (e.g., 3 angles for a Focused run). If the question has
    fewer real facets, use fewer. Beyond ±1 from the default, re-scope rather than stretching --
    the scope was probably wrong. Each angle must be independent and substantial enough to
    warrant a dedicated researcher; "I can think of another angle" isn't sufficient.
    
    For compound questions, map sub-questions to angles. Multiple sub-questions can share an
    angle if closely related; a single sub-question can span multiple angles if it has distinct
    facets.
    
    **Step 3: Confirm if high-investment.** For compound, contested, or Comprehensive-scope
    questions, present the research plan for user approval before spawning researchers:
    
    > Research plan for "{question}":
    >
    > - Type: {type} | Scope: {scope} | {N} researchers, {M} rounds
    > - Angles: {list of planned angles}
    > - [If compound] Sub-question → angle mapping: ...
    >
    > Proceed, or adjust?
    
    For clear, low-scope questions, skip confirmation and proceed.
    
    ### Phase 2: Calibrate Effort
    
    Select the scope tier based on question type defaults from `references/question-types.md`,
    then apply the scope modifiers from that file (de-escalation and escalation signals).
    Also adjust for:
    
    - User's explicit preference (if stated)
    - Structural complexity (compound questions with 3+ sub-types escalate)
    
    Announce the plan: "Starting {scope} team research with {N} researchers."
    
    ### Phase 3: Team Setup
    
    **Step 1: Create the output directory**
    
    ```bash
    mkdir -p {output_dir}/{topic-slug}
    ```
    
    **Step 2: Create the team**
    
    ```
    TeamCreate:
      team_name: "deep-research-{topic-slug}"
      description: "Researching {topic} in {scope} scope"
    ```
    
    **Step 3: Create initial tasks**
    
    One `TaskCreate` per research angle:
    
    ```
    TaskCreate:
      subject: "Investigate {angle title}"
      description: |
        Research angle: {angle description}
        Topic context: {brief topic summary}
        Question type: {type from Phase 0}
        Focus: {what specifically to investigate}
        Return structured findings via SendMessage to the lead.
      activeForm: "Investigating {angle title}"
    ```
    
    **Task brief clarity:** Make scope boundaries explicit between researchers to avoid overlap
    and gaps. Flag name ambiguities (e.g., multiple products sharing a name). Mark optional
    sub-tasks clearly (e.g., "if time permits" vs required).
    
    **Step 4: Spawn researchers**
    
    Launch ALL researchers in a SINGLE message. Each researcher gets:
    
    ```
    Task:
      subagent_type: "general-purpose"
      name: "researcher-{letter}"
      team_name: "deep-research-{topic-slug}"
      model: "sonnet"  (or "opus" for Comprehensive)
      description: "Spawn researcher {letter}"
      prompt: |
        You are a research agent on a team. Your job is to investigate research tasks
        by searching the web, evaluating sources, and reporting structured findings.
    
        FIRST: Read your methodology at: {absolute path to references/researcher-prompt.md}
    
        Question type: {type from Phase 0}
        Output directory: {output_dir}/{topic-slug}
        Your researcher letter: {letter} (use LOWERCASE in filenames: researcher-{lowercase letter})
        Lead name: team-lead (send all findings to this name via SendMessage)
    
        Your assigned task is #{id}: "{subject}"
        Use TaskGet for full details, then begin investigation.
    
        After completing your task, go idle. The lead will message you directly
        when new tasks are available.
    ```
    
    Include the task ID, subject, question type, output directory, researcher letter,
    and lead name directly in each researcher's spawn prompt.
    
    ### Phase 4: Investigation Loop
    
    This is the core research cycle. Each iteration is a **round**: researchers investigate,
    the lead triages, then either dispatches follow-ups (another round) or exits to synthesis.
    
    #### Round structure
    
    **Investigate:** Researchers work on their assigned tasks. Each researcher will:
    
    1. Read the methodology reference file
    2. Load web search and content extraction tools via ToolSearch
    3. Execute the investigation loop (search -> evaluate -> reflect -> decide)
    4. Write findings to `{output_dir}/{topic-slug}/researcher-{letter}-findings.md`
    5. Notify the lead via `SendMessage` with the file path (not the full findings --
       avoids doubling output tokens)
    6. Mark their task as completed via `TaskUpdate`
    7. Go idle and wait for the lead to dispatch follow-up tasks via `SendMessage`
    
    **Monitoring**: The lead reads each researcher's findings file after receiving
    their notification. No polling needed.
    
    **Handling partial results**: If a researcher reports rate limit issues or thin coverage,
    note the gap for triage rather than immediately spawning replacements.
    
    **Triage:** After all tasks for the current round complete, systematically review findings.
    
    *Step 1: Extract and cross-reference claims*
    
    For each significant claim across all researcher findings:
    
    - How many **independent sources** support it? (Different researchers finding the same
      source counts as one source, not two.)
    - **HIGH confidence**: 3+ independent sources of different types (e.g., paper + dataset +
      practitioner account), no credible dissent
    - **MEDIUM confidence**: 2 independent sources, or multiple sources of the same type
    - **LOW confidence**: Single source, or multiple sources that trace back to one original
    
    *Step 2: Identify gaps and conflicts*
    
    - What angles remain uncovered?
    - Where do researchers contradict each other?
    - What findings are surprising and deserve deeper investigation?
    - Which claims rest on a single source?
    
    *Step 3: Decide whether to continue or exit*
    
    Exit to Phase 5 (Synthesis) if findings have converged -- another round wouldn't change the
    report's conclusions. Continue if significant gaps, conflicts, or single-source high-impact
    claims remain and the scope's round budget allows.
    
    If findings reveal more complexity than anticipated (e.g., Broad scope uncovering deeply
    contested claims requiring steelmanning), escalate: spawn an additional researcher or add
    a round beyond the default budget.
    
    *Step 4: Persist triage state*
    
    Write or update `{output_dir}/{topic-slug}/state.md`:
    
    ```markdown
    # Research State: {topic}
    
    ## Status: TRIAGE_COMPLETE (Round {N})
    
    ## Question Type: {type}
    
    ## Scope: {scope}
    
    ## Round {N} Summary
    
    {brief cross-reference of key findings, gaps, conflicts}
    
    ## Follow-up Plan
    
    {list of planned follow-up tasks with rationale, or "Proceeding to synthesis"}
    ```
    
    #### Dispatching follow-up tasks
    
    If continuing, create targeted tasks based on what triage revealed. These are all just
    task types -- they use the same dispatch mechanism:
    
    > **Gap-fill**: "No researcher covered {aspect}. Investigate {specific question}."
    
    > **Conflict-resolution**: "One source says X, another says Y. Search for additional
    > sources that clarify which is accurate and why they might differ."
    
    > **Deep-dive**: "Initial findings revealed {unexpected thing}. Investigate further:
    > {specific follow-up questions}."
    
    > **Verification** (Comprehensive scope): Use when high-impact claims rest on a single
    > source, factual conflicts remain unresolved, or claims are in specialized/niche domains
    > where citation error rates are higher. Assign to a researcher who did NOT make the
    > original claim. Task description contains only the claim and its source URL:
    >
    > ```
    > TaskCreate:
    >   subject: "Verify: {claim summary}"
    >   description: |
    >     Verification task. Search for ADDITIONAL sources (not the original) and determine
    >     if they support, contradict, or add nuance to this claim.
    >
    >     CLAIM: {specific factual claim}
    >     ORIGINAL SOURCE: {URL}
    >
    >     Report your verdict as: SUPPORTED, SUPPORTED WITH NUANCE, CONTESTED, or UNCHANGED
    >     Include the additional sources you found and any important nuance.
    >   activeForm: "Verifying claim about {topic}"
    > ```
    
    **Critical**: Follow-up task descriptions contain just enough context without revealing
    other researchers' full conclusions. This preserves independence.
    
    **Assignment strategy:** The lead assigns follow-up tasks directly via `SendMessage` rather
    than relying on researchers to self-claim. Choose assignees based on task type:
    
    - **Deep-dives and gap-fills**: Assign to the researcher who covered the related angle
      (continuity -- they have context on what was already found).
    - **Conflict resolution and verification**: Assign to a researcher who did NOT cover either
      side (fresh perspective avoids confirmation bias).
    - **If researchers outnumber tasks**: Idle researchers wait or are shut down early.
    
    ```
    SendMessage:
      type: "message"
      recipient: "researcher-{letter}"
      content: "New task available: #{id} -- {subject}. Please claim it and begin."
      summary: "Follow-up task assignment"
    ```
    
    Then loop back to **Investigate** above.
    
    **Interpreting verification results:** When a verification task returns, adjust confidence:
    
    - **SUPPORTED**: Upgrade confidence; note additional sources
    - **SUPPORTED WITH NUANCE**: Directionally correct but specific details differ or require
      qualification. Upgrade confidence for the general claim; add caveats for specifics.
    - **CONTESTED**: Flag explicitly; present both sides with evidence
    - **UNCHANGED**: Keep original confidence level
    
    ### Phase 5: Synthesize (Type-Aware)
    
    Combine all findings from all rounds into a coherent report. Select the synthesis template
    matching the question type from Phase 0. For compound questions, use the template for each
    sub-question's type, then add an overall synthesis section.
    
    **End-of-sequence awareness:** Draft Confidence Assessment and Limitations sections **early**,
    not last. Review final paragraphs specifically for unsourced claims.
    
    #### Template Selection
    
    Read the template file matching the question type from Phase 0. Each template includes the
    full report structure (executive summary, type-specific body, confidence assessment,
    limitations, sources). For compound questions, read the template for each sub-question's
    type and add an overall synthesis section.
    
    | Question Type         | Template File                                |
    | --------------------- | -------------------------------------------- |
    | **Factual**           | `references/templates/factual.md`            |
    | **Scientific/Health** | `references/templates/scientific-health.md`  |
    | **Consumer**          | `references/templates/consumer.md`           |
    | **Technical**         | `references/templates/technical.md`          |
    | **Opinion/Sentiment** | `references/templates/opinion-sentiment.md`  |
    | **Contested**         | `references/templates/contested.md`          |
    | **Emerging/Frontier** | `references/templates/emerging-frontier.md`  |
    
    ### Phase 6: Persist Report
    
    Write the final report:
    
    ```
    Write: {output_dir}/{topic-slug}/report.md
    ```
    
    Update `state.md` status to `COMPLETE`:
    
    ```
    Edit: {output_dir}/{topic-slug}/state.md
      old_string: "## Status: TRIAGE_COMPLETE"
      new_string: "## Status: COMPLETE"
    ```
    
    **Format output files (optional):** If `prettier` is available, run it on all markdown files
    in the output directory to normalize formatting:
    
    ```bash
    prettier --write --prose-wrap preserve "{output_dir}/{topic-slug}/**/*.md"
    ```
    
    If prettier is not installed, skip this step silently -- it is cosmetic, not functional.
    
    Inform the user: "Report saved to `{output_dir}/{topic-slug}/report.md`."
    
    ### Phase 7: Cleanup
    
    Shut down the team cleanly.
    
    **Step 1: Shut down researchers**
    
    Send `shutdown_request` to each researcher via `SendMessage`:
    
    ```
    SendMessage:
      type: "shutdown_request"
      recipient: "researcher-a"
      content: "Research complete. Shutting down."
    ```
    
    Repeat for each researcher. Wait for shutdown responses.
    
    **Step 2: Read researcher feedback**
    
    After all researchers have shut down, read any feedback files written to
    `{output_dir}/{topic-slug}/researcher-{letter}-feedback.md`. These contain notes on
    tool usage (Exa parameters, Firecrawl usage), issues encountered (400 errors, rate limits),
    and suggestions. Use this feedback to identify patterns for skill improvement.
    
    **Step 3: Delete team**
    
    ```
    TeamDelete
    ```
    
    ## Writing Standards
    
    - Prose paragraphs, not bullet lists (bullets only for distinct enumerations)
    - Specific data: "increased 23%" not "increased significantly"
    - Cite inline with markdown footnotes: "The market grew 15%[^1]" not "The market grew.[^1]"
    - Each finding: 2-4 paragraphs with evidence
    - Distinguish FACTS (cited) from ANALYSIS (synthesis)
    
    ## Anti-Hallucination Protocol
    
    - Every factual claim must cite a source immediately
    - Mark synthesis distinctly: "This suggests..." or "Synthesizing these findings..."
    - If uncertain, say so: "Sources disagree on..." or "Limited evidence for..."
    - Never fabricate sources -- all citations come from researcher findings
    
    ## Additional Resources
    
    ### Reference Files
    
    - **`references/question-types.md`** -- 7-type taxonomy, signals, decomposition rules,
      type-to-scope defaults. Read in Phase 0.
    - **`references/researcher-prompt.md`** -- Investigation methodology, type-aware source
      evaluation, output format. Path provided to researchers in spawn prompt. (Tool guidance
      extracted to the standalone `search-tips` skill, which researchers load as their first step.)
    - **`references/templates/`** -- Type-specific synthesis templates. Read the relevant
      template(s) in Phase 5. See Template Selection table above.
    
    ### Scripts
    
    - **`scripts/analyze-transcripts.py`** -- Post-hoc analysis of researcher tool usage.
      Extracts MCP tool call parameters from subagent JSONL transcripts and produces a
      compliance report. Usage:
      - `python3 scripts/analyze-transcripts.py --session ${CLAUDE_SESSION_ID}` -- current session
      - `python3 scripts/analyze-transcripts.py "topic keyword"` -- auto-detect session by keyword
      - `python3 scripts/analyze-transcripts.py --list` -- list recent sessions with subagents
    
    ### Development History
    
    - **`dev/RESEARCH.md`** -- Design rationale with 50+ sources justifying the architecture
    - **`dev/ITERATION-LOG.md`** -- 13 iterations of improvement with backlog
    - **`dev/iterations/`** -- Detailed notes for each iteration
    
    ## Quick Reference
    
    0. **Classify** question type (silent) and check for resume
    1. **Clarify** -- understand the question, resolve ambiguity, confirm plan if high-investment
    2. **Calibrate** effort: announce scope and researcher count
    3. **Setup**: create output dir, TeamCreate, TaskCreate per angle, spawn researchers
    4. **Investigation loop**: investigate → triage → dispatch follow-ups or exit. Repeat until converged.
    5. **Synthesize**: type-aware template, calibrated confidence language
    6. **Persist**: write report.md, update state.md to COMPLETE, tell user file location
    7. **Cleanup**: shutdown_request to each researcher, then TeamDelete
    
    **Context budget:**
    
    - Lead context: reserve for triage + synthesis
    - Researcher contexts: handle all search/scrape operations
    - Researcher spawn prompts: compact (~30 lines), point to reference file
    - Researcher findings: structured summaries (~120 lines each)
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related