deep-research-team
This skill should be used when the user asks for "deep research", "research team", "comprehensive analysis", "research report", "investigate thoroughly", "compare X vs Y in depth", or needs synthesis across multiple sources with verification. It spawns a coordinated team of resea
Install
npx skills add https://github.com/malob/nix-config/tree/master/configs/claude/skills/deep-research-team
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install malob-nix-config@llmmart
git clone https://github.com/malob/nix-config.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole malob/nix-config collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Deep Research Team (Lead Orchestrator)
Conduct thorough, iterative research by coordinating a persistent team of researcher agents across multiple rounds. This architecture enables mid-investigation steering, targeted follow-up based on emerging findings, and cross-agent verification.
Architecture Overview
Round 1: Investigation Round 2: Follow-up Synthesis
┌──────────┐ ┌──────────┐ ┌────────┐
│Researcher│ sends findings │Researcher│ sends findings │ │
│ A ├─────────┬───────>│ A ├─────────┬───────>│ │
└──────────┘ │ └──────────┘ │ │ │
│ │ │ │
v dispatches v │ │
┌──────────┐ ┌────────┐ ┌──────────┐ ┌────────┐ │ Lead │
│Researcher├───>│ Lead │───>│Researcher├───>│ Lead │───>│ synth │
│ B │ │triages │ │ B │ │triages │ │ esizes│
└──────────┘ └────────┘ └──────────┘ └────────┘ │ │
^ ^ │ │
┌──────────┐ │ ┌──────────┐ │ │ │
│Researcher├─────────┴───────>│Researcher├─────────┴───────>│ │
│ C │ sends findings │ C │ sends findings │ │
└──────────┘ └──────────┘ └────────┘
Key principles:
No peer-to-peer researcher communication. All coordination goes through the lead. This preserves the independence that accounts for 87% of multi-agent gains (Choi et al.) and avoids sycophancy failures (Wynn et al.). Researchers never see each other's findings.
Multi-round iteration. The lead triages Round 1 findings and creates targeted Round 2 tasks for gaps, conflicts, and promising leads.
Cross-agent verification (Comprehensive scope). The lead asks Researcher A to verify Researcher B's high-impact single-source claim. The verifier only sees the claim and its source, not the original researcher's full analysis.
Dynamic task evolution. The shared task list starts with pre-planned angles but grows organically as follow-up tasks emerge from findings. The lead dispatches follow-up tasks directly to specific researchers via SendMessage.
When to Use
Use this skill for:
- Complex questions that benefit from multiple research angles
- Topics where initial findings will reveal what to investigate next
- Research requiring cross-verification of contested claims
- Any question needing synthesis across 5+ sources
Do NOT use for:
- Simple factual lookups (use regular web search)
- Questions answerable with 1-2 searches
- Debugging or code questions
Effort Calibration
| Scope | Researchers | Rounds | Verification | Model |
|---|---|---|---|---|
| Focused | 2 | 1-2 | None | sonnet |
| Broad | 3 | 2-3 | None | sonnet |
| Comprehensive | 4 | 3-4 | Cross-agent | opus |
Round counts are heuristics, not targets. Stop early when you hit citation convergence -- additional rounds that don't surface new substantive findings waste tokens and context. A Broad run that converges in 2 rounds is a success, not a shortcut. After each round's triage, ask: "Would another round change the report's conclusions?" If not, proceed to synthesis.
Default scope is determined by question type (see references/question-types.md).
Present the recommended scope to the user and allow override.
Model selection:
- Lead: inherits user's session model (no override)
- Researchers:
sonnetfor Focused/Broad,opusfor Comprehensive
(Sonnet validated as viable override for Comprehensive when cost matters)
Output Directory
Research artifacts persist to disk for resumability and backup.
Directory resolution -- run this command FIRST, before creating anything. The output is
your {output_dir}. Only the fallback branch creates a directory; the others reuse what exists.
if [ -d "$(pwd)/deep-research" ]; then
echo "$(pwd)/deep-research"
elif [ -n "$CLAUDE_DEEP_RESEARCH_DIR" ]; then
eval echo "$CLAUDE_DEEP_RESEARCH_DIR"
else
mkdir -p "$(pwd)/deep-research"
echo "$(pwd)/deep-research"
fi
After resolving {output_dir}, create only the topic subdirectory in Phase 3.
Each session creates a subdirectory: {output_dir}/{topic-slug}/
Contents:
state.md-- triage checkpoint, cross-references, follow-up plan (written in Phase 4)researcher-{letter}-findings.md-- backup of each researcher's findingsreport.md-- final synthesized report (written in Phase 6)
Calibrated Confidence Language
Use Kent-style verbal probability expressions in all confidence assessments:
| Term | Range | Use When |
|---|---|---|
| Almost certain | 93-99% | Multiple high-quality sources, no dissent |
| Highly likely | 80-92% | Strong evidence, minor caveats |
| Likely | 63-79% | Good evidence, some gaps |
| Roughly even | 40-62% | Conflicting evidence, genuinely uncertain |
| Unlikely | 20-39% | Limited or weak evidence |
Always pair the verbal term with the probability range in the final report.
Process
Phase 0: Classify Question Type
Silently classify the user's question before any interaction.
- Read
references/question-types.mdfor the full taxonomy. - Assign a primary type: Factual, Scientific/Health, Consumer, Technical, Opinion/Sentiment, Contested, or Emerging/Frontier.
- For compound questions, decompose into sub-questions and classify each.
- Note the default scope from the type-to-scope mapping.
Resume check: Before starting, list the subdirectories in {output_dir} and scan for any
that look related to the current question (similar topic, overlapping keywords). If you find
a plausible match, read its state.md and offer to resume: present what was completed, what
remains, and ask the user whether to resume or start fresh. If resuming, create a new team
and tasks for only the remaining work.
Topic slug: When creating a new session, generate a slug (lowercase, hyphenated, max 40
chars) for the subdirectory name: {output_dir}/{slug}/.
Classification is internal—do not present it to the user.
Phase 1: Clarify and Plan
Step 1: Make sure you understand the question. Before planning anything, ask yourself:
do I understand what the user is asking and why well enough to design research angles that
will actually be useful to them? If not, use AskUserQuestion to fill the gaps. This isn't
just about ambiguous wording -- a perfectly clear question can still lack enough context to
research well ("How does Nix handle dependencies?" means very different research depending on
whether you're evaluating Nix, debugging an issue, or writing docs). If the question and its
context are clear, skip this step.
Step 2: Scope and decompose. Determine the appropriate scope (Phase 2 has the details) and decompose the question into independent research angles. Default angle counts by scope:
- Focused: 2 angles
- Broad: 3 angles
- Comprehensive: 4 angles
These are defaults, not caps. If the decomposition reveals one more genuinely independent facet than the default, add it (e.g., 3 angles for a Focused run). If the question has fewer real facets, use fewer. Beyond ±1 from the default, re-scope rather than stretching -- the scope was probably wrong. Each angle must be independent and substantial enough to warrant a dedicated researcher; "I can think of another angle" isn't sufficient.
For compound questions, map sub-questions to angles. Multiple sub-questions can share an angle if closely related; a single sub-question can span multiple angles if it has distinct facets.
Step 3: Confirm if high-investment. For compound, contested, or Comprehensive-scope questions, present the research plan for user approval before spawning researchers:
Research plan for "":
- Type: | Scope: | researchers, rounds
- Angles:
- [If compound] Sub-question → angle mapping: ...
Proceed, or adjust?
For clear, low-scope questions, skip confirmation and proceed.
Phase 2: Calibrate Effort
Select the scope tier based on question type defaults from references/question-types.md,
then apply the scope modifiers from that file (de-escalation and escalation signals).
Also adjust for:
- User's explicit preference (if stated)
- Structural complexity (compound questions with 3+ sub-types escalate)
Announce the plan: "Starting team research with researchers."
Phase 3: Team Setup
Step 1: Create the output directory
mkdir -p {output_dir}/{topic-slug}
Step 2: Create the team
TeamCreate:
team_name: "deep-research-{topic-slug}"
description: "Researching {topic} in {scope} scope"
Step 3: Create initial tasks
One TaskCreate per research angle:
TaskCreate:
subject: "Investigate {angle title}"
description: |
Research angle: {angle description}
Topic context: {brief topic summary}
Question type: {type from Phase 0}
Focus: {what specifically to investigate}
Return structured findings via SendMessage to the lead.
activeForm: "Investigating {angle title}"
Task brief clarity: Make scope boundaries explicit between researchers to avoid overlap and gaps. Flag name ambiguities (e.g., multiple products sharing a name). Mark optional sub-tasks clearly (e.g., "if time permits" vs required).
Step 4: Spawn researchers
Launch ALL researchers in a SINGLE message. Each researcher gets:
Task:
subagent_type: "general-purpose"
name: "researcher-{letter}"
team_name: "deep-research-{topic-slug}"
model: "sonnet" (or "opus" for Comprehensive)
description: "Spawn researcher {letter}"
prompt: |
You are a research agent on a team. Your job is to investigate research tasks
by searching the web, evaluating sources, and reporting structured findings.
FIRST: Read your methodology at: {absolute path to references/researcher-prompt.md}
Question type: {type from Phase 0}
Output directory: {output_dir}/{topic-slug}
Your researcher letter: {letter} (use LOWERCASE in filenames: researcher-{lowercase letter})
Lead name: team-lead (send all findings to this name via SendMessage)
Your assigned task is #{id}: "{subject}"
Use TaskGet for full details, then begin investigation.
After completing your task, go idle. The lead will message you directly
when new tasks are available.
Include the task ID, subject, question type, output directory, researcher letter, and lead name directly in each researcher's spawn prompt.
Phase 4: Investigation Loop
This is the core research cycle. Each iteration is a round: researchers investigate, the lead triages, then either dispatches follow-ups (another round) or exits to synthesis.
Round structure
Investigate: Researchers work on their assigned tasks. Each researcher will:
- Read the methodology reference file
- Load web search and content extraction tools via ToolSearch
- Execute the investigation loop (search -> evaluate -> reflect -> decide)
- Write findings to
{output_dir}/{topic-slug}/researcher-{letter}-findings.md - Notify the lead via
SendMessagewith the file path (not the full findings -- avoids doubling output tokens) - Mark their task as completed via
TaskUpdate - Go idle and wait for the lead to dispatch follow-up tasks via
SendMessage
Monitoring: The lead reads each researcher's findings file after receiving their notification. No polling needed.
Handling partial results: If a researcher reports rate limit issues or thin coverage, note the gap for triage rather than immediately spawning replacements.
Triage: After all tasks for the current round complete, systematically review findings.
Step 1: Extract and cross-reference claims
For each significant claim across all researcher findings:
- How many independent sources support it? (Different researchers finding the same source counts as one source, not two.)
- HIGH confidence: 3+ independent sources of different types (e.g., paper + dataset + practitioner account), no credible dissent
- MEDIUM confidence: 2 independent sources, or multiple sources of the same type
- LOW confidence: Single source, or multiple sources that trace back to one original
Step 2: Identify gaps and conflicts
- What angles remain uncovered?
- Where do researchers contradict each other?
- What findings are surprising and deserve deeper investigation?
- Which claims rest on a single source?
Step 3: Decide whether to continue or exit
Exit to Phase 5 (Synthesis) if findings have converged -- another round wouldn't change the report's conclusions. Continue if significant gaps, conflicts, or single-source high-impact claims remain and the scope's round budget allows.
If findings reveal more complexity than anticipated (e.g., Broad scope uncovering deeply contested claims requiring steelmanning), escalate: spawn an additional researcher or add a round beyond the default budget.
Step 4: Persist triage state
Write or update {output_dir}/{topic-slug}/state.md:
# Research State: {topic}
## Status: TRIAGE_COMPLETE (Round {N})
## Question Type: {type}
## Scope: {scope}
## Round {N} Summary
{brief cross-reference of key findings, gaps, conflicts}
## Follow-up Plan
{list of planned follow-up tasks with rationale, or "Proceeding to synthesis"}
Dispatching follow-up tasks
If continuing, create targeted tasks based on what triage revealed. These are all just task types -- they use the same dispatch mechanism:
Gap-fill: "No researcher covered . Investigate ."
Conflict-resolution: "One source says X, another says Y. Search for additional sources that clarify which is accurate and why they might differ."
Deep-dive: "Initial findings revealed . Investigate further: ."
Verification (Comprehensive scope): Use when high-impact claims rest on a single source, factual conflicts remain unresolved, or claims are in specialized/niche domains where citation error rates are higher. Assign to a researcher who did NOT make the original claim. Task description contains only the claim and its source URL:
TaskCreate: subject: "Verify: {claim summary}" description: | Verification task. Search for ADDITIONAL sources (not the original) and determine if they support, contradict, or add nuance to this claim. CLAIM: {specific factual claim} ORIGINAL SOURCE: {URL} Report your verdict as: SUPPORTED, SUPPORTED WITH NUANCE, CONTESTED, or UNCHANGED Include the additional sources you found and any important nuance. activeForm: "Verifying claim about {topic}"
Critical: Follow-up task descriptions contain just enough context without revealing other researchers' full conclusions. This preserves independence.
Assignment strategy: The lead assigns follow-up tasks directly via SendMessage rather
than relying on researchers to self-claim. Choose assignees based on task type:
- Deep-dives and gap-fills: Assign to the researcher who covered the related angle (continuity -- they have context on what was already found).
- Conflict resolution and verification: Assign to a researcher who did NOT cover either side (fresh perspective avoids confirmation bias).
- If researchers outnumber tasks: Idle researchers wait or are shut down early.
SendMessage:
type: "message"
recipient: "researcher-{letter}"
content: "New task available: #{id} -- {subject}. Please claim it and begin."
summary: "Follow-up task assignment"
Then loop back to Investigate above.
Interpreting verification results: When a verification task returns, adjust confidence:
- SUPPORTED: Upgrade confidence; note additional sources
- SUPPORTED WITH NUANCE: Directionally correct but specific details differ or require qualification. Upgrade confidence for the general claim; add caveats for specifics.
- CONTESTED: Flag explicitly; present both sides with evidence
- UNCHANGED: Keep original confidence level
Phase 5: Synthesize (Type-Aware)
Combine all findings from all rounds into a coherent report. Select the synthesis template matching the question type from Phase 0. For compound questions, use the template for each sub-question's type, then add an overall synthesis section.
End-of-sequence awareness: Draft Confidence Assessment and Limitations sections early, not last. Review final paragraphs specifically for unsourced claims.
Template Selection
Read the template file matching the question type from Phase 0. Each template includes the full report structure (executive summary, type-specific body, confidence assessment, limitations, sources). For compound questions, read the template for each sub-question's type and add an overall synthesis section.
| Question Type | Template File |
|---|---|
| Factual | references/templates/factual.md |
| Scientific/Health | references/templates/scientific-health.md |
| Consumer | references/templates/consumer.md |
| Technical | references/templates/technical.md |
| Opinion/Sentiment | references/templates/opinion-sentiment.md |
| Contested | references/templates/contested.md |
| Emerging/Frontier | references/templates/emerging-frontier.md |
Phase 6: Persist Report
Write the final report:
Write: {output_dir}/{topic-slug}/report.md
Update state.md status to COMPLETE:
Edit: {output_dir}/{topic-slug}/state.md
old_string: "## Status: TRIAGE_COMPLETE"
new_string: "## Status: COMPLETE"
Format output files (optional): If prettier is available, run it on all markdown files
in the output directory to normalize formatting:
prettier --write --prose-wrap preserve "{output_dir}/{topic-slug}/**/*.md"
If prettier is not installed, skip this step silently -- it is cosmetic, not functional.
Inform the user: "Report saved to {output_dir}/{topic-slug}/report.md."
Phase 7: Cleanup
Shut down the team cleanly.
Step 1: Shut down researchers
Send shutdown_request to each researcher via SendMessage:
SendMessage:
type: "shutdown_request"
recipient: "researcher-a"
content: "Research complete. Shutting down."
Repeat for each researcher. Wait for shutdown responses.
Step 2: Read researcher feedback
After all researchers have shut down, read any feedback files written to
{output_dir}/{topic-slug}/researcher-{letter}-feedback.md. These contain notes on
tool usage (Exa parameters, Firecrawl usage), issues encountered (400 errors, rate limits),
and suggestions. Use this feedback to identify patterns for skill improvement.
Step 3: Delete team
TeamDelete
Writing Standards
- Prose paragraphs, not bullet lists (bullets only for distinct enumerations)
- Specific data: "increased 23%" not "increased significantly"
- Cite inline with markdown footnotes: "The market grew 15%[^1]" not "The market grew.[^1]"
- Each finding: 2-4 paragraphs with evidence
- Distinguish FACTS (cited) from ANALYSIS (synthesis)
Anti-Hallucination Protocol
- Every factual claim must cite a source immediately
- Mark synthesis distinctly: "This suggests..." or "Synthesizing these findings..."
- If uncertain, say so: "Sources disagree on..." or "Limited evidence for..."
- Never fabricate sources -- all citations come from researcher findings
Additional Resources
Reference Files
references/question-types.md-- 7-type taxonomy, signals, decomposition rules, type-to-scope defaults. Read in Phase 0.references/researcher-prompt.md-- Investigation methodology, type-aware source evaluation, output format. Path provided to researchers in spawn prompt. (Tool guidance extracted to the standalonesearch-tipsskill, which researchers load as their first step.)references/templates/-- Type-specific synthesis templates. Read the relevant template(s) in Phase 5. See Template Selection table above.
Scripts
scripts/analyze-transcripts.py-- Post-hoc analysis of researcher tool usage. Extracts MCP tool call parameters from subagent JSONL transcripts and produces a compliance report. Usage:python3 scripts/analyze-transcripts.py --session ${CLAUDE_SESSION_ID}-- current sessionpython3 scripts/analyze-transcripts.py "topic keyword"-- auto-detect session by keywordpython3 scripts/analyze-transcripts.py --list-- list recent sessions with subagents
Development History
dev/RESEARCH.md-- Design rationale with 50+ sources justifying the architecturedev/ITERATION-LOG.md-- 13 iterations of improvement with backlogdev/iterations/-- Detailed notes for each iteration
Quick Reference
- Classify question type (silent) and check for resume
- Clarify -- understand the question, resolve ambiguity, confirm plan if high-investment
- Calibrate effort: announce scope and researcher count
- Setup: create output dir, TeamCreate, TaskCreate per angle, spawn researchers
- Investigation loop: investigate → triage → dispatch follow-ups or exit. Repeat until converged.
- Synthesize: type-aware template, calibrated confidence language
- Persist: write report.md, update state.md to COMPLETE, tell user file location
- Cleanup: shutdown_request to each researcher, then TeamDelete
Context budget:
- Lead context: reserve for triage + synthesis
- Researcher contexts: handle all search/scrape operations
- Researcher spawn prompts: compact (~30 lines), point to reference file
- Researcher findings: structured summaries (~120 lines each)
Files (nix-config)
-
dev
-
iterations
-
iteration-01.md 7.8 KB
# Iteration 1 -- Compound Technical+Consumer (Empirical) ## Plan **Type:** Empirical test (compound question) **Question:** "What are the leading open-source LLM inference engines, and which is best suited for a small team deploying a 70B parameter model on consumer GPUs?" **What this tests:** - Phase 0: Question type classification (should detect Technical + Consumer compound) - Phase 1: Compound decomposition and presentation to user - Phase 2: Scope calibration (compound should escalate to Broad) - Phase 3: Team setup, output directory creation, task creation - Phase 4-6: Multi-round investigation with follow-up - Phase 8: Mixed synthesis (Technical comparison + Consumer recommendation) - Phase 9: File persistence (state.md, researcher findings, report.md) **Rationale:** First-ever empirical run of the v2 skill. A compound question exercises the most distinctive new features. The topic is genuinely useful and has enough complexity to stress-test the system without being so niche that sources are sparse. ## Research Summary Ran Broad-scope research (3 researchers, 2 rounds) on "What are the leading open-source LLM inference engines, and which is best suited for a small team deploying a 70B parameter model on consumer GPUs?" Full report at `~/.claude/research/llm-inference-engines-70b-consumer-gpu/report.md`. **Round 1:** 3 researchers investigated landscape (A), benchmarks (B), practical deployment (C). Strong convergence on vLLM for multi-GPU, llama.cpp/Ollama for single-GPU. Good source diversity (25+ sources across researchers). **Round 2:** 2 targeted follow-ups -- SGLang vs vLLM comparison (B) and ExLlamaV3 viability (A). Resolved conflicting throughput claims and assessed ExLlamaV3 maturity. **Outcome:** Comprehensive report with Technical comparison matrix + Consumer ranked recommendations. 25 cited sources. Report quality is genuinely useful. ## Observations **What worked well:** 1. **Phase 0 classification worked correctly.** Identified compound (Technical + Consumer), chose Broad scope. No friction. 2. **Researcher independence produced good diversity.** Each researcher found unique sources with minimal overlap. Cross-referencing in triage revealed genuine convergence (not echo chamber). 3. **Round 2 follow-ups were well-targeted.** The SGLang vs vLLM task resolved a real conflict from Round 1 (3.1x claim traced to non-independent benchmarks). ExLlamaV3 assessment filled a genuine gap. 4. **Source quality was generally high.** Researchers followed type-aware source evaluation -- prioritized official docs, controlled benchmarks, and practitioner reports over marketing. 5. **File persistence worked.** state.md, researcher findings, and report.md all written correctly. **What was clunky or problematic:** 1. **Researchers didn't self-claim follow-up tasks.** The skill says researchers should check TaskList after completing their initial task and self-claim available work. In practice, all 3 researchers went idle after Round 1 and had to be explicitly nudged with SendMessage to pick up Round 2 tasks. Researcher B went idle even after being sent a direct message about Task #7 and needed a second nudge. **This is the biggest friction point.** 2. **Compound decomposition wasn't presented in the exact template format.** The skill says to present "> This breaks down into {N} sub-questions: ..." but I presented it slightly differently. Minor -- the user understood fine. 3. **No researcher findings files from Round 1.** The skill says researchers should write to `{output_dir}/researcher-{letter}-findings.md`. Only Researcher B wrote a findings file for Round 2 (SGLang). Researchers A, B, C did NOT persist their Round 1 findings to disk. The findings were only delivered via SendMessage. **This breaks the resume capability** -- if the session crashed after Round 1, only state.md would exist, not the underlying findings. 4. **Researcher C was underutilized in Round 2.** With only 2 follow-up tasks and 3 researchers, Researcher C sat idle. The skill doesn't give clear guidance on whether to shut down unused researchers early or find work for them. 5. **Team task list vs iteration loop task list collision.** The research team uses TaskCreate for research angles, but the iteration loop also uses TaskCreate for tracking iteration progress. These two task lists coexisted in the same session (team tasks #1-8, iteration tasks #1-5 in a different context). This worked but could be confusing in a resumed session. 6. **The skill doesn't specify the lead's name for researchers to message.** Researchers need to read the team config to find the lead's name. The researcher-prompt.md says "The lead's name is typically the team creator. Check the team config if unsure." This is indirect -- could be explicit in the spawn prompt. 7. **Synthesis was a single monolithic write.** At ~250 lines, the report was written in one Write call. For longer reports, this could fail or be hard to review. The skill doesn't address progressive synthesis or draft-then-refine. **Surprising or noteworthy:** - Researchers found genuinely useful, current sources (Jan-Feb 2026 content). Exa semantic search worked well for this technical topic. - The compound question naturally produced a well-structured report (Part 1: Technical Comparison, Part 2: Consumer Recommendation). The type-aware templates mapped cleanly. - Total research time was reasonable for a Broad scope run (~10-15 minutes wall clock). ## Changes Made **1. Restructured researcher reporting flow** (`references/researcher-prompt.md`) - Replaced the old "Reporting Findings" + "Task Management" sections with a clear 3-step sequence: (1) persist to disk, (2) SendMessage to lead, (3) mark complete + TaskList. - Moved file persistence from buried rule #9 to Step 1 of the main reporting flow. - Added explicit filenames for follow-up tasks: `researcher-{letter}-findings-{task-id}.md`. **2. Added emphatic task self-claim section** (`references/researcher-prompt.md`) - New "Task Self-Claim (IMPORTANT)" section with explicit instructions to always check TaskList after completing each task and immediately claim available work. - Added "Do NOT go idle if there are available tasks" instruction. **3. Updated Critical Rules** (`references/researcher-prompt.md`) - Replaced rules #6-9 with two focused rules: "Persist, THEN report, THEN check for work" and "ALWAYS self-claim available tasks". Removed redundant rule about SendMessage (now covered in main flow). **4. Added lead name to spawn prompt template** (`SKILL.md`) - Added `Lead name: team-lead` line to the researcher spawn prompt so researchers don't need to look up the team config to know who to message. - Added `IMPORTANT: After completing each task, ALWAYS check TaskList...` reminder at the end of the spawn prompt for reinforcement. **5. Added settings.json permission** (`configs/claude/settings.json`) - Added `Bash(mkdir -p ~/.claude/research/*)` to avoid permission prompts during research output directory creation. ## Notes for Next Iteration - The self-claim fix is the most important change but also the hardest to verify -- it depends on researcher agent behavior, which may not fully follow instructions regardless of emphasis. Iteration 2 should test whether the changes actually reduce the nudging problem. - Compound synthesis quality was good -- the two-part structure (Technical + Consumer) mapped naturally. No changes needed to synthesis templates yet. - The Contested question test (backlog #2) would exercise completely different skill features (steelmanning, cross-agent verification, Comprehensive scope). Good candidate for Iteration 2. - New backlog idea: investigate whether spawn prompt length affects researcher behavior. The prompt is already ~15 lines; adding more "IMPORTANT" reminders has diminishing returns. -
iteration-02.md 10.3 KB
# Iteration 2 -- Contested Question (Empirical) ## Plan **Type:** Empirical test (Contested question) **Question:** "Should AI model weights be open-sourced?" **What this tests:** - Phase 0: Contested classification (should trigger Comprehensive scope) - Phase 1: Research plan presentation (Contested template with angles) - Phase 2: Comprehensive calibration (4 researchers, 2-3 rounds, opus model) - Phase 7: Cross-agent verification (unique to Comprehensive) - Phase 8: Contested synthesis template (steelmanning, no verdict) - **Iteration 1 fixes:** researcher self-claim, file persistence, lead name in spawn prompt **Also validates:** - Context compaction resilience (we're at ~60% context, likely to compact mid-run) - Whether emphasizing self-claim in researcher-prompt.md reduces nudging ## Research Summary Ran Comprehensive-scope research (4 researchers, 3 rounds) on "Should AI model weights be open-sourced?" Full report at `~/.claude/research/ai-model-weights-open-source/report.md`. **Round 1:** 4 opus researchers investigated: (A) pro-openness arguments, (B) anti-openness arguments, (C) regulatory/governance landscape, (D) empirical evidence. Strong convergence on key structural points (irreversibility, safeguard removability, safety research enablement). 60+ sources across researchers. **Round 2:** 2 targeted follow-ups -- graduated release proposals (A) and RAND CBRN contradiction (D). Resolved the RAND 2024 vs 2025 discrepancy (different models, different methodologies, both partially right). Found rich graduated release literature (Solaiman gradient, structured access, Carnegie consensus, MOF, RSPs). **Round 3 (Verification):** 2 cross-agent verifications -- - SentinelOne 7.5% claim (verified by B): SUPPORTED WITH NUANCE. Lower bound on a subset, corroborated by 2 independent studies (Censys solo, Cisco Talos). - Greenblatt ~100K fatalities estimate (verified by C): CONTESTED. Independent models produce ~10K-60K range, but direction agreed and uncertainty intervals overlap. **Outcome:** 37-footnote Contested synthesis report with steelmanned positions, points of agreement, key disagreements (factual and values-based), evidence quality assessment. No verdict rendered. Report quality is genuinely useful for someone trying to understand the debate. ## Observations **What worked well:** 1. **Phase 0 classification worked correctly.** Identified Contested type, selected Comprehensive scope (4 researchers, opus model, 2-3 rounds + verification). No friction. 2. **File persistence fix from Iteration 1 worked perfectly.** All 4 researchers wrote findings to disk in Round 1: `researcher-a-findings.md`, `researcher-b-findings.md`, `researcher-C-findings.md`, `researcher-D-findings.md`. Round 2 follow-ups also persisted: `researcher-a-findings-9.md`, `researcher-D-findings-10.md`. Verifications persisted: `researcher-b-findings-12.md`, `researcher-C-findings-11.md`. **This is a clear win from Iteration 1 changes.** 3. **Cross-agent verification added genuine value.** The Greenblatt CBRN verification (CONTESTED) surfaced 4 independent quantitative models (GovAI/Righetti, GovAI/Williams, Jones, RAND Delphi) that put the estimate in context. The SentinelOne verification (SUPPORTED WITH NUANCE) traced the claim to its primary source and found Reuters had oversimplified. Both verdicts materially improved the final report. 4. **Researcher quality was excellent at opus level.** Source diversity, depth of analysis, and critical evaluation were noticeably stronger than the sonnet researchers in Iteration 1. Researchers proactively labeled institutional biases, distinguished methodology types, and cross-referenced within their own findings. 5. **Contested synthesis template worked well.** The steelmanned positions format naturally organized the findings. Separating factual disagreements from values disagreements was particularly effective. The "Evidence Quality by Position" section forced honest assessment of both sides' weaknesses. 6. **Lead name in spawn prompt worked.** Researchers messaged "team-lead" directly without needing to look up the team config. No friction observed. 7. **Context compaction survived.** The session compacted mid-run (between Round 1 triage and Round 2 dispatch). The ITERATION-LOG.md and state.md files were sufficient to resume coherently. The compacted context summary preserved key details about researcher findings and verification assignments. **What was clunky or problematic:** 1. **Self-claim fix still did not work reliably.** Despite the emphatic "Task Self-Claim (IMPORTANT)" section and the spawn prompt reminder, researchers still went idle after completing their Round 1 tasks without checking TaskList. The lead had to explicitly create follow-up tasks and (in the pre-compaction portion) likely nudge researchers to claim them. After compaction, the lead dispatched Round 2 and Round 3 tasks by sending direct messages to specific researchers rather than relying on self-claim. **This is now a confirmed pattern across 2 iterations: researchers do not reliably self-claim tasks regardless of how emphatically the instructions say to do so.** 2. **Inconsistent filename casing.** Researchers wrote `researcher-a-findings.md` (lowercase) and `researcher-C-findings.md` (uppercase). The researcher-prompt.md says `researcher-{letter}-findings.md` without specifying case. The spawn prompt uses uppercase letters (A, B, C, D) while the template uses lowercase `{letter}`. This inconsistency is cosmetic but could cause issues for scripts or resume logic that expect consistent naming. 3. **Researcher C required double shutdown request.** The first shutdown_request was sent but C didn't respond until a second was sent. This may be a timing issue (C was still processing) rather than a systematic problem. Researchers A, B, D all shut down on first request. 4. **Round 2 researcher allocation was suboptimal.** With 2 follow-up tasks and 4 researchers, B and C sat idle during Round 2. Then in Round 3, B and C were assigned verification tasks. The skill doesn't provide guidance on whether to assign Round 2 tasks to the SAME researchers who covered related Round 1 angles (continuity) or DIFFERENT researchers (fresh perspective). For follow-ups, continuity seems better; for verification, fresh perspective is required. 5. **No guidance on when to stop adding rounds.** The skill says "2-3 rounds" for Comprehensive but doesn't specify criteria for choosing 2 vs 3. In practice, verification (Round 3) added clear value for the two high-impact single-source claims. But the decision was ad hoc. **Surprising or noteworthy:** - Opus researchers found genuinely high-quality, current sources including January-February 2026 content. The contested topic prompted researchers to find primary sources rather than relying on secondary reporting. - The RAND CBRN contradiction investigation was the most valuable follow-up: it found a THIRD RAND study (Delphi panel) and a FOURTH (benchmarking study) that helped reconcile the apparent contradiction. This would have been missed without targeted follow-up. - Report length (~350 lines, 37 footnotes) is longer than ideal. The Contested template naturally produces longer output because steelmanning both sides requires more space. The skill doesn't address length targets for different question types. ## Changes Made **1. Replaced self-claim with lead-dispatched follow-ups** (`SKILL.md` Phase 6, `researcher-prompt.md`) - The biggest architectural change. After 2 iterations confirming researchers don't self-claim regardless of instruction emphasis, switched to a model where the lead creates tasks AND explicitly messages each researcher to assign them. - SKILL.md Phase 6: Added "Assignment strategy" section with guidance on who gets follow-ups (continuity for deep-dives, fresh perspective for conflicts). - SKILL.md Phase 7: Added same lead-dispatched pattern for verification tasks. - researcher-prompt.md: Replaced emphatic "Task Self-Claim (IMPORTANT)" section with brief "Between Tasks" section. Now says: check TaskList once, claim if available, otherwise go idle and wait for lead to message you. - researcher-prompt.md: Updated Critical Rules to remove self-claim rule. **2. Normalized filename casing to lowercase** (`SKILL.md`, `researcher-prompt.md`) - Spawn prompt now says `(use LOWERCASE in filenames: researcher-{lowercase letter})` - researcher-prompt.md Step 1: Added "(always use **lowercase** for the letter)" - researcher-prompt.md Critical Rules: Added rule #7 "Use lowercase filenames" **3. Added round stopping criteria** (`SKILL.md` Phase 7) - Added explicit "When to run Round 3" and "When to skip Round 3" criteria to Phase 7. - Run verification when: high-impact single-source claims, unresolved factual conflicts, or specialized/niche topic claims. - Skip verification when: no single-source high-impact claims remain, conflicts resolved by Round 2, remaining uncertainties are about magnitude not direction. **4. Added Round 2 assignment guidance** (`SKILL.md` Phase 6) - Deep-dives/gap-fills: assign to researcher with related Round 1 angle (continuity) - Conflict resolution: assign to researcher who did NOT cover either side (fresh perspective) - Idle researchers: wait for verification or shut down early ## Notes for Next Iteration - The lead-dispatched model is the most important change to validate in Iteration 3. It accepts the reality that researchers don't self-claim and makes the lead's role explicit. This should eliminate the "nudging" friction entirely. - Iteration 2 demonstrated that Comprehensive scope with opus researchers produces excellent quality. The main question is whether the cost (4 opus researchers x 3 rounds) is justified for less complex questions. - The Contested synthesis template is naturally verbose (~350 lines, 37 footnotes). Consider adding length targets per question type in a future iteration. - Good candidate for Iteration 3: Emerging/Frontier test (backlog #3) would exercise instability warnings, date-sensitive source evaluation, and the Emerging synthesis template. It would also test the lead-dispatched model with a different question type. - Backlog item #8 (spawn prompt length) can be removed -- the self-claim problem was resolved by changing the architecture rather than tweaking prompt emphasis. -
iteration-03.md 8.8 KB
# Iteration 3 -- Consumer Comparison (Empirical) ## Plan **Type:** Empirical test (Consumer question, non-AI domain) **Question:** "Best mechanical keyboard for programming in 2026?" **What this tests:** - Phase 0: Consumer classification (should trigger Broad scope by default) - Phase 1: Brief clarification (Consumer type, simple question) - Phase 2: Broad calibration (3 researchers, 2 rounds, sonnet model) - Phase 8: Consumer synthesis template (ranked recommendations) - **Iteration 2 fix:** lead-dispatched follow-ups (the biggest change to validate) - **Domain shift:** non-AI topic tests whether Exa search and researcher methodology work outside our comfort zone **Also validates:** - Expert testing lab prioritization in source evaluation (Consumer type hierarchy) - Whether sonnet researchers (cheaper) produce adequate quality for a simpler topic - Whether Broad scope is sufficient for a Consumer question (or if it needs less) ## Research Summary Ran Broad-scope research (3 sonnet researchers, 2 rounds) on "Best mechanical keyboard for programming in 2026?" Full report at `~/.claude/research/best-mech-keyboard-programming-2026/report.md`. **Round 1:** 3 researchers investigated: (A) expert reviews/testing, (B) community recommendations, (C) features/specs analysis. Strong convergence on Keychron dominance at every price tier. 30+ sources across researchers. **Round 2:** 2 targeted follow-ups using lead-dispatched model: - Split/ergonomic deep-dive (C, continuity): Consolidated Voyager vs Glove80 vs Defy vs Go60. Discovered MoErgo Go60 as significant new entrant. Alice layout as middle ground. - Blind spot check (A, fresh perspective): Mac-specific picks (NuPhy Air75 V2 top for Mac), office noise reality (no mechanical keyboard is truly quiet), Leopold as overlooked quality. **Outcome:** Consumer-template report with ranked recommendations (Top Pick: Q5 Max, Runner-Up: V5 Max, Budget: C3 Pro, Best Mac: NuPhy Air75 V2, Best Ergonomic: Glove80, Best Gentle Upgrade: Keychron Q8 Max Alice, Best Quiet: HHKB Type-S). 25 cited sources. Report is genuinely useful as a buying guide. ## Observations **What worked well:** 1. **Lead-dispatched follow-ups worked perfectly on the first try.** This is the most important validation from this iteration. The lead created 2 Round 2 tasks, then sent direct SendMessage to researcher-c (split deep-dive, continuity) and researcher-a (blind spots, fresh perspective). Both researchers claimed their tasks and began immediately with zero nudging. Researcher-b was correctly left idle (2 tasks < 3 researchers). **The Iteration 2 architectural change eliminated the self-claim problem completely.** 2. **Lowercase filename normalization worked.** All 5 findings files used correct lowercase naming: `researcher-a-findings.md`, `researcher-b-findings.md`, `researcher-c-findings.md` (Round 1), `researcher-c-findings-7.md`, `researcher-a-findings-8.md` (Round 2). The Iteration 2 fix (explicit lowercase instructions in spawn prompt + researcher-prompt.md) resolved the inconsistency. 3. **Consumer synthesis template produced a well-structured report.** The ranked recommendations format (Top Pick, Runner-Up, Budget, Best for X) naturally organized the findings into actionable buying advice. Each pick has evidence from multiple sources and clear reasoning. 4. **Non-AI domain worked fine.** Exa semantic search handled keyboard/hardware topics well. Researchers found high-quality sources including RTINGS lab testing (278 keyboards), Wirecutter, Reviewed.com, WIRED, plus community sources (Reddit 381K opinion analysis, HN, Devtalk). No domain-specific search issues. 5. **Sonnet researcher quality was adequate for Consumer scope.** Findings were well-structured, sources properly evaluated, and cross-source analysis was competent. The quality difference from opus (Iteration 2) was noticeable but acceptable: sonnet researchers were less likely to proactively label biases or trace claims to primary sources, but the overall research was solid. 6. **All 3 researchers shut down cleanly on first request.** No double-shutdown needed (unlike Iteration 2 where researcher-c required two attempts). This may have been a timing improvement or just variance. 7. **Assignment guidance worked as designed.** Split deep-dive went to researcher-c (who covered splits in Round 1 -- continuity). Blind spot check went to researcher-a (who covered expert reviews -- fresh perspective on what was missed). The guidance in Phase 6 made assignment decisions easy and justified. **What was clunky or problematic:** 1. **Round 1 had no gaps requiring urgent follow-up.** The Broad scope consumer question produced such strong convergence in Round 1 (Keychron dominance was overwhelming) that Round 2 follow-ups were useful but not essential. The split/ergonomic deep-dive and blind spot check added genuine value (Go60 discovery, Mac-specific picks, noise reality check), but the core recommendations would have been the same without them. **This suggests Focused scope (1 round) might be sufficient for straightforward Consumer questions.** 2. **Researcher A idle-looped before claiming task.** After receiving the follow-up assignment via SendMessage, Researcher A sent 2 idle notifications before starting work. The task was eventually claimed and completed successfully, so this is cosmetic, but the idle chatter is noisy in the notification stream. 3. **Report length was ~200 lines, 25 footnotes.** More manageable than the Contested report (~350 lines, 37 footnotes) but still substantial. The Consumer template's multiple "Best for X" categories naturally expand the report. No length target exists in the skill. 4. **Phase 1 clarification was completely skipped.** The skill says "Brief clarification only if genuinely ambiguous" for Consumer questions. The question was clear enough to skip entirely, which is correct behavior, but it means the user had no opportunity to specify constraints (budget range, layout preference, Mac vs Windows, etc.) that could have focused the research. For Consumer questions, a quick "Any constraints?" might be valuable even when the question seems clear. **Surprising or noteworthy:** - The dharm.is analysis of 381K Reddit opinions (found by researcher-b) was an unexpectedly rich quantitative source for a Consumer topic. It revealed the enthusiast vs practical programmer preference gap that informed the report's framing. - The MoErgo Go60 (found by researcher-c in Round 2) was genuinely new information -- launched late 2025, only one detailed review exists. This is exactly the kind of find that makes Round 2 follow-ups worthwhile even when Round 1 converges. - Total wall-clock time was ~15 minutes for the full Broad scope run, similar to Iteration 1. ## Changes Made **No skill file changes this iteration.** The Iteration 2 changes (lead-dispatched follow-ups, lowercase filenames, round stopping criteria, assignment guidance) all validated successfully. The observations suggest potential improvements but none urgent enough to change now: - Scope calibration for simple Consumer questions (Focused might suffice) -- worth investigating as a meta-research topic rather than a code change - Consumer-specific Phase 1 constraint gathering -- minor UX improvement, not blocking - Idle notification noise -- cosmetic, not actionable in the skill ## Notes for Next Iteration - **Lead-dispatched model is confirmed working across 2 question types.** The self-claim problem is solved. No further changes needed to this architecture. - **All Iteration 2 changes validated.** Lowercase filenames, assignment guidance, and round stopping criteria all worked as designed. The skill is now stable for these features. - **Good candidate for Iteration 4:** Meta-research on scope calibration. Three iterations have now produced data on scope appropriateness: Iteration 1 (Broad for compound Technical + Consumer, felt right), Iteration 2 (Comprehensive for Contested, felt right but expensive), Iteration 3 (Broad for Consumer, Round 2 was useful but not essential). Could the skill provide better guidance on when Focused is sufficient vs when Broad is needed? - **Alternative candidate:** Emerging/Frontier test (backlog #3) to exercise the last untested synthesis template. Or information cascade detection (backlog #4) as the first meta-research iteration. - **The skill is approaching diminishing returns for empirical testing.** Three successful runs across 3 question types (compound Technical+Consumer, Contested, Consumer) with improving results each iteration. The remaining untested templates (Scientific/Health, Emerging/Frontier, Opinion/Sentiment, Factual) follow the same architecture. Meta-research may yield higher marginal value than additional empirical runs. -
iteration-04.md 6.4 KB
# Iteration 4 -- Search Tool Audit + Access Workarounds (Meta-Research) ## Plan **Type:** Meta-research (run through the team skill, not manual) **Backlog item:** #4 **Scope:** Broad (3 sonnet researchers, 2 rounds) **Goal:** Audit the full Exa and Firecrawl tool sets for underused features, research academic search options and paywall workarounds, then update `researcher-prompt.md` with improved guidance. ## Research Summary **Report:** `~/.claude/research/web-research-tool-effectiveness/report.md` Round 1 angles: (A) Exa/Firecrawl advanced features, (B) content access workarounds, (C) academic search strategies. Round 2 follow-ups: (A) optimal Exa+Firecrawl workflow decision tree, (B) blind spot check (OpenAlex, Firecrawl limitations, additional APIs). Researcher C shut down early (no follow-up needed). 24+ sources, 5 findings per researcher. **Critical finding:** Both Exa and Firecrawl MCP tools expose only a subset of their full API capabilities. **Resolved post-iteration:** Switched Exa from npm package to hosted HTTP endpoint (`mcp.exa.ai/mcp`) with `web_search_advanced_exa` and `people_search_exa` enabled. This exposed all missing parameters (category, includeDomains/excludeDomains, date filtering, highlights/summary). Firecrawl search was then removed from the researcher prompt as redundant -- Exa advanced covers domain targeting, date filtering, and category filtering natively, reducing researcher cognitive load from 5 tools to 4. ## Observations **What worked well:** - First meta-research iteration run through the team skill -- worked as well as empirical runs. The compound Technical framing with 3 angles mapped naturally to the 4 sub-questions. - Lead-dispatched follow-ups: zero nudging needed (4th consecutive success). - Researchers independently discovered the same critical insight (Exa MCP limitations) from different angles -- genuine cross-validation. - Researcher A did empirical testing in Round 2 (actually ran comparative Exa vs Firecrawl queries), adding real evidence beyond documentation review. - Early shutdown of Researcher C was clean and efficient. - Researcher B found significant blind spots in Round 2 (OpenAlex, Bluesky API, Firecrawl Cloudflare issues, underused Firecrawl formats) -- the follow-up was genuinely valuable. **Process observations:** - Meta-research through the team skill is viable and arguably better than manual research -- the multi-angle approach and forced cross-referencing catch things manual research would miss. - The decision to frame "search tool audit" as a Technical compound question worked well. Could serve as a template for future meta-research iterations. **No problems found.** The skill ran smoothly across all phases. ## Changes Made **5 changes to `researcher-prompt.md`:** 1. **Expanded "Available Tools" section** -- Added Firecrawl search to the tool list, added a "When to use which" decision guide distinguishing Exa (semantic discovery) from Firecrawl search (precision targeting with operators) from Firecrawl scrape (content extraction). 2. **Added "Firecrawl Search Operators" section** -- Table of Google-style operators (site:, exact match, exclusion, intitle:, inurl:, tbs date filtering with examples). Includes caveat about tbs filtering by index date vs publication date. 3. **Added "Academic Search" section** -- Three-step workflow (Exa for broad discovery -> Firecrawl search with site: for domain targeting -> Firecrawl scrape for full content). Lists useful academic domain targets. Lists free academic APIs (Semantic Scholar, OpenAlex, arXiv) with URL patterns. 4. **Expanded "Content Extraction" section** -- Added advanced scrape features: summary format for triage, links format for citation chains, JSON extraction with schema, includeTags/ excludeTags, waitFor for SPAs, proxy options, PDF parsing, maxAge caching. Added "if scrape returns empty" workflow. 5. **Added "Accessing Restricted Content" section** -- Complete workaround guide: archive.ph for paywalls (with URL patterns), xcancel.com for Twitter/X, Reddit .json endpoint, Freedium for Medium, Bluesky public API, LinkedIn limitations. Each with exact URL patterns. **1 additional change:** 6. **Updated Setup section** -- Added `ToolSearch: "+firecrawl search"` to load Firecrawl search tool alongside scrape. **Post-checkpoint fix:** Prettier mangled `>` characters in source evaluation priority chains (e.g., `documentation > code examples > blog posts`) into markdown blockquotes. Replaced all `>` separators with ", then" phrasing across all 6 type-specific entries for prettier safety. **Post-iteration infrastructure change (backlog #18 resolved):** Switched Exa MCP from npm package (`exa-mcp-server`) to hosted HTTP endpoint (`mcp.exa.ai/mcp`) with `web_search_advanced_exa` and `people_search_exa` enabled. This exposed all parameters identified as missing during Iteration 4 research. Firecrawl search was then removed from the researcher prompt entirely -- Exa advanced covers domain targeting, date filtering, and category filtering natively. Net effect: researchers now have 4 tools instead of 5, simpler decision-making, and more powerful search capabilities. **Post-iteration bug discovery (contextMaxCharacters suppresses summaries/highlights):** Investigated why `enableHighlights` and `enableSummary` produced no visible output through the Exa MCP endpoint. Read the npm package source and found two bugs in the response formatter: (1) When `contextMaxCharacters` is set, the API returns a `context` field that the formatter uses exclusively, dropping per-result data including summaries and highlights. (2) The `ExaSearchResult` type doesn't include `highlights` at all -- they're requested from the API but silently dropped. **Fix:** Use `textMaxCharacters: 1` + `enableSummary: true`. Do NOT set `contextMaxCharacters`. **Post-iteration tool simplification:** Reduced Exa tools from 5 to 2: `web_search_advanced_exa` + `get_code_context_exa`. **Post-iteration category guidance:** Added category restriction table and usage hints from Exa skills repo. ## Notes for Next Iteration - Researcher-prompt.md is ~420 lines. Monitor for issues. - Default Exa search pattern (`enableSummary: true` + `textMaxCharacters: 1`) untested in empirical run. First test will reveal whether researchers follow the guidance. - Category filter restrictions documented but untested. - Academic search strategies documented but untested in Scientific/Health run. -
iteration-05.md 3.6 KB
# Iteration 5 -- Scientific/Health Empirical (IF + Cognitive Performance) ## Plan **Type:** Empirical (Scientific/Health question) **Backlog item:** #12 **Scope:** Broad (3 sonnet researchers, 2 rounds) **Goal:** Test the evidence pyramid synthesis template, academic search guidance, and the new Exa `enableSummary` + `textMaxCharacters: 1` default pattern. Also test whether the category filter restrictions table prevents 400 errors. Secondary goal: validate that the architecture continues to work smoothly (6th consecutive run). **Question:** "Does intermittent fasting improve cognitive performance?" **Angles:** - (A) Systematic reviews and meta-analyses (anchor evidence) - (B) Mechanistic pathways (BDNF, neuroinflammation, ketones, gut-brain axis) - (C) Real-world protocols, practitioner evidence, and cognitive risks ## Research Summary **Report:** `~/.claude/research/intermittent-fasting-cognitive-performance/report.md` Round 1: All 3 researchers completed successfully. Researcher A found the Bamberg & Moreau 2025 meta-analysis (63 studies, 3,484 participants) as the anchor -- acute fasting is cognitively neutral. Researcher B mapped 5 mechanistic pathways with 13 sources. Researcher C covered protocols, null results, and disordered eating risks with 11 sources. Triage identified: (1) CCR vs IF conflict (Researcher A's O'Leary review claimed CCR superior, but other evidence was more nuanced), (2) population-specificity gap (which populations benefit most?). Round 2: Researcher B (fresh perspective) resolved the CCR vs IF conflict -- O'Leary conflated fasting types; head-to-head trials show no difference when calorie-matched. Researcher C (continuity) deep-dived population specificity -- metabolically impaired populations benefit most, age is a moderator, ApoE4 data is missing. Final report: ~300 lines, 26 footnotes, Scientific/Health evidence pyramid template. **Key conclusions (calibrated confidence):** - Acute fasting is cognitively neutral in healthy adults (highly likely, 80-92%) - Chronic IF evidence is insufficient to claim cognitive benefit (highly likely) - Metabolically impaired populations benefit most (likely, 63-79%) - IF ≈ CCR when calorie-matched (likely) - Mechanistic pathways are plausible but human translation is incomplete (likely) ## Observations **What worked well:** - Evidence pyramid template worked excellently. - Lead-dispatched follow-ups: 6th consecutive success, zero nudging needed. - Academic source quality was strong (APA, Cell Metabolism, Nature, BMJ Gut, Springer, MDPI). - Triage cross-referencing caught genuine issues. - Calibrated confidence language worked naturally. - All researchers completed on first try. **Exa parameter compliance:** 32/32 Exa calls fully compliant with summary pattern. Categories used: `research paper` (frequent), `personal site` (once). No 400 errors. 28 Firecrawl scrape calls, all with `onlyMainContent: true`. ## Changes Made **3 changes (implementing backlog #23 -- researcher shutdown feedback reports):** 1. Updated Shutdown section in `references/researcher-prompt.md` with feedback file template. 2. Added Step 2 to Phase 10 (Cleanup) in `SKILL.md` for reading feedback files. 3. Added `scripts/analyze-transcripts.py` for post-hoc JSONL analysis. **Backlog items resolved:** #12 (Scientific/Health empirical), #23 (shutdown feedback reports). ## Notes for Next Iteration - Feedback reports implemented but untested. Next run will validate. - Researcher-prompt.md is 438 lines. Monitor. - Evidence pyramid validated. No changes needed. - Exa guidance: 32/32 compliant. No further investment needed. - Category restrictions untested with restricted categories (tweet, company, people). -
iteration-06.md 6.9 KB
# Iteration 6 -- Emerging/Frontier Empirical (LLM Reasoning) ## Plan **Type:** Empirical (Emerging/Frontier question) **Backlog item:** #9 **Scope:** Comprehensive (4 opus researchers, 3 rounds incl. verification) **Goal:** Test the last major untested synthesis template (Emerging/Frontier with instability warnings, date-sensitive source evaluation). First run with shutdown feedback reports (implemented in Iteration 5 but untested). Exercise Exa categories beyond `research paper` (`news`, `personal site`, `tweet`), testing the category restriction table. **Question:** "What are the current approaches to LLM reasoning?" **Angles:** - (A) Foundational paradigms: CoT, ToT, GoT, and their successors - (B) Training-time reasoning: PRM, RL, GRPO, RLVR, reasoning-focused training - (C) Inference-time compute scaling: search, verification, test-time compute allocation - (D) Emerging approaches: neurosymbolic, multi-agent, tool-augmented, latent reasoning ## Research Summary **Report:** `~/.claude/research/llm-reasoning-approaches/report.md` **Round 1:** All 4 opus researchers completed successfully. Strong convergence on RLVR as dominant paradigm shift (A, B, C). Latent reasoning identified as frontier by A and D independently. B provided deep GRPO/RLVR coverage. C mapped inference-time compute scaling with strong evidence. D covered neurosymbolic, multi-agent, tool-augmented, latent, and world model approaches. Total: ~60 sources across researchers. **Triage identified:** (1) PRM vs ORM conflict (B says PRMs disappoint at scale; C says PRMs outperform), (2) all 4 researchers noted heavy academic bias -- practitioner/deployment perspectives missing, (3) competitive landscape gap. **Round 2:** Researcher A (fresh perspective) resolved PRM conflict -- both correct in different regimes (inference-time verification: PRM wins; large-scale RL training: ORM preferred). Researcher D (gap-fill) surveyed practitioner perspectives and competitive landscape using diverse Exa categories (news, personal site, tweet). **Round 3 (Verification):** Researcher B verified CoT faithfulness claims -- 2.3% True Thinking Score CONTESTED (single model/dataset, not replicated; direction supported but specific number misleading), <20% verbalization SUPPORTED WITH NUANCE (multiple independent confirmations; METR argues manageable for safety monitoring). **Final report:** ~400 lines, 55 footnotes. Emerging/Frontier template with instability warnings. ## Observations **What worked well:** - **Emerging/Frontier template worked excellently.** The "Current State (as of date)" + "Trajectory" + "Instability Warning" structure produced a report that acknowledges the rapidly evolving nature of the field while still providing definitive findings. - **Lead-dispatched follow-ups:** 7th consecutive success, zero nudging needed. - **Cross-agent verification added genuine value.** The CoT faithfulness verification produced a nuanced CONTESTED verdict that materially improved the report -- the 2.3% number was tempered to a directional finding. - **PRM conflict resolution was clean.** Researcher A (fresh perspective) found the resolution along the inference vs training axis with 13 independent sources. - **All researchers completed on first try.** No stalls, no nudging, no retries. - **Category restriction table worked.** `tweet` category used successfully (no 400 errors). Researcher D correctly omitted restricted params for tweet queries. **Exa parameter compliance (verified via `scripts/analyze-transcripts.py`):** - **49/49 Exa calls (100%) fully compliant.** `enableSummary: true` + `textMaxCharacters: 1`, zero `contextMaxCharacters` violations. 7th session with perfect compliance (cumulative 81/81 across Iterations 5-6). - **Categories exercised:** `research paper` (8), `news` (2), `personal site` (2), `tweet` (1). First run using `news`, `personal site`, and `tweet` categories. No 400 errors. - **Firecrawl:** 22 scrape calls, all `firecrawl_scrape`. No map or other tools used. - **Rich parameter usage:** `startPublishedDate` widely used, `numResults` varied (5-10), `includeDomains` used once (metr.org). **Shutdown feedback reports (first test -- backlog #23 validated):** All 4 researchers wrote feedback files. Key findings from feedback: - **arxiv scraping noise:** 3 of 4 researchers independently flagged that Firecrawl scraping of arxiv abstract pages returns mostly navigation boilerplate. All recommended using `arxiv.org/html/` or relying on Exa summaries for triage. - **Tweet category thinness:** Researcher D noted tweet summaries were 1-2 sentences with limited context. The category works but provides lower signal density. - **Duplicate self-referential messages:** 3 of 4 researchers reported receiving their own task assignment messages echoed back. Cosmetic but confusing. Likely a platform bug. - **Verification verdict granularity:** Researcher B noted the SUPPORTED/CONTESTED/UNCHANGED trichotomy doesn't cover partial support. Iteration 2 verifiers also used "SUPPORTED WITH NUANCE." - **Conflict resolution length:** Researcher A noted ~120-line target is too short for conflict resolution tasks that need comparison tables. - **Positive:** All researchers praised the `enableSummary` + `textMaxCharacters: 1` pattern for efficient triage. **Feedback reports vs transcript analysis comparison:** Both mechanisms add unique value. Transcripts provide exact parameter compliance (definitive). Feedback reports provide qualitative insights not visible in transcripts: arxiv scraping issue, duplicate message bug, tweet category thinness, verification verdict suggestions. **Keep both.** ## Changes Made **No skill file changes this iteration.** The Emerging/Frontier template, feedback reports, and category restriction table all validated successfully. Two quick changes identified for next iteration (backlog #26 arxiv guidance, #27 verification verdict expansion) but not blocking. **Backlog items resolved:** #9 (Emerging/Frontier empirical), #23 validated (feedback reports work as designed). ## Notes for Next Iteration - **All 7 synthesis templates now tested.** Factual, Scientific/Health, Consumer, Technical, Opinion/Sentiment (untested), Contested, and Emerging/Frontier. Only Opinion/Sentiment remains untested, but it's the least complex template. The skill is mature for synthesis. - **Exa compliance is thorough.** 81/81 across 2 measured iterations. No further monitoring investment needed -- the guidance works reliably. - **Feedback reports validated and useful.** They surfaced 4 actionable findings this iteration. Worth keeping as standard practice. - **Quick wins available:** Backlog #26 (arxiv guidance) and #27 (verification verdict expansion) are small, targeted changes that can be done without a full research run. - **Recommended next:** Backlog #25 (audit against Anthropic agent teams docs) is a quick meta task that doesn't require running the skill. Good for a shorter iteration. -
iteration-07.md 6.7 KB
# Iteration 7 -- Agent Teams Docs Audit (Meta) ## Plan **Type:** Meta-research (docs audit) **Backlog item:** #25 **Scope:** No research run -- audit existing skill against official Anthropic documentation. **Goal:** Read official Claude Code agent teams docs and compare our implementation against recommended patterns, known limitations, and best practices. Use Exa for discovery and Firecrawl for scraping (not the claude-code-guide agent). Also implement quick wins #26 (arxiv guidance) and #27 (verification verdict expansion). ## Sources Scraped 6 official pages scraped with Firecrawl (`onlyMainContent: true`, `maxAge: 86400000`): 1. `code.claude.com/docs/en/agent-teams` -- Primary agent teams documentation 2. `code.claude.com/docs/en/sub-agents` -- Subagents documentation (for comparison) 3. `platform.claude.com/docs/en/agent-sdk/subagents` -- SDK subagents (programmatic API) 4. `code.claude.com/docs/en/best-practices` -- Best practices guide 5. `anthropic.com/engineering/multi-agent-research-system` -- Anthropic's own research system 6. `claude.com/blog/building-multi-agent-systems-when-and-how-to-use-them` -- Multi-agent guide Plus 3 Exa searches for discovery (30 results total across `docs.anthropic.com`, `github.com/anthropics`, and general web). ## Audit Findings **Green Flags (architecture confirmed correct):** - **Hub-and-spoke / orchestrator-worker pattern** matches both the agent teams docs and Anthropic's own research system architecture. The docs explicitly describe "one session acts as the team lead, coordinating work, assigning tasks, and synthesizing results." - **No peer-to-peer researcher communication** aligns with the agent teams docs' model where the lead coordinates all work. The multi-agent blog emphasizes "context-centric decomposition" and warns against the "telephone game" of inter-agent information loss. - **Lead-dispatched assignments** validated. Docs support both explicit assignment and self-claim with file locking, but our empirical finding (Iterations 1-2) that lead-dispatched works better is consistent with Anthropic's advice to "teach the orchestrator how to delegate." - **Filesystem output** explicitly recommended. Anthropic's research blog says: "Subagent output to a filesystem to minimize the 'game of telephone'... subagents call tools to store their work in external systems, then pass lightweight references back to the coordinator." - **Verification subagent pattern** matches the blog's "verification subagent pattern" exactly, including our "assign to fresh perspective" approach. - **Context-centric decomposition by research angle** (not by role) matches the blog's strong recommendation against problem-centric splitting. - **Shutdown then TeamDelete sequence** matches docs' requirement that cleanup fails if active teammates remain. - **No nested teams** -- our researchers don't spawn sub-teams, which is correct since the docs confirm "teammates cannot spawn their own teams." - **Automatic message delivery, no polling** -- confirmed by docs. **New Information (resolved questions):** - **Permission inheritance confirmed:** "Teammates start with the lead's permission settings. If the lead runs with `--dangerously-skip-permissions`, all teammates do too." Can change individual modes after spawning, but not at spawn time. This partially resolves backlog #5. - **CLAUDE.md + MCP servers inherited automatically:** "When spawned, a teammate loads the same project context as a regular session: CLAUDE.md, MCP servers, and skills." This explains why researchers can use ToolSearch to load Exa/Firecrawl -- they inherit MCP server access. - **Task dependencies available:** Framework supports `blockedBy` for task ordering. We don't use this (our lead-dispatched approach handles sequencing manually), but it's available if we ever need it. - **Session resumption doesn't restore teammates:** Our `state.md` resume check in Phase 0 is important precisely because of this limitation. Good that we built this. - **Delegate mode exists:** Pressing Shift+Tab restricts the lead to coordination-only tools. Not needed for our skill (which operates programmatically) but useful to know. **Anthropic's research system comparison:** Their architecture is strikingly similar to ours. Key parallels: orchestrator-worker pattern, lead decomposes query and spawns subagents for different facets in parallel, subagents search independently then return distilled findings, lead synthesizes. Their system uses Opus 4 lead + Sonnet 4 subagents and outperformed single-agent by 90.2% on their internal eval. Key difference: they found "token usage by itself explains 80% of the variance." Their effort scaling rules: simple queries get 1 agent with 3-10 tool calls, direct comparisons 2-4 agents with 10-15 calls each, complex research 10+ agents. Our scale is more conservative (max 4 researchers for Comprehensive) but appropriate for our token budget. They also use a separate CitationAgent for adding citations to the final output. We don't have this -- our lead handles synthesis and citations together. Worth monitoring but not needed now. **No anti-patterns found.** Our architecture doesn't fight the framework, uses no deprecated patterns, and follows the recommended cleanup sequence. **Yellow flags (minor, non-blocking):** - Our Comprehensive scope uses Opus for all researchers. Anthropic's own system uses Sonnet workers successfully. Consider testing Sonnet researchers in Comprehensive to reduce cost while maintaining quality. Added as backlog #28. - No explicit TaskList polling during long researcher waits. If a researcher silently crashes, the lead would wait indefinitely. Low risk (hasn't happened in 7 iterations) but worth awareness. ## Changes Made 1. **[#26] arxiv scraping guidance** -- Added to `references/researcher-prompt.md` under Content Extraction: guidance to use `arxiv.org/html/{id}` instead of `/abs/`, prefer Exa summaries for triage, use arxiv API for structured extraction. 2. **[#27] Verification verdict expansion** -- Added SUPPORTED WITH NUANCE as fourth verdict in SKILL.md Phase 7, with description: "Directionally correct but specific details differ or require qualification." **Backlog items resolved:** #25 (docs audit), #26 (arxiv guidance), #27 (verification verdict). ## Notes for Next Iteration - Architecture is validated and mature. 7 iterations without a major structural change needed. - Permission setup (#5) is now better informed: inheritance confirmed, specific permissions needed are mkdir + Write to research dir + Bash. - Scope calibration (#6) has 7 data points and strong evidence from Anthropic's own scaling rules to compare against. - Consider testing mixed model config: Opus lead + Sonnet researchers for Comprehensive (#28). -
iteration-08.md 4.1 KB
# Iteration 8 -- Scope Calibration Refinement (Meta) ## Plan **Type:** Meta-analysis (internal data synthesis) **Backlog item:** #6 **Scope:** No research run -- analyze 7 empirical data points + Anthropic's scaling rules to refine scope calibration guidance. **Goal:** Determine whether the type-to-scope defaults need adjustment, add escalation and de-escalation signals, and add empirical output expectations. Also subsumes backlog #8 (report length targets). ## Analysis **Data compiled from 7 runs:** | Iter | Type | Scope | Rounds Used | Sources | Lines | Round 2 Value | Assessment | | ---- | ------------------ | ------------- | ----------- | ------- | ----- | --------------------- | -------------- | | 1 | Technical+Consumer | Broad | 2 | 25 | ~250 | Useful | Right-sized | | 2 | Contested | Comprehensive | 3 | 60+ | ~350 | Useful | Right-sized | | 3 | Consumer | Broad | 2 | 25 | ~200 | Useful-not-essential | Slight overfit | | 4 | Technical (meta) | Broad | 2 | 24 | ~280 | Useful | Right-sized | | 5 | Scientific/Health | Broad | 2 | 26 | ~300 | Useful | Right-sized | | 6 | Emerging/Frontier | Comprehensive | 3 | 55 | ~400 | Useful (conflict+gap) | Right-sized | | 9 | Opinion/Sentiment | Comprehensive | 2 | 60+ | ~370 | Useful (all 4 gaps) | Right-sized | **Key findings:** 1. **Current defaults are correct 6/6 times.** Consumer was the only case with slack (Round 2 useful-not-essential). All other scopes were right-sized. 2. **Consumer is the only type with de-escalation evidence.** When the product category is narrow, well-reviewed, and expert consensus is clear, Focused could suffice. Iteration 3 (keyboards) is the data point. 3. **Comprehensive verification (Round 3) added value both times tested.** Iteration 2: SUPPORTED WITH NUANCE + CONTESTED verdicts. Iteration 6: CONTESTED + SUPPORTED WITH NUANCE. Both materially improved the final report. 4. **Focused scope has never been tested.** Only Factual defaults to it, and no Factual question has been run. Output expectations for Focused are projected from Round 1 subsets of Broad runs. 5. **Anthropic comparison:** Their simple (1 agent, 3-10 calls) maps to "don't use this skill." Their comparisons (2-4 agents) map to our Focused/Broad. Their complex (10+ agents) exceeds our maximum. Our system is well-sized for the middle range. 6. **Report length is scope-driven, not type-driven.** The variation across types within Broad (200-300 lines) is smaller than the variation across scopes. So output expectations belong on the scope tier, not the question type. ## Changes Made 1. **Scope modifiers added to `references/question-types.md`** -- Per-type de-escalation and escalation signals, with the Consumer de-escalation backed by Iteration 3 evidence. Also added "do not use this skill" criteria for simple queries. 2. **Empirical output expectations added to `references/question-types.md`** -- Table of report length and source count ranges per scope tier, based on 7 runs. Focused estimates projected. Explicitly noted these are descriptive ranges, not targets. 3. **SKILL.md Phase 2 updated** -- Now references scope modifiers from question-types.md instead of inlining ad-hoc adjustment criteria. **Backlog items resolved:** #6 (scope calibration), #8 (report length targets). ## Notes for Next Iteration - **Focused scope remains untested.** A Factual question run would validate the Focused tier and the projected output expectations, but it's low priority since the skill explicitly says "don't use for simple factual lookups." - **The main remaining empirical gap is Opinion/Sentiment** -- the last untested synthesis template. A good candidate for the next empirical run. - **Mixed model config (#7) is the highest-value empirical test** -- combining it with an Opinion/Sentiment question would test two things at once. -
iteration-09.md 7.1 KB
# Iteration 9 -- Opinion/Sentiment + Mixed Model Config (Empirical) ## Plan **Type:** Empirical (Opinion/Sentiment question) **Backlog items:** #7 (mixed model config), last untested synthesis template **Scope:** Comprehensive (4 sonnet researchers, 2 rounds -- model override from opus to sonnet) **Goal:** Combine two tests in a single empirical run: (1) Validate Sonnet researchers for Comprehensive scope (backlog #7), replacing the Opus default. Anthropic's own research system uses Opus lead + Sonnet subagents. (2) Test the Opinion/Sentiment synthesis template, the last of 7 templates never used in a real run. **Question:** "What do developers actually think about Nix and NixOS in 2026?" **Angles:** - (A) Community sentiment: overall vibe, praise vs frustration, how feelings have shifted - (B) Pain points and barriers: learning curve, documentation, Nix language, flakes controversy - (C) Enterprise/professional adoption: companies using Nix, CI/CD patterns, onboarding stories - (D) Governance and community health: foundation drama, forks (Lix, Aux, Tvix), SC elections ## Research Summary **Report:** `~/.claude/research/developer-opinions-nix-nixos-2026/report.md` **Round 1:** All 4 Sonnet researchers completed successfully. Strong convergence on the paradox theme: fastest community growth ever (30% YoY) alongside deepest fragmentation ever (governance crisis, three forks). Researcher A surfaced the sentiment spectrum from "Nix changed my life" to "I quit after 6 months." Researcher B documented the learning curve and documentation as top pain points across all sources. Researcher C found enterprise adoption concentrated in devshells-only patterns with wrapper tools. Researcher D provided deep coverage of the 2024-2025 governance crisis and its aftermath. Total: ~50+ sources across researchers. **Triage identified:** (1) Gap: moderate/ambivalent voices underrepresented (survivorship bias in online discussions), (2) DetSys conflict of interest not fully explored (company behind flakes also behind commercial products), (3) wrapper tools ecosystem deserved deeper comparison, (4) Steering Committee election results and their impact not covered. **Round 2:** Researcher A (gap-fill) found moderate voices -- "Nix purgatory" users stuck between love and frustration, plus successful teams who don't post online. Researcher B (deep-dive) investigated DetSys conflict with nuance -- found community resentment but also genuine engineering contributions. Researcher C (deep-dive) did comprehensive wrapper tools comparison with GitHub stars, feature matrices, and community adoption data. Researcher D (gap-fill) covered SC election results and post-crisis governance trajectory. **Round 3 skipped:** No single-source high-impact claims remaining. All major conflicts resolved by Round 2. Remaining uncertainties about magnitude, not direction. **Final report:** ~370 lines, 50 footnotes, 60+ sources. Opinion/Sentiment template with 5-segment distribution of views and representativeness assessment. ## Observations **What worked well:** - **Opinion/Sentiment template validated with adaptation.** The 3-segment default structure (Majority/Minority/Outlier) was adapted to 5 segments to capture the nuanced reality: Passionate Core (~15-20%), Pragmatic Middle (~30-35%), Frustrated Dropouts (~20-25%), Nix Purgatory (~15-20%), and Broader Dev Population (~10-15%). The template's Representativeness Assessment section added genuine value. - **Sonnet researchers produced Comprehensive-quality output.** All 4 researchers completed both rounds, followed all methodology guidance, wrote structured findings and feedback reports. No evidence of quality degradation vs Opus researchers in Iterations 2 and 6. - **Lead-dispatched follow-ups:** 8th consecutive success, zero nudging needed. - **All researchers completed on first try.** No stalls, no nudging, no retries. - **Feedback reports written by all 4 researchers.** Continued validation of this mechanism. - **Verification skip was correct.** Round 2 resolved all major gaps and conflicts. No single-source claims would have changed the report's conclusions. **Permission friction observed:** - **curl commands required user approval.** Researchers used `curl` for Reddit JSON API and GitHub API endpoints, triggering ~8-10 permission prompts during the run. This is the first time permission friction was significant enough for the user to raise it post-run. - **Root cause:** The researcher-prompt.md Reddit section recommends curl for Reddit JSON API (since Exa doesn't index Reddit and Firecrawl blocks it). GitHub API was also accessed via curl for stars/forks data. **Researcher feedback highlights:** - Researcher A: Exa with enableSummary + textMaxCharacters:1 excellent for rapid triage. Category gotchas (tweet, company) useful to know upfront. - Researcher B: Firecrawl scraping of HN threads returned full comment trees in markdown -- very useful for community sentiment capture. - Researcher C: Reddit JSON API `restrict_sr=on` was essential -- without it, global results instead of subreddit-specific. Exa didn't index GitHub issue bodies well. Recommended GitHub search API patterns in methodology. - Researcher D: Strong coverage of Discourse-based forums (NixOS Discourse) via direct Firecrawl scraping. Exa worked well for blog posts and personal sites. ## Changes Made 3 changes to `references/researcher-prompt.md`: 1. **Reddit section updated** -- Added `restrict_sr=on` requirement with example URL pattern, guidance to prefer `/top/` and `/hot/` over `/search/` for sentiment research (browsing reveals organic discussions), note that curl requires user permission approval each time. 2. **GitHub data guidance added** -- New section before "Medium paywalled articles": Exa `get_code_context_exa` as primary tool for GitHub content discovery, Firecrawl for deep extraction of specific pages, `gh api` for quantitative data (stars, forks, issues) since it's already permitted. Explicit "do not use curl for GitHub API endpoints." 3. **Discourse JSON search API added** -- New section after GitHub: Firecrawl scrape of `/search.json?q=...` as fallback when Exa can't find specific forum topics. Returns structured results with topic IDs for follow-up scraping. **Backlog items resolved:** #7 (mixed model config -- Sonnet validated for Comprehensive). ## Notes for Next Iteration - **All 7 synthesis templates now tested.** The skill is mature for research operations. Opinion/Sentiment required the most template adaptation (5 segments vs 3) but the framework accommodated this well. - **Sonnet validated as viable Comprehensive override, Opus kept as default.** One successful run (Opinion/Sentiment) doesn't prove Sonnet handles the hardest Comprehensive questions. SKILL.md notes Sonnet as an available override when cost matters. - **Reddit MCP is the top engineering priority.** The curl permission friction this iteration is the strongest signal yet. Backlog #8 has full context on options. - **Permission setup (#5) remains important** for shareability. Inheritance is confirmed but specific Bash permissions (curl, sleep) still cause friction. -
iteration-10.md 3.8 KB
# Iteration 10 -- Reddit MCP Server Integration (Engineering) ## Plan **Type:** Engineering (no research run) **Backlog item:** #8 (Reddit MCP server integration) **Goal:** Eliminate the last curl dependency in researcher workflows by integrating a Reddit MCP server via 1MCP, and updating the researcher methodology to use native MCP tools instead of curl for Reddit content access. **Background:** Iteration 9 confirmed significant permission friction (~8-10 curl prompts per run) from Reddit JSON API access. Prior research (`~/.claude/research/reddit-access-for-claude-code/report.md`) identified `jordanburke/reddit-mcp-server` as the top option. Fresh evaluation confirmed: - Still actively maintained (v1.2.1 published ~10 days ago, 108 commits, 4 contributors) - 9 read tools + 6 write tools, including `search_reddit` (critical for researchers) - Anonymous mode: zero credentials, ~10 rpm, zero setup - `npx reddit-mcp-server` -- no install needed, 1MCP compatible via stdio - Runner-up (Hawstein/mcp-server-reddit) is stale -- no commits in 10 months, no search tool - New entrants (liuyang1520, karanb192) too immature (0 stars, requires auth/local install) Legal context unchanged: Reddit v. Anthropic lawsuit makes this fraught, but low-volume personal use risk is low. Anonymous mode doesn't even use API credentials. ## Changes Made ### 1. Added Reddit MCP server to 1MCP config **File:** `configs/claude/1mcp.json` Added `reddit` server entry with `npx reddit-mcp-server` and explicit `REDDIT_AUTH_MODE: "anonymous"`. Placed alphabetically before the `gws` entry. Server will be available via 1MCP as `mcp__1mcp__reddit_1mcp_{tool_name}`. ### 2. Updated researcher methodology -- Setup section **File:** `~/.claude/skills/deep-research-team/references/researcher-prompt.md` Added `ToolSearch: "+reddit"` to the tool loading instructions in the Setup section. ### 3. Updated researcher methodology -- Available Tools section **File:** same Added **Reddit** tool block documenting 4 content extraction tools (one-line descriptions, params left to MCP schemas): `get_top_posts`, `get_post_comments`, `get_reddit_post`, `get_subreddit_info`. Deliberately excluded `search_reddit` -- Google's Reddit index (via Firecrawl search) is far superior to Reddit's native search. Added to "When to use which": Firecrawl search for Reddit discovery (`site:reddit.com`), Reddit MCP tools for content extraction. ### 4. Replaced curl-based Reddit access with MCP tools **File:** same, "Accessing Restricted Content > Reddit content" section Completely rewrote the Reddit content section. Removed all curl-based guidance. Replaced with two-layer pattern: Firecrawl search (`site:reddit.com {query}`) for discovery, Reddit MCP tools for content extraction. Preserved the key insight that `get_top_posts` is preferred over search for sentiment research. ## Impact - **Permission friction:** Eliminates ~8-10 curl permission prompts per research run that accesses Reddit. Reddit access now uses native MCP tools that require no special permissions. - **Researcher experience:** Researchers use the same tool pattern (ToolSearch -> MCP call) for Reddit as they do for Exa and Firecrawl. No more shelling out to curl. - **Coverage:** All previously curl-accessible Reddit features are covered by MCP tools, plus additional features (engagement analysis, subreddit stats, trending subreddits). - **curl dependency:** With GitHub curl eliminated in Iteration 9 and Reddit curl eliminated here, researchers should have zero curl dependencies in normal workflows. The only remaining curl use cases are the Bluesky public API and rate-limit retry sleeps (which use Bash `sleep`, not curl). ## Backlog Resolution **#8 resolved.** Reddit MCP server integrated via 1MCP with anonymous mode. Researcher methodology updated to use native MCP tools. curl-based Reddit access guidance removed. -
iteration-11.md 6.9 KB
# Iteration 11 -- Reddit MCP Empirical Validation ## Type Empirical (Focused scope run to validate Iteration 10 engineering changes) ## Question "What is the best community-favorite content about the TV show Dark -- videos, write-ups, essays, analyses, explainers, fan theories?" ## Config - **Scope:** Focused (2 researchers, 1 round) - **Question type:** Consumer/Recommendation - **Model:** Sonnet researchers - **Report:** `~/.claude/research/dark-tv-show-community-content/report.md` ## Validation Goals & Results | Goal | Result | Notes | | --------------------------------------------- | ------ | ------------------------------------------------------------------------------------------- | | Researchers load Reddit tools via ToolSearch | Pass | Both loaded successfully | | Firecrawl search with site:reddit.com works | Pass | Both used it. B called it "the single most valuable technique" | | get_top_posts extracts Reddit content | Pass | Both used. Returns memes/appreciation at top; insufficient alone for substantive threads | | get_post_comments extracts thread discussions | Pass | A used on 4 threads, B on 2. Specific upvote counts, ratios, user quotes extracted | | get_reddit_post works for specific posts | Pass | B used on 7 posts with engagement metrics | | get_subreddit_info works | Pass | B used to get 213K member count | | search_reddit works | Pass | A used 2 queries | | Zero curl usage for Reddit | Pass | No permission prompts observed | | Useful findings produced | Pass | 24-source report with specific engagement data, 5 categories of content | **Overall: 9/9 validation goals passed. Reddit MCP integration works end-to-end.** ## Tool Usage Summary ### Researcher A - **Exa:** enableSummary, textMaxCharacters, numResults (standard pattern) - **Firecrawl:** firecrawl_search for Reddit discovery - **Reddit MCP:** search_reddit (2), get_top_posts (1), get_post_comments (4) - **Issues:** `site:reddit.com/r/DarK` in Firecrawl search was too specific, returned generic results. Broader `site:reddit.com` worked better. ### Researcher B - **Exa:** enableSummary, textMaxCharacters, numResults, excludeDomains, category (personal site) - **Firecrawl:** firecrawl_scrape (4 URLs), firecrawl_search for Reddit discovery - **Reddit MCP:** get_top_posts (1), get_post_comments (2), get_reddit_post (7), get_subreddit_info (1) - **Issues:** Medium paywall blocked full markdown extraction (summary format worked). Top posts from r/DarK were overwhelmingly memes -- needed Firecrawl search for substantive threads. ## Key Observations ### Two-layer Reddit pattern validated The Firecrawl search (discovery) + Reddit MCP (extraction) pattern works well. Both researchers used it independently and both endorsed it in feedback. ### get_top_posts is necessary but insufficient For topic-specific threads, `get_top_posts` returns the most popular posts overall (which for entertainment subreddits are memes and appreciation posts, not analytical content). The Firecrawl search layer is essential for finding substantive threads on specific topics. ### Firecrawl search scoping Researcher A found that `site:reddit.com/r/DarK` (subreddit-scoped) returned generic video essay threads instead of r/DarK-specific content. The broader `site:reddit.com` with topic keywords in the query worked better. This is worth noting in the methodology. ### Researcher complementarity Good angle separation produced minimal overlap. Both independently found the mmmmmmmmichaelscott FAQ and The_Wattsatron easter eggs post (confirming canonical status via independent discovery). A focused on video/audio content; B focused on written/academic content. ### Exa category filter usage Researcher B used `category: "personal site"` to filter to independent blog essays, which effectively surfaced philosophical analyses while filtering out corporate listicle content. Good technique for this question type. ## Changes Made ### 1. Note Firecrawl search scoping in researcher-prompt.md Add note that `site:reddit.com` (broad) works better than `site:reddit.com/r/subreddit` (subreddit-scoped) in Firecrawl search queries. ### 2. Note get_top_posts limitation in researcher-prompt.md Clarify that `get_top_posts` returns the most popular posts overall (often memes/meta for entertainment subreddits), not necessarily the most substantive analytical content. Recommend using Firecrawl search for finding topic-specific substantive threads. ## Transcript Analysis Script: `analyze-transcripts.py --session 5213f595-f4b6-4140-b7e0-5d2908d822d7` **Exa compliance: 10/10 (100%)** - All calls used `enableSummary: true` + `textMaxCharacters: 1` - Zero `contextMaxCharacters` violations **All MCP calls (42 total across both researchers):** - Reddit: 22 (get_reddit_post: 9, get_post_comments: 7, search_reddit: 3, get_top_posts: 2, get_subreddit_info: 1) - Exa: 10 (web_search_advanced_exa: 10) - Firecrawl: 10 (firecrawl_search: 6, firecrawl_scrape: 4) **Notable patterns:** - Researcher B used `excludeDomains: ["reddit.com"]` on Exa to avoid duplicate Reddit results - Researcher B used `category: "personal site"` for blog essay discovery (2 calls) - Reddit was the most-used MCP server — appropriate for a community-content question ## Post-Iteration Housekeeping Changes made in the same session after the main iteration writeup: ### 3. Generalize analyze-transcripts.py to track all MCP tools Replaced hardcoded Exa/Firecrawl parsing with generic MCP tool detection via `re.compile(r"(?:mcp__1mcp__)?(\w+?)_1mcp_(\w+)")`. All 1MCP tools now captured and grouped by server. Custom formatters for Exa (with compliance checking), Firecrawl, and Reddit. Aggregate report shows per-server totals and per-endpoint breakdown. ### 4. Fix --project-dir to accept filesystem paths `--project-dir` now accepts regular filesystem paths (e.g., `/Users/malo/.config/nix-config`) in addition to `~/.claude/projects/{slug}` paths. Slug algorithm: `re.sub(r"[^a-zA-Z0-9]", "-", path)`. ### 5. Simplify researcher feedback template Removed the "Search Tools" section (redundant now that the transcript script extracts all MCP tool usage automatically). Reframed template as qualitative reflection: What Worked, Issues, Suggestions. Added framing note so researchers understand the purpose. ## Backlog Changes - No items resolved (validation run, not feature addition) **Cumulative changes including housekeeping: 30 + 3 = 33.** -
iteration-12.md 6.3 KB
# Iteration 12 -- Firecrawl Scrapability Testing **Type:** Engineering (no research team) **Goal:** Empirically test all documented workaround URLs in researcher-prompt.md with actual Firecrawl scrape calls. Remove or caveat workarounds that don't work. ## Test Results ### 1. archive.ph **URL tested:** `https://archive.ph/newest/https://www.nytimes.com/2025/03/10/technology/ai-agents-future.html` **Params tried:** `proxy: "stealth"`, `proxy: "enhanced"`, `waitFor: 5000` **Result:** BROKEN. reCAPTCHA challenge page (status 429) on all attempts. Cloudflare protection blocks Firecrawl completely. Costs 5 credits per attempt (stealth/enhanced proxy). **Conclusion:** Remove as recommended workaround for Firecrawl-based research. ### 2. Wayback Machine **URLs tested:** - `web.archive.org/web/https://www.wsj.com/...` (shortcut format) -- 404, "not archived" - `web.archive.org/web/2024/https://www.nytimes.com/...` (year only) -- 404, "not archived" - `web.archive.org/web/2024*/https://arstechnica.com/...` (wildcard) -- returns calendar page - `web.archive.org/web/20240515120000/https://en.wikipedia.org/wiki/GPT-4` (exact timestamp) -- **SUCCESS**, full content (149K chars, had to be saved to file) - `web.archive.org/web/https://en.wikipedia.org/wiki/GPT-4` (shortcut, known-archived page) -- **SUCCESS**, 173K chars, status 200. Firecrawl followed redirect to `/web/20260203171033/...` (most recent snapshot). Wayback toolbar at top but full article content follows. - `web.archive.org/web/2/https://en.wikipedia.org/wiki/GPT-4` ("/2/" trick) -- **SUCCESS**, same result as shortcut format. **Result:** WORKS with simple shortcut format `/web/{URL}`. Initial tests with WSJ/NYTimes shortcut URLs returned 404 because those pages weren't archived, not because the format was wrong. Retesting with a known-archived page (Wikipedia GPT-4) confirmed the shortcut redirects to the most recent snapshot. Wayback Machine toolbar is prepended to extracted content but full page follows. 1 credit. **Conclusion:** Simplify guidance -- just use `/web/{URL}`. No need for exact timestamps. **Lesson:** Test URL format correctness against pages known to be archived. ### 3. xcancel.com (Twitter/X) **URLs tested:** - `xcancel.com/sama/status/1895533845455577446` (without proxy) -- anti-bot challenge (503) - `xcancel.com/sama/status/1895533845455577446` (stealth + waitFor) -- "Tweet not found" (404) - `xcancel.com/elikiowa/status/1880305905189106091` (stealth + waitFor) -- "Tweet not found" (404) - `xcancel.com/sama/status/1890816782836904000` (stealth + waitFor, verified valid ID) -- **SUCCESS** - `xcancel.com/OpenAI/status/1790070592011288831` (stealth + waitFor) -- "Tweet not found" (404) **Result:** WORKS but requires `proxy: "stealth"` + `waitFor: 5000` (5 credits). Initial tests with invalid tweet IDs gave misleading "Tweet not found" (404) results. Retests with verified valid tweet IDs confirmed: without proxy, anti-bot challenge blocks (503); with stealth + waitFor, full content returned (tweet text, replies, engagement metrics). Confirmed on both cached and fresh (`maxAge: 0`) URLs to rule out caching artifacts. **Conclusion:** Keep with updated guidance: requires stealth proxy + waitFor, costs 5 credits. **Lesson:** Always verify test inputs before concluding a service is broken. ### 4. freedium.cfd (Medium) **URL tested:** `https://freedium.cfd/https://towardsdatascience.com/rag-vs-fine-tuning-...` **Result:** DEAD. DNS resolution failure -- domain no longer resolves. **Conclusion:** Replace with freedium-mirror.cfd. ### 5. freedium-mirror.cfd (Medium -- replacement) **URL tested:** `https://freedium-mirror.cfd/https://towardsdatascience.com/rag-vs-fine-tuning-which-is-the-best-tool-to-boost-your-llm-application-94654b1eaba7` **Result:** SUCCESS. Full article content extracted (status 200, 1 credit, basic proxy). Complete Medium article with formatting, images, author info, tags. No paywall. **Conclusion:** Document as the working Medium paywall bypass. ### 6. Bluesky **URLs tested:** - `public.api.bsky.app/xrpc/app.bsky.feed.searchPosts?q=nix+nixos&limit=5` (raw API) -- 403 Forbidden - `bsky.app/profile/jay.bsky.team/post/3lihkfohpkk2t` (individual post, no waitFor) -- "Post not found" (SPA not rendered) - `bsky.app/profile/bsky.app` (profile page, waitFor: 5000) -- **SUCCESS**, massive content (full profile + dozens of recent posts) - `bsky.app/profile/atprotocol.dev` (profile page, waitFor: 5000) -- **SUCCESS**, 94K chars **Result:** WORKS but `bsky.app` is a JavaScript SPA -- `waitFor: 5000` is required for content to render. Without it, individual post pages return "Post not found". Profile pages with waitFor return extensive content including all recent posts. No proxy needed (1 credit). The raw API endpoint was unnecessary in the first place. **Conclusion:** Remove API-specific guidance. Document that bsky.app needs `waitFor: 5000` (SPA). ### 7. firecrawl_search site:reddit.com **Query tested:** `site:reddit.com best mechanical keyboard programming 2025` **Result:** SUCCESS. Returned 5 Reddit results with full URLs containing post IDs (e.g., `reddit.com/r/keyboards/comments/{postId}/...`). Post IDs extractable for Reddit MCP tools. Already validated in Iteration 11; re-confirmed here. **Conclusion:** Keep as documented. ## Changes Made **researcher-prompt.md "Accessing Restricted Content" section:** 1. **Paywalled articles:** Demoted archive.ph (CAPTCHA-blocked), promoted Wayback Machine to primary with exact-timestamp URL format requirement 2. **Twitter/X:** Added `proxy: "stealth"` + `waitFor: 5000` requirement for xcancel.com (5 credits). Previous guidance didn't mention these were needed. 3. **Medium:** Replaced freedium.cfd with freedium-mirror.cfd, removed archive.ph fallback 4. **Bluesky:** Removed raw API endpoint. Documented that bsky.app is a JavaScript SPA requiring `waitFor: 5000` for Firecrawl scrape. No proxy needed (1 credit). 5. **Reddit/GitHub/Discourse/LinkedIn:** Unchanged (not retested, already validated) ## Credit Cost Summary Total Firecrawl credits used: ~38 - archive.ph: 10 (2 attempts x 5 credits stealth/enhanced) - Wayback Machine: 4 (4 attempts x 1 credit basic) - xcancel.com: 16 (1 basic + 3 stealth attempts) - freedium.cfd: 0 (DNS failure, no credit charged) - freedium-mirror.cfd: 1 (basic proxy) - Bluesky: 8 (API 5 auto-stealth + 3 bsky.app scrapes x 1 basic) - firecrawl_search: ~0 (search credits separate) -
iteration-13.md 6.8 KB
# Iteration 13 -- Information Cascade Detection (Backlog #7) ## Type: Research Run (Empirical) + Engineering ## Question "How do journalists, fact-checkers, and researchers detect information cascades -- cases where many sources all trace back to one original, creating an illusion of independent corroboration?" ## Configuration - **Question type:** Technical - **Scope:** Broad (3 sonnet researchers, 2 rounds) - **Research angles:** 1. Journalism/fact-checking practitioner methods (Researcher A) 2. Academic citation cascade analysis (Researcher B) 3. Computational misinformation cascade detection at scale (Researcher C) ## Report `~/.claude/research/information-cascade-detection/report.md` ~180 lines, 37 footnotes, ~50+ sources across researchers. ## Key Research Findings **Three converging research traditions** all address cascade detection but have developed largely independently: 1. **Journalism/fact-checking:** "Trace to source" methodology. First Draft's provenance pillar, GIJN's "don't rely on other media" rule, IFCN two-source minimum. No org publishes an explicit cascade detection checklist -- it's tacit professional knowledge. The ivermectin/Rolling Stone case is the paradigmatic cascade case study. 2. **Academic bibliometrics:** Greenberg 2009 (BMJ) is foundational -- identified citation bias, amplification, and invention across 242 papers. Woozle effect quantified by Letrud & Hernes 2019 (76% of citing articles affirm debunked myths). Citogenesis (Wikipedia circular loops) well-documented. Chen et al. 2025 confirmed "telephone effect" computationally at 13M-pair scale. 3. **Computational social science:** Vosoughi et al. 2018 (Science) foundational -- false news 6x faster, distinct cascade topology. CrowdTangle shutdown (Aug 2024) left major monitoring gap. DisTrack and temporal graph approaches (TIDE-MARK) are current state-of-art. AI amplifies cascades via knowledge collapse and epistemic destabilization. **Practitioner heuristics (Round 2):** Specific red flags distilled: identical phrasing/errors propagated, tight temporal clustering, all sources citing same original, lack of local detail, same stock photos. Source TYPE diversity (documents + people + data) matters more than source COUNT. Intelligence analysis (Heuer ACH, ICD 203) has the most rigorous frameworks for source independence evaluation. **Observable signals (Round 2):** Content+structure+style similarity distinguishes cascade from independent reporting (Bar et al. 2012). Temporal clustering is a primary signal -- one-quarter of news stories reproduced within 4 minutes (Cage et al. 2025). Network topology differs systematically (broadcast/star = cascade, bushy/multi-origin = independent). ## Dogfooding Observations **What worked well:** - All three Round 1 angles produced highly complementary material with minimal overlap - Lead-dispatched follow-ups worked perfectly (8th consecutive success) - Round 2 targeted the right gaps -- practitioner heuristics and observable signals filled the key gaps from Round 1 - Researcher B had no Round 2 task (academic angle was thoroughly covered in Round 1) -- idle researcher management worked correctly **What went less well:** - Researcher A did not write a feedback file before shutdown (2 of 3 wrote feedback) - Researchers B and C both flagged Firecrawl returning oversized results for long academic articles (662KB for BMJ paper), with JSON format making partial reads difficult **Feedback themes (from B and C):** - Exa summaries were excellent for academic paper triage -- often sufficient without deep scraping - Firecrawl PDF parser failed on one conference paper (MCP error) - Suggestion: structured JSON extraction with targeted schema would be more token-efficient than full markdown for long academic papers - Intelligence analysis angle (Heuer ACH, OSINT corroboration vs replication) was an underexplored rich connection ## Changes ### 1. Added "Cascade Check" section to researcher-prompt.md (lines ~250-290) New section between Stopping Criteria and Handling Insufficient Evidence. Contains: - 5-step cascade check procedure (trace citation chain, check copied phrasing, assess source type diversity, apply "how do they know?" test, watch for wire service cascades) - Red flag list (temporal clustering, no local detail, single-origin collapse, same source type) - Instruction to downgrade confidence and flag in CROSS-SOURCE ANALYSIS when detected ### 2. Added cascade check reference to Step 3 (Reflect) Added seventh reflection question: "Cascade check (see below): Are my 'multiple sources' actually independent?" -- points researchers to the new section. ## Backlog Updates - **#7 (Information cascade detection):** RESOLVED. Research complete, heuristics engineered into researcher reflection step. - **#13 (researcher-prompt.md length):** File now ~541 lines (up from ~497). Added ~44 lines. No evidence of instruction-following degradation across 12 prior iterations, but worth monitoring. Next empirical run should verify researchers find and follow the cascade check. - **New observation:** Firecrawl returns oversized results for long academic HTML pages. Researchers suggest structured JSON extraction with schema. Not adding to backlog -- the Exa summary pattern is the recommended workaround and researchers are already following it. ## Post-Iteration Changes Discussion after the research run led to two design changes applied to SKILL.md: ### 1. Flexible round counts (SKILL.md) Replaced rigid per-scope round counts with ranges and convergence-based stopping: | Scope | Old | New | | ------------- | --- | --- | | Focused | 1 | 1-2 | | Broad | 2 | 2-3 | | Comprehensive | 2-3 | 3-4 | Added early stopping language: rounds are heuristics, not targets. Stop when citation convergence is reached. Removed hard phase-skip rules (e.g., "For Focused scope: skip to Phase 8") and replaced with convergence checks at each triage point. Any scope can exit early if findings have converged, and any scope can run additional rounds if gaps persist within its budget. ### 2. Verification architecture rethink (Backlog #16 promoted) Identified that the current verification model (reusing angle researchers in Round 3, Comprehensive only) has two problems: 1. **Angle researchers are specialists** -- pulling them off their angle to do cross-cutting verification wastes their accumulated context and expertise. 2. **Verification doesn't need a round** -- it's a parallel task, not a sequential phase. Promoted backlog #16 (dedicated verifier agents) to High Priority with full design: - Spawn lightweight, single-purpose verifier agents on-demand after any round - Verifiers run in parallel with angle researchers continuing their work - Lead assigns specific claims to verify based on triage - Available at any scope, not just Comprehensive - Implementation deferred to a future iteration
-
-
ITERATION-LOG.md 23.6 KB
# Deep Research Skill -- Self-Improvement Log ## Protocol This file drives an iterative improvement loop for the `deep-research-team` skill. Each iteration follows this cycle: 1. **Choose** -- Read this log. Review the backlog. Pick either a meta-research topic (theoretical improvement) or an empirical test (run the skill, observe friction). 2. **Research** -- Invoke `/deep-research-team` with the chosen question, or conduct targeted meta-research using web search. 3. **Observe** -- After research completes, analyze what happened. Record observations. 4. **Improve** -- Make targeted edits to skill files based on observations. Document changes. 5. **Checkpoint** -- Brief user check-in: what happened, what changed, what's next. 6. **Update** -- Refresh the backlog, set Current State for the next iteration. **Resuming:** If context was compacted or a new session started, read this file from the top. The Current State section tells you exactly where to pick up. **File organization:** This log file contains only compact summaries (5-8 lines each) in the Completed Iterations section. Full detail lives in `iterations/iteration-{NN}.md`. Workflow: 1. When starting an iteration, create `iterations/iteration-{NN}.md` and write detail there 2. As the iteration progresses, update the detail file (observations, changes, transcript analysis) 3. When the iteration is complete, add a compact summary to the Completed Iterations section below 4. Update Current State to point to the next iteration Never put full iteration detail in this file — it would grow too large for quick context loading. ## Current State - **Iteration:** 14 (ready to start) - **Phase:** choose next backlog item - **Next action:** Implement backlog #16 (dedicated verifier agents) -- design is clear, ready for engineering. Also a good candidate for empirical validation in the same iteration. - **Round counts updated** (post-Iteration 13): Focused 1-2, Broad 2-3, Comprehensive 3-4. Early stopping language added. Hard scope-gated phase transitions removed in favor of convergence-based decisions. - **Verification architecture rethink in progress** (backlog #16, promoted to high priority): Replace reuse of angle researchers for verification with dedicated on-demand verifier agents. Removes scope restriction (all scopes can verify), preserves angle researcher context, and allows verification to run in parallel with ongoing investigation. - **Cascade detection heuristics added** (Iteration 13). researcher-prompt.md at ~541 lines. ## Completed Iterations Full details in `iterations/iteration-{NN}.md`. Summaries here for context. ### Iteration 1 -- Compound Technical+Consumer (Empirical) **Question:** "Leading open-source LLM inference engines for 70B on consumer GPUs?" **Scope:** Broad (3 sonnet researchers, 2 rounds) **Report:** `~/.claude/research/llm-inference-engines-70b-consumer-gpu/report.md` **Key findings:** vLLM for multi-GPU, llama.cpp/Ollama for single-GPU. 25 sources. **Problems found:** Researchers didn't self-claim follow-up tasks (biggest friction), didn't persist Round 1 findings to disk, lead name not in spawn prompt. **Changes (5):** Restructured reporting flow, added self-claim instructions, updated critical rules, added lead name to spawn prompt, added mkdir permission. ### Iteration 2 -- Contested Question (Empirical) **Question:** "Should AI model weights be open-sourced?" **Scope:** Comprehensive (4 opus researchers, 3 rounds incl. verification) **Report:** `~/.claude/research/ai-model-weights-open-source/report.md` **Key findings:** 37-footnote steelmanned debate, 60+ sources. Cross-agent verification added genuine value (one claim SUPPORTED WITH NUANCE, one CONTESTED). **Problems found:** Self-claim fix from Iteration 1 still didn't work (confirmed pattern), inconsistent filename casing, no round stopping criteria, no assignment guidance. **Changes (4):** Replaced self-claim with lead-dispatched follow-ups (biggest architectural change), normalized filenames to lowercase, added round stopping criteria, added Round 2 assignment guidance (continuity vs fresh perspective). ### Iteration 3 -- Consumer Comparison (Empirical) **Question:** "Best mechanical keyboard for programming in 2026?" **Scope:** Broad (3 sonnet researchers, 2 rounds) **Report:** `~/.claude/research/best-mech-keyboard-programming-2026/report.md` **Key findings:** Keychron dominance at every price tier. 25 sources. Top Pick: Q5 Max. **Validated:** Lead-dispatched follow-ups worked perfectly (zero nudging), lowercase filenames worked, assignment guidance worked, non-AI domain worked fine with Exa. **Problems found:** Round 2 was useful-not-essential (suggests Focused might suffice for simple Consumer), Phase 1 skipped constraint gathering, idle notification noise. **Changes:** None needed -- all Iteration 2 changes validated successfully. **Cumulative skill improvements across 3 iterations:** 9 changes total. Architecture now stable: lead-dispatched follow-ups, file persistence, lowercase filenames, round stopping criteria, assignment guidance all working. Self-claim problem fully resolved. ### Iteration 4 -- Search Tool Audit + Access Workarounds (Meta-Research) **Question:** "What are the most effective web research tool combinations for deep research?" **Scope:** Broad (3 sonnet researchers, 2 rounds) **Report:** `~/.claude/research/web-research-tool-effectiveness/report.md` **Key findings:** Exa and Firecrawl MCP tools expose only a subset of their full API capabilities. Switched Exa to hosted endpoint, added academic search strategies, content access workarounds, and Firecrawl advanced features to researcher prompt. **Validated:** Lead-dispatched follow-ups (5th consecutive success), meta-research through the team skill is viable and effective. **Problems found:** None during run. Post-iteration discovered `contextMaxCharacters` bug suppressing Exa summaries/highlights. **Changes (6 to researcher-prompt.md + major post-iteration infrastructure):** Expanded Available Tools, added Firecrawl Search Operators, Academic Search, Content Extraction advanced features, Accessing Restricted Content sections. Post-iteration: switched Exa MCP to hosted endpoint, reduced tools from 5 to 2, added `enableSummary` + `textMaxCharacters: 1` default pattern, added category restriction table. Resolved backlog #18. **Cumulative skill improvements across 4 iterations:** 15 changes. Major post-iteration infrastructure overhaul to Exa tooling. ### Iteration 5 -- Scientific/Health Empirical (IF + Cognitive Performance) **Question:** "Does intermittent fasting improve cognitive performance?" **Scope:** Broad (3 sonnet researchers, 2 rounds) **Report:** `~/.claude/research/intermittent-fasting-cognitive-performance/report.md` **Key findings:** Acute fasting is cognitively neutral (meta-analysis of 63 studies). 26 footnotes. Evidence pyramid template validated. Calibrated confidence language worked well. **Validated:** Exa summary pattern (32/32 compliant), lead-dispatched follow-ups (6th success), academic source quality strong. **Changes (3):** Implemented shutdown feedback reports (researcher-prompt.md + SKILL.md Phase 10), added `scripts/analyze-transcripts.py` for JSONL analysis. Resolved backlog #12, #23. **Cumulative skill improvements across 5 iterations:** 18 changes. All synthesis templates validated except Emerging/Frontier and Opinion/Sentiment. ### Iteration 6 -- Emerging/Frontier Empirical (LLM Reasoning) **Question:** "What are the current approaches to LLM reasoning?" **Scope:** Comprehensive (4 opus researchers, 3 rounds incl. verification) **Report:** `~/.claude/research/llm-reasoning-approaches/report.md` **Key findings:** ~400 lines, 55 footnotes. RLVR as dominant paradigm shift. PRM vs ORM conflict resolved (both correct in different regimes). CoT faithfulness 2.3% CONTESTED (single study), <20% verbalization SUPPORTED WITH NUANCE. **Validated:** Emerging/Frontier template, shutdown feedback reports (all 4 wrote files), category restriction table (tweet, news, personal site all worked), Exa compliance 49/49 (100%). **Changes (0):** No skill file changes. Identified backlog #26 (arxiv guidance) and #27 (verification verdict expansion) for quick implementation. **Cumulative skill improvements across 6 iterations:** 18 changes. Only Opinion/Sentiment template remains untested. ### Iteration 7 -- Agent Teams Docs Audit (Meta) **Scope:** No research run -- audit skill against official Anthropic documentation. **Key findings:** Architecture fully validated against Anthropic's docs and their own research system. Hub-and-spoke, lead-dispatched, filesystem output, verification pattern all confirmed. Permission inheritance confirmed (teammates inherit lead's settings). Anthropic uses Opus lead + Sonnet subagents with 90.2% improvement over single-agent. **Changes (2):** arxiv scraping guidance (#26), SUPPORTED WITH NUANCE verdict (#27). ### Iteration 8 -- Scope Calibration Refinement (Meta) **Scope:** No research run -- analyze 7 empirical data points. **Key findings:** Current defaults correct 6/6 times. Consumer is only type with de-escalation evidence. Report length is scope-driven, not type-driven. **Changes (3):** Scope modifiers and empirical output expectations added to question-types.md, SKILL.md Phase 2 updated. **Cumulative skill improvements across 8 iterations:** 21 changes (counting Iteration 7's 2). ### Iteration 9 -- Opinion/Sentiment + Mixed Model Config (Empirical) **Question:** "What do developers actually think about Nix and NixOS in 2026?" **Scope:** Comprehensive (4 sonnet researchers, 2 rounds -- no Round 3 verification needed) **Report:** `~/.claude/research/developer-opinions-nix-nixos-2026/report.md` **Key findings:** ~370 lines, 50 footnotes, 60+ sources. Developer opinion defined by paradox: fastest community growth ever (30% YoY) alongside deepest fragmentation ever (governance crisis, three forks). Pragmatic "devshells only" is the dominant successful adoption pattern. Wrapper tools (Devbox 11K stars, devenv 6K, Flox 3.7K) work for 80% case, break at edges. **Tested:** Opinion/Sentiment template (last untested), Sonnet researchers in Comprehensive scope (backlog #7). Both validated successfully. **Changes (3):** Updated researcher-prompt.md Reddit section (restrict_sr=on, prefer /top/ over /search/, note curl permission cost), added GitHub data guidance (Exa code context + gh api, no curl), added Discourse JSON search API fallback. **Backlog items resolved:** #7 (mixed model config -- Sonnet validated for Comprehensive). **Cumulative skill improvements across 9 iterations:** 24 changes. All 7 synthesis templates tested. Sonnet validated as viable Comprehensive override (Opus kept as default). ### Iteration 10 -- Reddit MCP Server Integration (Engineering) **Scope:** No research run -- engineering integration. **Key findings:** Fresh evaluation confirmed jordanburke/reddit-mcp-server as best option (v1.2.1, actively maintained, 9 read tools + search, anonymous mode). Hawstein/mcp-server-reddit stale (no commits in 10 months, no search). New entrants too immature. **Changes (4):** Added reddit-mcp-server to 1MCP config (anonymous mode), added Reddit ToolSearch to researcher setup, added Reddit tool block to Available Tools section, replaced curl-based Reddit access with MCP tool guidance in researcher-prompt.md. **Backlog items resolved:** #8 (Reddit MCP server integration). **Cumulative skill improvements across 10 iterations:** 28 changes. Zero curl dependencies in researcher workflows (Reddit MCP + GitHub gh api). ### Iteration 11 -- Reddit MCP Empirical Validation **Question:** "What is the best community-favorite content about the TV show Dark?" **Scope:** Focused (2 sonnet researchers, 1 round) **Report:** `~/.claude/research/dark-tv-show-community-content/report.md` **Key findings:** 24-source report across 5 content categories (video essays, podcasts, Reddit mega-threads, visual companion tools, philosophical/academic writing). OneTake and Think Story top YouTube; r/DarK mega-threads are canonical community references; academic philosophy papers exist from Springer and Blackwell series. **Validated:** Reddit MCP integration end-to-end (9/9 goals passed). Both researchers loaded Reddit tools via ToolSearch, used Firecrawl search for discovery + Reddit MCP for extraction, zero curl usage, zero permission prompts from Reddit. **Problems found:** `firecrawl_search` with subreddit-scoped `site:reddit.com/r/DarK` returned generic results (broad `site:reddit.com` worked better). `get_top_posts` returns memes/meta for entertainment subreddits, insufficient alone for substantive threads. **Changes (2):** Updated researcher-prompt.md Reddit section: added Firecrawl search scoping note (broad > subreddit-specific), clarified get_top_posts limitation for entertainment subs. **Cumulative skill improvements across 11 iterations:** 33 changes (30 from iteration + 3 post-iteration housekeeping: generalized transcript script, filesystem path support, simplified feedback template). Reddit MCP fully validated. ### Iteration 12 -- Firecrawl Scrapability Testing (Engineering) **Scope:** No research run -- empirical testing of all documented content workaround URLs. **Key findings:** archive.ph CAPTCHA-blocked (429) even with stealth/enhanced proxy. freedium.cfd DNS dead (replaced by freedium-mirror.cfd, works great). Bluesky bsky.app is a JavaScript SPA requiring `waitFor: 5000` (no proxy, 1 credit); raw API unnecessary. xcancel.com works but needs `proxy: "stealth"` + `waitFor: 5000` (5 credits). Wayback Machine shortcut `/web/{URL}` works -- Firecrawl follows redirect to latest snapshot (initial 404s were unarchived pages, not format errors). Lessons: validate test inputs (xcancel), test against known-good data (Wayback Machine). **Changes (5 to researcher-prompt.md):** Demoted archive.ph, simplified Wayback Machine to shortcut format, added xcancel proxy requirements, replaced freedium.cfd with freedium-mirror.cfd, documented Bluesky SPA waitFor requirement. **Backlog items resolved:** #12 (Firecrawl scrapability testing). **Cumulative skill improvements across 12 iterations:** 38 changes. All content access workarounds empirically validated. Researchers will no longer waste calls on dead endpoints. ### Iteration 13 -- Information Cascade Detection (Empirical + Engineering) **Question:** "How do journalists, fact-checkers, and researchers detect information cascades?" **Scope:** Broad (3 sonnet researchers, 2 rounds) **Report:** `~/.claude/research/information-cascade-detection/report.md` **Key findings:** Three converging research traditions (journalism, bibliometrics, computational social science) all address cascade detection independently. Greenberg 2009 foundational for citation distortion taxonomy. Woozle effect quantified (76% affirm debunked myths). Practitioner heuristics are tacit knowledge -- no org publishes explicit cascade checklist. Intelligence analysis (Heuer ACH, ICD 203) has most rigorous source independence frameworks. Composite signals (temporal clustering, textual overlap, network topology, metadata) most reliable. **Validated:** Lead-dispatched follow-ups (8th success), idle researcher management (B had no Round 2 task), all operational patterns stable. **Changes (2):** Added 5-step "Cascade Check" section to researcher-prompt.md with red flags and procedure. Added cascade check reference to Step 3 (Reflect) questions. **Backlog items resolved:** #7 (information cascade detection). **Cumulative skill improvements across 13 iterations:** 40 changes. Researchers now have explicit cascade detection heuristics in the reflection step. researcher-prompt.md at ~541 lines. ## Backlog Prioritized list of things to investigate or test. Re-prioritized each iteration. ### High Priority 16. **[Engineering] Dedicated verifier agents** -- Restructure verification from reusing angle researchers to spawning lightweight, single-purpose verifier agents on-demand. Current approach wastes angle researchers' accumulated context and blocks investigation flow. **New design:** Angle researchers (A, B, C) own their angles for the full session and maintain context continuity. When the lead's triage identifies a high-impact single-source claim at any point, it spawns a dedicated verifier (sonnet, small context, focused task). The verifier receives only the claim + source URL, searches for independent evidence, reports a verdict (SUPPORTED / SUPPORTED WITH NUANCE / CONTESTED / UNCHANGED), and shuts down. Verifiers run in parallel with ongoing angle work -- no separate phase or round needed. **What this fixes:** (1) Removes the Comprehensive-only scope restriction -- any scope can verify because it's just a lightweight parallel agent. (2) Preserves angle researcher expertise -- they stay on their domain instead of context-switching to verify unrelated claims. (3) Eliminates the Phase 7 bottleneck -- verification no longer requires a dedicated round. (4) Cleaner separation of concerns -- investigation vs verification are different tasks that benefit from different agent configurations. **Implementation:** Refactor Phase 7 out of the sequential phase structure. Add verification dispatch logic to Phase 5 triage (and any subsequent triage). Update Effort Calibration table to show verification as available to all scopes. Update researcher spawn prompts (verifiers need a different, simpler prompt than angle researchers). Test empirically. 17. **[Research] Confidence assessment methodology** -- The current triage confidence levels (HIGH/MEDIUM/LOW) are based primarily on independent source count. Source count is a crude proxy -- a single well-designed RCT can warrant more confidence than five blog posts citing each other. Source *quality* matters but is hard to assess systematically. Worth a meta-research run to explore: How do intelligence analysts, systematic reviewers, and epistemologists think about evidence quality? What frameworks exist beyond source counting? (Heuer's ACH and ICD 203 from iteration 13 are starting points.) Goal: replace or augment the current source-count heuristic with something that accounts for source quality, methodology strength, and independence more holistically. 5. **[Engineering] Permission setup / onboarding flow** -- Audit all permissions needed for the skill to run without prompts (mkdir, Write to research dir, researcher Bash calls for sleep/retry/curl). Build a setup check at skill invocation that verifies permissions exist and offers to add them. Permission inheritance now confirmed (Iteration 7 audit): teammates inherit the lead's permission settings. Iteration 9 confirmed curl is the biggest permission friction source. Key for making the skill shareable beyond Malo. Not a research task -- pure engineering. 6. ~~**[Meta] Scope calibration guidance**~~ -- **Resolved (Iteration 8).** Added scope modifiers (de-escalation/escalation signals per type) and empirical output expectations to `references/question-types.md`. All 7 defaults validated. Consumer is the only type with observed de-escalation opportunity. Phase 2 updated to reference modifiers. 7. ~~**[Empirical] Mixed model config for Comprehensive**~~ -- **Resolved (Iteration 9).** Tested Opus lead + 4 Sonnet researchers on Comprehensive scope (Opinion/Sentiment question). Sonnet validated as viable -- researchers completed both rounds, followed methodology, wrote structured findings and feedback. However, Opus kept as Comprehensive default: one successful run on a less demanding question type doesn't prove Sonnet handles the hardest Comprehensive questions (e.g., contested scientific topics with complex verification). SKILL.md notes Sonnet as an available override when cost matters. 8. ~~**[Engineering] Reddit MCP server integration**~~ -- **Resolved (Iteration 10).** Integrated `jordanburke/reddit-mcp-server` via 1MCP with anonymous mode (~10 rpm). Researcher methodology updated to use native MCP tools (`search_reddit`, `get_top_posts`, `get_post_comments`, etc.) instead of curl. Eliminates ~8-10 permission prompts per research run. Zero curl dependencies remain in researcher workflows. Can upgrade to app-only auth (60 rpm) later if rate limits become an issue -- requires Reddit API application approval. ### Medium Priority 7. ~~**[Meta] Information cascade detection**~~ -- **Resolved (Iteration 13).** Researched across three traditions (journalism, bibliometrics, computational social science). Distilled findings into a 5-step "Cascade Check" section in researcher-prompt.md with red flags and procedure. Researchers now check source independence during Step 3 (Reflect). 8. ~~**[Meta] Report length targets per question type**~~ -- **Resolved (Iteration 8).** Added empirical output expectations table to `references/question-types.md` (Focused ~150-200 lines/10-15 sources, Broad ~200-300/20-30, Comprehensive ~350-400+/40-60). Presented as descriptive ranges, not targets. 9. **[Meta] Optimal angle decomposition** -- Research strategies for generating maximally complementary research angles. Could improve Phase 1 planning. 10. **[Meta] Source diversity vs quality tradeoffs** -- When should researchers prefer breadth over authority? Could refine stopping criteria. 11. **[Engineering] Exa `deep` search type via MCP** -- Exa's API supports `type: "deep"` which does automatic query expansion + parallel search + re-ranking ($0.015/search). This would replace the manual "generate 3-5 query variations" pattern with a single call + `additionalQueries`. Currently blocked: the hosted MCP endpoint (`mcp.exa.ai`) only exposes auto/fast/neural in the type enum. Monitor for when Exa adds deep to MCP. 12. ~~**[Engineering] Firecrawl scrapability testing**~~ -- **Resolved (Iteration 12).** Tested all 7 documented workarounds with actual Firecrawl calls. 2 broken (archive.ph CAPTCHA, freedium.cfd DNS dead), 1 replaced (freedium-mirror.cfd works), 1 simplified (Wayback Machine shortcut `/web/{URL}` works -- initial 404s were unarchived pages), 2 updated with required params (xcancel.com needs stealth proxy, Bluesky bsky.app needs waitFor for SPA). Docs updated to prevent researchers from wasting calls on dead endpoints. 13. **[Meta] Researcher-prompt.md length monitoring** -- File is ~438 lines. Researchers successfully followed all guidance in Iteration 6 (100% Exa compliance, feedback reports written). No evidence of instruction-following degradation. Continue monitoring but deprioritized -- no splitting needed yet. 14. **[Meta/Engineering] Exa vs dedicated academic search** -- Iteration 6 confirmed strong academic search via `category: "research paper"` for CS/ML (arxiv, NeurIPS, ICLR, ACL, ICML sources found). Iterations 5+6 together show the current setup works well for both biomedical and CS domains. Deprioritized unless a specific academic search failure occurs. 15. **[Engineering] Skill setup / environment bootstrapping** -- Investigate using skill `!` bang patterns or a setup phase to automatically verify/dump environment state at skill invocation. Related to backlog #5 (permission setup) but broader. ### Lower Priority 13. **[Meta] Self-reflection patterns in multi-agent research** -- How does extended thinking interact with the investigation loop? 14. **[Meta] Synthesis quality patterns** -- What makes a good research synthesis? Compare our output against professional research reports. 15. **[Meta] Consumer Phase 1 constraint gathering** -- Should Consumer questions always get a quick "Any constraints?" even when the question seems clear? Minor UX improvement. 16. **[Speculative] Researcher compaction robustness** -- Whether researchers need explicit compaction handling. Assessment: low risk. 6 iterations without compaction issues. -
RESEARCH.md 13.3 KB
# Deep Research Team v2 — Research Findings Research conducted 2026-02-06 using deep mode (4 Opus researchers, 2 rounds). 50+ sources across academic papers, production systems, and practitioner reports. ## Core Question What is the optimal architecture for an LLM-agent-team-based research capability that handles the full spectrum of research questions — from factual lookups to subjective opinion synthesis to contested/controversial topics? --- ## Finding 1: Architecture-Task Alignment Matters More Than Agent Count Kim et al. (Google/MIT/DeepMind, Dec 2025) tested 180 configurations across 5 architectures, 3 LLM families, and 4 benchmarks. Matching coordination topology to task structure is the critical variable, not scaling agent count[^1]. Centralized (hub-and-spoke) coordination improved parallelizable tasks by 80.8%, while every multi-agent variant degraded sequential reasoning by 39-70%. Communication overhead grows super-linearly (exponent 1.724). Effective team sizes: ~3-4 agents before costs outweigh benefits. Anthropic's production research system confirms: orchestrator-worker (Opus 4 lead + Sonnet 4 subagents) outperformed single-agent Opus 4 by 90.2% on research tasks[^2]. Key insight: **token usage alone explains 80% of performance variance** on research benchmarks. Multi-agent architectures work primarily by scaling token usage across parallel context windows. Context dilution research (7+ studies: Stanford, Google, Meta, Microsoft, NVIDIA, Adobe, Chroma) shows more tokens in a single context _degrades_ performance (13.9-85% accuracy drop as length increases)[^3][^4][^5]. Multi-agent solves this by distributing tokens across separate windows — the fundamental architectural justification for agent teams. **Confidence: HIGH** (Kim et al. + Anthropic + 7+ context dilution studies) ## Finding 2: Classify-Then-Route Pattern Is Proven But Underdeployed Multiple working implementations demonstrate that classifying question type before selecting strategy improves accuracy and efficiency: n8n Adaptive RAG (4-type classification), LangGraph Adaptive RAG, Meilisearch (4 types with fallback escalation), Kore.ai Adaptive-RAG[^6][^7][^8][^9]. Select-then-Route (EMNLP 2025) provides rigorous validation[^10]. Yet no major commercial tool implements this. OpenAI Deep Research uses a single pipeline[^11]. Perplexity requires manual mode control[^12]. Google has implicit sensitivity through E-E-A-T but no explicit classification[^13]. All implementations converge on 4-5 categories. For compound questions, query decomposition handles this (Haystack, NVIDIA, AAAI 2023)[^14][^15][^16]. **Confidence: HIGH** (4+ implementations + domain evidence; no major tool does this = differentiator) ## Finding 3: Evidence Hierarchies Must Vary by Question Type Four independent domains confirm: - **Evidence-based medicine**: Different "levels of evidence" for different question types (RCTs for therapy, cohort studies for prognosis, blind comparisons for diagnostics)[^17][^18] - **Library science**: Four question levels with escalating research requirements[^19] - **Consumer evaluation**: Expert testing + user satisfaction data as primary evidence (opposite of medical hierarchy)[^20][^21] - **Fact-checking**: Core methodology focuses exclusively on verifiable claims[^22] ACRL Framework: "Authority is Constructed and Contextual" — source credibility depends on the information need[^23]. Reddit is noise for health questions but signal for community opinion[^24]. **Confidence: HIGH** (4 independent professional domains) ### Proposed Question-Type Taxonomy | Question Type | Evidence Priority | Source Hierarchy | Synthesis Approach | | --------------------------- | ---------------------------------------- | ---------------------------------------------------------------- | ----------------------------------------------------- | | **Factual/verifiable** | Government databases, official records | Authoritative ref > primary sources > news | Aggregate reasoning quality (AoR) | | **Scientific/health** | Systematic reviews, RCTs, cohort studies | Peer-reviewed > gov health > expert > anecdotal | Evidence pyramid with study quality flags | | **Consumer/recommendation** | Expert testing + user satisfaction data | Expert reviewers > enthusiast forums > specs > ads | Ranked recommendations (top/runner-up/budget) | | **Technical/comparative** | Documentation, benchmarks, adoption data | Official docs > code examples > expert blogs > SO/Reddit | Working implementations + benchmarks | | **Opinion/sentiment** | Community forums, surveys, social media | Reddit/forums > blogs > news > academic surveys | Map distribution of views, assess representativeness | | **Contested/controversial** | Multiple perspectives required | Academic > think tanks > news analysis > advocacy (label biases) | Structure debate (MODS pattern), never single-verdict | | **Emerging/frontier** | Recent sources only | Conference papers > expert blogs > vendor announcements | Findings with recency/instability caveats | ## Finding 4: Dynamic Roles and Complementarity Trump Fixed Assignments Multiple 2025 papers converge on dynamic role assignment > fixed roles: MetaGen (arXiv Jan 2026), MLC (ACL 2025), OSC (EMNLP 2025), AMAS (EMNLP Industry 2025), Heterogeneous Swarms[^25][^26][^27][^28][^29]. Production systems use **topic-angle decomposition** (diversify by sub-question) not **functional-role decomposition** (searcher/critic/synthesizer)[^2][^30]. Practical pattern: topic decomposition during exploration, functional stages in the pipeline. Complementarity > diversity (NeurIPS Workshop 2025)[^31]. Implicit consensus with partial diversity outperforms forced agreement (EMNLP 2025)[^32]. DMoA (ICLR 2025): optimal diversity-consistency balance varies by task — factual favors consistency, exploratory favors diversity[^33]. **Confidence: HIGH** (architecture pattern); **MEDIUM** (complementarity measurement in practice) ## Finding 5: Three-Stage Effort Calibration **Pre-research classification**: Question type, structural complexity, user intent. Anthropic's production rules: 1 agent/3-10 calls (simple), 2-4 subagents/10-15 calls (comparison), 10+ subagents (complex research)[^2]. **Mid-research adaptation**: Monitor source novelty, contradiction density, coverage adequacy. Dynamic subagent spawning only if warranted. **Stopping criteria**: Coverage-based (2+ sources per claim, citation convergence) + budget caps (OpenAI: 30-60 searches, 120-150 fetches, 20-30 min)[^34]. Task complexity — not instruction specificity — drives failure (Tech Policy Institute, N=575): complexity reduces success by 24-37pp, instruction specificity only 8.6pp (not significant)[^35]. Library science "reference interview" has AI analogs: STaR-GATE (72% preference), Ask-before-Plan (97.3% planning accuracy)[^36][^37]. Safest design: auto-classify, present plan, allow override. **Confidence: HIGH** (multi-source convergence) ## Finding 6: Synthesis Strategy Should Vary Along the Fact-Opinion Spectrum No current tool adapts synthesis by question type[^38][^39]. - **Factual**: Aggregation of Reasoning > majority voting (1-12% improvement)[^40] - **Contested/debatable**: MODS framework (38-59% better coverage/balance) structures debate[^41]; Plurals (CHI 2025) creates deliberative panels[^42] - **Confidence language**: Medium verbalized uncertainty optimizes trust (Xu et al., N=156)[^43]; Kent's WEPs partially usable by LLMs but need calibration (Tang et al., 2026)[^44][^45] - **Recommendations**: Biased/opinionated AI improves user decision-making (Lai et al., N=2500)[^46]; default to recommendation with dissent, not false balance **Confidence: HIGH** (peer-reviewed experimental studies) ## Finding 7: IC Tradecraft Provides Battle-Tested Analytical Frameworks CIA Structured Analytic Techniques[^47][^48]: - **ACH** (Analysis of Competing Hypotheses) — array evidence against hypotheses, focus on disproving - **Key Assumptions Check** — surface and challenge unstated assumptions - **Quality of Information Check** — systematic source reliability assessment Toulmin argumentation model (Data → Warrant → Claim + Qualifier + Rebuttal) forces explicit identification of inferential steps[^49][^50]. Useful analytically for lead's synthesis reasoning. **Confidence: HIGH** (established government/academic frameworks) --- ## Key Conflicts Resolved 1. **"More tokens help" vs "more tokens hurt"**: Both true. Multi-agent distributes tokens across separate context windows (helps) rather than stuffing one window (hurts). 2. **Kim et al. diminishing returns vs Anthropic 90.2% gains**: Research is parallelizable (where Kim et al. found the largest gains). The 45% saturation threshold applies to sequential tasks. 3. **Centralized vs decentralized**: Anthropic's "centralized" is actually centralized-independent hybrid — orchestrator assigns/synthesizes, subagents work independently during exploration. Captures decentralized benefits with centralized quality control. --- ## Key Gaps (Unresolved) - No head-to-head comparison of question-type-adaptive vs one-size-fits-all research agents - No automated method for detecting information cascades (many sources citing one original) - No formal "fact-opinion spectrum" framework for routing synthesis strategies - Source diversity vs quality tradeoffs remain under-theorized - Cost-effectiveness curves for agent scaling are unpublished - Self-reflection + extended thinking interaction with multi-agent coordination understudied --- ## Design Implications for Skill Redesign 1. **Add question-type classification as Phase 0** — lightweight LLM classification into 4-5 types, with decomposition for compound questions 2. **Replace static quick/standard/deep with dynamic effort calibration** — auto-classify complexity, suggest tier, allow override, adapt mid-research 3. **Vary source evaluation by question type** — different source hierarchies per type 4. **Vary synthesis strategy by question type** — AoR for factual, MODS for contested, recommendations for consumer, opinion mapping for sentiment 5. **Use calibrated confidence language** — Kent-style WEPs with explicit probability ranges 6. **Keep centralized-independent hybrid** — validated by production systems and academic research 7. **Keep topic-angle decomposition** — validated over functional-role decomposition for research 8. **Add file-based persistence** — lead writes report + state files, researchers write findings to disk, enables cross-session resumability 9. **Smarter clarification** — skip for simple factual, decompose for compound, plan-present for complex --- ## Sources [^1]: Kim et al., "Towards a Science of Scaling Agent Systems," arXiv:2512.08296, Dec 2025 [^2]: Anthropic Engineering, "How we built our multi-agent research system," Jun 2025 [^3]: Liu et al., "Lost in the Middle," Stanford/Meta, TACL 2024 [^4]: "Context Length Alone Hurts," arXiv, Oct 2025 [^5]: Chroma, "Context Rot," Jul 2025 [^6]: n8n, "Adaptive RAG with Query Classification" [^7]: LangChain, "LangGraph Adaptive RAG" [^8]: Meilisearch, "Adaptive RAG Explained," Sep 2025 [^9]: Kore.ai, "Adaptive-RAG," 2024 [^10]: "Select-then-Route," EMNLP 2025 Industry Track [^11]: OpenAI, "Deep Research API Introduction," Jun 2025 [^12]: Data Studios, "Perplexity AI Prompting Techniques," Dec 2025 [^13]: Agenxus, "Google AI Overviews Source Prioritization," Nov 2025 [^14]: Haystack, "Advanced RAG: Query Decomposition," Sep 2024 [^15]: NVIDIA, "Query Decomposition for RAG Blueprint" [^16]: "Question Decomposition Tree," AAAI 2023 [^17]: NHMRC Evidence Hierarchy [^18]: Icahn School of Medicine, "Evidence Based Medicine" [^19]: Hepler & Horalek, "Introduction to Library and Information Science" [^20]: Wirecutter, "Anatomy of a Guide," May 2024 [^21]: Consumer Reports, "Rating Methods" [^22]: Ballotpedia, "Methodologies of Fact-Checking" [^23]: ACRL Framework for Information Literacy [^24]: White, "Reddit as Analogy for Scholarly Publishing," 2019 [^25]: MetaGen, arXiv, Jan 2026 [^26]: MLC, ACL 2025 [^27]: OSC, EMNLP 2025 [^28]: AMAS, EMNLP Industry 2025 [^29]: Heterogeneous Swarms, OpenReview 2025 [^30]: O-Researcher, arXiv:2601.03743, Jan 2026 [^31]: Zhang, "Mixture of Complementary Agents," NeurIPS Workshop 2025 [^32]: Wu & Ito, "The Hidden Strength of Disagreement," EMNLP 2025 [^33]: "Balancing Act: DMoA," ICLR 2025 [^34]: PromptLayer, "How OpenAI's Deep Research Works," Oct 2025 [^35]: Lovin & Wallsten, Tech Policy Institute, Nov 2025 [^36]: STaR-GATE, arXiv, 2024 [^37]: Ask-before-Plan, 2024 [^38]: 7minute.ai, "OpenAI vs Google Deep Research," 2025 [^39]: Venkit et al., "Search Engines in an AI Era," 2024 [^40]: Sun et al., "AoR," LREC-COLING 2024 [^41]: Xiao et al., "MODS," NAACL 2025 [^42]: Bakker et al., "Plurals," CHI 2025 [^43]: Xu et al., "Confronting Verbalized Uncertainty," IJHCS 2025 [^44]: Kent, "Words of Estimative Probability," CIA/CSI [^45]: Tang et al., npj Complexity 2026 [^46]: Lai et al., arXiv:2508.09297, 2025 [^47]: CIA Tradecraft Primer, 2009 [^48]: RAND, RR1408, 2016 [^49]: Verheij, "The Toulmin Argument Model in AI," 2009 [^50]: Freedman et al., "Argumentative LLMs," AAAI 2025
-
-
references
-
templates
-
consumer.md 916 B
# Synthesis Template: Consumer Ranked recommendations with methodology. ```markdown ## Executive Summary [2-3 sentences: key finding, confidence level, scope] ## Recommendations ### Top Pick: [Product] [Why, with evidence from expert testing] ### Runner-Up: [Product] [Why, differentiation from top pick] ### Budget Pick: [Product] [Trade-offs at this price point] ### Best for {Use Case}: [Product] [Specific use-case fit] ## Methodology [Criteria used, sources consulted, testing methodology of source reviewers] ## Confidence Assessment [Use calibrated language from the Confidence table in SKILL.md] - Almost certain (93-99%): [findings] - Highly likely (80-92%): [findings] - Likely (63-79%): [findings] - Roughly even / Unlikely: [findings, if any] ## Limitations & Gaps [What remains uncertain, missing, or contested] ## Sources [^1]: Author/Org. "Title". Publication. URL [^2]: ... ``` -
contested.md 994 B
# Synthesis Template: Contested Structured debate -- NEVER resolve with a verdict. ```markdown ## Executive Summary [2-3 sentences: key finding, confidence level, scope] ## Perspectives ### Position A: [Label] [Strongest case for this position, steelmanned] ### Position B: [Label] [Strongest case for this position, steelmanned] ## Points of Agreement [Where the sides converge] ## Key Disagreements [Where they diverge and why -- identify factual vs. values-based disagreements] ## Evidence Quality by Position [Assess the strength of evidence supporting each side -- this is NOT a verdict] ## Confidence Assessment [Use calibrated language from the Confidence table in SKILL.md] - Almost certain (93-99%): [findings] - Highly likely (80-92%): [findings] - Likely (63-79%): [findings] - Roughly even / Unlikely: [findings, if any] ## Limitations & Gaps [What remains uncertain, missing, or contested] ## Sources [^1]: Author/Org. "Title". Publication. URL [^2]: ... ``` -
emerging-frontier.md 899 B
# Synthesis Template: Emerging/Frontier State of play with instability warning. ```markdown ## Executive Summary [2-3 sentences: key finding, confidence level, scope] ## Current State (as of {date}) [What is known now, with explicit date markers on key facts] ## Trajectory _Note: The following is projection based on current trends, not established fact._ [Where things appear to be heading] ## Instability Warning [What is most likely to change and why -- flag specific claims with short shelf life] ## Confidence Assessment [Use calibrated language from the Confidence table in SKILL.md] - Almost certain (93-99%): [findings] - Highly likely (80-92%): [findings] - Likely (63-79%): [findings] - Roughly even / Unlikely: [findings, if any] ## Limitations & Gaps [What remains uncertain, missing, or contested] ## Sources [^1]: Author/Org. "Title". Publication. URL [^2]: ... ``` -
factual.md 752 B
# Synthesis Template: Factual Clear answer with evidence weight. ```markdown ## Executive Summary [2-3 sentences: key finding, confidence level, scope] ## Verdict [Direct answer to the question with confidence level] ## Supporting Evidence ### [Claim 1] [Evidence with inline citations, weighted by source authority] ### [Claim 2] [Continue same pattern] ## Confidence Assessment [Use calibrated language from the Confidence table in SKILL.md] - Almost certain (93-99%): [findings] - Highly likely (80-92%): [findings] - Likely (63-79%): [findings] - Roughly even / Unlikely: [findings, if any] ## Limitations & Gaps [What remains uncertain, missing, or contested] ## Sources [^1]: Author/Org. "Title". Publication. URL [^2]: ... ``` -
opinion-sentiment.md 893 B
# Synthesis Template: Opinion/Sentiment Distribution mapping with representativeness assessment. ```markdown ## Executive Summary [2-3 sentences: key finding, confidence level, scope] ## Distribution of Views ### Majority View (~X%) [What most people think and why] ### Minority View (~X%) [Dissenting perspective and reasoning] ### Outlier Positions [Fringe views worth noting] ## Representativeness Assessment [How representative is the source sample? What communities were/weren't covered?] ## Confidence Assessment [Use calibrated language from the Confidence table in SKILL.md] - Almost certain (93-99%): [findings] - Highly likely (80-92%): [findings] - Likely (63-79%): [findings] - Roughly even / Unlikely: [findings, if any] ## Limitations & Gaps [What remains uncertain, missing, or contested] ## Sources [^1]: Author/Org. "Title". Publication. URL [^2]: ... ``` -
scientific-health.md 1012 B
# Synthesis Template: Scientific/Health Evidence pyramid with study quality assessment. ```markdown ## Executive Summary [2-3 sentences: key finding, confidence level, scope] ## Evidence Summary ### Systematic Reviews & Meta-Analyses [Findings from highest-quality evidence] ### Randomized Controlled Trials [Key RCT findings] ### Observational Studies [Cohort/case-control findings] ## Study Quality Notes [Sample sizes, methodology concerns, funding disclosures] ## Practical Implications _Note: The following reflects interpretation of the evidence, not sourced fact._ [What the evidence means in practice] ## Confidence Assessment [Use calibrated language from the Confidence table in SKILL.md] - Almost certain (93-99%): [findings] - Highly likely (80-92%): [findings] - Likely (63-79%): [findings] - Roughly even / Unlikely: [findings, if any] ## Limitations & Gaps [What remains uncertain, missing, or contested] ## Sources [^1]: Author/Org. "Title". Publication. URL [^2]: ... ``` -
technical.md 942 B
# Synthesis Template: Technical Comparison with context and benchmarks. ```markdown ## Executive Summary [2-3 sentences: key finding, confidence level, scope] ## Comparison Matrix | Criterion | Option A | Option B | Option C | | ----------- | -------- | -------- | -------- | | [criterion] | ... | ... | ... | ## Detailed Analysis ### [Criterion 1] [Evidence with benchmarks and citations] ### [Criterion 2] [Continue same pattern] ## Recommendation [Context-dependent recommendation with caveats about when each option wins] ## Confidence Assessment [Use calibrated language from the Confidence table in SKILL.md] - Almost certain (93-99%): [findings] - Highly likely (80-92%): [findings] - Likely (63-79%): [findings] - Roughly even / Unlikely: [findings, if any] ## Limitations & Gaps [What remains uncertain, missing, or contested] ## Sources [^1]: Author/Org. "Title". Publication. URL [^2]: ... ```
-
-
question-types.md 5.3 KB
# Question Type Classification Reference Classify the user's question into one or more types before planning research. For compound questions, decompose into sub-questions and classify each independently. ## Taxonomy | Type | Signals | Example Questions | | --------------------- | ---------------------------------------------------------- | ------------------------------------------------------------ | | **Factual** | Verifiable claims, dates, numbers, "what is", "how many" | "What year was GDPR enacted?" "How many GPUs did GPT-4 use?" | | **Scientific/Health** | Medical, biological, clinical, "is X safe", evidence-based | "Does intermittent fasting reduce inflammation?" | | **Consumer** | "Best X for Y", recommendations, buying decisions | "Best noise-cancelling headphones under $300 for travel?" | | **Technical** | APIs, tools, benchmarks, "how to", architecture decisions | "Redis vs Memcached for session caching at 10k RPS?" | | **Opinion/Sentiment** | "What do people think", community views, satisfaction | "How do developers feel about Tailwind CSS in 2026?" | | **Contested** | Political, ethical, actively debated, no consensus | "Should AI models be open-sourced?" | | **Emerging/Frontier** | Very recent, rapidly evolving, limited sources | "What are the leading approaches to LLM reasoning in 2026?" | ## Compound Question Decomposition Many real questions span multiple types. Decompose and classify each part: > "What's the best RAG framework, and is RAG even the right approach for my use case?" - Sub-question 1: "What's the best RAG framework?" -> **Technical** (comparison) - Sub-question 2: "Is RAG the right approach for [use case]?" -> **Technical** (architecture decision) > "Is creatine safe, and which brand should I buy?" - Sub-question 1: "Is creatine safe?" -> **Scientific/Health** - Sub-question 2: "Which brand should I buy?" -> **Consumer** **Ordering rule for compound questions:** Investigate higher-complexity sub-questions first (Contested > Emerging > Scientific > Technical > Consumer > Opinion > Factual). Later sub-questions often depend on earlier answers. ## Type-to-Scope Defaults Each type has a natural default scope. Present the recommended scope with brief rationale and let the user confirm. Then apply the scope modifiers below to adjust. | Type | Default Scope | Rationale | | --------------------- | ------------- | ------------------------------------------------ | | **Factual** | Focused | Few authoritative sources suffice | | **Scientific/Health** | Broad | Need evidence hierarchy across study types | | **Consumer** | Broad | Need expert reviews + user satisfaction + specs | | **Technical** | Broad | Need docs + benchmarks + practitioner experience | | **Opinion/Sentiment** | Broad | Need representative sample across communities | | **Contested** | Comprehensive | Must steelman all major positions | | **Emerging/Frontier** | Comprehensive | Sources are sparse; need breadth to find them | ### Scope Modifiers After identifying the default scope, check for these signals to adjust up or down one tier. **De-escalate one tier when:** - **Consumer**: Narrow category with fewer than 5 viable options and established expert reviews (e.g., "best USB-C hub for MacBook"). When expert consensus is clear, Round 2 adds only marginal value. - **Scientific/Health**: Single well-established finding with existing meta-analyses and no active controversy over mechanisms. - **Technical**: Well-documented comparison with published benchmarks and clear community consensus (e.g., mature tools with head-to-head comparisons already available). - **Opinion/Sentiment**: Question targets a single community with low controversy. **Escalate one tier when:** - **Factual**: Preliminary search reveals conflicting sources or the fact is embedded in a contested narrative. - **Consumer**: Emerging product category with sparse or unreliable reviews. - **Technical**: Novel or niche tools with sparse documentation, or the question involves architectural tradeoffs without clear benchmarks. - **Opinion/Sentiment**: Cross-community disagreement, politically charged, or culture-war adjacent topics. - **Any type**: Compound question with 3+ sub-questions of different types. **Do not use this skill when:** - The question is answerable with 1-2 web searches. - The question is a simple factual lookup with a single authoritative source. - The question is about debugging, code, or tasks better handled inline. ## Classification Procedure 1. Read the question. Identify surface signals (keywords, phrasing, domain). 2. Check for compound structure. If present, decompose into sub-questions. 3. Assign a primary type to each sub-question (or the whole question if simple). 4. Note the default scope from the table above. 5. Proceed to Phase 1 (Clarify) with the classification ready. Classification is silent -- do not present it to the user. It informs scope defaults, source evaluation priorities, and synthesis template selection. -
researcher-prompt.md 13.1 KB
# Deep Researcher Methodology You are a researcher on a coordinated team. Your job is to investigate assigned tasks by searching the web, evaluating sources, and reporting structured findings back to the team lead. ## Setup (Do This First) Before starting any research: 1. **Load your search toolkit** by invoking the `search-tips` skill using the Skill tool. This provides search strategy, tool guidance, and site workarounds. 2. **Read your assigned task** using `TaskGet` with the task ID from your spawn prompt. 3. **Begin investigation** following the loop below. ## Investigation Loop For each assigned task: 1. **Search**: Formulate 2-4 queries using Exa (with summaries enabled) to discover and triage sources. Then scrape the promising ones with Firecrawl to read the actual content. **Summaries are for triage only** — any claim you include in your findings must be based on text you read from the source itself, not from any summary (AI-generated or otherwise) 2. **Evaluate**: Assess findings for credibility, recency, depth, and bias 3. **Reflect** (mandatory before synthesizing): Ask yourself: - Do my key claims have 2+ independent sources? - Are there contradictions between sources I haven't resolved? - Did a source reference a primary source I should find directly? - Am I missing a perspective (e.g., only found supporters, no critics)? - Did I find something surprising that deserves deeper investigation? - Am I relying on search snippets, or did I read the full content? - **Cascade check** (see below): Are my "multiple sources" actually independent? 4. **Decide** based on reflection: - Gaps, contradictions, or thin coverage found -> formulate follow-up queries -> step 1 - Citation convergence + good multi-source coverage -> step 5 5. **Synthesize**: Structure findings into ~120 lines following the output format 6. **Report**: Write findings to disk, notify the lead with the file path, mark task complete **Search budget:** Aim for 2-3 search rounds. Do not stop after round 1 unless you genuinely hit citation convergence. Do not exceed 4 rounds. ## Source Evaluation Evaluate each source for credibility, recency, depth, and bias. Then apply the type-specific hierarchy below. The question type is provided in your spawn prompt. ### Type-Specific Source Priorities **Factual/Verifiable:** Prioritize government databases, official records, authoritative references. Deprioritize blog posts, opinion pieces, secondary news coverage. A single authoritative primary source can be sufficient. **Scientific/Health:** Prioritize systematic reviews, then RCTs, then cohort studies, then case reports, then expert opinion. Deprioritize anecdotal evidence, supplement vendor sites, non-peer-reviewed claims. Note sample sizes, funding sources, and methodology quality. **Consumer/Recommendation:** Prioritize expert testing labs (Wirecutter, RTINGS, Consumer Reports), then enthusiast forums with hands-on experience, then manufacturer specs. Deprioritize paid reviews, affiliate-heavy listicles, AI-generated roundups with no testing methodology. For consumer research specifically: start Reddit early — practitioner threads surface real failure modes and durability data that reviews miss. Try cross-domain search terms when obvious ones return poor results (e.g., "studio rack" instead of just "server rack"). **Technical/Comparative:** Prioritize official documentation, then working code examples, then expert blog posts with benchmarks, then Stack Overflow accepted answers. Deprioritize outdated tutorials, vendor marketing, theoretical comparisons without benchmarks. **Opinion/Sentiment:** Prioritize community forums (Reddit, HN, specialized forums), then surveys with methodology, then blog posts from practitioners. Deprioritize cherry-picked quotes, astroturfing indicators, single-voice opinion pieces presented as consensus. **Contested/Controversial:** Prioritize peer-reviewed research, then think tank analyses (label institutional bias), then long-form journalism, then advocacy organizations (label explicitly). Deprioritize social media hot takes, anonymous claims, sources that only present one side. **Emerging/Frontier:** Prioritize conference papers and preprints (with recency weight), then expert blogs from practitioners, then vendor announcements (label as interested party). Deprioritize older sources (>6 months may be outdated), speculative commentary without evidence. ## Cascade Check (Source Independence) Before claiming "2+ sources agree," verify they are genuinely independent. Information cascades occur when many sources trace back to one original, creating an illusion of corroboration. **Run this check during Step 3 (Reflect) for every key claim:** 1. **Trace the citation chain.** For each source supporting a claim, ask: where did *they* get this? If two articles both cite the same original study, wire report, or press release, you have one source, not two. Follow citations backward until you hit primary evidence. 2. **Check for copied phrasing.** Identical wording, the same errors/typos, or verbatim quotes across sources that claim independent reporting is the strongest cascade signal. If three articles contain the same unusual phrasing, they likely share a common upstream source. 3. **Assess source type diversity.** Two blog posts citing the same press release is replication, not corroboration. True independence requires different *types* of evidence converging: e.g., an empirical study + practitioner experience + official data. Multiple sources of the same type (all news articles, all blog posts) with the same factual claims are suspect. 4. **Apply the "how do they know?" test.** For each source, classify their knowledge as: - **First-hand**: direct observation, original research, primary data - **Second-hand**: told by a first-hand source, reporting on primary evidence - **Third-hand**: citing other secondary sources, "reports say" Only first-hand sources with independent access to the underlying reality count as genuine corroboration. Multiple second-hand sources tracing to the same first-hand source are a cascade. 5. **Watch for wire service cascades.** When the same claim appears across many news outlets simultaneously with identical framing, check whether they all republished AP/Reuters wire copy. This is legitimate distribution but counts as one source, not many. **Red flags that suggest a cascade rather than independent corroboration:** - All sources appeared within a short time window (hours) with no independent reporting - No source adds local detail, original quotes, or independent verification - Removing one original source would collapse the entire evidence base - Sources are all the same type (all news, all blog posts, all citing one study) - A claim is widely repeated but every version traces to one interview, press release, or study **When you detect a cascade:** Downgrade confidence to single-source, note it in your CROSS-SOURCE ANALYSIS as "apparent multi-source but single origin," and search for genuinely independent evidence. If none exists, flag it in GAPS. **SEO content farms and information cascades:** AI-generated listicles, vendor marketing disguised as independent reviews, and content farms create an illusion of source breadth. When multiple sources make the same claim, verify they aren't all citing the same upstream source. Five blog posts citing one Bloomberg article is one source, not five. Deprioritize "Complete Guide 2026" articles with no original testing or methodology. For vendor-vs-vendor comparisons, independent benchmarks with published methodology outweigh vendor blog posts. ## Stopping Criteria **Continue searching if:** - Important claims rest on a single source - Sources contradict each other without resolution - You found only secondary sources -- hunt for the primary - Critical subtopics or perspectives remain uncovered - New searches still yield novel sources **Stop searching if:** - **Citation convergence**: same authoritative sources across multiple queries - Same URLs keep appearing in results - 3+ rounds without new substantive sources - 2+ credible sources for each major claim **Be explicit in reasoning:** - "Round 1 found 3 sources but all vendor blogs -> targeting independent reviews" - "Sources A and B contradict on pricing -> searching specifically for current pricing" - "Last two searches returned known URLs -> convergence reached, synthesizing" ## Handling Insufficient Evidence - **Explicitly state** "Insufficient evidence found for [X]" rather than inferring - **Do NOT synthesize** plausible-sounding content to fill gaps - **Flag gaps prominently** in your GAPS section - **Return fewer, well-sourced findings** rather than padding with uncertain claims - When the absence of evidence *is* the finding (e.g., "no major vendor supports this feature"), report it with the same rigor as positive findings: what you searched for, how many sources you checked, and why the absence is meaningful - When researching non-English-speaking subjects or regions, flag the language limitation in GAPS — key sources may exist in languages the current toolset cannot search ## Team Coordination ### Reporting Findings After completing investigation, do ALL THREE of these steps in order: **Step 1: Persist findings to disk.** Write your structured findings to `{output_dir}/researcher-{letter}-findings.md` using the Write tool (always use **lowercase** for the letter, e.g., `researcher-a-findings.md`). The output directory and your researcher letter are provided in your spawn prompt. For follow-up tasks, append with a suffix: `researcher-{letter}-findings-{task-id}.md`. This backup enables cross-session resume and survives if team coordination fails. **Step 2: Notify the lead.** Send the file path, not the full findings (avoids doubling output tokens). The lead's name is provided in your spawn prompt. ``` SendMessage: type: "message" recipient: "{lead-name}" content: "Findings written to {output_dir}/researcher-{letter}-findings.md" summary: "Findings ready: {angle topic}" ``` **Step 3: Mark task complete.** ``` TaskUpdate: { taskId: "{id}", status: "completed" } ``` Then go idle. The lead will message you with follow-up tasks if needed. ### Between Tasks After completing a task, go idle. The lead dispatches follow-up tasks directly via `SendMessage` -- watch for incoming messages. When you receive a task assignment, claim it with `TaskUpdate` and begin immediately. Note: you may receive echo notifications of your own task completions — ignore these. ### Shutdown When you receive a `shutdown_request`, write a brief feedback file BEFORE approving shutdown: **Step 1: Write feedback file** to `{output_dir}/researcher-{letter}-feedback.md`: Reflect on your research experience — what would help the next researcher on a similar task? Focus on qualitative observations, not tool call logs (those are extracted automatically from transcripts). Keep it brief (5-15 lines). ```markdown ## What Worked - [What strategies or tools were most effective for this task?] ## Issues - [What was frustrating, broken, or confusing? Tool failures, unclear instructions, dead ends?] ## Suggestions - [What would have made this easier? Missing tools, better guidance, different approach?] ``` **Step 2: Approve shutdown:** ``` SendMessage: type: "shutdown_response" request_id: "{from the request}" approve: true ``` ## Output Format (STRICT) Your findings MUST follow this structure. Aim for ~120 lines. ```markdown ## ANGLE: [The specific question/angle you investigated] ## FINDINGS ### 1. [Claim title] [2-3 sentences with evidence] -- Source: [Author/Org], [Date]. [URL] Credibility: [High/Medium/Low] ### 2. [Claim title] [Continue same pattern -- aim for 3-5 findings total] ## CROSS-SOURCE ANALYSIS - Consensus (2+ sources agree): [list claims] - Conflicts: [list with both sides] - Single-source (lower confidence): [list claims] ## CONFIDENCE - High: [findings] - Medium: [findings] - Low/Uncertain: [findings] ## GAPS [What couldn't you find? What remains unanswered?] ## SOURCES 1. [Title](URL) -- [credibility note] 2. [Continue...] ``` ## Critical Rules 1. **NEVER return raw scraped content** -- only structured findings 2. **ALWAYS cite sources inline** -- every factual claim needs attribution 3. **ALWAYS assess confidence** -- distinguish multi-source from single-source claims 4. **Aim for ~120 lines** -- compress aggressively but preserve important nuance 5. **Flag contradictions explicitly** -- never silently pick one side 6. **Persist, THEN notify, THEN go idle** -- follow the 3-step sequence in "Reporting Findings" every time: write to disk, notify lead with file path, mark task complete 7. **Use lowercase filenames** -- always `researcher-a-findings.md`, never uppercase letters 8. **Stay on-angle** -- investigate the specific topic in your assigned task, not adjacent topics 9. **Verify claims against source text** -- summaries (AI-generated or human-written) are for triage and discovery only. Every factual claim in your findings must be grounded in text you read directly from the source. If you can't scrape the source, say so in GAPS rather than reporting the summary as fact
-
-
scripts
-
analyze-transcripts.py 14.6 KB
#!/usr/bin/env python3 """Analyze subagent transcripts from a deep-research-team session. Extracts MCP tool call parameters from researcher subagent JSONL transcripts to verify guidance compliance and understand search patterns. Usage: # Auto-detect: finds sessions with subagents matching a keyword python3 analyze-transcripts.py "intermittent fasting" # Explicit session ID python3 analyze-transcripts.py --session <session-id> # List recent sessions that have subagents python3 analyze-transcripts.py --list Output: Per-researcher summary of all MCP tool calls grouped by server, plus an aggregate compliance report for the Exa summary pattern. """ import argparse import json import os import re import sys from pathlib import Path # MCP tool name pattern: mcp__1mcp__{server}_1mcp_{tool} # After ToolSearch loading, tools appear as: mcp__1mcp__{server}_1mcp_{tool} MCP_PATTERN = re.compile(r"(?:mcp__1mcp__)?(\w+?)_1mcp_(\w+)") def parse_mcp_tool_name(name): """Extract server and tool from an MCP tool name. Returns (server, tool) or None if not an MCP tool. """ m = MCP_PATTERN.match(name) if m: return m.group(1), m.group(2) return None def find_project_dir(path=None): """Find the Claude project directory. Args: path: Optional path override. Accepts either: - A filesystem path (e.g., /Users/malo/.config/nix-config) which gets converted to the ~/.claude/projects/{slug} format - A direct ~/.claude/projects/{slug} path (used as-is if it exists) - None to auto-detect from cwd """ base = Path.home() / ".claude" / "projects" if path: p = Path(path).expanduser().resolve() # If it's already a valid project dir, use it directly if p.exists() and p.parent == base: return p # Otherwise treat it as a filesystem path and convert to slug source = str(p) else: source = os.getcwd() # Claude Code slugs: replace all non-alphanumeric characters with dashes slug = re.sub(r"[^a-zA-Z0-9]", "-", source) candidates = list(base.glob(f"{slug}*")) if len(candidates) == 1: return candidates[0] elif len(candidates) > 1: # Pick the one matching most closely for c in candidates: if c.name == slug: return c return candidates[0] return None def list_sessions(project_dir): """List sessions that have subagent directories, sorted by recency.""" sessions = [] for d in project_dir.iterdir(): if d.is_dir() and (d / "subagents").is_dir(): subagent_count = len(list((d / "subagents").glob("agent-*.jsonl"))) if subagent_count > 0: mtime = max(f.stat().st_mtime for f in (d / "subagents").glob("*.jsonl")) sessions.append((d.name, subagent_count, mtime)) sessions.sort(key=lambda x: x[2], reverse=True) return sessions def find_session_by_keyword(project_dir, keyword): """Find sessions whose subagents mention a keyword in their spawn message.""" matches = [] for d in project_dir.iterdir(): subdir = d / "subagents" if not d.is_dir() or not subdir.is_dir(): continue for f in subdir.glob("agent-*.jsonl"): if "compact" in f.name: continue try: with open(f) as fh: first = json.loads(fh.readline()) content = str(first.get("message", {}).get("content", "")) if keyword.lower() in content.lower(): matches.append(d.name) break except (json.JSONDecodeError, OSError): continue return list(set(matches)) def parse_subagent(path): """Parse a subagent JSONL file, extracting identity and tool calls.""" researcher = None task_subject = None question_type = None # mcp_calls: {server: [{tool, params}]} mcp_calls = {} with open(path) as f: for line in f: try: d = json.loads(line) except json.JSONDecodeError: continue msg = d.get("message", {}) content = msg.get("content", "") # Extract identity from spawn message or follow-up context if isinstance(content, str): m = re.search(r"researcher letter: (\w)", content, re.IGNORECASE) if m: researcher = m.group(1).upper() m2 = re.search(r'task is #\d+: "([^"]+)"', content) if m2: task_subject = m2.group(1) # Follow-up tasks have the subject in a different format if not task_subject: m2b = re.search(r'#\d+ -- "([^"]+)"', content) if m2b: task_subject = m2b.group(1) m3 = re.search(r"Question type: (\w+)", content) if m3: question_type = m3.group(1) # Extract identity from Write calls (follow-up agents write # to researcher-{letter}-findings-{id}.md) if isinstance(content, list): for block in content: if ( isinstance(block, dict) and block.get("type") == "tool_use" and block.get("name") == "Write" and not researcher ): fp = block.get("input", {}).get("file_path", "") m = re.search(r"researcher-([a-z])-findings", fp) if m: researcher = m.group(1).upper() # Extract MCP tool calls if isinstance(content, list): for block in content: if not isinstance(block, dict) or block.get("type") != "tool_use": continue name = block.get("name", "") inp = block.get("input", {}) parsed = parse_mcp_tool_name(name) if not parsed: continue server, tool = parsed mcp_calls.setdefault(server, []).append( {"tool": tool, "params": dict(inp)} ) return { "researcher": researcher, "task": task_subject, "question_type": question_type, "mcp_calls": mcp_calls, } def format_exa_call(call): """Format an Exa call with compliance checking.""" p = call["params"] q = p.get("query", "")[:80] has_summary = p.get("enableSummary") is True has_text_max = p.get("textMaxCharacters") == 1 has_ctx_max = "contextMaxCharacters" in p compliance = "" if has_summary and has_text_max and not has_ctx_max: compliance = " [COMPLIANT]" elif not has_summary or not has_text_max: compliance = " [NON-COMPLIANT]" if has_ctx_max: compliance += " [ctxMaxChars SET - BAD]" interesting = { k: v for k, v in p.items() if v and k not in ("query", "type") } param_str = f" {interesting}" if interesting else " (no params)" return f" Q: {q}\n{param_str}{compliance}" def format_firecrawl_call(call): """Format a Firecrawl call.""" url = call["params"].get("url", "")[:80] return f" {call['tool']}: {url}" def format_reddit_call(call): """Format a Reddit call.""" p = call["params"] tool = call["tool"] if tool == "get_top_posts": sub = p.get("subreddit", "?") time = p.get("time_filter", "all") return f" {tool}: r/{sub} ({time})" elif tool == "get_post_comments": post = p.get("post_id", p.get("url", "?")) depth = p.get("depth", "?") return f" {tool}: {post} (depth={depth})" elif tool in ("get_reddit_post", "get_subreddit_info"): target = p.get("post_id", p.get("url", p.get("subreddit", "?"))) return f" {tool}: {target}" elif tool == "search_reddit": q = p.get("query", "")[:60] sub = p.get("subreddit", "") suffix = f" in r/{sub}" if sub else "" return f" {tool}: \"{q}\"{suffix}" else: return f" {tool}: {json.dumps(p, default=str)[:80]}" def format_generic_call(call): """Format a generic MCP call.""" p = call["params"] summary = json.dumps(p, default=str) if len(summary) > 80: summary = summary[:77] + "..." return f" {call['tool']}: {summary}" SERVER_FORMATTERS = { "exa": format_exa_call, "firecrawl": format_firecrawl_call, "reddit": format_reddit_call, } def print_report(agents): """Print a formatted report of tool usage across all researchers.""" researchers = [a for a in agents if a["mcp_calls"]] if not researchers: print("No researcher MCP tool calls found in this session's subagents.") return # Collect all servers seen all_servers = set() for r in researchers: all_servers.update(r["mcp_calls"].keys()) # Per-researcher detail for r in researchers: letter = r["researcher"] or "?" task = r["task"] or "(follow-up or unidentified)" print(f"\n{'='*70}") print(f"Researcher {letter}: {task}") if r["question_type"]: print(f"Question type: {r['question_type']}") print(f"{'='*70}") for server in sorted(r["mcp_calls"].keys()): calls = r["mcp_calls"][server] print(f"\n{server} calls: {len(calls)}") formatter = SERVER_FORMATTERS.get(server, format_generic_call) for call in calls: print(formatter(call)) # Aggregate report print(f"\n{'='*70}") print("AGGREGATE REPORT") print(f"{'='*70}") # Per-server totals server_totals = {} server_tools = {} for r in researchers: for server, calls in r["mcp_calls"].items(): server_totals[server] = server_totals.get(server, 0) + len(calls) for c in calls: key = f"{server}.{c['tool']}" server_tools[key] = server_tools.get(key, 0) + 1 print("\nTool calls by server:") for server in sorted(server_totals, key=lambda s: -server_totals[s]): print(f" {server}: {server_totals[server]}") print("\nTool calls by endpoint:") for key in sorted(server_tools, key=lambda k: -server_tools[k]): print(f" {key}: {server_tools[key]}") # Exa compliance section exa_calls = [ c for r in researchers for c in r["mcp_calls"].get("exa", []) ] if exa_calls: total = len(exa_calls) summary_ok = sum(1 for c in exa_calls if c["params"].get("enableSummary") is True) text_max_ok = sum(1 for c in exa_calls if c["params"].get("textMaxCharacters") == 1) ctx_max_bad = sum(1 for c in exa_calls if "contextMaxCharacters" in c["params"]) full_ok = sum( 1 for c in exa_calls if c["params"].get("enableSummary") is True and c["params"].get("textMaxCharacters") == 1 and "contextMaxCharacters" not in c["params"] ) pct = lambda n: f"{100*n//total}%" if total else "0%" print(f"\nExa compliance ({total} calls):") print(f" enableSummary: true {summary_ok}/{total} ({pct(summary_ok)})") print(f" textMaxCharacters: 1 {text_max_ok}/{total} ({pct(text_max_ok)})") print(f" contextMaxCharacters: {ctx_max_bad} violations (should be 0)") print(f" Full compliance: {full_ok}/{total} ({pct(full_ok)})") categories = {} domains = {} for c in exa_calls: cat = c["params"].get("category") if cat: categories[cat] = categories.get(cat, 0) + 1 for dom in c["params"].get("includeDomains", []): domains[dom] = domains.get(dom, 0) + 1 if categories: print(f"\n Categories used:") for cat, count in sorted(categories.items(), key=lambda x: -x[1]): print(f" {cat}: {count}") if domains: print(f"\n includeDomains used:") for dom, count in sorted(domains.items(), key=lambda x: -x[1]): print(f" {dom}: {count}") def main(): parser = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter) parser.add_argument("keyword", nargs="?", help="Keyword to find in subagent spawn messages") parser.add_argument("--session", help="Explicit session ID") parser.add_argument("--list", action="store_true", help="List recent sessions with subagents") parser.add_argument("--project-dir", help="Override project directory path (filesystem or slug)") args = parser.parse_args() project_dir = find_project_dir(args.project_dir) if not project_dir or not project_dir.exists(): print("Could not find Claude project directory. Use --project-dir.", file=sys.stderr) sys.exit(1) if args.list: sessions = list_sessions(project_dir) if not sessions: print("No sessions with subagents found.") return print(f"Sessions with subagents in {project_dir.name}:\n") for sid, count, _ in sessions: print(f" {sid} ({count} subagents)") return if args.session: session_dir = project_dir / args.session if not (session_dir / "subagents").is_dir(): print(f"No subagents directory in session {args.session}", file=sys.stderr) sys.exit(1) session_ids = [args.session] elif args.keyword: session_ids = find_session_by_keyword(project_dir, args.keyword) if not session_ids: print(f"No sessions found matching '{args.keyword}'", file=sys.stderr) sys.exit(1) if len(session_ids) > 1: print(f"Multiple sessions match '{args.keyword}':") for sid in session_ids: print(f" {sid}") print("\nUse --session to pick one.") return else: parser.print_help() return for session_id in session_ids: subdir = project_dir / session_id / "subagents" print(f"Session: {session_id}") print(f"Path: {subdir}\n") agents = [] for f in sorted(subdir.glob("agent-*.jsonl")): if "compact" in f.name: continue if f.stat().st_size < 5000: continue # Skip tiny files (shutdown messages, nudges) result = parse_subagent(f) if result["mcp_calls"]: agents.append(result) print_report(agents) if __name__ == "__main__": main()
-
-
SKILL.md 23.3 KB
--- name: deep-research-team description: > This skill should be used when the user asks for "deep research", "research team", "comprehensive analysis", "research report", "investigate thoroughly", "compare X vs Y in depth", or needs synthesis across multiple sources with verification. It spawns a coordinated team of researcher agents across multiple rounds, with the lead triaging findings and creating targeted follow-up tasks. Scales from Focused (2 researchers, 1-2 rounds) to Comprehensive (4 researchers, 3-4 rounds with cross-verification). Do NOT use for simple lookups, debugging, or questions answerable with 1-2 searches. --- # Deep Research Team (Lead Orchestrator) Conduct thorough, iterative research by coordinating a persistent team of researcher agents across multiple rounds. This architecture enables mid-investigation steering, targeted follow-up based on emerging findings, and cross-agent verification. ## Architecture Overview ``` Round 1: Investigation Round 2: Follow-up Synthesis ┌──────────┐ ┌──────────┐ ┌────────┐ │Researcher│ sends findings │Researcher│ sends findings │ │ │ A ├─────────┬───────>│ A ├─────────┬───────>│ │ └──────────┘ │ └──────────┘ │ │ │ │ │ │ │ v dispatches v │ │ ┌──────────┐ ┌────────┐ ┌──────────┐ ┌────────┐ │ Lead │ │Researcher├───>│ Lead │───>│Researcher├───>│ Lead │───>│ synth │ │ B │ │triages │ │ B │ │triages │ │ esizes│ └──────────┘ └────────┘ └──────────┘ └────────┘ │ │ ^ ^ │ │ ┌──────────┐ │ ┌──────────┐ │ │ │ │Researcher├─────────┴───────>│Researcher├─────────┴───────>│ │ │ C │ sends findings │ C │ sends findings │ │ └──────────┘ └──────────┘ └────────┘ ``` **Key principles:** 1. **No peer-to-peer researcher communication.** All coordination goes through the lead. This preserves the independence that accounts for 87% of multi-agent gains (Choi et al.) and avoids sycophancy failures (Wynn et al.). Researchers never see each other's findings. 2. **Multi-round iteration.** The lead triages Round 1 findings and creates targeted Round 2 tasks for gaps, conflicts, and promising leads. 3. **Cross-agent verification (Comprehensive scope).** The lead asks Researcher A to verify Researcher B's high-impact single-source claim. The verifier only sees the claim and its source, not the original researcher's full analysis. 4. **Dynamic task evolution.** The shared task list starts with pre-planned angles but grows organically as follow-up tasks emerge from findings. The lead dispatches follow-up tasks directly to specific researchers via SendMessage. ## When to Use **Use this skill for:** - Complex questions that benefit from multiple research angles - Topics where initial findings will reveal what to investigate next - Research requiring cross-verification of contested claims - Any question needing synthesis across 5+ sources **Do NOT use for:** - Simple factual lookups (use regular web search) - Questions answerable with 1-2 searches - Debugging or code questions ## Effort Calibration | Scope | Researchers | Rounds | Verification | Model | | ----------------- | ----------- | ------ | ------------ | ------ | | **Focused** | 2 | 1-2 | None | sonnet | | **Broad** | 3 | 2-3 | None | sonnet | | **Comprehensive** | 4 | 3-4 | Cross-agent | opus | Round counts are heuristics, not targets. Stop early when you hit citation convergence -- additional rounds that don't surface new substantive findings waste tokens and context. A Broad run that converges in 2 rounds is a success, not a shortcut. After each round's triage, ask: "Would another round change the report's conclusions?" If not, proceed to synthesis. Default scope is determined by question type (see `references/question-types.md`). Present the recommended scope to the user and allow override. **Model selection:** - **Lead**: inherits user's session model (no override) - **Researchers**: `sonnet` for Focused/Broad, `opus` for Comprehensive (Sonnet validated as viable override for Comprehensive when cost matters) ## Output Directory Research artifacts persist to disk for resumability and backup. **Directory resolution** -- run this command FIRST, before creating anything. The output is your `{output_dir}`. Only the fallback branch creates a directory; the others reuse what exists. ```bash if [ -d "$(pwd)/deep-research" ]; then echo "$(pwd)/deep-research" elif [ -n "$CLAUDE_DEEP_RESEARCH_DIR" ]; then eval echo "$CLAUDE_DEEP_RESEARCH_DIR" else mkdir -p "$(pwd)/deep-research" echo "$(pwd)/deep-research" fi ``` After resolving `{output_dir}`, create only the topic subdirectory in Phase 3. Each session creates a subdirectory: `{output_dir}/{topic-slug}/` **Contents:** - `state.md` -- triage checkpoint, cross-references, follow-up plan (written in Phase 4) - `researcher-{letter}-findings.md` -- backup of each researcher's findings - `report.md` -- final synthesized report (written in Phase 6) ## Calibrated Confidence Language Use Kent-style verbal probability expressions in all confidence assessments: | Term | Range | Use When | | -------------- | ------ | ----------------------------------------- | | Almost certain | 93-99% | Multiple high-quality sources, no dissent | | Highly likely | 80-92% | Strong evidence, minor caveats | | Likely | 63-79% | Good evidence, some gaps | | Roughly even | 40-62% | Conflicting evidence, genuinely uncertain | | Unlikely | 20-39% | Limited or weak evidence | Always pair the verbal term with the probability range in the final report. ## Process ### Phase 0: Classify Question Type Silently classify the user's question before any interaction. 1. Read `references/question-types.md` for the full taxonomy. 2. Assign a primary type: Factual, Scientific/Health, Consumer, Technical, Opinion/Sentiment, Contested, or Emerging/Frontier. 3. For compound questions, decompose into sub-questions and classify each. 4. Note the default scope from the type-to-scope mapping. **Resume check:** Before starting, list the subdirectories in `{output_dir}` and scan for any that look related to the current question (similar topic, overlapping keywords). If you find a plausible match, read its `state.md` and offer to resume: present what was completed, what remains, and ask the user whether to resume or start fresh. If resuming, create a new team and tasks for only the remaining work. **Topic slug:** When creating a new session, generate a slug (lowercase, hyphenated, max 40 chars) for the subdirectory name: `{output_dir}/{slug}/`. Classification is internal—do not present it to the user. ### Phase 1: Clarify and Plan **Step 1: Make sure you understand the question.** Before planning anything, ask yourself: do I understand what the user is asking and *why* well enough to design research angles that will actually be useful to them? If not, use `AskUserQuestion` to fill the gaps. This isn't just about ambiguous wording -- a perfectly clear question can still lack enough context to research well ("How does Nix handle dependencies?" means very different research depending on whether you're evaluating Nix, debugging an issue, or writing docs). If the question and its context are clear, skip this step. **Step 2: Scope and decompose.** Determine the appropriate scope (Phase 2 has the details) and decompose the question into independent research angles. Default angle counts by scope: - Focused: 2 angles - Broad: 3 angles - Comprehensive: 4 angles These are defaults, not caps. If the decomposition reveals one more genuinely independent facet than the default, add it (e.g., 3 angles for a Focused run). If the question has fewer real facets, use fewer. Beyond ±1 from the default, re-scope rather than stretching -- the scope was probably wrong. Each angle must be independent and substantial enough to warrant a dedicated researcher; "I can think of another angle" isn't sufficient. For compound questions, map sub-questions to angles. Multiple sub-questions can share an angle if closely related; a single sub-question can span multiple angles if it has distinct facets. **Step 3: Confirm if high-investment.** For compound, contested, or Comprehensive-scope questions, present the research plan for user approval before spawning researchers: > Research plan for "{question}": > > - Type: {type} | Scope: {scope} | {N} researchers, {M} rounds > - Angles: {list of planned angles} > - [If compound] Sub-question → angle mapping: ... > > Proceed, or adjust? For clear, low-scope questions, skip confirmation and proceed. ### Phase 2: Calibrate Effort Select the scope tier based on question type defaults from `references/question-types.md`, then apply the scope modifiers from that file (de-escalation and escalation signals). Also adjust for: - User's explicit preference (if stated) - Structural complexity (compound questions with 3+ sub-types escalate) Announce the plan: "Starting {scope} team research with {N} researchers." ### Phase 3: Team Setup **Step 1: Create the output directory** ```bash mkdir -p {output_dir}/{topic-slug} ``` **Step 2: Create the team** ``` TeamCreate: team_name: "deep-research-{topic-slug}" description: "Researching {topic} in {scope} scope" ``` **Step 3: Create initial tasks** One `TaskCreate` per research angle: ``` TaskCreate: subject: "Investigate {angle title}" description: | Research angle: {angle description} Topic context: {brief topic summary} Question type: {type from Phase 0} Focus: {what specifically to investigate} Return structured findings via SendMessage to the lead. activeForm: "Investigating {angle title}" ``` **Task brief clarity:** Make scope boundaries explicit between researchers to avoid overlap and gaps. Flag name ambiguities (e.g., multiple products sharing a name). Mark optional sub-tasks clearly (e.g., "if time permits" vs required). **Step 4: Spawn researchers** Launch ALL researchers in a SINGLE message. Each researcher gets: ``` Task: subagent_type: "general-purpose" name: "researcher-{letter}" team_name: "deep-research-{topic-slug}" model: "sonnet" (or "opus" for Comprehensive) description: "Spawn researcher {letter}" prompt: | You are a research agent on a team. Your job is to investigate research tasks by searching the web, evaluating sources, and reporting structured findings. FIRST: Read your methodology at: {absolute path to references/researcher-prompt.md} Question type: {type from Phase 0} Output directory: {output_dir}/{topic-slug} Your researcher letter: {letter} (use LOWERCASE in filenames: researcher-{lowercase letter}) Lead name: team-lead (send all findings to this name via SendMessage) Your assigned task is #{id}: "{subject}" Use TaskGet for full details, then begin investigation. After completing your task, go idle. The lead will message you directly when new tasks are available. ``` Include the task ID, subject, question type, output directory, researcher letter, and lead name directly in each researcher's spawn prompt. ### Phase 4: Investigation Loop This is the core research cycle. Each iteration is a **round**: researchers investigate, the lead triages, then either dispatches follow-ups (another round) or exits to synthesis. #### Round structure **Investigate:** Researchers work on their assigned tasks. Each researcher will: 1. Read the methodology reference file 2. Load web search and content extraction tools via ToolSearch 3. Execute the investigation loop (search -> evaluate -> reflect -> decide) 4. Write findings to `{output_dir}/{topic-slug}/researcher-{letter}-findings.md` 5. Notify the lead via `SendMessage` with the file path (not the full findings -- avoids doubling output tokens) 6. Mark their task as completed via `TaskUpdate` 7. Go idle and wait for the lead to dispatch follow-up tasks via `SendMessage` **Monitoring**: The lead reads each researcher's findings file after receiving their notification. No polling needed. **Handling partial results**: If a researcher reports rate limit issues or thin coverage, note the gap for triage rather than immediately spawning replacements. **Triage:** After all tasks for the current round complete, systematically review findings. *Step 1: Extract and cross-reference claims* For each significant claim across all researcher findings: - How many **independent sources** support it? (Different researchers finding the same source counts as one source, not two.) - **HIGH confidence**: 3+ independent sources of different types (e.g., paper + dataset + practitioner account), no credible dissent - **MEDIUM confidence**: 2 independent sources, or multiple sources of the same type - **LOW confidence**: Single source, or multiple sources that trace back to one original *Step 2: Identify gaps and conflicts* - What angles remain uncovered? - Where do researchers contradict each other? - What findings are surprising and deserve deeper investigation? - Which claims rest on a single source? *Step 3: Decide whether to continue or exit* Exit to Phase 5 (Synthesis) if findings have converged -- another round wouldn't change the report's conclusions. Continue if significant gaps, conflicts, or single-source high-impact claims remain and the scope's round budget allows. If findings reveal more complexity than anticipated (e.g., Broad scope uncovering deeply contested claims requiring steelmanning), escalate: spawn an additional researcher or add a round beyond the default budget. *Step 4: Persist triage state* Write or update `{output_dir}/{topic-slug}/state.md`: ```markdown # Research State: {topic} ## Status: TRIAGE_COMPLETE (Round {N}) ## Question Type: {type} ## Scope: {scope} ## Round {N} Summary {brief cross-reference of key findings, gaps, conflicts} ## Follow-up Plan {list of planned follow-up tasks with rationale, or "Proceeding to synthesis"} ``` #### Dispatching follow-up tasks If continuing, create targeted tasks based on what triage revealed. These are all just task types -- they use the same dispatch mechanism: > **Gap-fill**: "No researcher covered {aspect}. Investigate {specific question}." > **Conflict-resolution**: "One source says X, another says Y. Search for additional > sources that clarify which is accurate and why they might differ." > **Deep-dive**: "Initial findings revealed {unexpected thing}. Investigate further: > {specific follow-up questions}." > **Verification** (Comprehensive scope): Use when high-impact claims rest on a single > source, factual conflicts remain unresolved, or claims are in specialized/niche domains > where citation error rates are higher. Assign to a researcher who did NOT make the > original claim. Task description contains only the claim and its source URL: > > ``` > TaskCreate: > subject: "Verify: {claim summary}" > description: | > Verification task. Search for ADDITIONAL sources (not the original) and determine > if they support, contradict, or add nuance to this claim. > > CLAIM: {specific factual claim} > ORIGINAL SOURCE: {URL} > > Report your verdict as: SUPPORTED, SUPPORTED WITH NUANCE, CONTESTED, or UNCHANGED > Include the additional sources you found and any important nuance. > activeForm: "Verifying claim about {topic}" > ``` **Critical**: Follow-up task descriptions contain just enough context without revealing other researchers' full conclusions. This preserves independence. **Assignment strategy:** The lead assigns follow-up tasks directly via `SendMessage` rather than relying on researchers to self-claim. Choose assignees based on task type: - **Deep-dives and gap-fills**: Assign to the researcher who covered the related angle (continuity -- they have context on what was already found). - **Conflict resolution and verification**: Assign to a researcher who did NOT cover either side (fresh perspective avoids confirmation bias). - **If researchers outnumber tasks**: Idle researchers wait or are shut down early. ``` SendMessage: type: "message" recipient: "researcher-{letter}" content: "New task available: #{id} -- {subject}. Please claim it and begin." summary: "Follow-up task assignment" ``` Then loop back to **Investigate** above. **Interpreting verification results:** When a verification task returns, adjust confidence: - **SUPPORTED**: Upgrade confidence; note additional sources - **SUPPORTED WITH NUANCE**: Directionally correct but specific details differ or require qualification. Upgrade confidence for the general claim; add caveats for specifics. - **CONTESTED**: Flag explicitly; present both sides with evidence - **UNCHANGED**: Keep original confidence level ### Phase 5: Synthesize (Type-Aware) Combine all findings from all rounds into a coherent report. Select the synthesis template matching the question type from Phase 0. For compound questions, use the template for each sub-question's type, then add an overall synthesis section. **End-of-sequence awareness:** Draft Confidence Assessment and Limitations sections **early**, not last. Review final paragraphs specifically for unsourced claims. #### Template Selection Read the template file matching the question type from Phase 0. Each template includes the full report structure (executive summary, type-specific body, confidence assessment, limitations, sources). For compound questions, read the template for each sub-question's type and add an overall synthesis section. | Question Type | Template File | | --------------------- | -------------------------------------------- | | **Factual** | `references/templates/factual.md` | | **Scientific/Health** | `references/templates/scientific-health.md` | | **Consumer** | `references/templates/consumer.md` | | **Technical** | `references/templates/technical.md` | | **Opinion/Sentiment** | `references/templates/opinion-sentiment.md` | | **Contested** | `references/templates/contested.md` | | **Emerging/Frontier** | `references/templates/emerging-frontier.md` | ### Phase 6: Persist Report Write the final report: ``` Write: {output_dir}/{topic-slug}/report.md ``` Update `state.md` status to `COMPLETE`: ``` Edit: {output_dir}/{topic-slug}/state.md old_string: "## Status: TRIAGE_COMPLETE" new_string: "## Status: COMPLETE" ``` **Format output files (optional):** If `prettier` is available, run it on all markdown files in the output directory to normalize formatting: ```bash prettier --write --prose-wrap preserve "{output_dir}/{topic-slug}/**/*.md" ``` If prettier is not installed, skip this step silently -- it is cosmetic, not functional. Inform the user: "Report saved to `{output_dir}/{topic-slug}/report.md`." ### Phase 7: Cleanup Shut down the team cleanly. **Step 1: Shut down researchers** Send `shutdown_request` to each researcher via `SendMessage`: ``` SendMessage: type: "shutdown_request" recipient: "researcher-a" content: "Research complete. Shutting down." ``` Repeat for each researcher. Wait for shutdown responses. **Step 2: Read researcher feedback** After all researchers have shut down, read any feedback files written to `{output_dir}/{topic-slug}/researcher-{letter}-feedback.md`. These contain notes on tool usage (Exa parameters, Firecrawl usage), issues encountered (400 errors, rate limits), and suggestions. Use this feedback to identify patterns for skill improvement. **Step 3: Delete team** ``` TeamDelete ``` ## Writing Standards - Prose paragraphs, not bullet lists (bullets only for distinct enumerations) - Specific data: "increased 23%" not "increased significantly" - Cite inline with markdown footnotes: "The market grew 15%[^1]" not "The market grew.[^1]" - Each finding: 2-4 paragraphs with evidence - Distinguish FACTS (cited) from ANALYSIS (synthesis) ## Anti-Hallucination Protocol - Every factual claim must cite a source immediately - Mark synthesis distinctly: "This suggests..." or "Synthesizing these findings..." - If uncertain, say so: "Sources disagree on..." or "Limited evidence for..." - Never fabricate sources -- all citations come from researcher findings ## Additional Resources ### Reference Files - **`references/question-types.md`** -- 7-type taxonomy, signals, decomposition rules, type-to-scope defaults. Read in Phase 0. - **`references/researcher-prompt.md`** -- Investigation methodology, type-aware source evaluation, output format. Path provided to researchers in spawn prompt. (Tool guidance extracted to the standalone `search-tips` skill, which researchers load as their first step.) - **`references/templates/`** -- Type-specific synthesis templates. Read the relevant template(s) in Phase 5. See Template Selection table above. ### Scripts - **`scripts/analyze-transcripts.py`** -- Post-hoc analysis of researcher tool usage. Extracts MCP tool call parameters from subagent JSONL transcripts and produces a compliance report. Usage: - `python3 scripts/analyze-transcripts.py --session ${CLAUDE_SESSION_ID}` -- current session - `python3 scripts/analyze-transcripts.py "topic keyword"` -- auto-detect session by keyword - `python3 scripts/analyze-transcripts.py --list` -- list recent sessions with subagents ### Development History - **`dev/RESEARCH.md`** -- Design rationale with 50+ sources justifying the architecture - **`dev/ITERATION-LOG.md`** -- 13 iterations of improvement with backlog - **`dev/iterations/`** -- Detailed notes for each iteration ## Quick Reference 0. **Classify** question type (silent) and check for resume 1. **Clarify** -- understand the question, resolve ambiguity, confirm plan if high-investment 2. **Calibrate** effort: announce scope and researcher count 3. **Setup**: create output dir, TeamCreate, TaskCreate per angle, spawn researchers 4. **Investigation loop**: investigate → triage → dispatch follow-ups or exit. Repeat until converged. 5. **Synthesize**: type-aware template, calibrated confidence language 6. **Persist**: write report.md, update state.md to COMPLETE, tell user file location 7. **Cleanup**: shutdown_request to each researcher, then TeamDelete **Context budget:** - Lead context: reserve for triage + synthesis - Researcher contexts: handle all search/scrape operations - Researcher spawn prompts: compact (~30 lines), point to reference file - Researcher findings: structured summaries (~120 lines each)
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.