Claude Skill

search-tips

This skill should be used when performing web research beyond a simple single search -- looking into topics, comparing options, investigating questions, finding recommendations, or any task where effective use of Exa, Firecrawl, and Reddit tools matters. Triggers on "research", "

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download malob-nix-config-configs_claude_skills_search-tips-500b474.zip · 19 KB
Part of malob/nix-config — 13 skills

Install

skills CLI npx skills add https://github.com/malob/nix-config/tree/master/configs/claude/skills/search-tips
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install malob-nix-config@llmmart
Git git clone https://github.com/malob/nix-config.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole malob/nix-config collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Search Tips

Accumulated guidance for web research using Exa, Firecrawl CLI, and Reddit MCP tools. These are starting points, not rigid rules -- think strategically about each situation and adapt. If a different approach makes more sense for what you're trying to do, go with it. Run npx firecrawl-cli <command> --help to check available options beyond what's documented here. Reference files cover tool-specific deep dives -- see bottom of this file.

Setup

Before starting research, load the required MCP tools using ToolSearch:

  1. Exa tools -- web_search_advanced_exa and get_code_context_exa
  2. Reddit tools -- get_top_posts, get_post_comments, get_reddit_post, get_subreddit_info

Firecrawl CLI (npx firecrawl-cli) runs via Bash -- no MCP setup needed. Load only what the task requires.

The Research Cycle

Prefer Exa and Firecrawl over built-in WebSearch/WebFetch.

Research alternates between searching (discovering sources) and fetching (extracting content from them). Find promising leads, read the best ones, refine your understanding, search again.

Searching

Finding sources you don't have yet.

  • Exa search (web_search_advanced_exa) -- primary tool for web discovery. Natural language queries, add filters as needed (domains, dates, categories).
  • Exa code context (get_code_context_exa) -- programming topics. Worth trying before general Exa search for technical/code tasks -- surfaces repos, packages, and docs.
  • Firecrawl CLI search (npx firecrawl-cli search) -- Google-powered keyword search. Useful when keyword matching works better than Exa's semantic approach, for site-scoped queries (site:reddit.com {query}), and for content-type filtering (--categories research for academic, --sources news, --tbs qdr:w for time).

Default Exa search pattern: Default to enableHighlights: true and textMaxCharacters: 1. This returns quoted passages from actual page text while preventing the MCP server from flooding context with full text. Use highlightsPerUrl and highlightsNumSentences to control volume if needed.

Fetching

Extracting content from a source you've identified.

  • Firecrawl CLI scrape (npx firecrawl-cli scrape "<url>" --only-main-content) -- primary tool for reading a known URL. The flag strips nav/sidebars to save tokens.
  • Reddit MCP (get_post_comments, get_reddit_post) -- for reading Reddit threads. Firecrawl can't scrape reddit.com directly.
  • Firecrawl CLI map (npx firecrawl-cli map "<url>" --search "query") -- discover URLs on a site (useful when you need to find the right page, or when scrape returns empty).

Adapting the Workflow

The defaults above won't always be right. Some common deviations:

  • Exa full text as a scraping fallback -- some sites are blocked or inaccessible via Firecrawl (LinkedIn, Twitter/X, etc.), but Exa often has the full page text in its index. Drop both enableHighlights and textMaxCharacters: 1 to get the complete text. Be aware this can produce large responses.
  • --only-main-content can strip too much -- if you got empty or partial results, retry without the flag. Known to fail on Future plc sites, Blogspot, and GDPR-heavy sites. See Content Extraction below.

The reference files cover more edge cases -- scraping issues, category restrictions, and academic search.

Search Strategy

How Exa Works

Exa is a neural/semantic search engine. It uses embeddings to understand meaning.

  • Natural questions or statements work best -- Exa finds pages that answer them
  • Longer, more specific queries work BETTER -- unlike keyword-based search
  • Keyword lists tend to confuse the semantic model

Good: "What do professional reviewers say are the most reliable dishwasher brands in 2025?" Bad: "best dishwasher 2025 reliable"

Query Reformulation

For broad topics, Exa's additionalQueries parameter can automate this -- it bundles query variations in a single call at no extra cost (see references/exa-tips.md). For manual reformulation, try generating 3-5 query variations:

Technique What It Does Example
Paraphrase Same meaning, different words "RAG failures" -> "problems in RAG systems"
Decompose Break into sub-questions "Why fail?" -> "Why return irrelevant docs?"
Scope shift Broader context or narrower specifics "Challenges in production AI search"
Perspective shift Different viewpoints User vs expert vs critic view
Temporal framing Target different time periods "Recent 2024-2025" vs "foundational"

Domain Filtering

Try includeDomains when the authoritative site for a topic is known -- faster and less noisy than broad search. Try excludeDomains to suppress sites that keep appearing but aren't useful (e.g., exclude youtube.com when video pages crowd out needed editorial content about YouTube creators/content).

Searching by Content Type

Match your research target to the right approach. Reference files have full strategies.

  • Social sentiment / opinions -- Exa tweet for Twitter; site:reddit.com via Firecrawl search + Reddit MCP for discussions. See references/twitter.md, references/reddit.md.
  • People / companies -- Exa people or company category for discovery, then broaden. See references/people-companies.md.
  • Academic papers -- Exa research paper category; academic APIs for structured data. See references/academic-search.md.
  • Code / GitHub -- get_code_context_exa for code; gh api for repo data. See references/code-github.md.
  • News -- Exa news category with date filters; Firecrawl CLI search --sources news --tbs qdr:w for Google News with time filtering.
  • Financial reports -- Exa financial report with domain/date filtering.
  • Personal blogs / independent takes -- Exa personal site for practitioner opinions, blog posts, and independent analysis. Full parameter support (unlike most specialized categories). See references/personal-sites.md.

Some categories reject certain parameters (400/500 errors). See references/exa-tips.md.

Content Extraction

Always use --only-main-content by default -- it strips nav, sidebars, and footers, saving significant tokens. If you get empty or partial results, retry without the flag; it's known to strip article bodies on Future plc sites (iMore, Pocket-lint), GDPR-heavy sites (StorageReview), and Blogspot blogs -- see references/scraping-issues.md.

For even narrower extraction, use --include-tags (e.g., --include-tags "article", --include-tags ".post-content").

Scraping Issues

When Firecrawl scrape fails (policy blocks, paywalls, SPA rendering), check references/scraping-issues.md for per-site workarounds. General fallback: try Exa full text (drop enableHighlights and textMaxCharacters). For interactive pages that need clicks or form fills, escalate to npx firecrawl-cli browser (see references/firecrawl-tips.md).

Operational Notes

Parallel call failures: If any tool call in a parallel batch fails, all sibling calls fail too. Retry individually. Keep failure-prone calls (Reddit MCP) in their own batch.

Rate limits: Retry once after a short pause. If still blocked, try a different tool for the same intent before giving up -- Exa rate-limited? Try Firecrawl search. Reddit MCP throttled? Try Exa with includeDomains: ["reddit.com"]. Only note the gap and move on after both the original tool and an alternative have failed.

Reference Files

Content-type strategies

  • references/twitter.md -- Exa tweet category, restrictions, query tips, livecrawl
  • references/reddit.md -- Discovery via Firecrawl search, Reddit MCP for reading, batching gotchas, rate limits
  • references/people-companies.md -- Exa people/company categories, LinkedIn, multi- category approach
  • references/code-github.md -- Code search, GitHub repos/issues, gh api, raw URLs
  • references/personal-sites.md -- Independent blogs, practitioner opinions, full parameter support, portfolio exploration
  • references/academic-search.md -- Academic domains, free APIs, Firecrawl CLI

Tool reference

  • references/exa-tips.md -- Category restrictions, additionalQueries, highlights/summaries
  • references/firecrawl-tips.md -- CLI commands (search, scrape, map, crawl, download, browser), PDF scraping, arxiv extraction, MCP fallback notes

Troubleshooting

  • references/scraping-issues.md -- main-content detection fixes, paywalls, Discourse JSON, policy-blocked sites, API workarounds
Files (nix-config)
  • references
    • academic-search.md 2.9 KB
      # Academic Search
      
      ## When to Use
      
      Literature reviews, evidence gathering, prior art research, finding methodology details,
      systematic exploration of a research field or topic.
      
      ## Tools and Approach
      
      **Primary tool:** Exa `category: "research paper"` -- surfaces arXiv papers, conference
      proceedings, and journal articles for research-oriented queries.
      
      **Key parameters:**
      
      - `includeDomains` for specific venues: `arxiv.org`, `openreview.net`, `nature.com`,
        `semanticscholar.org`, `pubmed.ncbi.nlm.nih.gov`
      - `startPublishedDate` / `endPublishedDate` for temporal scope (e.g., last 2 years for
        state-of-the-art, broader for foundational work)
      - `highlightsQuery` for targeted extraction of methodology, findings, or specific aspects
      - `additionalQueries` for phrasing variations (different technical terms for the same concept)
      
      The default search pattern (`enableHighlights: true`, `textMaxCharacters: 1`) works well for
      triage -- highlights extract abstracts and key passages effectively.
      
      The undocumented `pdf` category also works -- returns papers from arxiv, researchgate, icml,
      aclanthology, and similar sources. Useful as an alternative when `research paper` misses
      relevant results.
      
      ## Academic Domains
      
      Useful values for `includeDomains`:
      
      - `arxiv.org` -- preprints (CS, physics, math, biology, economics)
      - `openreview.net` -- peer review and conference papers (ICLR, NeurIPS)
      - `semanticscholar.org` -- 214M papers across all disciplines
      - `pubmed.ncbi.nlm.nih.gov` -- medical/biomedical literature
      - `biorxiv.org` -- biology preprints
      - `nature.com`, `science.org` -- high-impact journals
      
      ## Free APIs
      
      These return clean structured data via Firecrawl scrape or `curl` (no auth needed):
      
      - **Semantic Scholar:** `https://api.semanticscholar.org/graph/v1/paper/search?query={query}`
      - **OpenAlex:** `https://api.openalex.org/works?search={query}`
      - **arXiv:** `http://export.arxiv.org/api/query?search_query=all:{query}`
      
      For structured metadata (citation counts, authors, abstracts), these APIs are more reliable
      than scraping paper pages.
      
      ## Firecrawl CLI Academic Search
      
      `npx firecrawl-cli search --categories research "{query}"` provides Google Scholar-like
      academic search. Returns results from arxiv, nature.com, and similar sources.
      
      ```bash
      npx firecrawl-cli search --categories research "transformer attention mechanisms" --scrape --limit 5
      npx firecrawl-cli search --categories research "RLHF alignment" --tbs qdr:y  # past year
      ```
      
      The `--categories research` flag is CLI-only (not available via MCP). See
      `firecrawl-tips.md` for more CLI commands.
      
      ## Gotchas
      
      - **arxiv `/abs/` pages** return abstract + metadata mixed with navigation chrome. Better to
        scrape `arxiv.org/html/{paper-id}` for the full paper or use Exa highlights for triage.
        See `firecrawl-tips.md` for details.
      - **Paywalled journals** (Nature, Science, Elsevier) may return only abstracts. Check for
        preprint versions on arxiv or author pages. The Wayback Machine sometimes has cached copies.
      
    • code-github.md 3.1 KB
      # Code and GitHub
      
      ## When to Use
      
      Code examples, library documentation, package APIs, GitHub repos/issues/discussions, open-source
      project research, technical implementation patterns.
      
      ## Code Search
      
      **Primary tool:** `get_code_context_exa` -- searches code-related content including repos,
      packages, documentation, Stack Overflow, and technical blogs.
      
      **Query tips:**
      
      - Always include the **programming language** ("Go generics" not just "generics")
      - Include framework + version when applicable ("Next.js 15", "React 19", "Python 3.12")
      - Use exact identifiers: function names, error messages, class names, config keys
      - More specific queries perform better than broad ones (Exa is semantic, not keyword-based)
      
      Good: "How to use React Server Components with Next.js 15 app router for data fetching"
      Bad: "react server components nextjs"
      
      **`tokensNum` tuning:** Controls how much context is returned per result.
      
      - Focused snippet needed: 1000-3000
      - Most tasks (default): 5000
      - Complex integration patterns: 10000-20000
      
      ## GitHub Discovery
      
      `get_code_context_exa` also searches GitHub repos, issues, discussions, and READMEs. This is
      the fastest path for finding relevant GitHub content -- no special permissions needed.
      
      ## GitHub Deep Read
      
      When you need the full content of a specific GitHub page:
      
      - **Firecrawl scrape** (`npx firecrawl-cli scrape "<url>"`) works well for GitHub issues,
        discussions, READMEs, and wiki pages
      - **Raw file content:** `raw.githubusercontent.com/{owner}/{repo}/{branch}/{path}` for clean
        file content without UI noise. Scrape this URL via Firecrawl.
      
      ## GitHub Data via `gh` CLI
      
      Prefer `gh` subcommands over `gh api` -- subcommands are auto-approved while `gh api` with
      write flags (`-X`, `-f`, `-F`) requires user approval.
      
      - `gh repo view {owner}/{repo} --json stargazerCount,forkCount,pushedAt,licenseInfo`
      - `gh issue list -R {owner}/{repo} --state open --json number,title,labels,createdAt`
      - `gh issue view {number} -R {owner}/{repo} --json body,comments`
      - `gh pr list -R {owner}/{repo} --json number,title,state` -- check for pending fixes
      - `gh pr view {number} -R {owner}/{repo} --json body,comments,files`
      - `gh search repos "{query}" --json fullName,stargazersCount,description`
      - `gh search code "{pattern}" --repo {owner}/{repo}`
      - `gh release list -R {owner}/{repo}`
      - `gh release view {tag} -R {owner}/{repo}`
      
      For data without subcommand equivalents, `gh api` GET requests are also auto-approved:
      - `gh api repos/{owner}/{repo}/stats/participation` -- weekly commit activity
      - `gh api repos/{owner}/{repo}/contributors` -- contributors with commit counts
      
      ## Gotchas
      
      - **GitHub rate limits:** The `gh api` commands use your authenticated token, so limits are
        generous (5000/hour). Unauthenticated API calls are 60/hour.
      - **Large repos:** READMEs and issue threads can be very long. Use
        `npx firecrawl-cli scrape "<url>" --only-main-content` to get the main content without
        sidebar noise.
      - **GitHub search via Exa vs `gh`:** Exa is better for semantic discovery ("libraries for X").
        `gh search repos "query"` is better for exact keyword matches and quantitative filtering
        (stars, language, license).
      
    • exa-tips.md 3.5 KB
      # Exa Tips
      
      Deep-dive reference for Exa search parameters. For everyday usage, see `SKILL.md`.
      
      ## Category Restrictions
      
      Some categories reject certain parameters with 400/500 errors. Others silently ignore them
      (no error, but filter not applied).
      
      | Category           | Unsupported parameters                                                                          |
      | ------------------ | ----------------------------------------------------------------------------------------------- |
      | `tweet`            | `includeText`, `excludeText`, `includeDomains`, `excludeDomains`, `moderation` (causes 500)     |
      | `company`          | All date filters, `includeText`, `excludeText`, `includeDomains`, `excludeDomains`              |
      | `people`           | All date filters, `includeText`, `excludeText`, `excludeDomains`. `includeDomains` LinkedIn only |
      | `financial report` | `excludeText` (causes 400). `includeText` works (single-item only)                              |
      | `research paper`   | Full support                                                                                    |
      | `personal site`    | Full support                                                                                    |
      | `news`             | Full support                                                                                    |
      | `pdf`              | Undocumented category -- available in MCP, returns academic papers (arxiv, researchgate, etc.)   |
      | `github`           | Undocumented category -- available in MCP, returns GitHub repos                                 |
      
      **`company` + `includeDomains`** silently ignores the domain filter (returns results from any
      domain). No error thrown.
      
      ## includeText / excludeText Limits
      
      - Only accept **single-item arrays** (one string each)
      - Each string must be **5 words or fewer**
      - Multi-item arrays or strings longer than 5 words cause 400 errors
      - Put additional filtering terms in the `query` string instead
      
      ## additionalQueries
      
      `additionalQueries: ["variation 1", "variation 2"]` bundles query variations in a single call.
      Same cost as a single query. Works with `auto` mode (our MCP tool exposes `auto`/`fast`/`neural`).
      
      - **Broad topics** (e.g., "LLM hallucination mitigation"): genuine diversity gain, surfaces
        results the primary query misses. Worth using.
      - **Niche/specific topics** (e.g., "Nix flakes best practices"): results largely converge.
        May surface a few extra sources but core results stay the same.
      
      Worth trying for research tasks where comprehensive coverage of a broad topic matters. Less
      useful for targeted lookups. Can substitute for manual query reformulation on broad topics.
      
      ## Highlights (Default Search Pattern)
      
      `enableHighlights: true` + `textMaxCharacters: 1` is the recommended default for Exa search.
      Highlights are extracted server-side before text truncation, so you get actual quoted passages
      from the page without the full-text context flood.
      
      - `highlightsPerUrl` -- control how many highlight blocks per result
      - `highlightsNumSentences` -- control sentences per highlight
      - `highlightsQuery` -- focus highlights on a specific aspect of the page
      - Without `textMaxCharacters: 1`, you get full text (7K-32K chars) alongside highlights --
        best to pair them.
      
      ## Summaries
      
      `enableSummary: true` + `textMaxCharacters: 1` generates AI summaries per result. These are
      concise but **can hallucinate** -- they sometimes fabricate details not present on the page.
      Highlights are preferred because they return actual page text. Summaries exist as an option
      but there's rarely a reason to use them over highlights.
      
    • firecrawl-tips.md 5.9 KB
      # Firecrawl Tips
      
      Deep-dive reference for Firecrawl CLI. For everyday usage, see `SKILL.md`.
      
      Run `npx firecrawl-cli <command> --help` to see all available options for a command. The
      common flags are documented below; `--help` covers the rest.
      
      ## Commands
      
      Firecrawl CLI is primarily for **fetching content** from known or discovered URLs. For
      search/discovery, Exa is the primary tool (see `SKILL.md`). Firecrawl search complements
      Exa for keyword matching, site-scoped queries, and content-type filtering.
      
      For fetching, escalate as needed: **scrape** (have a URL) -> **map + scrape** (large site,
      need to find the right page) -> **crawl** (need many pages) -> **browser** (interactive).
      
      ## search
      
      Google-powered keyword search. Complements Exa's semantic search for cases where keyword
      matching, site-scoping, or content-type filtering (`--categories`, `--tbs`) works better.
      The `--scrape` flag fetches full page content for each result.
      
      ```bash
      npx firecrawl-cli search "your query" --scrape --limit 5
      npx firecrawl-cli search "your query" --categories research      # academic sources
      npx firecrawl-cli search "your query" --sources news --tbs qdr:w  # news, past week
      npx firecrawl-cli search "site:docs.example.com auth API"        # site-scoped
      ```
      
      Key flags:
      
      - `--scrape` -- fetch full content for each result
      - `--limit <n>` -- number of results (default 5)
      - `--categories <github|research|pdf>` -- `research` gives Google Scholar-like results
        (not available via MCP server)
      - `--tbs <qdr:h|d|w|m|y>` -- time filter (hour/day/week/month/year)
      - `--sources <web|news|images>` -- source type
      
      Supports Google operators in the query string (`site:`, `intitle:`, etc.).
      
      ## scrape
      
      Extract content from one or more URLs. Multiple URLs are scraped concurrently.
      
      ```bash
      npx firecrawl-cli scrape "https://example.com/page" --only-main-content
      npx firecrawl-cli scrape "https://example.com/page" --wait-for 3000
      npx firecrawl-cli scrape "https://url1" "https://url2" "https://url3" --only-main-content
      ```
      
      Key flags:
      
      - `--only-main-content` -- strip nav, sidebars, footers. **Always use by default** to save
        tokens. If results come back empty or partial, retry without the flag -- it's too aggressive
        on some sites (see `scraping-issues.md` for known false positives).
      - `--wait-for <ms>` -- wait for JS rendering
      - `--include-tags <selector>` / `--exclude-tags <selector>` -- CSS selector targeting
      - `--format <markdown|links|html|screenshot>` -- output format (default: markdown)
      
      Always quote URLs -- shell interprets `?` and `&` as special characters.
      
      ## map
      
      Discover URLs on a site. Use `--search` to filter by keyword.
      
      ```bash
      npx firecrawl-cli map "https://docs.example.com" --search "authentication"
      npx firecrawl-cli map "https://example.com" --limit 100
      ```
      
      Useful when you need a specific subpage on a large site, or when scrape returns empty
      (SPA routing issues).
      
      ## crawl
      
      Bulk-extract from a site section. Set a low `--limit` (5-10) to avoid runaway crawls and
      credit burn.
      
      ```bash
      npx firecrawl-cli crawl "https://docs.example.com" --include-paths /api --limit 10 --wait
      ```
      
      Key flags: `--include-paths`, `--exclude-paths`, `--max-depth <n>`, `--limit <n>`, `--wait`
      
      **Usually better:** `map --search` to find the specific page you need, then `scrape` it.
      More controlled, fewer wasted credits.
      
      ## download
      
      Convenience command: map + scrape in one step. Saves pages as local files.
      
      ```bash
      npx firecrawl-cli download "https://docs.example.com" --include-paths /api --limit 20 -y
      ```
      
      The `-y` flag skips the confirmation prompt. Useful for bulk documentation extraction.
      
      ## browser
      
      Cloud Chromium for interactive pages. Last-resort escalation when scrape fails due to
      content behind clicks, forms, or pagination. Not limited by SAFE_MODE (unlike MCP server).
      
      ```bash
      npx firecrawl-cli browser "open https://example.com"
      npx firecrawl-cli browser "snapshot -i"     # see interactive elements with @ref IDs
      npx firecrawl-cli browser "click @e5"       # interact by ref
      npx firecrawl-cli browser "scrape"          # extract content
      npx firecrawl-cli browser close
      ```
      
      Run `npx firecrawl-cli browser --help` for the full command set (fill, type, eval, scroll,
      screenshot, sessions, profiles).
      
      ## arxiv Extraction
      
      Firecrawl scraping of `arxiv.org/abs/` pages returns abstract + metadata mixed with
      navigation chrome. Better approaches:
      
      - **Triage:** Exa highlights with the default search pattern capture the abstract well
      - **Full paper:** Scrape the HTML version at `arxiv.org/html/{paper-id}` (not `/abs/`)
      - **Structured metadata:** The arxiv API at
        `http://export.arxiv.org/api/query?id_list={id}` returns clean XML
      - Only scrape `/abs/` if just the abstract is needed and Exa highlights are insufficient
      
      ## PDF Scraping
      
      Firecrawl auto-detects PDF URLs and extracts text without any special parameters -- just
      scrape the URL normally. Output is clean markdown with headings, tables, LaTeX math, figure
      captions, and references. Credits scale 1:1 with pages extracted.
      
      For large documents, the API supports a `parsers` parameter to limit pages
      (`parsers: [{"type": "pdf", "maxPages": 10}]`), but this is only available via the MCP
      tool, not CLI flags. Without limits, Firecrawl processes the entire PDF, which can flood
      context and burn credits on long documents.
      
      PDF Parser v2 (Feb 2025) uses a fast Rust-based extractor with automatic OCR fallback for
      scanned documents. No configuration needed.
      
      ## MCP Server
      
      The Firecrawl MCP tools (`firecrawl_scrape`, `firecrawl_search`, `firecrawl_map`, etc.)
      remain available as fallbacks. Main differences from CLI:
      
      - **SAFE_MODE** limits browser actions to `wait`, `screenshot`, `scroll`, `scrape`
      - **No `--categories`** on search (academic search not available)
      - **No `--tbs`** on search (time filtering not available)
      - **`firecrawl_extract`** -- structured data extraction from multiple URLs with a schema.
        Expensive (20-50 credits) and unreliable on thin pages; prefer scrape for single URLs.
      
      Use MCP tools when CLI is unavailable or when you need `firecrawl_extract`.
      
    • people-companies.md 2.9 KB
      # People and Companies
      
      ## When to Use
      
      Finding people (professionals, researchers, founders), researching companies or organizations,
      professional profiles, entity discovery with metadata (headcount, location, funding).
      
      ## People Search
      
      **Primary tool:** Exa `category: "people"` -- surfaces LinkedIn profiles and professional bios.
      
      **Restrictions (severe):**
      
      - No date filters (`startPublishedDate`, `endPublishedDate`)
      - No `includeText` / `excludeText`
      - No `excludeDomains`
      - `includeDomains` accepts **LinkedIn only**
      
      **Workaround:** Put all filtering terms directly in the query string. Include job title,
      location, industry, or organization name for better targeting.
      
      **LinkedIn access:** Exa `people` with `includeDomains: ["linkedin.com"]` is the only reliable
      path. Direct LinkedIn page scraping is unreliable -- treat as effectively inaccessible.
      
      **Full LinkedIn profile text:** Drop the default search pattern (`enableHighlights` and
      `textMaxCharacters: 1`) to get the complete profile from Exa's index -- full work history
      with descriptions, education, skills, and structured entity data. This is the same "Exa full
      text as scraping fallback" pattern used for other inaccessible sites.
      
      ## Company Search
      
      **Primary tool:** Exa `category: "company"` -- returns company homepages with metadata
      (headcount, location, funding stage).
      
      **Restrictions:**
      
      - No date filters
      - No `includeText` / `excludeText`
      - No `includeDomains` / `excludeDomains` (silently ignored -- no error, but filter not applied)
      
      **Workaround:** Put sector, size, geography, and other filtering terms in the query string.
      
      ## Multi-Category Strategy
      
      Specialized categories are best for initial discovery, but each has blind spots. Broaden after
      the initial pass:
      
      1. **Start with specialized category** -- `people` or `company` for targeted discovery
      2. **`personal site`** -- blogs, portfolios, personal pages (full parameter support)
      3. **`news`** -- press mentions, announcements, interviews
      4. **No category** -- full web search for anything the specialized categories miss
      
      Use `additionalQueries` for phrasing variations (e.g., different job titles, company name
      variations).
      
      **Deep-diving a specific person or company:** Drop the category and use `livecrawl: "fallback"`
      for broader results. Example: searching `"Dario Amodei Anthropic CEO background"` with no
      category + `livecrawl: "fallback"` surfaces general web pages, interviews, and analysis that
      the specialized categories miss.
      
      ## Gotchas
      
      - **`company` + `includeDomains`** silently ignores the domain filter. Results come from any
        domain regardless of what you specify. No error thrown.
      - **People search returns profiles, not content.** For someone's actual writing or opinions,
        search without a category and target their name + topic.
      - **LinkedIn is the only scrapeable professional network.** Other professional platforms
        (AngelList, Crunchbase) may work via `npx firecrawl-cli scrape` for individual pages but have no
        specialized Exa category.
      
    • personal-sites.md 2.6 KB
      # Personal Sites and Blogs
      
      ## When to Use
      
      Independent opinions, practitioner experiences, technical blog posts, tutorials, lessons-learned
      content, portfolio sites. Use when you want individual perspectives rather than corporate content
      (news), community consensus (Reddit), or academic rigor (research papers).
      
      Complements Reddit well: Reddit gives you community discussion and quick takes; personal sites
      give you longer-form, more considered individual perspectives on the same topics.
      
      ## Search
      
      **Primary tool:** Exa `category: "personal site"`.
      
      **Full parameter support:** Unlike most specialized categories (`tweet`, `people`, `company`),
      `personal site` supports all parameters -- date filters, `includeText`/`excludeText`, domain
      filters, `additionalQueries`, `livecrawl`, and `subpages`. This makes it one of the most
      flexible categories for targeted search.
      
      **Query tips:**
      
      - Natural language works well: "experienced developer thoughts on switching from React to Svelte"
      - Include the domain or expertise level you want: "senior engineer", "founder", "researcher"
      - Date filters are useful for recent takes on evolving topics (frameworks, AI tools, etc.)
      - `additionalQueries` for phrasing variations (e.g., "lessons learned" vs "things I wish I knew"
        vs "mistakes I made")
      
      ## Domain Filtering
      
      Use `includeDomains` or `excludeDomains` when you have a specific reason:
      
      - **Include specific platforms:** `includeDomains: ["substack.com"]` for newsletter-style writing
      - **Exclude when results are noisy:** if a particular platform is dominating results with
        low-quality content for your query, exclude it and try again
      - **Self-hosted blogs:** no built-in filter for this, but excluding major platforms
        (medium.com, substack.com, dev.to) narrows to people running their own domains
      
      Don't exclude platforms by default -- Medium, Substack, and dev.to host plenty of high-quality
      practitioner writing.
      
      ## Exploring Portfolio Sites
      
      The `subpages` and `subpageTarget` parameters let you explore a site after finding it:
      
      - `subpages: 3` returns up to 3 subpages from each result
      - `subpageTarget: ["projects", "writing"]` targets specific sections
      
      Useful when you find a relevant person and want to see more of their work beyond the landing
      page.
      
      ## Gotchas
      
      - **`includeText`/`excludeText` single-item arrays only.** Multi-item arrays cause 400 errors.
        Put additional filtering terms in the query string instead.
      - **Overlaps with no-category search.** If `personal site` results are thin, try the same query
        without a category -- the broader search may surface blog posts that Exa didn't classify as
        personal sites.
      
    • reddit.md 2.1 KB
      # Reddit
      
      ## When to Use
      
      Community opinions, practitioner views, product recommendations, troubleshooting threads,
      real-world experience reports. Reddit is often the richest source of practitioner views for
      opinion/sentiment and consumer research.
      
      ## Discovery
      
      Finding relevant threads before reading them.
      
      **Primary: Firecrawl search** -- `npx firecrawl-cli search "site:reddit.com {query}"`. Google's
      Reddit index is very good at matching specific entities and surfacing canonical "which X is
      best" threads. Broad `site:reddit.com` with topic keywords works best -- scoping to a specific
      subreddit (`site:reddit.com/r/subreddit`) often returns generic results instead.
      
      **Fallback: Exa** -- `web_search_advanced_exa` with `includeDomains: ["reddit.com"]`. Less
      precise keyword matching but useful when Firecrawl search is rate-limited or returns thin
      results.
      
      ## Reading Threads
      
      Once you have thread URLs from discovery:
      
      - **`get_post_comments`** -- primary tool for discussion content. Set `depth` to control comment
        tree depth. Extract the post ID from the URL found via discovery.
      - **`get_reddit_post`** -- post content + engagement metrics (upvotes, comment count) by ID or URL.
      
      ## Browsing Subreddits
      
      - **`get_top_posts`** with time filter (`week`, `month`, `year`) -- useful for community
        temperature. But top posts are often memes or meta content.
      - For substantive threads, Firecrawl search with topic-specific keywords finds better results
        than browsing top posts.
      
      ## Gotchas
      
      - **Firecrawl blocks all reddit.com domains.** Firecrawl scrape fails on reddit.com, including
        `old.reddit.com` and the Reddit JSON API (`/top.json`, etc.). Use discovery tools to find
        threads, then Reddit MCP tools to read them.
      - **Batch Reddit MCP calls separately.** If any tool call in a parallel batch fails, all sibling
        calls fail too. Keep Reddit MCP calls in their own batch, separate from Exa/Firecrawl calls.
      - **Rate limit:** ~10 requests/minute in anonymous mode. Space out Reddit calls if doing many.
      - **`get_subreddit_info`** provides subscriber counts and description -- useful for gauging
        community size and relevance before diving in.
      
    • scraping-issues.md 6.3 KB
      # Scraping Issues
      
      Troubleshooting guide for Firecrawl scrape failures (CLI or MCP). When the default scrape
      doesn't work, try these approaches before giving up.
      
      Last verified: 2026-02-28.
      
      ## General Troubleshooting
      
      Before consulting per-site workarounds, try these in order:
      
      1. **Retry without `--only-main-content`** (CLI) or with `onlyMainContent: false` (MCP) --
         since `--only-main-content` is the recommended default, this is the first thing to try when
         scrape returns empty or partial results. The detection strips article bodies on some sites
         (Future plc, Blogspot, GDPR-consent-heavy sites).
      2. **Try `--wait-for 5000`** (CLI) or `waitFor: 5000` (MCP) -- helps with JS-rendered SPAs
         (though less often than expected).
      3. **Check for an API** -- many sites have JSON APIs that return cleaner data than scraping.
      4. **Wayback Machine** -- `https://web.archive.org/web/{URL}` for archived/paywalled content.
         Firecrawl follows the redirect to the most recent snapshot automatically. Content
         extraction can be incomplete (some pages render as iframes). Returns 404 if never captured.
      5. **Exa as last resort** -- when scraping fails entirely, Exa often has the page text in
         its index. Drop both `enableHighlights` and `textMaxCharacters: 1` to get the complete
         text.
      
      ## Main-Content Detection False Positives
      
      These sites return HTTP 200 but Firecrawl's content detection strips the article body. Fix
      by omitting `--only-main-content` (CLI) or setting `onlyMainContent: false` (MCP).
      
      ### Future plc sites (Pocket-lint, iMore, TechRadar, PC Gamer, etc.)
      
      - **Failure with main-content detection:** Returns 70-500 chars of skip-links, video player
        embeds, or "Submit a Thread" modals. Article body is stripped entirely.
      - **Without the flag (CLI) / `onlyMainContent: false` (MCP):** Full article content returned
        (may be large, 60KB+).
      - **Note:** Tom's Guide and ZDNET are also Future plc but are policy-blocked (see below).
      
      ### Blogspot / Google Blogger blogs
      
      - **Failure with `--only-main-content`:** Returns only the sidebar "Useful Links" section.
        Blog post feed is stripped as non-main content.
      - **Without the flag:** Full blog feed with titles, dates, summaries, and links.
      - **Affected sites:** Google Workspace Updates blog, and other `*.blogspot.com` / Blogger sites.
      
      ### StorageReview
      
      - **Failure with `--only-main-content`:** Returns GDPR consent wall HTML (cookie consent
        modal rendered inline). Article content behind the overlay is stripped.
      - **Without the flag:** Full review with benchmark tables, specs, and images.
      
      ## Firecrawl Policy-Blocked Sites
      
      These sites are on Firecrawl's infrastructure blocklist. All proxy tiers fail with "we do not
      support this site." No scraping workaround exists.
      
      ### Tom's Guide, ZDNET (Ziff Davis / Future plc network)
      
      - **Failure:** Hard block before any HTTP request is made.
      - **Workaround:** Exa highlights or full text for triage. For deep content, check Wayback
        Machine. No Firecrawl approach works (plain, enhanced all fail identically).
      
      ## Paywalled Sites
      
      ### Medium
      
      - **Workaround:** `https://freedium-mirror.cfd/{FULL_MEDIUM_URL}` via Firecrawl scrape.
        Returns full article content including images.
      - **Note:** The original `freedium.cfd` is dead (DNS failure). The mirror may have intermittent
        uptime. No proxy needed (1 credit).
      
      ### Consumer Reports
      
      - **Free tier:** Category overview pages are accessible (`pagePayState: free` in metadata).
        Product names, images, and recent news articles are visible.
      - **Paywalled:** Ratings pages load the full page structure but silently replace score numbers
        with "log in or sign up" prompts. No HTTP error, no redirect -- the paywall is invisible
        from a scraping perspective.
      - **Workaround:** Note the gap. For product recommendations, use RTINGS (fully open),
        HouseFresh, or AirPurifierFirst as alternatives.
      
      ## Discourse Forums
      
      - Append `.json` to any topic URL for clean structured data (full post content).
      - Search API: `https://{domain}/search.json?q={query}` returns topic IDs to scrape directly.
      - Works well via Firecrawl scrape. Tested on discourse.nixos.org and other Discourse-hosted
        forums.
      
      ## Bluesky
      
      - Bluesky content is on the public web (`bsky.app` URLs). Exa can discover posts.
      - **Profile pages** (`bsky.app/profile/{handle}`) scrape well without `waitFor` -- they return
        bio, follower counts, and recent posts. No proxy needed (1 credit).
      - **Individual post URLs** are unreliable via Firecrawl -- even with `waitFor: 5000`, posts
        often return "Post not found" due to SPA rendering. Use Exa discovery instead.
      
      ## API Workarounds
      
      When scraping fails, these API endpoints return clean structured data:
      
      | Use Case           | API Endpoint                                               | Notes                   |
      | ------------------ | ---------------------------------------------------------- | ----------------------- |
      | npm download stats | `api.npmjs.org/downloads/point/last-week/{pkg}`            | 90 chars, clean JSON    |
      | Discourse search   | `https://{domain}/search.json?q={query}`                   | Returns topic IDs       |
      | Discourse topic    | Append `.json` to any topic URL                            | Full post content       |
      | GitHub file content | `raw.githubusercontent.com/{owner}/{repo}/{branch}/{path}` | Clean, no UI noise     |
      | GitHub repo stats  | `gh api repos/{owner}/{repo}`                              | Stars, forks, last push |
      | arxiv full paper   | `arxiv.org/html/{paper-id}` (not `/abs/`)                  | Large for long papers   |
      | arxiv metadata     | `export.arxiv.org/api/query?id_list={id}`                  | Structured XML          |
      
      ## Sites That Actually Work
      
      These were reported as problematic in researcher feedback but work fine as of 2026-02-28:
      
      | Site              | Status | Notes                                               |
      | ----------------- | ------ | --------------------------------------------------- |
      | npm package pages | Works  | SSR'd, plain scrape gets full README + metadata     |
      | APC/Schneider     | Works  | Redirects to se.com transparently; specs all present |
      | llm-stats.com     | Works  | Next.js SSR renders full leaderboard                |
      | RTINGS            | Works  | Fully open, no paywall; "Insider" only gates voting |
      | rfc-editor.org    | Works  | Returns full RFC (may be very large for long RFCs)  |
      
    • twitter.md 1.7 KB
      # Twitter / X
      
      ## When to Use
      
      Social sentiment, developer opinions, expert threads, trending topics, public reactions to
      announcements. Twitter is informal and context-dependent -- useful for gauging community
      temperature, not for authoritative claims.
      
      ## Tools and Approach
      
      **Primary tool:** Exa `category: "tweet"`. This is the only reliable access path -- x.com is
      behind auth walls and cannot be scraped by Firecrawl or any other tool.
      
      **Key parameters:**
      
      - Date filters work normally (`startPublishedDate`, `endPublishedDate`)
      - `livecrawl: "preferred"` for recent tweets (within hours/days)
      - `additionalQueries` for phrasing variations (e.g., different hashtags, @handles)
      - Consider omitting `enableHighlights` and `textMaxCharacters` -- highlights may not extract
        well from short tweet text. Raw tweet content is usually small enough to return in full.
      
      ## Restrictions
      
      The `tweet` category rejects these parameters with 500 errors:
      
      - `includeText` / `excludeText`
      - `includeDomains` / `excludeDomains`
      - `moderation`
      
      **Workaround:** Put all filtering keywords directly in the query string. Use @handles and
      #hashtags for targeting. For example, instead of `includeText: ["open source"]`, use
      `query: "launching announcing new open source release"`.
      
      ## Gotchas
      
      - **No profile/follower data.** Exa returns tweet content only. Treat profile-level data
        (follower counts, bios) as inaccessible and note the gap.
      - **Tweets are informal.** Context-dependent, sarcastic, abbreviated. Weight accordingly --
        a single tweet is an anecdote, not evidence. Look for convergence across many tweets.
      - **No scraping fallback.** If Exa tweet search doesn't surface what you need, the content
        is effectively inaccessible. Note the gap rather than guessing.
      
  • SKILL.md 9.4 KB
    ---
    name: search-tips
    description: >
      This skill should be used when performing web research beyond a simple single search -- looking
      into topics, comparing options, investigating questions, finding recommendations, or any task
      where effective use of Exa, Firecrawl, and Reddit tools matters. Triggers on "research",
      "look into", "investigate", "compare", "find out about", "search for", "find information",
      "what do people think about", "what are the best", "look up", or multi-source search tasks.
      Also invocable explicitly by deep-research team members via the Skill tool.
    ---
    
    # Search Tips
    
    Accumulated guidance for web research using Exa, Firecrawl CLI, and Reddit MCP tools. These
    are **starting points, not rigid rules** -- think strategically about each situation and
    adapt. If a different approach makes more sense for what you're trying to do, go with it.
    Run `npx firecrawl-cli <command> --help` to check available options beyond what's documented
    here. Reference files cover tool-specific deep dives -- see bottom of this file.
    
    ## Setup
    
    Before starting research, load the required MCP tools using ToolSearch:
    
    1. **Exa tools** -- `web_search_advanced_exa` and `get_code_context_exa`
    2. **Reddit tools** -- `get_top_posts`, `get_post_comments`, `get_reddit_post`, `get_subreddit_info`
    
    Firecrawl CLI (`npx firecrawl-cli`) runs via Bash -- no MCP setup needed. Load only what
    the task requires.
    
    ## The Research Cycle
    
    Prefer Exa and Firecrawl over built-in WebSearch/WebFetch.
    
    Research alternates between **searching** (discovering sources) and **fetching** (extracting
    content from them). Find promising leads, read the best ones, refine your understanding,
    search again.
    
    ### Searching
    
    Finding sources you don't have yet.
    
    - **Exa search** (`web_search_advanced_exa`) -- primary tool for web discovery. Natural
      language queries, add filters as needed (domains, dates, categories).
    - **Exa code context** (`get_code_context_exa`) -- programming topics. Worth trying before
      general Exa search for technical/code tasks -- surfaces repos, packages, and docs.
    - **Firecrawl CLI search** (`npx firecrawl-cli search`) -- Google-powered keyword search.
      Useful when keyword matching works better than Exa's semantic approach, for site-scoped
      queries (`site:reddit.com {query}`), and for content-type filtering (`--categories research`
      for academic, `--sources news`, `--tbs qdr:w` for time).
    
    **Default Exa search pattern:** Default to `enableHighlights: true` and `textMaxCharacters: 1`.
    This returns quoted passages from actual page text while preventing the MCP server from
    flooding context with full text. Use `highlightsPerUrl` and `highlightsNumSentences` to
    control volume if needed.
    
    ### Fetching
    
    Extracting content from a source you've identified.
    
    - **Firecrawl CLI scrape** (`npx firecrawl-cli scrape "<url>" --only-main-content`) -- primary
      tool for reading a known URL. The flag strips nav/sidebars to save tokens.
    - **Reddit MCP** (`get_post_comments`, `get_reddit_post`) -- for reading Reddit threads.
      Firecrawl can't scrape reddit.com directly.
    - **Firecrawl CLI map** (`npx firecrawl-cli map "<url>" --search "query"`) -- discover URLs
      on a site (useful when you need to find the right page, or when scrape returns empty).
    
    ### Adapting the Workflow
    
    The defaults above won't always be right. Some common deviations:
    
    - **Exa full text as a scraping fallback** -- some sites are blocked or inaccessible via
      Firecrawl (LinkedIn, Twitter/X, etc.), but Exa often has the full page text in its index.
      Drop both `enableHighlights` and `textMaxCharacters: 1` to get the complete text. Be aware
      this can produce large responses.
    - **`--only-main-content` can strip too much** -- if you got empty or partial results, retry
      without the flag. Known to fail on Future plc sites, Blogspot, and GDPR-heavy sites. See
      Content Extraction below.
    
    The reference files cover more edge cases -- scraping issues, category restrictions, and
    academic search.
    
    ## Search Strategy
    
    ### How Exa Works
    
    Exa is a **neural/semantic search engine**. It uses embeddings to understand meaning.
    
    - **Natural questions or statements work best** -- Exa finds pages that answer them
    - **Longer, more specific queries work BETTER** -- unlike keyword-based search
    - **Keyword lists tend to confuse** the semantic model
    
    Good: "What do professional reviewers say are the most reliable dishwasher brands in 2025?"
    Bad: "best dishwasher 2025 reliable"
    
    ### Query Reformulation
    
    For broad topics, Exa's `additionalQueries` parameter can automate this -- it bundles query
    variations in a single call at no extra cost (see `references/exa-tips.md`). For manual
    reformulation, try generating 3-5 query variations:
    
    | Technique             | What It Does                          | Example                                      |
    | --------------------- | ------------------------------------- | -------------------------------------------- |
    | **Paraphrase**        | Same meaning, different words         | "RAG failures" -> "problems in RAG systems"  |
    | **Decompose**         | Break into sub-questions              | "Why fail?" -> "Why return irrelevant docs?" |
    | **Scope shift**       | Broader context or narrower specifics | "Challenges in production AI search"         |
    | **Perspective shift** | Different viewpoints                  | User vs expert vs critic view                |
    | **Temporal framing**  | Target different time periods         | "Recent 2024-2025" vs "foundational"         |
    
    ### Domain Filtering
    
    Try `includeDomains` when the authoritative site for a topic is known -- faster and less noisy
    than broad search. Try `excludeDomains` to suppress sites that keep appearing but aren't
    useful (e.g., exclude `youtube.com` when video pages crowd out needed editorial content about 
    YouTube creators/content).
    
    ### Searching by Content Type
    
    Match your research target to the right approach. Reference files have full strategies.
    
    - **Social sentiment / opinions** -- Exa `tweet` for Twitter; `site:reddit.com` via Firecrawl
      search + Reddit MCP for discussions. See `references/twitter.md`, `references/reddit.md`.
    - **People / companies** -- Exa `people` or `company` category for discovery, then broaden.
      See `references/people-companies.md`.
    - **Academic papers** -- Exa `research paper` category; academic APIs for structured data.
      See `references/academic-search.md`.
    - **Code / GitHub** -- `get_code_context_exa` for code; `gh api` for repo data.
      See `references/code-github.md`.
    - **News** -- Exa `news` category with date filters; Firecrawl CLI
      `search --sources news --tbs qdr:w` for Google News with time filtering.
    - **Financial reports** -- Exa `financial report` with domain/date filtering.
    - **Personal blogs / independent takes** -- Exa `personal site` for practitioner opinions,
      blog posts, and independent analysis. Full parameter support (unlike most specialized
      categories). See `references/personal-sites.md`.
    
    Some categories reject certain parameters (400/500 errors). See `references/exa-tips.md`.
    
    ## Content Extraction
    
    Always use `--only-main-content` by default -- it strips nav, sidebars, and footers, saving
    significant tokens. If you get empty or partial results, retry without the flag; it's known
    to strip article bodies on Future plc sites (iMore, Pocket-lint), GDPR-heavy sites
    (StorageReview), and Blogspot blogs -- see `references/scraping-issues.md`.
    
    For even narrower extraction, use `--include-tags` (e.g., `--include-tags "article"`,
    `--include-tags ".post-content"`).
    
    ## Scraping Issues
    
    When Firecrawl scrape fails (policy blocks, paywalls, SPA rendering), check
    `references/scraping-issues.md` for per-site workarounds. General fallback: try Exa full text
    (drop `enableHighlights` and `textMaxCharacters`). For interactive pages that need clicks or
    form fills, escalate to `npx firecrawl-cli browser` (see `references/firecrawl-tips.md`).
    
    ## Operational Notes
    
    **Parallel call failures:** If any tool call in a parallel batch fails, all sibling calls
    fail too. Retry individually. Keep failure-prone calls (Reddit MCP) in their own batch.
    
    **Rate limits:** Retry once after a short pause. If still blocked, try a different tool for the
    same intent before giving up -- Exa rate-limited? Try Firecrawl search. Reddit MCP throttled?
    Try Exa with `includeDomains: ["reddit.com"]`. Only note the gap and move on after both the
    original tool and an alternative have failed.
    
    ## Reference Files
    
    ### Content-type strategies
    - **`references/twitter.md`** -- Exa tweet category, restrictions, query tips, livecrawl
    - **`references/reddit.md`** -- Discovery via Firecrawl search, Reddit MCP for reading,
      batching gotchas, rate limits
    - **`references/people-companies.md`** -- Exa people/company categories, LinkedIn, multi-
      category approach
    - **`references/code-github.md`** -- Code search, GitHub repos/issues, gh api, raw URLs
    - **`references/personal-sites.md`** -- Independent blogs, practitioner opinions, full
      parameter support, portfolio exploration
    - **`references/academic-search.md`** -- Academic domains, free APIs, Firecrawl CLI
    
    ### Tool reference
    - **`references/exa-tips.md`** -- Category restrictions, additionalQueries,
      highlights/summaries
    - **`references/firecrawl-tips.md`** -- CLI commands (search, scrape, map, crawl, download,
      browser), PDF scraping, arxiv extraction, MCP fallback notes
    
    ### Troubleshooting
    - **`references/scraping-issues.md`** -- main-content detection fixes, paywalls, Discourse
      JSON, policy-blocked sites, API workarounds
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related