search-tips
This skill should be used when performing web research beyond a simple single search -- looking into topics, comparing options, investigating questions, finding recommendations, or any task where effective use of Exa, Firecrawl, and Reddit tools matters. Triggers on "research", "
Install
npx skills add https://github.com/malob/nix-config/tree/master/configs/claude/skills/search-tips
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install malob-nix-config@llmmart
git clone https://github.com/malob/nix-config.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole malob/nix-config collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Search Tips
Accumulated guidance for web research using Exa, Firecrawl CLI, and Reddit MCP tools. These
are starting points, not rigid rules -- think strategically about each situation and
adapt. If a different approach makes more sense for what you're trying to do, go with it.
Run npx firecrawl-cli <command> --help to check available options beyond what's documented
here. Reference files cover tool-specific deep dives -- see bottom of this file.
Setup
Before starting research, load the required MCP tools using ToolSearch:
- Exa tools --
web_search_advanced_exaandget_code_context_exa - Reddit tools --
get_top_posts,get_post_comments,get_reddit_post,get_subreddit_info
Firecrawl CLI (npx firecrawl-cli) runs via Bash -- no MCP setup needed. Load only what
the task requires.
The Research Cycle
Prefer Exa and Firecrawl over built-in WebSearch/WebFetch.
Research alternates between searching (discovering sources) and fetching (extracting content from them). Find promising leads, read the best ones, refine your understanding, search again.
Searching
Finding sources you don't have yet.
- Exa search (
web_search_advanced_exa) -- primary tool for web discovery. Natural language queries, add filters as needed (domains, dates, categories). - Exa code context (
get_code_context_exa) -- programming topics. Worth trying before general Exa search for technical/code tasks -- surfaces repos, packages, and docs. - Firecrawl CLI search (
npx firecrawl-cli search) -- Google-powered keyword search. Useful when keyword matching works better than Exa's semantic approach, for site-scoped queries (site:reddit.com {query}), and for content-type filtering (--categories researchfor academic,--sources news,--tbs qdr:wfor time).
Default Exa search pattern: Default to enableHighlights: true and textMaxCharacters: 1.
This returns quoted passages from actual page text while preventing the MCP server from
flooding context with full text. Use highlightsPerUrl and highlightsNumSentences to
control volume if needed.
Fetching
Extracting content from a source you've identified.
- Firecrawl CLI scrape (
npx firecrawl-cli scrape "<url>" --only-main-content) -- primary tool for reading a known URL. The flag strips nav/sidebars to save tokens. - Reddit MCP (
get_post_comments,get_reddit_post) -- for reading Reddit threads. Firecrawl can't scrape reddit.com directly. - Firecrawl CLI map (
npx firecrawl-cli map "<url>" --search "query") -- discover URLs on a site (useful when you need to find the right page, or when scrape returns empty).
Adapting the Workflow
The defaults above won't always be right. Some common deviations:
- Exa full text as a scraping fallback -- some sites are blocked or inaccessible via
Firecrawl (LinkedIn, Twitter/X, etc.), but Exa often has the full page text in its index.
Drop both
enableHighlightsandtextMaxCharacters: 1to get the complete text. Be aware this can produce large responses. --only-main-contentcan strip too much -- if you got empty or partial results, retry without the flag. Known to fail on Future plc sites, Blogspot, and GDPR-heavy sites. See Content Extraction below.
The reference files cover more edge cases -- scraping issues, category restrictions, and academic search.
Search Strategy
How Exa Works
Exa is a neural/semantic search engine. It uses embeddings to understand meaning.
- Natural questions or statements work best -- Exa finds pages that answer them
- Longer, more specific queries work BETTER -- unlike keyword-based search
- Keyword lists tend to confuse the semantic model
Good: "What do professional reviewers say are the most reliable dishwasher brands in 2025?" Bad: "best dishwasher 2025 reliable"
Query Reformulation
For broad topics, Exa's additionalQueries parameter can automate this -- it bundles query
variations in a single call at no extra cost (see references/exa-tips.md). For manual
reformulation, try generating 3-5 query variations:
| Technique | What It Does | Example |
|---|---|---|
| Paraphrase | Same meaning, different words | "RAG failures" -> "problems in RAG systems" |
| Decompose | Break into sub-questions | "Why fail?" -> "Why return irrelevant docs?" |
| Scope shift | Broader context or narrower specifics | "Challenges in production AI search" |
| Perspective shift | Different viewpoints | User vs expert vs critic view |
| Temporal framing | Target different time periods | "Recent 2024-2025" vs "foundational" |
Domain Filtering
Try includeDomains when the authoritative site for a topic is known -- faster and less noisy
than broad search. Try excludeDomains to suppress sites that keep appearing but aren't
useful (e.g., exclude youtube.com when video pages crowd out needed editorial content about
YouTube creators/content).
Searching by Content Type
Match your research target to the right approach. Reference files have full strategies.
- Social sentiment / opinions -- Exa
tweetfor Twitter;site:reddit.comvia Firecrawl search + Reddit MCP for discussions. Seereferences/twitter.md,references/reddit.md. - People / companies -- Exa
peopleorcompanycategory for discovery, then broaden. Seereferences/people-companies.md. - Academic papers -- Exa
research papercategory; academic APIs for structured data. Seereferences/academic-search.md. - Code / GitHub --
get_code_context_exafor code;gh apifor repo data. Seereferences/code-github.md. - News -- Exa
newscategory with date filters; Firecrawl CLIsearch --sources news --tbs qdr:wfor Google News with time filtering. - Financial reports -- Exa
financial reportwith domain/date filtering. - Personal blogs / independent takes -- Exa
personal sitefor practitioner opinions, blog posts, and independent analysis. Full parameter support (unlike most specialized categories). Seereferences/personal-sites.md.
Some categories reject certain parameters (400/500 errors). See references/exa-tips.md.
Content Extraction
Always use --only-main-content by default -- it strips nav, sidebars, and footers, saving
significant tokens. If you get empty or partial results, retry without the flag; it's known
to strip article bodies on Future plc sites (iMore, Pocket-lint), GDPR-heavy sites
(StorageReview), and Blogspot blogs -- see references/scraping-issues.md.
For even narrower extraction, use --include-tags (e.g., --include-tags "article",
--include-tags ".post-content").
Scraping Issues
When Firecrawl scrape fails (policy blocks, paywalls, SPA rendering), check
references/scraping-issues.md for per-site workarounds. General fallback: try Exa full text
(drop enableHighlights and textMaxCharacters). For interactive pages that need clicks or
form fills, escalate to npx firecrawl-cli browser (see references/firecrawl-tips.md).
Operational Notes
Parallel call failures: If any tool call in a parallel batch fails, all sibling calls fail too. Retry individually. Keep failure-prone calls (Reddit MCP) in their own batch.
Rate limits: Retry once after a short pause. If still blocked, try a different tool for the
same intent before giving up -- Exa rate-limited? Try Firecrawl search. Reddit MCP throttled?
Try Exa with includeDomains: ["reddit.com"]. Only note the gap and move on after both the
original tool and an alternative have failed.
Reference Files
Content-type strategies
references/twitter.md-- Exa tweet category, restrictions, query tips, livecrawlreferences/reddit.md-- Discovery via Firecrawl search, Reddit MCP for reading, batching gotchas, rate limitsreferences/people-companies.md-- Exa people/company categories, LinkedIn, multi- category approachreferences/code-github.md-- Code search, GitHub repos/issues, gh api, raw URLsreferences/personal-sites.md-- Independent blogs, practitioner opinions, full parameter support, portfolio explorationreferences/academic-search.md-- Academic domains, free APIs, Firecrawl CLI
Tool reference
references/exa-tips.md-- Category restrictions, additionalQueries, highlights/summariesreferences/firecrawl-tips.md-- CLI commands (search, scrape, map, crawl, download, browser), PDF scraping, arxiv extraction, MCP fallback notes
Troubleshooting
references/scraping-issues.md-- main-content detection fixes, paywalls, Discourse JSON, policy-blocked sites, API workarounds
Files (nix-config)
-
references
-
academic-search.md 2.9 KB
# Academic Search ## When to Use Literature reviews, evidence gathering, prior art research, finding methodology details, systematic exploration of a research field or topic. ## Tools and Approach **Primary tool:** Exa `category: "research paper"` -- surfaces arXiv papers, conference proceedings, and journal articles for research-oriented queries. **Key parameters:** - `includeDomains` for specific venues: `arxiv.org`, `openreview.net`, `nature.com`, `semanticscholar.org`, `pubmed.ncbi.nlm.nih.gov` - `startPublishedDate` / `endPublishedDate` for temporal scope (e.g., last 2 years for state-of-the-art, broader for foundational work) - `highlightsQuery` for targeted extraction of methodology, findings, or specific aspects - `additionalQueries` for phrasing variations (different technical terms for the same concept) The default search pattern (`enableHighlights: true`, `textMaxCharacters: 1`) works well for triage -- highlights extract abstracts and key passages effectively. The undocumented `pdf` category also works -- returns papers from arxiv, researchgate, icml, aclanthology, and similar sources. Useful as an alternative when `research paper` misses relevant results. ## Academic Domains Useful values for `includeDomains`: - `arxiv.org` -- preprints (CS, physics, math, biology, economics) - `openreview.net` -- peer review and conference papers (ICLR, NeurIPS) - `semanticscholar.org` -- 214M papers across all disciplines - `pubmed.ncbi.nlm.nih.gov` -- medical/biomedical literature - `biorxiv.org` -- biology preprints - `nature.com`, `science.org` -- high-impact journals ## Free APIs These return clean structured data via Firecrawl scrape or `curl` (no auth needed): - **Semantic Scholar:** `https://api.semanticscholar.org/graph/v1/paper/search?query={query}` - **OpenAlex:** `https://api.openalex.org/works?search={query}` - **arXiv:** `http://export.arxiv.org/api/query?search_query=all:{query}` For structured metadata (citation counts, authors, abstracts), these APIs are more reliable than scraping paper pages. ## Firecrawl CLI Academic Search `npx firecrawl-cli search --categories research "{query}"` provides Google Scholar-like academic search. Returns results from arxiv, nature.com, and similar sources. ```bash npx firecrawl-cli search --categories research "transformer attention mechanisms" --scrape --limit 5 npx firecrawl-cli search --categories research "RLHF alignment" --tbs qdr:y # past year ``` The `--categories research` flag is CLI-only (not available via MCP). See `firecrawl-tips.md` for more CLI commands. ## Gotchas - **arxiv `/abs/` pages** return abstract + metadata mixed with navigation chrome. Better to scrape `arxiv.org/html/{paper-id}` for the full paper or use Exa highlights for triage. See `firecrawl-tips.md` for details. - **Paywalled journals** (Nature, Science, Elsevier) may return only abstracts. Check for preprint versions on arxiv or author pages. The Wayback Machine sometimes has cached copies. -
code-github.md 3.1 KB
# Code and GitHub ## When to Use Code examples, library documentation, package APIs, GitHub repos/issues/discussions, open-source project research, technical implementation patterns. ## Code Search **Primary tool:** `get_code_context_exa` -- searches code-related content including repos, packages, documentation, Stack Overflow, and technical blogs. **Query tips:** - Always include the **programming language** ("Go generics" not just "generics") - Include framework + version when applicable ("Next.js 15", "React 19", "Python 3.12") - Use exact identifiers: function names, error messages, class names, config keys - More specific queries perform better than broad ones (Exa is semantic, not keyword-based) Good: "How to use React Server Components with Next.js 15 app router for data fetching" Bad: "react server components nextjs" **`tokensNum` tuning:** Controls how much context is returned per result. - Focused snippet needed: 1000-3000 - Most tasks (default): 5000 - Complex integration patterns: 10000-20000 ## GitHub Discovery `get_code_context_exa` also searches GitHub repos, issues, discussions, and READMEs. This is the fastest path for finding relevant GitHub content -- no special permissions needed. ## GitHub Deep Read When you need the full content of a specific GitHub page: - **Firecrawl scrape** (`npx firecrawl-cli scrape "<url>"`) works well for GitHub issues, discussions, READMEs, and wiki pages - **Raw file content:** `raw.githubusercontent.com/{owner}/{repo}/{branch}/{path}` for clean file content without UI noise. Scrape this URL via Firecrawl. ## GitHub Data via `gh` CLI Prefer `gh` subcommands over `gh api` -- subcommands are auto-approved while `gh api` with write flags (`-X`, `-f`, `-F`) requires user approval. - `gh repo view {owner}/{repo} --json stargazerCount,forkCount,pushedAt,licenseInfo` - `gh issue list -R {owner}/{repo} --state open --json number,title,labels,createdAt` - `gh issue view {number} -R {owner}/{repo} --json body,comments` - `gh pr list -R {owner}/{repo} --json number,title,state` -- check for pending fixes - `gh pr view {number} -R {owner}/{repo} --json body,comments,files` - `gh search repos "{query}" --json fullName,stargazersCount,description` - `gh search code "{pattern}" --repo {owner}/{repo}` - `gh release list -R {owner}/{repo}` - `gh release view {tag} -R {owner}/{repo}` For data without subcommand equivalents, `gh api` GET requests are also auto-approved: - `gh api repos/{owner}/{repo}/stats/participation` -- weekly commit activity - `gh api repos/{owner}/{repo}/contributors` -- contributors with commit counts ## Gotchas - **GitHub rate limits:** The `gh api` commands use your authenticated token, so limits are generous (5000/hour). Unauthenticated API calls are 60/hour. - **Large repos:** READMEs and issue threads can be very long. Use `npx firecrawl-cli scrape "<url>" --only-main-content` to get the main content without sidebar noise. - **GitHub search via Exa vs `gh`:** Exa is better for semantic discovery ("libraries for X"). `gh search repos "query"` is better for exact keyword matches and quantitative filtering (stars, language, license). -
exa-tips.md 3.5 KB
# Exa Tips Deep-dive reference for Exa search parameters. For everyday usage, see `SKILL.md`. ## Category Restrictions Some categories reject certain parameters with 400/500 errors. Others silently ignore them (no error, but filter not applied). | Category | Unsupported parameters | | ------------------ | ----------------------------------------------------------------------------------------------- | | `tweet` | `includeText`, `excludeText`, `includeDomains`, `excludeDomains`, `moderation` (causes 500) | | `company` | All date filters, `includeText`, `excludeText`, `includeDomains`, `excludeDomains` | | `people` | All date filters, `includeText`, `excludeText`, `excludeDomains`. `includeDomains` LinkedIn only | | `financial report` | `excludeText` (causes 400). `includeText` works (single-item only) | | `research paper` | Full support | | `personal site` | Full support | | `news` | Full support | | `pdf` | Undocumented category -- available in MCP, returns academic papers (arxiv, researchgate, etc.) | | `github` | Undocumented category -- available in MCP, returns GitHub repos | **`company` + `includeDomains`** silently ignores the domain filter (returns results from any domain). No error thrown. ## includeText / excludeText Limits - Only accept **single-item arrays** (one string each) - Each string must be **5 words or fewer** - Multi-item arrays or strings longer than 5 words cause 400 errors - Put additional filtering terms in the `query` string instead ## additionalQueries `additionalQueries: ["variation 1", "variation 2"]` bundles query variations in a single call. Same cost as a single query. Works with `auto` mode (our MCP tool exposes `auto`/`fast`/`neural`). - **Broad topics** (e.g., "LLM hallucination mitigation"): genuine diversity gain, surfaces results the primary query misses. Worth using. - **Niche/specific topics** (e.g., "Nix flakes best practices"): results largely converge. May surface a few extra sources but core results stay the same. Worth trying for research tasks where comprehensive coverage of a broad topic matters. Less useful for targeted lookups. Can substitute for manual query reformulation on broad topics. ## Highlights (Default Search Pattern) `enableHighlights: true` + `textMaxCharacters: 1` is the recommended default for Exa search. Highlights are extracted server-side before text truncation, so you get actual quoted passages from the page without the full-text context flood. - `highlightsPerUrl` -- control how many highlight blocks per result - `highlightsNumSentences` -- control sentences per highlight - `highlightsQuery` -- focus highlights on a specific aspect of the page - Without `textMaxCharacters: 1`, you get full text (7K-32K chars) alongside highlights -- best to pair them. ## Summaries `enableSummary: true` + `textMaxCharacters: 1` generates AI summaries per result. These are concise but **can hallucinate** -- they sometimes fabricate details not present on the page. Highlights are preferred because they return actual page text. Summaries exist as an option but there's rarely a reason to use them over highlights. -
firecrawl-tips.md 5.9 KB
# Firecrawl Tips Deep-dive reference for Firecrawl CLI. For everyday usage, see `SKILL.md`. Run `npx firecrawl-cli <command> --help` to see all available options for a command. The common flags are documented below; `--help` covers the rest. ## Commands Firecrawl CLI is primarily for **fetching content** from known or discovered URLs. For search/discovery, Exa is the primary tool (see `SKILL.md`). Firecrawl search complements Exa for keyword matching, site-scoped queries, and content-type filtering. For fetching, escalate as needed: **scrape** (have a URL) -> **map + scrape** (large site, need to find the right page) -> **crawl** (need many pages) -> **browser** (interactive). ## search Google-powered keyword search. Complements Exa's semantic search for cases where keyword matching, site-scoping, or content-type filtering (`--categories`, `--tbs`) works better. The `--scrape` flag fetches full page content for each result. ```bash npx firecrawl-cli search "your query" --scrape --limit 5 npx firecrawl-cli search "your query" --categories research # academic sources npx firecrawl-cli search "your query" --sources news --tbs qdr:w # news, past week npx firecrawl-cli search "site:docs.example.com auth API" # site-scoped ``` Key flags: - `--scrape` -- fetch full content for each result - `--limit <n>` -- number of results (default 5) - `--categories <github|research|pdf>` -- `research` gives Google Scholar-like results (not available via MCP server) - `--tbs <qdr:h|d|w|m|y>` -- time filter (hour/day/week/month/year) - `--sources <web|news|images>` -- source type Supports Google operators in the query string (`site:`, `intitle:`, etc.). ## scrape Extract content from one or more URLs. Multiple URLs are scraped concurrently. ```bash npx firecrawl-cli scrape "https://example.com/page" --only-main-content npx firecrawl-cli scrape "https://example.com/page" --wait-for 3000 npx firecrawl-cli scrape "https://url1" "https://url2" "https://url3" --only-main-content ``` Key flags: - `--only-main-content` -- strip nav, sidebars, footers. **Always use by default** to save tokens. If results come back empty or partial, retry without the flag -- it's too aggressive on some sites (see `scraping-issues.md` for known false positives). - `--wait-for <ms>` -- wait for JS rendering - `--include-tags <selector>` / `--exclude-tags <selector>` -- CSS selector targeting - `--format <markdown|links|html|screenshot>` -- output format (default: markdown) Always quote URLs -- shell interprets `?` and `&` as special characters. ## map Discover URLs on a site. Use `--search` to filter by keyword. ```bash npx firecrawl-cli map "https://docs.example.com" --search "authentication" npx firecrawl-cli map "https://example.com" --limit 100 ``` Useful when you need a specific subpage on a large site, or when scrape returns empty (SPA routing issues). ## crawl Bulk-extract from a site section. Set a low `--limit` (5-10) to avoid runaway crawls and credit burn. ```bash npx firecrawl-cli crawl "https://docs.example.com" --include-paths /api --limit 10 --wait ``` Key flags: `--include-paths`, `--exclude-paths`, `--max-depth <n>`, `--limit <n>`, `--wait` **Usually better:** `map --search` to find the specific page you need, then `scrape` it. More controlled, fewer wasted credits. ## download Convenience command: map + scrape in one step. Saves pages as local files. ```bash npx firecrawl-cli download "https://docs.example.com" --include-paths /api --limit 20 -y ``` The `-y` flag skips the confirmation prompt. Useful for bulk documentation extraction. ## browser Cloud Chromium for interactive pages. Last-resort escalation when scrape fails due to content behind clicks, forms, or pagination. Not limited by SAFE_MODE (unlike MCP server). ```bash npx firecrawl-cli browser "open https://example.com" npx firecrawl-cli browser "snapshot -i" # see interactive elements with @ref IDs npx firecrawl-cli browser "click @e5" # interact by ref npx firecrawl-cli browser "scrape" # extract content npx firecrawl-cli browser close ``` Run `npx firecrawl-cli browser --help` for the full command set (fill, type, eval, scroll, screenshot, sessions, profiles). ## arxiv Extraction Firecrawl scraping of `arxiv.org/abs/` pages returns abstract + metadata mixed with navigation chrome. Better approaches: - **Triage:** Exa highlights with the default search pattern capture the abstract well - **Full paper:** Scrape the HTML version at `arxiv.org/html/{paper-id}` (not `/abs/`) - **Structured metadata:** The arxiv API at `http://export.arxiv.org/api/query?id_list={id}` returns clean XML - Only scrape `/abs/` if just the abstract is needed and Exa highlights are insufficient ## PDF Scraping Firecrawl auto-detects PDF URLs and extracts text without any special parameters -- just scrape the URL normally. Output is clean markdown with headings, tables, LaTeX math, figure captions, and references. Credits scale 1:1 with pages extracted. For large documents, the API supports a `parsers` parameter to limit pages (`parsers: [{"type": "pdf", "maxPages": 10}]`), but this is only available via the MCP tool, not CLI flags. Without limits, Firecrawl processes the entire PDF, which can flood context and burn credits on long documents. PDF Parser v2 (Feb 2025) uses a fast Rust-based extractor with automatic OCR fallback for scanned documents. No configuration needed. ## MCP Server The Firecrawl MCP tools (`firecrawl_scrape`, `firecrawl_search`, `firecrawl_map`, etc.) remain available as fallbacks. Main differences from CLI: - **SAFE_MODE** limits browser actions to `wait`, `screenshot`, `scroll`, `scrape` - **No `--categories`** on search (academic search not available) - **No `--tbs`** on search (time filtering not available) - **`firecrawl_extract`** -- structured data extraction from multiple URLs with a schema. Expensive (20-50 credits) and unreliable on thin pages; prefer scrape for single URLs. Use MCP tools when CLI is unavailable or when you need `firecrawl_extract`. -
people-companies.md 2.9 KB
# People and Companies ## When to Use Finding people (professionals, researchers, founders), researching companies or organizations, professional profiles, entity discovery with metadata (headcount, location, funding). ## People Search **Primary tool:** Exa `category: "people"` -- surfaces LinkedIn profiles and professional bios. **Restrictions (severe):** - No date filters (`startPublishedDate`, `endPublishedDate`) - No `includeText` / `excludeText` - No `excludeDomains` - `includeDomains` accepts **LinkedIn only** **Workaround:** Put all filtering terms directly in the query string. Include job title, location, industry, or organization name for better targeting. **LinkedIn access:** Exa `people` with `includeDomains: ["linkedin.com"]` is the only reliable path. Direct LinkedIn page scraping is unreliable -- treat as effectively inaccessible. **Full LinkedIn profile text:** Drop the default search pattern (`enableHighlights` and `textMaxCharacters: 1`) to get the complete profile from Exa's index -- full work history with descriptions, education, skills, and structured entity data. This is the same "Exa full text as scraping fallback" pattern used for other inaccessible sites. ## Company Search **Primary tool:** Exa `category: "company"` -- returns company homepages with metadata (headcount, location, funding stage). **Restrictions:** - No date filters - No `includeText` / `excludeText` - No `includeDomains` / `excludeDomains` (silently ignored -- no error, but filter not applied) **Workaround:** Put sector, size, geography, and other filtering terms in the query string. ## Multi-Category Strategy Specialized categories are best for initial discovery, but each has blind spots. Broaden after the initial pass: 1. **Start with specialized category** -- `people` or `company` for targeted discovery 2. **`personal site`** -- blogs, portfolios, personal pages (full parameter support) 3. **`news`** -- press mentions, announcements, interviews 4. **No category** -- full web search for anything the specialized categories miss Use `additionalQueries` for phrasing variations (e.g., different job titles, company name variations). **Deep-diving a specific person or company:** Drop the category and use `livecrawl: "fallback"` for broader results. Example: searching `"Dario Amodei Anthropic CEO background"` with no category + `livecrawl: "fallback"` surfaces general web pages, interviews, and analysis that the specialized categories miss. ## Gotchas - **`company` + `includeDomains`** silently ignores the domain filter. Results come from any domain regardless of what you specify. No error thrown. - **People search returns profiles, not content.** For someone's actual writing or opinions, search without a category and target their name + topic. - **LinkedIn is the only scrapeable professional network.** Other professional platforms (AngelList, Crunchbase) may work via `npx firecrawl-cli scrape` for individual pages but have no specialized Exa category. -
personal-sites.md 2.6 KB
# Personal Sites and Blogs ## When to Use Independent opinions, practitioner experiences, technical blog posts, tutorials, lessons-learned content, portfolio sites. Use when you want individual perspectives rather than corporate content (news), community consensus (Reddit), or academic rigor (research papers). Complements Reddit well: Reddit gives you community discussion and quick takes; personal sites give you longer-form, more considered individual perspectives on the same topics. ## Search **Primary tool:** Exa `category: "personal site"`. **Full parameter support:** Unlike most specialized categories (`tweet`, `people`, `company`), `personal site` supports all parameters -- date filters, `includeText`/`excludeText`, domain filters, `additionalQueries`, `livecrawl`, and `subpages`. This makes it one of the most flexible categories for targeted search. **Query tips:** - Natural language works well: "experienced developer thoughts on switching from React to Svelte" - Include the domain or expertise level you want: "senior engineer", "founder", "researcher" - Date filters are useful for recent takes on evolving topics (frameworks, AI tools, etc.) - `additionalQueries` for phrasing variations (e.g., "lessons learned" vs "things I wish I knew" vs "mistakes I made") ## Domain Filtering Use `includeDomains` or `excludeDomains` when you have a specific reason: - **Include specific platforms:** `includeDomains: ["substack.com"]` for newsletter-style writing - **Exclude when results are noisy:** if a particular platform is dominating results with low-quality content for your query, exclude it and try again - **Self-hosted blogs:** no built-in filter for this, but excluding major platforms (medium.com, substack.com, dev.to) narrows to people running their own domains Don't exclude platforms by default -- Medium, Substack, and dev.to host plenty of high-quality practitioner writing. ## Exploring Portfolio Sites The `subpages` and `subpageTarget` parameters let you explore a site after finding it: - `subpages: 3` returns up to 3 subpages from each result - `subpageTarget: ["projects", "writing"]` targets specific sections Useful when you find a relevant person and want to see more of their work beyond the landing page. ## Gotchas - **`includeText`/`excludeText` single-item arrays only.** Multi-item arrays cause 400 errors. Put additional filtering terms in the query string instead. - **Overlaps with no-category search.** If `personal site` results are thin, try the same query without a category -- the broader search may surface blog posts that Exa didn't classify as personal sites. -
reddit.md 2.1 KB
# Reddit ## When to Use Community opinions, practitioner views, product recommendations, troubleshooting threads, real-world experience reports. Reddit is often the richest source of practitioner views for opinion/sentiment and consumer research. ## Discovery Finding relevant threads before reading them. **Primary: Firecrawl search** -- `npx firecrawl-cli search "site:reddit.com {query}"`. Google's Reddit index is very good at matching specific entities and surfacing canonical "which X is best" threads. Broad `site:reddit.com` with topic keywords works best -- scoping to a specific subreddit (`site:reddit.com/r/subreddit`) often returns generic results instead. **Fallback: Exa** -- `web_search_advanced_exa` with `includeDomains: ["reddit.com"]`. Less precise keyword matching but useful when Firecrawl search is rate-limited or returns thin results. ## Reading Threads Once you have thread URLs from discovery: - **`get_post_comments`** -- primary tool for discussion content. Set `depth` to control comment tree depth. Extract the post ID from the URL found via discovery. - **`get_reddit_post`** -- post content + engagement metrics (upvotes, comment count) by ID or URL. ## Browsing Subreddits - **`get_top_posts`** with time filter (`week`, `month`, `year`) -- useful for community temperature. But top posts are often memes or meta content. - For substantive threads, Firecrawl search with topic-specific keywords finds better results than browsing top posts. ## Gotchas - **Firecrawl blocks all reddit.com domains.** Firecrawl scrape fails on reddit.com, including `old.reddit.com` and the Reddit JSON API (`/top.json`, etc.). Use discovery tools to find threads, then Reddit MCP tools to read them. - **Batch Reddit MCP calls separately.** If any tool call in a parallel batch fails, all sibling calls fail too. Keep Reddit MCP calls in their own batch, separate from Exa/Firecrawl calls. - **Rate limit:** ~10 requests/minute in anonymous mode. Space out Reddit calls if doing many. - **`get_subreddit_info`** provides subscriber counts and description -- useful for gauging community size and relevance before diving in. -
scraping-issues.md 6.3 KB
# Scraping Issues Troubleshooting guide for Firecrawl scrape failures (CLI or MCP). When the default scrape doesn't work, try these approaches before giving up. Last verified: 2026-02-28. ## General Troubleshooting Before consulting per-site workarounds, try these in order: 1. **Retry without `--only-main-content`** (CLI) or with `onlyMainContent: false` (MCP) -- since `--only-main-content` is the recommended default, this is the first thing to try when scrape returns empty or partial results. The detection strips article bodies on some sites (Future plc, Blogspot, GDPR-consent-heavy sites). 2. **Try `--wait-for 5000`** (CLI) or `waitFor: 5000` (MCP) -- helps with JS-rendered SPAs (though less often than expected). 3. **Check for an API** -- many sites have JSON APIs that return cleaner data than scraping. 4. **Wayback Machine** -- `https://web.archive.org/web/{URL}` for archived/paywalled content. Firecrawl follows the redirect to the most recent snapshot automatically. Content extraction can be incomplete (some pages render as iframes). Returns 404 if never captured. 5. **Exa as last resort** -- when scraping fails entirely, Exa often has the page text in its index. Drop both `enableHighlights` and `textMaxCharacters: 1` to get the complete text. ## Main-Content Detection False Positives These sites return HTTP 200 but Firecrawl's content detection strips the article body. Fix by omitting `--only-main-content` (CLI) or setting `onlyMainContent: false` (MCP). ### Future plc sites (Pocket-lint, iMore, TechRadar, PC Gamer, etc.) - **Failure with main-content detection:** Returns 70-500 chars of skip-links, video player embeds, or "Submit a Thread" modals. Article body is stripped entirely. - **Without the flag (CLI) / `onlyMainContent: false` (MCP):** Full article content returned (may be large, 60KB+). - **Note:** Tom's Guide and ZDNET are also Future plc but are policy-blocked (see below). ### Blogspot / Google Blogger blogs - **Failure with `--only-main-content`:** Returns only the sidebar "Useful Links" section. Blog post feed is stripped as non-main content. - **Without the flag:** Full blog feed with titles, dates, summaries, and links. - **Affected sites:** Google Workspace Updates blog, and other `*.blogspot.com` / Blogger sites. ### StorageReview - **Failure with `--only-main-content`:** Returns GDPR consent wall HTML (cookie consent modal rendered inline). Article content behind the overlay is stripped. - **Without the flag:** Full review with benchmark tables, specs, and images. ## Firecrawl Policy-Blocked Sites These sites are on Firecrawl's infrastructure blocklist. All proxy tiers fail with "we do not support this site." No scraping workaround exists. ### Tom's Guide, ZDNET (Ziff Davis / Future plc network) - **Failure:** Hard block before any HTTP request is made. - **Workaround:** Exa highlights or full text for triage. For deep content, check Wayback Machine. No Firecrawl approach works (plain, enhanced all fail identically). ## Paywalled Sites ### Medium - **Workaround:** `https://freedium-mirror.cfd/{FULL_MEDIUM_URL}` via Firecrawl scrape. Returns full article content including images. - **Note:** The original `freedium.cfd` is dead (DNS failure). The mirror may have intermittent uptime. No proxy needed (1 credit). ### Consumer Reports - **Free tier:** Category overview pages are accessible (`pagePayState: free` in metadata). Product names, images, and recent news articles are visible. - **Paywalled:** Ratings pages load the full page structure but silently replace score numbers with "log in or sign up" prompts. No HTTP error, no redirect -- the paywall is invisible from a scraping perspective. - **Workaround:** Note the gap. For product recommendations, use RTINGS (fully open), HouseFresh, or AirPurifierFirst as alternatives. ## Discourse Forums - Append `.json` to any topic URL for clean structured data (full post content). - Search API: `https://{domain}/search.json?q={query}` returns topic IDs to scrape directly. - Works well via Firecrawl scrape. Tested on discourse.nixos.org and other Discourse-hosted forums. ## Bluesky - Bluesky content is on the public web (`bsky.app` URLs). Exa can discover posts. - **Profile pages** (`bsky.app/profile/{handle}`) scrape well without `waitFor` -- they return bio, follower counts, and recent posts. No proxy needed (1 credit). - **Individual post URLs** are unreliable via Firecrawl -- even with `waitFor: 5000`, posts often return "Post not found" due to SPA rendering. Use Exa discovery instead. ## API Workarounds When scraping fails, these API endpoints return clean structured data: | Use Case | API Endpoint | Notes | | ------------------ | ---------------------------------------------------------- | ----------------------- | | npm download stats | `api.npmjs.org/downloads/point/last-week/{pkg}` | 90 chars, clean JSON | | Discourse search | `https://{domain}/search.json?q={query}` | Returns topic IDs | | Discourse topic | Append `.json` to any topic URL | Full post content | | GitHub file content | `raw.githubusercontent.com/{owner}/{repo}/{branch}/{path}` | Clean, no UI noise | | GitHub repo stats | `gh api repos/{owner}/{repo}` | Stars, forks, last push | | arxiv full paper | `arxiv.org/html/{paper-id}` (not `/abs/`) | Large for long papers | | arxiv metadata | `export.arxiv.org/api/query?id_list={id}` | Structured XML | ## Sites That Actually Work These were reported as problematic in researcher feedback but work fine as of 2026-02-28: | Site | Status | Notes | | ----------------- | ------ | --------------------------------------------------- | | npm package pages | Works | SSR'd, plain scrape gets full README + metadata | | APC/Schneider | Works | Redirects to se.com transparently; specs all present | | llm-stats.com | Works | Next.js SSR renders full leaderboard | | RTINGS | Works | Fully open, no paywall; "Insider" only gates voting | | rfc-editor.org | Works | Returns full RFC (may be very large for long RFCs) | -
twitter.md 1.7 KB
# Twitter / X ## When to Use Social sentiment, developer opinions, expert threads, trending topics, public reactions to announcements. Twitter is informal and context-dependent -- useful for gauging community temperature, not for authoritative claims. ## Tools and Approach **Primary tool:** Exa `category: "tweet"`. This is the only reliable access path -- x.com is behind auth walls and cannot be scraped by Firecrawl or any other tool. **Key parameters:** - Date filters work normally (`startPublishedDate`, `endPublishedDate`) - `livecrawl: "preferred"` for recent tweets (within hours/days) - `additionalQueries` for phrasing variations (e.g., different hashtags, @handles) - Consider omitting `enableHighlights` and `textMaxCharacters` -- highlights may not extract well from short tweet text. Raw tweet content is usually small enough to return in full. ## Restrictions The `tweet` category rejects these parameters with 500 errors: - `includeText` / `excludeText` - `includeDomains` / `excludeDomains` - `moderation` **Workaround:** Put all filtering keywords directly in the query string. Use @handles and #hashtags for targeting. For example, instead of `includeText: ["open source"]`, use `query: "launching announcing new open source release"`. ## Gotchas - **No profile/follower data.** Exa returns tweet content only. Treat profile-level data (follower counts, bios) as inaccessible and note the gap. - **Tweets are informal.** Context-dependent, sarcastic, abbreviated. Weight accordingly -- a single tweet is an anecdote, not evidence. Look for convergence across many tweets. - **No scraping fallback.** If Exa tweet search doesn't surface what you need, the content is effectively inaccessible. Note the gap rather than guessing.
-
-
SKILL.md 9.4 KB
--- name: search-tips description: > This skill should be used when performing web research beyond a simple single search -- looking into topics, comparing options, investigating questions, finding recommendations, or any task where effective use of Exa, Firecrawl, and Reddit tools matters. Triggers on "research", "look into", "investigate", "compare", "find out about", "search for", "find information", "what do people think about", "what are the best", "look up", or multi-source search tasks. Also invocable explicitly by deep-research team members via the Skill tool. --- # Search Tips Accumulated guidance for web research using Exa, Firecrawl CLI, and Reddit MCP tools. These are **starting points, not rigid rules** -- think strategically about each situation and adapt. If a different approach makes more sense for what you're trying to do, go with it. Run `npx firecrawl-cli <command> --help` to check available options beyond what's documented here. Reference files cover tool-specific deep dives -- see bottom of this file. ## Setup Before starting research, load the required MCP tools using ToolSearch: 1. **Exa tools** -- `web_search_advanced_exa` and `get_code_context_exa` 2. **Reddit tools** -- `get_top_posts`, `get_post_comments`, `get_reddit_post`, `get_subreddit_info` Firecrawl CLI (`npx firecrawl-cli`) runs via Bash -- no MCP setup needed. Load only what the task requires. ## The Research Cycle Prefer Exa and Firecrawl over built-in WebSearch/WebFetch. Research alternates between **searching** (discovering sources) and **fetching** (extracting content from them). Find promising leads, read the best ones, refine your understanding, search again. ### Searching Finding sources you don't have yet. - **Exa search** (`web_search_advanced_exa`) -- primary tool for web discovery. Natural language queries, add filters as needed (domains, dates, categories). - **Exa code context** (`get_code_context_exa`) -- programming topics. Worth trying before general Exa search for technical/code tasks -- surfaces repos, packages, and docs. - **Firecrawl CLI search** (`npx firecrawl-cli search`) -- Google-powered keyword search. Useful when keyword matching works better than Exa's semantic approach, for site-scoped queries (`site:reddit.com {query}`), and for content-type filtering (`--categories research` for academic, `--sources news`, `--tbs qdr:w` for time). **Default Exa search pattern:** Default to `enableHighlights: true` and `textMaxCharacters: 1`. This returns quoted passages from actual page text while preventing the MCP server from flooding context with full text. Use `highlightsPerUrl` and `highlightsNumSentences` to control volume if needed. ### Fetching Extracting content from a source you've identified. - **Firecrawl CLI scrape** (`npx firecrawl-cli scrape "<url>" --only-main-content`) -- primary tool for reading a known URL. The flag strips nav/sidebars to save tokens. - **Reddit MCP** (`get_post_comments`, `get_reddit_post`) -- for reading Reddit threads. Firecrawl can't scrape reddit.com directly. - **Firecrawl CLI map** (`npx firecrawl-cli map "<url>" --search "query"`) -- discover URLs on a site (useful when you need to find the right page, or when scrape returns empty). ### Adapting the Workflow The defaults above won't always be right. Some common deviations: - **Exa full text as a scraping fallback** -- some sites are blocked or inaccessible via Firecrawl (LinkedIn, Twitter/X, etc.), but Exa often has the full page text in its index. Drop both `enableHighlights` and `textMaxCharacters: 1` to get the complete text. Be aware this can produce large responses. - **`--only-main-content` can strip too much** -- if you got empty or partial results, retry without the flag. Known to fail on Future plc sites, Blogspot, and GDPR-heavy sites. See Content Extraction below. The reference files cover more edge cases -- scraping issues, category restrictions, and academic search. ## Search Strategy ### How Exa Works Exa is a **neural/semantic search engine**. It uses embeddings to understand meaning. - **Natural questions or statements work best** -- Exa finds pages that answer them - **Longer, more specific queries work BETTER** -- unlike keyword-based search - **Keyword lists tend to confuse** the semantic model Good: "What do professional reviewers say are the most reliable dishwasher brands in 2025?" Bad: "best dishwasher 2025 reliable" ### Query Reformulation For broad topics, Exa's `additionalQueries` parameter can automate this -- it bundles query variations in a single call at no extra cost (see `references/exa-tips.md`). For manual reformulation, try generating 3-5 query variations: | Technique | What It Does | Example | | --------------------- | ------------------------------------- | -------------------------------------------- | | **Paraphrase** | Same meaning, different words | "RAG failures" -> "problems in RAG systems" | | **Decompose** | Break into sub-questions | "Why fail?" -> "Why return irrelevant docs?" | | **Scope shift** | Broader context or narrower specifics | "Challenges in production AI search" | | **Perspective shift** | Different viewpoints | User vs expert vs critic view | | **Temporal framing** | Target different time periods | "Recent 2024-2025" vs "foundational" | ### Domain Filtering Try `includeDomains` when the authoritative site for a topic is known -- faster and less noisy than broad search. Try `excludeDomains` to suppress sites that keep appearing but aren't useful (e.g., exclude `youtube.com` when video pages crowd out needed editorial content about YouTube creators/content). ### Searching by Content Type Match your research target to the right approach. Reference files have full strategies. - **Social sentiment / opinions** -- Exa `tweet` for Twitter; `site:reddit.com` via Firecrawl search + Reddit MCP for discussions. See `references/twitter.md`, `references/reddit.md`. - **People / companies** -- Exa `people` or `company` category for discovery, then broaden. See `references/people-companies.md`. - **Academic papers** -- Exa `research paper` category; academic APIs for structured data. See `references/academic-search.md`. - **Code / GitHub** -- `get_code_context_exa` for code; `gh api` for repo data. See `references/code-github.md`. - **News** -- Exa `news` category with date filters; Firecrawl CLI `search --sources news --tbs qdr:w` for Google News with time filtering. - **Financial reports** -- Exa `financial report` with domain/date filtering. - **Personal blogs / independent takes** -- Exa `personal site` for practitioner opinions, blog posts, and independent analysis. Full parameter support (unlike most specialized categories). See `references/personal-sites.md`. Some categories reject certain parameters (400/500 errors). See `references/exa-tips.md`. ## Content Extraction Always use `--only-main-content` by default -- it strips nav, sidebars, and footers, saving significant tokens. If you get empty or partial results, retry without the flag; it's known to strip article bodies on Future plc sites (iMore, Pocket-lint), GDPR-heavy sites (StorageReview), and Blogspot blogs -- see `references/scraping-issues.md`. For even narrower extraction, use `--include-tags` (e.g., `--include-tags "article"`, `--include-tags ".post-content"`). ## Scraping Issues When Firecrawl scrape fails (policy blocks, paywalls, SPA rendering), check `references/scraping-issues.md` for per-site workarounds. General fallback: try Exa full text (drop `enableHighlights` and `textMaxCharacters`). For interactive pages that need clicks or form fills, escalate to `npx firecrawl-cli browser` (see `references/firecrawl-tips.md`). ## Operational Notes **Parallel call failures:** If any tool call in a parallel batch fails, all sibling calls fail too. Retry individually. Keep failure-prone calls (Reddit MCP) in their own batch. **Rate limits:** Retry once after a short pause. If still blocked, try a different tool for the same intent before giving up -- Exa rate-limited? Try Firecrawl search. Reddit MCP throttled? Try Exa with `includeDomains: ["reddit.com"]`. Only note the gap and move on after both the original tool and an alternative have failed. ## Reference Files ### Content-type strategies - **`references/twitter.md`** -- Exa tweet category, restrictions, query tips, livecrawl - **`references/reddit.md`** -- Discovery via Firecrawl search, Reddit MCP for reading, batching gotchas, rate limits - **`references/people-companies.md`** -- Exa people/company categories, LinkedIn, multi- category approach - **`references/code-github.md`** -- Code search, GitHub repos/issues, gh api, raw URLs - **`references/personal-sites.md`** -- Independent blogs, practitioner opinions, full parameter support, portfolio exploration - **`references/academic-search.md`** -- Academic domains, free APIs, Firecrawl CLI ### Tool reference - **`references/exa-tips.md`** -- Category restrictions, additionalQueries, highlights/summaries - **`references/firecrawl-tips.md`** -- CLI commands (search, scrape, map, crawl, download, browser), PDF scraping, arxiv extraction, MCP fallback notes ### Troubleshooting - **`references/scraping-issues.md`** -- main-content detection fixes, paywalls, Discourse JSON, policy-blocked sites, API workarounds
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.