Doc Scraper

Go web crawler to scrape documentation sites and convert content to clean Markdown for LLM ingestion (RAG, training data).

LLM Mart 0 views 1 listing impressions
Transport
Not stated
Package
—
Registry id
—

No install snippet on purpose. A working MCP config is a command, its arguments and an environment block — the last two are where API keys live, so this catalogue never stores them and cannot publish them. Follow the link above for the authors' own instructions.

mcp

A configurable, concurrent, and resumable web crawler written in Go. Specifically designed to scrape technical documentation websites, extract core content, convert it cleanly to Markdown format suitable for ingestion by Large Language Models (LLMs), and save the results locally.

doc-scraper crawling a docs site and answering search queries offline

Overview

This project provides a powerful command-line tool to crawl documentation sites based on settings defined in a config.yaml file. It navigates the site structure, extracts content from specified HTML sections using CSS selectors, and converts it into clean Markdown files.

Why Use This Tool?

  • Built for LLM Training & RAG Systems - Creates clean, consistent Markdown optimized for ingestion
  • Preserves Documentation Structure - Maintains the original site hierarchy for context preservation
  • Production-Ready Features - Offers resumable crawls, rate limiting, and graceful error handling
  • High Performance - Uses Go's concurrency model for efficient parallel processing

Goal: Preparing Documentation for LLMs

The main objective of this tool is to automate the often tedious process of gathering and cleaning web-based documentation for use with Large Language Models. By converting structured web content into clean Markdown, it aims to provide a dataset that is:

  • Text-Focused: Prioritizes the textual content extracted via CSS selectors
  • Structured: Maintains the directory hierarchy of the original documentation site, preserving context
  • Cleaned: Converts HTML to Markdown, removing web-specific markup and clutter
  • Locally Accessible: Provides the content as local files for easier processing and pipeline integration

Key Features

Feature Description
Configurable Crawling Uses YAML for global and site-specific settings
Scope Control Limits crawling by domain, path prefix, and disallowed path patterns (regex)
Content Extraction Extracts main content using CSS selectors
HTML-to-Markdown Converts extracted HTML to clean GitHub-Flavored Markdown (tables, task lists, strikethrough)
Image Handling Opt-in downloading and local rewriting of image links with domain and size filtering (disabled by default; doc-scraper is text-first)
Link Rewriting Rewrites internal links to relative paths for local structure
JSONL Output Optional one-record-per-page JSONL with a trailing crawl-summary record, for RAG ingestion
Concurrency Configurable worker pools and semaphore-based request limits (global and per-host)
Rate Limiting Configurable per-host delays with jitter
Robots.txt & Sitemaps Respects robots.txt and processes discovered sitemaps
State Persistence Uses BadgerDB for state; supports resuming crawls via crawl --resume
Graceful Shutdown Handles SIGINT/SIGTERM with proper cleanup
HTTP Retries Exponential backoff with jitter for transient errors
Observability Structured logging (log/slog); optional pprof endpoint (build with -tags pprof)
Modular Code Organized into packages for clarity and maintainability
CLI Utilities Built-in config validate and config list commands for configuration management
MCP Server Mode Expose as Model Context Protocol server for Claude Code/Cursor integration
Full-Text Search Offline BM25 search over crawled docs (SQLite FTS5) via the search_docs MCP tool
Auto Content Detection Automatic detection of 30+ documentation frameworks (Docusaurus, MkDocs, Sphinx, VitePress, GitBook, and more) with readability fallback
Parallel Site Crawling Crawl multiple sites concurrently with shared resource management
Watch Mode Scheduled periodic re-crawling with state persistence

Getting Started

Prerequisites

From the project's README.

Related servers

MCP server for Geargrafx PC Engine / TurboGrafx-16 emulator

15 views

Read-only discovery for NeuralNg Angular components, APIs, packages, icons and theme recipes.

14 views

Umami v3 MCP for Cloud or self-hosted analytics, with read-only, privacy-conscious defaults.

12 views

Read and write RAGE Package Format archives

12 views