Claude
Skill
pdf-extractor
This skill should be used when the user asks to "extract text from PDF", "convert PDF to text", "parse PDF", "read PDF contents", "extract data from documents", "batch PDF extraction", "PDF to markdown", "OCR PDF", "get text from PDF files", "I have a PDF", "can you read this PDF
Virus-scanned
Reviewed automatically before listing.
Download
ahundt-autorun-plugins_pdf-extractor_skills_pdf-extractor-6fb6027.zip · 7 KB
Install
skills CLI
npx skills add https://github.com/ahundt/autorun/tree/main/plugins/pdf-extractor/skills/pdf-extractor
Claude Code
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install ahundt-autorun@llmmart
Git
git clone https://github.com/ahundt/autorun.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole ahundt/autorun collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
PDF Data Extraction
Files (autorun)
-
references
-
backends.md 7.2 KB
# PDF Extraction Backend Reference Detailed comparison and selection guide for the 9 PDF extraction backends. ## Modern Backends (2024-2025) ### MarkItDown (Microsoft) **License:** MIT **Dependencies:** `markitdown>=0.1.0` **Models:** None (lightweight) **Strengths:** - Fast, lightweight extraction - Good for general documents - No model downloads required - MIT license allows commercial use **Weaknesses:** - Limited OCR capability - May miss complex layouts **Best for:** General text documents, forms, simple layouts **Usage:** ```python from markitdown import MarkItDown md = MarkItDown() result = md.convert('/path/to/document.pdf') text = result.text_content ``` ### Docling (IBM) **License:** MIT **Dependencies:** `docling>=2.0.0` **Models:** ~500MB (downloaded on first use) **Strengths:** - Excellent layout analysis - Table detection and extraction - Figure handling - GPU acceleration available **Weaknesses:** - Large model download - Slower than lightweight backends - Requires more memory **Best for:** Complex layouts, academic papers, reports with figures/tables **Usage:** ```python from docling.document_converter import DocumentConverter converter = DocumentConverter() result = converter.convert('/path/to/document.pdf') markdown = result.document.export_to_markdown() ``` ### Marker (datalab-to) **License:** GPL-3.0 (copyleft) **Dependencies:** separately managed `marker-pdf` installation **Models:** vision models downloaded by marker Marker is not selected by a published extra because its supported-platform dependency graph pins Pillow below the first fully patched release. **Strengths:** - Best for scanned documents - Vision-based approach - Handles complex layouts - GPU-accelerated **Weaknesses:** - GPL license restricts commercial use - Large model download - Slowest backend - Requires significant GPU memory **Best for:** Scanned documents, PDFs from images, complex layouts **Usage:** ```python from marker.converters.pdf import PdfConverter from marker.models import create_model_dict converter = PdfConverter(artifact_dict=create_model_dict()) result = converter('/path/to/document.pdf') markdown = result.markdown ``` ## Traditional Backends ### PDFPlumber **License:** MIT **Dependencies:** `pdfplumber>=0.10.0` **Models:** None **Strengths:** - Excellent table extraction - Precise character positioning - Good for structured documents - MIT license **Weaknesses:** - Can be slow on large documents - Limited formatting preservation **Best for:** Tables, forms, invoices, structured data **Usage:** ```python import pdfplumber with pdfplumber.open('/path/to/document.pdf') as pdf: text = '\n'.join([page.extract_text() for page in pdf.pages]) ``` ### PDFMiner **License:** MIT **Dependencies:** `pdfminer.six>=20221105` **Models:** None **Strengths:** - Pure Python, always available - Reliable text extraction - Good layout analysis - MIT license **Weaknesses:** - No table extraction - Older codebase **Best for:** Simple text documents, fallback extraction **Usage:** ```python from pdfminer.high_level import extract_text text = extract_text('/path/to/document.pdf') ``` ### pypdf (`pypdf2` CLI backend id) **License:** BSD-3 **Dependencies:** `pypdf>=6.0.0` **Models:** None **Strengths:** - Included in the `cpu` extra - Fast - BSD license (permissive) - Good for encryption detection **Weaknesses:** - Basic text extraction only - Poor formatting preservation - May miss text in complex layouts **Best for:** Last-resort fallback, encryption checking **Usage:** ```python import pypdf with open('/path/to/document.pdf', 'rb') as f: reader = pypdf.PdfReader(f) text = ''.join([page.extract_text() for page in reader.pages]) ``` ### PyMuPDF4LLM **License:** AGPL-3.0 (copyleft) **Dependencies:** `pymupdf4llm>=0.0.1` **Models:** None **Strengths:** - LLM-optimized markdown output - Fast extraction - Good structure preservation **Weaknesses:** - AGPL license limits commercial use - Less tested than other backends **Best for:** LLM pipelines, document Q&A systems **Usage:** ```python import pymupdf4llm markdown = pymupdf4llm.to_markdown('/path/to/document.pdf') ``` ### PDFBox **License:** Apache-2.0 **Dependencies:** `pdfbox>=0.1.0` (Java wrapper) **Models:** None **Strengths:** - Mature, well-tested - Good for tables - Apache license **Weaknesses:** - Requires Java runtime - Installation can be problematic - Slower due to JVM overhead **Best for:** Table extraction when pdfplumber fails **Usage:** ```python from pdfbox import PDFBox pdfbox = PDFBox() pdfbox.extract_text('/path/to/document.pdf', '/path/to/output.txt') ``` ### Pdftotext (CLI) **License:** GPL (Poppler) **Dependencies:** System `pdftotext` command **Models:** None **Strengths:** - Very fast - Good text preservation - Layout mode available **Weaknesses:** - Requires system installation - Not available everywhere - GPL license **Best for:** Simple text extraction, system scripts **Usage:** ```bash pdftotext -layout document.pdf output.txt ``` ## Selection Decision Tree ``` Start ├── Is PDF scanned/image-based? │ ├── Yes → Use marker (GPU) or docling │ └── No → Continue ├── Does PDF have tables? │ ├── Yes → Use pdfplumber or pdfbox │ └── No → Continue ├── Is speed critical? │ ├── Yes → Use markitdown or pdfminer │ └── No → Continue ├── Is commercial use required? │ ├── Yes → Avoid marker (GPL), pymupdf4llm (AGPL) │ └── No → Any backend works └── Default → markitdown → pdfplumber → pdfminer → pypdf2 ``` ## License Summary | Backend | License | Commercial Use | |---------|---------|----------------| | markitdown | MIT | Yes | | docling | MIT | Yes | | marker | GPL-3.0 | No (copyleft) | | pymupdf4llm | AGPL-3.0 | No (copyleft) | | pdfplumber | MIT | Yes | | pdfminer | MIT | Yes | | pypdf2 | BSD-3 | Yes | | pdfbox | Apache-2.0 | Yes | | pdftotext | GPL | No (copyleft) | ## Performance Benchmarks Approximate extraction times for a 10-page document: | Backend | Time | Memory | |---------|------|--------| | pypdf2 | 0.5s | Low | | pdfminer | 1s | Low | | pdftotext | 0.3s | Low | | markitdown | 1s | Low | | pdfplumber | 2s | Medium | | pdfbox | 3s | Medium | | docling | 10s | High | | marker | 30s | Very High | | pymupdf4llm | 1s | Medium | Note: GPU-accelerated backends (docling, marker) are much faster with CUDA. ## Troubleshooting ### Backend Not Available ```python # Check if backend is importable try: from markitdown import MarkItDown print("markitdown available") except ImportError: print("markitdown not installed") ``` ### Installation Issues **pdfbox:** Requires Java 8+ ```bash java -version # Check Java installed pip install pdfbox ``` **marker:** Separately managed only. It is intentionally excluded from the published extras while its dependency graph requires an unpatched Pillow. **docling:** Large download ```bash pip install docling # First run downloads ~500MB models ``` ### Common Errors **"No module named 'X'"**: Backend not installed ```bash pip install X ``` **"Java not found"**: PDFBox needs Java ```bash brew install openjdk # macOS apt install default-jre # Ubuntu ``` **"CUDA out of memory"**: Reduce batch size or use CPU ```python backends = ['markitdown', 'pdfplumber'] # CPU-only ```
-
-
SKILL.md 12.6 KB
--- name: pdf-extractor description: This skill should be used when the user asks to "extract text from PDF", "convert PDF to text", "parse PDF", "read PDF contents", "extract data from documents", "batch PDF extraction", "PDF to markdown", "OCR PDF", "get text from PDF files", "I have a PDF", "can you read this PDF", "what's in this PDF", "summarize this PDF", "open PDF file", "extract from [filename].pdf", or needs to process PDF documents for data extraction. Handles single-file extraction, batch processing, and OCR for scanned documents with automatic backend selection. metadata: version: 1.0.0rc3 example-prompt: "Extract text from document.pdf" --- # PDF Data Extraction <purpose> Extract text and structured data from PDF documents using a multi-backend approach with automatic fallback. ## Overview This skill provides PDF text extraction with 9 different backends, automatic GPU detection, and intelligent backend selection. The extraction system tries backends in order until one succeeds, producing markdown output optimized for further processing. </purpose> <workflow> ## Quick Start Workflow To extract text from PDFs: 1. **Single file extraction (installed CLI - recommended):** ```bash extract-pdfs /path/to/document.pdf ``` Output: Creates `document.md` in the same directory. 2. **Batch extraction (directory):** ```bash extract-pdfs /path/to/pdfs/ /path/to/output/ ``` Output: Creates `.md` files for all PDFs in output directory. 3. **Custom output file:** ```bash extract-pdfs document.pdf output.md ``` 4. **Specific backends:** ```bash extract-pdfs document.pdf --backends markitdown pdfplumber ``` 5. **List available backends:** ```bash extract-pdfs --list-backends ``` Output: Shows available backends and GPU status. ### Alternative Execution Methods If the `extract-pdfs` CLI isn't installed, install it first (recommended): ```bash # Install as global UV tool (from repo root); the code ships in the autorun-ai distribution: cd "${CLAUDE_PLUGIN_ROOT}/../.." && uv tool install --force --editable "./plugins/autorun[pdf]" extract-pdfs --list-backends # verify ``` Or use these fallback methods without installing: ```bash # From a source checkout, without installing (the package lives in plugins/autorun): uv run --project plugins/autorun --extra pdf python -m pdf_extraction document.pdf # Module execution, if the console script is not on PATH python -m pdf_extraction document.pdf ``` ## Backend Selection Guide ### Custom Backend Ordering Specify backends in any order with `--backends`. The system tries each in order, stopping on first success: ```bash # Tables first, then general extraction extract-pdfs document.pdf --backends pdfplumber markitdown pdfminer # Scanned documents: vision-based first extract-pdfs scanned.pdf --backends docling markitdown # Most permissive fallback order (handles problematic PDFs) extract-pdfs document.pdf --backends pdfminer pypdf2 markitdown # Single backend only (no fallback) extract-pdfs document.pdf --backends markitdown ``` ### CPU-Only Systems (Default) For systems without GPU, the recommended backend order: - `markitdown` - Microsoft's lightweight converter (MIT, fast, no models) - `pdfplumber` - Excellent for tables (MIT) - `pdfminer` - Pure Python, reliable (MIT) - `pypdf2` - Basic extraction through maintained `pypdf` (BSD-3; `pdf` extra) ### GPU Systems For systems with CUDA-enabled GPU: - `docling` - IBM layout analysis (MIT, downloads models on first use) - Plus all CPU backends as fallback The `marker` backend is recognized only for separately managed installations; it is not selected by a published extra because its dependency graph pins an unpatched Pillow release. ### Backend Comparison | Backend | License | Models | Best For | Speed | |---------|---------|--------|----------|-------| | markitdown | MIT | None | General text, forms | Fast | | pdfplumber | MIT | None | Tables, structured data | Fast | | pdfminer | MIT | None | Simple text documents | Fast | | pypdf2 | BSD-3 | None | Basic extraction | Fast | | docling | MIT | ~500MB | Layout analysis | Medium | | marker | GPL-3.0 | ~1GB | Scanned documents | Slow | | pymupdf4llm | AGPL-3.0 | None | LLM-optimized output | Fast | | pdfbox | Apache-2.0 | None | Tables (Java-based) | Medium | | pdftotext | System | None | Simple text (CLI) | Fast | ### Backend Decision Matrix | Document Type | Recommended Backend(s) | Why | |---------------|------------------------|-----| | Digital text PDF (default) | markitdown, pdfplumber | Fast, accurate | | PDF with tables/invoices | pdfplumber, pdfbox | Best table structure | | Complex layouts/columns | docling (GPU) | Layout analysis | | Scanned documents/images | marker, docling (GPU) | OCR/vision required | | Insurance policies/forms | markitdown, pdfplumber | Handles form fields | | Academic papers | docling | Equations, figures | | Maximum compatibility | pdfminer, pypdf2 | Fewest dependencies | | Commercial use required | markitdown, pdfplumber | MIT license | ## Programmatic Usage To use the extraction library directly in Python code: ```python from pdf_extraction import extract_single_pdf, pdf_to_txt, detect_gpu_availability # Check available backends gpu_info = detect_gpu_availability() print(f"Recommended backends: {gpu_info['recommended_backends']}") # Extract single file result = extract_single_pdf( input_file='/path/to/document.pdf', output_file='/path/to/output.md', backends=['markitdown', 'pdfplumber'] ) if result['success']: print(f"Extracted with {result['backend_used']}") print(f"Quality metrics: {result['quality_metrics']}") # Batch extract directory output_files, metadata = pdf_to_txt( input_dir='/path/to/pdfs/', output_dir='/path/to/output/', resume=True, # Skip already-extracted files return_metadata=True ) ``` </workflow> <reference> ## Extraction Metadata Every extraction returns metadata for quality assessment: ```python { 'success': True, 'backend_used': 'markitdown', 'extraction_time_seconds': 2.5, 'output_size_bytes': 15234, 'quality_metrics': { 'char_count': 15234, 'line_count': 450, 'word_count': 2800, 'table_markers': 12, # Count of | (tables) 'has_structure': True # Has markdown structure }, 'encrypted': False, 'error': None } ``` ## Handling Common Scenarios ### Encrypted PDFs The system detects encrypted PDFs and reports them: ```python if result['encrypted']: print("PDF is password-protected") ``` Encrypted PDFs cannot be extracted without the password. ### Empty or Failed Extractions When all backends fail: 1. Check if PDF is encrypted 2. Try with `--backends pdfminer pypdf2` (most permissive) 3. Check PDF isn't corrupted 4. Consider OCR-based backends for scanned documents ### Resume Batch Processing To continue interrupted batch extraction: ```bash extract-pdfs /path/to/pdfs/ /path/to/output/ ``` The `resume=True` default skips already-extracted files. To force re-extraction: ```bash extract-pdfs /path/to/pdfs/ --no-resume ``` ### Tables and Structured Data For PDFs with tables, prioritize: ```bash extract-pdfs document.pdf --backends pdfplumber markitdown ``` The output will contain markdown tables when detected: ```markdown | Column1 | Column2 | Column3 | |---------|---------|---------| | Data | Data | Data | ``` ## Module Structure Reference ### Source Code Layout **Location:** `plugins/pdf-extractor/src/pdf_extraction/` in a source checkout, and the importable `pdf_extraction` package once `autorun` is installed. This plugin directory holds the manifest, command, and this skill; the code ships inside the `autorun-ai` distribution behind its `pdf` extra. | File | Purpose | |------|---------| | `__init__.py` | Package exports (extract_single_pdf, pdf_to_txt, etc.) | | `__main__.py` | Support for `python -m pdf_extraction` | | `cli.py` | CLI entry point with argparse | | `backends.py` | BackendExtractor base class + 9 backend implementations | | `extractors.py` | extract_single_pdf(), pdf_to_txt() functions | | `utils.py` | GPU detection, quality metrics, encryption check | ### Key Classes and Functions | Component | Location | Purpose | |-----------|----------|---------| | `BackendExtractor` | backends.py:35-123 | Base class with Template Method pattern | | `DoclingExtractor` | backends.py:130-142 | IBM Docling backend (MIT, GPU) | | `MarkerExtractor` | backends.py:145-158 | Vision-based marker backend (GPL-3.0, GPU) | | `MarkItDownExtractor` | backends.py:161-173 | Microsoft MarkItDown (MIT, CPU) | | `PdfplumberExtractor` | backends.py:244-253 | Table-focused extraction (MIT) | | `PdfminerExtractor` | backends.py:219-226 | Pure Python fallback (MIT) | | `Pypdf2Extractor` | backends.py:229-241 | Basic extraction through optional `pypdf` (BSD-3) | | `BACKEND_REGISTRY` | backends.py:279-292 | Dict mapping backend names to factories | | `detect_gpu_availability()` | utils.py:9-40 | Auto-detect GPU and recommend backends | | `extract_single_pdf()` | extractors.py:13-80 | Extract one PDF with backend fallback | | `pdf_to_txt()` | extractors.py:83-170 | Batch extract directory with resume | **Key implementation details:** - Backend fallback loop: `extractors.py:55-78` - Tries each backend in order, stops on first success - Lazy initialization: `backends.py:77-79` - Converters created only when first used - Quality metrics: `utils.py:43-76` - Calculates char/word/table counts ## Additional Resources ### Reference Files For detailed backend documentation and advanced patterns: - **`references/backends.md`** - Detailed backend comparison and selection guide ### Example Usage Working examples in the insurance analysis that prompted this skill: - Extracted 21 PDFs from mortgage statements and insurance policies - Used markitdown backend for fast extraction - Parsed structured data (dates, amounts, policy numbers) </reference> <troubleshooting> ## Error Handling The extraction system handles errors gracefully: 1. **Backend failures**: Automatically tries next backend 2. **Import errors**: Skips unavailable backends 3. **File errors**: Reports specific error message 4. **Partial success**: Continues with remaining files in batch All errors are captured in metadata rather than raising exceptions. ## Dependencies The base package has no required Python dependencies. Select extras for the backends you need: - `pdf`: markitdown, pdfplumber, pdfminer.six, and maintained `pypdf` (the CLI backend id stays `pypdf2`) - `pdf-gpu`: docling on Linux/Windows; empty on macOS while docling's model stack selects an advisory-affected transformers 4.x release - `pdf-llm`: pymupdf4llm - `pdf-progress`: tqdm - `pdf-all`: every extra above Install CPU dependencies: ```bash uv pip install "markitdown>=0.1.0" "pdfplumber>=0.10.0" "pdfminer.six>=20221105" "pypdf>=6.0.0" tqdm ``` For the supported GPU extra on Linux or Windows: ```bash uv pip install "docling>=2.94.0" ``` The `marker` backend remains discoverable for separately managed installs, but marker-pdf is excluded from published extras because its supported-platform dependency graph pins Pillow below the first fully patched release. ## Troubleshooting ### `extract-pdfs: command not found` ```bash # Install as global UV tool from repo root: uv tool install --force --editable "./plugins/autorun[pdf]" extract-pdfs --list-backends # verify ``` ### `ModuleNotFoundError: No module named 'pdf_extraction'` (or 'markitdown', 'pdfplumber') ```bash # Re-install with all base dependencies: uv tool install --force --editable "./plugins/autorun[pdf]" # Or install explicitly: uv pip install "markitdown>=0.1.0" "pdfplumber>=0.10.0" "pdfminer.six>=20221105" "pypdf>=6.0.0" tqdm ``` ### GPU backend (docling) not available ```bash # Requires PyTorch; install the GPU extra: uv tool install --force --editable "./plugins/autorun[pdf,pdf-gpu]" extract-pdfs --list-backends # verify docling appears # Note: docling downloads models on first use. ``` ### Empty output from scanned PDF (image-only document) ```bash # Scanned PDFs require OCR; docling is in the supported GPU extra: extract-pdfs scanned.pdf --backends docling # If GPU unavailable, try pdftotext (system tool): brew install poppler # macOS # apt install poppler-utils # Ubuntu/Debian extract-pdfs scanned.pdf --backends pdftotext ``` ### pdfminer import error (package name confusion) ```bash # Install correct package (name has .six suffix): uv pip install "pdfminer.six>=20221105" # Import is still: from pdfminer.high_level import extract_text (no .six) ``` ### markitdown version conflict ```bash # API changed significantly in 0.1.0; ensure correct version: uv pip install "markitdown>=0.1.0" ``` </troubleshooting>
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.