Claude Skill

pdf-extractor

This skill should be used when the user asks to "extract text from PDF", "convert PDF to text", "parse PDF", "read PDF contents", "extract data from documents", "batch PDF extraction", "PDF to markdown", "OCR PDF", "get text from PDF files", "I have a PDF", "can you read this PDF

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download ahundt-autorun-plugins_pdf-extractor_skills_pdf-extractor-6fb6027.zip · 7 KB
Part of ahundt/autorun — 19 skills

Install

skills CLI npx skills add https://github.com/ahundt/autorun/tree/main/plugins/pdf-extractor/skills/pdf-extractor
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install ahundt-autorun@llmmart
Git git clone https://github.com/ahundt/autorun.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole ahundt/autorun collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

PDF Data Extraction

Files (autorun)
  • references
    • backends.md 7.2 KB
      # PDF Extraction Backend Reference
      
      Detailed comparison and selection guide for the 9 PDF extraction backends.
      
      ## Modern Backends (2024-2025)
      
      ### MarkItDown (Microsoft)
      
      **License:** MIT
      **Dependencies:** `markitdown>=0.1.0`
      **Models:** None (lightweight)
      
      **Strengths:**
      - Fast, lightweight extraction
      - Good for general documents
      - No model downloads required
      - MIT license allows commercial use
      
      **Weaknesses:**
      - Limited OCR capability
      - May miss complex layouts
      
      **Best for:** General text documents, forms, simple layouts
      
      **Usage:**
      ```python
      from markitdown import MarkItDown
      md = MarkItDown()
      result = md.convert('/path/to/document.pdf')
      text = result.text_content
      ```
      
      ### Docling (IBM)
      
      **License:** MIT
      **Dependencies:** `docling>=2.0.0`
      **Models:** ~500MB (downloaded on first use)
      
      **Strengths:**
      - Excellent layout analysis
      - Table detection and extraction
      - Figure handling
      - GPU acceleration available
      
      **Weaknesses:**
      - Large model download
      - Slower than lightweight backends
      - Requires more memory
      
      **Best for:** Complex layouts, academic papers, reports with figures/tables
      
      **Usage:**
      ```python
      from docling.document_converter import DocumentConverter
      converter = DocumentConverter()
      result = converter.convert('/path/to/document.pdf')
      markdown = result.document.export_to_markdown()
      ```
      
      ### Marker (datalab-to)
      
      **License:** GPL-3.0 (copyleft)
      **Dependencies:** separately managed `marker-pdf` installation
      **Models:** vision models downloaded by marker
      
      Marker is not selected by a published extra because its supported-platform
      dependency graph pins Pillow below the first fully patched release.
      
      **Strengths:**
      - Best for scanned documents
      - Vision-based approach
      - Handles complex layouts
      - GPU-accelerated
      
      **Weaknesses:**
      - GPL license restricts commercial use
      - Large model download
      - Slowest backend
      - Requires significant GPU memory
      
      **Best for:** Scanned documents, PDFs from images, complex layouts
      
      **Usage:**
      ```python
      from marker.converters.pdf import PdfConverter
      from marker.models import create_model_dict
      converter = PdfConverter(artifact_dict=create_model_dict())
      result = converter('/path/to/document.pdf')
      markdown = result.markdown
      ```
      
      ## Traditional Backends
      
      ### PDFPlumber
      
      **License:** MIT
      **Dependencies:** `pdfplumber>=0.10.0`
      **Models:** None
      
      **Strengths:**
      - Excellent table extraction
      - Precise character positioning
      - Good for structured documents
      - MIT license
      
      **Weaknesses:**
      - Can be slow on large documents
      - Limited formatting preservation
      
      **Best for:** Tables, forms, invoices, structured data
      
      **Usage:**
      ```python
      import pdfplumber
      with pdfplumber.open('/path/to/document.pdf') as pdf:
          text = '\n'.join([page.extract_text() for page in pdf.pages])
      ```
      
      ### PDFMiner
      
      **License:** MIT
      **Dependencies:** `pdfminer.six>=20221105`
      **Models:** None
      
      **Strengths:**
      - Pure Python, always available
      - Reliable text extraction
      - Good layout analysis
      - MIT license
      
      **Weaknesses:**
      - No table extraction
      - Older codebase
      
      **Best for:** Simple text documents, fallback extraction
      
      **Usage:**
      ```python
      from pdfminer.high_level import extract_text
      text = extract_text('/path/to/document.pdf')
      ```
      
      ### pypdf (`pypdf2` CLI backend id)
      
      **License:** BSD-3
      **Dependencies:** `pypdf>=6.0.0`
      **Models:** None
      
      **Strengths:**
      - Included in the `cpu` extra
      - Fast
      - BSD license (permissive)
      - Good for encryption detection
      
      **Weaknesses:**
      - Basic text extraction only
      - Poor formatting preservation
      - May miss text in complex layouts
      
      **Best for:** Last-resort fallback, encryption checking
      
      **Usage:**
      ```python
      import pypdf
      with open('/path/to/document.pdf', 'rb') as f:
          reader = pypdf.PdfReader(f)
          text = ''.join([page.extract_text() for page in reader.pages])
      ```
      
      ### PyMuPDF4LLM
      
      **License:** AGPL-3.0 (copyleft)
      **Dependencies:** `pymupdf4llm>=0.0.1`
      **Models:** None
      
      **Strengths:**
      - LLM-optimized markdown output
      - Fast extraction
      - Good structure preservation
      
      **Weaknesses:**
      - AGPL license limits commercial use
      - Less tested than other backends
      
      **Best for:** LLM pipelines, document Q&A systems
      
      **Usage:**
      ```python
      import pymupdf4llm
      markdown = pymupdf4llm.to_markdown('/path/to/document.pdf')
      ```
      
      ### PDFBox
      
      **License:** Apache-2.0
      **Dependencies:** `pdfbox>=0.1.0` (Java wrapper)
      **Models:** None
      
      **Strengths:**
      - Mature, well-tested
      - Good for tables
      - Apache license
      
      **Weaknesses:**
      - Requires Java runtime
      - Installation can be problematic
      - Slower due to JVM overhead
      
      **Best for:** Table extraction when pdfplumber fails
      
      **Usage:**
      ```python
      from pdfbox import PDFBox
      pdfbox = PDFBox()
      pdfbox.extract_text('/path/to/document.pdf', '/path/to/output.txt')
      ```
      
      ### Pdftotext (CLI)
      
      **License:** GPL (Poppler)
      **Dependencies:** System `pdftotext` command
      **Models:** None
      
      **Strengths:**
      - Very fast
      - Good text preservation
      - Layout mode available
      
      **Weaknesses:**
      - Requires system installation
      - Not available everywhere
      - GPL license
      
      **Best for:** Simple text extraction, system scripts
      
      **Usage:**
      ```bash
      pdftotext -layout document.pdf output.txt
      ```
      
      ## Selection Decision Tree
      
      ```
      Start
      ├── Is PDF scanned/image-based?
      │   ├── Yes → Use marker (GPU) or docling
      │   └── No → Continue
      ├── Does PDF have tables?
      │   ├── Yes → Use pdfplumber or pdfbox
      │   └── No → Continue
      ├── Is speed critical?
      │   ├── Yes → Use markitdown or pdfminer
      │   └── No → Continue
      ├── Is commercial use required?
      │   ├── Yes → Avoid marker (GPL), pymupdf4llm (AGPL)
      │   └── No → Any backend works
      └── Default → markitdown → pdfplumber → pdfminer → pypdf2
      ```
      
      ## License Summary
      
      | Backend | License | Commercial Use |
      |---------|---------|----------------|
      | markitdown | MIT | Yes |
      | docling | MIT | Yes |
      | marker | GPL-3.0 | No (copyleft) |
      | pymupdf4llm | AGPL-3.0 | No (copyleft) |
      | pdfplumber | MIT | Yes |
      | pdfminer | MIT | Yes |
      | pypdf2 | BSD-3 | Yes |
      | pdfbox | Apache-2.0 | Yes |
      | pdftotext | GPL | No (copyleft) |
      
      ## Performance Benchmarks
      
      Approximate extraction times for a 10-page document:
      
      | Backend | Time | Memory |
      |---------|------|--------|
      | pypdf2 | 0.5s | Low |
      | pdfminer | 1s | Low |
      | pdftotext | 0.3s | Low |
      | markitdown | 1s | Low |
      | pdfplumber | 2s | Medium |
      | pdfbox | 3s | Medium |
      | docling | 10s | High |
      | marker | 30s | Very High |
      | pymupdf4llm | 1s | Medium |
      
      Note: GPU-accelerated backends (docling, marker) are much faster with CUDA.
      
      ## Troubleshooting
      
      ### Backend Not Available
      
      ```python
      # Check if backend is importable
      try:
          from markitdown import MarkItDown
          print("markitdown available")
      except ImportError:
          print("markitdown not installed")
      ```
      
      ### Installation Issues
      
      **pdfbox:** Requires Java 8+
      ```bash
      java -version  # Check Java installed
      pip install pdfbox
      ```
      
      **marker:** Separately managed only. It is intentionally excluded from the
      published extras while its dependency graph requires an unpatched Pillow.
      
      **docling:** Large download
      ```bash
      pip install docling
      # First run downloads ~500MB models
      ```
      
      ### Common Errors
      
      **"No module named 'X'"**: Backend not installed
      ```bash
      pip install X
      ```
      
      **"Java not found"**: PDFBox needs Java
      ```bash
      brew install openjdk  # macOS
      apt install default-jre  # Ubuntu
      ```
      
      **"CUDA out of memory"**: Reduce batch size or use CPU
      ```python
      backends = ['markitdown', 'pdfplumber']  # CPU-only
      ```
      
  • SKILL.md 12.6 KB
    ---
    name: pdf-extractor
    description: This skill should be used when the user asks to "extract text from PDF", "convert PDF to text", "parse PDF", "read PDF contents", "extract data from documents", "batch PDF extraction", "PDF to markdown", "OCR PDF", "get text from PDF files", "I have a PDF", "can you read this PDF", "what's in this PDF", "summarize this PDF", "open PDF file", "extract from [filename].pdf", or needs to process PDF documents for data extraction. Handles single-file extraction, batch processing, and OCR for scanned documents with automatic backend selection.
    metadata:
      version: 1.0.0rc3
      example-prompt: "Extract text from document.pdf"
    ---
    
    # PDF Data Extraction
    
    <purpose>
    
    Extract text and structured data from PDF documents using a multi-backend approach with automatic fallback.
    
    ## Overview
    
    This skill provides PDF text extraction with 9 different backends, automatic GPU detection, and intelligent backend selection. The extraction system tries backends in order until one succeeds, producing markdown output optimized for further processing.
    
    </purpose>
    
    <workflow>
    
    ## Quick Start Workflow
    
    To extract text from PDFs:
    
    1. **Single file extraction (installed CLI - recommended):**
       ```bash
       extract-pdfs /path/to/document.pdf
       ```
       Output: Creates `document.md` in the same directory.
    
    2. **Batch extraction (directory):**
       ```bash
       extract-pdfs /path/to/pdfs/ /path/to/output/
       ```
       Output: Creates `.md` files for all PDFs in output directory.
    
    3. **Custom output file:**
       ```bash
       extract-pdfs document.pdf output.md
       ```
    
    4. **Specific backends:**
       ```bash
       extract-pdfs document.pdf --backends markitdown pdfplumber
       ```
    
    5. **List available backends:**
       ```bash
       extract-pdfs --list-backends
       ```
       Output: Shows available backends and GPU status.
    
    ### Alternative Execution Methods
    
    If the `extract-pdfs` CLI isn't installed, install it first (recommended):
    
    ```bash
    # Install as global UV tool (from repo root); the code ships in the autorun-ai distribution:
    cd "${CLAUDE_PLUGIN_ROOT}/../.." && uv tool install --force --editable "./plugins/autorun[pdf]"
    extract-pdfs --list-backends  # verify
    ```
    
    Or use these fallback methods without installing:
    
    ```bash
    # From a source checkout, without installing (the package lives in plugins/autorun):
    uv run --project plugins/autorun --extra pdf python -m pdf_extraction document.pdf
    
    # Module execution, if the console script is not on PATH
    python -m pdf_extraction document.pdf
    ```
    
    ## Backend Selection Guide
    
    ### Custom Backend Ordering
    
    Specify backends in any order with `--backends`. The system tries each in order, stopping on first success:
    
    ```bash
    # Tables first, then general extraction
    extract-pdfs document.pdf --backends pdfplumber markitdown pdfminer
    
    # Scanned documents: vision-based first
    extract-pdfs scanned.pdf --backends docling markitdown
    
    # Most permissive fallback order (handles problematic PDFs)
    extract-pdfs document.pdf --backends pdfminer pypdf2 markitdown
    
    # Single backend only (no fallback)
    extract-pdfs document.pdf --backends markitdown
    ```
    
    ### CPU-Only Systems (Default)
    
    For systems without GPU, the recommended backend order:
    - `markitdown` - Microsoft's lightweight converter (MIT, fast, no models)
    - `pdfplumber` - Excellent for tables (MIT)
    - `pdfminer` - Pure Python, reliable (MIT)
    - `pypdf2` - Basic extraction through maintained `pypdf` (BSD-3; `pdf` extra)
    
    ### GPU Systems
    
    For systems with CUDA-enabled GPU:
    - `docling` - IBM layout analysis (MIT, downloads models on first use)
    - Plus all CPU backends as fallback
    
    The `marker` backend is recognized only for separately managed installations;
    it is not selected by a published extra because its dependency graph pins an
    unpatched Pillow release.
    
    ### Backend Comparison
    
    | Backend | License | Models | Best For | Speed |
    |---------|---------|--------|----------|-------|
    | markitdown | MIT | None | General text, forms | Fast |
    | pdfplumber | MIT | None | Tables, structured data | Fast |
    | pdfminer | MIT | None | Simple text documents | Fast |
    | pypdf2 | BSD-3 | None | Basic extraction | Fast |
    | docling | MIT | ~500MB | Layout analysis | Medium |
    | marker | GPL-3.0 | ~1GB | Scanned documents | Slow |
    | pymupdf4llm | AGPL-3.0 | None | LLM-optimized output | Fast |
    | pdfbox | Apache-2.0 | None | Tables (Java-based) | Medium |
    | pdftotext | System | None | Simple text (CLI) | Fast |
    
    ### Backend Decision Matrix
    
    | Document Type | Recommended Backend(s) | Why |
    |---------------|------------------------|-----|
    | Digital text PDF (default) | markitdown, pdfplumber | Fast, accurate |
    | PDF with tables/invoices | pdfplumber, pdfbox | Best table structure |
    | Complex layouts/columns | docling (GPU) | Layout analysis |
    | Scanned documents/images | marker, docling (GPU) | OCR/vision required |
    | Insurance policies/forms | markitdown, pdfplumber | Handles form fields |
    | Academic papers | docling | Equations, figures |
    | Maximum compatibility | pdfminer, pypdf2 | Fewest dependencies |
    | Commercial use required | markitdown, pdfplumber | MIT license |
    
    ## Programmatic Usage
    
    To use the extraction library directly in Python code:
    
    ```python
    from pdf_extraction import extract_single_pdf, pdf_to_txt, detect_gpu_availability
    
    # Check available backends
    gpu_info = detect_gpu_availability()
    print(f"Recommended backends: {gpu_info['recommended_backends']}")
    
    # Extract single file
    result = extract_single_pdf(
        input_file='/path/to/document.pdf',
        output_file='/path/to/output.md',
        backends=['markitdown', 'pdfplumber']
    )
    
    if result['success']:
        print(f"Extracted with {result['backend_used']}")
        print(f"Quality metrics: {result['quality_metrics']}")
    
    # Batch extract directory
    output_files, metadata = pdf_to_txt(
        input_dir='/path/to/pdfs/',
        output_dir='/path/to/output/',
        resume=True,  # Skip already-extracted files
        return_metadata=True
    )
    ```
    
    </workflow>
    
    <reference>
    
    ## Extraction Metadata
    
    Every extraction returns metadata for quality assessment:
    
    ```python
    {
        'success': True,
        'backend_used': 'markitdown',
        'extraction_time_seconds': 2.5,
        'output_size_bytes': 15234,
        'quality_metrics': {
            'char_count': 15234,
            'line_count': 450,
            'word_count': 2800,
            'table_markers': 12,      # Count of | (tables)
            'has_structure': True     # Has markdown structure
        },
        'encrypted': False,
        'error': None
    }
    ```
    
    ## Handling Common Scenarios
    
    ### Encrypted PDFs
    
    The system detects encrypted PDFs and reports them:
    ```python
    if result['encrypted']:
        print("PDF is password-protected")
    ```
    
    Encrypted PDFs cannot be extracted without the password.
    
    ### Empty or Failed Extractions
    
    When all backends fail:
    1. Check if PDF is encrypted
    2. Try with `--backends pdfminer pypdf2` (most permissive)
    3. Check PDF isn't corrupted
    4. Consider OCR-based backends for scanned documents
    
    ### Resume Batch Processing
    
    To continue interrupted batch extraction:
    ```bash
    extract-pdfs /path/to/pdfs/ /path/to/output/
    ```
    The `resume=True` default skips already-extracted files.
    
    To force re-extraction:
    ```bash
    extract-pdfs /path/to/pdfs/ --no-resume
    ```
    
    ### Tables and Structured Data
    
    For PDFs with tables, prioritize:
    ```bash
    extract-pdfs document.pdf --backends pdfplumber markitdown
    ```
    
    The output will contain markdown tables when detected:
    ```markdown
    | Column1 | Column2 | Column3 |
    |---------|---------|---------|
    | Data    | Data    | Data    |
    ```
    
    ## Module Structure Reference
    
    ### Source Code Layout
    
    **Location:** `plugins/pdf-extractor/src/pdf_extraction/` in a source checkout, and the
    importable `pdf_extraction` package once `autorun` is installed. This plugin
    directory holds the manifest, command, and this skill; the code ships inside the
    `autorun-ai` distribution behind its `pdf` extra.
    
    | File | Purpose |
    |------|---------|
    | `__init__.py` | Package exports (extract_single_pdf, pdf_to_txt, etc.) |
    | `__main__.py` | Support for `python -m pdf_extraction` |
    | `cli.py` | CLI entry point with argparse |
    | `backends.py` | BackendExtractor base class + 9 backend implementations |
    | `extractors.py` | extract_single_pdf(), pdf_to_txt() functions |
    | `utils.py` | GPU detection, quality metrics, encryption check |
    
    ### Key Classes and Functions
    
    | Component | Location | Purpose |
    |-----------|----------|---------|
    | `BackendExtractor` | backends.py:35-123 | Base class with Template Method pattern |
    | `DoclingExtractor` | backends.py:130-142 | IBM Docling backend (MIT, GPU) |
    | `MarkerExtractor` | backends.py:145-158 | Vision-based marker backend (GPL-3.0, GPU) |
    | `MarkItDownExtractor` | backends.py:161-173 | Microsoft MarkItDown (MIT, CPU) |
    | `PdfplumberExtractor` | backends.py:244-253 | Table-focused extraction (MIT) |
    | `PdfminerExtractor` | backends.py:219-226 | Pure Python fallback (MIT) |
    | `Pypdf2Extractor` | backends.py:229-241 | Basic extraction through optional `pypdf` (BSD-3) |
    | `BACKEND_REGISTRY` | backends.py:279-292 | Dict mapping backend names to factories |
    | `detect_gpu_availability()` | utils.py:9-40 | Auto-detect GPU and recommend backends |
    | `extract_single_pdf()` | extractors.py:13-80 | Extract one PDF with backend fallback |
    | `pdf_to_txt()` | extractors.py:83-170 | Batch extract directory with resume |
    
    **Key implementation details:**
    - Backend fallback loop: `extractors.py:55-78` - Tries each backend in order, stops on first success
    - Lazy initialization: `backends.py:77-79` - Converters created only when first used
    - Quality metrics: `utils.py:43-76` - Calculates char/word/table counts
    
    ## Additional Resources
    
    ### Reference Files
    
    For detailed backend documentation and advanced patterns:
    - **`references/backends.md`** - Detailed backend comparison and selection guide
    
    ### Example Usage
    
    Working examples in the insurance analysis that prompted this skill:
    - Extracted 21 PDFs from mortgage statements and insurance policies
    - Used markitdown backend for fast extraction
    - Parsed structured data (dates, amounts, policy numbers)
    
    </reference>
    
    <troubleshooting>
    
    ## Error Handling
    
    The extraction system handles errors gracefully:
    
    1. **Backend failures**: Automatically tries next backend
    2. **Import errors**: Skips unavailable backends
    3. **File errors**: Reports specific error message
    4. **Partial success**: Continues with remaining files in batch
    
    All errors are captured in metadata rather than raising exceptions.
    
    ## Dependencies
    
    The base package has no required Python dependencies. Select extras for the
    backends you need:
    
    - `pdf`: markitdown, pdfplumber, pdfminer.six, and maintained `pypdf` (the CLI
      backend id stays `pypdf2`)
    - `pdf-gpu`: docling on Linux/Windows; empty on macOS while docling's model stack
      selects an advisory-affected transformers 4.x release
    - `pdf-llm`: pymupdf4llm
    - `pdf-progress`: tqdm
    - `pdf-all`: every extra above
    
    Install CPU dependencies:
    ```bash
    uv pip install "markitdown>=0.1.0" "pdfplumber>=0.10.0" "pdfminer.six>=20221105" "pypdf>=6.0.0" tqdm
    ```
    
    For the supported GPU extra on Linux or Windows:
    ```bash
    uv pip install "docling>=2.94.0"
    ```
    
    The `marker` backend remains discoverable for separately managed installs, but
    marker-pdf is excluded from published extras because its supported-platform
    dependency graph pins Pillow below the first fully patched release.
    
    ## Troubleshooting
    
    ### `extract-pdfs: command not found`
    ```bash
    # Install as global UV tool from repo root:
    uv tool install --force --editable "./plugins/autorun[pdf]"
    extract-pdfs --list-backends  # verify
    ```
    
    ### `ModuleNotFoundError: No module named 'pdf_extraction'` (or 'markitdown', 'pdfplumber')
    ```bash
    # Re-install with all base dependencies:
    uv tool install --force --editable "./plugins/autorun[pdf]"
    # Or install explicitly:
    uv pip install "markitdown>=0.1.0" "pdfplumber>=0.10.0" "pdfminer.six>=20221105" "pypdf>=6.0.0" tqdm
    ```
    
    ### GPU backend (docling) not available
    ```bash
    # Requires PyTorch; install the GPU extra:
    uv tool install --force --editable "./plugins/autorun[pdf,pdf-gpu]"
    extract-pdfs --list-backends  # verify docling appears
    # Note: docling downloads models on first use.
    ```
    
    ### Empty output from scanned PDF (image-only document)
    ```bash
    # Scanned PDFs require OCR; docling is in the supported GPU extra:
    extract-pdfs scanned.pdf --backends docling
    # If GPU unavailable, try pdftotext (system tool):
    brew install poppler        # macOS
    # apt install poppler-utils  # Ubuntu/Debian
    extract-pdfs scanned.pdf --backends pdftotext
    ```
    
    ### pdfminer import error (package name confusion)
    ```bash
    # Install correct package (name has .six suffix):
    uv pip install "pdfminer.six>=20221105"
    # Import is still: from pdfminer.high_level import extract_text  (no .six)
    ```
    
    ### markitdown version conflict
    ```bash
    # API changed significantly in 0.1.0; ensure correct version:
    uv pip install "markitdown>=0.1.0"
    ```
    
    </troubleshooting>
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related