Claude Skill

markdown-converter

Convert PDF, Office, HTML, data, media, ZIP to Markdown.

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download notque-vexjoy-agent-skills_research_markdown-converter-8ad6845.zip · 1 KB
Part of notque/vexjoy-agent — 69 skills

Install

skills CLI npx skills add https://github.com/notque/vexjoy-agent/tree/main/skills/research/markdown-converter
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install notque-vexjoy-agent@llmmart
Git git clone https://github.com/notque/vexjoy-agent.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole notque/vexjoy-agent collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Markdown Converter

Convert a file to Markdown with markitdown, zero install:

uvx 'markitdown[all]' input.pdf -o output.md   # to file
uvx 'markitdown[all]' input.docx               # to stdout
cat blob | uvx 'markitdown[all]' -x .pdf       # stdin, with extension hint

When uvx is missing, run pipx run 'markitdown[all]' … with the same arguments. First run downloads dependencies; later runs hit the cache. Output preserves headings, tables, lists, and links.

For video transcripts, use the video-transcript skill.

Formats

Input Notes
PDF, .docx, .pptx, .xlsx, .xls Document structure preserved
HTML, CSV, JSON, XML Structured Markdown
Images EXIF metadata + OCR text
Audio EXIF metadata + speech transcription
ZIP, EPub Iterates contents, converts each

Options

Flag Effect
-o FILE Write output to FILE
-x .EXT Extension hint for stdin input
-m MIME MIME-type hint
-c CHARSET Charset hint, e.g. UTF-8

Error handling

Garbled or empty text from a scanned PDF

Cause: page is an image; the base extractor reads text layers only. Solution: render pages to images (pdftoppm), then convert the images so OCR runs.

Files (vexjoy-agent)
  • SKILL.md 1.9 KB
    ---
    name: markdown-converter
    description: "Convert PDF, Office, HTML, data, media, ZIP to Markdown."
    user_invocable: false  # default -- router-dispatched, not user-typed
    agent: python-general-engineer
    allowed-tools:
      - Bash
      - Read
    routing:
      triggers:
        - "convert to markdown"
        - "markitdown"
        - "extract text from PDF"
        - "PDF to markdown"
        - "docx to markdown"
        - "ingest document"
        - "read this PDF"
        - "read this document"
        - "extract text from document"
        - "convert PDF"
        - "convert document"
        - "pptx to markdown"
        - "xlsx to markdown"
      category: research
      pairs_with:
        - research
        - domain
    ---
    
    # Markdown Converter
    
    Convert a file to Markdown with markitdown, zero install:
    
    ```bash
    uvx 'markitdown[all]' input.pdf -o output.md   # to file
    uvx 'markitdown[all]' input.docx               # to stdout
    cat blob | uvx 'markitdown[all]' -x .pdf       # stdin, with extension hint
    ```
    
    When `uvx` is missing, run `pipx run 'markitdown[all]' …` with the same arguments. First run downloads dependencies; later runs hit the cache. Output preserves headings, tables, lists, and links.
    
    For video transcripts, use the `video-transcript` skill.
    
    ## Formats
    
    | Input | Notes |
    |---|---|
    | PDF, .docx, .pptx, .xlsx, .xls | Document structure preserved |
    | HTML, CSV, JSON, XML | Structured Markdown |
    | Images | EXIF metadata + OCR text |
    | Audio | EXIF metadata + speech transcription |
    | ZIP, EPub | Iterates contents, converts each |
    
    ## Options
    
    | Flag | Effect |
    |---|---|
    | `-o FILE` | Write output to FILE |
    | `-x .EXT` | Extension hint for stdin input |
    | `-m MIME` | MIME-type hint |
    | `-c CHARSET` | Charset hint, e.g. UTF-8 |
    
    ## Error handling
    
    ### Garbled or empty text from a scanned PDF
    Cause: page is an image; the base extractor reads text layers only.
    Solution: render pages to images (`pdftoppm`), then convert the images so OCR runs.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related