Claude Skill

Apache Tika Content Extraction Hub

Extracts text and metadata from 1400+ file formats via Apache Tika Server REST API. Handles PDF, DOCX, PPTX, email archives, and embedded document extraction with MIME type detection.

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download agentskillexchange-skills-skills_apache-tika-content-extraction-hub-07beb56.zip · 0 KB
Part of agentskillexchange/skills — 249 skills

Install

skills CLI npx skills add https://github.com/agentskillexchange/skills/tree/main/skills/apache-tika-content-extraction-hub
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install agentskillexchange-skills@llmmart
Git git clone https://github.com/agentskillexchange/skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole agentskillexchange/skills collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Apache Tika Content Extraction Hub

Extracts text and metadata from 1400+ file formats via Apache Tika Server REST API. Handles PDF, DOCX, PPTX, email archives, and embedded document extraction with MIME type detection.

Installation

Requirements and caveats from upstream:

  • N.B. Docker is used for tests in tika-integration-tests. If Docker is not installed, those tests are skipped.

Basic usage or getting-started notes:

Source

Files (skills)
  • SKILL.md 1.3 KB
    ---
    name: "Apache Tika Content Extraction Hub"
    slug: "apache-tika-content-extraction-hub"
    description: "Extracts text and metadata from 1400+ file formats via Apache Tika Server REST API. Handles PDF, DOCX, PPTX, email archives, and embedded document extraction with MIME type detection."
    github_stars: 3703
    verification: "security_reviewed"
    source: "https://github.com/apache/tika"
    author: "The Apache Software Foundation"
    category: "Data Extraction & Transformation"
    framework: "Custom Agents"
    tool_ecosystem:
      github_repo: "apache/tika"
      github_stars: 3703
    ---
    
    # Apache Tika Content Extraction Hub
    
    Extracts text and metadata from 1400+ file formats via Apache Tika Server REST API. Handles PDF, DOCX, PPTX, email archives, and embedded document extraction with MIME type detection.
    
    ## Installation
    
    Requirements and caveats from upstream:
    - **N.B.** [Docker](https://www.docker.com/products/personal) is used for tests in tika-integration-tests. If Docker is not installed, those tests are skipped.
    
    Basic usage or getting-started notes:
    - ===========
    - **Parse a file in Java:**
    - java
    
    - Source: https://github.com/apache/tika
    - Extracted from upstream docs: https://raw.githubusercontent.com/apache/tika/HEAD/README.md
    
    ## Source
    
    - [Agent Skill Exchange](https://agentskillexchange.com/skills/apache-tika-content-extraction-hub/)
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related