Apache Tika Document Extractor
Wraps Apache Tika Server REST API for extracting structured text from PDFs, DOCX, PPTX, and 1,200+ file formats. Outputs clean markdown with metadata preservation using Tika /rmeta/text endpoint and recursive parsing mode.
Install
npx skills add https://github.com/agentskillexchange/skills/tree/main/skills/apache-tika-document-extractor
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install agentskillexchange-skills@llmmart
git clone https://github.com/agentskillexchange/skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole agentskillexchange/skills collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Apache Tika Document Extractor
Wraps Apache Tika Server REST API for extracting structured text from PDFs, DOCX, PPTX, and 1,200+ file formats. Outputs clean markdown with metadata preservation using Tika /rmeta/text endpoint and recursive parsing mode.
Installation
Requirements and caveats from upstream:
- N.B. Docker is used for tests in tika-integration-tests. If Docker is not installed, those tests are skipped.
Basic usage or getting-started notes:
===========
Parse a file in Java:
java
Source: https://github.com/apache/tika
Extracted from upstream docs: https://raw.githubusercontent.com/apache/tika/HEAD/README.md
Source
Files (skills)
-
SKILL.md 1.3 KB
--- name: "Apache Tika Document Extractor" slug: "apache-tika-document-extractor" description: "Wraps Apache Tika Server REST API for extracting structured text from PDFs, DOCX, PPTX, and 1,200+ file formats. Outputs clean markdown with metadata preservation using Tika /rmeta/text endpoint and recursive parsing mode." github_stars: 3695 verification: "security_reviewed" source: "https://github.com/apache/tika" category: "Data Extraction & Transformation" framework: "Codex" tool_ecosystem: github_repo: "apache/tika" github_stars: 3695 --- # Apache Tika Document Extractor Wraps Apache Tika Server REST API for extracting structured text from PDFs, DOCX, PPTX, and 1,200+ file formats. Outputs clean markdown with metadata preservation using Tika /rmeta/text endpoint and recursive parsing mode. ## Installation Requirements and caveats from upstream: - **N.B.** [Docker](https://www.docker.com/products/personal) is used for tests in tika-integration-tests. If Docker is not installed, those tests are skipped. Basic usage or getting-started notes: - =========== - **Parse a file in Java:** - java - Source: https://github.com/apache/tika - Extracted from upstream docs: https://raw.githubusercontent.com/apache/tika/HEAD/README.md ## Source - [Agent Skill Exchange](https://agentskillexchange.com/skills/apache-tika-document-extractor/)
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.