Apache Tika Document Parser
Extracts structured text, metadata, and embedded objects from PDFs, Office documents, and 1000+ file formats using the Apache Tika REST API. Outputs clean Markdown or JSON with XMP metadata preservation.
Install
npx skills add https://github.com/agentskillexchange/skills/tree/main/skills/apache-tika-document-parser
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install agentskillexchange-skills@llmmart
git clone https://github.com/agentskillexchange/skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole agentskillexchange/skills collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Apache Tika Document Parser
Extracts structured text, metadata, and embedded objects from PDFs, Office documents, and 1000+ file formats using the Apache Tika REST API. Outputs clean Markdown or JSON with XMP metadata preservation.
Installation
Requirements and caveats from upstream:
- N.B. Docker is used for tests in tika-integration-tests. If Docker is not installed, those tests are skipped.
Basic usage or getting-started notes:
===========
Parse a file in Java:
java
Source: https://github.com/apache/tika
Extracted from upstream docs: https://raw.githubusercontent.com/apache/tika/HEAD/README.md
Source
Files (skills)
-
SKILL.md 1.3 KB
--- name: "Apache Tika Document Parser" slug: "apache-tika-document-parser" description: "Extracts structured text, metadata, and embedded objects from PDFs, Office documents, and 1000+ file formats using the Apache Tika REST API. Outputs clean Markdown or JSON with XMP metadata preservation." github_stars: 3703 verification: "security_reviewed" source: "https://github.com/apache/tika" author: "The Apache Software Foundation" category: "Data Extraction & Transformation" framework: "Gemini" tool_ecosystem: github_repo: "apache/tika" github_stars: 3703 --- # Apache Tika Document Parser Extracts structured text, metadata, and embedded objects from PDFs, Office documents, and 1000+ file formats using the Apache Tika REST API. Outputs clean Markdown or JSON with XMP metadata preservation. ## Installation Requirements and caveats from upstream: - **N.B.** [Docker](https://www.docker.com/products/personal) is used for tests in tika-integration-tests. If Docker is not installed, those tests are skipped. Basic usage or getting-started notes: - =========== - **Parse a file in Java:** - java - Source: https://github.com/apache/tika - Extracted from upstream docs: https://raw.githubusercontent.com/apache/tika/HEAD/README.md ## Source - [Agent Skill Exchange](https://agentskillexchange.com/skills/apache-tika-document-parser/)
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.