Beautiful Soup Academic Paper Parser
Extracts structured citation data from academic repositories using BeautifulSoup4 with lxml parser. Parses DOI metadata, author affiliations, and reference lists from PubMed, arXiv, and Semantic Scholar HTML.
Install
npx skills add https://github.com/agentskillexchange/skills/tree/main/skills/beautifulsoup-academic-paper-parser
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install agentskillexchange-skills@llmmart
git clone https://github.com/agentskillexchange/skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole agentskillexchange/skills collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Beautiful Soup Academic Paper Parser
Extracts structured citation data from academic repositories using BeautifulSoup4 with lxml parser. Parses DOI metadata, author affiliations, and reference lists from PubMed, arXiv, and Semantic Scholar HTML.
Installation
Use the upstream install or setup path that matches your environment:
- pip install beautifulsoup4
- format. Run make html in that directory to create HTML
Requirements and caveats from upstream:
- Requires: Python >=3.7.0
- Python
- Python :: 3
Basic usage or getting-started notes:
-
from bs4 import BeautifulSoup
-
soup = BeautifulSoup("
SomebadHTML")
-
print(soup.prettify())
Source
Files (skills)
-
SKILL.md 1.3 KB
--- name: "Beautiful Soup Academic Paper Parser" slug: "beautifulsoup-academic-paper-parser" description: "Extracts structured citation data from academic repositories using BeautifulSoup4 with lxml parser. Parses DOI metadata, author affiliations, and reference lists from PubMed, arXiv, and Semantic Scholar HTML." verification: "security_reviewed" source: "https://pypi.org/project/beautifulsoup4/" category: "Research & Scraping" framework: "MCP" --- # Beautiful Soup Academic Paper Parser Extracts structured citation data from academic repositories using BeautifulSoup4 with lxml parser. Parses DOI metadata, author affiliations, and reference lists from PubMed, arXiv, and Semantic Scholar HTML. ## Installation Use the upstream install or setup path that matches your environment: - pip install beautifulsoup4 - format. Run make html in that directory to create HTML Requirements and caveats from upstream: - Requires: Python >=3.7.0 - Python - Python :: 3 Basic usage or getting-started notes: - >> from bs4 import BeautifulSoup - >> soup = BeautifulSoup("<p>Some<b>bad<i>HTML") - >> print(soup.prettify()) - Source: https://pypi.org/project/beautifulsoup4/ ## Source - [Agent Skill Exchange](https://agentskillexchange.com/skills/beautifulsoup-academic-paper-parser/)
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.