extracting-keywords
Extract keywords from documents using YAKE algorithm with support for 34 languages (Arabic to Chinese). Use when users request keyword extraction, key terms, topic identification, content summarization, or document analysis. Includes domain-specific stopwords for AI/ML and life s
Install
npx skills add https://github.com/oaustegard/claude-skills/tree/main/plugins/data-and-visualization/skills/extracting-keywords
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install oaustegard-claude-skills@llmmart
git clone https://github.com/oaustegard/claude-skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole oaustegard/claude-skills collection as a plugin from our marketplace. Git is the plain clone.
README
extracting-keywords
Extract keywords from documents using YAKE algorithm with support for 34 languages (Arabic to Chinese). Use when users request keyword extraction, key terms, topic identification, content summarization, or document analysis. Includes domain-specific stopwords for AI/ML and life sciences. Optional deeper extraction mode (n=2+n=3 combined) for comprehensive coverage.
Skill manifest
Extracting Keywords
Extract keywords from text using YAKE (Yet Another Keyword Extractor), an unsupervised statistical keyword extraction algorithm.
Installation
First time only: Install YAKE with optimized dependencies to avoid unnecessary downloads.
cd /home/claude
uv venv yake-venv --system-site-packages
uv pip install yake --python yake-venv/bin/python --no-deps
uv pip install jellyfish segtok regex --python yake-venv/bin/python
This reuses system packages (numpy, networkx) instead of downloading them (~0.08s vs ~5s).
Stopwords Configuration
Built-in YAKE stopwords (34 languages): Use lan="<code>" parameter
- See Parameters section below for all 34 supported language codes
- English (
lan="en") is the default
Custom domain stopwords (bundled in assets/):
AI/ML: stopwords_ai.txt
- English stopwords + 783 AI/ML domain-specific terms (1357 total)
- Filters AI/ML methodology noise (model, training, network, algorithm, parameter)
- Filters ML boilerplate (dataset, baseline, benchmark, experiment, evaluation)
- Filters technical terms (transformer, embedding, attention, optimization, inference)
- Includes full lemmatization (train/trains/trained/training/trainer)
- Use for AI/ML papers, technical reports, machine learning literature
- Performance impact: +4-5% runtime vs English stopwords
Life Sciences: stopwords_ls.txt
- English stopwords + 719 life sciences domain-specific terms (1293 total)
- Filters research methodology noise (study, results, analysis, significant, observed)
- Filters academic boilerplate (paper, manuscript, publication, review, editing)
- Filters statistical terms (correlation, distribution, deviation, variance)
- Filters clinical terms (patient, treatment, diagnosis, symptom, therapy)
- Filters biology/medicine (cell, tissue, protein, gene, organism)
- Includes full lemmatization (analyze/analyzes/analyzed/analyzing/analysis)
- Use for biomedical papers, clinical studies, research articles, scientific literature
- Performance impact: +4-5% runtime vs English stopwords
Basic Usage
import yake
# Read text
with open('document.txt', 'r') as f:
text = f.read()
# Extract with English stopwords (default)
kw_extractor = yake.KeywordExtractor(
lan="en", # Language code
n=3, # Max n-gram size (1-3 word phrases)
dedupLim=0.9, # Deduplication threshold (0-1)
top=20 # Number of keywords to return
)
keywords = kw_extractor.extract_keywords(text)
# Display results (lower score = more important)
for kw, score in keywords:
print(f"{score:.4f} {kw}")
Domain-Specific Extraction
Using Life Sciences Stopwords
Option 1: Install custom stopwords file
# Copy life sciences stopwords to YAKE package
cp assets/stopwords_ls.txt /home/claude/yake-venv/lib/python3.12/site-packages/yake/core/StopwordsList/stopwords_ls.txt
# Use with lan="ls"
kw_extractor = yake.KeywordExtractor(lan="ls", n=3, top=20)
Option 2: Load custom stopwords directly
# Load stopwords from file
with open('assets/stopwords_ls.txt', 'r') as f:
custom_stops = set(line.strip().lower() for line in f)
# Pass to extractor
kw_extractor = yake.KeywordExtractor(
stopwords=custom_stops,
n=3,
top=20
)
Using AI/ML Stopwords
# Load AI/ML stopwords
with open('/mnt/skills/user/extracting-keywords/assets/stopwords_ai.txt', 'r') as f:
ai_stops = set(line.strip().lower() for line in f)
# Extract with AI stopwords
kw_extractor = yake.KeywordExtractor(
stopwords=ai_stops,
n=3,
top=20
)
keywords = kw_extractor.extract_keywords(text)
Deeper Extraction (n=2 + n=3 Combined)
For more comprehensive extraction, run both n=2 and n=3 and consolidate results. This captures both focused phrases and broader context with ~100% time overhead (still <2s for large documents).
import yake
# Load domain stopwords
with open('/mnt/skills/user/extracting-keywords/assets/stopwords_ai.txt', 'r') as f:
stops = set(line.strip().lower() for line in f)
# Extract with n=2 (captures focused phrases)
kw_n2 = yake.KeywordExtractor(stopwords=stops, n=2, dedupLim=0.9, top=50)
results_n2 = kw_n2.extract_keywords(text)
# Extract with n=3 (captures broader context)
kw_n3 = yake.KeywordExtractor(stopwords=stops, n=3, dedupLim=0.9, top=50)
results_n3 = kw_n3.extract_keywords(text)
# Consolidate: union with score averaging for overlaps
combined = {}
for kw, score in results_n2:
combined[kw] = score
for kw, score in results_n3:
if kw in combined:
combined[kw] = (combined[kw] + score) / 2
else:
combined[kw] = score
# Sort by score (lower = more important)
consolidated = sorted(combined.items(), key=lambda x: x[1])
# Display top 30
for kw, score in consolidated[:30]:
print(f"{score:.4f} {kw}")
Benefits:
- n=2 extracts cleaner domain-specific phrases ("disk move", "error rate")
- n=3 captures contextual combinations ("Move disk 1", "per-step error rate")
- Consolidation provides richer keyword set for topic modeling or search indexing
Performance:
- Combined approach: ~2x runtime of single extraction
- Typical timing: 0.4s (small doc) to 1.0s (large doc)
- Use when quality matters more than speed
Parameters
lan (str): Language code for built-in stopwords
"en"- English (default)"ai"- AI/ML (if stopwords_ai.txt installed in YAKE)"ls"- Life sciences (if stopwords_ls.txt installed in YAKE)
Built-in YAKE languages (34 total):
"ar"- Arabic"bg"- Bulgarian"br"- Breton"cz"- Czech"da"- Danish"de"- German"el"- Greek"es"- Spanish"et"- Estonian"fa"- Farsi/Persian"fi"- Finnish"fr"- French"hi"- Hindi"hr"- Croatian"hu"- Hungarian"hy"- Armenian"id"- Indonesian"it"- Italian"ja"- Japanese"lt"- Lithuanian"lv"- Latvian"nl"- Dutch"no"- Norwegian"pl"- Polish"pt"- Portuguese"ro"- Romanian"ru"- Russian"sk"- Slovak"sl"- Slovenian"sv"- Swedish"tr"- Turkish"uk"- Ukrainian"zh"- Chinese
n (int): Maximum n-gram size (default: 3)
1- Single words only2- Up to 2-word phrases3- Up to 3-word phrases (recommended)4-5- May produce suboptimal results with YAKE's algorithm
dedupLim (float): Deduplication threshold (default: 0.9)
- Range: 0.0 to 1.0
- Higher values = more aggressive deduplication
- Controls handling of similar terms (e.g., "cancer cell" vs "cancer cells")
top (int): Number of keywords to return (default: 20)
stopwords (set): Custom stopwords set (overrides lan parameter)
Workflow Patterns
Single Document Analysis
import yake
# Read document
with open('/mnt/user-data/uploads/article.txt', 'r') as f:
text = f.read()
# Extract keywords
kw_extractor = yake.KeywordExtractor(lan="en", n=3, top=30)
keywords = kw_extractor.extract_keywords(text)
# Format results
results = []
for kw, score in keywords:
results.append(f"{score:.4f} {kw}")
print("\n".join(results))
Comparing Stopwords Strategies
import yake
# Load life sciences stopwords
with open('assets/stopwords_ls.txt', 'r') as f:
ls_stops = set(line.strip().lower() for line in f)
# Extract with English stopwords
kw_en = yake.KeywordExtractor(lan="en", n=3, top=20)
keywords_en = kw_en.extract_keywords(text)
# Extract with life sciences stopwords
kw_ls = yake.KeywordExtractor(stopwords=ls_stops, n=3, top=20)
keywords_ls = kw_ls.extract_keywords(text)
# Compare results
print("English stopwords:")
for kw, score in keywords_en:
print(f" {score:.4f} {kw}")
print("\nLife sciences stopwords:")
for kw, score in keywords_ls:
print(f" {score:.4f} {kw}")
Batch Processing
import yake
import os
# Initialize extractor
kw_extractor = yake.KeywordExtractor(lan="en", n=3, top=15)
# Process multiple files
results = {}
for filename in os.listdir('/mnt/user-data/uploads'):
if filename.endswith('.txt'):
with open(f'/mnt/user-data/uploads/{filename}', 'r') as f:
text = f.read()
keywords = kw_extractor.extract_keywords(text)
results[filename] = keywords
# Output results
for filename, keywords in results.items():
print(f"\n{filename}:")
for kw, score in keywords[:10]: # Top 10
print(f" {score:.4f} {kw}")
Multilingual Extraction
import yake
# French document
with open('/mnt/user-data/uploads/article_fr.txt', 'r') as f:
french_text = f.read()
# Extract with French stopwords
kw_fr = yake.KeywordExtractor(lan="fr", n=3, top=20)
keywords_fr = kw_fr.extract_keywords(french_text)
print("Mots-clés (French):")
for kw, score in keywords_fr:
print(f" {score:.4f} {kw}")
# German document
with open('/mnt/user-data/uploads/artikel_de.txt', 'r') as f:
german_text = f.read()
# Extract with German stopwords
kw_de = yake.KeywordExtractor(lan="de", n=3, top=20)
keywords_de = kw_de.extract_keywords(german_text)
print("\nSchlüsselwörter (German):")
for kw, score in keywords_de:
print(f" {score:.4f} {kw}")
Output Formats
Plain Text
for kw, score in keywords:
print(f"{kw}: {score:.4f}")
CSV
import csv
with open('/mnt/user-data/outputs/keywords.csv', 'w', newline='') as f:
writer = csv.writer(f)
writer.writerow(['Keyword', 'Score'])
writer.writerows(keywords)
JSON
import json
output = [{"keyword": kw, "score": score} for kw, score in keywords]
with open('/mnt/user-data/outputs/keywords.json', 'w') as f:
json.dump(output, f, indent=2)
Notes
- Lower scores indicate more important keywords
- YAKE is unsupervised - no training data required
- Supports 34 languages - built-in stopwords for Arabic, Bulgarian, Chinese, Czech, Danish, Dutch, English, Estonian, Farsi, Finnish, French, German, Greek, Hindi, Croatian, Hungarian, Armenian, Indonesian, Italian, Japanese, Lithuanian, Latvian, Norwegian, Polish, Portuguese, Romanian, Russian, Slovak, Slovenian, Spanish, Swedish, Turkish, Ukrainian, and more
- Optimal n-gram size is 2 or 3 for most use cases
- For longer technical phrases (4+ words), consider post-processing or ontology matching
- Always specify full venv path:
/home/claude/yake-venv/bin/python
Troubleshooting
Import errors: Verify venv installation
/home/claude/yake-venv/bin/python -c "import yake; print(yake.__version__)"
Empty results: Check text length (YAKE needs sufficient content, typically 100+ words)
Poor quality keywords: Adjust parameters:
- Increase
dedupLimfor more aggressive deduplication - Try domain-specific stopwords
- Increase
topto see more candidates
Generic terms appearing: Add custom stopwords for your domain:
with open('assets/stopwords_ls.txt', 'r') as f:
stops = set(line.strip().lower() for line in f)
# Add domain-specific terms
stops.update(['term1', 'term2', 'term3'])
kw_extractor = yake.KeywordExtractor(stopwords=stops, n=3, top=20)
Files (claude-skills)
-
assets
-
stopwords_ai.txt 10.7 KB
a a's abilities ability able about above according accordingly accuracy accurate accurately achieve achieved achievement achieves achieving across activation activations actually after afterwards again against agent agentic agents aggregate aggregated aggregates aggregating aggregation ai ain't algorithm algorithmic algorithms all allow allowed allowing allows almost alone along already also although always am among amongst amount amounts an analyses analysis analytical analyze analyzed analyzes analyzing and another any anybody anyhow anyone anything anyway anyways anywhere apart appear appendices appendix application applications applied applies apply applying appreciate approach approaches appropriate architectural architecture architectures are aren't around as aside ask asking aspect aspects associate associated associates associating association at attention attentional attentions attribute attributes available average averaged averaging away awfully b base based baseline baselines basis batch batches be became because become becomes becoming been before beforehand behind being believe below benchmark benchmarked benchmarking benchmarks beside besides best better between beyond bias biased biases both brief but by c c'mon c's calculate calculated calculates calculating calculation came can can't cannot cant capabilities capability capable case cases categorical categories category cause causes certain certainly change changed changes changing characteristic characteristics class classes classification classified classify clearly co collaborate collaborated collaborates collaborating collaboration collaborative com come comes common commonly communicate communicated communicates communicating communication comparative compare compared compares comparing comparison complete completely completion complex complexity component components compose composed composes composing composition compositional computation computational compute computed computer-vision computes computing concerning conclude concluded concludes concluding conclusion condition conditional conditions consecutive consecutively consequently consider considering consist consisted consisting consists contain contained containing contains context contexts contextual coordinate coordinated coordinates coordinating coordination coordinator correction corrections corrective corrector correctors correlate correlated correlating correlation correspond corresponded correspondence corresponding corresponds could couldn't course current currently d data dataset datasets decode decoded decoder decoders decodes decoding decompose decomposed decomposes decomposing decomposition decompositions decrease decreased decreases decreasing deep-learning definitely degree degrees delegate delegated delegates delegating delegation demonstrate demonstrated demonstrates demonstrating demonstration describe described describes describing description despite deviation did didn't difference differences different differently dimension dimensional dimensionality dimensions discuss discussed discusses discussing discussion distribute distributed distribution distributions dl do does doesn't doing don't done double down downwards dr dra during e each edu effective effectively effectiveness efficiency efficient efficiently eg eight either element elemental elements else elsewhere embedding embeddings employ employed employing employs enable enabled enables enabling encode encoded encoder encoders encodes encoding end-to-end enhance enhanced enhancement enhances enhancing enough entire entirely epoch epochs error errors especially et etc evaluate evaluated evaluates evaluating evaluation even ever every everybody everyone everything everywhere ex exactly example examples except executable executing execution executions executor executors exist existed existing exists experiment experimental experimented experimenting experiments extend extended extending extends extension extensions extensive extensively extreme extremely extremes extremity f factor factors far feature featured features few fifth figure figures fine-tune fine-tuned fine-tuning finetune finetuned finetuning first five focus focused focuses focusing followed following follows for form formed former formerly forming forms forth four framework frameworks from full fully function functional functions further furthermore g general generally generate generated generates generation generative get gets getting given gives go goes going gone got gotten gradient gradients greetings h had hadn't happens hardly has hasn't have haven't having he he's hello help hence her here here's hereafter hereby herein hereupon hers herself hi high higher highest him himself his hither hopefully how howbeit however hyperparameter hyperparameters i i'd i'll i'm i've ie if ignored illustrate illustrated illustrates illustrating illustration immediate implement implementation implemented implementing implements importance important importantly improve improved improvement improves improving in inasmuch inc include included includes including increase increased increases increasing indeed indicate indicated indicates indicating indication infer inference inferred infers inner input inputs insofar instance instances instead interact interacted interacting interaction interactions interactive interacts into introduce introduced introduces introducing introduction inward is isn't it it'd it'll it's iterate iterated iterates iterating iteration iterations its itself j just k keep keeps kept key kind kinds know known knows l language-model large large-language-model larger largest last lately later latter latterly layer layered layers learn learned learning learns least less lest let let's level levels like liked likely limit limitation limitations limited limiting limits little llm llms look looking looks loss losses low lower lowest ltd m machine-learning main mainly major manner manners many massive massively matrices matrix may maybe me mean meaning means meanwhile measure measured measurement measures measuring merely method methodologies methodology methods metric metrics microagent microagents might minor ml mode model modeling models modes modification modified modifies modify modifying modular modularity module modules more moreover most mostly mr ms much multiple must my myself n name namely natural-language nd near nearly necessary need needs neither network networks neural neural-network never nevertheless new next nine nlp no nobody non none noone nor normally not nothing novel novelty now nowhere number numbered numbering numbers o observation observations observe observed observes observing obtain obtained obtaining obtains obviously of off often oh ok okay old on once one ones only onto optimization optimize optimized optimizer optimizers optimizes optimizing or orchestrate orchestrated orchestrates orchestrating orchestration other others otherwise ought our ours ourselves out output outputs outside over overall own p paper papers parameter parameters parametric part partial partially particular particularly parts per perform performance performed performing performs perhaps piece pieces placed please plus possible pre-trained pre-training predict predicted prediction predictions predicts present presentation presented presenting presents presumably pretrain pretrained pretraining primary probabilistic probabilities probability probably problem problems process processed processes processing prompt prompted prompting prompts properties property proposal propose proposed proposes proposing provide provided provides providing q qualities quality que queried queries query querying quite qv r rate rates rather rd re really reason reasonably reasoned reasoning reasons recent recently reduce reduced reduces reducing reduction regarding regardless regards reinforcement-learning relate related relates relating relation relations relationship relationships relatively reliability reliable represent representation representations represented representing represents require required requirement requirements requires requiring research researcher researchers respectively respond responded responding responds response responses result resulting results reveal revealed revealing reveals right robust robustness s said same sample sampled samples sampling saw say saying says scalability scalable scale scaled scales scaling score scored scores scoring second secondary secondly section sections see seeing seem seemed seeming seems seen self selves sensible sent sequence sequences sequential serious seriously seven several shall she should shouldn't show showed showing shown shows significance significant significantly similar similarities similarity similarly simple simplicity simply since single six size sized sizes small smaller smallest so solution solutions solvable solve solved solver solvers solves solving some somebody somehow someone something sometime sometimes somewhat somewhere soon sorry specific specifically specification specified specify specifying stability stable stage stages standard standards state state-of-the-art stated states stating statistic statistical statistics step stepped stepping steps still studied studies study studying sub subtask subtasks such suggest suggested suggesting suggestion suggests sup sure system systematic systems t t's table tables take taken task tasks technical technique techniques tell tends tensor tensors test tested testing tests th than thank thanks thanx that that's thats the their theirs them themselves then thence there there's thereafter thereby therefore therein theres thereupon these they they'd they'll they're they've think third this thorough thoroughly those though three threshold thresholds through throughout thru thus time timed times timing to together token tokenization tokenize tokenized tokens too took total totaling totals toward towards train trained training trains transformer transformers tried tries triple truly try trying twice two type types typical typically u un under unfortunately unit units unless unlikely until unto up upon us usage use used useful uses using usually utilization utilize utilized utilizes utilizing uucp v validate validated validates validation value valued values variance variant variants variation variations variety various vector vectorize vectorized vectors version versions very via viz vote voted voter voters votes voting vs w want wants was wasn't way ways we we'd we'll we're we've weight weighted weights welcome well went were weren't what what's whatever when whence whenever where where's whereafter whereas whereby wherein whereupon wherever whether which while whither who who's whoever whole whom whose why will willing wish with within without won't wonder work worked working works worse worst would wouldn't x y yes yet you you'd you'll you're you've your yours yourself yourselves z zero -
stopwords_ls.txt 10 KB
a a's able abnormal abnormality abnormally about above absence absent abstract abstracts accept acceptance accepted accepting accepts access accessed accesses accessing according accordingly accuracy across activate activated activating activation active actively activities activity actually acute acutely affect affected affecting affects after afterwards again against aim aimed aiming aims ain't al all allow allows almost alone along already also although always am among amongst an analyses analysis analytical analyze analyzed analyzes analyzing and another any anybody anyhow anyone anything anyway anyways anywhere apart appear appendices appendix appreciate approach approached approaches approaching appropriate approximately are aren't around article articles as aside ask asking assay assayed assaying assays assess assessed assesses assessing assessment assessments associated association at atypical atypically availability available average averaged averages averaging away awfully b background backgrounds baseline baselines be became because become becomes becoming been before beforehand behind being believe below beside besides best better between beyond both brief but by c c'mon c's came can can't cannot cant carried cause causes cell cells cellular certain certainly changes chromatographic chromatography chronic chronically citation citations cite cited citing clearly clinic clinical clinically clinics co cohort cohorts com come comes common commonly comparative compare compared compares comparing comparison comparisons component components compound compounds concerning conclude concluded concludes concluding conclusion conclusions condition conditional conditions conduct conducted conducting conducts consequently consider consideration considered considering considers contain containing contains control controlled controlling controls correlate correlated correlating correlation correlations corresponding could couldn't course culture cultured cultures culturing currently d data dataset datasets decrease decreased decreases decreasing define defined defines defining definitely definition definitions degree degrees demonstrate demonstrated demonstrates demonstrating demonstration describe described describes describing description descriptions descriptive despite detect detected detecting detection detects determination determine determined determines determining deviate deviated deviating deviation deviations diagnose diagnosed diagnoses diagnosing diagnosis diagnostic diagnostics did didn't different direct directly discussion discussions disease diseased diseases disorder disorders distribute distributed distributing distribution distributions dna do document documentation documented documenting documents does doesn't doing don't done dosage dose dosed doses dosing down download downloaded downloading downloads downwards dr dra draft drafted drafting drafts drug drugs duration durations during e e.g. each edit edited editing editor editorial editors edits edu effect effective effectively effectiveness effects eg eight either element elemental elements elevate elevated elevates elevating elevation else elsewhere enough entirely enzymatic enzyme enzymes especially et etc evaluate evaluated evaluates evaluating evaluation evaluations even ever every everybody everyone everything everywhere ex exactly examination examine examined examines examining example except exclude excluded excludes excluding exclusion experiment experimental experimentation experiments f factor factors far few fifth fig figs figure figures find finding findings finds first five follow-up followed following follows followup for form formal formally formation formed former formerly forming forms forth found four frequencies frequency frequent frequently from function functional functionally functioning functions further furthermore g gene general generally genes genetic genetically genetics get gets getting given gives go goal goals goes going gone got gotten great greater greetings group grouped grouping groups h had hadn't happens hardly has hasn't have haven't having he he's hello help hence her here here's hereafter hereby herein hereupon hers herself hi high higher highest him himself his hither hopefully how howbeit however i i'd i'll i'm i've i.e. identification identified identifies identify identifying ie if ignored image imaged images imaging immediate impact impacted impacting impacts in inasmuch inc include included includes including inclusion increase increased increases increasing indeed indicate indicated indicates indicating indication indicative indirect indirectly influence influenced influences influencing influential informal informally inner insofar instead into introduction introductions introductory investigate investigated investigates investigating investigation investigations inward is isn't it it'd it'll it's its itself j just k keep keeps kept know known knows l last lately later latter latterly least less lesser lest let let's level leveled leveling levels like liked likely little look looking looks low lower lowest ltd m mainly major manuscript manuscripts many may maybe me mean means meanwhile measure measured measurement measurements measures measuring mechanism mechanisms mechanistic median medians merely method methodological methodologies methodology methods microscope microscopes microscopic microscopy might minor model modeled modeling modelled modelling models molecular molecule molecules more moreover most mostly mr ms much must my myself n name namely nd near nearly necessary need needs negative negatively neither never nevertheless new next nine no nobody non none noone nor normal normality normally not note noted notes nothing noting novel now nowhere number numbered numbering numbers numerical numerically o objective objectives observable observation observations observe observed observes observing obtain obtained obtaining obtains obviously of off often oh ok okay old on once one ones only onto or organ organism organismal organisms organs other others otherwise ought our ours ourselves out outside over overall own p page pages paper papers participant participants particular particularly pathological pathologies pathology pathway pathways patient patients pdf pdfs per percent percentage percentages perform performance performed performing performs perhaps period periodic periodically periods phase phases placed please plus population populations positive positively possible possibly potential potentially prediction presence present presented presenting presumably previous primarily primary probable probably procedural procedure procedures process processed processes processing proportion proportional proportionally proportions protein proteins protocol protocols provide provided provides providing publication publications publish published publisher publishers publishes publishing purpose purposes q qualitative qualitatively qualities quality quantification quantified quantify quantifying quantitative quantitatively que quite qv r range ranged ranges ranging rare rarely rate rated rates rather rating rd re really reasonably recent record recorded recording records reduce reduced reduces reducing reduction reductions reference referenced references referencing regarding regardless regards relatively report reported reporting reports require required requirement requirements requires requiring research researched researcher researchers researching respectively respond responded responding response responses responsive responsiveness result resulted resulting results retrieval retrieve retrieved retrieves retrieving reveal revealed revealing reveals review reviewed reviewer reviewers reviewing reviews revise revised revising revision revisions right rna role roles s said same sample sampled samples sampling saw say saying says score scored scores scoring screen screened screening screens second secondarily secondary secondly section sections see seeing seem seemed seeming seems seen select selected selecting selection selects self selves sensible sent sequence sequenced sequences sequencing serious seriously seven several shall she should shouldn't show showed showing shown shows significance significant significantly since six so some somebody somehow someone something sometime sometimes somewhat somewhere soon sorry species specific specifically specified specify specifying spectrometer spectrometric spectroscopy stage staged stages staging stain stained staining stains standard statistic statistical statistically statistics still studied studies study studying sub subject subjects submission submit submits submitted submitting such suggest suggested suggesting suggestion suggestions suggests sup sure symptom symptomatic symptoms syndrome syndromes t t's table tables take taken technical technically technique techniques tell tends test tested testing tests th than thank thanks thanx that that's thats the their theirs them themselves then thence therapeutic therapeutics therapies therapy there there's thereafter thereby therefore therein theres thereupon these they they'd they'll they're they've think third this thorough thoroughly those though three through throughout thru thus time timed times timing tissue tissues to together too took total totaled totaling totals toward towards treat treated treating treatment treatments trial trials tried tries truly try trying twice two type typed types typical typically typing u un uncommon under unfortunately unless unlikely until unto up upload uploaded uploading uploads upon us use used useful uses using usually uucp v valuation value valued values valuing variance variant variants variation variations various version versions versus very via viz vs w want wants was wasn't way we we'd we'll we're we've welcome well went were weren't what what's whatever when whence whenever where where's whereafter whereas whereby wherein whereupon wherever whether which while whither who who's whoever whole whom whose why will willing wish with within without won't wonder would wouldn't write writes writing written wrote x y yes yet you you'd you'll you're you've your yours yourself yourselves z zero zeroanalysis
-
-
README.md 391 B
# extracting-keywords Extract keywords from documents using YAKE algorithm with support for 34 languages (Arabic to Chinese). Use when users request keyword extraction, key terms, topic identification, content summarization, or document analysis. Includes domain-specific stopwords for AI/ML and life sciences. Optional deeper extraction mode (n=2+n=3 combined) for comprehensive coverage. -
SKILL.md 11.4 KB
--- name: extracting-keywords description: Extract keywords from documents using YAKE algorithm with support for 34 languages (Arabic to Chinese). Use when users request keyword extraction, key terms, topic identification, content summarization, or document analysis. Includes domain-specific stopwords for AI/ML and life sciences. Optional deeper extraction mode (n=2+n=3 combined) for comprehensive coverage. metadata: version: 0.2.1 --- # Extracting Keywords Extract keywords from text using YAKE (Yet Another Keyword Extractor), an unsupervised statistical keyword extraction algorithm. ## Installation **First time only:** Install YAKE with optimized dependencies to avoid unnecessary downloads. ```bash cd /home/claude uv venv yake-venv --system-site-packages uv pip install yake --python yake-venv/bin/python --no-deps uv pip install jellyfish segtok regex --python yake-venv/bin/python ``` This reuses system packages (numpy, networkx) instead of downloading them (~0.08s vs ~5s). ## Stopwords Configuration **Built-in YAKE stopwords (34 languages):** Use `lan="<code>"` parameter - See Parameters section below for all 34 supported language codes - English (`lan="en"`) is the default **Custom domain stopwords (bundled in `assets/`):** **AI/ML:** `stopwords_ai.txt` - English stopwords + 783 AI/ML domain-specific terms (1357 total) - Filters AI/ML methodology noise (model, training, network, algorithm, parameter) - Filters ML boilerplate (dataset, baseline, benchmark, experiment, evaluation) - Filters technical terms (transformer, embedding, attention, optimization, inference) - Includes full lemmatization (train/trains/trained/training/trainer) - Use for AI/ML papers, technical reports, machine learning literature - **Performance impact:** +4-5% runtime vs English stopwords **Life Sciences:** `stopwords_ls.txt` - English stopwords + 719 life sciences domain-specific terms (1293 total) - Filters research methodology noise (study, results, analysis, significant, observed) - Filters academic boilerplate (paper, manuscript, publication, review, editing) - Filters statistical terms (correlation, distribution, deviation, variance) - Filters clinical terms (patient, treatment, diagnosis, symptom, therapy) - Filters biology/medicine (cell, tissue, protein, gene, organism) - Includes full lemmatization (analyze/analyzes/analyzed/analyzing/analysis) - Use for biomedical papers, clinical studies, research articles, scientific literature - **Performance impact:** +4-5% runtime vs English stopwords ## Basic Usage ```python import yake # Read text with open('document.txt', 'r') as f: text = f.read() # Extract with English stopwords (default) kw_extractor = yake.KeywordExtractor( lan="en", # Language code n=3, # Max n-gram size (1-3 word phrases) dedupLim=0.9, # Deduplication threshold (0-1) top=20 # Number of keywords to return ) keywords = kw_extractor.extract_keywords(text) # Display results (lower score = more important) for kw, score in keywords: print(f"{score:.4f} {kw}") ``` ## Domain-Specific Extraction ### Using Life Sciences Stopwords **Option 1: Install custom stopwords file** ```bash # Copy life sciences stopwords to YAKE package cp assets/stopwords_ls.txt /home/claude/yake-venv/lib/python3.12/site-packages/yake/core/StopwordsList/stopwords_ls.txt # Use with lan="ls" kw_extractor = yake.KeywordExtractor(lan="ls", n=3, top=20) ``` **Option 2: Load custom stopwords directly** ```python # Load stopwords from file with open('assets/stopwords_ls.txt', 'r') as f: custom_stops = set(line.strip().lower() for line in f) # Pass to extractor kw_extractor = yake.KeywordExtractor( stopwords=custom_stops, n=3, top=20 ) ``` ### Using AI/ML Stopwords ```python # Load AI/ML stopwords with open('/mnt/skills/user/extracting-keywords/assets/stopwords_ai.txt', 'r') as f: ai_stops = set(line.strip().lower() for line in f) # Extract with AI stopwords kw_extractor = yake.KeywordExtractor( stopwords=ai_stops, n=3, top=20 ) keywords = kw_extractor.extract_keywords(text) ``` ## Deeper Extraction (n=2 + n=3 Combined) For more comprehensive extraction, run both n=2 and n=3 and consolidate results. This captures both focused phrases and broader context with ~100% time overhead (still <2s for large documents). ```python import yake # Load domain stopwords with open('/mnt/skills/user/extracting-keywords/assets/stopwords_ai.txt', 'r') as f: stops = set(line.strip().lower() for line in f) # Extract with n=2 (captures focused phrases) kw_n2 = yake.KeywordExtractor(stopwords=stops, n=2, dedupLim=0.9, top=50) results_n2 = kw_n2.extract_keywords(text) # Extract with n=3 (captures broader context) kw_n3 = yake.KeywordExtractor(stopwords=stops, n=3, dedupLim=0.9, top=50) results_n3 = kw_n3.extract_keywords(text) # Consolidate: union with score averaging for overlaps combined = {} for kw, score in results_n2: combined[kw] = score for kw, score in results_n3: if kw in combined: combined[kw] = (combined[kw] + score) / 2 else: combined[kw] = score # Sort by score (lower = more important) consolidated = sorted(combined.items(), key=lambda x: x[1]) # Display top 30 for kw, score in consolidated[:30]: print(f"{score:.4f} {kw}") ``` **Benefits:** - n=2 extracts cleaner domain-specific phrases ("disk move", "error rate") - n=3 captures contextual combinations ("Move disk 1", "per-step error rate") - Consolidation provides richer keyword set for topic modeling or search indexing **Performance:** - Combined approach: ~2x runtime of single extraction - Typical timing: 0.4s (small doc) to 1.0s (large doc) - Use when quality matters more than speed ## Parameters **lan** (str): Language code for built-in stopwords - `"en"` - English (default) - `"ai"` - AI/ML (if stopwords_ai.txt installed in YAKE) - `"ls"` - Life sciences (if stopwords_ls.txt installed in YAKE) **Built-in YAKE languages (34 total):** - `"ar"` - Arabic - `"bg"` - Bulgarian - `"br"` - Breton - `"cz"` - Czech - `"da"` - Danish - `"de"` - German - `"el"` - Greek - `"es"` - Spanish - `"et"` - Estonian - `"fa"` - Farsi/Persian - `"fi"` - Finnish - `"fr"` - French - `"hi"` - Hindi - `"hr"` - Croatian - `"hu"` - Hungarian - `"hy"` - Armenian - `"id"` - Indonesian - `"it"` - Italian - `"ja"` - Japanese - `"lt"` - Lithuanian - `"lv"` - Latvian - `"nl"` - Dutch - `"no"` - Norwegian - `"pl"` - Polish - `"pt"` - Portuguese - `"ro"` - Romanian - `"ru"` - Russian - `"sk"` - Slovak - `"sl"` - Slovenian - `"sv"` - Swedish - `"tr"` - Turkish - `"uk"` - Ukrainian - `"zh"` - Chinese **n** (int): Maximum n-gram size (default: 3) - `1` - Single words only - `2` - Up to 2-word phrases - `3` - Up to 3-word phrases (recommended) - `4-5` - May produce suboptimal results with YAKE's algorithm **dedupLim** (float): Deduplication threshold (default: 0.9) - Range: 0.0 to 1.0 - Higher values = more aggressive deduplication - Controls handling of similar terms (e.g., "cancer cell" vs "cancer cells") **top** (int): Number of keywords to return (default: 20) **stopwords** (set): Custom stopwords set (overrides lan parameter) ## Workflow Patterns ### Single Document Analysis ```python import yake # Read document with open('/mnt/user-data/uploads/article.txt', 'r') as f: text = f.read() # Extract keywords kw_extractor = yake.KeywordExtractor(lan="en", n=3, top=30) keywords = kw_extractor.extract_keywords(text) # Format results results = [] for kw, score in keywords: results.append(f"{score:.4f} {kw}") print("\n".join(results)) ``` ### Comparing Stopwords Strategies ```python import yake # Load life sciences stopwords with open('assets/stopwords_ls.txt', 'r') as f: ls_stops = set(line.strip().lower() for line in f) # Extract with English stopwords kw_en = yake.KeywordExtractor(lan="en", n=3, top=20) keywords_en = kw_en.extract_keywords(text) # Extract with life sciences stopwords kw_ls = yake.KeywordExtractor(stopwords=ls_stops, n=3, top=20) keywords_ls = kw_ls.extract_keywords(text) # Compare results print("English stopwords:") for kw, score in keywords_en: print(f" {score:.4f} {kw}") print("\nLife sciences stopwords:") for kw, score in keywords_ls: print(f" {score:.4f} {kw}") ``` ### Batch Processing ```python import yake import os # Initialize extractor kw_extractor = yake.KeywordExtractor(lan="en", n=3, top=15) # Process multiple files results = {} for filename in os.listdir('/mnt/user-data/uploads'): if filename.endswith('.txt'): with open(f'/mnt/user-data/uploads/{filename}', 'r') as f: text = f.read() keywords = kw_extractor.extract_keywords(text) results[filename] = keywords # Output results for filename, keywords in results.items(): print(f"\n{filename}:") for kw, score in keywords[:10]: # Top 10 print(f" {score:.4f} {kw}") ``` ### Multilingual Extraction ```python import yake # French document with open('/mnt/user-data/uploads/article_fr.txt', 'r') as f: french_text = f.read() # Extract with French stopwords kw_fr = yake.KeywordExtractor(lan="fr", n=3, top=20) keywords_fr = kw_fr.extract_keywords(french_text) print("Mots-clés (French):") for kw, score in keywords_fr: print(f" {score:.4f} {kw}") # German document with open('/mnt/user-data/uploads/artikel_de.txt', 'r') as f: german_text = f.read() # Extract with German stopwords kw_de = yake.KeywordExtractor(lan="de", n=3, top=20) keywords_de = kw_de.extract_keywords(german_text) print("\nSchlüsselwörter (German):") for kw, score in keywords_de: print(f" {score:.4f} {kw}") ``` ## Output Formats ### Plain Text ```python for kw, score in keywords: print(f"{kw}: {score:.4f}") ``` ### CSV ```python import csv with open('/mnt/user-data/outputs/keywords.csv', 'w', newline='') as f: writer = csv.writer(f) writer.writerow(['Keyword', 'Score']) writer.writerows(keywords) ``` ### JSON ```python import json output = [{"keyword": kw, "score": score} for kw, score in keywords] with open('/mnt/user-data/outputs/keywords.json', 'w') as f: json.dump(output, f, indent=2) ``` ## Notes - Lower scores indicate more important keywords - YAKE is unsupervised - no training data required - **Supports 34 languages** - built-in stopwords for Arabic, Bulgarian, Chinese, Czech, Danish, Dutch, English, Estonian, Farsi, Finnish, French, German, Greek, Hindi, Croatian, Hungarian, Armenian, Indonesian, Italian, Japanese, Lithuanian, Latvian, Norwegian, Polish, Portuguese, Romanian, Russian, Slovak, Slovenian, Spanish, Swedish, Turkish, Ukrainian, and more - Optimal n-gram size is 2 or 3 for most use cases - For longer technical phrases (4+ words), consider post-processing or ontology matching - Always specify full venv path: `/home/claude/yake-venv/bin/python` ## Troubleshooting **Import errors:** Verify venv installation ```bash /home/claude/yake-venv/bin/python -c "import yake; print(yake.__version__)" ``` **Empty results:** Check text length (YAKE needs sufficient content, typically 100+ words) **Poor quality keywords:** Adjust parameters: - Increase `dedupLim` for more aggressive deduplication - Try domain-specific stopwords - Increase `top` to see more candidates **Generic terms appearing:** Add custom stopwords for your domain: ```python with open('assets/stopwords_ls.txt', 'r') as f: stops = set(line.strip().lower() for line in f) # Add domain-specific terms stops.update(['term1', 'term2', 'term3']) kw_extractor = yake.KeywordExtractor(stopwords=stops, n=3, top=20) ```
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.