Claude Skill

extracting-keywords

Extract keywords from documents using YAKE algorithm with support for 34 languages (Arabic to Chinese). Use when users request keyword extraction, key terms, topic identification, content summarization, or document analysis. Includes domain-specific stopwords for AI/ML and life s

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download oaustegard-claude-skills-plugins_data-and-visualization_skills_extracting-keywords-e39c726.zip · 11 KB
Part of oaustegard/claude-skills — 39 skills

Install

skills CLI npx skills add https://github.com/oaustegard/claude-skills/tree/main/plugins/data-and-visualization/skills/extracting-keywords
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install oaustegard-claude-skills@llmmart
Git git clone https://github.com/oaustegard/claude-skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole oaustegard/claude-skills collection as a plugin from our marketplace. Git is the plain clone.

README

extracting-keywords

Extract keywords from documents using YAKE algorithm with support for 34 languages (Arabic to Chinese). Use when users request keyword extraction, key terms, topic identification, content summarization, or document analysis. Includes domain-specific stopwords for AI/ML and life sciences. Optional deeper extraction mode (n=2+n=3 combined) for comprehensive coverage.

Skill manifest

Extracting Keywords

Extract keywords from text using YAKE (Yet Another Keyword Extractor), an unsupervised statistical keyword extraction algorithm.

Installation

First time only: Install YAKE with optimized dependencies to avoid unnecessary downloads.

cd /home/claude
uv venv yake-venv --system-site-packages
uv pip install yake --python yake-venv/bin/python --no-deps
uv pip install jellyfish segtok regex --python yake-venv/bin/python

This reuses system packages (numpy, networkx) instead of downloading them (~0.08s vs ~5s).

Stopwords Configuration

Built-in YAKE stopwords (34 languages): Use lan="<code>" parameter

  • See Parameters section below for all 34 supported language codes
  • English (lan="en") is the default

Custom domain stopwords (bundled in assets/):

AI/ML: stopwords_ai.txt

  • English stopwords + 783 AI/ML domain-specific terms (1357 total)
  • Filters AI/ML methodology noise (model, training, network, algorithm, parameter)
  • Filters ML boilerplate (dataset, baseline, benchmark, experiment, evaluation)
  • Filters technical terms (transformer, embedding, attention, optimization, inference)
  • Includes full lemmatization (train/trains/trained/training/trainer)
  • Use for AI/ML papers, technical reports, machine learning literature
  • Performance impact: +4-5% runtime vs English stopwords

Life Sciences: stopwords_ls.txt

  • English stopwords + 719 life sciences domain-specific terms (1293 total)
  • Filters research methodology noise (study, results, analysis, significant, observed)
  • Filters academic boilerplate (paper, manuscript, publication, review, editing)
  • Filters statistical terms (correlation, distribution, deviation, variance)
  • Filters clinical terms (patient, treatment, diagnosis, symptom, therapy)
  • Filters biology/medicine (cell, tissue, protein, gene, organism)
  • Includes full lemmatization (analyze/analyzes/analyzed/analyzing/analysis)
  • Use for biomedical papers, clinical studies, research articles, scientific literature
  • Performance impact: +4-5% runtime vs English stopwords

Basic Usage

import yake

# Read text
with open('document.txt', 'r') as f:
    text = f.read()

# Extract with English stopwords (default)
kw_extractor = yake.KeywordExtractor(
    lan="en",           # Language code
    n=3,                # Max n-gram size (1-3 word phrases)
    dedupLim=0.9,       # Deduplication threshold (0-1)
    top=20              # Number of keywords to return
)

keywords = kw_extractor.extract_keywords(text)

# Display results (lower score = more important)
for kw, score in keywords:
    print(f"{score:.4f}  {kw}")

Domain-Specific Extraction

Using Life Sciences Stopwords

Option 1: Install custom stopwords file

# Copy life sciences stopwords to YAKE package
cp assets/stopwords_ls.txt /home/claude/yake-venv/lib/python3.12/site-packages/yake/core/StopwordsList/stopwords_ls.txt

# Use with lan="ls"
kw_extractor = yake.KeywordExtractor(lan="ls", n=3, top=20)

Option 2: Load custom stopwords directly

# Load stopwords from file
with open('assets/stopwords_ls.txt', 'r') as f:
    custom_stops = set(line.strip().lower() for line in f)

# Pass to extractor
kw_extractor = yake.KeywordExtractor(
    stopwords=custom_stops,
    n=3,
    top=20
)

Using AI/ML Stopwords

# Load AI/ML stopwords
with open('/mnt/skills/user/extracting-keywords/assets/stopwords_ai.txt', 'r') as f:
    ai_stops = set(line.strip().lower() for line in f)

# Extract with AI stopwords
kw_extractor = yake.KeywordExtractor(
    stopwords=ai_stops,
    n=3,
    top=20
)
keywords = kw_extractor.extract_keywords(text)

Deeper Extraction (n=2 + n=3 Combined)

For more comprehensive extraction, run both n=2 and n=3 and consolidate results. This captures both focused phrases and broader context with ~100% time overhead (still <2s for large documents).

import yake

# Load domain stopwords
with open('/mnt/skills/user/extracting-keywords/assets/stopwords_ai.txt', 'r') as f:
    stops = set(line.strip().lower() for line in f)

# Extract with n=2 (captures focused phrases)
kw_n2 = yake.KeywordExtractor(stopwords=stops, n=2, dedupLim=0.9, top=50)
results_n2 = kw_n2.extract_keywords(text)

# Extract with n=3 (captures broader context)
kw_n3 = yake.KeywordExtractor(stopwords=stops, n=3, dedupLim=0.9, top=50)
results_n3 = kw_n3.extract_keywords(text)

# Consolidate: union with score averaging for overlaps
combined = {}
for kw, score in results_n2:
    combined[kw] = score
for kw, score in results_n3:
    if kw in combined:
        combined[kw] = (combined[kw] + score) / 2
    else:
        combined[kw] = score

# Sort by score (lower = more important)
consolidated = sorted(combined.items(), key=lambda x: x[1])

# Display top 30
for kw, score in consolidated[:30]:
    print(f"{score:.4f}  {kw}")

Benefits:

  • n=2 extracts cleaner domain-specific phrases ("disk move", "error rate")
  • n=3 captures contextual combinations ("Move disk 1", "per-step error rate")
  • Consolidation provides richer keyword set for topic modeling or search indexing

Performance:

  • Combined approach: ~2x runtime of single extraction
  • Typical timing: 0.4s (small doc) to 1.0s (large doc)
  • Use when quality matters more than speed

Parameters

lan (str): Language code for built-in stopwords

  • "en" - English (default)
  • "ai" - AI/ML (if stopwords_ai.txt installed in YAKE)
  • "ls" - Life sciences (if stopwords_ls.txt installed in YAKE)

Built-in YAKE languages (34 total):

  • "ar" - Arabic
  • "bg" - Bulgarian
  • "br" - Breton
  • "cz" - Czech
  • "da" - Danish
  • "de" - German
  • "el" - Greek
  • "es" - Spanish
  • "et" - Estonian
  • "fa" - Farsi/Persian
  • "fi" - Finnish
  • "fr" - French
  • "hi" - Hindi
  • "hr" - Croatian
  • "hu" - Hungarian
  • "hy" - Armenian
  • "id" - Indonesian
  • "it" - Italian
  • "ja" - Japanese
  • "lt" - Lithuanian
  • "lv" - Latvian
  • "nl" - Dutch
  • "no" - Norwegian
  • "pl" - Polish
  • "pt" - Portuguese
  • "ro" - Romanian
  • "ru" - Russian
  • "sk" - Slovak
  • "sl" - Slovenian
  • "sv" - Swedish
  • "tr" - Turkish
  • "uk" - Ukrainian
  • "zh" - Chinese

n (int): Maximum n-gram size (default: 3)

  • 1 - Single words only
  • 2 - Up to 2-word phrases
  • 3 - Up to 3-word phrases (recommended)
  • 4-5 - May produce suboptimal results with YAKE's algorithm

dedupLim (float): Deduplication threshold (default: 0.9)

  • Range: 0.0 to 1.0
  • Higher values = more aggressive deduplication
  • Controls handling of similar terms (e.g., "cancer cell" vs "cancer cells")

top (int): Number of keywords to return (default: 20)

stopwords (set): Custom stopwords set (overrides lan parameter)

Workflow Patterns

Single Document Analysis

import yake

# Read document
with open('/mnt/user-data/uploads/article.txt', 'r') as f:
    text = f.read()

# Extract keywords
kw_extractor = yake.KeywordExtractor(lan="en", n=3, top=30)
keywords = kw_extractor.extract_keywords(text)

# Format results
results = []
for kw, score in keywords:
    results.append(f"{score:.4f}  {kw}")

print("\n".join(results))

Comparing Stopwords Strategies

import yake

# Load life sciences stopwords
with open('assets/stopwords_ls.txt', 'r') as f:
    ls_stops = set(line.strip().lower() for line in f)

# Extract with English stopwords
kw_en = yake.KeywordExtractor(lan="en", n=3, top=20)
keywords_en = kw_en.extract_keywords(text)

# Extract with life sciences stopwords
kw_ls = yake.KeywordExtractor(stopwords=ls_stops, n=3, top=20)
keywords_ls = kw_ls.extract_keywords(text)

# Compare results
print("English stopwords:")
for kw, score in keywords_en:
    print(f"  {score:.4f}  {kw}")

print("\nLife sciences stopwords:")
for kw, score in keywords_ls:
    print(f"  {score:.4f}  {kw}")

Batch Processing

import yake
import os

# Initialize extractor
kw_extractor = yake.KeywordExtractor(lan="en", n=3, top=15)

# Process multiple files
results = {}
for filename in os.listdir('/mnt/user-data/uploads'):
    if filename.endswith('.txt'):
        with open(f'/mnt/user-data/uploads/{filename}', 'r') as f:
            text = f.read()
        
        keywords = kw_extractor.extract_keywords(text)
        results[filename] = keywords

# Output results
for filename, keywords in results.items():
    print(f"\n{filename}:")
    for kw, score in keywords[:10]:  # Top 10
        print(f"  {score:.4f}  {kw}")

Multilingual Extraction

import yake

# French document
with open('/mnt/user-data/uploads/article_fr.txt', 'r') as f:
    french_text = f.read()

# Extract with French stopwords
kw_fr = yake.KeywordExtractor(lan="fr", n=3, top=20)
keywords_fr = kw_fr.extract_keywords(french_text)

print("Mots-clés (French):")
for kw, score in keywords_fr:
    print(f"  {score:.4f}  {kw}")

# German document
with open('/mnt/user-data/uploads/artikel_de.txt', 'r') as f:
    german_text = f.read()

# Extract with German stopwords
kw_de = yake.KeywordExtractor(lan="de", n=3, top=20)
keywords_de = kw_de.extract_keywords(german_text)

print("\nSchlüsselwörter (German):")
for kw, score in keywords_de:
    print(f"  {score:.4f}  {kw}")

Output Formats

Plain Text

for kw, score in keywords:
    print(f"{kw}: {score:.4f}")

CSV

import csv

with open('/mnt/user-data/outputs/keywords.csv', 'w', newline='') as f:
    writer = csv.writer(f)
    writer.writerow(['Keyword', 'Score'])
    writer.writerows(keywords)

JSON

import json

output = [{"keyword": kw, "score": score} for kw, score in keywords]
with open('/mnt/user-data/outputs/keywords.json', 'w') as f:
    json.dump(output, f, indent=2)

Notes

  • Lower scores indicate more important keywords
  • YAKE is unsupervised - no training data required
  • Supports 34 languages - built-in stopwords for Arabic, Bulgarian, Chinese, Czech, Danish, Dutch, English, Estonian, Farsi, Finnish, French, German, Greek, Hindi, Croatian, Hungarian, Armenian, Indonesian, Italian, Japanese, Lithuanian, Latvian, Norwegian, Polish, Portuguese, Romanian, Russian, Slovak, Slovenian, Spanish, Swedish, Turkish, Ukrainian, and more
  • Optimal n-gram size is 2 or 3 for most use cases
  • For longer technical phrases (4+ words), consider post-processing or ontology matching
  • Always specify full venv path: /home/claude/yake-venv/bin/python

Troubleshooting

Import errors: Verify venv installation

/home/claude/yake-venv/bin/python -c "import yake; print(yake.__version__)"

Empty results: Check text length (YAKE needs sufficient content, typically 100+ words)

Poor quality keywords: Adjust parameters:

  • Increase dedupLim for more aggressive deduplication
  • Try domain-specific stopwords
  • Increase top to see more candidates

Generic terms appearing: Add custom stopwords for your domain:

with open('assets/stopwords_ls.txt', 'r') as f:
    stops = set(line.strip().lower() for line in f)

# Add domain-specific terms
stops.update(['term1', 'term2', 'term3'])

kw_extractor = yake.KeywordExtractor(stopwords=stops, n=3, top=20)
Files (claude-skills)
  • assets
    • stopwords_ai.txt 10.7 KB
      a
      a's
      abilities
      ability
      able
      about
      above
      according
      accordingly
      accuracy
      accurate
      accurately
      achieve
      achieved
      achievement
      achieves
      achieving
      across
      activation
      activations
      actually
      after
      afterwards
      again
      against
      agent
      agentic
      agents
      aggregate
      aggregated
      aggregates
      aggregating
      aggregation
      ai
      ain't
      algorithm
      algorithmic
      algorithms
      all
      allow
      allowed
      allowing
      allows
      almost
      alone
      along
      already
      also
      although
      always
      am
      among
      amongst
      amount
      amounts
      an
      analyses
      analysis
      analytical
      analyze
      analyzed
      analyzes
      analyzing
      and
      another
      any
      anybody
      anyhow
      anyone
      anything
      anyway
      anyways
      anywhere
      apart
      appear
      appendices
      appendix
      application
      applications
      applied
      applies
      apply
      applying
      appreciate
      approach
      approaches
      appropriate
      architectural
      architecture
      architectures
      are
      aren't
      around
      as
      aside
      ask
      asking
      aspect
      aspects
      associate
      associated
      associates
      associating
      association
      at
      attention
      attentional
      attentions
      attribute
      attributes
      available
      average
      averaged
      averaging
      away
      awfully
      b
      base
      based
      baseline
      baselines
      basis
      batch
      batches
      be
      became
      because
      become
      becomes
      becoming
      been
      before
      beforehand
      behind
      being
      believe
      below
      benchmark
      benchmarked
      benchmarking
      benchmarks
      beside
      besides
      best
      better
      between
      beyond
      bias
      biased
      biases
      both
      brief
      but
      by
      c
      c'mon
      c's
      calculate
      calculated
      calculates
      calculating
      calculation
      came
      can
      can't
      cannot
      cant
      capabilities
      capability
      capable
      case
      cases
      categorical
      categories
      category
      cause
      causes
      certain
      certainly
      change
      changed
      changes
      changing
      characteristic
      characteristics
      class
      classes
      classification
      classified
      classify
      clearly
      co
      collaborate
      collaborated
      collaborates
      collaborating
      collaboration
      collaborative
      com
      come
      comes
      common
      commonly
      communicate
      communicated
      communicates
      communicating
      communication
      comparative
      compare
      compared
      compares
      comparing
      comparison
      complete
      completely
      completion
      complex
      complexity
      component
      components
      compose
      composed
      composes
      composing
      composition
      compositional
      computation
      computational
      compute
      computed
      computer-vision
      computes
      computing
      concerning
      conclude
      concluded
      concludes
      concluding
      conclusion
      condition
      conditional
      conditions
      consecutive
      consecutively
      consequently
      consider
      considering
      consist
      consisted
      consisting
      consists
      contain
      contained
      containing
      contains
      context
      contexts
      contextual
      coordinate
      coordinated
      coordinates
      coordinating
      coordination
      coordinator
      correction
      corrections
      corrective
      corrector
      correctors
      correlate
      correlated
      correlating
      correlation
      correspond
      corresponded
      correspondence
      corresponding
      corresponds
      could
      couldn't
      course
      current
      currently
      d
      data
      dataset
      datasets
      decode
      decoded
      decoder
      decoders
      decodes
      decoding
      decompose
      decomposed
      decomposes
      decomposing
      decomposition
      decompositions
      decrease
      decreased
      decreases
      decreasing
      deep-learning
      definitely
      degree
      degrees
      delegate
      delegated
      delegates
      delegating
      delegation
      demonstrate
      demonstrated
      demonstrates
      demonstrating
      demonstration
      describe
      described
      describes
      describing
      description
      despite
      deviation
      did
      didn't
      difference
      differences
      different
      differently
      dimension
      dimensional
      dimensionality
      dimensions
      discuss
      discussed
      discusses
      discussing
      discussion
      distribute
      distributed
      distribution
      distributions
      dl
      do
      does
      doesn't
      doing
      don't
      done
      double
      down
      downwards
      dr
      dra
      during
      e
      each
      edu
      effective
      effectively
      effectiveness
      efficiency
      efficient
      efficiently
      eg
      eight
      either
      element
      elemental
      elements
      else
      elsewhere
      embedding
      embeddings
      employ
      employed
      employing
      employs
      enable
      enabled
      enables
      enabling
      encode
      encoded
      encoder
      encoders
      encodes
      encoding
      end-to-end
      enhance
      enhanced
      enhancement
      enhances
      enhancing
      enough
      entire
      entirely
      epoch
      epochs
      error
      errors
      especially
      et
      etc
      evaluate
      evaluated
      evaluates
      evaluating
      evaluation
      even
      ever
      every
      everybody
      everyone
      everything
      everywhere
      ex
      exactly
      example
      examples
      except
      executable
      executing
      execution
      executions
      executor
      executors
      exist
      existed
      existing
      exists
      experiment
      experimental
      experimented
      experimenting
      experiments
      extend
      extended
      extending
      extends
      extension
      extensions
      extensive
      extensively
      extreme
      extremely
      extremes
      extremity
      f
      factor
      factors
      far
      feature
      featured
      features
      few
      fifth
      figure
      figures
      fine-tune
      fine-tuned
      fine-tuning
      finetune
      finetuned
      finetuning
      first
      five
      focus
      focused
      focuses
      focusing
      followed
      following
      follows
      for
      form
      formed
      former
      formerly
      forming
      forms
      forth
      four
      framework
      frameworks
      from
      full
      fully
      function
      functional
      functions
      further
      furthermore
      g
      general
      generally
      generate
      generated
      generates
      generation
      generative
      get
      gets
      getting
      given
      gives
      go
      goes
      going
      gone
      got
      gotten
      gradient
      gradients
      greetings
      h
      had
      hadn't
      happens
      hardly
      has
      hasn't
      have
      haven't
      having
      he
      he's
      hello
      help
      hence
      her
      here
      here's
      hereafter
      hereby
      herein
      hereupon
      hers
      herself
      hi
      high
      higher
      highest
      him
      himself
      his
      hither
      hopefully
      how
      howbeit
      however
      hyperparameter
      hyperparameters
      i
      i'd
      i'll
      i'm
      i've
      ie
      if
      ignored
      illustrate
      illustrated
      illustrates
      illustrating
      illustration
      immediate
      implement
      implementation
      implemented
      implementing
      implements
      importance
      important
      importantly
      improve
      improved
      improvement
      improves
      improving
      in
      inasmuch
      inc
      include
      included
      includes
      including
      increase
      increased
      increases
      increasing
      indeed
      indicate
      indicated
      indicates
      indicating
      indication
      infer
      inference
      inferred
      infers
      inner
      input
      inputs
      insofar
      instance
      instances
      instead
      interact
      interacted
      interacting
      interaction
      interactions
      interactive
      interacts
      into
      introduce
      introduced
      introduces
      introducing
      introduction
      inward
      is
      isn't
      it
      it'd
      it'll
      it's
      iterate
      iterated
      iterates
      iterating
      iteration
      iterations
      its
      itself
      j
      just
      k
      keep
      keeps
      kept
      key
      kind
      kinds
      know
      known
      knows
      l
      language-model
      large
      large-language-model
      larger
      largest
      last
      lately
      later
      latter
      latterly
      layer
      layered
      layers
      learn
      learned
      learning
      learns
      least
      less
      lest
      let
      let's
      level
      levels
      like
      liked
      likely
      limit
      limitation
      limitations
      limited
      limiting
      limits
      little
      llm
      llms
      look
      looking
      looks
      loss
      losses
      low
      lower
      lowest
      ltd
      m
      machine-learning
      main
      mainly
      major
      manner
      manners
      many
      massive
      massively
      matrices
      matrix
      may
      maybe
      me
      mean
      meaning
      means
      meanwhile
      measure
      measured
      measurement
      measures
      measuring
      merely
      method
      methodologies
      methodology
      methods
      metric
      metrics
      microagent
      microagents
      might
      minor
      ml
      mode
      model
      modeling
      models
      modes
      modification
      modified
      modifies
      modify
      modifying
      modular
      modularity
      module
      modules
      more
      moreover
      most
      mostly
      mr
      ms
      much
      multiple
      must
      my
      myself
      n
      name
      namely
      natural-language
      nd
      near
      nearly
      necessary
      need
      needs
      neither
      network
      networks
      neural
      neural-network
      never
      nevertheless
      new
      next
      nine
      nlp
      no
      nobody
      non
      none
      noone
      nor
      normally
      not
      nothing
      novel
      novelty
      now
      nowhere
      number
      numbered
      numbering
      numbers
      o
      observation
      observations
      observe
      observed
      observes
      observing
      obtain
      obtained
      obtaining
      obtains
      obviously
      of
      off
      often
      oh
      ok
      okay
      old
      on
      once
      one
      ones
      only
      onto
      optimization
      optimize
      optimized
      optimizer
      optimizers
      optimizes
      optimizing
      or
      orchestrate
      orchestrated
      orchestrates
      orchestrating
      orchestration
      other
      others
      otherwise
      ought
      our
      ours
      ourselves
      out
      output
      outputs
      outside
      over
      overall
      own
      p
      paper
      papers
      parameter
      parameters
      parametric
      part
      partial
      partially
      particular
      particularly
      parts
      per
      perform
      performance
      performed
      performing
      performs
      perhaps
      piece
      pieces
      placed
      please
      plus
      possible
      pre-trained
      pre-training
      predict
      predicted
      prediction
      predictions
      predicts
      present
      presentation
      presented
      presenting
      presents
      presumably
      pretrain
      pretrained
      pretraining
      primary
      probabilistic
      probabilities
      probability
      probably
      problem
      problems
      process
      processed
      processes
      processing
      prompt
      prompted
      prompting
      prompts
      properties
      property
      proposal
      propose
      proposed
      proposes
      proposing
      provide
      provided
      provides
      providing
      q
      qualities
      quality
      que
      queried
      queries
      query
      querying
      quite
      qv
      r
      rate
      rates
      rather
      rd
      re
      really
      reason
      reasonably
      reasoned
      reasoning
      reasons
      recent
      recently
      reduce
      reduced
      reduces
      reducing
      reduction
      regarding
      regardless
      regards
      reinforcement-learning
      relate
      related
      relates
      relating
      relation
      relations
      relationship
      relationships
      relatively
      reliability
      reliable
      represent
      representation
      representations
      represented
      representing
      represents
      require
      required
      requirement
      requirements
      requires
      requiring
      research
      researcher
      researchers
      respectively
      respond
      responded
      responding
      responds
      response
      responses
      result
      resulting
      results
      reveal
      revealed
      revealing
      reveals
      right
      robust
      robustness
      s
      said
      same
      sample
      sampled
      samples
      sampling
      saw
      say
      saying
      says
      scalability
      scalable
      scale
      scaled
      scales
      scaling
      score
      scored
      scores
      scoring
      second
      secondary
      secondly
      section
      sections
      see
      seeing
      seem
      seemed
      seeming
      seems
      seen
      self
      selves
      sensible
      sent
      sequence
      sequences
      sequential
      serious
      seriously
      seven
      several
      shall
      she
      should
      shouldn't
      show
      showed
      showing
      shown
      shows
      significance
      significant
      significantly
      similar
      similarities
      similarity
      similarly
      simple
      simplicity
      simply
      since
      single
      six
      size
      sized
      sizes
      small
      smaller
      smallest
      so
      solution
      solutions
      solvable
      solve
      solved
      solver
      solvers
      solves
      solving
      some
      somebody
      somehow
      someone
      something
      sometime
      sometimes
      somewhat
      somewhere
      soon
      sorry
      specific
      specifically
      specification
      specified
      specify
      specifying
      stability
      stable
      stage
      stages
      standard
      standards
      state
      state-of-the-art
      stated
      states
      stating
      statistic
      statistical
      statistics
      step
      stepped
      stepping
      steps
      still
      studied
      studies
      study
      studying
      sub
      subtask
      subtasks
      such
      suggest
      suggested
      suggesting
      suggestion
      suggests
      sup
      sure
      system
      systematic
      systems
      t
      t's
      table
      tables
      take
      taken
      task
      tasks
      technical
      technique
      techniques
      tell
      tends
      tensor
      tensors
      test
      tested
      testing
      tests
      th
      than
      thank
      thanks
      thanx
      that
      that's
      thats
      the
      their
      theirs
      them
      themselves
      then
      thence
      there
      there's
      thereafter
      thereby
      therefore
      therein
      theres
      thereupon
      these
      they
      they'd
      they'll
      they're
      they've
      think
      third
      this
      thorough
      thoroughly
      those
      though
      three
      threshold
      thresholds
      through
      throughout
      thru
      thus
      time
      timed
      times
      timing
      to
      together
      token
      tokenization
      tokenize
      tokenized
      tokens
      too
      took
      total
      totaling
      totals
      toward
      towards
      train
      trained
      training
      trains
      transformer
      transformers
      tried
      tries
      triple
      truly
      try
      trying
      twice
      two
      type
      types
      typical
      typically
      u
      un
      under
      unfortunately
      unit
      units
      unless
      unlikely
      until
      unto
      up
      upon
      us
      usage
      use
      used
      useful
      uses
      using
      usually
      utilization
      utilize
      utilized
      utilizes
      utilizing
      uucp
      v
      validate
      validated
      validates
      validation
      value
      valued
      values
      variance
      variant
      variants
      variation
      variations
      variety
      various
      vector
      vectorize
      vectorized
      vectors
      version
      versions
      very
      via
      viz
      vote
      voted
      voter
      voters
      votes
      voting
      vs
      w
      want
      wants
      was
      wasn't
      way
      ways
      we
      we'd
      we'll
      we're
      we've
      weight
      weighted
      weights
      welcome
      well
      went
      were
      weren't
      what
      what's
      whatever
      when
      whence
      whenever
      where
      where's
      whereafter
      whereas
      whereby
      wherein
      whereupon
      wherever
      whether
      which
      while
      whither
      who
      who's
      whoever
      whole
      whom
      whose
      why
      will
      willing
      wish
      with
      within
      without
      won't
      wonder
      work
      worked
      working
      works
      worse
      worst
      would
      wouldn't
      x
      y
      yes
      yet
      you
      you'd
      you'll
      you're
      you've
      your
      yours
      yourself
      yourselves
      z
      zero
      
    • stopwords_ls.txt 10 KB
      a
      a's
      able
      abnormal
      abnormality
      abnormally
      about
      above
      absence
      absent
      abstract
      abstracts
      accept
      acceptance
      accepted
      accepting
      accepts
      access
      accessed
      accesses
      accessing
      according
      accordingly
      accuracy
      across
      activate
      activated
      activating
      activation
      active
      actively
      activities
      activity
      actually
      acute
      acutely
      affect
      affected
      affecting
      affects
      after
      afterwards
      again
      against
      aim
      aimed
      aiming
      aims
      ain't
      al
      all
      allow
      allows
      almost
      alone
      along
      already
      also
      although
      always
      am
      among
      amongst
      an
      analyses
      analysis
      analytical
      analyze
      analyzed
      analyzes
      analyzing
      and
      another
      any
      anybody
      anyhow
      anyone
      anything
      anyway
      anyways
      anywhere
      apart
      appear
      appendices
      appendix
      appreciate
      approach
      approached
      approaches
      approaching
      appropriate
      approximately
      are
      aren't
      around
      article
      articles
      as
      aside
      ask
      asking
      assay
      assayed
      assaying
      assays
      assess
      assessed
      assesses
      assessing
      assessment
      assessments
      associated
      association
      at
      atypical
      atypically
      availability
      available
      average
      averaged
      averages
      averaging
      away
      awfully
      b
      background
      backgrounds
      baseline
      baselines
      be
      became
      because
      become
      becomes
      becoming
      been
      before
      beforehand
      behind
      being
      believe
      below
      beside
      besides
      best
      better
      between
      beyond
      both
      brief
      but
      by
      c
      c'mon
      c's
      came
      can
      can't
      cannot
      cant
      carried
      cause
      causes
      cell
      cells
      cellular
      certain
      certainly
      changes
      chromatographic
      chromatography
      chronic
      chronically
      citation
      citations
      cite
      cited
      citing
      clearly
      clinic
      clinical
      clinically
      clinics
      co
      cohort
      cohorts
      com
      come
      comes
      common
      commonly
      comparative
      compare
      compared
      compares
      comparing
      comparison
      comparisons
      component
      components
      compound
      compounds
      concerning
      conclude
      concluded
      concludes
      concluding
      conclusion
      conclusions
      condition
      conditional
      conditions
      conduct
      conducted
      conducting
      conducts
      consequently
      consider
      consideration
      considered
      considering
      considers
      contain
      containing
      contains
      control
      controlled
      controlling
      controls
      correlate
      correlated
      correlating
      correlation
      correlations
      corresponding
      could
      couldn't
      course
      culture
      cultured
      cultures
      culturing
      currently
      d
      data
      dataset
      datasets
      decrease
      decreased
      decreases
      decreasing
      define
      defined
      defines
      defining
      definitely
      definition
      definitions
      degree
      degrees
      demonstrate
      demonstrated
      demonstrates
      demonstrating
      demonstration
      describe
      described
      describes
      describing
      description
      descriptions
      descriptive
      despite
      detect
      detected
      detecting
      detection
      detects
      determination
      determine
      determined
      determines
      determining
      deviate
      deviated
      deviating
      deviation
      deviations
      diagnose
      diagnosed
      diagnoses
      diagnosing
      diagnosis
      diagnostic
      diagnostics
      did
      didn't
      different
      direct
      directly
      discussion
      discussions
      disease
      diseased
      diseases
      disorder
      disorders
      distribute
      distributed
      distributing
      distribution
      distributions
      dna
      do
      document
      documentation
      documented
      documenting
      documents
      does
      doesn't
      doing
      don't
      done
      dosage
      dose
      dosed
      doses
      dosing
      down
      download
      downloaded
      downloading
      downloads
      downwards
      dr
      dra
      draft
      drafted
      drafting
      drafts
      drug
      drugs
      duration
      durations
      during
      e
      e.g.
      each
      edit
      edited
      editing
      editor
      editorial
      editors
      edits
      edu
      effect
      effective
      effectively
      effectiveness
      effects
      eg
      eight
      either
      element
      elemental
      elements
      elevate
      elevated
      elevates
      elevating
      elevation
      else
      elsewhere
      enough
      entirely
      enzymatic
      enzyme
      enzymes
      especially
      et
      etc
      evaluate
      evaluated
      evaluates
      evaluating
      evaluation
      evaluations
      even
      ever
      every
      everybody
      everyone
      everything
      everywhere
      ex
      exactly
      examination
      examine
      examined
      examines
      examining
      example
      except
      exclude
      excluded
      excludes
      excluding
      exclusion
      experiment
      experimental
      experimentation
      experiments
      f
      factor
      factors
      far
      few
      fifth
      fig
      figs
      figure
      figures
      find
      finding
      findings
      finds
      first
      five
      follow-up
      followed
      following
      follows
      followup
      for
      form
      formal
      formally
      formation
      formed
      former
      formerly
      forming
      forms
      forth
      found
      four
      frequencies
      frequency
      frequent
      frequently
      from
      function
      functional
      functionally
      functioning
      functions
      further
      furthermore
      g
      gene
      general
      generally
      genes
      genetic
      genetically
      genetics
      get
      gets
      getting
      given
      gives
      go
      goal
      goals
      goes
      going
      gone
      got
      gotten
      great
      greater
      greetings
      group
      grouped
      grouping
      groups
      h
      had
      hadn't
      happens
      hardly
      has
      hasn't
      have
      haven't
      having
      he
      he's
      hello
      help
      hence
      her
      here
      here's
      hereafter
      hereby
      herein
      hereupon
      hers
      herself
      hi
      high
      higher
      highest
      him
      himself
      his
      hither
      hopefully
      how
      howbeit
      however
      i
      i'd
      i'll
      i'm
      i've
      i.e.
      identification
      identified
      identifies
      identify
      identifying
      ie
      if
      ignored
      image
      imaged
      images
      imaging
      immediate
      impact
      impacted
      impacting
      impacts
      in
      inasmuch
      inc
      include
      included
      includes
      including
      inclusion
      increase
      increased
      increases
      increasing
      indeed
      indicate
      indicated
      indicates
      indicating
      indication
      indicative
      indirect
      indirectly
      influence
      influenced
      influences
      influencing
      influential
      informal
      informally
      inner
      insofar
      instead
      into
      introduction
      introductions
      introductory
      investigate
      investigated
      investigates
      investigating
      investigation
      investigations
      inward
      is
      isn't
      it
      it'd
      it'll
      it's
      its
      itself
      j
      just
      k
      keep
      keeps
      kept
      know
      known
      knows
      l
      last
      lately
      later
      latter
      latterly
      least
      less
      lesser
      lest
      let
      let's
      level
      leveled
      leveling
      levels
      like
      liked
      likely
      little
      look
      looking
      looks
      low
      lower
      lowest
      ltd
      m
      mainly
      major
      manuscript
      manuscripts
      many
      may
      maybe
      me
      mean
      means
      meanwhile
      measure
      measured
      measurement
      measurements
      measures
      measuring
      mechanism
      mechanisms
      mechanistic
      median
      medians
      merely
      method
      methodological
      methodologies
      methodology
      methods
      microscope
      microscopes
      microscopic
      microscopy
      might
      minor
      model
      modeled
      modeling
      modelled
      modelling
      models
      molecular
      molecule
      molecules
      more
      moreover
      most
      mostly
      mr
      ms
      much
      must
      my
      myself
      n
      name
      namely
      nd
      near
      nearly
      necessary
      need
      needs
      negative
      negatively
      neither
      never
      nevertheless
      new
      next
      nine
      no
      nobody
      non
      none
      noone
      nor
      normal
      normality
      normally
      not
      note
      noted
      notes
      nothing
      noting
      novel
      now
      nowhere
      number
      numbered
      numbering
      numbers
      numerical
      numerically
      o
      objective
      objectives
      observable
      observation
      observations
      observe
      observed
      observes
      observing
      obtain
      obtained
      obtaining
      obtains
      obviously
      of
      off
      often
      oh
      ok
      okay
      old
      on
      once
      one
      ones
      only
      onto
      or
      organ
      organism
      organismal
      organisms
      organs
      other
      others
      otherwise
      ought
      our
      ours
      ourselves
      out
      outside
      over
      overall
      own
      p
      page
      pages
      paper
      papers
      participant
      participants
      particular
      particularly
      pathological
      pathologies
      pathology
      pathway
      pathways
      patient
      patients
      pdf
      pdfs
      per
      percent
      percentage
      percentages
      perform
      performance
      performed
      performing
      performs
      perhaps
      period
      periodic
      periodically
      periods
      phase
      phases
      placed
      please
      plus
      population
      populations
      positive
      positively
      possible
      possibly
      potential
      potentially
      prediction
      presence
      present
      presented
      presenting
      presumably
      previous
      primarily
      primary
      probable
      probably
      procedural
      procedure
      procedures
      process
      processed
      processes
      processing
      proportion
      proportional
      proportionally
      proportions
      protein
      proteins
      protocol
      protocols
      provide
      provided
      provides
      providing
      publication
      publications
      publish
      published
      publisher
      publishers
      publishes
      publishing
      purpose
      purposes
      q
      qualitative
      qualitatively
      qualities
      quality
      quantification
      quantified
      quantify
      quantifying
      quantitative
      quantitatively
      que
      quite
      qv
      r
      range
      ranged
      ranges
      ranging
      rare
      rarely
      rate
      rated
      rates
      rather
      rating
      rd
      re
      really
      reasonably
      recent
      record
      recorded
      recording
      records
      reduce
      reduced
      reduces
      reducing
      reduction
      reductions
      reference
      referenced
      references
      referencing
      regarding
      regardless
      regards
      relatively
      report
      reported
      reporting
      reports
      require
      required
      requirement
      requirements
      requires
      requiring
      research
      researched
      researcher
      researchers
      researching
      respectively
      respond
      responded
      responding
      response
      responses
      responsive
      responsiveness
      result
      resulted
      resulting
      results
      retrieval
      retrieve
      retrieved
      retrieves
      retrieving
      reveal
      revealed
      revealing
      reveals
      review
      reviewed
      reviewer
      reviewers
      reviewing
      reviews
      revise
      revised
      revising
      revision
      revisions
      right
      rna
      role
      roles
      s
      said
      same
      sample
      sampled
      samples
      sampling
      saw
      say
      saying
      says
      score
      scored
      scores
      scoring
      screen
      screened
      screening
      screens
      second
      secondarily
      secondary
      secondly
      section
      sections
      see
      seeing
      seem
      seemed
      seeming
      seems
      seen
      select
      selected
      selecting
      selection
      selects
      self
      selves
      sensible
      sent
      sequence
      sequenced
      sequences
      sequencing
      serious
      seriously
      seven
      several
      shall
      she
      should
      shouldn't
      show
      showed
      showing
      shown
      shows
      significance
      significant
      significantly
      since
      six
      so
      some
      somebody
      somehow
      someone
      something
      sometime
      sometimes
      somewhat
      somewhere
      soon
      sorry
      species
      specific
      specifically
      specified
      specify
      specifying
      spectrometer
      spectrometric
      spectroscopy
      stage
      staged
      stages
      staging
      stain
      stained
      staining
      stains
      standard
      statistic
      statistical
      statistically
      statistics
      still
      studied
      studies
      study
      studying
      sub
      subject
      subjects
      submission
      submit
      submits
      submitted
      submitting
      such
      suggest
      suggested
      suggesting
      suggestion
      suggestions
      suggests
      sup
      sure
      symptom
      symptomatic
      symptoms
      syndrome
      syndromes
      t
      t's
      table
      tables
      take
      taken
      technical
      technically
      technique
      techniques
      tell
      tends
      test
      tested
      testing
      tests
      th
      than
      thank
      thanks
      thanx
      that
      that's
      thats
      the
      their
      theirs
      them
      themselves
      then
      thence
      therapeutic
      therapeutics
      therapies
      therapy
      there
      there's
      thereafter
      thereby
      therefore
      therein
      theres
      thereupon
      these
      they
      they'd
      they'll
      they're
      they've
      think
      third
      this
      thorough
      thoroughly
      those
      though
      three
      through
      throughout
      thru
      thus
      time
      timed
      times
      timing
      tissue
      tissues
      to
      together
      too
      took
      total
      totaled
      totaling
      totals
      toward
      towards
      treat
      treated
      treating
      treatment
      treatments
      trial
      trials
      tried
      tries
      truly
      try
      trying
      twice
      two
      type
      typed
      types
      typical
      typically
      typing
      u
      un
      uncommon
      under
      unfortunately
      unless
      unlikely
      until
      unto
      up
      upload
      uploaded
      uploading
      uploads
      upon
      us
      use
      used
      useful
      uses
      using
      usually
      uucp
      v
      valuation
      value
      valued
      values
      valuing
      variance
      variant
      variants
      variation
      variations
      various
      version
      versions
      versus
      very
      via
      viz
      vs
      w
      want
      wants
      was
      wasn't
      way
      we
      we'd
      we'll
      we're
      we've
      welcome
      well
      went
      were
      weren't
      what
      what's
      whatever
      when
      whence
      whenever
      where
      where's
      whereafter
      whereas
      whereby
      wherein
      whereupon
      wherever
      whether
      which
      while
      whither
      who
      who's
      whoever
      whole
      whom
      whose
      why
      will
      willing
      wish
      with
      within
      without
      won't
      wonder
      would
      wouldn't
      write
      writes
      writing
      written
      wrote
      x
      y
      yes
      yet
      you
      you'd
      you'll
      you're
      you've
      your
      yours
      yourself
      yourselves
      z
      zero
      zeroanalysis
  • README.md 391 B
    # extracting-keywords
    
    Extract keywords from documents using YAKE algorithm with support for 34 languages (Arabic to Chinese). Use when users request keyword extraction, key terms, topic identification, content summarization, or document analysis. Includes domain-specific stopwords for AI/ML and life sciences. Optional deeper extraction mode (n=2+n=3 combined) for comprehensive coverage.
    
  • SKILL.md 11.4 KB
    ---
    name: extracting-keywords
    description: Extract keywords from documents using YAKE algorithm with support for 34 languages (Arabic to Chinese). Use when users request keyword extraction, key terms, topic identification, content summarization, or document analysis. Includes domain-specific stopwords for AI/ML and life sciences. Optional deeper extraction mode (n=2+n=3 combined) for comprehensive coverage.
    metadata:
      version: 0.2.1
    ---
    
    # Extracting Keywords
    
    Extract keywords from text using YAKE (Yet Another Keyword Extractor), an unsupervised statistical keyword extraction algorithm.
    
    ## Installation
    
    **First time only:** Install YAKE with optimized dependencies to avoid unnecessary downloads.
    
    ```bash
    cd /home/claude
    uv venv yake-venv --system-site-packages
    uv pip install yake --python yake-venv/bin/python --no-deps
    uv pip install jellyfish segtok regex --python yake-venv/bin/python
    ```
    
    This reuses system packages (numpy, networkx) instead of downloading them (~0.08s vs ~5s).
    
    ## Stopwords Configuration
    
    **Built-in YAKE stopwords (34 languages):** Use `lan="<code>"` parameter
    - See Parameters section below for all 34 supported language codes
    - English (`lan="en"`) is the default
    
    **Custom domain stopwords (bundled in `assets/`):**
    
    **AI/ML:** `stopwords_ai.txt`
    - English stopwords + 783 AI/ML domain-specific terms (1357 total)
    - Filters AI/ML methodology noise (model, training, network, algorithm, parameter)
    - Filters ML boilerplate (dataset, baseline, benchmark, experiment, evaluation)
    - Filters technical terms (transformer, embedding, attention, optimization, inference)
    - Includes full lemmatization (train/trains/trained/training/trainer)
    - Use for AI/ML papers, technical reports, machine learning literature
    - **Performance impact:** +4-5% runtime vs English stopwords
    
    **Life Sciences:** `stopwords_ls.txt`
    - English stopwords + 719 life sciences domain-specific terms (1293 total)
    - Filters research methodology noise (study, results, analysis, significant, observed)
    - Filters academic boilerplate (paper, manuscript, publication, review, editing)
    - Filters statistical terms (correlation, distribution, deviation, variance)
    - Filters clinical terms (patient, treatment, diagnosis, symptom, therapy)
    - Filters biology/medicine (cell, tissue, protein, gene, organism)
    - Includes full lemmatization (analyze/analyzes/analyzed/analyzing/analysis)
    - Use for biomedical papers, clinical studies, research articles, scientific literature
    - **Performance impact:** +4-5% runtime vs English stopwords
    
    ## Basic Usage
    
    ```python
    import yake
    
    # Read text
    with open('document.txt', 'r') as f:
        text = f.read()
    
    # Extract with English stopwords (default)
    kw_extractor = yake.KeywordExtractor(
        lan="en",           # Language code
        n=3,                # Max n-gram size (1-3 word phrases)
        dedupLim=0.9,       # Deduplication threshold (0-1)
        top=20              # Number of keywords to return
    )
    
    keywords = kw_extractor.extract_keywords(text)
    
    # Display results (lower score = more important)
    for kw, score in keywords:
        print(f"{score:.4f}  {kw}")
    ```
    
    ## Domain-Specific Extraction
    
    ### Using Life Sciences Stopwords
    
    **Option 1: Install custom stopwords file**
    
    ```bash
    # Copy life sciences stopwords to YAKE package
    cp assets/stopwords_ls.txt /home/claude/yake-venv/lib/python3.12/site-packages/yake/core/StopwordsList/stopwords_ls.txt
    
    # Use with lan="ls"
    kw_extractor = yake.KeywordExtractor(lan="ls", n=3, top=20)
    ```
    
    **Option 2: Load custom stopwords directly**
    
    ```python
    # Load stopwords from file
    with open('assets/stopwords_ls.txt', 'r') as f:
        custom_stops = set(line.strip().lower() for line in f)
    
    # Pass to extractor
    kw_extractor = yake.KeywordExtractor(
        stopwords=custom_stops,
        n=3,
        top=20
    )
    ```
    
    ### Using AI/ML Stopwords
    
    ```python
    # Load AI/ML stopwords
    with open('/mnt/skills/user/extracting-keywords/assets/stopwords_ai.txt', 'r') as f:
        ai_stops = set(line.strip().lower() for line in f)
    
    # Extract with AI stopwords
    kw_extractor = yake.KeywordExtractor(
        stopwords=ai_stops,
        n=3,
        top=20
    )
    keywords = kw_extractor.extract_keywords(text)
    ```
    
    ## Deeper Extraction (n=2 + n=3 Combined)
    
    For more comprehensive extraction, run both n=2 and n=3 and consolidate results. This captures both focused phrases and broader context with ~100% time overhead (still <2s for large documents).
    
    ```python
    import yake
    
    # Load domain stopwords
    with open('/mnt/skills/user/extracting-keywords/assets/stopwords_ai.txt', 'r') as f:
        stops = set(line.strip().lower() for line in f)
    
    # Extract with n=2 (captures focused phrases)
    kw_n2 = yake.KeywordExtractor(stopwords=stops, n=2, dedupLim=0.9, top=50)
    results_n2 = kw_n2.extract_keywords(text)
    
    # Extract with n=3 (captures broader context)
    kw_n3 = yake.KeywordExtractor(stopwords=stops, n=3, dedupLim=0.9, top=50)
    results_n3 = kw_n3.extract_keywords(text)
    
    # Consolidate: union with score averaging for overlaps
    combined = {}
    for kw, score in results_n2:
        combined[kw] = score
    for kw, score in results_n3:
        if kw in combined:
            combined[kw] = (combined[kw] + score) / 2
        else:
            combined[kw] = score
    
    # Sort by score (lower = more important)
    consolidated = sorted(combined.items(), key=lambda x: x[1])
    
    # Display top 30
    for kw, score in consolidated[:30]:
        print(f"{score:.4f}  {kw}")
    ```
    
    **Benefits:**
    - n=2 extracts cleaner domain-specific phrases ("disk move", "error rate")
    - n=3 captures contextual combinations ("Move disk 1", "per-step error rate")
    - Consolidation provides richer keyword set for topic modeling or search indexing
    
    **Performance:**
    - Combined approach: ~2x runtime of single extraction
    - Typical timing: 0.4s (small doc) to 1.0s (large doc)
    - Use when quality matters more than speed
    
    ## Parameters
    
    **lan** (str): Language code for built-in stopwords
    - `"en"` - English (default)
    - `"ai"` - AI/ML (if stopwords_ai.txt installed in YAKE)
    - `"ls"` - Life sciences (if stopwords_ls.txt installed in YAKE)
    
    **Built-in YAKE languages (34 total):**
    - `"ar"` - Arabic
    - `"bg"` - Bulgarian  
    - `"br"` - Breton
    - `"cz"` - Czech
    - `"da"` - Danish
    - `"de"` - German
    - `"el"` - Greek
    - `"es"` - Spanish
    - `"et"` - Estonian
    - `"fa"` - Farsi/Persian
    - `"fi"` - Finnish
    - `"fr"` - French
    - `"hi"` - Hindi
    - `"hr"` - Croatian
    - `"hu"` - Hungarian
    - `"hy"` - Armenian
    - `"id"` - Indonesian
    - `"it"` - Italian
    - `"ja"` - Japanese
    - `"lt"` - Lithuanian
    - `"lv"` - Latvian
    - `"nl"` - Dutch
    - `"no"` - Norwegian
    - `"pl"` - Polish
    - `"pt"` - Portuguese
    - `"ro"` - Romanian
    - `"ru"` - Russian
    - `"sk"` - Slovak
    - `"sl"` - Slovenian
    - `"sv"` - Swedish
    - `"tr"` - Turkish
    - `"uk"` - Ukrainian
    - `"zh"` - Chinese
    
    **n** (int): Maximum n-gram size (default: 3)
    - `1` - Single words only
    - `2` - Up to 2-word phrases
    - `3` - Up to 3-word phrases (recommended)
    - `4-5` - May produce suboptimal results with YAKE's algorithm
    
    **dedupLim** (float): Deduplication threshold (default: 0.9)
    - Range: 0.0 to 1.0
    - Higher values = more aggressive deduplication
    - Controls handling of similar terms (e.g., "cancer cell" vs "cancer cells")
    
    **top** (int): Number of keywords to return (default: 20)
    
    **stopwords** (set): Custom stopwords set (overrides lan parameter)
    
    ## Workflow Patterns
    
    ### Single Document Analysis
    
    ```python
    import yake
    
    # Read document
    with open('/mnt/user-data/uploads/article.txt', 'r') as f:
        text = f.read()
    
    # Extract keywords
    kw_extractor = yake.KeywordExtractor(lan="en", n=3, top=30)
    keywords = kw_extractor.extract_keywords(text)
    
    # Format results
    results = []
    for kw, score in keywords:
        results.append(f"{score:.4f}  {kw}")
    
    print("\n".join(results))
    ```
    
    ### Comparing Stopwords Strategies
    
    ```python
    import yake
    
    # Load life sciences stopwords
    with open('assets/stopwords_ls.txt', 'r') as f:
        ls_stops = set(line.strip().lower() for line in f)
    
    # Extract with English stopwords
    kw_en = yake.KeywordExtractor(lan="en", n=3, top=20)
    keywords_en = kw_en.extract_keywords(text)
    
    # Extract with life sciences stopwords
    kw_ls = yake.KeywordExtractor(stopwords=ls_stops, n=3, top=20)
    keywords_ls = kw_ls.extract_keywords(text)
    
    # Compare results
    print("English stopwords:")
    for kw, score in keywords_en:
        print(f"  {score:.4f}  {kw}")
    
    print("\nLife sciences stopwords:")
    for kw, score in keywords_ls:
        print(f"  {score:.4f}  {kw}")
    ```
    
    ### Batch Processing
    
    ```python
    import yake
    import os
    
    # Initialize extractor
    kw_extractor = yake.KeywordExtractor(lan="en", n=3, top=15)
    
    # Process multiple files
    results = {}
    for filename in os.listdir('/mnt/user-data/uploads'):
        if filename.endswith('.txt'):
            with open(f'/mnt/user-data/uploads/{filename}', 'r') as f:
                text = f.read()
            
            keywords = kw_extractor.extract_keywords(text)
            results[filename] = keywords
    
    # Output results
    for filename, keywords in results.items():
        print(f"\n{filename}:")
        for kw, score in keywords[:10]:  # Top 10
            print(f"  {score:.4f}  {kw}")
    ```
    
    ### Multilingual Extraction
    
    ```python
    import yake
    
    # French document
    with open('/mnt/user-data/uploads/article_fr.txt', 'r') as f:
        french_text = f.read()
    
    # Extract with French stopwords
    kw_fr = yake.KeywordExtractor(lan="fr", n=3, top=20)
    keywords_fr = kw_fr.extract_keywords(french_text)
    
    print("Mots-clés (French):")
    for kw, score in keywords_fr:
        print(f"  {score:.4f}  {kw}")
    
    # German document
    with open('/mnt/user-data/uploads/artikel_de.txt', 'r') as f:
        german_text = f.read()
    
    # Extract with German stopwords
    kw_de = yake.KeywordExtractor(lan="de", n=3, top=20)
    keywords_de = kw_de.extract_keywords(german_text)
    
    print("\nSchlüsselwörter (German):")
    for kw, score in keywords_de:
        print(f"  {score:.4f}  {kw}")
    ```
    
    ## Output Formats
    
    ### Plain Text
    ```python
    for kw, score in keywords:
        print(f"{kw}: {score:.4f}")
    ```
    
    ### CSV
    ```python
    import csv
    
    with open('/mnt/user-data/outputs/keywords.csv', 'w', newline='') as f:
        writer = csv.writer(f)
        writer.writerow(['Keyword', 'Score'])
        writer.writerows(keywords)
    ```
    
    ### JSON
    ```python
    import json
    
    output = [{"keyword": kw, "score": score} for kw, score in keywords]
    with open('/mnt/user-data/outputs/keywords.json', 'w') as f:
        json.dump(output, f, indent=2)
    ```
    
    ## Notes
    
    - Lower scores indicate more important keywords
    - YAKE is unsupervised - no training data required
    - **Supports 34 languages** - built-in stopwords for Arabic, Bulgarian, Chinese, Czech, Danish, Dutch, English, Estonian, Farsi, Finnish, French, German, Greek, Hindi, Croatian, Hungarian, Armenian, Indonesian, Italian, Japanese, Lithuanian, Latvian, Norwegian, Polish, Portuguese, Romanian, Russian, Slovak, Slovenian, Spanish, Swedish, Turkish, Ukrainian, and more
    - Optimal n-gram size is 2 or 3 for most use cases
    - For longer technical phrases (4+ words), consider post-processing or ontology matching
    - Always specify full venv path: `/home/claude/yake-venv/bin/python`
    
    ## Troubleshooting
    
    **Import errors:** Verify venv installation
    ```bash
    /home/claude/yake-venv/bin/python -c "import yake; print(yake.__version__)"
    ```
    
    **Empty results:** Check text length (YAKE needs sufficient content, typically 100+ words)
    
    **Poor quality keywords:** Adjust parameters:
    - Increase `dedupLim` for more aggressive deduplication
    - Try domain-specific stopwords
    - Increase `top` to see more candidates
    
    **Generic terms appearing:** Add custom stopwords for your domain:
    ```python
    with open('assets/stopwords_ls.txt', 'r') as f:
        stops = set(line.strip().lower() for line in f)
    
    # Add domain-specific terms
    stops.update(['term1', 'term2', 'term3'])
    
    kw_extractor = yake.KeywordExtractor(stopwords=stops, n=3, top=20)
    ```
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related