{"slug":"bioinformatics-installer","title":"bioinformatics-installer","summary":"Install bioinformatics tools for ENCODE data analysis. Covers CLI tools (BWA, STAR, samtools, MACS2), R/Bioconductor packages (DESeq2, Seurat, ChIPseeker), Python packages (Scanpy, deeptools), and Nextflow pipeline infrastructure. Generates conda environments, R install scripts, ","platform":"Claude","tags":[],"authorName":"LLM Mart","authorSlug":"llm-mart","score":0,"source":"github","price":null,"verified":false,"createdAt":"2026-09-22T13:25:31.676809Z","repo":{"url":"https://github.com/ammawla/encode-toolkit","stars":21,"forks":5,"license":"AGPL-3.0","updatedAt":"2026-09-27T01:22:42Z"},"bodyHtml":"<hr>\n<h2>name: bioinformatics-installer\ndescription: \"Install bioinformatics tools for ENCODE data analysis. Covers CLI tools (BWA, STAR, samtools, MACS2), R/Bioconductor packages (DESeq2, Seurat, ChIPseeker), Python packages (Scanpy, deeptools), and Nextflow pipeline infrastructure. Generates conda environments, R install scripts, and Python requirements. Use when the user needs to set up a bioinformatics workstation, install tools for a specific assay, create reproducible environments, or troubleshoot dependency issues. Trigger on: install tools, set up environment, conda create, bioinformatics setup, install R packages, install Bioconductor, install pipeline tools.\"</h2>\n<h1>Bioinformatics Installer for ENCODE Data Analysis</h1>\n<p>Install all bioinformatics tools needed for ENCODE data analysis, organized by assay type.\nThis skill provides ready-to-use conda environment definitions, R/Bioconductor install scripts,\nPython package lists, and Nextflow pipeline infrastructure setup. Every primary tool is\nversion-pinned for reproducibility; a few utility packages (<code>bedops</code>, <code>ucsc-bedgraphtobigwig</code>,\n<code>pigz</code>, <code>openjdk</code>, <code>r-base</code>) float so the solver can satisfy the pinned tools around them.</p>\n<h2>When to Use</h2>\n<ul>\n<li>User wants to install bioinformatics tools needed for ENCODE data analysis</li>\n<li>User asks about \"install tools\", \"conda environment\", \"setup bioinformatics\", or \"install HOMER/MACS2/deeptools\"</li>\n<li>User needs pre-configured conda environments for specific assay pipelines (ChIP-seq, ATAC-seq, RNA-seq, etc.)</li>\n<li>User wants to install R/Bioconductor packages (DESeq2, Seurat, ChIPseeker) or Python packages (Scanpy, pysam)</li>\n<li>Example queries: \"install tools for ChIP-seq analysis\", \"set up a conda environment for ATAC-seq\", \"install deeptools and bedtools\"</li>\n</ul>\n<h2>Overview</h2>\n<p>ENCODE data analysis requires a broad ecosystem of tools spanning command-line aligners, peak\ncallers, signal processors, statistical analysis frameworks in R, Python visualization and\nsingle-cell packages, and workflow engines. Setting up these tools correctly — with compatible\nversions, proper channel priorities, and no dependency conflicts — is a significant barrier for\nnew users and a reproducibility concern for experienced analysts.</p>\n<p>This skill solves that by providing:</p>\n<ul>\n<li><strong>7 assay-specific conda environments</strong> with pinned tool versions matching ENCODE pipeline standards</li>\n<li><strong>R/Bioconductor install script</strong> covering 47 packages across 8 categories</li>\n<li><strong>Python install script</strong> for single-cell, Hi-C, and genomics packages, locked by <code>scripts/constraints.txt</code></li>\n<li><strong>Nextflow install script + container checks</strong> for pipeline execution on local, HPC, and cloud platforms</li>\n</ul>\n<p>The environment files, <code>scripts/requirements.in</code> and <code>scripts/install-r-packages.R</code> are the\nauthoritative package lists; the tables below summarise them.</p>\n<p>All environments use the same channel priority (conda-forge &gt; bioconda). Every file is dry-run\nsolved for Linux x86_64 in CI, so the pinned versions exist and install together. Several tools\nhave no macOS arm64 build on bioconda; on Apple Silicon use the pipeline Docker images instead.</p>\n<p>For every tool that an environment and the matching <code>pipeline-*</code> Docker image both install, the\ntwo pin the same version, and CI fails if they drift. Some tools exist on only one side — for\nexample <code>phantompeakqualtools</code>, <code>salmon</code>, <code>subread</code> and the Hi-C <code>bedtools</code> are conda-only,\nwhile juicer_tools, SEACR, Hotspot2 and <code>modwt</code> are image-only because they are not conda\npackages. Those are noted in the sections below.</p>\n<h2>Quick Start</h2>\n<p>Install a complete environment for any assay type with a single command:</p>\n<pre><code># ChIP-seq (histone or TF)\nconda env create -f skills/bioinformatics-installer/environments/chipseq-env.yml\n\n# ATAC-seq\nconda env create -f skills/bioinformatics-installer/environments/atacseq-env.yml\n\n# RNA-seq\nconda env create -f skills/bioinformatics-installer/environments/rnaseq-env.yml\n\n# Hi-C\nconda env create -f skills/bioinformatics-installer/environments/hic-env.yml\n\n# Whole-Genome Bisulfite Sequencing (WGBS)\nconda env create -f skills/bioinformatics-installer/environments/wgbs-env.yml\n\n# DNase-seq\nconda env create -f skills/bioinformatics-installer/environments/dnaseseq-env.yml\n\n# CUT&amp;RUN / CUT&amp;Tag\nconda env create -f skills/bioinformatics-installer/environments/cutandrun-env.yml\n</code></pre>\n<p>Using mamba for faster solves (recommended):</p>\n<pre><code>mamba env create -f skills/bioinformatics-installer/environments/chipseq-env.yml\n</code></pre>\n<p>Install R and Python packages:</p>\n<pre><code># All R/Bioconductor packages\nRscript skills/bioinformatics-installer/scripts/install-r-packages.R --all\n\n# All Python packages\nbash skills/bioinformatics-installer/scripts/install-python-packages.sh --all\n\n# Install the pinned Nextflow release and check for a Docker runtime\nbash skills/bioinformatics-installer/scripts/install-nextflow.sh --docker\n</code></pre>\n<h2>Per-Assay Environments</h2>\n<h3>ChIP-seq Environment (<code>encode-chipseq</code>)</h3>\n<p>For histone modification and transcription factor ChIP-seq processing following ENCODE\nuniform pipeline standards (Landt et al. 2012, ENCODE Consortium 2020).</p>\n<table>\n<thead>\n<tr>\n<th>Tool</th>\n<th>Version</th>\n<th>Purpose</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>BWA-MEM</td>\n<td>0.7.18</td>\n<td>Read alignment to reference genome (Li &amp; Durbin 2009)</td>\n</tr>\n<tr>\n<td>samtools</td>\n<td>1.19</td>\n<td>BAM manipulation, sorting, indexing, flagstat (Li et al. 2009)</td>\n</tr>\n<tr>\n<td>MACS2</td>\n<td>2.2.9.1</td>\n<td>Peak calling for narrow (TF) and broad (histone) marks (Zhang et al. 2008)</td>\n</tr>\n<tr>\n<td>Picard</td>\n<td>3.1.1</td>\n<td>Duplicate marking and library complexity metrics (Broad Institute)</td>\n</tr>\n<tr>\n<td>phantompeakqualtools</td>\n<td>1.2.2</td>\n<td>Strand cross-correlation (NSC/RSC) quality metrics (Kharchenko et al. 2008)</td>\n</tr>\n<tr>\n<td>IDR</td>\n<td>2.0.4.2</td>\n<td>Irreproducible Discovery Rate for replicate consistency (Li et al. 2011)</td>\n</tr>\n<tr>\n<td>deeptools</td>\n<td>3.5.5</td>\n<td>Signal normalization (bamCoverage), fingerprint, correlation (Ramirez et al. 2016)</td>\n</tr>\n<tr>\n<td>bedtools</td>\n<td>2.31.0</td>\n<td>Interval operations, blacklist filtering (Quinlan &amp; Hall 2010)</td>\n</tr>\n<tr>\n<td>FastQC</td>\n<td>0.12.1</td>\n<td>Raw read quality assessment (Andrews 2010)</td>\n</tr>\n<tr>\n<td>Trim Galore</td>\n<td>0.6.10</td>\n<td>Adapter and quality trimming via Cutadapt (Krueger 2012)</td>\n</tr>\n<tr>\n<td>MultiQC</td>\n<td>1.21</td>\n<td>Aggregate QC report across all pipeline stages (Ewels et al. 2016)</td>\n</tr>\n<tr>\n<td>bedGraphToBigWig</td>\n<td>—</td>\n<td>Convert bedGraph signal to bigWig for genome browser viewing (Kent et al. 2010)</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Memory</strong>: BWA index for GRCh38 requires ~5.5 GB RAM. Peak calling with MACS2 typically requires\n4-8 GB. phantompeakqualtools loads full BAM into memory.</p>\n<p><strong>Environment file</strong>: <code>environments/chipseq-env.yml</code></p>\n<hr>\n<h3>ATAC-seq Environment (<code>encode-atacseq</code>)</h3>\n<p>For chromatin accessibility profiling via ATAC-seq following ENCODE standards\n(Buenrostro et al. 2013, Corces et al. 2017).</p>\n<table>\n<thead>\n<tr>\n<th>Tool</th>\n<th>Version</th>\n<th>Purpose</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Bowtie2</td>\n<td>2.5.4</td>\n<td>Alignment (preferred over BWA for ATAC-seq short fragments) (Langmead &amp; Salzberg 2012)</td>\n</tr>\n<tr>\n<td>MACS2</td>\n<td>2.2.9.1</td>\n<td>Peak calling (<code>pipeline-atacseq</code> calls it with <code>-f BAMPE</code> on Tn5-shifted reads) (Zhang et al. 2008)</td>\n</tr>\n<tr>\n<td>IDR</td>\n<td>2.0.4.2</td>\n<td>Irreproducible Discovery Rate for replicate consistency (Li et al. 2011)</td>\n</tr>\n<tr>\n<td>samtools</td>\n<td>1.19</td>\n<td>BAM manipulation, mitochondrial read filtering</td>\n</tr>\n<tr>\n<td>Picard</td>\n<td>3.1.1</td>\n<td>Duplicate marking, insert size metrics</td>\n</tr>\n<tr>\n<td>deeptools</td>\n<td>3.5.5</td>\n<td>alignmentSieve (Tn5 offset), bamCoverage (signal tracks), plotFingerprint</td>\n</tr>\n<tr>\n<td>bedtools</td>\n<td>2.31.0</td>\n<td>Blacklist filtering, interval operations</td>\n</tr>\n<tr>\n<td>FastQC</td>\n<td>0.12.1</td>\n<td>Raw read quality and adapter content assessment</td>\n</tr>\n<tr>\n<td>Trim Galore</td>\n<td>0.6.10</td>\n<td>Adapter trimming (Nextera adapters for ATAC-seq)</td>\n</tr>\n<tr>\n<td>MultiQC</td>\n<td>1.21</td>\n<td>Aggregate QC reporting</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Key ATAC-seq parameters</strong>: Tn5 transposase introduces a +4/-5 bp offset that must be corrected.\nFragment size distribution should show nucleosomal ladder (sub-nucleosomal, mono-, di-, tri-).\nTSS enrichment score should be &gt;= 5 (GRCh38), &gt;= 6 (hg19), or &gt;= 10 (mm10) for high-quality data (ENCODE data standards).</p>\n<p><strong>Environment file</strong>: <code>environments/atacseq-env.yml</code></p>\n<hr>\n<h3>RNA-seq Environment (<code>encode-rnaseq</code>)</h3>\n<p>For gene expression quantification following ENCODE RNA-seq standards\n(Conesa et al. 2016, ENCODE Consortium 2020).</p>\n<table>\n<thead>\n<tr>\n<th>Tool</th>\n<th>Version</th>\n<th>Purpose</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>STAR</td>\n<td>2.7.11b</td>\n<td>Splice-aware alignment with 2-pass mapping (Dobin et al. 2013)</td>\n</tr>\n<tr>\n<td>RSEM</td>\n<td>1.3.3</td>\n<td>Gene/transcript quantification with expectation-maximization (Li &amp; Dewey 2011)</td>\n</tr>\n<tr>\n<td>Kallisto</td>\n<td>0.50.1</td>\n<td>Pseudoalignment-based transcript quantification (Bray et al. 2016)</td>\n</tr>\n<tr>\n<td>Salmon</td>\n<td>1.10.3</td>\n<td>Quasi-mapping transcript quantification with GC bias correction (Patro et al. 2017)</td>\n</tr>\n<tr>\n<td>featureCounts (subread)</td>\n<td>2.0.6</td>\n<td>Gene-level read counting for count-based DE methods (Liao et al. 2014)</td>\n</tr>\n<tr>\n<td>samtools</td>\n<td>1.19</td>\n<td>BAM handling, flagstat, idxstats</td>\n</tr>\n<tr>\n<td>FastQC</td>\n<td>0.12.1</td>\n<td>Read quality assessment</td>\n</tr>\n<tr>\n<td>Trim Galore</td>\n<td>0.6.10</td>\n<td>Adapter and quality trimming</td>\n</tr>\n<tr>\n<td>MultiQC</td>\n<td>1.21</td>\n<td>Aggregate QC report</td>\n</tr>\n<tr>\n<td>RSeQC</td>\n<td>5.0.3</td>\n<td>RNA-seq-specific QC: gene body coverage, read distribution, inner distance (Wang et al. 2012)</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Memory</strong>: STAR genome generation requires 32+ GB RAM for human genome. STAR alignment requires\n~30 GB RAM. Kallisto and Salmon are memory-efficient alternatives (~4 GB).</p>\n<p><strong>Environment file</strong>: <code>environments/rnaseq-env.yml</code></p>\n<hr>\n<h3>Hi-C Environment (<code>encode-hic</code>)</h3>\n<p>For chromatin conformation capture processing following ENCODE Hi-C standards\n(Yardimci et al. 2019, Rao et al. 2014).</p>\n<table>\n<thead>\n<tr>\n<th>Tool</th>\n<th>Version</th>\n<th>Purpose</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>BWA-MEM</td>\n<td>0.7.18</td>\n<td>Chimeric read alignment (each mate aligned independently)</td>\n</tr>\n<tr>\n<td>pairtools</td>\n<td>1.1.2</td>\n<td>Parse, sort, deduplicate, filter contact pairs (Open2C)</td>\n</tr>\n<tr>\n<td>cooler</td>\n<td>0.9.3</td>\n<td>Multi-resolution contact matrix storage and balancing (Abdennur &amp; Mirny 2020)</td>\n</tr>\n<tr>\n<td>openjdk</td>\n<td>&gt;=11</td>\n<td>Java runtime for Juicer Tools (the jar itself is installed separately, see below)</td>\n</tr>\n<tr>\n<td>samtools</td>\n<td>1.19</td>\n<td>BAM handling for chimeric alignment parsing</td>\n</tr>\n<tr>\n<td>bedtools</td>\n<td>2.31.0</td>\n<td>Restriction fragment and TAD boundary operations</td>\n</tr>\n<tr>\n<td>FastQC</td>\n<td>0.12.1</td>\n<td>Read quality assessment</td>\n</tr>\n<tr>\n<td>Trim Galore</td>\n<td>0.6.10</td>\n<td>Adapter trimming</td>\n</tr>\n<tr>\n<td>MultiQC</td>\n<td>1.21</td>\n<td>Aggregate QC reporting</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Key Hi-C parameters</strong>: Cis/trans ratio &gt; 60%, long-range cis contacts (&gt; 20 kb) &gt; 40%.\nResolution depends on sequencing depth: ~1 billion valid pairs for 5 kb resolution on human.</p>\n<p><strong>Juicer Tools is not in this environment.</strong> The YAML installs only the Java runtime it needs.\nDownload <code>juicer_tools.2.20.00.jar</code> from the <code>aidenlab/Juicebox</code> GitHub releases and invoke it\nwith <code>java -jar</code>. The Hi-C pipeline image (<code>pipeline-hic/scripts/Dockerfile</code>) already contains it.</p>\n<p>The environment also installs <code>cooltools</code>, <code>hic-straw</code> and <code>pyGenomeTracks</code> from PyPI (unpinned).</p>\n<p><strong>Environment file</strong>: <code>environments/hic-env.yml</code></p>\n<hr>\n<h3>WGBS Environment (<code>encode-wgbs</code>)</h3>\n<p>For whole-genome bisulfite sequencing (DNA methylation) following ENCODE standards\n(Foox et al. 2021, Schultz et al. 2015).</p>\n<table>\n<thead>\n<tr>\n<th>Tool</th>\n<th>Version</th>\n<th>Purpose</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Bismark</td>\n<td>0.24.2</td>\n<td>Bisulfite-aware alignment and methylation extraction (Krueger &amp; Andrews 2011)</td>\n</tr>\n<tr>\n<td>MethylDackel</td>\n<td>0.6.1</td>\n<td>Fast methylation extraction from bisulfite BAMs (Ryan 2023)</td>\n</tr>\n<tr>\n<td>samtools</td>\n<td>1.19</td>\n<td>BAM manipulation, merge, index</td>\n</tr>\n<tr>\n<td>bedtools</td>\n<td>2.31.0</td>\n<td>Interval operations for DMR analysis</td>\n</tr>\n<tr>\n<td>FastQC</td>\n<td>0.12.1</td>\n<td>Read quality assessment (note: bisulfite libraries have biased base composition)</td>\n</tr>\n<tr>\n<td>Trim Galore</td>\n<td>0.6.10</td>\n<td>Adapter trimming with --rrbs or default mode</td>\n</tr>\n<tr>\n<td>MultiQC</td>\n<td>1.21</td>\n<td>Aggregate QC reporting with Bismark module</td>\n</tr>\n<tr>\n<td>htslib</td>\n<td>1.19</td>\n<td>Provides <code>tabix</code> and <code>bgzip</code> for indexed, block-gzipped methylation BED files</td>\n</tr>\n<tr>\n<td>Bowtie2</td>\n<td>2.5.4</td>\n<td>Backend aligner required by Bismark</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Key WGBS parameters</strong>: Bisulfite conversion rate ≥ 98% (check unmethylated spike-in lambda DNA).\nCpG coverage &gt;= 10x for reliable DMR calling. M-bias plots should be checked for end-repair artifacts.</p>\n<p><strong>Environment file</strong>: <code>environments/wgbs-env.yml</code></p>\n<hr>\n<h3>DNase-seq Environment (<code>encode-dnaseseq</code>)</h3>\n<p>For DNase I hypersensitive site mapping following ENCODE standards\n(Thurman et al. 2012, ENCODE Consortium 2020).</p>\n<table>\n<thead>\n<tr>\n<th>Tool</th>\n<th>Version</th>\n<th>Purpose</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>BWA-MEM</td>\n<td>0.7.18</td>\n<td>Read alignment to reference genome</td>\n</tr>\n<tr>\n<td>Picard</td>\n<td>3.1.1</td>\n<td>Duplicate marking and library complexity metrics</td>\n</tr>\n<tr>\n<td>BEDOPS</td>\n<td>unpinned</td>\n<td><code>sort-bed</code> and <code>unstarch</code> for the Hotspot2 <code>.starch</code> archives (Neph et al. 2012)</td>\n</tr>\n<tr>\n<td>HINT (RGT)</td>\n<td>1.0.2</td>\n<td>TF footprinting from DNase-seq data (Li et al. 2019); installed from PyPI</td>\n</tr>\n<tr>\n<td>F-Seq2</td>\n<td>2.0.3</td>\n<td>Feature density estimation for peak calling (Boyle et al. 2008, Zhao et al. 2020); installed from PyPI</td>\n</tr>\n<tr>\n<td>samtools</td>\n<td>1.19</td>\n<td>BAM handling and filtering</td>\n</tr>\n<tr>\n<td>bedtools</td>\n<td>2.31.0</td>\n<td>Interval operations, blacklist filtering</td>\n</tr>\n<tr>\n<td>FastQC</td>\n<td>0.12.1</td>\n<td>Read quality assessment</td>\n</tr>\n<tr>\n<td>Trim Galore</td>\n<td>0.6.10</td>\n<td>Adapter trimming</td>\n</tr>\n<tr>\n<td>MultiQC</td>\n<td>1.21</td>\n<td>Aggregate QC reporting</td>\n</tr>\n<tr>\n<td>bedGraphToBigWig</td>\n<td>unpinned</td>\n<td>Convert bedGraph signal to bigWig</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Hotspot2 2.1.2 and its <code>modwt</code> dependency are not in this environment</strong> — neither is packaged\nfor conda. Build both from source (<code>pipeline-dnaseseq/scripts/Dockerfile</code> shows the exact steps)\nor run the pipeline through that image, which is what <code>pipeline-dnaseseq</code> does.</p>\n<p><strong>Environment file</strong>: <code>environments/dnaseseq-env.yml</code></p>\n<hr>\n<h3>CUT&amp;RUN / CUT&amp;Tag Environment (<code>encode-cutandrun</code>)</h3>\n<p>For antibody-targeted chromatin profiling via CUT&amp;RUN (Skene &amp; Henikoff 2017) and\nCUT&amp;Tag (Kaya-Okur et al. 2019).</p>\n<table>\n<thead>\n<tr>\n<th>Tool</th>\n<th>Version</th>\n<th>Purpose</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Bowtie2</td>\n<td>2.5.4</td>\n<td>Alignment (recommended for shorter CUT&amp;RUN/Tag fragments)</td>\n</tr>\n<tr>\n<td>r-base</td>\n<td>&gt;=4.3</td>\n<td>R runtime that the SEACR shell script calls (SEACR itself is installed separately, see below)</td>\n</tr>\n<tr>\n<td>MACS2</td>\n<td>2.2.9.1</td>\n<td>Alternative peak calling with adjusted parameters</td>\n</tr>\n<tr>\n<td>samtools</td>\n<td>1.19</td>\n<td>BAM handling, spike-in alignment filtering</td>\n</tr>\n<tr>\n<td>Picard</td>\n<td>3.1.1</td>\n<td>Duplicate marking (low duplication expected for CUT&amp;RUN/Tag)</td>\n</tr>\n<tr>\n<td>deeptools</td>\n<td>3.5.5</td>\n<td>Signal tracks, heatmaps, spike-in normalization</td>\n</tr>\n<tr>\n<td>bedtools</td>\n<td>2.31.0</td>\n<td>Interval operations, suspect list filtering</td>\n</tr>\n<tr>\n<td>FastQC</td>\n<td>0.12.1</td>\n<td>Read quality assessment</td>\n</tr>\n<tr>\n<td>Trim Galore</td>\n<td>0.6.10</td>\n<td>Adapter trimming</td>\n</tr>\n<tr>\n<td>MultiQC</td>\n<td>1.21</td>\n<td>Aggregate QC reporting</td>\n</tr>\n</tbody>\n</table>\n<p><strong>SEACR 1.3 is not in this environment.</strong> It is a shell script plus an R script; download the\n<code>v1.3</code> tarball from <code>FredHutch/SEACR</code> and put both <code>SEACR_1.3.sh</code> and <code>SEACR_1.3.R</code> on the PATH,\nalongside the <code>r-base</code> this environment installs. The CUT&amp;RUN pipeline image\n(<code>pipeline-cutandrun/scripts/Dockerfile</code>) already contains it.</p>\n<p><strong>Key CUT&amp;RUN/Tag notes</strong>: These assays have inherently lower background than ChIP-seq. Do NOT\napply ChIP-seq quality thresholds — use CUT&amp;RUN-specific metrics (Nordin et al. 2023). Apply\nthe CUT&amp;RUN suspect list instead of the standard ENCODE blacklist. Spike-in normalization\n(E. coli DNA for CUT&amp;RUN, carry-over for CUT&amp;Tag) is strongly recommended for quantitative\ncomparisons.</p>\n<p><strong>Environment file</strong>: <code>environments/cutandrun-env.yml</code></p>\n<h2>R/Bioconductor Packages</h2>\n<p>Install all R packages needed for ENCODE downstream analysis. The install script at\n<code>scripts/install-r-packages.R</code> handles BiocManager setup, version locking, and\ncategory-based installation.</p>\n<h3>Core Genomic Infrastructure</h3>\n<p>These packages provide the foundation for all genomic data manipulation in R:</p>\n<table>\n<thead>\n<tr>\n<th>Package</th>\n<th>Purpose</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>GenomicRanges</td>\n<td>Interval arithmetic on genomic coordinates (Lawrence et al. 2013)</td>\n</tr>\n<tr>\n<td>GenomicFeatures</td>\n<td>Gene model and transcript annotation handling</td>\n</tr>\n<tr>\n<td>rtracklayer</td>\n<td>Import/export BED, bigWig, GFF, narrowPeak, broadPeak</td>\n</tr>\n<tr>\n<td>IRanges</td>\n<td>Integer range operations (underlying GenomicRanges)</td>\n</tr>\n<tr>\n<td>GenomeInfoDb</td>\n<td>Chromosome naming conventions (UCSC vs Ensembl vs NCBI)</td>\n</tr>\n<tr>\n<td>BiocGenerics</td>\n<td>Common S4 generics across Bioconductor</td>\n</tr>\n<tr>\n<td>S4Vectors</td>\n<td>S4 class infrastructure for Bioconductor objects</td>\n</tr>\n<tr>\n<td>AnnotationDbi</td>\n<td>Unified interface to annotation databases</td>\n</tr>\n<tr>\n<td>biomaRt</td>\n<td>Ensembl BioMart query interface for gene annotation (Durinck et al. 2009)</td>\n</tr>\n</tbody>\n</table>\n<h3>Differential Analysis</h3>\n<table>\n<thead>\n<tr>\n<th>Package</th>\n<th>Purpose</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>DESeq2</td>\n<td>Differential gene expression with shrinkage estimators (Love et al. 2014)</td>\n</tr>\n<tr>\n<td>edgeR</td>\n<td>Differential expression using empirical Bayes (Robinson et al. 2010)</td>\n</tr>\n<tr>\n<td>limma</td>\n<td>Linear models for microarray and RNA-seq data (Ritchie et al. 2015)</td>\n</tr>\n<tr>\n<td>DiffBind</td>\n<td>Differential binding analysis for ChIP-seq/ATAC-seq peaks (Stark &amp; Brown 2011)</td>\n</tr>\n<tr>\n<td>ChIPQC</td>\n<td>ChIP-seq quality control in R (Carroll et al. 2014)</td>\n</tr>\n<tr>\n<td>chromVAR</td>\n<td>Chromatin accessibility variation across single cells (Schep et al. 2017)</td>\n</tr>\n</tbody>\n</table>\n<h3>Annotation and Pathway Analysis</h3>\n<table>\n<thead>\n<tr>\n<th>Package</th>\n<th>Purpose</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ChIPseeker</td>\n<td>Peak annotation and visualization (Yu et al. 2015)</td>\n</tr>\n<tr>\n<td>annotatr</td>\n<td>Annotate genomic regions with CpG islands, genes, enhancers (Cavalcante &amp; Sartor 2017)</td>\n</tr>\n<tr>\n<td>clusterProfiler</td>\n<td>Gene ontology and KEGG pathway enrichment (Yu et al. 2012)</td>\n</tr>\n<tr>\n<td>org.Hs.eg.db</td>\n<td>Human gene annotation database</td>\n</tr>\n<tr>\n<td>org.Mm.eg.db</td>\n<td>Mouse gene annotation database</td>\n</tr>\n<tr>\n<td>TxDb.Hsapiens.UCSC.hg38.knownGene</td>\n<td>Human transcript models (GRCh38)</td>\n</tr>\n<tr>\n<td>TxDb.Mmusculus.UCSC.mm10.knownGene</td>\n<td>Mouse transcript models (mm10)</td>\n</tr>\n</tbody>\n</table>\n<h3>Single-Cell Analysis</h3>\n<table>\n<thead>\n<tr>\n<th>Package</th>\n<th>Purpose</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Seurat</td>\n<td>Comprehensive single-cell RNA-seq analysis (Hao et al. 2021)</td>\n</tr>\n<tr>\n<td>Signac</td>\n<td>Single-cell chromatin accessibility (ATAC-seq) analysis (Stuart et al. 2021)</td>\n</tr>\n<tr>\n<td>SingleCellExperiment</td>\n<td>Core Bioconductor container for single-cell data</td>\n</tr>\n<tr>\n<td>scater</td>\n<td>Single-cell QC, normalization, visualization (McCarthy et al. 2017)</td>\n</tr>\n<tr>\n<td>scran</td>\n<td>Single-cell normalization and feature selection (Lun et al. 2016)</td>\n</tr>\n</tbody>\n</table>\n<h3>Bulk-to-Single-Cell Deconvolution</h3>\n<table>\n<thead>\n<tr>\n<th>Package</th>\n<th>Purpose</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>BisqueRNA</td>\n<td>Reference-based and marker-based deconvolution (Jew et al. 2020)</td>\n</tr>\n<tr>\n<td>DWLS</td>\n<td>Dampened Weighted Least Squares deconvolution (Tsoucas et al. 2019)</td>\n</tr>\n<tr>\n<td>BayesPrism</td>\n<td>Bayesian deconvolution with scRNA-seq reference (Chu et al. 2022). <strong>GitHub only</strong> — the script prints the <code>devtools::install_github()</code> line, it does not install it</td>\n</tr>\n<tr>\n<td>InstaPrism</td>\n<td>Fast approximation of BayesPrism for large datasets (Wang et al. 2024). <strong>GitHub only</strong>, same as BayesPrism</td>\n</tr>\n</tbody>\n</table>\n<h3>DNA Methylation Analysis</h3>\n<table>\n<thead>\n<tr>\n<th>Package</th>\n<th>Purpose</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>DMRcate</td>\n<td>Differentially methylated region detection (Peters et al. 2021)</td>\n</tr>\n<tr>\n<td>bsseq</td>\n<td>Bisulfite sequencing data handling and smoothing (Hansen et al. 2012)</td>\n</tr>\n<tr>\n<td>methylKit</td>\n<td>Methylation analysis from bisulfite sequencing (Akalin et al. 2012)</td>\n</tr>\n</tbody>\n</table>\n<h3>Visualization</h3>\n<table>\n<thead>\n<tr>\n<th>Package</th>\n<th>Purpose</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ComplexHeatmap</td>\n<td>Publication-quality heatmaps with annotations (Gu et al. 2016)</td>\n</tr>\n<tr>\n<td>EnhancedVolcano</td>\n<td>Volcano plots for differential expression (Blighe et al. 2018)</td>\n</tr>\n<tr>\n<td>Gviz</td>\n<td>Genome browser-style track visualization (Hahne &amp; Ivanek 2016)</td>\n</tr>\n<tr>\n<td>ggplot2</td>\n<td>Grammar of graphics for all custom plots (Wickham 2016)</td>\n</tr>\n</tbody>\n</table>\n<h3>Statistics and Batch Correction</h3>\n<table>\n<thead>\n<tr>\n<th>Package</th>\n<th>Purpose</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>sva (ComBat)</td>\n<td>Surrogate variable analysis and batch correction (Leek et al. 2012)</td>\n</tr>\n<tr>\n<td>WGCNA</td>\n<td>Weighted Gene Co-expression Network Analysis (Langfelder &amp; Horvath 2008)</td>\n</tr>\n<tr>\n<td>ReactomePA</td>\n<td>Reactome pathway analysis (Yu &amp; He 2016)</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Install script</strong>: <code>scripts/install-r-packages.R</code></p>\n<pre><code># Install all categories\nRscript scripts/install-r-packages.R --all\n\n# Install only specific categories\nRscript scripts/install-r-packages.R --chipseq      # DiffBind, ChIPQC, ChIPseeker\nRscript scripts/install-r-packages.R --rnaseq       # DESeq2, edgeR, limma\nRscript scripts/install-r-packages.R --singlecell   # Seurat, Signac, scater, scran\nRscript scripts/install-r-packages.R --methylation   # DMRcate, bsseq, methylKit\nRscript scripts/install-r-packages.R --deconvolution # BisqueRNA, DWLS (+ GitHub lines for BayesPrism, InstaPrism)\nRscript scripts/install-r-packages.R --visualization # ComplexHeatmap, EnhancedVolcano, Gviz, ggplot2\nRscript scripts/install-r-packages.R --stats         # sva, WGCNA, ReactomePA\n</code></pre>\n<p>The tables above list the main packages per category. <code>scripts/install-r-packages.R</code> holds the\ncomplete, authoritative lists (47 packages across 8 categories), including supporting packages\nsuch as <code>TFBSTools</code>, <code>motifmatchr</code>, <code>tximport</code>, <code>tximeta</code>, <code>celda</code>, <code>pheatmap</code>, <code>RColorBrewer</code>\nand <code>viridis</code>.</p>\n<h2>Python Packages</h2>\n<p>Install Python packages for single-cell analysis, Hi-C processing, signal visualization,\nand genomic data manipulation.</p>\n<h3>Core Single-Cell Stack</h3>\n<table>\n<thead>\n<tr>\n<th>Package</th>\n<th>Purpose</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>scanpy</td>\n<td>Single-cell RNA-seq analysis framework (Wolf et al. 2018)</td>\n</tr>\n<tr>\n<td>anndata</td>\n<td>Annotated data matrix for single-cell (Virshup et al. 2021)</td>\n</tr>\n<tr>\n<td>scvi-tools</td>\n<td>Deep generative models for single-cell (Gayoso et al. 2022)</td>\n</tr>\n<tr>\n<td>numpy</td>\n<td>Numerical computing</td>\n</tr>\n<tr>\n<td>pandas</td>\n<td>Data manipulation and tabular operations</td>\n</tr>\n<tr>\n<td>scipy</td>\n<td>Scientific computing (sparse matrices, statistics)</td>\n</tr>\n<tr>\n<td>matplotlib</td>\n<td>Plotting foundation</td>\n</tr>\n<tr>\n<td>seaborn</td>\n<td>Statistical visualization</td>\n</tr>\n</tbody>\n</table>\n<h3>Genomics and Signal Processing</h3>\n<table>\n<thead>\n<tr>\n<th>Package</th>\n<th>Purpose</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>deeptools</td>\n<td>Signal tracks, heatmaps, correlation (also CLI; Ramirez et al. 2016)</td>\n</tr>\n<tr>\n<td>pyBigWig</td>\n<td>Read/write bigWig signal files (Ryan 2023)</td>\n</tr>\n<tr>\n<td>pysam</td>\n<td>Python interface to samtools/htslib (Li et al. 2009)</td>\n</tr>\n<tr>\n<td>pybedtools</td>\n<td>Python interface to bedtools (Dale et al. 2011)</td>\n</tr>\n</tbody>\n</table>\n<h3>Hi-C Analysis</h3>\n<table>\n<thead>\n<tr>\n<th>Package</th>\n<th>Purpose</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>cooler</td>\n<td>Multi-resolution contact matrices (Abdennur &amp; Mirny 2020)</td>\n</tr>\n<tr>\n<td>cooltools</td>\n<td>Analysis toolkit for cooler data: TADs, compartments, insulation</td>\n</tr>\n<tr>\n<td>hic-straw</td>\n<td>Read .hic files from Juicer/Juicebox (Durand et al. 2016)</td>\n</tr>\n<tr>\n<td>pyGenomeTracks</td>\n<td>Genome browser visualization including Hi-C tracks</td>\n</tr>\n</tbody>\n</table>\n<h3>Single-Cell QC and Integration</h3>\n<table>\n<thead>\n<tr>\n<th>Package</th>\n<th>Purpose</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>scrublet</td>\n<td>Doublet detection for scRNA-seq (Wolock et al. 2019)</td>\n</tr>\n<tr>\n<td>harmony-pytorch</td>\n<td>Batch integration via Harmony in PyTorch (Korsunsky et al. 2019)</td>\n</tr>\n<tr>\n<td>scanorama</td>\n<td>Panoramic stitching of scRNA-seq datasets (Hie et al. 2019)</td>\n</tr>\n<tr>\n<td>bbknn</td>\n<td>Batch-balanced KNN graph construction (Polanski et al. 2020)</td>\n</tr>\n</tbody>\n</table>\n<p>CellBender (ambient RNA removal, Fleming et al. 2023) is <strong>not</strong> installed by the script — it is\nGPU-oriented and is left to a manual <code>pip install cellbender</code>, which the script prints as a note.</p>\n<p><strong>Install script</strong>: <code>scripts/install-python-packages.sh</code></p>\n<pre><code># Install all Python packages\nbash scripts/install-python-packages.sh --all\n\n# Install only specific categories\nbash scripts/install-python-packages.sh --genomics     # numpy, pandas, scipy, matplotlib, seaborn\nbash scripts/install-python-packages.sh --singlecell   # scanpy, scvi-tools, harmony-pytorch\nbash scripts/install-python-packages.sh --hic          # cooler, cooltools, hic-straw, bioframe\nbash scripts/install-python-packages.sh --deeptools    # deeptools, pyBigWig, pysam, pybedtools\n</code></pre>\n<p>Every category except <code>--genomics</code> also installs the core genomics packages first. The complete,\nauthoritative list of direct dependencies is <code>scripts/requirements.in</code>; exact versions for the\nwhole dependency tree are locked in <code>scripts/constraints.txt</code>, which every install is constrained\nby.</p>\n<h2>Nextflow and Container Setup</h2>\n<p>ENCODE pipeline execution requires Nextflow DSL2 and a container runtime (Docker or Singularity).</p>\n<h3>Nextflow Installation</h3>\n<pre><code># Install the pinned Nextflow release (requires Java 17+) and check for Docker.\n# Use --singularity for HPC, or --both.\nbash scripts/install-nextflow.sh --docker\n\n# Verify\nnextflow -version\n</code></pre>\n<p>What the script does:</p>\n<ul>\n<li>Downloads the pinned, self-contained Nextflow release the pipelines are validated against\n(the version and its SHA-256 are at the top of <code>scripts/install-nextflow.sh</code>), verifies the\nchecksum, and only then installs it to <code>/usr/local/bin</code> or <code>~/.local/bin</code>.</li>\n<li>An existing Nextflow is accepted only if it is <strong>exactly</strong> that pinned release. Any other\nversion is left untouched; the pinned release is installed next to it and the script tells you\nto put its directory first on <code>PATH</code>.</li>\n<li>Docker and Singularity are <strong>checked, not installed</strong>: the script reports what it finds and\nprints the install commands for your platform. Run those yourself.</li>\n</ul>\n<h3>Docker (recommended for local/cloud)</h3>\n<pre><code># macOS\nbrew install --cask docker\n\n# Linux (Ubuntu/Debian)\nsudo apt-get update\nsudo apt-get install -y docker-ce docker-ce-cli containerd.io\n\n# Add current user to docker group (Linux)\nsudo usermod -aG docker $USER\n</code></pre>\n<h3>Singularity (for HPC clusters)</h3>\n<pre><code># Most HPC clusters have Singularity pre-installed\n# Check with: module load singularity &amp;&amp; singularity version\n\n# If not available, install via conda:\nconda install -c conda-forge singularity\n</code></pre>\n<h3>Nextflow Configuration Profiles</h3>\n<p>The pipeline skills (pipeline-chipseq, pipeline-atacseq, etc.) include <code>nextflow.config</code> files\nwith profiles for local, SLURM, GCP, and AWS execution. Select the appropriate profile:</p>\n<pre><code># Local with Docker\nnextflow run main.nf -profile local\n\n# HPC with Singularity\nnextflow run main.nf -profile slurm\n\n# Google Cloud\nnextflow run main.nf -profile gcp\n\n# AWS Batch\nnextflow run main.nf -profile aws\n</code></pre>\n<p><strong>Install script</strong>: <code>scripts/install-nextflow.sh</code></p>\n<h2>Motif Analysis Tools</h2>\n<p>For transcription factor binding motif discovery and scanning.</p>\n<table>\n<thead>\n<tr>\n<th>Tool</th>\n<th>Version</th>\n<th>Type</th>\n<th>Purpose</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>HOMER</td>\n<td>4.11</td>\n<td>CLI</td>\n<td>De novo and known motif discovery, annotation (Heinz et al. 2010)</td>\n</tr>\n<tr>\n<td>MEME Suite</td>\n<td>5.5.5</td>\n<td>CLI</td>\n<td>MEME, DREME, STREME de novo discovery; FIMO scanning; AME enrichment (Bailey et al. 2015)</td>\n</tr>\n<tr>\n<td>FIMO</td>\n<td>5.5.5</td>\n<td>CLI (part of MEME Suite)</td>\n<td>Motif occurrence scanning across sequences</td>\n</tr>\n<tr>\n<td>TFBSTools</td>\n<td>R</td>\n<td>R/Bioconductor</td>\n<td>JASPAR motif handling, PFM/PWM conversion, motif scanning in R (Tan &amp; Lenhard 2016)</td>\n</tr>\n</tbody>\n</table>\n<h3>HOMER Installation</h3>\n<pre><code># Download and configure HOMER\nmkdir -p ~/software/homer\ncd ~/software/homer\nwget http://homer.ucsd.edu/homer/configureHomer.pl\nperl configureHomer.pl -install homer\nperl configureHomer.pl -install hg38   # Human genome\nperl configureHomer.pl -install mm10   # Mouse genome\n\n# Add to PATH\nexport PATH=$PATH:~/software/homer/bin\n</code></pre>\n<h3>MEME Suite Installation</h3>\n<pre><code># Via conda (recommended)\nconda install -c bioconda meme\n\n# Or from source\nwget https://meme-suite.org/meme/meme-software/5.5.5/meme-5.5.5.tar.gz\ntar xzf meme-5.5.5.tar.gz\ncd meme-5.5.5\n./configure --prefix=$HOME/software/meme --enable-build-libxml2 --enable-build-libxslt\nmake &amp;&amp; make install\n</code></pre>\n<h2>Walkthrough: Setting Up a Complete ENCODE Analysis Environment</h2>\n<p><strong>Goal</strong>: Install all bioinformatics tools needed to process ENCODE data, from raw FASTQ files through peak calling, annotation, and visualization, using Conda environments.\n<strong>Context</strong>: ENCODE analysis requires dozens of specialized tools. This skill automates installation with pre-configured Conda environments for each pipeline stage.</p>\n<h3>Step 1: Determine required tools by experiment type</h3>\n<pre><code>encode_get_experiment(accession=\"ENCSR000AKA\")\n</code></pre>\n<p>Expected output:</p>\n<pre><code>{\n  \"accession\": \"ENCSR000AKA\",\n  \"assay_title\": \"Histone ChIP-seq\",\n  \"target\": \"H3K27ac\"\n}\n</code></pre>\n<p><strong>Interpretation</strong>: Histone ChIP-seq requires: BWA-MEM (alignment), SAMtools (BAM processing), MACS2 (peak calling), IDR (reproducibility), bedtools (interval operations), deepTools (signal visualization).</p>\n<h3>Step 2: Install the ChIP-seq Conda environment</h3>\n<pre><code># Using the pre-configured environment YAML\nconda env create -f skills/bioinformatics-installer/environments/chipseq-env.yml\nconda activate encode-chipseq\n</code></pre>\n<p><code>environments/chipseq-env.yml</code> is the authoritative list of what that environment installs —\nread it rather than retyping the versions. It covers alignment (BWA), BAM processing (samtools,\nPicard), peak calling (MACS2), replicate consistency (IDR), cross-correlation metrics\n(phantompeakqualtools), signal processing (deeptools), interval operations (bedtools), and\nQC/trimming (FastQC, Trim Galore, MultiQC). See the ChIP-seq table above for the pinned versions.</p>\n<h3>Step 3: Install additional tools for downstream analysis</h3>\n<p>For peak annotation and motif analysis:</p>\n<pre><code>conda create -n encode-annotation -c conda-forge -c bioconda \\\n  homer bedtools bioconductor-chipseeker bioconductor-clusterprofiler bioconductor-rgreat\nconda activate encode-annotation\n# Includes: HOMER, bedtools, R/Bioconductor (ChIPseeker, clusterProfiler, rGREAT for GREAT queries)\n</code></pre>\n<h3>Step 4: Verify installation</h3>\n<pre><code># Quick verification of key tools\nbwa 2&gt;&amp;1 | head -3\nsamtools --version | head -1\nmacs2 --version\nbedtools --version\n</code></pre>\n<h3>Step 5: Download reference data for ENCODE analysis</h3>\n<pre><code>encode_download_files(file_accessions=[\"ENCFF001ABC\"], download_dir=\"/data/references\")\n</code></pre>\n<p>Reference files needed:</p>\n<ul>\n<li>GRCh38 genome FASTA</li>\n<li>ENCODE blacklist v2 (Amemiya et al. 2019)</li>\n<li>Gene annotation GTF (GENCODE v36)</li>\n</ul>\n<h3>Integration with downstream skills</h3>\n<ul>\n<li>Installed tools are used by → <strong>pipeline-chipseq</strong> through <strong>pipeline-cutandrun</strong> for processing</li>\n<li>Reference data feeds into → <strong>download-encode</strong> for FASTQ retrieval</li>\n<li>Environment setup enables → <strong>quality-assessment</strong> tool execution</li>\n<li>Installed annotation tools support → <strong>peak-annotation</strong> and <strong>motif-analysis</strong></li>\n</ul>\n<h2>Code Examples</h2>\n<h3>1. Find experiments to identify required tools</h3>\n<pre><code>encode_search_experiments(\n  assay_title=\"ATAC-seq\",\n  organ=\"pancreas\"\n)\n</code></pre>\n<p>Expected output:</p>\n<pre><code>{\n  \"results\": [\n    {\n      \"accession\": \"ENCSR799GHJ\",\n      \"assay_title\": \"ATAC-seq\",\n      \"biosample_summary\": \"pancreatic islet tissue male adult (44 years)\",\n      \"organ\": \"pancreas\",\n      \"status\": \"released\"\n    }\n  ],\n  \"total\": 8,\n  \"limit\": 25,\n  \"offset\": 0,\n  \"has_more\": false,\n  \"next_offset\": null\n}\n</code></pre>\n<p><strong>Install decision</strong>: ATAC-seq requires the <code>atacseq-env.yml</code> conda environment (Bowtie2 + MACS2 + deeptools + samtools + bedtools).</p>\n<h3>2. Get file info to understand format requirements</h3>\n<pre><code>encode_get_file_info(accession=\"ENCFF001ABC\")\n</code></pre>\n<p>Expected output:</p>\n<pre><code>{\n  \"accession\": \"ENCFF001ABC\",\n  \"file_format\": \"fastq\",\n  \"file_type\": \"fastq\",\n  \"output_type\": \"reads\",\n  \"file_size_human\": \"4.4 GB\",\n  \"experiment_assay\": \"ATAC-seq\",\n  \"biological_replicates\": [1],\n  \"status\": \"released\"\n}\n</code></pre>\n<p><strong>Install decision</strong>: raw ATAC-seq reads need Bowtie2 (not BWA), Picard for duplicate marking, and samtools for BAM processing — the <code>atacseq-env.yml</code> environment. Whether the FASTQ is one mate of a pair is on the file's page on encodeproject.org (<code>paired_end</code>, <code>paired_with</code>), not in this response.</p>\n<h2>Pitfalls &amp; Edge Cases</h2>\n<ul>\n<li><strong>Conda solver conflicts</strong>: Large conda environments with many packages can take hours to solve. Use mamba instead of conda for faster dependency resolution, or install in smaller focused environments.</li>\n<li><strong>R/Bioconductor version mismatch</strong>: R packages from CRAN and Bioconductor must match the R version. Installing Bioconductor 3.18 packages with R 4.4 will fail silently or produce errors. Use BiocManager::install() to ensure version compatibility.</li>\n<li><strong>Python 2 vs Python 3</strong>: Some legacy bioinformatics tools (MACS 1.x, old HOMER) require Python 2. Never install Python 2 tools in the same environment as Python 3 tools — use separate conda environments.</li>\n<li><strong>ARM Mac (M1/M2/M3) compatibility</strong>: Many bioinformatics tools lack native ARM builds. Use <code>CONDA_SUBDIR=osx-64</code> or Rosetta 2 emulation for x86_64 packages. Some tools (samtools, BWA) have ARM-native builds.</li>\n<li><strong>Nextflow requires Java 17+</strong>: check <code>java -version</code> before running pipelines. Install Nextflow with <code>scripts/install-nextflow.sh</code>, which pins the release the pipelines are validated against and verifies its checksum; avoid piping a remote installer straight into a shell.</li>\n<li><strong>Docker vs Singularity on HPC</strong>: Most HPC clusters do not allow Docker (requires root). Use Singularity instead. The pipeline skills express this through the execution profile, not a runtime profile: <code>-profile local</code> enables Docker, <code>-profile slurm</code> enables Singularity. There is no <code>docker</code> or <code>singularity</code> profile. With <code>-profile slurm</code>, convert the image once and pass the file: <code>singularity build pipeline-chipseq.sif docker-daemon://encode-toolkit/pipeline-chipseq:1.0.0</code>, then <code>--container /path/to/pipeline-chipseq.sif</code>.</li>\n</ul>\n<h2>Literature Foundation</h2>\n<table>\n<thead>\n<tr>\n<th>#</th>\n<th>Reference</th>\n<th>Key Contribution</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>Li &amp; Durbin 2009, Bioinformatics, DOI:10.1093/bioinformatics/btp324 (~30,000 cit)</td>\n<td>BWA aligner</td>\n</tr>\n<tr>\n<td>2</td>\n<td>Langmead &amp; Salzberg 2012, Nat Methods, DOI:10.1038/nmeth.1923 (~25,000 cit)</td>\n<td>Bowtie2 aligner</td>\n</tr>\n<tr>\n<td>3</td>\n<td>Li et al. 2009, Bioinformatics, DOI:10.1093/bioinformatics/btp352 (~20,000 cit)</td>\n<td>SAMtools/BAM format</td>\n</tr>\n<tr>\n<td>4</td>\n<td>Zhang et al. 2008, Genome Biol, DOI:10.1186/gb-2008-9-9-r137 (~7,000 cit)</td>\n<td>MACS2 peak caller</td>\n</tr>\n<tr>\n<td>5</td>\n<td>Dobin et al. 2013, Bioinformatics, DOI:10.1093/bioinformatics/bts635 (~15,000 cit)</td>\n<td>STAR RNA-seq aligner</td>\n</tr>\n<tr>\n<td>6</td>\n<td>Love et al. 2014, Genome Biol, DOI:10.1186/s13059-014-0550-8 (~30,000 cit)</td>\n<td>DESeq2</td>\n</tr>\n<tr>\n<td>7</td>\n<td>Ramirez et al. 2016, Nucleic Acids Res, DOI:10.1093/nar/gkw257 (~3,000 cit)</td>\n<td>deeptools</td>\n</tr>\n<tr>\n<td>8</td>\n<td>Wolf et al. 2018, Genome Biol, DOI:10.1186/s13059-017-1382-0 (~5,000 cit)</td>\n<td>Scanpy</td>\n</tr>\n<tr>\n<td>9</td>\n<td>Hao et al. 2021, Cell, DOI:10.1016/j.cell.2021.04.048 (~8,000 cit)</td>\n<td>Seurat v4</td>\n</tr>\n<tr>\n<td>10</td>\n<td>Quinlan &amp; Hall 2010, Bioinformatics, DOI:10.1093/bioinformatics/btq033 (~10,000 cit)</td>\n<td>bedtools</td>\n</tr>\n<tr>\n<td>11</td>\n<td>Ewels et al. 2016, Bioinformatics, DOI:10.1093/bioinformatics/btw354 (~3,000 cit)</td>\n<td>MultiQC</td>\n</tr>\n<tr>\n<td>12</td>\n<td>Krueger &amp; Andrews 2011, Bioinformatics, DOI:10.1093/bioinformatics/btr167 (~5,000 cit)</td>\n<td>Bismark</td>\n</tr>\n<tr>\n<td>13</td>\n<td>Heinz et al. 2010, Molecular Cell, DOI:10.1016/j.molcel.2010.05.004 (~7,000 cit)</td>\n<td>HOMER motif analysis</td>\n</tr>\n<tr>\n<td>14</td>\n<td>Bailey et al. 2015, Nucleic Acids Res, DOI:10.1093/nar/gkv416 (~3,000 cit)</td>\n<td>MEME Suite</td>\n</tr>\n<tr>\n<td>15</td>\n<td>Meers et al. 2019, Epigenetics Chromatin, DOI:10.1186/s13072-019-0287-4 (~800 cit)</td>\n<td>SEACR for CUT&amp;RUN</td>\n</tr>\n<tr>\n<td>16</td>\n<td>Di Tommaso et al. 2017, Nat Biotechnol, DOI:10.1038/nbt.3820 (~2,500 cit)</td>\n<td>Nextflow</td>\n</tr>\n<tr>\n<td>17</td>\n<td>Landt et al. 2012, Genome Res, DOI:10.1101/gr.136184.111 (~4,000 cit)</td>\n<td>ENCODE ChIP-seq standards</td>\n</tr>\n<tr>\n<td>18</td>\n<td>ENCODE Consortium 2020, Nature, DOI:10.1038/s41586-020-2493-4 (~1,656 cit)</td>\n<td>ENCODE Phase 3</td>\n</tr>\n<tr>\n<td>19</td>\n<td>Amemiya et al. 2019, Sci Rep, DOI:10.1038/s41598-019-45839-z (~1,372 cit)</td>\n<td>ENCODE Blacklist v2</td>\n</tr>\n</tbody>\n</table>\n<h2>Integration</h2>\n<table>\n<thead>\n<tr>\n<th>This skill produces...</th>\n<th>Feed into...</th>\n<th>Purpose</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Conda environments</td>\n<td><strong>pipeline-chipseq</strong> through <strong>pipeline-cutandrun</strong></td>\n<td>Provide tool dependencies for all pipeline stages</td>\n</tr>\n<tr>\n<td>Installed reference data</td>\n<td><strong>download-encode</strong></td>\n<td>Reference genomes and annotations for alignment</td>\n</tr>\n<tr>\n<td>Tool version inventory</td>\n<td><strong>data-provenance</strong></td>\n<td>Record exact tool versions for reproducibility</td>\n</tr>\n<tr>\n<td>QC tool installations</td>\n<td><strong>quality-assessment</strong></td>\n<td>Enable FastQC, MultiQC, and ENCODE QC metric tools</td>\n</tr>\n<tr>\n<td>Annotation tool setup</td>\n<td><strong>peak-annotation</strong></td>\n<td>HOMER, ChIPseeker for peak-to-gene assignment</td>\n</tr>\n<tr>\n<td>Motif scanning tools</td>\n<td><strong>jaspar-motifs</strong></td>\n<td>MEME Suite for motif scanning against JASPAR</td>\n</tr>\n<tr>\n<td>Visualization tools</td>\n<td><strong>visualization-workflow</strong></td>\n<td>deepTools, IGV, R/ggplot2 for data visualization</td>\n</tr>\n<tr>\n<td>Liftover utilities</td>\n<td><strong>liftover-coordinates</strong></td>\n<td>UCSC liftOver binary for assembly conversion</td>\n</tr>\n</tbody>\n</table>\n<h2>Related Skills</h2>\n<ul>\n<li><strong>pipeline-guide</strong>: Parent skill for all pipeline execution; provides overview of available pipelines and tool selection guidance</li>\n<li><strong>pipeline-chipseq</strong>: Uses the ChIP-seq conda environment tools for FASTQ-to-peaks processing</li>\n<li><strong>pipeline-atacseq</strong>: Uses the ATAC-seq conda environment tools for accessibility analysis</li>\n<li><strong>pipeline-rnaseq</strong>: Uses the RNA-seq conda environment for expression quantification</li>\n<li><strong>pipeline-wgbs</strong>: Uses the WGBS conda environment for methylation analysis</li>\n<li><strong>pipeline-hic</strong>: Uses the Hi-C conda environment for contact matrix generation</li>\n<li><strong>pipeline-dnaseseq</strong>: Uses the DNase-seq conda environment for hotspot detection</li>\n<li><strong>pipeline-cutandrun</strong>: Uses the CUT&amp;RUN conda environment for CUT&amp;RUN/CUT&amp;Tag processing</li>\n<li><strong>quality-assessment</strong>: Quality metrics require properly installed tools to compute</li>\n<li><strong>setup</strong>: Initial ENCODE Toolkit server setup (MCP connection, not bioinformatics tools)</li>\n<li><strong>motif-analysis</strong>: Requires HOMER and MEME Suite from this installer</li>\n<li><strong>visualization-workflow</strong>: Uses deeptools, pyGenomeTracks, and R visualization packages from this installer</li>\n<li><strong>single-cell-encode</strong>: Uses Seurat, Signac, Scanpy from this installer</li>\n<li><strong>publication-trust</strong>: Assess scientific integrity of publications before relying on their methods or findings</li>\n</ul>\n<h2>Presenting Results</h2>\n<ul>\n<li>Present installed tools as a checklist table: tool | version | status (installed/failed/skipped). Group by assay environment. Suggest: \"Would you like to verify the installation by running a quick test on sample ENCODE data?\"</li>\n<li>If any installation fails, provide the exact error and a targeted fix. Common fixes: update conda, set channel priority, install system dependencies.</li>\n</ul>\n<h2>For the request: \"$ARGUMENTS\"</h2>\n","files":[{"path":"environments/atacseq-env.yml","sizeBytes":974,"isText":true},{"path":"environments/chipseq-env.yml","sizeBytes":996,"isText":true},{"path":"environments/cutandrun-env.yml","sizeBytes":979,"isText":true},{"path":"environments/dnaseseq-env.yml","sizeBytes":1154,"isText":true},{"path":"environments/hic-env.yml","sizeBytes":856,"isText":true},{"path":"environments/rnaseq-env.yml","sizeBytes":808,"isText":true},{"path":"environments/wgbs-env.yml","sizeBytes":917,"isText":true},{"path":"references/literature.md","sizeBytes":8961,"isText":true},{"path":"scripts/constraints.txt","sizeBytes":23854,"isText":true},{"path":"scripts/install-nextflow.sh","sizeBytes":7182,"isText":true},{"path":"scripts/install-python-packages.sh","sizeBytes":3287,"isText":true},{"path":"scripts/install-r-packages.R","sizeBytes":5361,"isText":false},{"path":"scripts/requirements.in","sizeBytes":512,"isText":false},{"path":"SKILL.md","sizeBytes":36699,"isText":true}],"reviewScore":null,"reviewSummary":null,"trust":{"provenance":"trusted-source-unreviewed","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow.","bodySource":null},"bodyLocked":false,"purchaseUrl":null,"sourceUrl":null,"report":{"provenance":"trusted-source-unreviewed","screen":{"ran":true,"outcome":"clean","suspicious":0,"notes":0,"hiddenCharacters":false},"virusScan":{"engine":"clamav","status":"clean","scannedAt":"2026-09-22T13:25:41.468922Z","sha256":"D9B93DA4F49BED517E07184301B3EFEBC18E8B772968020503452B7D7BAE1CC7","sizeBytes":32110},"review":null,"source":{"repositoryUrl":"https://github.com/ammawla/encode-toolkit","path":"plugin/skills/bioinformatics-installer","license":"AGPL-3.0","commit":"36836c8725fd4d20d9c851ce314f5151aea5f57c","subtreeSha":"24F6CEC220ABB29F20175FBC01ED36DC30A869F93082767D9D593F786A2E5C21","lastSyncedAt":"2026-09-29T20:56:54.045383Z"},"reviewedAt":"2026-09-22T13:25:50.126698Z","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow."},"install":[{"target":"skills-cli","command":"npx skills add https://github.com/ammawla/encode-toolkit/tree/main/plugin/skills/bioinformatics-installer"},{"target":"claude-code","command":"claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install ammawla-encode-toolkit@llmmart"},{"target":"git","command":"git clone https://github.com/ammawla/encode-toolkit.git"}]}