{"slug":"pipeline-hic-2","title":"pipeline-hic","summary":"Execute ENCODE Hi-C pipeline from FASTQ to contact matrices and loop calls. Child of pipeline-guide. Provides Nextflow execution with Docker and cloud deployment. Use when processing Hi-C data, generating contact matrices, or calling loops. Trigger on: Hi-C pipeline, chromatin co","platform":"Claude","tags":[],"authorName":"LLM Mart","authorSlug":"llm-mart","score":0,"source":"github","price":null,"verified":false,"createdAt":"2026-09-22T13:25:44.533978Z","repo":{"url":"https://github.com/ammawla/encode-toolkit","stars":21,"forks":5,"license":"AGPL-3.0","updatedAt":"2026-09-27T01:22:42Z"},"bodyHtml":"<hr>\n<h2>name: pipeline-hic\ndescription: \"Execute ENCODE Hi-C pipeline from FASTQ to contact matrices and loop calls. Child of pipeline-guide. Provides Nextflow execution with Docker and cloud deployment. Use when processing Hi-C data, generating contact matrices, or calling loops. Trigger on: Hi-C pipeline, chromatin conformation, contact matrix, loop calling, TAD detection, Juicer, HiCCUPS, 3D genome.\"</h2>\n<h1>ENCODE Hi-C Pipeline: FASTQ to Contact Matrices and Loops</h1>\n<h2>When to Use</h2>\n<ul>\n<li>User wants to run a Hi-C processing pipeline from FASTQ to contact matrices and loop calls</li>\n<li>User asks about \"Hi-C pipeline\", \"contact matrix\", \"loop calling\", \"Juicer\", \"HiCCUPS\", or \"TAD detection\"</li>\n<li>User needs to process Hi-C data for 3D genome structure analysis</li>\n<li>Example queries: \"process my Hi-C FASTQs\", \"generate contact matrices from Hi-C\", \"call chromatin loops with HiCCUPS\"</li>\n</ul>\n<p>Execute the ENCODE Hi-C pipeline for chromatin conformation capture data,\nproducing multi-resolution contact matrices and loop calls.</p>\n<h2>Pipeline Overview</h2>\n<pre><code>FASTQ -&gt; FastQC (raw reads)\n      -&gt; bwa mem -SP5M (both mates in one call) -&gt; {sample}.paired.bam\n         -&gt; pairtools parse -&gt; sort -&gt; dedup -&gt; select UU\n              |\n              +-&gt; Juicer pre -&gt; .hic -&gt; HiCCUPS -&gt; loops (BEDPE)\n              |\n              +-&gt; cooler cload + zoomify -&gt; .mcool\n</code></pre>\n<h3>Not run by this workflow</h3>\n<p>Adapter trimming, TAD calling, A/B compartment calling, and every cooltools\nanalysis are outside this workflow. They are documented as manual, optional\nsteps only: trimming in <code>references/01-qc-trimming.md</code>, distance decay and\ncompartments in <code>references/04-matrix-generation.md</code>. cooltools and bedtools\nare not installed in the container image, so those manual commands need the\nconda environment (<code>environments/hic-env.yml</code> in the <code>bioinformatics-installer</code>\nskill) or a separate install.</p>\n<h3>ENCODE Repository</h3>\n<ul>\n<li><strong>GitHub</strong>: <code>ENCODE-DCC/hic-pipeline</code></li>\n<li><strong>Container</strong>: built from <code>scripts/Dockerfile</code> in this skill (<code>docker build -t encode-toolkit/pipeline-hic:1.0.0 scripts/</code>); override with <code>--container</code></li>\n<li><strong>WDL</strong>: Available for Cromwell execution</li>\n<li><strong>This skill</strong>: Nextflow DSL2 reimplementation for portability</li>\n</ul>\n<h2>Core Tools and Versions</h2>\n<p>Versions are those installed by <code>scripts/Dockerfile</code>, which is what the\nworkflow runs.</p>\n<table>\n<thead>\n<tr>\n<th>Tool</th>\n<th>Version</th>\n<th>Purpose</th>\n<th>Citation</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>BWA-MEM</td>\n<td>0.7.18</td>\n<td>Alignment (both mates, <code>-SP5M</code>)</td>\n<td>Li &amp; Durbin 2009</td>\n</tr>\n<tr>\n<td>pairtools</td>\n<td>1.1.2</td>\n<td>Pair classification, dedup</td>\n<td>Open2C</td>\n</tr>\n<tr>\n<td>Juicer tools</td>\n<td>2.20.00</td>\n<td>.hic generation, HiCCUPS</td>\n<td>Durand et al. 2016</td>\n</tr>\n<tr>\n<td>cooler</td>\n<td>0.9.3</td>\n<td>.cool/.mcool generation</td>\n<td>Abdennur &amp; Mirny 2020</td>\n</tr>\n<tr>\n<td>samtools</td>\n<td>1.19</td>\n<td>BAM operations</td>\n<td>Li et al. 2009</td>\n</tr>\n<tr>\n<td>FastQC</td>\n<td>0.12.1</td>\n<td>Read quality</td>\n<td>Andrews (Babraham)</td>\n</tr>\n<tr>\n<td>MultiQC</td>\n<td>1.21</td>\n<td>Aggregated QC</td>\n<td>Ewels et al. 2016</td>\n</tr>\n</tbody>\n</table>\n<p>The conda alternative (<code>hic-env.yml</code>) ships no juicer_tools jar, only a JRE, so\n<code>.hic</code> generation and HiCCUPS are unavailable on that route; conversely it\nprovides cooltools and bedtools, which the container image does not.</p>\n<h2>Key Literature</h2>\n<ol>\n<li><p><strong>Rao et al. 2014</strong> - \"A 3D Map of the Human Genome at Kilobase Resolution\nReveals Principles of Chromatin Looping\" (Cell, ~5,000 citations)\nDOI: 10.1016/j.cell.2014.11.021</p>\n</li>\n<li><p><strong>Lieberman-Aiden et al. 2009</strong> - \"Comprehensive Mapping of Long-Range\nInteractions Reveals Folding Principles of the Human Genome\" (Science, ~6,000 citations)\nDOI: 10.1126/science.1181369</p>\n</li>\n<li><p><strong>Durand et al. 2016</strong> - \"Juicer Provides a One-Click System for Analyzing\nLoop-Resolution Hi-C Experiments\" (Cell Systems, ~2,000 citations)\nDOI: 10.1016/j.cels.2016.07.002</p>\n</li>\n<li><p><strong>Abdennur &amp; Mirny 2020</strong> - \"Cooler: scalable storage for Hi-C data and\nother genomically labeled arrays\" (Bioinformatics)\nDOI: 10.1093/bioinformatics/btz540</p>\n</li>\n<li><p><strong>Amemiya et al. 2019</strong> - \"The ENCODE Blacklist\" (Scientific Reports, ~1,372 citations)\nDOI: 10.1038/s41598-019-45839-z</p>\n</li>\n</ol>\n<h2>Execution</h2>\n<h3>Quick Start (Local)</h3>\n<pre><code>nextflow run scripts/main.nf \\\n    -profile local \\\n    --reads '/data/fastq/*_R{1,2}.fastq.gz' \\\n    --bwa_index '/ref/bwa_index/GRCh38.fa' \\\n    --chrom_sizes '/ref/hg38.chrom.sizes' \\\n    --outdir results/ \\\n    -resume\n</code></pre>\n<h3>SLURM HPC</h3>\n<pre><code>nextflow run scripts/main.nf \\\n    -profile slurm \\\n    --container /path/to/pipeline-hic.sif \\\n    --reads '/data/fastq/*_R{1,2}.fastq.gz' \\\n    --bwa_index '/ref/bwa_index/GRCh38.fa' \\\n    --chrom_sizes '/ref/hg38.chrom.sizes' \\\n    --outdir results/ \\\n    -resume\n</code></pre>\n<h3>Cloud (GCP / AWS)</h3>\n<pre><code># Google Cloud Batch\nnextflow run scripts/main.nf -profile gcp \\\n    --container us-docker.pkg.dev/&lt;project&gt;/&lt;repo&gt;/pipeline-hic:1.0.0 \\\n    --gcp_project &lt;project&gt; \\\n    --gcp_workdir gs://&lt;bucket&gt;/work \\\n    --reads 'gs://&lt;bucket&gt;/fastq/*_R{1,2}.fastq.gz' \\\n    --bwa_index gs://&lt;bucket&gt;/ref/GRCh38.fa \\\n    --chrom_sizes gs://&lt;bucket&gt;/ref/hg38.chrom.sizes \\\n    --outdir gs://&lt;bucket&gt;/results\n\n# AWS Batch\nnextflow run scripts/main.nf -profile aws \\\n    --container &lt;account&gt;.dkr.ecr.&lt;region&gt;.amazonaws.com/pipeline-hic:1.0.0 \\\n    --aws_queue &lt;job-queue&gt; \\\n    --aws_workdir s3://&lt;bucket&gt;/work \\\n    --reads 's3://&lt;bucket&gt;/fastq/*_R{1,2}.fastq.gz' \\\n    --bwa_index s3://&lt;bucket&gt;/ref/GRCh38.fa \\\n    --chrom_sizes s3://&lt;bucket&gt;/ref/hg38.chrom.sizes \\\n    --outdir s3://&lt;bucket&gt;/results\n</code></pre>\n<p><code>--outdir</code> only sets where results are published; Google Batch and AWS Batch\nstage every task through the work directory, and the workflow stops with an\nerror if it or the project/queue is missing.</p>\n<h2>Resource Requirements</h2>\n<table>\n<thead>\n<tr>\n<th>Step</th>\n<th>CPUs</th>\n<th>RAM</th>\n<th>Time (2B contacts)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>BWA alignment</td>\n<td>8</td>\n<td>16 GB</td>\n<td>4-6 hours</td>\n</tr>\n<tr>\n<td>pairtools parse + sort</td>\n<td>4</td>\n<td>16 GB</td>\n<td>2-3 hours</td>\n</tr>\n<tr>\n<td>pairtools dedup</td>\n<td>4</td>\n<td>16 GB</td>\n<td>1-2 hours</td>\n</tr>\n<tr>\n<td>Juicer pre + hic</td>\n<td>4</td>\n<td>64 GB</td>\n<td>2-4 hours</td>\n</tr>\n<tr>\n<td>HiCCUPS</td>\n<td>4</td>\n<td>16 GB</td>\n<td>1-2 hours</td>\n</tr>\n<tr>\n<td><strong>Total</strong></td>\n<td><strong>8</strong></td>\n<td><strong>64 GB</strong></td>\n<td><strong>8-16 hours</strong></td>\n</tr>\n</tbody>\n</table>\n<p>The RAM column is each step's first-attempt request. Every process asks for that\nmuch memory per attempt, so a task killed for exceeding it is retried with more\n(at most two retries, capped by <code>--max_memory</code>). Failures with any other exit\nstatus stop the run.</p>\n<h2>Pipeline Parameters</h2>\n<table>\n<thead>\n<tr>\n<th>Parameter</th>\n<th>Default</th>\n<th>Description</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>--reads</code></td>\n<td>required</td>\n<td>Glob pattern to paired FASTQ files</td>\n</tr>\n<tr>\n<td><code>--bwa_index</code></td>\n<td>required</td>\n<td>BWA index prefix: the genome FASTA path whose <code>.amb .ann .bwt .pac .sa</code> files sit beside it (every file starting with this prefix is staged)</td>\n</tr>\n<tr>\n<td><code>--chrom_sizes</code></td>\n<td>required</td>\n<td>Chromosome sizes file</td>\n</tr>\n<tr>\n<td><code>--outdir</code></td>\n<td><code>./results</code></td>\n<td>Output directory</td>\n</tr>\n<tr>\n<td><code>--resolutions</code></td>\n<td><code>1000,5000,10000,25000,50000,100000,250000,500000,1000000</code></td>\n<td>Matrix resolutions for <code>juicer_tools pre</code> and <code>cooler zoomify</code>: a comma-separated list of positive integers. The smallest value is the cooler base bin, and every other value must be a multiple of it; the workflow stops with an error otherwise</td>\n</tr>\n<tr>\n<td><code>--hiccups_resolutions</code></td>\n<td><code>5000,10000,25000</code></td>\n<td>Resolutions HiCCUPS calls loops at. Only 5000, 10000 and 25000 are accepted, and each must also be listed in <code>--resolutions</code>; the workflow stops with an error otherwise. Peak width (<code>-p</code>), window width (<code>-i</code>), merge radius (<code>-d</code>) and FDR (<code>-f</code>) follow Juicer's published per-resolution defaults, one value per resolution</td>\n</tr>\n<tr>\n<td><code>--min_mapq</code></td>\n<td><code>30</code></td>\n<td>Minimum MAPQ passed to <code>pairtools parse</code></td>\n</tr>\n<tr>\n<td><code>--hiccups_gpu</code></td>\n<td><code>false</code></td>\n<td>Run HiCCUPS on an NVIDIA GPU. By default the CPU mode is used, which only searches within 8 Mb of the diagonal</td>\n</tr>\n<tr>\n<td><code>--assembly</code></td>\n<td><code>hg38</code></td>\n<td>Assembly name recorded in the <code>.mcool</code> metadata (<code>cooler cload --assembly</code>). The <code>.hic</code> file is built from the chrom.sizes file only</td>\n</tr>\n</tbody>\n</table>\n<h3>Infrastructure parameters (<code>nextflow.config</code>)</h3>\n<table>\n<thead>\n<tr>\n<th>Parameter</th>\n<th>Default</th>\n<th>Description</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>--container</code></td>\n<td><code>encode-toolkit/pipeline-hic:1.0.0</code></td>\n<td>Image built from <code>scripts/Dockerfile</code>. Pass a registry image for <code>gcp</code>/<code>aws</code>, or a <code>.sif</code> file for <code>slurm</code></td>\n</tr>\n<tr>\n<td><code>--max_cpus</code>, <code>--max_memory</code>, <code>--max_time</code></td>\n<td><code>16</code>, <code>128.GB</code>, <code>48.h</code></td>\n<td>Upper bounds applied to every process</td>\n</tr>\n<tr>\n<td><code>--slurm_queue</code>, <code>--slurm_account</code></td>\n<td><code>normal</code>, none</td>\n<td>SLURM partition and account</td>\n</tr>\n<tr>\n<td><code>--gcp_project</code>, <code>--gcp_workdir</code></td>\n<td>none (both required for <code>-profile gcp</code>)</td>\n<td>Google Cloud project and <code>gs://</code> work directory</td>\n</tr>\n<tr>\n<td><code>--gcp_location</code>, <code>--gcp_disk</code></td>\n<td><code>us-central1</code>, <code>200.GB</code></td>\n<td>Google Batch region and per-task disk</td>\n</tr>\n<tr>\n<td><code>--aws_queue</code>, <code>--aws_workdir</code></td>\n<td>none (both required for <code>-profile aws</code>)</td>\n<td>AWS Batch job queue and <code>s3://</code> work directory</td>\n</tr>\n<tr>\n<td><code>--aws_region</code>, <code>--aws_cli_path</code></td>\n<td><code>us-east-1</code>, <code>/home/ec2-user/miniconda/bin/aws</code></td>\n<td>AWS region, and the AWS CLI path inside the Batch AMI</td>\n</tr>\n</tbody>\n</table>\n<h2>Output Files</h2>\n<pre><code>results/\n  fastqc/\n    *_fastqc.html                 # Raw read quality\n    *_fastqc.zip\n  alignment/\n    {sample}.paired.bam           # Both mates from one bwa mem -SP5M call\n  pairs/\n    {sample}.parse_stats.txt      # pairtools parse stats (pair-type breakdown)\n    {sample}.dedup.pairs.gz       # Classified, sorted, deduplicated pairs\n    {sample}.dedup_stats.txt      # pairtools dedup stats (duplication, complexity)\n  matrices/\n    {sample}.hic                  # Juicer .hic file (primary output)\n    {sample}.mcool                # Cooler multi-resolution matrix\n  loops/\n    {sample}.hiccups_loops.bedpe  # HiCCUPS merged_loops.bedpe, renamed\n  qc/\n    {sample}.contact_stats.txt    # pairtools stats on the selected UU pairs\n  multiqc/\n    multiqc_report.html\n  pipeline_info/\n    timeline.html\n    report.html\n    trace.txt\n</code></pre>\n<p>The UU-selected pairs file and the <code>pairtools stats</code> run on it are intermediate:\nonly <code>qc/{sample}.contact_stats.txt</code> is published, not the selected pairs.</p>\n<h3>.hic File Format</h3>\n<p>The .hic format (Juicer) stores multi-resolution contact matrices with\nnormalization vectors. Can be visualized in Juicebox and loaded by\n<code>hic-straw</code> in Python/R.</p>\n<h3>.mcool File Format</h3>\n<p>The .mcool format (cooler) is an HDF5-based multi-resolution contact matrix.\nWidely supported by <code>cooler</code>, <code>cooltools</code>, <code>HiGlass</code>, and <code>FAN-C</code>.</p>\n<h2>QC Thresholds (ENCODE Standards)</h2>\n<p>This is the only QC threshold table for this skill; the reference files point\nback to it.</p>\n<table>\n<thead>\n<tr>\n<th>Metric</th>\n<th>Pass</th>\n<th>Warning</th>\n<th>Fail</th>\n<th>Computed from</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Valid (UU) pair fraction</td>\n<td>&gt;40%</td>\n<td>25-40%</td>\n<td>&lt;25%</td>\n<td><code>pairs/{sample}.parse_stats.txt</code></td>\n</tr>\n<tr>\n<td>Cis contacts (&gt;20kb)</td>\n<td>&gt;40%</td>\n<td>25-40%</td>\n<td>&lt;25%</td>\n<td><code>qc/{sample}.contact_stats.txt</code></td>\n</tr>\n<tr>\n<td>Cis/trans ratio</td>\n<td>&gt;1.5</td>\n<td>1.0-1.5</td>\n<td>&lt;1.0</td>\n<td><code>qc/{sample}.contact_stats.txt</code></td>\n</tr>\n<tr>\n<td>Library complexity (unique/total)</td>\n<td>&gt;0.7</td>\n<td>0.5-0.7</td>\n<td>&lt;0.5</td>\n<td><code>pairs/{sample}.dedup_stats.txt</code></td>\n</tr>\n</tbody>\n</table>\n<p><code>qc/{sample}.contact_stats.txt</code> is computed after UU selection, so its pair-type\nbreakdown is 100% UU by construction. Read pair types from\n<code>pairs/{sample}.parse_stats.txt</code> instead.</p>\n<h3>Resolution vs Depth Requirements</h3>\n<table>\n<thead>\n<tr>\n<th>Resolution</th>\n<th>Minimum Contacts Needed</th>\n<th>Typical Depth</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1 kb</td>\n<td>&gt;2 billion</td>\n<td>Very deep</td>\n</tr>\n<tr>\n<td>5 kb</td>\n<td>&gt;500 million</td>\n<td>Deep</td>\n</tr>\n<tr>\n<td>10 kb</td>\n<td>&gt;200 million</td>\n<td>Standard</td>\n</tr>\n<tr>\n<td>25 kb</td>\n<td>&gt;50 million</td>\n<td>Moderate</td>\n</tr>\n<tr>\n<td>100 kb</td>\n<td>&gt;10 million</td>\n<td>Low</td>\n</tr>\n</tbody>\n</table>\n<h2>Pair Classification</h2>\n<p>pairtools assigns each read pair a two-letter code (one letter per side:\nU unique, R rescued, M multi, N null/unmapped, W walk, D duplicate, X corrupt):</p>\n<table>\n<thead>\n<tr>\n<th>Category</th>\n<th>Description</th>\n<th>Use</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>UU</td>\n<td>Both sides uniquely mapped</td>\n<td>Valid contact -- the only type this workflow keeps</td>\n</tr>\n<tr>\n<td>UR / RU</td>\n<td>One unique, one rescued</td>\n<td>Valid but not selected here</td>\n</tr>\n<tr>\n<td>NU</td>\n<td>One unique, one unmapped</td>\n<td>Not used</td>\n</tr>\n<tr>\n<td>NM</td>\n<td>One unmapped, one multi-mapped</td>\n<td>Not used</td>\n</tr>\n<tr>\n<td>MM</td>\n<td>Both multi-mapped</td>\n<td>Not used</td>\n</tr>\n<tr>\n<td>WW</td>\n<td>Complex walk (multiple ligation events), masked by <code>--walks-policy mask</code></td>\n<td>Not used</td>\n</tr>\n<tr>\n<td>DD</td>\n<td>Duplicate</td>\n<td>Removed by <code>pairtools dedup</code></td>\n</tr>\n<tr>\n<td>XX</td>\n<td>Corrupt record</td>\n<td>Not used</td>\n</tr>\n</tbody>\n</table>\n<p>This workflow selects UU only\n(<code>pairtools select '(pair_type == \"UU\")'</code>) before matrix generation.</p>\n<h2>Critical Pitfalls</h2>\n<h3>Restriction Enzyme Choice</h3>\n<p>The restriction enzyme determines fragment size and resolution:</p>\n<ul>\n<li><strong>MboI/DpnII</strong> (GATC): 4-cutter, ~256 bp average fragment -- higher resolution</li>\n<li><strong>HindIII</strong> (AAGCTT): 6-cutter, ~4 kb average fragment -- lower resolution</li>\n<li><strong>Arima</strong> (proprietary): Two enzymes, ~160 bp average -- highest resolution</li>\n<li>Always verify which enzyme was used before interpreting resolution</li>\n<li>The workflow itself is enzyme-agnostic: pairtools works at read-pair level and the <code>.hic</code>\nfile is built without a restriction-site file, so there is no enzyme parameter to set</li>\n</ul>\n<h3>Normalization Method</h3>\n<p>Different normalization methods yield different results:</p>\n<ul>\n<li><strong>KR</strong> (Knight-Ruiz): built by <code>juicer_tools pre -k KR,VC,VC_SQRT</code> and used by HiCCUPS (<code>-k KR</code>)</li>\n<li><strong>ICE</strong> (Imakaev et al.): applied to the .mcool by <code>cooler zoomify --balance</code></li>\n<li><strong>VC</strong> (Vanilla Coverage): simple coverage normalization, also built into the .hic</li>\n<li>Always document which normalization a downstream analysis read.</li>\n</ul>\n<h3>Resolution Depends on Depth</h3>\n<p>Do not call features at resolutions unsupported by sequencing depth:</p>\n<ul>\n<li>Calling 1 kb loops from 100M contacts will produce noise</li>\n<li>Check the Juicer resolution QC to determine achievable resolution</li>\n<li>HiCCUPS runs at the resolutions in <code>--hiccups_resolutions</code> (default 5 kb, 10 kb\nand 25 kb); each of them must also be listed in <code>--resolutions</code>, because HiCCUPS\nreads them out of the <code>.hic</code> file</li>\n</ul>\n<h3>Ligation Artifacts</h3>\n<p>Monitor the pair-type breakdown in <code>pairs/{sample}.parse_stats.txt</code>:</p>\n<ul>\n<li>WW pairs are complex walks (more than one ligation in a read); <code>--walks-policy mask</code>\nmasks them so they never reach the contact matrix</li>\n<li>A large unmapped/multi-mapped fraction points at poor library or the wrong genome</li>\n<li>The workflow produces no re-ligation distance plot; derive one manually from the\npublished pairs file if needed</li>\n</ul>\n<h2>Provenance Integration</h2>\n<p>After pipeline completion, log all outputs:</p>\n<pre><code>encode_log_derived_file(\n    file_path=\"/results/matrices/sample1.hic\",\n    source_accessions=[\"ENCSR...\", \"ENCFF...\"],\n    description=\"Hi-C contact matrix from ENCODE Hi-C pipeline\",\n    file_type=\"hic\",\n    tool_used=\"BWA 0.7.18 + pairtools 1.1.2 + Juicer 2.20.00\",\n    parameters=\"--min_mapq 30, UU pairs only, KR/VC/VC_SQRT normalization, resolutions 1kb-1Mb\"\n)\n</code></pre>\n<h2>Reference Files</h2>\n<p>Detailed step-by-step documentation is provided in the <code>references/</code> directory:</p>\n<ol>\n<li><code>01-qc-trimming.md</code> -- Read QC (trimming is a manual option, not run here)</li>\n<li><code>02-alignment.md</code> -- BWA <code>-SP5M</code> alignment of both mates in one call</li>\n<li><code>03-pair-processing.md</code> -- pairtools parse, sort, dedup, and select</li>\n<li><code>04-matrix-generation.md</code> -- Juicer .hic and cooler .mcool generation; manual cooltools analyses</li>\n<li><code>05-loop-calling.md</code> -- HiCCUPS loop detection and QC</li>\n</ol>\n<h2>Walkthrough: Processing ENCODE Hi-C from FASTQ to Contact Maps and Loops</h2>\n<p><strong>Goal</strong>: Process raw Hi-C FASTQ files through the ENCODE pipeline to generate contact matrices and chromatin loop calls.\n<strong>Context</strong>: Hi-C captures 3D chromatin organization. The pipeline uses BWA for chimeric read alignment, pairtools for pair processing, and Juicer/HiCCUPS for loop calling.</p>\n<h3>Step 1: Find Hi-C experiment</h3>\n<pre><code>encode_get_experiment(accession=\"ENCSR000AKA\")\n</code></pre>\n<p>Expected output:</p>\n<pre><code>{\n  \"accession\": \"ENCSR000AKA\",\n  \"assay_title\": \"Hi-C\",\n  \"biosample_summary\": \"GM12878\",\n  \"bio_replicate_count\": 2,\n  \"status\": \"released\"\n}\n</code></pre>\n<h3>Step 2: List FASTQ files</h3>\n<pre><code>encode_list_files(experiment_accession=\"ENCSR000AKA\", file_format=\"fastq\")\n</code></pre>\n<p>Expected output (a JSON array of file records; fields abridged):</p>\n<pre><code>[\n  {\"accession\": \"ENCFF500HI1\", \"file_format\": \"fastq\", \"output_type\": \"reads\", \"biological_replicates\": [1], \"file_size_human\": \"34.2 GB\", \"status\": \"released\"},\n  {\"accession\": \"ENCFF501HI2\", \"file_format\": \"fastq\", \"output_type\": \"reads\", \"biological_replicates\": [1], \"file_size_human\": \"35.2 GB\", \"status\": \"released\"}\n]\n</code></pre>\n<p><strong>Interpretation</strong>: Hi-C paired-end reads represent chimeric ligation junctions. Each read pair captures a 3D contact.</p>\n<h3>Step 3: Name the files so a read-pair glob can find them</h3>\n<p>ENCODE FASTQs are named by accession, so the two mates of a pair share no\nprefix, and the workflow matches file pairs with a <code>{1,2}</code> glob. Which mate a\nfile is comes from its page on encodeproject.org (<code>paired_end</code> 1 or 2, and\n<code>paired_with</code> naming the other accession), not from any tool here. Link the\nfiles into the shape the glob expects:</p>\n<pre><code>mkdir -p fastq\nln -s \"$PWD/ENCFF500HI1.fastq.gz\" fastq/ENCSR000AKA_R1.fastq.gz\nln -s \"$PWD/ENCFF501HI2.fastq.gz\" fastq/ENCSR000AKA_R2.fastq.gz\n</code></pre>\n<h3>Step 4: Run the Hi-C pipeline</h3>\n<pre><code>nextflow run scripts/main.nf \\\n  -profile local \\\n  --reads 'fastq/ENCSR000AKA_R{1,2}.fastq.gz' \\\n  --bwa_index '/ref/bwa_index/GRCh38.fa' \\\n  --chrom_sizes '/ref/hg38.chrom.sizes' \\\n  --outdir results/ \\\n  -resume\n</code></pre>\n<p>Key pipeline steps:</p>\n<ol>\n<li>FastQC on the raw reads</li>\n<li>BWA-MEM <code>-SP5M</code> alignment of both mates in one call (chimeric read handling)</li>\n<li>pairtools parse + sort (classify pairs, MAPQ 30, mask walks)</li>\n<li>pairtools dedup (remove PCR duplicates), then select UU pairs</li>\n<li>Contact matrix generation (<code>.hic</code> via Juicer, <code>.mcool</code> via cooler)</li>\n<li>Loop calling (HiCCUPS at the <code>--hiccups_resolutions</code>, by default 5 kb, 10 kb\nand 25 kb, merged into one BEDPE)</li>\n</ol>\n<h3>Step 5: Validate output quality</h3>\n<p>Use the QC threshold table above with <code>pairs/{sample}.parse_stats.txt</code>,\n<code>pairs/{sample}.dedup_stats.txt</code> and <code>qc/{sample}.contact_stats.txt</code>.</p>\n<h3>Step 6: Identify significant loops</h3>\n<p>Download loop calls for downstream analysis:</p>\n<pre><code>encode_list_files(experiment_accession=\"ENCSR000AKA\", file_format=\"bedpe\", assembly=\"GRCh38\")\n</code></pre>\n<h3>Integration with downstream skills</h3>\n<ul>\n<li>Loop calls (BEDPE) feed into -&gt; <strong>hic-aggregation</strong> for cross-tissue loop catalog</li>\n<li>Loop anchors feed into -&gt; <strong>peak-annotation</strong> for enhancer-promoter assignment</li>\n<li>Contact data integrates with -&gt; <strong>visualization-workflow</strong> for 3D genome display</li>\n<li>Pipeline provenance logged by -&gt; <strong>data-provenance</strong></li>\n</ul>\n<h2>Code Examples</h2>\n<h3>1. Find Hi-C data for 3D genome analysis</h3>\n<pre><code>encode_search_experiments(\n  assay_title=\"Hi-C\",\n  organ=\"heart\"\n)\n</code></pre>\n<p>Expected output:</p>\n<pre><code>{\n  \"results\": [\n    {\n      \"accession\": \"ENCSR654HRT\",\n      \"assay_title\": \"Hi-C\",\n      \"biosample_summary\": \"heart left ventricle tissue male adult (51 years)\",\n      \"status\": \"released\"\n    }\n  ],\n  \"total\": 4,\n  \"limit\": 25,\n  \"offset\": 0,\n  \"has_more\": false,\n  \"next_offset\": null\n}\n</code></pre>\n<h3>2. Get experiment details for pipeline configuration</h3>\n<pre><code>encode_get_experiment(accession=\"ENCSR654HRT\")\n</code></pre>\n<p>Expected output:</p>\n<pre><code>{\n  \"accession\": \"ENCSR654HRT\",\n  \"assay_title\": \"Hi-C\",\n  \"bio_replicate_count\": 2,\n  \"biosample_summary\": \"heart left ventricle tissue male adult (51 years)\",\n  \"assembly\": [\"GRCh38\"],\n  \"audit_error_count\": 0,\n  \"audit_warning_count\": 1\n}\n</code></pre>\n<h2>Integration</h2>\n<table>\n<thead>\n<tr>\n<th>This skill produces...</th>\n<th>Feed into...</th>\n<th>Purpose</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Chromatin loops (BEDPE)</td>\n<td><strong>hic-aggregation</strong></td>\n<td>Cross-tissue loop catalog</td>\n</tr>\n<tr>\n<td>Loop anchors (BED)</td>\n<td><strong>peak-annotation</strong></td>\n<td>Assign genes to loop-connected enhancers</td>\n</tr>\n<tr>\n<td>Contact matrices (.hic / .mcool)</td>\n<td><strong>visualization-workflow</strong></td>\n<td>3D genome visualization</td>\n</tr>\n<tr>\n<td>Loop-disrupting coordinates</td>\n<td><strong>variant-annotation</strong></td>\n<td>Identify variants breaking chromatin contacts</td>\n</tr>\n<tr>\n<td>QC metrics</td>\n<td><strong>quality-assessment</strong></td>\n<td>Validate Hi-C library quality</td>\n</tr>\n<tr>\n<td>Pipeline parameters</td>\n<td><strong>data-provenance</strong></td>\n<td>Record BWA/pairtools/Juicer versions</td>\n</tr>\n<tr>\n<td>Loop anchor regions</td>\n<td><strong>motif-analysis</strong></td>\n<td>Discover CTCF motifs at loop anchors</td>\n</tr>\n</tbody>\n</table>\n<h2>Related Skills</h2>\n<ul>\n<li><code>pipeline-guide</code> -- Parent skill with compute resource assessment and cloud setup</li>\n<li><code>hic-aggregation</code> -- Aggregate Hi-C loops across samples/tissues</li>\n<li><code>quality-assessment</code> -- Evaluate pipeline output quality metrics</li>\n<li><code>data-provenance</code> -- Track all pipeline inputs, outputs, and parameters</li>\n<li><code>download-encode</code> -- Download ENCODE Hi-C FASTQ files for pipeline input</li>\n<li><code>publication-trust</code> -- Verify literature claims backing analytical decisions</li>\n</ul>\n<h2>Presenting Results</h2>\n<p>When reporting Hi-C pipeline results:</p>\n<ul>\n<li><strong>Valid pair count</strong>: Report the UU pair count and its fraction of all parsed pairs from <code>pairs/{sample}.parse_stats.txt</code>. UU is the only pair type this workflow carries forward</li>\n<li><strong>Cis/trans ratio</strong>: Report the cis/trans contact ratio (&gt;1.5 pass) and long-range cis fraction (&gt;20kb, &gt;40% expected) from <code>qc/{sample}.contact_stats.txt</code>. These are the primary Hi-C quality indicators</li>\n<li><strong>Contact matrix resolution</strong>: Report the achievable resolution based on sequencing depth (e.g., \"500M valid pairs supports 5kb resolution\") and list the <code>--resolutions</code> actually generated</li>\n<li><strong>Loop counts</strong>: Report the number of loops in <code>loops/{sample}.hiccups_loops.bedpe</code> (the file starts with a <code>#chr1 ...</code> header line, so exclude it from the count). HiCCUPS searches the resolutions in <code>--hiccups_resolutions</code> (default 5 kb, 10 kb and 25 kb) and merges them into that single file; per-resolution files stay in the Nextflow work directory</li>\n<li><strong>Matrix paths</strong>: Provide paths to the .hic file (Juicebox-compatible) and .mcool file (cooler/HiGlass-compatible)</li>\n<li><strong>Key QC metrics</strong>: Present library complexity (unique/total &gt;0.7, from <code>pairs/{sample}.dedup_stats.txt</code>) and the pair-type breakdown from <code>pairs/{sample}.parse_stats.txt</code> in a summary table</li>\n<li><strong>Normalization</strong>: Note that <code>.hic</code> carries KR, VC and VC_SQRT vectors (HiCCUPS uses KR) and the <code>.mcool</code> is ICE-balanced by <code>cooler zoomify --balance</code></li>\n<li><strong>Not produced here</strong>: TAD calls, A/B compartments and cooltools outputs are not generated by this workflow; say so rather than implying they are missing</li>\n<li><strong>Next steps</strong>: Suggest <code>hic-aggregation</code> for cross-sample loop catalogs, or <code>visualization-workflow</code> for Juicebox/HiGlass session setup</li>\n</ul>\n<h2>For the request: \"$ARGUMENTS\"</h2>\n","files":[{"path":"references/01-qc-trimming.md","sizeBytes":2904,"isText":true},{"path":"references/02-alignment.md","sizeBytes":3285,"isText":true},{"path":"references/03-pair-processing.md","sizeBytes":4616,"isText":true},{"path":"references/04-matrix-generation.md","sizeBytes":4874,"isText":true},{"path":"references/05-loop-calling.md","sizeBytes":6660,"isText":true},{"path":"references/literature.md","sizeBytes":11456,"isText":true},{"path":"scripts/Dockerfile","sizeBytes":2123,"isText":false},{"path":"scripts/main.nf","sizeBytes":12829,"isText":false},{"path":"scripts/nextflow.config","sizeBytes":4243,"isText":false},{"path":"SKILL.md","sizeBytes":21796,"isText":true}],"reviewScore":null,"reviewSummary":null,"trust":{"provenance":"trusted-source-unreviewed","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow.","bodySource":null},"bodyLocked":false,"purchaseUrl":null,"sourceUrl":null,"report":{"provenance":"trusted-source-unreviewed","screen":{"ran":true,"outcome":"clean","suspicious":0,"notes":0,"hiddenCharacters":false},"virusScan":{"engine":"clamav","status":"clean","scannedAt":"2026-09-22T13:27:29.388843Z","sha256":"07A3B81605664EA1285DD42198723E079A52A4263DF5338CC864B50074D2332E","sizeBytes":30572},"review":null,"source":{"repositoryUrl":"https://github.com/ammawla/encode-toolkit","path":"skills/pipeline-hic","license":"AGPL-3.0","commit":"36836c8725fd4d20d9c851ce314f5151aea5f57c","subtreeSha":"9B3D93461C8FCA8376ED946D22ECDB20573F056B051BD7A1F82181C1AD5E8F56","lastSyncedAt":"2026-09-29T20:56:54.045383Z"},"reviewedAt":"2026-09-22T13:30:50.41976Z","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow."},"install":[{"target":"skills-cli","command":"npx skills add https://github.com/ammawla/encode-toolkit/tree/main/skills/pipeline-hic"},{"target":"claude-code","command":"claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install ammawla-encode-toolkit@llmmart"},{"target":"git","command":"git clone https://github.com/ammawla/encode-toolkit.git"}]}