{"slug":"hic-aggregation","title":"hic-aggregation","summary":"Build comprehensive chromatin contact maps by aggregating Hi-C loop calls (BEDPE) across multiple ENCODE experiments, donors, and labs. Use when the user wants to answer \"what regions are in 3D contact in my tissue?\" by creating a union catalog of chromatin loops. Handles resolut","platform":"Claude","tags":[],"authorName":"LLM Mart","authorSlug":"llm-mart","score":0,"source":"github","price":null,"verified":false,"createdAt":"2026-09-22T13:25:34.440435Z","repo":{"url":"https://github.com/ammawla/encode-toolkit","stars":21,"forks":5,"license":"AGPL-3.0","updatedAt":"2026-09-27T01:22:42Z"},"bodyHtml":"<hr>\n<h2>name: hic-aggregation\ndescription: Build comprehensive chromatin contact maps by aggregating Hi-C loop calls (BEDPE) across multiple ENCODE experiments, donors, and labs. Use when the user wants to answer \"what regions are in 3D contact in my tissue?\" by creating a union catalog of chromatin loops. Handles resolution-aware anchor matching, cross-lab variation, and Hi-C-specific quality metrics.</h2>\n<h1>Aggregate Hi-C Chromatin Contacts Across Studies</h1>\n<h2>When to Use</h2>\n<ul>\n<li>User wants to build a comprehensive catalog of chromatin loops from multiple Hi-C experiments</li>\n<li>User asks \"what regions are in 3D contact in my tissue?\" or \"aggregate loop calls across donors\"</li>\n<li>User needs a union catalog of BEDPE loops with resolution-aware anchor matching</li>\n<li>User wants to identify high-confidence loops supported by multiple experiments</li>\n<li>Example queries: \"aggregate Hi-C loops for K562\", \"combine chromatin contacts across labs\", \"find consensus TAD boundaries in liver\"</li>\n</ul>\n<p>Build a comprehensive catalog of chromatin loops for a tissue/cell type by merging BEDPE loop calls from multiple ENCODE Hi-C experiments.</p>\n<h2>Scientific Rationale</h2>\n<p><strong>The question</strong>: \"What regions are in 3D physical contact in my tissue?\"</p>\n<p>Like histone marks and accessibility, chromatin loops are a <strong>detection question</strong>. If a loop between Region A and Region B is detected in one donor but not another, the contact is still real — individual variation, sequencing depth, and computational resolution explain absence. We want the <strong>union of all detected contacts</strong>.</p>\n<h3>Key Concepts</h3>\n<p><strong>Hi-C data</strong> measures pairwise chromatin interactions genome-wide. After processing:</p>\n<ul>\n<li><strong>Contact matrix</strong> (<code>.hic</code> file): Genome-wide interaction frequencies at multiple resolutions</li>\n<li><strong>Loop calls</strong> (BEDPE): Statistically significant point interactions (loops) identified by algorithms like HICCUPS or Juicer</li>\n<li><strong>TAD boundaries</strong>: Topologically associating domain boundaries</li>\n<li><strong>Compartments</strong>: A/B compartment assignments</li>\n</ul>\n<p><strong>BEDPE format</strong> (Paired-End BED):</p>\n<pre><code>chr1  start1  end1  chr2  start2  end2  name  score  strand1  strand2\n</code></pre>\n<p>Each row represents a contact between two genomic anchor regions.</p>\n<h3>Literature Support</h3>\n<ul>\n<li><strong>Loop Catalog</strong> (Reyna et al. 2025, Nucleic Acids Research): Created a union catalog of 4.19M unique loops across 1,089 Hi-C datasets. Demonstrated that union approach captures tissue-specific and constitutive loops. Used resolution-aware merging at 5kb, 10kb, and 25kb bins.</li>\n<li><strong>AQuA Tools</strong> (Chakraborty et al. 2025): Toolkit for BEDPE intersection, union, and annotation. Handles paired-region arithmetic.</li>\n<li><strong>mariner</strong> (Flores et al. 2024, Bioinformatics): R/Bioconductor package for BEDPE manipulation including merging loops across experiments with configurable anchor tolerance.</li>\n<li><strong>ENCODE Phase 3</strong> (Gorkin et al. 2020, Nature, 301 citations): Integrated Hi-C data across tissues to define regulatory loops connecting enhancers to promoters.</li>\n<li><strong>ENCODE Blacklist</strong> (Amemiya et al. 2019, Scientific Reports, 1,372 citations): Problematic genomic regions to filter from loop anchors. <a href=\"https://doi.org/10.1038/s41598-019-45839-z\">DOI</a></li>\n<li><strong>Mustache</strong> (Roayaei Ardakany et al. 2020, Genome Biology, 165 citations): Multi-scale loop caller that recovers more validated loops than HICCUPS. Different callers produce discordant loop sets.</li>\n<li><strong>Wolff et al. 2022</strong> (GigaScience): Benchmark showing loop callers intersect by <strong>~50% at most</strong> — critical context for why union approach is necessary.</li>\n</ul>\n<h2>Step 1: Find All Available Hi-C Data</h2>\n<pre><code>encode_search_experiments(\n    assay_title=\"Hi-C\",\n    organ=\"pancreas\",           # user's tissue of interest\n    biosample_type=\"tissue\",\n    limit=100\n)\n</code></pre>\n<p>Present a summary to the user:</p>\n<ul>\n<li>Total Hi-C experiments</li>\n<li>Labs represented</li>\n<li>Unique donors/biosamples</li>\n<li>Resolution(s) available (check experiment metadata)</li>\n</ul>\n<p>Use <code>encode_get_facets</code> to check availability:</p>\n<pre><code>encode_get_facets(assay_title=\"Hi-C\", organ=\"pancreas\")\n</code></pre>\n<p><strong>Note</strong>: Hi-C data is computationally expensive to produce, so there are typically fewer experiments per tissue than ChIP-seq or ATAC-seq. Even 2-3 experiments can be valuable for union catalogs.</p>\n<h2>Step 2: Quality-Gate Each Experiment</h2>\n<pre><code>encode_get_experiment(accession=\"ENCSR...\")\n</code></pre>\n<h3>Hi-C Quality Checks</h3>\n<ul>\n<li>Audit status: no ERROR flags</li>\n<li><strong>Sequencing depth</strong>: 400M+ valid read pairs for loop calling (ENCODE standard)</li>\n<li><strong>Cis/trans ratio</strong>: &gt;60% cis contacts expected (low cis suggests noisy library)</li>\n<li><strong>Hi-C-specific QC</strong>: Library complexity, PCR duplicate rate</li>\n<li>Has loop calls (BEDPE output) — not all Hi-C experiments have called loops</li>\n<li>Resolution: at least 5-10kb resolution for loop detection</li>\n</ul>\n<h3>Include if:</h3>\n<ul>\n<li>Has BEDPE loop calls at consistent resolution</li>\n<li>Passes ENCODE audit (no ERROR flags)</li>\n<li>Adequate sequencing depth for loop resolution</li>\n</ul>\n<h3>Exclude if:</h3>\n<ul>\n<li>ERROR audit flags</li>\n<li>Only contact matrices without loop calls</li>\n<li>Very low sequencing depth (&lt;200M valid pairs — insufficient for loop calling)</li>\n</ul>\n<p>Track all included experiments:</p>\n<pre><code>encode_track_experiment(accession=\"ENCSR...\")\n</code></pre>\n<h2>Step 3: Download Loop Call Files</h2>\n<p>For each experiment, get BEDPE loop calls:</p>\n<pre><code># Search for loop/interaction files\nencode_list_files(\n    experiment_accession=\"ENCSR...\",\n    file_format=\"bedpe\",\n    assembly=\"GRCh38\"\n)\n\n# Or ask for loop calls by output type\nencode_list_files(\n    experiment_accession=\"ENCSR...\",\n    output_type=\"loops\",\n    assembly=\"GRCh38\"\n)\n\n# Or contact domains\nencode_list_files(\n    experiment_accession=\"ENCSR...\",\n    output_type=\"contact domains\",\n    assembly=\"GRCh38\"\n)\n</code></pre>\n<p><strong>File selection priority:</strong></p>\n<ol>\n<li><strong>Chromatin interactions</strong> (loop calls from HICCUPS or similar)</li>\n<li><strong>Contact domains</strong> (TADs — different analysis, handle separately)</li>\n<li><strong>Replicated loops</strong> (if available)</li>\n</ol>\n<p>Prefer <code>preferred_default=True</code> files when available.</p>\n<pre><code>encode_download_files(\n    file_accessions=[\"ENCFF...\", ...],\n    download_dir=\"/path/to/data/hic_loops\",\n    organize_by=\"flat\"\n)\n</code></pre>\n<p>Validate the downloaded BEDPE files before filtering. The report gives the detected\nanchor resolution, which Step 4 needs; gzipped inputs are read directly.</p>\n<pre><code>python3 scripts/validate_loops.py sample.bedpe [--min-distance 20000] [--expected-resolution 10000]\n</code></pre>\n<h2>Step 4: Understanding Hi-C Resolution and Anchors</h2>\n<h3>Critical: Resolution-Aware Processing</h3>\n<p>Hi-C loop anchors are <strong>binned regions</strong>, not precise positions. The resolution determines anchor size:</p>\n<table>\n<thead>\n<tr>\n<th>Resolution</th>\n<th>Anchor Width</th>\n<th>Best For</th>\n<th>Typical Loop Count</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>5 kb</td>\n<td>5,000 bp</td>\n<td>Fine-scale promoter-enhancer loops</td>\n<td>More loops</td>\n</tr>\n<tr>\n<td>10 kb</td>\n<td>10,000 bp</td>\n<td>Standard analysis</td>\n<td>Moderate</td>\n</tr>\n<tr>\n<td>25 kb</td>\n<td>25,000 bp</td>\n<td>Large-scale domain contacts</td>\n<td>Fewer loops</td>\n</tr>\n</tbody>\n</table>\n<p><strong>All loops being merged must be at the same resolution</strong>, or anchors must be harmonized to a common resolution.</p>\n<h3>Harmonizing Resolution</h3>\n<p>If experiments have loops called at different resolutions:</p>\n<pre><code># Expand 5kb anchors to 10kb resolution\nawk -v res=10000 'BEGIN{OFS=\"\\t\"} {\n    # Bin anchor 1\n    bin1_start = int($2/res) * res\n    bin1_end = bin1_start + res\n    # Bin anchor 2\n    bin2_start = int($5/res) * res\n    bin2_end = bin2_start + res\n    print $1, bin1_start, bin1_end, $4, bin2_start, bin2_end, $7, $8, $9, $10\n}' fine_res_loops.bedpe &gt; harmonized_loops.bedpe\n</code></pre>\n<h2>Step 5: Per-Sample Filtering</h2>\n<h3>5a. ENCODE Blocklist Filtering (Amemiya et al. 2019)</h3>\n<p>Remove loops with anchors in artifact-prone regions (download from <a href=\"https://github.com/Boyle-Lab/Blacklist/blob/master/lists/hg38-blacklist.v2.bed.gz\">https://github.com/Boyle-Lab/Blacklist/blob/master/lists/hg38-blacklist.v2.bed.gz</a>):</p>\n<pre><code># Filter loops where EITHER anchor overlaps a blocklist region\ngunzip -k hg38-blacklist.v2.bed.gz\n\n# First, extract anchor 1 and anchor 2 as separate BED files\nawk 'BEGIN{OFS=\"\\t\"} {print $1,$2,$3,NR}' sample.bedpe &gt; anchors1.bed\nawk 'BEGIN{OFS=\"\\t\"} {print $4,$5,$6,NR}' sample.bedpe &gt; anchors2.bed\n\n# Find anchor rows NOT in blocklist\nbedtools intersect -a anchors1.bed -b hg38-blacklist.v2.bed -v | cut -f4 &gt; clean_rows_1.txt\nbedtools intersect -a anchors2.bed -b hg38-blacklist.v2.bed -v | cut -f4 &gt; clean_rows_2.txt\n\n# Keep only rows where BOTH anchors pass\ncomm -12 &lt;(sort clean_rows_1.txt) &lt;(sort clean_rows_2.txt) &gt; clean_rows.txt\nawk 'NR==FNR{a[$1];next} FNR in a' clean_rows.txt sample.bedpe &gt; sample.filtered.bedpe\n</code></pre>\n<h3>5b. Score Filtering</h3>\n<p>Filter by interaction score/significance:</p>\n<pre><code># If BEDPE has a score column (col 8), filter to significant interactions\n# Keep top 75% by score (true distribution quantile, not range-based)\nTOTAL=$(wc -l &lt; sample.filtered.bedpe)\nLINE_25=$(echo \"$TOTAL\" | awk '{printf \"%d\", $1 * 0.25}')\nTHRESHOLD=$(sort -k8,8n sample.filtered.bedpe | awk -v line=\"$LINE_25\" 'NR==line{print $8}')\nawk -v t=\"$THRESHOLD\" '$8 &gt;= t' sample.filtered.bedpe &gt; sample.qfiltered.bedpe\n</code></pre>\n<h3>5c. Remove Self-Ligation Artifacts</h3>\n<p>Loops where both anchors are very close are likely artifacts:</p>\n<pre><code># Remove loops where anchors are on same chromosome and &lt; 20kb apart\nawk '{\n    if ($1 != $4) print $0;  # inter-chromosomal: keep (rare but real)\n    else if (($5 - $3) &gt;= 20000) print $0;  # &gt; 20kb apart: keep\n}' sample.qfiltered.bedpe &gt; sample.clean.bedpe\n</code></pre>\n<h2>Step 6: Union Merge of Loops</h2>\n<h3>The Paired-Region Matching Problem</h3>\n<p>Unlike peaks (single regions), loops are <strong>pairs of regions</strong>. Two loops match if <strong>both anchors overlap</strong>:</p>\n<pre><code>Loop 1:  [anchor1A]--------[anchor1B]\nLoop 2:    [anchor2A]------[anchor2B]\n</code></pre>\n<p>These should merge if anchor1A overlaps anchor2A AND anchor1B overlaps anchor2B.</p>\n<h3>Method A: bedtools pairToPair (Recommended for simple union)</h3>\n<pre><code># Concatenate all filtered loops\ncat sample1.clean.bedpe sample2.clean.bedpe ... &gt; all_loops.bedpe\n\n# Sort by anchor 1 coordinates\nsort -k1,1 -k2,2n -k4,4 -k5,5n all_loops.bedpe &gt; all_loops.sorted.bedpe\n\n# Use a custom merge approach:\n# 1. Bin anchors to resolution, creating a loop ID\n# 2. Group by loop ID\n# 3. Count support\n\nawk -v res=10000 'BEGIN{OFS=\"\\t\"} {\n    # Create binned anchor coordinates as loop identifier\n    a1_bin = $1 \":\" int($2/res)*res\n    a2_bin = $4 \":\" int($5/res)*res\n    # Canonical order (smaller coordinate first) to handle orientation\n    if (a1_bin &lt; a2_bin) loop_id = a1_bin \"-\" a2_bin\n    else loop_id = a2_bin \"-\" a1_bin\n    print loop_id, $0\n}' all_loops.sorted.bedpe | \\\nsort -k1,1 | \\\nawk 'BEGIN{OFS=\"\\t\"} {\n    if ($1 != prev_id) {\n        if (NR &gt; 1) print chr1, start1, end1, chr2, start2, end2, count, max_score\n        prev_id = $1\n        chr1=$2; start1=$3; end1=$4; chr2=$5; start2=$6; end2=$7\n        count = 1; max_score = $9\n    } else {\n        count++\n        if ($9 &gt; max_score) max_score = $9\n        # Expand anchors to encompass all overlapping calls\n        if ($3 &lt; start1) start1 = $3\n        if ($4 &gt; end1) end1 = $4\n        if ($6 &lt; start2) start2 = $6\n        if ($7 &gt; end2) end2 = $7\n    }\n} END {\n    print chr1, start1, end1, chr2, start2, end2, count, max_score\n}' &gt; union_loops.bedpe\n</code></pre>\n<h3>Method B: Resolution-Binned Approach (Loop Catalog method)</h3>\n<p>Following the Loop Catalog (Reyna et al. 2025) approach:</p>\n<pre><code># Bin all loop anchors to a fixed resolution\nawk -v res=10000 'BEGIN{OFS=\"\\t\"} {\n    a1_start = int($2/res) * res\n    a1_end = a1_start + res\n    a2_start = int($5/res) * res\n    a2_end = a2_start + res\n    # Canonical ordering\n    if ($1 &lt; $4 || ($1 == $4 &amp;&amp; a1_start &lt;= a2_start))\n        print $1, a1_start, a1_end, $4, a2_start, a2_end\n    else\n        print $4, a2_start, a2_end, $1, a1_start, a1_end\n}' all_loops.sorted.bedpe | \\\nsort -u | \\\nsort -k1,1 -k2,2n -k4,4 -k5,5n | \\\nuniq -c | \\\nawk 'BEGIN{OFS=\"\\t\"} {print $2,$3,$4,$5,$6,$7,$1}' &gt; union_loops_binned.bedpe\n# Columns: chr1, start1, end1, chr2, start2, end2, n_supporting_samples\n</code></pre>\n<h3>Method C: Using Specialized Tools</h3>\n<p><strong>mariner</strong> (R/Bioconductor):</p>\n<pre><code>library(mariner)\n# Read BEDPE files as GInteractions\nloops &lt;- lapply(bedpe_files, read.table)\n# Convert to GInteractions and merge\ngi &lt;- as_ginteractions(loops)\nmerged &lt;- mergePairs(gi, radius = 10000)  # 10kb tolerance\n</code></pre>\n<p><strong>AQuA Tools</strong> (Python):</p>\n<pre><code># BEDPE union with anchor overlap tolerance\naqua bedpe-union -i sample1.bedpe sample2.bedpe -o union.bedpe --slop 5000\n</code></pre>\n<h2>Step 7: Confidence Annotation</h2>\n<p>Given N total experiments:</p>\n<table>\n<thead>\n<tr>\n<th>Confidence</th>\n<th>Criteria</th>\n<th>Interpretation</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>High</strong></td>\n<td>Detected in &gt;=50% of samples</td>\n<td>Constitutive loop, present across individuals</td>\n</tr>\n<tr>\n<td><strong>Supported</strong></td>\n<td>Detected in 2+ samples</td>\n<td>Likely real, some variation</td>\n</tr>\n<tr>\n<td><strong>Singleton</strong></td>\n<td>Detected in 1 sample only</td>\n<td>May be individual-specific or depth-dependent</td>\n</tr>\n</tbody>\n</table>\n<pre><code>awk -v N=4 '{\n    if ($7 &gt;= N*0.5) conf=\"HIGH\";\n    else if ($7 &gt;= 2) conf=\"SUPPORTED\";\n    else conf=\"SINGLETON\";\n    print $0\"\\t\"conf\"\\t\"$7\"/\"N\n}' union_loops_binned.bedpe &gt; union_loops.annotated.bedpe\n</code></pre>\n<p><strong>Context for singletons</strong>: Hi-C loop detection is very sensitive to sequencing depth. Many singletons may simply be under-powered in other samples rather than biologically absent. The Loop Catalog found that a union approach captures ~3x more loops than any individual experiment.</p>\n<h2>Step 8: Separate Analysis for TADs and Compartments</h2>\n<p><strong>TAD boundaries</strong> and <strong>A/B compartments</strong> require different aggregation than loops:</p>\n<h3>TAD Boundaries</h3>\n<p>TAD boundaries are single genomic positions. Aggregate like narrow peaks:</p>\n<pre><code># Extract TAD boundary BED from contact domain files\n# Each boundary is a narrow region\ncat tad_boundaries_sample*.bed | \\\nbedtools sort -i - | \\\nbedtools merge -i - -d 40000 -c 1 -o count &gt; union_tad_boundaries.bed\n# 40kb gap tolerance because TAD boundaries are resolution-dependent\n</code></pre>\n<h3>A/B Compartments</h3>\n<p>Compartment calls (eigenvector sign at each bin) should be aggregated by majority vote:</p>\n<pre><code># For each resolution bin, assign A or B based on majority of samples\n# This is more complex and typically done in R/Python\n</code></pre>\n<h2>Step 9: Log Provenance</h2>\n<pre><code>encode_log_derived_file(\n    file_path=\"/path/to/union_loops.annotated.bedpe\",\n    source_accessions=[\"ENCSR...\", \"ENCSR...\", ...],\n    description=\"Union chromatin loops across N pancreas Hi-C experiments\",\n    file_type=\"aggregated_loops\",\n    tool_used=\"bedtools + custom merge at 10kb resolution\",\n    parameters=\"blocklist filtered, score &gt;= 25th pctl, self-ligation &gt;= 20kb removed, 10kb resolution binning\"\n)\n</code></pre>\n<h2>Step 10: Summary Statistics</h2>\n<p>Report to the user:</p>\n<ul>\n<li>Total input experiments: N</li>\n<li>Experiments passing QC: M</li>\n<li>Resolution used: Xkb</li>\n<li>Total loops before merge: X</li>\n<li>Union loops after merge: Y</li>\n<li>High-confidence loops: Z (≥50% support)</li>\n<li>Supported loops: W (2+ support)</li>\n<li>Singleton loops: V (1 sample only)</li>\n<li>Distance distribution: median and range of loop sizes (anchor-to-anchor)</li>\n<li>Inter-chromosomal loops: count (expect very few)</li>\n</ul>\n<h2>Pitfalls Specific to Hi-C Data</h2>\n<ol>\n<li><p><strong>Resolution mismatch</strong>: Loop calls at 5kb vs 25kb resolution will have very different anchor sizes. Always harmonize to a common resolution before merging.</p>\n</li>\n<li><p><strong>Sequencing depth sensitivity</strong>: Loop calling requires deep sequencing (400M+ valid pairs). Shallowly sequenced experiments will call far fewer loops — this is under-detection, not absence.</p>\n</li>\n<li><p><strong>Algorithm differences are LARGE</strong>: Wolff et al. 2022 (GigaScience) found that HICCUPS, Mustache, Fit-Hi-C, and HiCExplorer loop callers <strong>intersect by ~50% at most</strong>. Mustache tends to recover more validated loops (Roayaei Ardakany et al. 2020). If mixing callers, note this in provenance — and this discordance is itself a reason to prefer the union approach.</p>\n</li>\n<li><p><strong>Orientation matters</strong>: BEDPE anchors should be canonically ordered (anchor1 &lt; anchor2 by genomic coordinate) before merging to avoid duplicate counting.</p>\n</li>\n<li><p><strong>Inter-chromosomal contacts</strong>: These are rare but real. Handle separately — they cannot be distance-filtered.</p>\n</li>\n<li><p><strong>Distance distribution</strong>: Most loops are 100kb-2Mb. Very short-range contacts (&lt;20kb) are often noise from undigested chromatin. Very long-range (&gt;10Mb) are rare.</p>\n</li>\n<li><p><strong>Do NOT mix assemblies</strong>: All files must be GRCh38 or all hg19. Hi-C resolution binning makes liftOver of loops particularly error-prone.</p>\n</li>\n<li><p><strong>TADs vs loops</strong>: These are different features. TADs are domains (regions), loops are point contacts (pairs). Do not mix them in the same union.</p>\n</li>\n<li><p><strong>Micro-C as complement</strong>: Micro-C achieves higher resolution than Hi-C and can detect sub-TAD loops. Treat Micro-C loops as compatible with Hi-C loops in a union (Mustache works on both).</p>\n</li>\n</ol>\n<h2>Walkthrough: Building a Cross-Tissue Loop Catalog for the MYC Locus</h2>\n<p><strong>Goal</strong>: Aggregate Hi-C chromatin loops across tissues to identify conserved and tissue-specific 3D contacts at the MYC gene locus.\n<strong>Context</strong>: Cancer research — MYC is regulated by distal enhancers via chromatin looping.</p>\n<h3>Step 1: Find Hi-C experiments across tissues</h3>\n<pre><code>encode_search_experiments(assay_title=\"Hi-C\", organism=\"Homo sapiens\", limit=50)\n</code></pre>\n<p>Expected output:</p>\n<pre><code>{\n  \"results\": [\n    {\"accession\": \"ENCSR000AKA\", \"assay_title\": \"Hi-C\", \"biosample_summary\": \"GM12878\", \"status\": \"released\"},\n    {\"accession\": \"ENCSR489OCU\", \"assay_title\": \"Hi-C\", \"biosample_summary\": \"K562\", \"status\": \"released\"},\n    {\"accession\": \"ENCSR382RFU\", \"assay_title\": \"Hi-C\", \"biosample_summary\": \"liver\", \"status\": \"released\"}\n  ],\n  \"total\": 89,\n  \"limit\": 50,\n  \"offset\": 0,\n  \"has_more\": true,\n  \"next_offset\": 50\n}\n</code></pre>\n<p><strong>Interpretation</strong>: 89 Hi-C experiments available. Select 5–10 spanning diverse tissue types for cross-tissue comparison.</p>\n<h3>Step 2: List loop files for each experiment</h3>\n<pre><code>encode_list_files(experiment_accession=\"ENCSR000AKA\", file_format=\"bedpe\", assembly=\"GRCh38\")\n</code></pre>\n<p>Expected output (a JSON array of file records; fields abridged):</p>\n<pre><code>[\n  {\"accession\": \"ENCFF001ABC\", \"output_type\": \"contact domains\", \"file_format\": \"bedpe\", \"file_size_human\": \"2.4 MB\"},\n  {\"accession\": \"ENCFF002DEF\", \"output_type\": \"loops\", \"file_format\": \"bedpe\", \"file_size_human\": \"1.8 MB\"}\n]\n</code></pre>\n<p><strong>Interpretation</strong>: Use \"loops\" files for loop aggregation. Contact domains are TADs, not loops.</p>\n<h3>Step 3: Download loop files</h3>\n<pre><code>encode_download_files(file_accessions=[\"ENCFF002DEF\", \"ENCFF003GHI\", \"ENCFF004JKL\"], download_dir=\"/data/hic_loops\")\n</code></pre>\n<p>Expected output (one of the three <code>downloaded</code> entries shown):</p>\n<pre><code>{\n  \"downloaded\": [\n    {\n      \"accession\": \"ENCFF002DEF\",\n      \"file_path\": \"/data/hic_loops/ENCFF002DEF.bedpe\",\n      \"file_size\": 1887436,\n      \"file_size_human\": \"1.8 MB\",\n      \"success\": true,\n      \"error\": \"\",\n      \"md5_verified\": true\n    }\n  ],\n  \"errors\": [],\n  \"summary\": {\n    \"total_requested\": 3,\n    \"successful\": 3,\n    \"failed\": 0,\n    \"total_size\": 5872025,\n    \"total_size_human\": \"5.6 MB\"\n  }\n}\n</code></pre>\n<h3>Step 4: Aggregate loops with resolution-aware anchor matching</h3>\n<p>Apply union merge across tissues:</p>\n<ul>\n<li>Expand loop anchors by ±resolution (e.g., ±5kb for 5kb resolution data)</li>\n<li>Merge overlapping anchors using bedtools pairToPair</li>\n<li>Assign tissue support counts to each union loop</li>\n<li>Filter: require ≥2 tissue support for conserved loops</li>\n</ul>\n<h3>Step 5: Filter to MYC locus</h3>\n<pre><code># MYC locus: chr8:127,700,000-128,000,000\nawk '$1==\"chr8\" &amp;&amp; $2&gt;=127700000 &amp;&amp; $3&lt;=128000000' union_loops.bedpe &gt; myc_loops.bedpe\n</code></pre>\n<p><strong>Interpretation</strong>: Loops anchored at the MYC promoter connecting to distal enhancers. Conserved loops (≥3 tissues) likely represent fundamental regulatory architecture; tissue-specific loops may drive context-dependent MYC activation.</p>\n<h3>Integration with downstream skills</h3>\n<ul>\n<li>Feed loop anchors into → <strong>peak-annotation</strong> for gene assignment at anchor regions</li>\n<li>Overlay with → <strong>histone-aggregation</strong> H3K27ac peaks to identify active enhancer-promoter loops</li>\n<li>Cross-reference loop-disrupting variants via → <strong>variant-annotation</strong></li>\n<li>Visualize in → <strong>ucsc-browser</strong> as interaction tracks</li>\n</ul>\n<h2>Code Examples</h2>\n<h3>1. Survey available Hi-C data by tissue</h3>\n<pre><code>encode_get_facets(assay_title=\"Hi-C\", organism=\"Homo sapiens\")\n</code></pre>\n<p>Expected output (facet field names are the top-level keys):</p>\n<pre><code>{\n  \"biosample_ontology.organ_slims\": [\n    {\"term\": \"brain\", \"count\": 24},\n    {\"term\": \"blood\", \"count\": 15}\n  ]\n}\n</code></pre>\n<h3>2. Get details for a specific Hi-C experiment</h3>\n<pre><code>encode_get_experiment(accession=\"ENCSR000AKA\")\n</code></pre>\n<p>Expected output (fields abridged):</p>\n<pre><code>{\n  \"accession\": \"ENCSR000AKA\",\n  \"assay_title\": \"Hi-C\",\n  \"biosample_summary\": \"GM12878\",\n  \"assembly\": [\"GRCh38\"],\n  \"bio_replicate_count\": 2,\n  \"status\": \"released\",\n  \"lab\": \"Erez Lieberman Aiden, Baylor\",\n  \"audit_error_count\": 0,\n  \"audit_not_compliant_count\": 0,\n  \"audit_warning_count\": 1,\n  \"audit_internal_action_count\": 0\n}\n</code></pre>\n<h3>3. Compare loop sets between two cell types</h3>\n<p>Both experiments must already be tracked; otherwise the tool returns <code>{\"error\": \"Experiment ... not tracked. Track it first.\"}</code>.</p>\n<pre><code>encode_compare_experiments(accession1=\"ENCSR000AKA\", accession2=\"ENCSR489OCU\")\n</code></pre>\n<p>Expected output:</p>\n<pre><code>{\n  \"experiment_1\": {\"accession\": \"ENCSR000AKA\", \"assay\": \"Hi-C\", \"biosample\": \"GM12878\"},\n  \"experiment_2\": {\"accession\": \"ENCSR489OCU\", \"assay\": \"Hi-C\", \"biosample\": \"K562\"},\n  \"verdict\": \"COMPATIBLE_WITH_CAVEATS\",\n  \"recommendation\": \"These experiments can be compared, but the warnings should be addressed in your analysis.\",\n  \"compatible_aspects\": [\n    \"Same organism: Homo sapiens\",\n    \"Same assembly: GRCh38\",\n    \"Same assay: Hi-C\",\n    \"Same biosample type: cell line\"\n  ],\n  \"issues\": [],\n  \"warnings\": [\n    \"Different labs: Erez Lieberman Aiden, Baylor vs Job Dekker, UMass. Batch effects possible.\"\n  ]\n}\n</code></pre>\n<h2>Integration</h2>\n<table>\n<thead>\n<tr>\n<th>This skill produces...</th>\n<th>Feed into...</th>\n<th>Purpose</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Union loop catalog (BEDPE)</td>\n<td><strong>peak-annotation</strong></td>\n<td>Assign genes to loop anchors</td>\n</tr>\n<tr>\n<td>Conserved loop coordinates</td>\n<td><strong>histone-aggregation</strong></td>\n<td>Overlay H3K27ac at anchors to find active enhancer-promoter loops</td>\n</tr>\n<tr>\n<td>Tissue-specific loops</td>\n<td><strong>accessibility-aggregation</strong></td>\n<td>Check if loop anchors overlap open chromatin</td>\n</tr>\n<tr>\n<td>Loop anchor BED intervals</td>\n<td><strong>variant-annotation</strong></td>\n<td>Find GWAS/clinical variants disrupting loop anchors</td>\n</tr>\n<tr>\n<td>Loop anchor coordinates</td>\n<td><strong>liftover-coordinates</strong></td>\n<td>Convert hg19 loops to GRCh38</td>\n</tr>\n<tr>\n<td>Aggregated loop statistics</td>\n<td><strong>visualization-workflow</strong></td>\n<td>Generate loop frequency heatmaps</td>\n</tr>\n<tr>\n<td>Loop-gene assignments</td>\n<td><strong>disease-research</strong></td>\n<td>Connect loop disruptions to disease phenotypes</td>\n</tr>\n</tbody>\n</table>\n<h2>Related Skills</h2>\n<ul>\n<li><strong>histone-aggregation</strong>: Loop anchors often overlap with H3K27ac/H3K4me1 peaks — integrate with histone union sets to annotate loop function</li>\n<li><strong>accessibility-aggregation</strong>: Loop anchors frequently coincide with accessible chromatin — validate loops by requiring anchor accessibility</li>\n<li><strong>regulatory-elements</strong>: Use loops to connect distal enhancers (H3K27ac) to target promoters (H3K4me3)</li>\n<li><strong>epigenome-profiling</strong>: Loops add 3D context to 1D chromatin state maps</li>\n<li><strong>pipeline-hic</strong>: Process raw Hi-C data through the full ENCODE-aligned pipeline</li>\n<li><strong>batch-analysis</strong>: Batch processing workflows for systematic Hi-C loop aggregation</li>\n<li><strong>publication-trust</strong>: Verify literature claims backing analytical decisions</li>\n</ul>\n<h2>Presenting Results</h2>\n<ul>\n<li>Present aggregated loops as: chr | anchor1_start | anchor1_end | anchor2_start | anchor2_end | sample_count | resolution. Show loop statistics. Suggest: \"Would you like to check if any GWAS variants overlap loop anchors?\"</li>\n</ul>\n<h2>For the request: \"$ARGUMENTS\"</h2>\n","files":[{"path":"references/literature.md","sizeBytes":7356,"isText":true},{"path":"references/loop-caller-comparison.md","sizeBytes":6845,"isText":true},{"path":"scripts/validate_loops.py","sizeBytes":11930,"isText":true},{"path":"SKILL.md","sizeBytes":23379,"isText":true}],"reviewScore":null,"reviewSummary":null,"trust":{"provenance":"trusted-source-unreviewed","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow.","bodySource":null},"bodyLocked":false,"purchaseUrl":null,"sourceUrl":null,"report":{"provenance":"trusted-source-unreviewed","screen":{"ran":true,"outcome":"clean","suspicious":0,"notes":0,"hiddenCharacters":false},"virusScan":{"engine":"clamav","status":"clean","scannedAt":"2026-09-22T13:25:59.188332Z","sha256":"50DA749F8776646C91D7F4DDA8E43B569595969D1EE26906031AF8B23532F428","sizeBytes":18981},"review":null,"source":{"repositoryUrl":"https://github.com/ammawla/encode-toolkit","path":"plugin/skills/hic-aggregation","license":"AGPL-3.0","commit":"36836c8725fd4d20d9c851ce314f5151aea5f57c","subtreeSha":"DC2F069407FEE8F561C9C32CAF37086E685FBE373725D6C8F026746C5167C70A","lastSyncedAt":"2026-09-29T20:56:54.045383Z"},"reviewedAt":"2026-09-22T13:26:50.233209Z","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow."},"install":[{"target":"skills-cli","command":"npx skills add https://github.com/ammawla/encode-toolkit/tree/main/plugin/skills/hic-aggregation"},{"target":"claude-code","command":"claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install ammawla-encode-toolkit@llmmart"},{"target":"git","command":"git clone https://github.com/ammawla/encode-toolkit.git"}]}