{"slug":"liftover-coordinates-2","title":"liftover-coordinates","summary":"Convert genomic coordinates between assembly versions (GRCh37/hg19 to GRCh38/hg38, mm9 to mm10). Guides UCSC liftOver for BED files, CrossMap for VCF/bigWig, and handles unmapped regions with provenance logging.","platform":"Claude","tags":[],"authorName":"LLM Mart","authorSlug":"llm-mart","score":0,"source":"github","price":null,"verified":false,"createdAt":"2026-09-22T13:25:42.928061Z","repo":{"url":"https://github.com/ammawla/encode-toolkit","stars":21,"forks":5,"license":"AGPL-3.0","updatedAt":"2026-09-27T01:22:42Z"},"bodyHtml":"<hr>\n<h2>name: liftover-coordinates\ndescription: Convert genomic coordinates between assembly versions (GRCh37/hg19 to GRCh38/hg38, mm9 to mm10). Guides UCSC liftOver for BED files, CrossMap for VCF/bigWig, and handles unmapped regions with provenance logging.</h2>\n<h1>Convert Genomic Coordinates Between Assembly Versions</h1>\n<h2>When to Use</h2>\n<ul>\n<li>User needs to convert genomic coordinates between assemblies (hg19↔hg38, mm9↔mm10)</li>\n<li>User asks about \"liftover\", \"coordinate conversion\", \"assembly mismatch\", or \"CrossMap\"</li>\n<li>User has data in hg19/GRCh37 that needs conversion to GRCh38 (or vice versa) before integration</li>\n<li>User wants to use UCSC liftOver or CrossMap for BED, VCF, bigWig, or BAM files</li>\n<li>Example queries: \"convert my hg19 peaks to hg38\", \"liftover coordinates for integration with ENCODE\", \"my data is in mm9, how do I convert to mm10?\"</li>\n</ul>\n<p>Guide coordinate liftover between genome assemblies using UCSC liftOver, CrossMap, Ensembl REST API, and rtracklayer. Assembly conversion is one of the most common pitfalls in genomics — this skill provides the definitive workflow for safe, reproducible liftover with full provenance tracking.</p>\n<h2>Scientific Rationale</h2>\n<p><strong>The question</strong>: \"How do I safely convert my genomic coordinates from one assembly to another without losing data or introducing errors?\"</p>\n<p>Assembly conversion is referenced as a critical step in 10+ other ENCODE Toolkit skills because ENCODE spans multiple data releases: some experiments were processed against hg19/GRCh37, while most current data uses GRCh38/hg38. Combining data across assemblies without proper liftover is one of the most common and most dangerous errors in computational genomics — coordinates that look valid in both assemblies may refer to completely different genomic locations.</p>\n<h3>The Core Problem</h3>\n<p>Genome assemblies are updated to fix errors, fill gaps, add alternative haplotypes, and improve centromeric/telomeric sequence. Between hg19 and hg38, approximately 1,000 sequence gaps were closed, 8% of the genome was modified, and several regions were rearranged. A coordinate like chr17:41,197,694 in hg19 (BRCA1) maps to chr17:43,044,295 in GRCh38 — a shift of nearly 2 Mb. Using the wrong assembly silently produces incorrect results.</p>\n<h3>When to Liftover</h3>\n<p>Common scenarios requiring coordinate conversion:</p>\n<ul>\n<li><strong>Combining ENCODE data from different releases</strong>: Some hg19, some GRCh38 — must unify before intersection</li>\n<li><strong>Integrating GWAS Catalog results</strong>: Many GWAS hits are still reported in hg19/GRCh37 coordinates</li>\n<li><strong>Using gnomAD</strong>: gnomAD v4 uses GRCh38; older v2 datasets use GRCh37</li>\n<li><strong>Cross-species comparison</strong>: Mouse data across mm9/mm10/GRCm39</li>\n<li><strong>Legacy datasets</strong>: Published supplementary files often use older assemblies</li>\n<li><strong>ClinVar integration</strong>: Some ClinVar entries reference GRCh37 positions</li>\n<li><strong>GTEx cross-reference</strong>: GTEx v8 uses GRCh38, earlier versions used GRCh37</li>\n</ul>\n<h3>Literature Support</h3>\n<ul>\n<li><strong>Kent et al. 2002</strong> (Genome Research, ~5,000 citations): UCSC Genome Browser and the liftOver tool. The original chain/net alignment framework for coordinate conversion between genome assemblies. <a href=\"https://doi.org/10.1101/gr.229102\">DOI</a></li>\n<li><strong>Zhao et al. 2014</strong> (Bioinformatics, ~800 citations): CrossMap — a versatile tool for coordinate conversion between genome assemblies. Handles VCF, BAM, bigWig, GFF, and Wiggle formats that UCSC liftOver cannot process natively. <a href=\"https://doi.org/10.1093/bioinformatics/btt730\">DOI</a></li>\n<li><strong>Hinrichs et al. 2006</strong> (Nucleic Acids Research, ~1,200 citations): UCSC genome browser chain/net alignment methodology. Defines the reciprocal-best chain alignment that underpins coordinate conversion. <a href=\"https://doi.org/10.1093/nar/gkj144\">DOI</a></li>\n<li><strong>Kuhn et al. 2013</strong> (Nucleic Acids Research, ~600 citations): Assembly updates and the implications for re-annotation. Documents the biological impact of assembly changes on gene models and regulatory element coordinates. <a href=\"https://doi.org/10.1093/nar/gks1195\">DOI</a></li>\n<li><strong>Schneider et al. 2017</strong> (Genome Research, ~400 citations): GRCh38 improvements over GRCh37 — gap closures, centromere models, alternative haplotypes. Quantifies what changed and why liftover is necessary. <a href=\"https://doi.org/10.1101/gr.213611.116\">DOI</a></li>\n<li><strong>Amemiya et al. 2019</strong> (Scientific Reports, ~1,372 citations): ENCODE Blacklist regions — some blacklisted regions are assembly-specific. Liftover of blacklist files must use the correct version. <a href=\"https://doi.org/10.1038/s41598-019-45839-z\">DOI</a></li>\n</ul>\n<h2>Assembly Version Mapping</h2>\n<table>\n<thead>\n<tr>\n<th>Common Name</th>\n<th>UCSC Name</th>\n<th>NCBI/GRC Name</th>\n<th>Species</th>\n<th>Release Year</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>hg19</td>\n<td>hg19</td>\n<td>GRCh37</td>\n<td>Human</td>\n<td>2009</td>\n</tr>\n<tr>\n<td>hg38</td>\n<td>hg38</td>\n<td>GRCh38</td>\n<td>Human</td>\n<td>2013</td>\n</tr>\n<tr>\n<td>mm9</td>\n<td>mm9</td>\n<td>MGSCv37</td>\n<td>Mouse</td>\n<td>2007</td>\n</tr>\n<tr>\n<td>mm10</td>\n<td>mm10</td>\n<td>GRCm38</td>\n<td>Mouse</td>\n<td>2012</td>\n</tr>\n<tr>\n<td>mm39</td>\n<td>mm39</td>\n<td>GRCm39</td>\n<td>Mouse</td>\n<td>2020</td>\n</tr>\n</tbody>\n</table>\n<h3>Naming Convention Alert</h3>\n<p>The same assembly has different names depending on the source:</p>\n<ul>\n<li><strong>UCSC convention</strong>: <code>hg19</code>, <code>hg38</code>, <code>mm10</code> — used in filenames, chromosome prefixes (<code>chr1</code>)</li>\n<li><strong>NCBI/GRC convention</strong>: <code>GRCh37</code>, <code>GRCh38</code>, <code>GRCm38</code> — used in publications, Ensembl</li>\n<li><strong>Ensembl convention</strong>: Chromosomes without <code>chr</code> prefix (<code>1</code> instead of <code>chr1</code>)</li>\n</ul>\n<p>Always verify which naming convention your data uses. Mixing <code>chr1</code> (UCSC) with <code>1</code> (Ensembl) causes silent failures in bedtools intersection and peak overlap analysis.</p>\n<h2>Chain Files</h2>\n<p>Chain files encode the alignment between assemblies and are the essential input for liftover.</p>\n<h3>Source: UCSC (Recommended)</h3>\n<pre><code>https://hgdownload.soe.ucsc.edu/goldenPath/{from}/liftOver/{from}To{To}.over.chain.gz\n</code></pre>\n<p>Common chain files:</p>\n<table>\n<thead>\n<tr>\n<th>Conversion</th>\n<th>Chain File</th>\n<th>URL</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>hg19 to hg38</td>\n<td>hg19ToHg38.over.chain.gz</td>\n<td><code>https://hgdownload.soe.ucsc.edu/goldenPath/hg19/liftOver/hg19ToHg38.over.chain.gz</code></td>\n</tr>\n<tr>\n<td>hg38 to hg19</td>\n<td>hg38ToHg19.over.chain.gz</td>\n<td><code>https://hgdownload.soe.ucsc.edu/goldenPath/hg38/liftOver/hg38ToHg19.over.chain.gz</code></td>\n</tr>\n<tr>\n<td>mm9 to mm10</td>\n<td>mm9ToMm10.over.chain.gz</td>\n<td><code>https://hgdownload.soe.ucsc.edu/goldenPath/mm9/liftOver/mm9ToMm10.over.chain.gz</code></td>\n</tr>\n<tr>\n<td>mm10 to mm39</td>\n<td>mm10ToMm39.over.chain.gz</td>\n<td><code>https://hgdownload.soe.ucsc.edu/goldenPath/mm10/liftOver/mm10ToMm39.over.chain.gz</code></td>\n</tr>\n<tr>\n<td>mm10 to hg38</td>\n<td>mm10ToHg38.over.chain.gz</td>\n<td><code>https://hgdownload.soe.ucsc.edu/goldenPath/mm10/liftOver/mm10ToHg38.over.chain.gz</code></td>\n</tr>\n</tbody>\n</table>\n<h3>Source: Ensembl</h3>\n<pre><code>ftp://ftp.ensembl.org/pub/assembly_mapping/\n</code></pre>\n<p>Ensembl provides chain files for their coordinate system (without <code>chr</code> prefix). Useful when working with Ensembl VEP output or Ensembl gene annotations.</p>\n<h3>Source: NCBI Remap</h3>\n<p>NCBI Genome Remapping Service: <code>https://www.ncbi.nlm.nih.gov/genome/tools/remap</code></p>\n<ul>\n<li>Web-based and API access</li>\n<li>Handles complex remapping with alignment-based and annotation-based methods</li>\n<li>Useful for non-standard assemblies or patch-level conversions</li>\n</ul>\n<h3>Chain File Verification</h3>\n<p>Always verify chain file integrity after download:</p>\n<pre><code>wget https://hgdownload.soe.ucsc.edu/goldenPath/hg19/liftOver/hg19ToHg38.over.chain.gz\nmd5sum hg19ToHg38.over.chain.gz\n# Verify against UCSC md5sum.txt in the same directory\ngunzip -t hg19ToHg38.over.chain.gz  # Test archive integrity\n</code></pre>\n<h2>Tool Guide: UCSC liftOver (BED Files)</h2>\n<p>The standard tool for BED-format coordinate conversion.</p>\n<h3>Basic Usage</h3>\n<pre><code>liftOver input.bed hg19ToHg38.over.chain.gz output.bed unmapped.bed\n</code></pre>\n<h3>Parameters</h3>\n<table>\n<thead>\n<tr>\n<th>Parameter</th>\n<th>Default</th>\n<th>Description</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>-minMatch</code></td>\n<td>0.95</td>\n<td>Minimum ratio of bases that must remap (0.0–1.0)</td>\n</tr>\n<tr>\n<td><code>-minBlocks</code></td>\n<td>1</td>\n<td>Minimum number of alignment blocks</td>\n</tr>\n<tr>\n<td><code>-fudgeThick</code></td>\n<td>off</td>\n<td>If thickStart/thickEnd not mapped, use mapped region</td>\n</tr>\n<tr>\n<td><code>-multiple</code></td>\n<td>off</td>\n<td>Allow mapping to multiple output regions</td>\n</tr>\n<tr>\n<td><code>-minChainT</code></td>\n<td>0</td>\n<td>Minimum chain target coverage</td>\n</tr>\n<tr>\n<td><code>-minChainQ</code></td>\n<td>0</td>\n<td>Minimum chain query coverage</td>\n</tr>\n</tbody>\n</table>\n<h3>Recommended Settings by Data Type</h3>\n<table>\n<thead>\n<tr>\n<th>Data Type</th>\n<th><code>-minMatch</code></th>\n<th>Notes</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>SNP positions (1bp)</td>\n<td>0.95 (default)</td>\n<td>Point coordinates almost always map cleanly</td>\n</tr>\n<tr>\n<td>Narrow peaks (100–500bp)</td>\n<td>0.95 (default)</td>\n<td>Short regions map well</td>\n</tr>\n<tr>\n<td>Broad peaks (1–50kb)</td>\n<td>0.50–0.80</td>\n<td>Large regions may partially overlap rearrangements</td>\n</tr>\n<tr>\n<td>Regulatory elements</td>\n<td>0.90</td>\n<td>Balance between completeness and accuracy</td>\n</tr>\n<tr>\n<td>TAD boundaries (5–50kb)</td>\n<td>0.50</td>\n<td>Large-scale organization is approximate anyway</td>\n</tr>\n</tbody>\n</table>\n<h3>Handling narrowPeak Format</h3>\n<p>UCSC liftOver expects standard BED (3–12 columns). For narrowPeak files (BED6+4):</p>\n<pre><code># Step 1: Extract BED6 columns + preserve extra columns as name\nawk 'BEGIN{OFS=\"\\t\"} {print $1, $2, $3, $4, $5, $6, $7, $8, $9, $10}' input.narrowPeak &gt; input_full.bed\n\n# Step 2: Liftover (liftOver handles extra columns)\nliftOver input_full.bed hg19ToHg38.over.chain.gz output.bed unmapped.bed\n\n# Step 3: Verify column count is preserved\nawk '{print NF}' output.bed | sort -u\n</code></pre>\n<p><strong>Peak summit recalculation</strong>: After liftover, the summit position (column 10 in narrowPeak = offset from start) may no longer accurately represent the signal maximum. For critical analyses, re-calculate summits from signal data in the new assembly rather than relying on lifted summit positions.</p>\n<h3>Checking Unmapped Regions</h3>\n<pre><code># Count unmapped regions\nwc -l unmapped.bed  # Note: comment lines start with #\n\n# Calculate loss rate\ntotal=$(wc -l &lt; input.bed)\nunmapped=$(grep -v '^#' unmapped.bed | wc -l)\nloss_pct=$(echo \"scale=2; $unmapped * 100 / $total\" | bc)\necho \"Lost $unmapped of $total regions ($loss_pct%)\"\n\n# Investigate reasons for unmapping\ngrep '^#' unmapped.bed | sort | uniq -c | sort -rn\n# Common reasons:\n# \"Partially deleted in new\" — region spans a deletion\n# \"Deleted in new\" — region fully removed\n# \"Split in new\" — region maps to multiple locations\n</code></pre>\n<h2>Tool Guide: CrossMap (VCF, bigWig, BAM, GFF)</h2>\n<p>CrossMap (Zhao et al. 2014) handles file formats that UCSC liftOver cannot process natively.</p>\n<h3>VCF Conversion</h3>\n<pre><code>CrossMap vcf hg19ToHg38.over.chain.gz input.vcf hg38.fa output.vcf\n</code></pre>\n<p><strong>Critical VCF considerations</strong>:</p>\n<ul>\n<li>CrossMap updates coordinates AND checks REF alleles against the new reference</li>\n<li>Variants where the REF allele changes between assemblies are flagged</li>\n<li>Always re-validate variant calls after liftover using <code>bcftools norm</code></li>\n<li>Multi-allelic variants may need special handling</li>\n</ul>\n<pre><code># Post-liftover VCF validation\nbcftools norm -f hg38.fa -c ws output.vcf -o output.normalized.vcf 2&gt; norm_warnings.log\n# -c ws: warn about and set incorrect REF alleles\n</code></pre>\n<h3>bigWig Conversion</h3>\n<pre><code>CrossMap bigwig hg19ToHg38.over.chain.gz input.bw output.bw\n</code></pre>\n<p><strong>Signal track caveats</strong>:</p>\n<ul>\n<li>Resolution is reduced during conversion (interpolation at boundaries)</li>\n<li>Regions that split during liftover lose signal accuracy</li>\n<li>For quantitative analysis, re-generate signal tracks from re-aligned reads when possible</li>\n</ul>\n<h3>BAM Conversion</h3>\n<pre><code>CrossMap bam hg19ToHg38.over.chain.gz input.bam output.bam\n</code></pre>\n<p><strong>BAM liftover is generally NOT recommended</strong>:</p>\n<ul>\n<li>Read mapping quality is meaningless after coordinate shifting</li>\n<li>Paired-end relationships may break</li>\n<li>Duplicate marking becomes invalid</li>\n<li><strong>Best practice</strong>: Re-align from FASTQ to the new reference genome</li>\n</ul>\n<h3>GFF/GTF Conversion</h3>\n<pre><code>CrossMap gff hg19ToHg38.over.chain.gz input.gff output.gff\n</code></pre>\n<p>Useful for lifting gene annotations, but prefer downloading the native annotation for the target assembly from GENCODE or Ensembl.</p>\n<h2>Tool Guide: Ensembl REST API (Single Coordinates)</h2>\n<p>For programmatic conversion of individual coordinates without installing local tools.</p>\n<h3>API Endpoint</h3>\n<pre><code>GET https://rest.ensembl.org/map/human/GRCh37/{region}/GRCh38?content-type=application/json\n</code></pre>\n<h3>Example</h3>\n<pre><code>import requests\n\ndef liftover_ensembl(chrom, start, end, source=\"GRCh37\", target=\"GRCh38\", species=\"human\"):\n    \"\"\"Convert coordinates using Ensembl REST API.\"\"\"\n    region = f\"{chrom}:{start}..{end}:1\"\n    url = f\"https://rest.ensembl.org/map/{species}/{source}/{region}/{target}\"\n    headers = {\"Content-Type\": \"application/json\"}\n    response = requests.get(url, headers=headers)\n    if response.status_code == 200:\n        mappings = response.json()[\"mappings\"]\n        return mappings\n    return None\n\n# Example: BRCA1 region\nmappings = liftover_ensembl(\"17\", 41197694, 41276113)\nfor m in mappings:\n    mapped = m[\"mapped\"]\n    print(f\"  {mapped['seq_region_name']}:{mapped['start']}-{mapped['end']}\")\n</code></pre>\n<h3>Rate Limits</h3>\n<ul>\n<li>15 requests per second (without registered email)</li>\n<li>50 requests per second (with registered email in User-Agent header)</li>\n<li>For batch conversion, use local liftOver or CrossMap instead</li>\n</ul>\n<h3>Ensembl Chromosome Naming</h3>\n<p>Ensembl uses chromosomes WITHOUT <code>chr</code> prefix:</p>\n<ul>\n<li>Ensembl: <code>17:41197694-41276113</code></li>\n<li>UCSC: <code>chr17:41197694-41276113</code></li>\n</ul>\n<p>Convert between conventions:</p>\n<pre><code># Add 'chr' prefix (Ensembl to UCSC)\nsed 's/^/chr/' input.bed &gt; input_ucsc.bed\n\n# Remove 'chr' prefix (UCSC to Ensembl)\nsed 's/^chr//' input.bed &gt; input_ensembl.bed\n</code></pre>\n<h2>Tool Guide: R (rtracklayer)</h2>\n<p>For R-based workflows, rtracklayer provides native liftover support.</p>\n<pre><code>library(rtracklayer)\nlibrary(GenomicRanges)\n\n# Import chain file\nchain &lt;- import.chain(\"hg19ToHg38.over.chain\")\n\n# Create GRanges object from your coordinates\ngr &lt;- GRanges(\n    seqnames = c(\"chr17\", \"chr7\", \"chr1\"),\n    ranges = IRanges(\n        start = c(41197694, 55086725, 11873),\n        end = c(41276113, 55275031, 14409)\n    ),\n    name = c(\"BRCA1\", \"EGFR\", \"DDX11L1\")\n)\n\n# Perform liftover\nlifted &lt;- liftOver(gr, chain)\n\n# liftOver returns a GRangesList (1:many mapping possible)\n# Convert to GRanges (keeping only 1:1 mappings)\nlifted_1to1 &lt;- unlist(lifted[elementNROWS(lifted) == 1])\n\n# Check for unmapped\nn_unmapped &lt;- sum(elementNROWS(lifted) == 0)\nn_multimapped &lt;- sum(elementNROWS(lifted) &gt; 1)\ncat(sprintf(\"Mapped: %d, Unmapped: %d, Multi-mapped: %d\\n\",\n    length(lifted_1to1), n_unmapped, n_multimapped))\n</code></pre>\n<h3>Bioconductor Packages for Liftover</h3>\n<table>\n<thead>\n<tr>\n<th>Package</th>\n<th>Purpose</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>rtracklayer</code></td>\n<td>Core liftover functionality</td>\n</tr>\n<tr>\n<td><code>liftOver</code> (AnnotationHub)</td>\n<td>Pre-packaged chain files</td>\n</tr>\n<tr>\n<td><code>GenomicRanges</code></td>\n<td>GRanges manipulation pre/post liftover</td>\n</tr>\n<tr>\n<td><code>VariantAnnotation</code></td>\n<td>VCF-aware liftover</td>\n</tr>\n</tbody>\n</table>\n<h2>Expected Loss Rates</h2>\n<table>\n<thead>\n<tr>\n<th>Conversion</th>\n<th>Typical Loss</th>\n<th>High-Loss Regions</th>\n<th>Notes</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>hg19 to hg38</td>\n<td>1–3%</td>\n<td>Centromeric, telomeric, segmental duplications</td>\n<td>Most reliable conversion</td>\n</tr>\n<tr>\n<td>hg38 to hg19</td>\n<td>2–5%</td>\n<td>New alt haplotypes, gap-filled regions in hg38</td>\n<td>Higher loss due to new hg38 sequences</td>\n</tr>\n<tr>\n<td>mm9 to mm10</td>\n<td>3–5%</td>\n<td>Significant rearrangements on multiple chromosomes</td>\n<td>Document chromosome-level changes</td>\n</tr>\n<tr>\n<td>mm10 to mm39</td>\n<td>1–2%</td>\n<td>Minor scaffold updates</td>\n<td>Relatively clean conversion</td>\n</tr>\n<tr>\n<td>mm10 to hg38</td>\n<td>N/A</td>\n<td>Cross-species: use synteny, not liftover</td>\n<td>Requires different approach (e.g., UCSC synteny maps)</td>\n</tr>\n</tbody>\n</table>\n<h3>When Loss Rates Are Concerning</h3>\n<ul>\n<li><strong>&lt;2% loss</strong>: Normal, proceed with analysis</li>\n<li><strong>2–5% loss</strong>: Acceptable for most analyses, document in methods</li>\n<li><strong>5–10% loss</strong>: Investigate — may indicate problematic input regions (many centromeric/repeat-rich regions)</li>\n<li><strong>&gt;10% loss</strong>: Something is wrong — check assembly mismatch, chromosome naming, or chain file version</li>\n</ul>\n<h2>Pitfalls &amp; Edge Cases</h2>\n<ul>\n<li><strong>Unmapped regions are expected</strong>: 1-5% of coordinates typically fail to lift over. Regions near centromeres, telomeres, and assembly gaps are most affected. Always check the unmapped file and report the loss rate.</li>\n<li><strong>Many-to-one mapping</strong>: Some hg19 regions map to multiple hg38 locations due to assembly improvements. UCSC liftOver reports only one mapping by default — use <code>-multiple</code> flag to detect split mappings.</li>\n<li><strong>Peak coordinates may shift asymmetrically</strong>: Peak summits can shift by different amounts than peak boundaries after liftover. Re-center peaks on summits after conversion rather than trusting the lifted boundaries.</li>\n<li><strong>Chain file source matters</strong>: Only use chain files from UCSC or Ensembl. Third-party chain files may have different coordinate conventions or incomplete mappings. Verify chain file checksums.</li>\n<li><strong>VCF liftover requires reference allele check</strong>: After lifting VCF coordinates, the reference allele may no longer match the new assembly. CrossMap handles this with <code>--refgenome</code> but UCSC liftOver does not — always validate.</li>\n<li><strong>Assembly detection is unreliable from filenames</strong>: File names like \"peaks.bed\" give no assembly hint. Check the actual coordinate ranges against known chromosome sizes. chrM length differs between hg19 (16571) and hg38 (16569).</li>\n</ul>\n<h2>Provenance Integration</h2>\n<p>Log every liftover operation with <code>encode_log_derived_file</code> for full reproducibility:</p>\n<pre><code>encode_log_derived_file(\n    file_path=\"/path/to/lifted_peaks_hg38.bed\",\n    source_accessions=[\"ENCSR...\", \"ENCFF...\"],\n    description=\"Lifted [N] narrowPeak regions from hg19 to GRCh38. [X] unmapped ([Y]% loss). Original: [source description]\",\n    file_type=\"lifted_coordinates\",\n    tool_used=\"UCSC liftOver v377\",\n    parameters=\"minMatch=0.95, chain=hg19ToHg38.over.chain.gz (MD5: abc123...), unmapped=[X]/[N] ([Y]%)\"\n)\n</code></pre>\n<h3>Provenance Checklist</h3>\n<p>Every liftover log entry should include:</p>\n<ul>\n<li>Source file path and assembly</li>\n<li>Chain file used with MD5 checksum</li>\n<li>Tool name and version (e.g., liftOver v377, CrossMap v0.6.4)</li>\n<li><code>-minMatch</code> or equivalent parameter</li>\n<li>Total input regions/variants</li>\n<li>Successfully mapped count</li>\n<li>Unmapped count and percentage</li>\n<li>Multi-mapped count (if applicable)</li>\n<li>Output file path and assembly</li>\n</ul>\n<h2>Workflow Summary</h2>\n<pre><code>1. Identify assemblies:  Check input assembly → check target assembly\n2. Get chain file:       Download from UCSC → verify MD5\n3. Select tool:          BED → liftOver | VCF → CrossMap | single → Ensembl API | R → rtracklayer\n4. Convert:              Run liftover with appropriate parameters\n5. Check loss:           Count unmapped, flag if &gt;5%\n6. Validate:             Verify output assembly, check chromosome names\n7. Post-process:         Re-center peaks, normalize VCF, re-annotate genes\n8. Log provenance:       Record all parameters, tools, loss rates\n</code></pre>\n<h2>Walkthrough: Converting ENCODE Peak Coordinates Between Genome Assemblies</h2>\n<p><strong>Goal</strong>: Convert ENCODE BED peak files from hg19 to GRCh38 (or vice versa) for cross-study integration when experiments use different genome builds.\n<strong>Context</strong>: Older ENCODE experiments may be aligned to hg19, while newer ones use GRCh38. LiftOver enables coordinate conversion for combined analysis.</p>\n<h3>Step 1: Identify experiments needing liftover</h3>\n<pre><code>encode_search_experiments(assay_title=\"Histone ChIP-seq\", organ=\"liver\", target=\"H3K27ac\", organism=\"Homo sapiens\")\n</code></pre>\n<p>Expected output:</p>\n<pre><code>{\n  \"results\": [\n    {\"accession\": \"ENCSR100OLD\", \"assay_title\": \"Histone ChIP-seq\", \"assembly\": [\"hg19\"]},\n    {\"accession\": \"ENCSR200NEW\", \"assay_title\": \"Histone ChIP-seq\", \"assembly\": [\"GRCh38\"]}\n  ],\n  \"total\": 8,\n  \"limit\": 25,\n  \"offset\": 0,\n  \"has_more\": false,\n  \"next_offset\": null\n}\n</code></pre>\n<p><strong>Interpretation</strong>: ENCSR100OLD uses hg19 — needs liftover before merging with ENCSR200NEW (GRCh38).</p>\n<h3>Step 2: Download the hg19 peak file</h3>\n<pre><code>encode_list_files(experiment_accession=\"ENCSR100OLD\", file_format=\"bed\", assembly=\"hg19\")\n</code></pre>\n<h3>Step 3: Run UCSC liftOver</h3>\n<pre><code>liftOver ENCSR100OLD_peaks.bed hg19ToHg38.over.chain.gz peaks_GRCh38.bed unmapped.bed\n</code></pre>\n<h3>Step 4: Check conversion results</h3>\n<p>Count converted vs. unmapped:</p>\n<ul>\n<li>If &gt;95% convert successfully → proceed with analysis</li>\n<li>If &gt;5% unmapped → investigate (regions may be in assembly-specific contigs)</li>\n</ul>\n<h3>Step 5: Log the conversion provenance</h3>\n<pre><code>encode_log_derived_file(\n  source_accessions=[\"ENCFF100OLD\"],\n  file_path=\"/data/peaks_GRCh38.bed\",\n  description=\"Lifted from hg19 to GRCh38 using UCSC liftOver\",\n  tool_used=\"liftOver (UCSC, chain: hg19ToHg38.over.chain.gz)\"\n)\n</code></pre>\n<h3>Integration with downstream skills</h3>\n<ul>\n<li>Lifted peaks feed into → <strong>histone-aggregation</strong> for cross-assembly union merge</li>\n<li>Conversion provenance logged by → <strong>data-provenance</strong></li>\n<li>UCSC chain files accessed via → <strong>ucsc-browser</strong> REST API</li>\n<li>Lifted coordinates used by → <strong>variant-annotation</strong> for position-dependent annotation</li>\n</ul>\n<h2>Code Examples</h2>\n<h3>1. Check file assembly before liftover</h3>\n<pre><code>encode_get_file_info(accession=\"ENCFF100OLD\")\n</code></pre>\n<p>Expected output:</p>\n<pre><code>{\n  \"accession\": \"ENCFF100OLD\",\n  \"file_format\": \"bed\",\n  \"file_type\": \"bed narrowPeak\",\n  \"output_type\": \"IDR thresholded peaks\",\n  \"assembly\": \"hg19\",\n  \"file_size\": 1258291,\n  \"file_size_human\": \"1.2 MB\",\n  \"status\": \"released\"\n}\n</code></pre>\n<h3>2. Find GRCh38 version of same experiment</h3>\n<pre><code>encode_list_files(experiment_accession=\"ENCSR100OLD\", file_format=\"bed\", assembly=\"GRCh38\")\n</code></pre>\n<p>Expected output (a JSON array of file records — empty here):</p>\n<pre><code>[]\n</code></pre>\n<p><strong>Interpretation</strong>: No GRCh38 files available — liftover is required.</p>\n<h3>3. Log liftover provenance</h3>\n<pre><code>encode_log_derived_file(\n  source_accessions=[\"ENCFF100OLD\"],\n  file_path=\"/data/peaks_GRCh38.bed\",\n  description=\"hg19→GRCh38 liftOver\",\n  tool_used=\"UCSC liftOver\"\n)\n</code></pre>\n<p>Expected output:</p>\n<pre><code>{\n  \"success\": true,\n  \"record_id\": 7,\n  \"file_path\": \"/data/peaks_GRCh38.bed\",\n  \"source_accessions\": [\"ENCFF100OLD\"],\n  \"message\": \"Provenance logged. Use encode_get_provenance to view the full chain.\"\n}\n</code></pre>\n<h2>Related Skills</h2>\n<ul>\n<li><code>variant-annotation</code> — Variants often need liftover before annotation with ENCODE data (hg19 GWAS variants to GRCh38)</li>\n<li><code>gwas-catalog</code> — GWAS Catalog coordinates may be in GRCh37; liftover needed for ENCODE GRCh38 integration</li>\n<li><code>ensembl-annotation</code> — Ensembl REST API provides coordinate mapping; Ensembl uses non-chr chromosome naming</li>\n<li><code>ucsc-browser</code> — UCSC provides chain files and the liftOver tool; retrieve assembly-specific tracks</li>\n<li><code>gnomad-variants</code> — gnomAD v4 uses GRCh38; v2 uses GRCh37; liftover needed for cross-version analysis</li>\n<li><code>histone-aggregation</code> — Aggregating peaks across samples requires all peaks in the same assembly</li>\n<li><code>accessibility-aggregation</code> — ATAC-seq/DNase-seq peak union requires assembly-consistent coordinates</li>\n<li><code>data-provenance</code> — Every liftover operation must be logged with chain file, tool version, and loss rate</li>\n<li><code>publication-trust</code> — Verify literature claims backing analytical decisions</li>\n</ul>\n<h2>Presenting Results</h2>\n<p>When reporting liftover results, always present:</p>\n<ul>\n<li><strong>Input summary</strong>: Number of regions/variants, source assembly</li>\n<li><strong>Output summary</strong>: Number successfully mapped, target assembly</li>\n<li><strong>Loss report</strong>: Unmapped count and percentage, with breakdown by reason if available</li>\n<li><strong>Multi-mapping report</strong>: Number of regions mapping to multiple locations and how they were handled</li>\n<li><strong>Assembly confirmation</strong>: Explicit statement of output assembly (e.g., \"All coordinates are now in GRCh38/hg38\")</li>\n<li><strong>Flag if loss &gt;5%</strong>: Warn the user and investigate the cause (centromeric regions, assembly-specific sequences, or input errors)</li>\n<li><strong>Chain file version</strong>: Which chain file was used and its source</li>\n</ul>\n<p>Example output summary:</p>\n<pre><code>Liftover: hg19 -&gt; GRCh38\nInput:    45,231 narrowPeak regions\nMapped:   44,012 (97.3%)\nUnmapped: 1,219 (2.7%) — 847 partially deleted, 312 split, 60 fully deleted\nChain:    hg19ToHg38.over.chain.gz (UCSC, MD5: 7a42e...)\nTool:     UCSC liftOver v377, minMatch=0.95\n</code></pre>\n<h2>For the request: \"$ARGUMENTS\"</h2>\n","files":[{"path":"references/literature.md","sizeBytes":6444,"isText":true},{"path":"SKILL.md","sizeBytes":23330,"isText":true}],"reviewScore":null,"reviewSummary":null,"trust":{"provenance":"trusted-source-unreviewed","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow.","bodySource":null},"bodyLocked":false,"purchaseUrl":null,"sourceUrl":null,"report":{"provenance":"trusted-source-unreviewed","screen":{"ran":true,"outcome":"clean","suspicious":0,"notes":0,"hiddenCharacters":false},"virusScan":{"engine":"clamav","status":"clean","scannedAt":"2026-09-22T13:27:14.185029Z","sha256":"D94B34E70572B27BDA4E0597A38387CF885E4029BF6A545DE7251592FE073792","sizeBytes":11974},"review":null,"source":{"repositoryUrl":"https://github.com/ammawla/encode-toolkit","path":"skills/liftover-coordinates","license":"AGPL-3.0","commit":"36836c8725fd4d20d9c851ce314f5151aea5f57c","subtreeSha":"2D08AA98BABBFE03E4B57A15BE515280220C35B61935F028ECE315EC35E2188D","lastSyncedAt":"2026-09-29T20:56:54.045383Z"},"reviewedAt":"2026-09-22T13:30:10.403552Z","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow."},"install":[{"target":"skills-cli","command":"npx skills add https://github.com/ammawla/encode-toolkit/tree/main/skills/liftover-coordinates"},{"target":"claude-code","command":"claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install ammawla-encode-toolkit@llmmart"},{"target":"git","command":"git clone https://github.com/ammawla/encode-toolkit.git"}]}