{"slug":"msa-structure-prediction-pipeline","title":"msa-structure-prediction-pipeline","summary":"NOTE: your protein sequence and the retrieved MSA alignment are transmitted to external NVIDIA-hosted APIs (health.api.nvidia.com) on every call. Use local NIM containers for confidential or proprietary sequences. Run a complete protein structure prediction pipeline using NVIDIA ","platform":"Claude","tags":[],"authorName":"LLM Mart","authorSlug":"llm-mart","score":0,"source":"github","price":null,"verified":false,"createdAt":"2026-09-26T16:30:33.235373Z","repo":{"url":"https://github.com/NVIDIA/skills","stars":3445,"forks":416,"license":"Apache-2.0","updatedAt":"2026-09-25T03:14:56Z"},"bodyHtml":"<hr>\n<h2>name: msa-structure-prediction-pipeline\ndescription: &gt;\nNOTE: your protein sequence and the retrieved MSA alignment are transmitted to\nexternal NVIDIA-hosted APIs (health.api.nvidia.com) on every call. Use local\nNIM containers for confidential or proprietary sequences.\nRun a complete protein structure prediction pipeline using NVIDIA BioNeMo NIMs:\nsearch for MSA alignments with MSA-Search (ColabFold), then predict the structure\nwith OpenFold3 using the retrieved alignments. Use this skill whenever the user wants\nto predict a protein structure with maximum accuracy using MSA context, run the\nfull AlphaFold3-style pipeline, generate MSA-informed structure predictions, or\nimprove structure prediction accuracy by providing evolutionary information.\nTriggers on: MSA structure prediction pipeline, structure prediction pipeline, MSA-informed prediction, OpenFold3,\nColabFold MSA, AlphaFold3 pipeline, protein structure, homology search, a3m alignment,\nUniRef30, NIM microservice. This pipeline chains MSA-Search and OpenFold3.\nlicense: Apache-2.0 AND CC-BY-4.0\nallowed-tools: Bash, Read, Write, AskUserQuestion</h2>\n<h1>MSA Structure Prediction Pipeline</h1>\n<p>Predict protein structures with high accuracy by chaining two BioNeMo NIMs:</p>\n<pre><code>Step 1: MSA-Search  →  Step 2: OpenFold3\n(Search homologs)       (Predict structure with MSA)\n</code></pre>\n<hr>\n<h2>Overview</h2>\n<p>Why chain these NIMs?</p>\n<ul>\n<li><strong>MSA-Search</strong> finds evolutionary homologs in UniRef30 and ColabFold databases using GPU-accelerated MMSeqs2. The resulting alignment provides crucial evolutionary information.</li>\n<li><strong>OpenFold3</strong> uses the MSA to improve structure prediction accuracy — especially for sequences where no close homolog exists in PDB.</li>\n<li>Running MSA-Search first means OpenFold3 gets the full evolutionary context rather than a single-sequence prediction.</li>\n</ul>\n<hr>\n<h2>Before you start</h2>\n<p>Confirm with the user:</p>\n<ol>\n<li><strong>Query sequence</strong>: amino acid sequence to predict</li>\n<li><strong>MSA depth</strong>: how many sequences to retrieve (default 500; more = slower but more context)</li>\n<li><strong>API mode</strong>: hosted or local Docker?</li>\n</ol>\n<p>Note: local MSA-Search requires 1.4 TB of database storage — strongly recommend hosted unless the user has that infrastructure.</p>\n<p>For local Docker, do not assume MSA-Search and OpenFold3 are both on\n<code>localhost:8000</code> concurrently. Run one container at a time and hand off the A3M\nfile, or start each NIM on a distinct host port and set the URLs explicitly.</p>\n<hr>\n<h2>Step 1: Search for MSA with MSA-Search</h2>\n<pre><code>import requests, json, os\nfrom pathlib import Path\n\nNGC_API_KEY = os.getenv(\"NGC_API_KEY\")\nHOSTED = True\n\nquery_sequence = \"&lt;YOUR_PROTEIN_SEQUENCE&gt;\"\n\nif HOSTED:\n    msa_url = \"https://health.api.nvidia.com/v1/biology/colabfold/msa-search/predict\"\n    headers = {\"Content-Type\": \"application/json\",\n               \"Authorization\": f\"Bearer {NGC_API_KEY}\"}\nelse:\n    msa_url = \"http://localhost:8000/biology/colabfold/msa-search/predict\"\n    headers = {\"Content-Type\": \"application/json\"}\n\npayload = {\n    \"sequence\": query_sequence,\n    \"databases\": [\"Uniref30_2302\", \"colabfold_envdb_202108\"],\n    \"e_value\": 0.0001,\n    \"output_alignment_formats\": [\"a3m\"],\n}\n\nr = requests.post(msa_url, headers=headers, json=payload)\nr.raise_for_status()\nmsa_result = r.json()\n\n# Extract the A3M alignment\na3m_alignment = msa_result[\"alignments\"][\"Uniref30_2302\"][\"a3m\"][\"alignment\"]\n\n# Save for reference\nwith open(\"query_msa.a3m\", \"w\") as f:\n    f.write(a3m_alignment)\n\n# Count sequences in alignment\nn_seqs = a3m_alignment.count(\"&gt;\")\nprint(f\"Step 1 complete: found {n_seqs} homologous sequences\")\nprint(f\"MSA saved to query_msa.a3m\")\n</code></pre>\n<hr>\n<h2>Step 2: Predict structure with OpenFold3</h2>\n<p>Pass the MSA directly into OpenFold3's <code>msa</code> field:</p>\n<pre><code>if HOSTED:\n    of3_url = \"https://health.api.nvidia.com/v1/biology/openfold/openfold3/predict\"\nelse:\n    of3_url = \"http://localhost:8000/biology/openfold/openfold3/predict\"\n\n# Build the OpenFold3 MSA structure from the retrieved alignment\nmsa_data = {\n    \"uniref30\": {\n        \"a3m\": {\n            \"alignment\": a3m_alignment,\n            \"format\": \"a3m\"\n        }\n    }\n}\n\n# Optionally also include colabfold_envdb alignment if requested\n# env_alignment = msa_result[\"alignments\"][\"colabfold_envdb\"][\"a3m\"][\"alignment\"]\n# msa_data[\"colabfold_env\"] = {\"a3m\": {\"alignment\": env_alignment, \"format\": \"a3m\"}}\n\npayload = {\n    \"inputs\": [{\n        \"input_id\": \"prediction_with_msa\",\n        \"output_format\": \"pdb\",\n        \"molecules\": [\n            {\n                \"type\": \"protein\",\n                \"sequence\": query_sequence,\n                \"diffusion_samples\": 1,\n                \"msa\": msa_data\n            }\n        ]\n    }]\n}\n\nr = requests.post(of3_url, headers=headers, json=payload, timeout=300)\nr.raise_for_status()\nresult = r.json()\n\noutput = result[\"outputs\"][0]\nfor i, sample in enumerate(output[\"structures_with_scores\"]):\n    fmt = sample[\"format\"]\n    filename = f\"predicted_structure_{i+1}.{fmt}\"\n    with open(filename, \"w\") as f:\n        f.write(sample[\"structure\"])\n    print(f\"\\nStep 2 complete: {filename} saved\")\n    print(f\"  Confidence:  {sample['confidence_score']:.4f}\")\n    print(f\"  pLDDT:       {sample['complex_plddt_score']:.4f}\")\n    print(f\"  pTM:         {sample['ptm_score']:.4f}\")\n</code></pre>\n<hr>\n<h2>Comparing single-sequence vs MSA-informed prediction</h2>\n<p>If the user wants to see the impact of MSA, run OpenFold3 twice — once with the full MSA and once with just the query sequence as a minimal alignment:</p>\n<pre><code># Minimal MSA (single sequence — same as no MSA context):\nminimal_msa = {\n    \"main\": {\n        \"a3m\": {\n            \"alignment\": f\"&gt;query\\n{query_sequence}\",\n            \"format\": \"a3m\"\n        }\n    }\n}\n</code></pre>\n<p>A larger, higher-quality MSA typically yields higher pLDDT and lower pDE, especially for proteins with many known homologs.</p>\n<hr>\n<h2>For protein complexes</h2>\n<p>Use the <code>/paired/predict</code> endpoint of MSA-Search to get paired alignments for multi-chain complexes, then pass each chain's alignment into the corresponding molecule's <code>msa</code> field and <code>paired_msa</code> fields:</p>\n<pre><code># Paired MSA search endpoint for complexes:\nmsa_paired_url = \"https://health.api.nvidia.com/v1/biology/colabfold/msa-search/paired/predict\"\npaired_payload = {\n    \"sequences\": [chain_A_sequence, chain_B_sequence],\n    \"e_value\": 0.0001,\n}\n</code></pre>\n<hr>\n<h2>Quick reference — skill dependencies</h2>\n<table>\n<thead>\n<tr>\n<th>Step</th>\n<th>Skill</th>\n<th>Key endpoint</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>MSA search</td>\n<td><code>msa-search-nim</code></td>\n<td><code>/biology/colabfold/msa-search/predict</code></td>\n</tr>\n<tr>\n<td>Structure prediction</td>\n<td><code>openfold3-nim</code></td>\n<td><code>/biology/openfold/openfold3/predict</code></td>\n</tr>\n</tbody>\n</table>\n","files":[{"path":"BENCHMARK.md","sizeBytes":7893,"isText":true},{"path":"evals/evals.json","sizeBytes":8006,"isText":true},{"path":"skill-card.md","sizeBytes":4153,"isText":true},{"path":"SKILL.md","sizeBytes":6570,"isText":true},{"path":"skill.oms.sig","sizeBytes":4609,"isText":false}],"reviewScore":null,"reviewSummary":null,"trust":{"provenance":"trusted-source-unreviewed","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow.","bodySource":null},"bodyLocked":false,"purchaseUrl":null,"sourceUrl":null,"report":{"provenance":"trusted-source-unreviewed","screen":{"ran":true,"outcome":"clean","suspicious":0,"notes":0,"hiddenCharacters":false},"virusScan":{"engine":"clamav","status":"clean","scannedAt":"2026-09-26T16:30:59.020786Z","sha256":"051B2E98E76C582EC5C1BBB7FBA08C939FACE49FF42AC305938AB61A6DFF8FC9","sizeBytes":13002},"review":null,"source":{"repositoryUrl":"https://github.com/NVIDIA/skills","path":"skills/bionemo-msa-structure-prediction-pipeline","license":"Apache-2.0","commit":"d8519c57da6db5d9bea274ec1724a4a7a56a3dee","subtreeSha":"CFB8F4DDA16F909BD2F5A78BE7490AE75C8127270A69FD8FDAF184C767BB8E1D","lastSyncedAt":"2026-09-26T16:30:30.547995Z"},"reviewedAt":"2026-09-26T16:31:47.781649Z","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow."},"install":[{"target":"skills-cli","command":"npx skills add https://github.com/NVIDIA/skills/tree/main/skills/bionemo-msa-structure-prediction-pipeline"},{"target":"claude-code","command":"claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install nvidia-skills@llmmart"},{"target":"git","command":"git clone https://github.com/NVIDIA/skills.git"}]}