{"slug":"alterlab-bindingdb","title":"alterlab-bindingdb","summary":"Query BindingDB for measured protein-ligand binding affinities (Ki, Kd, IC50, EC50) via its keyless REST API or the full TSV download, searching by target (UniProt ID), compound (SMILES), or pathogen. Use when looking up experimental binding constants, profiling inhibitors of a p","platform":"Claude","tags":[],"authorName":"LLM Mart","authorSlug":"llm-mart","score":0,"source":"github","price":null,"verified":false,"createdAt":"2026-09-23T18:57:07.939819Z","repo":{"url":"https://github.com/AlterLab-IEU/AlterLab-Academic-Skills","stars":68,"forks":13,"license":"MIT","updatedAt":"2026-09-23T13:42:59Z"},"bodyHtml":"<hr>\n<h2>name: alterlab-bindingdb\ndescription: Query BindingDB for measured protein-ligand binding affinities (Ki, Kd, IC50, EC50) via its keyless REST API or the full TSV download, searching by target (UniProt ID), compound (SMILES), or pathogen. Use when looking up experimental binding constants, profiling inhibitors of a protein target, doing lead optimization, polypharmacology analysis, or structure-activity relationship (SAR) studies; for curated bioactivity mining or drug-like compound library screening at scale prefer alterlab-chembl instead. Part of the AlterLab Academic Skills suite.\nlicense: CC-BY-3.0\nallowed-tools: Read WebFetch Bash(curl:<em>) Bash(python:</em>)\ncompatibility: Keyless public BindingDB web services (no authentication required)\nmetadata:\nskill-author: AlterLab\nversion: \"1.0.1\"\nlast_updated: \"2026-09-23\"</h2>\n<h1>BindingDB Database</h1>\n<h2>Overview</h2>\n<p>BindingDB (<a href=\"https://www.bindingdb.org/\">https://www.bindingdb.org/</a>) is the primary public database of measured drug-protein binding affinities. It contains roughly 3.2 million binding data records for ~1.4 million compounds tested against ~11,500 protein targets (homepage figures, 2026-09), curated from scientific literature and patent literature. BindingDB stores quantitative binding measurements (Ki, Kd, IC50, EC50) essential for drug discovery, pharmacology, and computational chemistry research.</p>\n<p><strong>Key resources:</strong></p>\n<ul>\n<li>BindingDB website: <a href=\"https://www.bindingdb.org/\">https://www.bindingdb.org/</a></li>\n<li>REST API base: <a href=\"https://bindingdb.org/rest/\">https://bindingdb.org/rest/</a> (no key; default response is XML, append <code>response=application/json</code>)</li>\n<li>Downloads page: <a href=\"https://www.bindingdb.org/rwd/bind/chemsearch/marvin/Download.jsp\">https://www.bindingdb.org/rwd/bind/chemsearch/marvin/Download.jsp</a> (the full TSV is the dated <code>BindingDB_All_&lt;YYYYMM&gt;_tsv.zip</code>, ~560 MB zipped, refreshed monthly)</li>\n</ul>\n<h2>When to Use This Skill</h2>\n<p>Use BindingDB when:</p>\n<ul>\n<li><strong>Target-based drug discovery</strong>: What known compounds bind to a target protein? What are their affinities?</li>\n<li><strong>SAR analysis</strong>: How do structural modifications affect binding affinity for a series of analogs?</li>\n<li><strong>Lead compound profiling</strong>: What targets does a compound bind (selectivity/polypharmacology)?</li>\n<li><strong>Benchmark datasets</strong>: Obtain curated protein-ligand affinity data for ML model training</li>\n<li><strong>Repurposing analysis</strong>: Does an approved drug bind to an unintended target?</li>\n<li><strong>Competitive analysis</strong>: What is the best reported affinity for a target class?</li>\n<li><strong>Fragment screening</strong>: Find validated binding data for fragments against a target</li>\n</ul>\n<h3>Does NOT Trigger</h3>\n<table>\n<thead>\n<tr>\n<th>Scenario</th>\n<th>Use Instead</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Curated bioactivity mining at scale, assay metadata, drug mechanisms</td>\n<td><code>alterlab-chembl</code></td>\n</tr>\n<tr>\n<td>Compound identifiers/properties by name or CID, PubChem BioAssay</td>\n<td><code>alterlab-pubchem</code></td>\n</tr>\n<tr>\n<td>Purchasable analogs or docking-ready 3D libraries</td>\n<td><code>alterlab-zinc-db</code></td>\n</tr>\n<tr>\n<td>Docking a ligand into a receptor structure</td>\n<td><code>alterlab-diffdock</code></td>\n</tr>\n</tbody>\n</table>\n<h2>Core Capabilities</h2>\n<h3>1. BindingDB REST API</h3>\n<p>Base URL: <code>https://bindingdb.org/rest</code></p>\n<pre><code>import requests\n\nBASE_URL = \"https://bindingdb.org/rest\"\n\ndef bindingdb_query(method, params):\n    \"\"\"Query the BindingDB REST API.\"\"\"\n    url = f\"{BASE_URL}/{method}\"\n    response = requests.get(url, params=params, headers={\"Accept\": \"application/json\"})\n    response.raise_for_status()\n    return response.json()\n</code></pre>\n<h3>2. Query by Target (UniProt ID)</h3>\n<p>The <code>getLigandsByUniprot</code> endpoint takes a single <code>uniprot</code> parameter formatted as\n<code>&lt;UniProt accession&gt;;&lt;cutoff in nM&gt;</code>, e.g. <code>P00519;10000</code>.</p>\n<pre><code>def get_ligands_for_target(uniprot_id, cutoff=10000):\n    \"\"\"\n    Get all ligands with measured affinity for a UniProt target.\n\n    Args:\n        uniprot_id: UniProt accession (e.g., \"P00519\" for ABL1)\n        cutoff: Maximum affinity value to return (in nM)\n    \"\"\"\n    params = {\n        \"uniprot\": f\"{uniprot_id};{cutoff}\",\n        \"response\": \"application/json\",\n    }\n    return bindingdb_query(\"getLigandsByUniprot\", params)\n\n# Example: Get all compounds binding ABL1 (imatinib target) at &lt;=100 nM\nligands = get_ligands_for_target(\"P00519\", cutoff=100)\n\n# Response shape (verified 2026-09): one top-level key, spelled\n# \"getLindsByUniprotResponse\" (sic — getTargetByCompound uses the same key),\n# holding \"bdb.hit\" (count, as a string) and \"bdb.affinities\": a list of\n# {\"bdb.monomerid\", \"bdb.smile\", \"bdb.affinity_type\", \"bdb.affinity\"}.\n# Affinities are strings, may carry leading spaces or \"&gt;\"/\"&lt;\" qualifiers.\nresp = next(iter(ligands.values()))\nrows = resp.get(\"bdb.affinities\", [])\n</code></pre>\n<h3>3. Query by SMILES (structural similarity)</h3>\n<pre><code>def search_by_smiles(smiles, cutoff=0.85):\n    \"\"\"\n    Search BindingDB by SMILES string (structural-similarity search).\n\n    Args:\n        smiles: SMILES string of the compound\n        cutoff: Tanimoto similarity threshold (0.0-1.0; e.g. 0.85)\n    \"\"\"\n    params = {\n        \"smiles\": smiles,\n        \"cutoff\": cutoff,\n        \"response\": \"application/json\",\n    }\n    return bindingdb_query(\"getTargetByCompound\", params)\n\n# Example: structural-similarity search for imatinib's binding targets\nresult = search_by_smiles(\"Cc1ccc(NC(=O)c2ccc(CN3CCN(C)CC3)cc2)cc1Nc1nccc(-c2cccnc2)n1\")\n</code></pre>\n<p>To query by a PubChem CID, first convert the CID to a SMILES string via PubChem\nPUG-REST, then pass it to <code>search_by_smiles</code>. There is no by-name or by-CID\nBindingDB REST endpoint.</p>\n<h3>4. Download-Based Analysis (Recommended for Large Queries)</h3>\n<p>For comprehensive analyses, download BindingDB data directly:</p>\n<pre><code>import pandas as pd\n\ndef load_bindingdb(filepath=\"BindingDB_All.tsv\"):\n    \"\"\"\n    Load BindingDB TSV file (the unzipped BindingDB_All_&lt;YYYYMM&gt;.tsv).\n    Download the dated BindingDB_All_&lt;YYYYMM&gt;_tsv.zip from:\n    https://www.bindingdb.org/rwd/bind/chemsearch/marvin/Download.jsp\n    \"\"\"\n    # Key columns\n    usecols = [\n        \"BindingDB Reactant_set_id\",\n        \"Ligand SMILES\",\n        \"Ligand InChI\",\n        \"Ligand InChI Key\",\n        \"BindingDB Target Chain  Sequence\",\n        \"PDB ID(s) for Ligand-Target Complex\",\n        \"UniProt (SwissProt) Entry Name of Target Chain\",\n        \"UniProt (SwissProt) Primary ID of Target Chain\",\n        \"UniProt (TrEMBL) Primary ID of Target Chain\",\n        \"Ki (nM)\",\n        \"IC50 (nM)\",\n        \"Kd (nM)\",\n        \"EC50 (nM)\",\n        \"kon (M-1-s-1)\",\n        \"koff (s-1)\",\n        \"Target Name\",\n        \"Target Source Organism According to Curator or DataSource\",\n        \"Number of Protein Chains in Target (&gt;1 implies a multichain complex)\",\n        \"PubChem CID\",\n        \"PubChem SID\",\n        \"ChEMBL ID of Ligand\",\n        \"DrugBank ID of Ligand\",\n    ]\n\n    # Exact TSV headers drift between monthly releases (some contain double\n    # spaces), so intersect with the actual header rather than hard-failing.\n    header = pd.read_csv(filepath, sep=\"\\t\", nrows=0).columns\n    keep = [c for c in usecols if c in header]\n    df = pd.read_csv(filepath, sep=\"\\t\", usecols=keep,\n                     low_memory=False, on_bad_lines='skip')\n\n    # Convert affinity columns to numeric\n    for col in [\"Ki (nM)\", \"IC50 (nM)\", \"Kd (nM)\", \"EC50 (nM)\"]:\n        if col in df.columns:\n            df[col] = pd.to_numeric(df[col], errors='coerce')\n\n    return df\n\ndef query_target_affinity(df, uniprot_id, affinity_types=None, max_nm=10000):\n    \"\"\"Query loaded BindingDB for a specific target.\"\"\"\n    if affinity_types is None:\n        affinity_types = [\"Ki (nM)\", \"IC50 (nM)\", \"Kd (nM)\"]\n\n    # Filter by UniProt ID\n    mask = df[\"UniProt (SwissProt) Primary ID of Target Chain\"] == uniprot_id\n    target_df = df[mask].copy()\n\n    # Filter by affinity cutoff\n    has_affinity = pd.Series(False, index=target_df.index)\n    for col in affinity_types:\n        if col in target_df.columns:\n            has_affinity |= target_df[col] &lt;= max_nm\n\n    result = target_df[has_affinity][[\"Ligand SMILES\"] + affinity_types +\n                                      [\"PubChem CID\", \"ChEMBL ID of Ligand\"]].dropna(how='all')\n    return result.sort_values(affinity_types[0])\n</code></pre>\n<h3>5. SAR Analysis</h3>\n<pre><code>import pandas as pd\n\ndef sar_analysis(df, target_uniprot, affinity_col=\"IC50 (nM)\"):\n    \"\"\"\n    Structure-activity relationship analysis for a target.\n    Retrieves all compounds with affinity data and ranks by potency.\n    \"\"\"\n    target_data = query_target_affinity(df, target_uniprot, [affinity_col])\n\n    if target_data.empty:\n        return target_data\n\n    # Add pIC50 (negative log of IC50 in molar)\n    if affinity_col in target_data.columns:\n        target_data = target_data[target_data[affinity_col].notna()].copy()\n        target_data[\"pAffinity\"] = -((target_data[affinity_col] * 1e-9).apply(\n            lambda x: __import__('math').log10(x)\n        ))\n        target_data = target_data.sort_values(\"pAffinity\", ascending=False)\n\n    return target_data\n\n# Most potent compounds against EGFR (P00533)\n# sar = sar_analysis(df, \"P00533\", \"IC50 (nM)\")\n# print(sar.head(20))\n</code></pre>\n<h3>6. Polypharmacology Profile</h3>\n<pre><code>def polypharmacology_profile(df, ligand_smiles, affinity_cutoff_nM=1000):\n    \"\"\"\n    Find all targets a compound binds to, by exact SMILES match.\n    For tolerance to tautomers/salts/charge states, match on the InChIKey\n    skeleton (first 14 chars of \"Ligand InChI Key\") instead of raw SMILES.\n    \"\"\"\n    # Search by ligand SMILES (exact string match)\n    mask = df[\"Ligand SMILES\"] == ligand_smiles\n\n    ligand_data = df[mask].copy()\n\n    # Filter by affinity\n    aff_cols = [\"Ki (nM)\", \"IC50 (nM)\", \"Kd (nM)\"]\n    has_aff = pd.Series(False, index=ligand_data.index)\n    for col in aff_cols:\n        if col in ligand_data.columns:\n            has_aff |= ligand_data[col] &lt;= affinity_cutoff_nM\n\n    result = ligand_data[has_aff][\n        [\"Target Name\", \"UniProt (SwissProt) Primary ID of Target Chain\"] + aff_cols\n    ].dropna(how='all')\n\n    # Rank by the tightest measured constant per row (NaNs ignored)\n    result = result.assign(best_nM=result[aff_cols].min(axis=1))\n    return result.sort_values(\"best_nM\")\n</code></pre>\n<h2>Query Workflows</h2>\n<h3>Workflow 1: Find Best Inhibitors for a Target</h3>\n<pre><code>import pandas as pd\n\ndef find_best_inhibitors(uniprot_id, affinity_type=\"IC50 (nM)\", top_n=20):\n    \"\"\"Find the most potent inhibitors for a target in BindingDB.\"\"\"\n    df = load_bindingdb(\"BindingDB_All.tsv\")  # Load once and reuse\n    result = query_target_affinity(df, uniprot_id, [affinity_type])\n\n    if result.empty:\n        print(f\"No data found for {uniprot_id}\")\n        return result\n\n    result = result.sort_values(affinity_type).head(top_n)\n    print(f\"Top {top_n} inhibitors for {uniprot_id} by {affinity_type}:\")\n    for _, row in result.iterrows():\n        print(f\"  {row['PubChem CID']}: {row[affinity_type]:.1f} nM | SMILES: {row['Ligand SMILES'][:40]}...\")\n    return result\n</code></pre>\n<h3>Workflow 2: Selectivity Profiling</h3>\n<ol>\n<li>Get all affinity data for your compound across all targets</li>\n<li>Compare affinity ratios between on-target and off-targets</li>\n<li>Identify selectivity cliffs (structural changes that improve selectivity)</li>\n<li>Cross-reference with ChEMBL for additional selectivity data</li>\n</ol>\n<h3>Workflow 3: Machine Learning Dataset Preparation</h3>\n<pre><code>def prepare_ml_dataset(df, uniprot_ids, affinity_col=\"IC50 (nM)\",\n                        max_affinity_nM=100000, min_count=50):\n    \"\"\"Prepare BindingDB data for ML model training.\"\"\"\n    records = []\n    for uid in uniprot_ids:\n        target_df = query_target_affinity(df, uid, [affinity_col], max_affinity_nM)\n        if len(target_df) &gt;= min_count:\n            target_df = target_df.copy()\n            target_df[\"target\"] = uid\n            records.append(target_df)\n\n    if not records:\n        return pd.DataFrame()\n\n    combined = pd.concat(records)\n    # Add pAffinity (normalized)\n    combined[\"pAffinity\"] = -((combined[affinity_col] * 1e-9).apply(\n        lambda x: __import__('math').log10(max(x, 1e-12))\n    ))\n    return combined[[\"Ligand SMILES\", \"target\", \"pAffinity\", affinity_col]].dropna()\n</code></pre>\n<h2>Key Data Fields</h2>\n<table>\n<thead>\n<tr>\n<th>Field</th>\n<th>Description</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>Ligand SMILES</code></td>\n<td>2D structure of the compound</td>\n</tr>\n<tr>\n<td><code>Ligand InChI Key</code></td>\n<td>Unique chemical identifier</td>\n</tr>\n<tr>\n<td><code>Ki (nM)</code></td>\n<td>Inhibition constant (equilibrium, functional)</td>\n</tr>\n<tr>\n<td><code>Kd (nM)</code></td>\n<td>Dissociation constant (thermodynamic, binding)</td>\n</tr>\n<tr>\n<td><code>IC50 (nM)</code></td>\n<td>Half-maximal inhibitory concentration</td>\n</tr>\n<tr>\n<td><code>EC50 (nM)</code></td>\n<td>Half-maximal effective concentration</td>\n</tr>\n<tr>\n<td><code>kon (M-1-s-1)</code></td>\n<td>Association rate constant</td>\n</tr>\n<tr>\n<td><code>koff (s-1)</code></td>\n<td>Dissociation rate constant</td>\n</tr>\n<tr>\n<td><code>UniProt (SwissProt) Primary ID</code></td>\n<td>Target UniProt accession</td>\n</tr>\n<tr>\n<td><code>Target Name</code></td>\n<td>Protein name</td>\n</tr>\n<tr>\n<td><code>PDB ID(s) for Ligand-Target Complex</code></td>\n<td>Crystal structures</td>\n</tr>\n<tr>\n<td><code>PubChem CID</code></td>\n<td>PubChem compound ID</td>\n</tr>\n<tr>\n<td><code>ChEMBL ID of Ligand</code></td>\n<td>ChEMBL compound ID</td>\n</tr>\n</tbody>\n</table>\n<h2>Affinity Interpretation</h2>\n<table>\n<thead>\n<tr>\n<th>Affinity</th>\n<th>Classification</th>\n<th>Drug-likeness</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>&lt; 1 nM</td>\n<td>Sub-nanomolar</td>\n<td>Very potent (picomolar range)</td>\n</tr>\n<tr>\n<td>1–10 nM</td>\n<td>Nanomolar</td>\n<td>Potent, typical for approved drugs</td>\n</tr>\n<tr>\n<td>10–100 nM</td>\n<td>Moderate</td>\n<td>Common lead compounds</td>\n</tr>\n<tr>\n<td>100–1000 nM</td>\n<td>Weak</td>\n<td>Fragment/starting point</td>\n</tr>\n<tr>\n<td>&gt; 1000 nM</td>\n<td>Very weak</td>\n<td>Generally below drug-relevance threshold</td>\n</tr>\n</tbody>\n</table>\n<h2>Best Practices</h2>\n<ul>\n<li><strong>Use Ki for direct binding</strong>: Ki reflects true binding affinity independent of enzymatic mechanism</li>\n<li><strong>IC50 context-dependency</strong>: IC50 values depend on substrate concentration (Cheng-Prusoff equation)</li>\n<li><strong>Watch for qualifier prefixes</strong>: TSV affinity cells often carry <code>&gt;</code>, <code>&lt;</code>, or <code>&gt;=</code> prefixes (e.g. <code>&gt;10000</code>) for censored measurements. <code>pd.to_numeric(errors='coerce')</code> silently turns these into NaN — strip the prefix first (e.g. <code>df[col].astype(str).str.lstrip(\"&lt;&gt;= \")</code>) and decide explicitly whether to keep, drop, or treat censored values as inequalities before modeling</li>\n<li><strong>Normalize units</strong>: BindingDB reports in nM; verify units when comparing across studies</li>\n<li><strong>Filter by target organism</strong>: Use <code>Target Source Organism</code> to ensure human protein data</li>\n<li><strong>Handle missing values</strong>: Not all compounds have all measurement types</li>\n<li><strong>Cross-reference with ChEMBL</strong>: ChEMBL has more curated activity data for medicinal chemistry</li>\n</ul>\n<h2>Additional Resources</h2>\n<ul>\n<li><strong>BindingDB website</strong>: <a href=\"https://www.bindingdb.org/\">https://www.bindingdb.org/</a></li>\n<li><strong>Data downloads</strong>: <a href=\"https://www.bindingdb.org/rwd/bind/chemsearch/marvin/Download.jsp\">https://www.bindingdb.org/rwd/bind/chemsearch/marvin/Download.jsp</a></li>\n<li><strong>REST API documentation</strong>: <a href=\"https://www.bindingdb.org/rwd/bind/BindingDBRESTfulAPI.jsp\">https://www.bindingdb.org/rwd/bind/BindingDBRESTfulAPI.jsp</a> (REST base: <a href=\"https://bindingdb.org/rest\">https://bindingdb.org/rest</a>)</li>\n<li><strong>Citation</strong>: Liu T et al. \"BindingDB in 2024: a FAIR knowledgebase of protein-small molecule binding data.\" Nucleic Acids Research 2025;53(D1):D1633-D1644. doi:10.1093/nar/gkae1075 (earlier: Gilson MK et al., NAR 2016;44(D1):D1045-53, doi:10.1093/nar/gkv1072)</li>\n<li><strong>Related resources</strong>: ChEMBL (<a href=\"https://www.ebi.ac.uk/chembl/\">https://www.ebi.ac.uk/chembl/</a>), PubChem BioAssay</li>\n</ul>\n<h2>Scripts</h2>\n<p><code>scripts/query_bindingdb.py</code> — runnable helper for the BindingDB REST API (no key):</p>\n<pre><code>python scripts/query_bindingdb.py uniprot P00519 --cutoff 10000\npython scripts/query_bindingdb.py pdb 1Q0L,3ANM --cutoff 100 --identity 92\npython scripts/query_bindingdb.py compound \"&lt;SMILES&gt;\" --cutoff 0.85\n</code></pre>\n","files":[{"path":"evals/evals.json","sizeBytes":4213,"isText":true},{"path":"references/affinity_queries.md","sizeBytes":7692,"isText":true},{"path":"scripts/query_bindingdb.py","sizeBytes":2737,"isText":true},{"path":"SKILL.md","sizeBytes":14927,"isText":true}],"reviewScore":null,"reviewSummary":null,"trust":{"provenance":"trusted-source-unreviewed","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow.","bodySource":null},"bodyLocked":false,"purchaseUrl":null,"sourceUrl":null,"report":{"provenance":"trusted-source-unreviewed","screen":{"ran":true,"outcome":"clean","suspicious":0,"notes":0,"hiddenCharacters":false},"virusScan":{"engine":"clamav","status":"clean","scannedAt":"2026-09-23T18:59:19.300626Z","sha256":"2D04AFD6B2A04C81AD83AFDE24EFCDAC8701DD646B63A272EB54B5C74ED3721F","sizeBytes":11918},"review":null,"source":{"repositoryUrl":"https://github.com/AlterLab-IEU/AlterLab-Academic-Skills","path":"skills/databases/alterlab-bindingdb","license":"MIT","commit":"e4836c08a20da195a11f30f203a8cf23ec30aa95","subtreeSha":"6AC6798BFDBF1769D3B9771A534A2B2F4E223804E51D20389B9F02C618FBF200","lastSyncedAt":"2026-09-23T18:56:52.297238Z"},"reviewedAt":"2026-09-23T19:03:30.357758Z","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow."},"install":[{"target":"skills-cli","command":"npx skills add https://github.com/AlterLab-IEU/AlterLab-Academic-Skills/tree/main/skills/databases/alterlab-bindingdb"},{"target":"claude-code","command":"claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install alterlab-ieu-alterlab-academic-skills@llmmart"},{"target":"git","command":"git clone https://github.com/AlterLab-IEU/AlterLab-Academic-Skills.git"}]}