{"slug":"kermt-add-cmim-pretrain","title":"kermt-add-cmim-pretrain","summary":"Convert a grover_base checkpoint (encoder-only or encoder + vocab heads) into a hybrid checkpoint by adding a randomly-initialized cMIM decoder + latent_dist, then continue pretraining on the user's corpus as hybrid (vocab + contrast). Effectively kermt-continue-pretrain with a o","platform":"Claude","tags":[],"authorName":"LLM Mart","authorSlug":"llm-mart","score":0,"source":"github","price":null,"verified":false,"createdAt":"2026-09-26T16:30:32.069608Z","repo":{"url":"https://github.com/NVIDIA/skills","stars":3445,"forks":416,"license":"Apache-2.0","updatedAt":"2026-09-25T03:14:56Z"},"bodyHtml":"<hr>\n<p>name: kermt-add-cmim-pretrain\ndescription: Convert a grover_base checkpoint (encoder-only or encoder + vocab heads) into a hybrid checkpoint by adding a randomly-initialized cMIM decoder + latent_dist, then continue pretraining on the user's corpus as hybrid (vocab + contrast). Effectively kermt-continue-pretrain with a one-time ckpt-conversion step prepended.\nlicense: Apache-2.0\ncompatibility: Requires docker, nvidia-container-toolkit, and a CUDA-capable NVIDIA GPU. Designed for Claude Code, Codex, and Nemotron.\nmetadata:\nowner: evax@nvidia.com\nclassification: workflow-skill\nrisk_tier: skill</p>\n<h1>Line/token budget: ~165 lines, ~1900 tokens — well within the</h1>\n<h1>500-line / 5000-token cap for skill files.</h1>\n<hr>\n<h1>kermt-add-cmim-pretrain</h1>\n<p>Convert a grover_base checkpoint (legacy original-GROVER <code>grover.encoders.*</code>\nor modern <code>kermt.encoders.*</code>, with or without vocab heads) into a fully-formed\nhybrid (cMIM + vocab) checkpoint, then continue pretraining on the user's\ncorpus as hybrid.</p>\n<p>This is a thin wrapper: <code>upgrade_to_hybrid.py</code> produces a new ckpt that\nclassifies as <code>model_type: hybrid</code> via <code>check_checkpoint.py</code>, and the rest of\nthe workflow is identical to <code>kermt-continue-pretrain</code>.</p>\n<blockquote>\n<p><strong>Status: experimental.</strong> This workflow is functional end-to-end but has not\nbeen benchmarked against the manuscript's from-scratch hybrid training (which\nproduces the released checkpoint). Use as an experimental alternative to\n<code>kermt-pretrain-scratch</code> when you want to extend an existing grover_base\ncheckpoint rather than restart from random init. Validate downstream\nperformance on your own benchmark before relying on the upgraded ckpt for\nproduction work.</p>\n</blockquote>\n<h2>Skill and runtime paths</h2>\n<p>Set <code>SKILL_DIR</code> to the absolute path of this installed skill directory. Export\n<code>KERMT_REPO</code> as the absolute path to the KERMT checkout used for model\nexecution. The bundled container helper mounts that checkout at\n<code>/workspace</code> and this skill at <code>/skill</code> (read-only). Commands inside\nthe container use <code>/skill/scripts/</code>; defaults are bundled in <code>config/</code>.</p>\n<h2>Hardware requirements</h2>\n<p>Same as <code>kermt-continue-pretrain</code> (the cMIM decoder adds parameters but not\nsubstantially; VRAM headroom should be fine). The upgrade step itself is\nfast (~5 s) and CPU-only — only the subsequent continue-pretrain consumes\nGPU.</p>\n<h2>When to invoke</h2>\n<ul>\n<li>User has a grover_base checkpoint (encoder-only or with vocab heads) and\nwants to extend it into a hybrid (vocab + cMIM contrastive) pretrain.</li>\n<li>Useful for adding the SMILES-reconstruction contrastive objective to a\npretrained encoder without restarting pretraining from scratch (which\n<code>kermt-pretrain-scratch</code> would do at days-scale).</li>\n</ul>\n<p>For continuing an existing hybrid or cmim ckpt: use <code>kermt-continue-pretrain</code>\ndirectly. For training a fresh model on a custom corpus: use\n<code>kermt-pretrain-scratch</code>.</p>\n<h2>Inputs</h2>\n<p>Required:</p>\n<ul>\n<li><code>--ckpt &lt;path&gt;</code> — grover_base ckpt to upgrade. Validated via\n<code>check_checkpoint.py --mode upgrade_to_hybrid</code>; rejected if the ckpt\nalready has a contrast head or task FFN.</li>\n<li><code>--csv &lt;path&gt;</code> — pretrain corpus CSV. Same shape as\n<code>kermt-continue-pretrain</code>'s <code>--csv</code> input.</li>\n</ul>\n<p>Optional (same as <code>kermt-continue-pretrain</code>):</p>\n<ul>\n<li><code>--val-csv &lt;path&gt;</code> — separate validation CSV. Without it, prepare_data\nauto-splits by <code>--val-frac 0.1</code>.</li>\n<li>Training-hyperparameter overrides (<code>--epochs N</code>, <code>--batch-size N</code>, lr triple,\n<code>--warmup-epochs F</code>, etc.).</li>\n<li><code>--vocab-loss-weight F</code> / <code>--latent-dim N</code> / <code>--contrastive-temperature F</code>.</li>\n<li><code>--wandb-project NAME</code> / <code>--wandb-run-name NAME</code> — optional Weights &amp; Biases\nlogging (run name honored only alongside a project). Off by default.</li>\n<li><code>--gpus 0,2</code>.</li>\n</ul>\n<h2>Workflow</h2>\n<p>Let <code>$KERMT_REPO</code> be the path to your kermt repo checkout.</p>\n<ol>\n<li><p><strong>Pre-flight: check_system</strong> (same as <code>kermt-continue-pretrain</code> step 1).</p>\n</li>\n<li><p><strong>Compute run directory:</strong></p>\n<pre><code>RUN_DIR=$KERMT_REPO/runs/add-cmim-pretrain_$(date -u +%Y-%m-%dT%H-%M-%SZ)\n</code></pre>\n</li>\n<li><p><strong>Validate the input ckpt with <code>check_checkpoint --mode upgrade_to_hybrid</code>.</strong>\nAbort on <code>ok: false</code>. The validator rejects ckpts that already have\ncontrast head (suggest <code>kermt-continue-pretrain</code>) or task FFN heads\n(the ckpt has been finetuned; suggest using the original pretrain\ncheckpoint).</p>\n</li>\n<li><p><strong>Validate the corpus</strong> via <code>check_data --mode pretrain</code>. Abort on\n<code>ok: false</code>.</p>\n</li>\n<li><p><strong>Prepare the data</strong> with <code>--mode pretrain</code> — <em>without</em> <code>--vocab-dir</code>.\nThe upgrade builds fresh vocab heads sized to the corpus's vocab, so we\nwant <code>prepare_data</code> to produce a new vocab from the corpus rather than\npassing through the ckpt's old vocab (which may not even exist for\nencoder-only legacy grover_base ckpts):</p>\n<pre><code>\"$SKILL_DIR/scripts/kermt_container.sh\" run --data &lt;user-csv&gt; --run-dir $RUN_DIR -- \\\n    \"python /skill/scripts/prepare_data.py --mode pretrain \\\\\n         --csv /data/&lt;basename&gt; --out /runs/data \\\\\n         [--val-csv /data/&lt;val-basename&gt;] [--val-frac 0.1] [--seed 0]\"\n</code></pre>\n<p>The output manifest has <code>vocab_source: \"built_fresh\"</code> and includes a\n<code>smiles_vocab</code> (built from the corpus, needed for the new decoder).</p>\n</li>\n<li><p><strong>Upgrade the ckpt.</strong></p>\n<pre><code>\"$SKILL_DIR/scripts/kermt_container.sh\" run --ckpt &lt;user-ckpt&gt; --run-dir $RUN_DIR -- \\\n    \"python /skill/scripts/upgrade_to_hybrid.py \\\\\n         --ckpt /ckpt \\\\\n         --prepare-manifest /runs/data/prepare_data.json \\\\\n         --out /runs/upgraded.pt\"\n</code></pre>\n<p>Surface the JSON summary to the user — especially <code>warnings[]</code>, which\nincludes any encoder-arch drift notes (e.g. legacy GROVER had two extra\n<code>act_func_*</code> keys that modern KERMTEmbedding doesn't) and the\npretrain_ddp.py <code>--backbone</code> argparse-restriction note if the upgraded\nckpt's backbone is anything other than <code>gtrans</code>.</p>\n</li>\n<li><p><strong>Estimate runtime + confirm with the user.</strong> Same heuristic as\n<code>kermt-continue-pretrain</code> (corpus size × epochs × GPU count → wall time).</p>\n</li>\n<li><p><strong>Launch the runner detached.</strong></p>\n<pre><code>\"$SKILL_DIR/scripts/kermt_container.sh\" run_detached \\\\\n    --name kermt-add-cmim-pretrain-&lt;ts&gt; \\\\\n    --run-dir $RUN_DIR -- \\\\\n    \"python /skill/scripts/run_pretrain_local.py \\\\\n         --ckpt /runs/upgraded.pt \\\\\n         --prepare-manifest /runs/data/prepare_data.json \\\\\n         --out /runs \\\\\n         [--epochs N --batch-size N ...]\"\n</code></pre>\n<p>The runner sees the upgraded ckpt as <code>model_type: hybrid</code>, so it auto-dispatches\n<code>--pretrain_mode hybrid --vocab_loss_weight 1.0</code> with smiles_vocab plumbed\nthrough.</p>\n</li>\n<li><p><strong>Report to the user</strong> with the upgraded ckpt path + the same run.json\npointer / log path / tensorboard URL pattern as <code>kermt-continue-pretrain</code>.</p>\n</li>\n</ol>\n<h2>Hard rules</h2>\n<ul>\n<li><strong>Never modify the user's input ckpt.</strong> The upgrade writes a new file at\n<code>&lt;run_dir&gt;/upgraded.pt</code>; the source ckpt stays untouched.</li>\n<li><strong>Vocab heads are always fresh.</strong> Even if the input grover_base has vocab\nheads, they're discarded and rebuilt sized to the new corpus's vocab.\nContinue-pretraining the upgraded ckpt will train those new heads alongside\nthe decoder.</li>\n<li><strong>Don't auto-relax <code>--backbone</code> choices.</strong> If the upgrade warning fires\nbecause the input ckpt's backbone isn't <code>gtrans</code> (e.g. legacy <code>dualtrans</code>),\nsurface the warning and ask the user. Do NOT silently modify parsing.py to\nadd the legacy backbone to the choices list.</li>\n</ul>\n<h2>Common errors</h2>\n<ul>\n<li><code>check_checkpoint rejected the ckpt</code> with model_type=hybrid or cmim →\nuser's ckpt already has a contrast head. Redirect to\n<code>kermt-continue-pretrain</code>.</li>\n<li><code>check_checkpoint rejected the ckpt</code> with task_ffn=true → the ckpt has\nbeen finetuned. The upgrade workflow only supports pretrain checkpoints.</li>\n<li><code>prepare manifest missing smiles_vocab</code> → prepare_data was invoked with\n<code>--skip-vocab</code> or some equivalent that omitted the smiles vocab. Re-run\nprepare without those flags.</li>\n<li><code>unexpected key(s) in encoder load</code> warning → legacy GROVER architectures\nsaved a couple of <code>act_func_*</code> weights that modern KERMTEmbedding doesn't\nuse. Benign; the rest of the encoder loaded correctly.</li>\n</ul>\n<h2>What's in <code>run.json</code> after a successful run</h2>\n<p>Same reproducibility fields as <code>kermt-continue-pretrain</code>, plus the upgrade step's\n<code>summary.json</code> is captured under the <code>inputs.upgrade_summary</code> path so the\nprovenance of the upgraded ckpt is auditable.</p>\n<h2>Replayability</h2>\n<p>Same as <code>kermt-continue-pretrain</code>: <code>cmd_replay</code> rebuilds the\n<code>run_pretrain_local.py --ckpt &lt;upgraded.pt&gt; ...</code> invocation. To redo the\nfull add-cmim flow end-to-end, the user also needs the input grover_base\nckpt and the corpus — both are captured in the prepare_data and upgrade\nmanifests by absolute path.</p>\n","files":[{"path":"BENCHMARK.md","sizeBytes":7784,"isText":true},{"path":"config/defaults_pretrain.json","sizeBytes":2699,"isText":true},{"path":"evals/evals.json","sizeBytes":5784,"isText":true},{"path":"scripts/check_checkpoint.py","sizeBytes":20605,"isText":true},{"path":"scripts/check_data.py","sizeBytes":12185,"isText":true},{"path":"scripts/kermt_container.sh","sizeBytes":19370,"isText":true},{"path":"scripts/prepare_data.py","sizeBytes":37383,"isText":true},{"path":"scripts/run_pretrain_local.py","sizeBytes":35597,"isText":true},{"path":"scripts/upgrade_to_hybrid.py","sizeBytes":18647,"isText":true},{"path":"scripts/_utils.py","sizeBytes":14483,"isText":true},{"path":"skill-card.md","sizeBytes":4474,"isText":true},{"path":"SKILL.md","sizeBytes":8631,"isText":true},{"path":"skill.oms.sig","sizeBytes":6497,"isText":false}],"reviewScore":null,"reviewSummary":null,"trust":{"provenance":"trusted-source-unreviewed","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow.","bodySource":null},"bodyLocked":false,"purchaseUrl":null,"sourceUrl":null,"report":{"provenance":"trusted-source-unreviewed","screen":{"ran":true,"outcome":"clean","suspicious":0,"notes":0,"hiddenCharacters":false},"virusScan":{"engine":"clamav","status":"clean","scannedAt":"2026-09-26T16:30:46.878534Z","sha256":"D784F620542F78E65FCB813D8DDA25DA76A4C8494424CD9FA5737B7D26787C8D","sizeBytes":65854},"review":null,"source":{"repositoryUrl":"https://github.com/NVIDIA/skills","path":"skills/bionemo-kermt-add-cmim-pretrain","license":"Apache-2.0","commit":"d8519c57da6db5d9bea274ec1724a4a7a56a3dee","subtreeSha":"45FC28CA8F3CDFF20AAF1B884D820877C09A7434A71481AE4C4C47F5FC64C8CF","lastSyncedAt":"2026-09-26T16:30:30.547995Z"},"reviewedAt":"2026-09-26T16:31:27.673737Z","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow."},"install":[{"target":"skills-cli","command":"npx skills add https://github.com/NVIDIA/skills/tree/main/skills/bionemo-kermt-add-cmim-pretrain"},{"target":"claude-code","command":"claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install nvidia-skills@llmmart"},{"target":"git","command":"git clone https://github.com/NVIDIA/skills.git"}]}