{"slug":"kermt-monitor","title":"kermt-monitor","summary":"Check progress for a detached KERMT run (pretrain, finetune, or any kermt_run_detached invocation). Reads run.json, queries docker for container state, tails the pretrain/finetune log, and parses progress lines (epoch, step, val loss).","platform":"Claude","tags":[],"authorName":"LLM Mart","authorSlug":"llm-mart","score":0,"source":"github","price":null,"verified":false,"createdAt":"2026-09-26T16:30:32.89339Z","repo":{"url":"https://github.com/NVIDIA/skills","stars":3528,"forks":432,"license":"Apache-2.0","updatedAt":"2026-10-06T07:06:10Z"},"bodyHtml":"<hr>\n<p>name: kermt-monitor\ndescription: Check progress for a detached KERMT run (pretrain, finetune, or any kermt_run_detached invocation). Reads run.json, queries docker for container state, tails the pretrain/finetune log, and parses progress lines (epoch, step, val loss).\nlicense: Apache-2.0\ncompatibility: Requires docker and jq. Designed for Claude Code, Codex, and Nemotron.\nmetadata:\nowner: evax@nvidia.com\nclassification: atomic-skill\nrisk_tier: skill</p>\n<h1>Line/token budget: this file is targeted at ~120 lines / ~1500 tokens — well</h1>\n<h1>within the 500-line / 5000-token cap for skill files.</h1>\n<hr>\n<h1>kermt-monitor</h1>\n<p>Companion skill for any KERMT workflow that runs detached: the three pretrain\nskills (<code>kermt-continue-pretrain</code>, <code>kermt-pretrain-scratch</code>,\n<code>kermt-add-cmim-pretrain</code>) plus <code>kermt-finetune</code>. <code>kermt-infer</code> and\n<code>kermt-embed</code> run blocking by default and don't need this skill, but if a\nuser launches them detached on purpose the monitor still works (the\nworkflow-dispatch in step 4 handles unknown workflows by tailing the\nmost-recent log file in the run dir). Reads the run directory's <code>run.json</code>,\nqueries docker for the container's state, surfaces the latest progress,\nand either tails or follows the log.</p>\n<h2>Hardware requirements</h2>\n<p>None. This skill only reads disk + queries docker; no GPU compute.</p>\n<h2>Inputs</h2>\n<p>One of:</p>\n<ul>\n<li><code>&lt;run-dir&gt;</code> — a positional argument pointing at the directory containing\n<code>run.json</code> (e.g. <code>runs/continue-pretrain_2026-05-17T10-23Z</code>). Preferred.</li>\n<li><code>--container &lt;name-or-id&gt;</code> — direct container reference; the skill still\nreads <code>run.json</code> from the run dir referenced inside the container's\ninspect output if available, but works degraded-mode without it.</li>\n</ul>\n<p>Optional:</p>\n<ul>\n<li><code>--lines N</code> — number of trailing log lines to print (default 50).</li>\n<li><code>--follow</code> — stream <code>docker logs -f</code> until ^C. Useful for \"watch the\nloss\". Without it, the skill is one-shot and exits.</li>\n<li><code>--json</code> — emit a structured status report instead of human-readable text.\nUseful when the parent agent wants to take downstream action.</li>\n</ul>\n<h2>Workflow</h2>\n<p>Let <code>RUN_DIR=$1</code> (or whatever path the user supplies).</p>\n<ol>\n<li><p><strong>Locate the manifest.</strong></p>\n<pre><code>MANIFEST=$RUN_DIR/run.json\n</code></pre>\n<p>Refuse to proceed if it doesn't exist; surface a helpful message\npointing the user at the run-dir convention (<code>runs/&lt;workflow&gt;_&lt;ts&gt;/</code>).</p>\n</li>\n<li><p><strong>Parse the manifest</strong> (Python helper):</p>\n<pre><code>workflow=$(jq -r .workflow $MANIFEST)\ncontainer_name=...   # not directly in run.json today; the skill that\n                     # launched stored it in run.json under\n                     # container.name during launch (see below note).\nlogs_dir=$(jq -r .logs_dir $MANIFEST)\nimage_tag=$(jq -r .container.image_tag $MANIFEST)\nstarted_at=$(jq -r .started_at $MANIFEST)\n</code></pre>\n</li>\n<li><p><strong>Query docker for container state.</strong></p>\n<pre><code>docker ps --filter \"name=$container_name\" --format \\\n    '{{.ID}}\\t{{.Status}}\\t{{.CreatedAt}}'\n</code></pre>\n<p>If absent, fall back to <code>docker inspect $container_name --format '{{.State.Status}} (exit {{.State.ExitCode}})'</code> to see whether the\ncontainer exited (ok or failed) or was removed (<code>--rm</code> after exit).</p>\n</li>\n<li><p><strong>Find the live log file.</strong></p>\n<pre><code>case \"$workflow\" in\n  continue-pretrain|pretrain-scratch)  LOG=$logs_dir/pretrain_ddp.log ;;\n  finetune)                            LOG=$logs_dir/finetune.log ;;\n  *)                                   LOG=$(ls -1t $logs_dir/*.log 2&gt;/dev/null | head -n 1) ;;\nesac\n</code></pre>\n<p>The manifest's <code>workflow</code> field disambiguates pretrain (<code>pretrain_ddp.log</code>)\nfrom finetune (<code>finetune.log</code>). Other workflows fall back to the\nmost-recently-modified <code>.log</code> in <code>$logs_dir</code>.</p>\n</li>\n<li><p><strong>Show the latest progress.</strong></p>\n<ul>\n<li><code>tail -n $LINES $LOG</code> for the raw recent output.</li>\n<li>Parse the last few progress lines and surface a human-friendly\nsummary. The format differs per workflow:\n<ul>\n<li>Pretrain: epoch / step / val_loss\n<pre><code>Current epoch: 12/100  step: 4523/9000  val_loss: 0.832 (best 0.821 @ step 4100)\n</code></pre>\n</li>\n<li>Finetune: fold / epoch / val_</li>\n</ul>\n<pre><code>Wall-clock: 1h 23m since started_at; ETA ~6h remaining.\n</code></pre>\n</li>\n</ul>\n</li>\n<li><p><strong>Final test-metrics block (finetune, on completion).</strong> If <code>workflow</code> is\n<code>finetune</code> AND the container has exited cleanly (<code>State.Status=exited</code>,\n<code>ExitCode=0</code>) AND <code>$RUN_DIR/ckpt/fold_*/test_result.csv</code> exists, parse it\nand emit a per-task metric table:</p>\n<pre><code>Final test metrics (per task):\n  Target              MAE\n  HLM_clearance       0.187\n  RLM_clearance       0.213\n  MDR1-MDCK_efflux    0.241\n  solubility_pH6.8    0.156\n</code></pre>\n<p>The metric column matches <code>args_applied.metric</code> (mae for regression, auc\nfor classification, etc.). For multi-fold or ensemble runs, average across\nfolds/models and note <code>± std</code> if std &gt; 0. Skip silently if no\n<code>test_result.csv</code> exists (run incomplete or no test split was emitted).</p>\n</li>\n<li><p><strong>If <code>--follow</code>, stream live logs.</strong></p>\n<pre><code>docker logs -f $container_name\n</code></pre>\n<p>Wraps until ^C.</p>\n</li>\n<li><p><strong>Stop / cleanup hints</strong> (printed at end of one-shot mode):</p>\n<pre><code>To stop:        docker stop $container_name\nTo remove:     docker rm $container_name\nTo re-run:    `$(jq -r .cmd_replay $MANIFEST)`\n</code></pre>\n</li>\n</ol>\n<h2>Hard rules</h2>\n<ul>\n<li><strong>Read-only on the user's data.</strong> Never modify <code>run.json</code>, never touch the\ncontainer's checkpoint dir. The monitor only inspects.</li>\n<li><strong>Don't kill the container without explicit user instruction.</strong> If the\nuser asks to stop, run <code>docker stop</code>; if they ask to abandon, leave it\nrunning and just exit.</li>\n<li><strong>Don't pull or modify the kermt image.</strong> The monitor only reads.</li>\n<li><strong>JSON output mode is non-interactive.</strong> Skip the \"press ^C to exit\"\nprompts and emit a single JSON document so the parent agent can pipe it.</li>\n</ul>\n<h2>Note on container_name plumbing</h2>\n<p>The run.json schema as currently written does not yet include the launched\ncontainer name — <code>kermt_run_detached</code> prints it to stdout but the runner\nscript doesn't capture it into run.json. The monitor falls back to a\nfilesystem-based lookup: list <code>runs/&lt;workflow&gt;_*/</code> directories and match by\nmtime; or accept <code>--container &lt;name&gt;</code> explicitly. Follow-up: have the\nlaunching skill record container name into run.json before exiting.</p>\n<h2>Output (text mode, default)</h2>\n<pre><code>KERMT continue-pretrain · runs/continue-pretrain_2026-05-17T10-23Z\n  Container : kermt-continue-pretrain-…  (Up 1 hour, status: running)\n  Image     : kermt:latest@sha256:…\n  Repo      : 2fe00f9 (clean)\n  Started   : 2026-05-17T10:23:14Z (1h 23m ago)\n  Workflow  : continue-pretrain, pretrain_mode=hybrid, world_size=2\n\n  Latest log (last 50 lines from $LOG):\n    [Epoch 12/100] step 4523/9000 loss 0.832 lr 1.2e-4\n    [val] step 4100 val_loss 0.821 (new best)\n    ...\n\n  Progress: epoch 12/100, ~12% done. ETA ~6h.\n  TensorBoard: tensorboard --logdir $RUN_DIR/logs/tb\n  Replay command: $(jq -r .cmd_replay $RUN_DIR/run.json)\n</code></pre>\n","files":[{"path":"BENCHMARK.md","sizeBytes":7637,"isText":true},{"path":"evals/evals.json","sizeBytes":4669,"isText":true},{"path":"skill-card.md","sizeBytes":4422,"isText":true},{"path":"SKILL.md","sizeBytes":7118,"isText":true},{"path":"skill.oms.sig","sizeBytes":4581,"isText":false}],"reviewScore":null,"reviewSummary":null,"trust":{"provenance":"trusted-source-unreviewed","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow.","bodySource":null},"bodyLocked":false,"purchaseUrl":null,"sourceUrl":null,"report":{"provenance":"trusted-source-unreviewed","screen":{"ran":true,"outcome":"clean","suspicious":0,"notes":0,"hiddenCharacters":false},"virusScan":{"engine":"clamav","status":"clean","scannedAt":"2026-09-26T16:30:58.098533Z","sha256":"EABAFB11403A5F6E115D1DD300B9DB35EFAC223C0616269AA019B70F011C0220","sizeBytes":12833},"review":null,"source":{"repositoryUrl":"https://github.com/NVIDIA/skills","path":"skills/bionemo-kermt-monitor","license":"Apache-2.0","commit":"0e0d506f4eb67204a62586ac5f19df3cb7ad9b1f","subtreeSha":"FA7E3532389BDDCEA02B7D354133B2A1C885D45C30947768E5CF4FB368A8470B","lastSyncedAt":"2026-10-06T15:23:58.01595Z"},"reviewedAt":"2026-09-26T16:31:47.632423Z","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow."},"install":[{"target":"skills-cli","command":"npx skills add https://github.com/NVIDIA/skills/tree/main/skills/bionemo-kermt-monitor"},{"target":"claude-code","command":"claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install nvidia-skills@llmmart"},{"target":"git","command":"git clone https://github.com/NVIDIA/skills.git"}]}