spark-memory-thermal-ops
Manage unified memory and thermals during long-running ML jobs on NVIDIA DGX Spark. Use when planning memory headroom for a training run on GB10, when a job OOMs on unified memory, or when monitoring temperature and power during multi-hour training.
Install
npx skills add https://github.com/wshobson/agents/tree/main/plugins/dgx-spark-ops/skills/spark-memory-thermal-ops
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install wshobson-agents@llmmart
git clone https://github.com/wshobson/agents.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole wshobson/agents collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Spark Memory & Thermal Ops
DGX Spark's GB10 chip has one 128GB unified
memory (UMA) pool shared by CPU and GPU, and a
sustained power ceiling well below its rated
figure. Both break discrete-GPU assumptions:
headroom isn't what nvidia-smi reports, and a
run that starts fast will slow down mid-job
with nothing misconfigured. This skill covers
planning memory headroom, working an actual
OOM, and watching thermals across a long job.
For launch-time failure modes (ABI mismatches,
flash-attn, playbook breakage), see
spark-training-gotchas — this skill assumes
the job starts.
Common Issues Quick Reference
| Situation | Do this |
|---|---|
| Planning headroom before launch | Budget against free -g, not nvidia-smi — see UMA Memory Model |
| Job OOMs on unified memory | Work the OOM Ladder in order: flush, then batch/pack, then method downgrade |
| Throughput drops mid-run | Check the power/temp log before assuming a config bug — see Thermal Monitoring |
| Trainer + inference server both wanted | Run one at a time — see Concurrent Workloads |
When to Use This Skill
- Sizing a training run against the 128GB pool before launch — will this model, method, and batch/pack combination fit.
- A run OOMs mid-load or mid-step and the remediation order matters — what to try first, second, third.
- Watching temperature and power during a multi-hour job, deciding whether a slowdown is thermal throttling or something else.
- Planning to run a trainer alongside an inference server (vLLM, Ollama) on the same box.
UMA Memory Model
Spark has no separate GPU VRAM — the GPU and CPU share one 128GB pool. Two consequences:
nvidia-smiandcudaMemGetInfounderreport pressure — or report nothing at all. Both report CUDA-allocator-visible memory, not the pool's actual state — a box can show headroom innvidia-smiand still OOM, because page-cache and mmap'd pages the allocator doesn't see consume the same pool. On some driver/setups, the memory query returns[N/A], [N/A]outright instead of a number — a script grepping for a numeric value there gets nothing, not a misleading undercount (seespark-training-gotchasgotcha G3).Model load is a transient peak, not the steady state. Loading safetensors weights mmaps the file, then copies into CUDA tensors — for a window during load, both the mmap'd pages and the CUDA copy count against the pool at once. A model that fits while training can still OOM during load if headroom was sized for the post-load footprint instead of this doubled transient.
Plan and diagnose with free -g, not
nvidia-smi:
free -g | awk 'NR==2 {print "free:", $4, "GB"}'
Rule of thumb: take that free figure, subtract a
few GB for OS/driver overhead, and budget against
the result — not the 128GB spec number.
The worksheet in references/uma-accounting.md
accepts parameter count, dtype, and method as
input, and returns a memory estimate to compare
against known anchors.
Planning Sequence
Before launch, work through these in order:
- Read
free -g; subtract OS/driver overhead for the budget. - Estimate weights + optimizer + gradients +
activations from
references/uma-accounting.md. - Compare against the closest anchor (70B QLoRA, 27B LoRA, 9B full FT), not the estimate alone.
- If the estimate is close to the budget, start with shorter packing or a smaller batch — cheaper than hitting the OOM Ladder mid-run.
Example: Sizing a 70B QLoRA Run
A sanity check of the worksheet formula against the ≈40GB anchor:
params = 70e9
weights_gb = params * 0.5 / 1e9 # NF4, step 1
adapter_gb = 0.5 # step 5, negligible
total_gb = weights_gb + adapter_gb # + activations
print(f"{total_gb:.0f}GB before activations")
Weights alone land near the ≈40GB anchor — a plan estimating far above that for the same model class is a signal to recheck dtype and method.
The OOM Ladder
When a job OOMs on unified memory, work this ladder in order. Each step is more disruptive than the last — don't skip ahead: reducing batch size is never step 1.
Flush the buffer cache. Page cache from a previous run or a large dataset read often accounts for GB of the "missing" headroom. This costs nothing but a rerun and doesn't touch the job's configuration:
sync; echo 3 > /proc/sys/vm/drop_cachesNeeds root; a between-run reset, not a mid-training step. See
spark-training-gotchas(gotcha G3) for the full diagnostic behind this step.Reduce batch size or packing length. Only after a flush fails to free enough headroom, cut batch size or packing length — the first step that changes what the run does. Prefer packing length first; it drives activation footprint more directly at long context.
Downgrade the method: bf16 LoRA before QLoRA. If flushing and shrinking batch/pack still OOM, drop the method a tier — bf16 LoRA is next, not the reverse. QLoRA's bitsandbytes dequantization buffers are transient CUDA-side allocations that can OOM before an equivalent bf16 LoRA run would, even though QLoRA's steady-state footprint is smaller. A QLoRA OOM is not proof the model doesn't fit.
Fall back further (smaller model, multi-Spark) only after all three steps and the job still won't fit.
Thermal Monitoring
Multi-hour runs push into Spark's sustained power ceiling, well under the rated figure — expected platform behavior, not a symptom to explain away:
Sample temperature and power alongside the training logs, not after a slowdown is noticed — every 30-60 seconds correlates a throughput drop with a thermal event. Keep the CSV output format
assets/thermal-sample.shwrites, so timestamps line up against the log:bash assets/thermal-sample.sh 30 thermal.logA sustained ~100W power draw is the platform cap, not a configuration bug. Don't re-tune batch size or precision to "fix" a plateau that's the box behaving normally under load. If temperature climbs while power stays flat under the rated 240W figure, that's the signature to recognize.
Log throttle events explicitly instead of letting a run silently slow down unrecorded. A run whose per-step time doubles two hours in should show that in the log, correlated against the thermal sample at that timestamp. Full throttling diagnostics:
spark-training-gotchas(gotcha G4).
Concurrent Workloads
Because the 128GB pool is global, eviction happens without either process's logs showing an OOM:
The one-heavy-job rule applies to uncapped or near-capacity workloads — an uncapped trainer and inference server (vLLM, Ollama) compete for the same pool. A small, capped workload doesn't: a <4GB LoRA fine-tune coexists fine alongside vLLM capped at
gpu-memory-utilization<=0.5— check the other process's cap, not just its presence, before stopping it.Inference servers evict trainer pages silently under uncapped/near-capacity contention, and vice versa — neither logs an error, so a slow run or lost KV cache is a contention symptom to check for. Stop unrelated uncapped servers before a long or full-pool run.
Check for GPU-resident processes first:
ps aux | grep -E 'vllm|ollama|trl|axolotl' | grep -v grep
This procedure complements spark-training-gotchas
(gotchas G3, G4, G6) — that skill covers launch-time
failures; this one, the running job.
Memory math worksheets:
references/uma-accounting.md.
Files (agents)
-
assets
-
thermal-sample.sh 1017 B
#!/usr/bin/env bash # Background thermal/power sampler for a long training run. # Output contract: CSV, one line per sample, matching # `nvidia-smi --query-gpu` field order — pipe-compatible with # the training log for later correlation by timestamp. # Usage: bash thermal-sample.sh [interval_seconds] [logfile] set -uo pipefail command -v nvidia-smi >/dev/null || { echo "nvidia-smi not found" >&2; exit 1; } INTERVAL="${1:-30}" LOGFILE="${2:-thermal.log}" PIDFILE="${LOGFILE}.pid" if [ -f "$PIDFILE" ] && kill -0 "$(cat "$PIDFILE")" 2>/dev/null; then echo "Stopping existing sampler (pid $(cat "$PIDFILE"))" kill "$(cat "$PIDFILE")" fi echo "timestamp,temperature.gpu,power.draw" > "$LOGFILE" nvidia-smi --query-gpu=timestamp,temperature.gpu,power.draw \ --format=csv,noheader -l "$INTERVAL" >> "$LOGFILE" & PID=$! echo "$PID" > "$PIDFILE" echo "Sampling every ${INTERVAL}s into ${LOGFILE} (pid $PID)" echo "A sustained ~100W reading is the platform cap, not a bug — see SKILL.md Thermal Monitoring."
-
-
references
-
uma-accounting.md 4.3 KB
Last verified: 2026-07-13 — refresh when a new model-size anchor is validated on Spark or the DGX Spark playbooks change quantization defaults. # UMA Memory Accounting A worksheet for estimating whether a model/method/batch combination fits DGX Spark's 128GB unified pool before a run, and for sanity checking a plan against known-working anchors. This is planning math, not a guarantee — always leave headroom rather than sizing to the byte; see the OOM Ladder in `SKILL.md` for what to do when the estimate turns out optimistic. ## The Four Terms Total footprint ≈ **weights + optimizer states + gradients + activations**, plus a near-zero term for LoRA/QLoRA adapters. Work each term from parameter count and dtype, then sum. ### 1. Weights `params × bytes/param`, by dtype: | dtype | bytes/param | |---|---| | fp32 | 4 | | bf16 / fp16 | 2 | | int8 | 1 | | int4 (QLoRA NF4) | 0.5 | A 70B model in bf16 is ~140GB — already over the pool before anything else loads. The same model in 4-bit (QLoRA) is ~35GB, which is why QLoRA, not bf16, is what makes 70B-class models reachable on Spark at all. ### 2. Optimizer states Full fine-tuning carries optimizer state for every trainable parameter; LoRA and QLoRA carry it only for the adapter parameters, which is why this term is negligible for them regardless of base model size. | Optimizer | bytes/param (trainable only) | |---|---| | AdamW, fp32 states | 8 (4B momentum + 4B variance) | | AdamW 8-bit (bitsandbytes) | ≈2 (quantized momentum + variance) | `adamw_8bit` is the Unsloth default for a reason on a 128GB box — the fp32 variant roughly quadruples this term for any run that isn't LoRA/QLoRA-adapter-only. ### 3. Gradients Same dtype as the compute precision — typically bf16, so 2 bytes/param — and, like optimizer state, only for trainable parameters. Full fine-tuning pays this for every weight; LoRA and QLoRA pay it only for the adapter, since the frozen base weights never accumulate a gradient. ### 4. Activations The hardest term to pin to a single number — it scales with batch size, sequence/packing length, and architecture (attention variant, hidden size, layer count), not just parameter count. Two levers matter more than exact estimation: - **Gradient checkpointing** trades recompute for memory: expect roughly **30% savings** on this term versus no checkpointing, at the cost of a recompute pass per checkpointed segment. Unsloth's `use_gradient_checkpointing="unsloth"` is the default for this reason. - Packing/sequence length is the more direct lever than batch size for this term — see the OOM Ladder in `SKILL.md` for why packing length is the preferred first cut over batch size. ### 5. LoRA/QLoRA adapter overhead A rank-`r` adapter on a linear layer adds `r × (in + out)` parameters — `A` is `r×in` and `B` is `out×r`, so the two matrices together contribute `r·in + r·out`. At the rank sizes in normal use (1–32 for RL, up to ~256 for SFT-at-scale), this is a small fraction of a percent of base model size — round it to zero in the worksheet unless an unusually high rank is in play. ## Anchors Known-working combinations on a single Spark, to sanity check a new plan against rather than trusting the formula in isolation: | Model class | Method | Observed total | Notes | |---|---|---|---| | 70B | QLoRA | ≈40GB | 30–48h for 3 epochs; the reference point for "70B fits via QLoRA, not bf16." | | a ~120B-class MoE model | NVFP4-native LoRA | ≈68GB (per Unsloth's official DGX Spark tutorial, unsloth.ai docs, 2025-12) | Community recipe (`nvfp4-lora-spark`); experimental, not the default assumption for other 100B+ MoE models. | | 27B | LoRA | fits at pack ≤1024 | The LoRA ceiling on a single Spark — larger dense models need multi-Spark or a smaller method. | | 9B | Full fine-tune | fits comfortably | The full-FT ceiling — above this, full FT needs LoRA/QLoRA or multi-Spark instead. | Treat "ceiling" entries as the largest class that fit in practice, not a hard architectural limit — a smaller batch, shorter packing, or a leaner optimizer can sometimes push slightly past one of these, and a heavier configuration of the same model class can fail well under it. Re-derive from the four terms above for anything outside these four reference points, and confirm with the OOM Ladder in `SKILL.md` if the estimate turns out optimistic.
-
-
SKILL.md 7.8 KB
--- name: spark-memory-thermal-ops description: Manage unified memory and thermals during long-running ML jobs on NVIDIA DGX Spark. Use when planning memory headroom for a training run on GB10, when a job OOMs on unified memory, or when monitoring temperature and power during multi-hour training. --- # Spark Memory & Thermal Ops DGX Spark's GB10 chip has one 128GB unified memory (UMA) pool shared by CPU and GPU, and a sustained power ceiling well below its rated figure. Both break discrete-GPU assumptions: headroom isn't what `nvidia-smi` reports, and a run that starts fast will slow down mid-job with nothing misconfigured. This skill covers planning memory headroom, working an actual OOM, and watching thermals across a long job. For launch-time failure modes (ABI mismatches, flash-attn, playbook breakage), see `spark-training-gotchas` — this skill assumes the job starts. ## Common Issues Quick Reference | Situation | Do this | |---|---| | Planning headroom before launch | Budget against `free -g`, not `nvidia-smi` — see UMA Memory Model | | Job OOMs on unified memory | Work the OOM Ladder in order: flush, then batch/pack, then method downgrade | | Throughput drops mid-run | Check the power/temp log before assuming a config bug — see Thermal Monitoring | | Trainer + inference server both wanted | Run one at a time — see Concurrent Workloads | ## When to Use This Skill - Sizing a training run against the 128GB pool before launch — will this model, method, and batch/pack combination fit. - A run OOMs mid-load or mid-step and the remediation order matters — what to try first, second, third. - Watching temperature and power during a multi-hour job, deciding whether a slowdown is thermal throttling or something else. - Planning to run a trainer alongside an inference server (vLLM, Ollama) on the same box. ## UMA Memory Model Spark has no separate GPU VRAM — the GPU and CPU share one 128GB pool. Two consequences: - **`nvidia-smi` and `cudaMemGetInfo` underreport pressure — or report nothing at all.** Both report CUDA-allocator-visible memory, not the pool's actual state — a box can show headroom in `nvidia-smi` and still OOM, because page-cache and mmap'd pages the allocator doesn't see consume the same pool. On some driver/setups, the memory query returns `[N/A], [N/A]` outright instead of a number — a script grepping for a numeric value there gets nothing, not a misleading undercount (see `spark-training-gotchas` gotcha G3). - **Model load is a transient peak, not the steady state.** Loading safetensors weights mmaps the file, then copies into CUDA tensors — for a window during load, both the mmap'd pages and the CUDA copy count against the pool at once. A model that fits while training can still OOM during load if headroom was sized for the post-load footprint instead of this doubled transient. Plan and diagnose with `free -g`, not `nvidia-smi`: ```bash free -g | awk 'NR==2 {print "free:", $4, "GB"}' ``` Rule of thumb: take that free figure, subtract a few GB for OS/driver overhead, and budget against the result — not the 128GB spec number. The worksheet in `references/uma-accounting.md` accepts parameter count, dtype, and method as input, and returns a memory estimate to compare against known anchors. ### Planning Sequence Before launch, work through these in order: 1. Read `free -g`; subtract OS/driver overhead for the budget. 2. Estimate weights + optimizer + gradients + activations from `references/uma-accounting.md`. 3. Compare against the closest anchor (70B QLoRA, 27B LoRA, 9B full FT), not the estimate alone. 4. If the estimate is close to the budget, start with shorter packing or a smaller batch — cheaper than hitting the OOM Ladder mid-run. ### Example: Sizing a 70B QLoRA Run A sanity check of the worksheet formula against the ≈40GB anchor: ```python params = 70e9 weights_gb = params * 0.5 / 1e9 # NF4, step 1 adapter_gb = 0.5 # step 5, negligible total_gb = weights_gb + adapter_gb # + activations print(f"{total_gb:.0f}GB before activations") ``` Weights alone land near the ≈40GB anchor — a plan estimating far above that for the same model class is a signal to recheck dtype and method. ## The OOM Ladder When a job OOMs on unified memory, work this ladder in order. Each step is more disruptive than the last — don't skip ahead: **reducing batch size is never step 1.** 1. **Flush the buffer cache.** Page cache from a previous run or a large dataset read often accounts for GB of the "missing" headroom. This costs nothing but a rerun and doesn't touch the job's configuration: ```bash sync; echo 3 > /proc/sys/vm/drop_caches ``` Needs root; a between-run reset, not a mid-training step. See `spark-training-gotchas` (gotcha G3) for the full diagnostic behind this step. 2. **Reduce batch size or packing length.** Only after a flush fails to free enough headroom, cut batch size or packing length — the first step that changes what the run does. Prefer packing length first; it drives activation footprint more directly at long context. 3. **Downgrade the method: bf16 LoRA before QLoRA.** If flushing and shrinking batch/pack still OOM, drop the method a tier — bf16 LoRA is next, not the reverse. QLoRA's bitsandbytes dequantization buffers are transient CUDA-side allocations that can OOM before an equivalent bf16 LoRA run would, even though QLoRA's steady-state footprint is smaller. A QLoRA OOM is not proof the model doesn't fit. Fall back further (smaller model, multi-Spark) only after all three steps and the job still won't fit. ## Thermal Monitoring Multi-hour runs push into Spark's sustained power ceiling, well under the rated figure — expected platform behavior, not a symptom to explain away: - Sample temperature and power alongside the training logs, not after a slowdown is noticed — every 30-60 seconds correlates a throughput drop with a thermal event. Keep the CSV output format `assets/thermal-sample.sh` writes, so timestamps line up against the log: ```bash bash assets/thermal-sample.sh 30 thermal.log ``` - **A sustained ~100W power draw is the platform cap, not a configuration bug.** Don't re-tune batch size or precision to "fix" a plateau that's the box behaving normally under load. If temperature climbs while power stays flat under the rated 240W figure, that's the signature to recognize. - Log throttle events explicitly instead of letting a run silently slow down unrecorded. A run whose per-step time doubles two hours in should show that in the log, correlated against the thermal sample at that timestamp. Full throttling diagnostics: `spark-training-gotchas` (gotcha G4). ## Concurrent Workloads Because the 128GB pool is global, eviction happens without either process's logs showing an OOM: - The one-heavy-job rule applies to **uncapped or near-capacity** workloads — an uncapped trainer and inference server (vLLM, Ollama) compete for the same pool. A small, capped workload doesn't: a <4GB LoRA fine-tune coexists fine alongside vLLM capped at `gpu-memory-utilization<=0.5` — check the other process's cap, not just its presence, before stopping it. - Inference servers evict trainer pages silently under uncapped/near-capacity contention, and vice versa — neither logs an error, so a slow run or lost KV cache is a contention symptom to check for. Stop unrelated *uncapped* servers before a long or full-pool run. Check for GPU-resident processes first: ```bash ps aux | grep -E 'vllm|ollama|trl|axolotl' | grep -v grep ``` This procedure complements `spark-training-gotchas` (gotchas G3, G4, G6) — that skill covers launch-time failures; this one, the running job. Memory math worksheets: `references/uma-accounting.md`.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.