{"slug":"dataset-curation","title":"dataset-curation","summary":"Prepare, format, and validate datasets for supervised fine-tuning and preference training. Use when converting raw data into training format, applying chat templates, configuring sequence packing, generating synthetic training data, or writing a dataset card before a run.","platform":"Claude","tags":[],"authorName":"LLM Mart","authorSlug":"llm-mart","score":0,"source":"github","price":null,"verified":false,"createdAt":"2026-09-01T18:59:30.011487Z","repo":{"url":"https://github.com/wshobson/agents","stars":40003,"forks":4267,"license":"MIT","updatedAt":"2026-09-26T19:54:17Z"},"bodyHtml":"<hr>\n<h2>name: dataset-curation\ndescription: Prepare, format, and validate datasets for supervised fine-tuning and preference training. Use when converting raw data into training format, applying chat templates, configuring sequence packing, generating synthetic training data, or writing a dataset card before a run.</h2>\n<h1>Dataset Curation</h1>\n<p>This skill assumes <code>finetuning-method-selection</code>\nalready routed here — the next step is preparing\ndata, not choosing a method. What follows: format\nselection by target method, the template/packing\nmechanics behind the most common silent training\nfailures, rules for mixing in synthetic data\nwithout collapse, and the dataset card that closes\nout Phase 2 before a run starts.</p>\n<p><strong>Input:</strong> raw examples (demonstrations, preference\njudgments, or task prompts) plus a routing decision\nfrom <code>finetuning-method-selection</code>.\n<strong>Output format:</strong> a formatted, packed, validated\nJSONL dataset plus a completed dataset card — the\nPhase 2 artifact <code>/finetune</code> checks before launching\ntraining.</p>\n<h2>Format Selection</h2>\n<table>\n<thead>\n<tr>\n<th>Method</th>\n<th>Shape</th>\n<th>Rows</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>SFT, single-turn</td>\n<td>Instruct (<code>instruction</code>/<code>response</code> or <code>prompt</code>/<code>completion</code>)</td>\n<td>~1,000+ floor</td>\n</tr>\n<tr>\n<td>SFT, multi-turn</td>\n<td>Conversation / ChatML <code>messages</code> list</td>\n<td>~1,000+ floor</td>\n</tr>\n<tr>\n<td>DPO / ORPO</td>\n<td>Preference pair (<code>prompt</code>, <code>chosen</code>, <code>rejected</code>)</td>\n<td>Method-dependent, see <code>preference-optimization</code></td>\n</tr>\n<tr>\n<td>KTO</td>\n<td>Unpaired (<code>prompt</code>, <code>completion</code>, <code>label</code>)</td>\n<td>Method-dependent, see <code>preference-optimization</code></td>\n</tr>\n<tr>\n<td>GRPO / RLVR</td>\n<td>Prompt-only (<code>prompt</code> + verifier metadata)</td>\n<td>Method-dependent, see <code>grpo-rlvr-training</code></td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li><p><strong>~1,000+ rows is the recommended floor for SFT</strong>,\nnot a target. Below it, a handful of low-quality\nor duplicate examples can dominate the gradient;\nabove it, <strong>quality over quantity</strong> — a smaller\nverified, deduplicated set beats a larger noisy one.</p>\n</li>\n<li><p>The ChatML shape, for orientation; the other four\nformats plus a ShareGPT conversion note live in\n<code>references/formats-and-templates.md</code>:</p>\n<pre><code>{\"messages\": [\n  {\"role\": \"user\", \"content\": \"...\"},\n  {\"role\": \"assistant\", \"content\": \"...\"}\n]}\n</code></pre>\n</li>\n</ul>\n<h2>Chat Templates and Loss Masking</h2>\n<p>Apply the target model's chat template <strong>before</strong>\nany concatenation or packing, never after — packing\nraw text and templating the packed blob afterward\ncorrupts turn boundaries, landing role markers in\nthe wrong place relative to each example.</p>\n<ul>\n<li><p><strong>Train on assistant responses only.</strong> Mask the\nloss (<code>-100</code> in the labels tensor) over system/user\nturns and the template's own role markers — only\nassistant-turn content tokens contribute to loss.</p>\n</li>\n<li><p><strong>Template/tokenizer mismatches are a top silent\nfailure mode.</strong> A model trained against one chat\ntemplate but served or evaluated with a different\none degrades without erroring. Verify the same\ntemplate string used in training is applied at\ninference and eval time.</p>\n</li>\n<li><p><strong>Keep the dataset in <code>messages</code> shape</strong> and let\nthe trainer template and mask it\n(<code>assistant_only_loss=True</code> in current TRL) —\npre-rendering to a flat text field destroys the\nturn boundaries masking needs. Full code sketch:\n<code>references/formats-and-templates.md</code>. Sanity-check\nbefore training — decode only unmasked positions;\nexpect only assistant text:</p>\n<pre><code>keep = batch[\"labels\"][0] != -100\nprint(tokenizer.decode(batch[\"input_ids\"][0][keep]))\n</code></pre>\n</li>\n</ul>\n<h2>Packing</h2>\n<p><strong>Without packing, 40–70% of compute is spent on\npadding</strong> — variable-length examples batched at a\nfixed sequence length waste the gap between each\nexample's length and the batch's max. Packing\nconcatenates multiple examples into one sequence\nup to the max length, cutting most of that waste.</p>\n<ul>\n<li><p><strong>Packing changes batch semantics.</strong> A packed\nsequence can contain several original examples, so\n\"steps per epoch\" and any LR schedule keyed to\nexample count shift once packing is on — recompute\nschedule milestones against packed-sequence count.</p>\n</li>\n<li><p><strong>MANDATORY: decode and manually inspect 5–10\npacked sequences before scaling to a full run.</strong>\nConfirm example boundaries land where expected,\ntemplate markers are intact per sub-example, and\nthe loss mask is still assistant-only within each\npacked sequence. Not optional — packing bugs are\nsilent (the loss curve looks normal) and only\nsurface in eval quality, hours later:</p>\n<pre><code>for seq in packed_dataset.select(range(10)):\n    print(tokenizer.decode(seq[\"input_ids\"]))\n</code></pre>\n</li>\n</ul>\n<h2>Synthetic Data Rules</h2>\n<ul>\n<li><strong>Keep ≥25% real data as a collapse guard.</strong>\nTraining on a growing share of model-generated\ndata without a real-data floor drives measurable\nquality collapse over successive generations —\n25% real is the minimum that holds the line.\n<strong>General-domain replay rows\ncount toward this floor</strong> —\n\"real\" means \"not generated\nfor this task from this\nstudent,\" not \"human-authored.\"\nAn all-synthetic-by-construction\ndataset can meet the ≥25% floor\nthrough replay alone (see\n<code>references/synthetic-data.md</code>'s\nReplay-Mix Construction recipe);\nstate which rows count as \"real\"\nin the dataset card rather than\nleaving the floor structurally\nunmeetable.</li>\n<li><strong>Magpie and rejection sampling are the\nworkhorses.</strong> Magpie extracts prompts from the\nmodel's own template prior; rejection sampling\ngenerates several candidates per prompt and keeps\nonly the ones a filter passes. Both beat naive\nsingle-shot generation.</li>\n<li><strong>Targeted, student-aware generation beats static\ngeneration by 1.3–2x sample efficiency</strong> — aiming\nat the student's actual failure modes hits a\nquality bar with fewer filtered examples.</li>\n<li><strong>Typical accept rates after filtering run\n10–30%.</strong> Plan volume accordingly — a 10,000-row\ntarget at 15% accept needs ~65,000+ raw generations.</li>\n<li>Generation-method ranking, filter funnel, replay-\nmix construction, and distillation pattern:\n<code>references/synthetic-data.md</code>.</li>\n</ul>\n<h2>The Dataset Card</h2>\n<p>Every dataset that reaches training gets a card —\nthe required Phase 2 artifact <code>/finetune</code> checks\nbefore launching. The card is not free-form\ndocumentation; it MUST carry these fields:</p>\n<ul>\n<li><strong>Provenance</strong> — where every row came from (real\nsource(s), synthetic method(s), or both),\ntraceable to <code>trace-to-training-data</code> output.</li>\n<li><strong>Counts</strong> — total rows, and rows per split\n(train/eval/held-out) if split.</li>\n<li><strong>Synthetic/real ratio</strong> — the measured ratio,\nchecked against the ≥25% real floor above.</li>\n<li><strong>Dedup method</strong> — exact-match, semantic\n(embedding threshold), or both; see the filter\nfunnel in <code>references/synthetic-data.md</code>.</li>\n<li><strong>Template used</strong> — the exact chat template\nstring/identifier, kept consistent through\ninference and eval — this is what ties an\n<code>eval-harness-first</code> run back to the checkpoint.</li>\n<li><strong>Packing config</strong> — whether packing was used,\nmax sequence length, and confirmation the\n5–10-sequence manual inspection above was done.</li>\n</ul>\n<p>A dataset missing any of these six fields isn't\nready for <code>/finetune</code> — the card is a gate, not a\nsummary written after the fact.</p>\n<h3>Phase 2 Exit Checklist</h3>\n<p>Before handing off to <code>/finetune</code>, confirm:</p>\n<ol>\n<li>Format matches the method (table above).</li>\n<li>Template applied before concatenation.</li>\n<li>Loss masked to assistant turns only.</li>\n<li>5–10 packed sequences decoded and read.</li>\n<li>≥25% real data in the final mix.</li>\n<li>Dataset card complete — all six fields.</li>\n</ol>\n<h2>References</h2>\n<ul>\n<li><code>references/formats-and-templates.md</code> — JSONL\nexamples per format, current-TRL masking code,\nand the ShareGPT conversion note.</li>\n<li><code>references/synthetic-data.md</code> — generation-method\nranking, filter funnel, replay-mix construction,\nand teacher→student distillation pattern.</li>\n</ul>\n<p>Related skills: <code>finetuning-method-selection</code> routes\nhere; <code>lora-qlora-recipes</code>, <code>vision-sft</code>, and\n<code>preference-optimization</code> consume the datasets this\nskill produces; <code>trace-to-training-data</code> is the\nprovenance source for graded-trajectory datasets;\n<code>eval-harness-first</code> grades the resulting checkpoint.</p>\n","files":[{"path":"references/formats-and-templates.md","sizeBytes":6509,"isText":true},{"path":"references/synthetic-data.md","sizeBytes":12834,"isText":true},{"path":"SKILL.md","sizeBytes":7991,"isText":true}],"reviewScore":null,"reviewSummary":null,"trust":{"provenance":"trusted-source-unreviewed","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow.","bodySource":null},"bodyLocked":false,"purchaseUrl":null,"sourceUrl":null,"report":{"provenance":"trusted-source-unreviewed","screen":{"ran":true,"outcome":"clean","suspicious":0,"notes":0,"hiddenCharacters":false},"virusScan":{"engine":"clamav","status":"clean","scannedAt":"2026-09-01T19:00:34.614056Z","sha256":"A64D5D6D6899E3DBC4F2A9ABBF6F1020C38C3F3AB670ED0577FD7142547261CF","sizeBytes":12378},"review":null,"source":{"repositoryUrl":"https://github.com/wshobson/agents","path":"plugins/llm-finetuning/skills/dataset-curation","license":"MIT","commit":"9b15b34b0bfc13a815cbfc2366e14ea549e09422","subtreeSha":"6E6D58B93D6F9460536EFE6EBA2BAF1AC44EA06A4C48F349F0548C10AB067C43","lastSyncedAt":"2026-09-26T23:12:03.520842Z"},"reviewedAt":"2026-09-01T19:03:23.703071Z","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow."},"install":[{"target":"skills-cli","command":"npx skills add https://github.com/wshobson/agents/tree/main/plugins/llm-finetuning/skills/dataset-curation"},{"target":"claude-code","command":"claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install wshobson-agents@llmmart"},{"target":"git","command":"git clone https://github.com/wshobson/agents.git"}]}