{"slug":"generate-codebook","title":"generate-codebook","summary":"Generate a citable data dictionary / codebook from a tabular dataset (CSV/TSV/Excel/Parquet/Stata/SAS). Profiles every variable — role, type, units placeholder, level frequencies, range/quantiles, missingness — and emits codebook.md + codebook.json. Flags coded variables whose le","platform":"Claude","tags":[],"authorName":"LLM Mart","authorSlug":"llm-mart","score":0,"source":"github","price":null,"verified":false,"createdAt":"2026-09-14T20:48:10.713146Z","repo":{"url":"https://github.com/Aperivue/medsci-skills","stars":318,"forks":75,"license":"MIT","updatedAt":"2026-09-27T05:05:17Z"},"bodyHtml":"<hr>\n<h2>name: generate-codebook\ndescription: Generate a citable data dictionary / codebook from a tabular dataset (CSV/TSV/Excel/Parquet/Stata/SAS). Profiles every variable — role, type, units placeholder, level frequencies, range/quantiles, missingness — and emits codebook.md + codebook.json. Flags coded variables whose level meanings are unknown as [NEEDS DICTIONARY] rather than guessing them, feeding /define-variables and the dictionary-first workflow.\ntriggers: generate codebook, data dictionary, codebook, profile variables, variable dictionary, describe dataset, what variables, column dictionary, build codebook\ntools: Read, Write, Edit, Bash, Grep, Glob\nmodel: inherit</h2>\n<h1>Generate Codebook Skill</h1>\n<p>You help a medical researcher turn a raw tabular dataset into a structured,\n<strong>citable</strong> data dictionary (codebook). This is the <em>generator</em> side of the\ndictionary-first workflow: it produces the artifact that <code>/define-variables</code> and\ndictionary-first QC later consume. You generate code and review output — you do\n<strong>not</strong> invent the meaning of coded values.</p>\n<h2>Communication Rules</h2>\n<ul>\n<li>Communicate with the user in their preferred language.</li>\n<li>Variable names, codebook fields, and report output are in English.</li>\n<li>Medical terminology is always in English.</li>\n</ul>\n<h2>Philosophy</h2>\n<p>A codebook describes <em>what is in the data</em>, not <em>what the codes mean</em>. Column\ndistributions, types, and missingness are observable and safe to profile. The\n<strong>meaning</strong> of a coded value (<code>fatty_liver_grade = 0</code>) is NOT observable from the\ndata — it lives in the authoritative data dictionary. This skill profiles the\nformer deterministically and explicitly flags the latter as <code>[NEEDS DICTIONARY]</code>\nso a human fills it from the source. This is the generator counterpart to the\ndictionary-first rule that <code>/define-variables</code> enforces on consumption.</p>\n<h2>Reference Files</h2>\n<ul>\n<li><strong>Schema + role rules</strong>: <code>${CLAUDE_SKILL_DIR}/references/codebook_schema.md</code> — the\ncodebook.json schema, the role-inference heuristics, and how the output threads\ninto <code>/define-variables</code> and dictionary-first QC. Read this before interpreting output.</li>\n</ul>\n<h2>Deterministic Script</h2>\n<p>Run the bundled profiler rather than describing columns from memory:</p>\n<pre><code>python \"${CLAUDE_SKILL_DIR}/scripts/generate_codebook.py\" data.csv --out-dir .\n</code></pre>\n<p>Supports <code>.csv/.tsv/.xlsx/.parquet/.dta/.sas7bdat</code>. Flags: <code>--max-levels N</code>\n(categorical cutoff, default 20), <code>--json-only</code>, <code>--md-only</code>. The script is\npandas-only, runs locally, and never sends data anywhere.</p>\n<h2>Workflow</h2>\n<h3>Step 1: Profile (deterministic)</h3>\n<p>Run <code>generate_codebook.py</code> on the dataset. It writes <code>codebook.json</code> (machine-\nreadable) and <code>codebook.md</code> (review table), reporting per variable: role\n(id / continuous / categorical / binary / date / text), dtype, missingness,\nunique count, level frequencies or quantile summary, and a <code>needs_dictionary</code> flag.</p>\n<h3>Step 2: Review with the researcher (gate)</h3>\n<p>Present <code>codebook.md</code> and walk the user through it. <strong>Gate:</strong> the user confirms\nthe inferred roles (e.g., an integer-coded scale mis-read as continuous, or an id\ncolumn). Do not proceed to definition work until the user approves the role\nassignments.</p>\n<h3>Step 3: Resolve [NEEDS DICTIONARY] items (gate)</h3>\n<p>For every variable flagged <code>needs_dictionary: true</code>, the level codes are\nuninterpretable without the authoritative source. <strong>Gate:</strong> ask the user to\nsupply the meaning of each code from the real data dictionary (file/sheet/row),\nor to confirm none exists. Fill <code>label</code>, <code>units</code>, and per-level meanings into the\ncodebook <strong>only</strong> from that source — never from inference. If the user cannot\nsupply it, leave the <code>[NEEDS DICTIONARY]</code> marker in place; do not erase it.</p>\n<h3>Step 4: Hand off</h3>\n<p>The completed <code>codebook.json</code> becomes the input dictionary for <code>/define-variables</code>\n(operationalization) and the citation source for dictionary-first QC. <strong>Gate:</strong>\nconfirm with the user that no <code>needs_dictionary</code> flags remain unresolved before\nthe codebook is treated as authoritative for downstream analysis.</p>\n<h2>Scope Limitations</h2>\n<h3>Supported</h3>\n<ul>\n<li>Tabular files: CSV, TSV, Excel, Parquet, Stata (<code>.dta</code>), SAS (<code>.sas7bdat</code>).</li>\n<li>Per-variable profiling, role inference, missingness, level/range summaries.</li>\n</ul>\n<h3>NOT Supported</h3>\n<ul>\n<li>Inventing or guessing the meaning of coded values (that is <code>[NEEDS DICTIONARY]</code>).</li>\n<li>Cleaning or transforming data — use <code>/clean-data</code>.</li>\n<li>De-identification — use <code>/deidentify</code> before sharing.</li>\n<li>Operationalizing exposure/outcome definitions — use <code>/define-variables</code> (this skill feeds it).</li>\n</ul>\n<h2>Cross-Skill Integration</h2>\n<ul>\n<li><strong>/define-variables</strong> consumes <code>codebook.json</code> as its data dictionary input.</li>\n<li><strong>/clean-data</strong> profiles + cleans; this skill produces a durable dictionary artifact instead.</li>\n<li><strong>/deidentify</strong> should run on the raw data before a codebook is shared externally.</li>\n</ul>\n<h2>Output Format</h2>\n<p><code>codebook.json</code> (schema in references) and <code>codebook.md</code> (review table with a\n\"Columns requiring dictionary lookup\" section). Summarize the counts\n(rows, columns, <code>needs_dictionary_count</code>) in chat; do not paste the full JSON.</p>\n<h2>Worked Example</h2>\n<p>Input <code>cohort.csv</code>:</p>\n<pre><code>patient_id,age,sex,fatty_liver_grade,smoking_status,visit_date\n1001,54,1,0,never,2023-01-15\n1002,61,2,2,former,2023-02-03\n</code></pre>\n<p>Run:</p>\n<pre><code>python \"${CLAUDE_SKILL_DIR}/scripts/generate_codebook.py\" cohort.csv --out-dir .\n# -&gt; {\"n_rows\": ..., \"n_columns\": 6, \"needs_dictionary_count\": 2, \"outputs\": [...]}\n</code></pre>\n<p><code>codebook.md</code> (excerpt):</p>\n<pre><code>| Variable            | Role        | Missing % | Unique | Needs dictionary |\n| `patient_id`        | id          | 0.0       | N      |                  |\n| `age`               | continuous  | 0.0       | ...    |                  |\n| `sex`               | binary      | 0.0       | 2      | ⚠️ YES           |\n| `fatty_liver_grade` | categorical | 0.0       | 5      | ⚠️ YES           |\n| `smoking_status`    | categorical | 0.0       | 3      |                  |\n| `visit_date`        | date        | 0.0       | ...    |                  |\n</code></pre>\n<p><code>sex</code> and <code>fatty_liver_grade</code> are flagged because their levels are bare codes\n(<code>1/2</code>, <code>0..4</code>). <code>smoking_status</code> is <strong>not</strong> flagged — its levels are already\nhuman-readable. The reviewer then:</p>\n<ol>\n<li>Opens the project's authoritative data dictionary.</li>\n<li>Fills <code>sex</code>: <code>1 = male, 2 = female</code> and <code>fatty_liver_grade</code>: <code>0 = none … 4 = suspected</code>\ninto the codebook <strong>from that source</strong> (citing file &gt; sheet &gt; row).</li>\n<li>Confirms no <code>[NEEDS DICTIONARY]</code> flags remain, then hands <code>codebook.json</code> to\n<code>/define-variables</code>.</li>\n</ol>\n<p>What the skill must <strong>never</strong> do: write <code>sex: 1 = male</code> because \"that is the\nusual coding.\" If the dictionary is unavailable, the flag stays.</p>\n<h2>Anti-Hallucination</h2>\n<ul>\n<li>Never invent a variable's label, units, or the meaning of any coded level.</li>\n<li>Coded categorical/binary columns with bare codes are flagged <code>[NEEDS DICTIONARY]</code>;\nthe meaning is filled only from the authoritative data dictionary, then cited.</li>\n<li>Role inference is a heuristic — surface it for user confirmation, do not assert it as ground truth.</li>\n<li>The profiler reads values locally; no data is sent to any model or network.</li>\n</ul>\n","files":[{"path":"references/codebook_schema.md","sizeBytes":3435,"isText":true},{"path":"scripts/generate_codebook.py","sizeBytes":10982,"isText":true},{"path":"SKILL.md","sizeBytes":7129,"isText":true},{"path":"skill.yml","sizeBytes":1237,"isText":true},{"path":"tests/test_generate_codebook.sh","sizeBytes":3500,"isText":true}],"reviewScore":null,"reviewSummary":null,"trust":{"provenance":"trusted-source-unreviewed","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow.","bodySource":null},"bodyLocked":false,"purchaseUrl":null,"sourceUrl":null,"report":{"provenance":"trusted-source-unreviewed","screen":{"ran":true,"outcome":"clean","suspicious":0,"notes":0,"hiddenCharacters":false},"virusScan":{"engine":"clamav","status":"clean","scannedAt":"2026-09-14T20:48:41.168236Z","sha256":"71990F5ABEE9D742F5916EF3DE15093533E306E436AAD2366AD8942446AF0D73","sizeBytes":11220},"review":null,"source":{"repositoryUrl":"https://github.com/Aperivue/medsci-skills","path":"skills/generate-codebook","license":"MIT","commit":"5599b724675a1d788e03cd58dabd3db7c68ca86b","subtreeSha":"8DC72C7A8E86CF71E280AFED5143C70AF328B201C30179B981D0D58BAF567B98","lastSyncedAt":"2026-09-27T19:46:33.449845Z"},"reviewedAt":"2026-09-14T20:53:35.165377Z","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow."},"install":[{"target":"skills-cli","command":"npx skills add https://github.com/Aperivue/medsci-skills/tree/main/skills/generate-codebook"},{"target":"claude-code","command":"claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install aperivue-medsci-skills@llmmart"},{"target":"git","command":"git clone https://github.com/Aperivue/medsci-skills.git"}]}