{"slug":"obsidian-paper-vault","title":"obsidian-paper-vault","summary":"Turn a folder of research PDFs into an Obsidian knowledge vault — consistently formatted literature notes with frontmatter, PDF embed links, and cross-referenced atomic concept notes. Use whenever the user wants PDFs converted to Obsidian notes, a batch of papers summarized into ","platform":"Claude","tags":[],"authorName":"LLM Mart","authorSlug":"llm-mart","score":0,"source":"github","price":null,"verified":false,"createdAt":"2026-09-14T20:48:09.269323Z","repo":{"url":"https://github.com/Aperivue/medsci-skills","stars":318,"forks":75,"license":"MIT","updatedAt":"2026-09-27T05:05:17Z"},"bodyHtml":"<hr>\n<h2>name: obsidian-paper-vault\ndescription: Turn a folder of research PDFs into an Obsidian knowledge vault — consistently formatted literature notes with frontmatter, PDF embed links, and cross-referenced atomic concept notes. Use whenever the user wants PDFs converted to Obsidian notes, a batch of papers summarized into a common template, a research \"second brain\" built or extended, or concepts extracted across accumulated notes — even if they never say \"Obsidian\". Pairs with /lit-sync, which owns the same vault folders from the Zotero/BibTeX side.\ntriggers: obsidian-paper-vault, paper vault, second brain, PDF를 Obsidian 노트로, 논문 요약 노트, 논문 노트 만들어줘, 이 폴더의 PDF 정리해줘, batch process papers, add papers to vault, extract concepts from papers, literature vault\ntools: Read, Write, Edit, Bash, Grep, Glob\nmodel: inherit</h2>\n<h1>Obsidian Paper Vault</h1>\n<p>Converts a folder of research PDFs into a two-layer Obsidian vault: <strong>literature notes</strong>\n(one per paper, templated) and <strong>atomic concept notes</strong> (synthesized across papers).</p>\n<p>The rules below are not style preferences. Each one is here because its absence produced a\nspecific, silent failure — a fabricated patient count, a broken PDF link, an empty Dataview\ntable — in a vault of 100+ papers.</p>\n<h2>Relationship to /lit-sync</h2>\n<p>Both skills write literature and concept notes into the same vault folders. They enter from\nopposite ends and must not overwrite each other.</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th><code>/lit-sync</code></th>\n<th><code>obsidian-paper-vault</code></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Input</td>\n<td>Zotero collection / <code>refs.bib</code></td>\n<td>a folder of PDFs</td>\n</tr>\n<tr>\n<td>Note key</td>\n<td>citekey</td>\n<td>short descriptive title</td>\n</tr>\n<tr>\n<td>Owns</td>\n<td><code>manuscript/_src/refs.bib</code></td>\n<td>extracted-text cache</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Never overwrite an existing note.</strong> If a note for the paper already exists (by title, DOI,\nor citekey), report it and skip. When both skills are in play, <code>/lit-sync</code> notes are the\nbibliographic spine; this skill's notes are the read-through summaries.</p>\n<h2>Step 0: Resolve the vault layout — ask, do not assume</h2>\n<p>Establish three paths before writing anything:</p>\n<ol>\n<li><strong>Vault root</strong> — from the user, or <code>$OBSIDIAN_VAULT</code>. Never guess a home-directory path.</li>\n<li><strong>Literature notes folder</strong> — default <code>Literature/</code>. If the vault already has a folder\nserving this role (<code>02_research/논문/</code>, <code>Papers/</code>, <code>文献/</code>), <strong>honor the existing layout</strong>\nrather than imposing the default.</li>\n<li><strong>Concept notes folder</strong> — default <code>Concepts/</code>, same honor-what-exists rule.</li>\n</ol>\n<p>For a vault whose structure is Korean, see <code>references/locale/ko/note_templates.md</code> — folder\nnames and note headings in Korean, opt-in.</p>\n<p>Also confirm PyMuPDF is available: <code>python3 -c \"import fitz; print(fitz.__version__)\"</code>\n(install with <code>pip install PyMuPDF</code>). Text extraction caches to\n<code>~/.local/cache/paper-vault-texts/</code> unless the user names another location.</p>\n<h2>Step 1: Pre-extract PDF text — always, for any batch</h2>\n<pre><code>python3 scripts/extract_pdfs.py &lt;pdf_folder&gt; &lt;text_cache_folder&gt; [max_pages]\n</code></pre>\n<p>Defaults to 12 pages, which covers abstract through discussion for most papers.</p>\n<p><strong>Never hand a PDF path to a subagent.</strong> A subagent that cannot open a file does not report\nfailure — it writes the note from training data, and the result is a plausible note with\ninvented numbers. Pass the <code>.txt</code> paths instead. Single-paper interactive work may read the\nPDF directly (the Read tool handles PDFs); batches may not.</p>\n<h2>Step 2: Launch subagents in parallel</h2>\n<p>Five subagents × 5–6 papers is the working batch size: enough parallelism to clear 25 papers\nin one pass, small enough that per-agent quality holds. Group papers thematically per agent\nso each one can spot recurring concepts.</p>\n<p>Give each subagent: its assigned text-file paths with destination filenames, the template\nfrom <code>references/templates.md</code> verbatim, the list of concept notes that already exist, and\nthe prohibition on inventing anything. <code>references/subagent-prompt.md</code> holds the full prompt.</p>\n<h2>Step 3: Track progress in a queue file</h2>\n<p>Keep <code>PAPER_QUEUE.md</code> in the vault with per-paper status (done / pending / skipped / in\nprogress) so a 200-paper vault survives across sessions. Update it after each batch.</p>\n<h2>Step 4: Extract concept notes once 10+ literature notes exist</h2>\n<p>A phrase earns a concept note when it appears in 3+ notes, carries pedagogical value, and is\ntreated differently by different papers. Model names, datasets, and journals are entities,\nnot concepts. See <code>references/concept-extraction.md</code> for the full criteria, the frequency\nscan, and the seedling/growing/mature lifecycle.</p>\n<p>Roughly one new concept note per 5–7 literature notes is healthy. Faster than that is concept\ninflation, and it shows up as dozens of stub notes the user never edits.</p>\n<h2>Anti-Hallucination rules (non-negotiable)</h2>\n<p>A note that is fluent, correctly formatted, and wrong in its numbers is worse than no note:\nthe user cites it. These three rules exist to make that failure impossible rather than\nunlikely.</p>\n<ol>\n<li><strong>Numbers, authors, and dates come from the extracted text only</strong> — never from model\nknowledge, however familiar the paper. Well-known papers drift between versions, and that\nis exactly where invented values look most plausible.</li>\n<li><strong>Subagents receive <code>.txt</code> paths, never PDF paths.</strong> A subagent that cannot open a file\ndoes not report the failure; it writes from training data. The text indirection is the\nonly reliable guard.</li>\n<li><strong>What the text does not state, the note does not claim.</strong> Write \"not stated in the\nextracted text\" instead of filling the gap.</li>\n</ol>\n<p><strong>Gate — before a batch is accepted</strong>: spot-check two notes against their text files (one\nsample size, one effect estimate). If either value is absent from the text, stop the batch\nand report it rather than continuing. This gate is the user's call to waive, not the skill's.</p>\n<h2>Structural rules</h2>\n<ol start=\"4\">\n<li><strong>Preserve the PDF filename exactly</strong> in <code>![[filename.pdf]]</code>. Obsidian embeds are\nsensitive to case, spaces, and punctuation — take the text filename and swap <code>.txt</code> for\n<code>.pdf</code>, character for character.</li>\n<li><strong>Match the frontmatter field names</strong> in <code>references/templates.md</code>. Dataview queries break\non a renamed field, and they break by returning an empty table, not an error.</li>\n<li><strong>Use the existing tag vocabulary</strong> (<code>references/tag-vocabulary.md</code>) rather than inventing\ntop-level tags.</li>\n<li><strong>Name notes with 3–5 keyword concepts</strong>, not the PDF's full title and not <code>paper_001</code>.</li>\n<li><strong>Never overwrite an existing note</strong> — see the <code>/lit-sync</code> boundary above.</li>\n</ol>\n<p><strong>Gate — before concept notes are presented as done</strong>: concept notes ship as \uD83C\uDF31Seedling with\nthe definition marked as a placeholder, and require user review before they count as the\nreader's own. Say so explicitly when handing them over.</p>\n<h2>Reference files</h2>\n<table>\n<thead>\n<tr>\n<th>File</th>\n<th>Read it when</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>references/templates.md</code></td>\n<td>writing any literature or concept note</td>\n</tr>\n<tr>\n<td><code>references/subagent-prompt.md</code></td>\n<td>launching a batch</td>\n</tr>\n<tr>\n<td><code>references/concept-extraction.md</code></td>\n<td>extracting concepts across notes</td>\n</tr>\n<tr>\n<td><code>references/tag-vocabulary.md</code></td>\n<td>choosing tags</td>\n</tr>\n<tr>\n<td><code>references/workflow.md</code></td>\n<td>the user asks how the layers fit together</td>\n</tr>\n<tr>\n<td><code>references/locale/ko/note_templates.md</code></td>\n<td>the vault is Korean-structured</td>\n</tr>\n<tr>\n<td><code>assets/example_paper_note.md</code></td>\n<td>unsure about literature-note formatting</td>\n</tr>\n<tr>\n<td><code>assets/example_concept_note.md</code></td>\n<td>unsure about concept-note formatting</td>\n</tr>\n</tbody>\n</table>\n<h2>Troubleshooting</h2>\n<table>\n<thead>\n<tr>\n<th>Symptom</th>\n<th>Cause</th>\n<th>Fix</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Subagent says it could not read the PDF</td>\n<td>it was given a PDF path</td>\n<td>run <code>extract_pdfs.py</code>, pass <code>.txt</code> paths</td>\n</tr>\n<tr>\n<td>Note reads plausibly but numbers are wrong</td>\n<td>subagent wrote from training data</td>\n<td>re-run against the text file; verify n, CI, p-values</td>\n</tr>\n<tr>\n<td>Dataview table is empty</td>\n<td>frontmatter field renamed</td>\n<td>match <code>references/templates.md</code> exactly</td>\n</tr>\n<tr>\n<td>PDF embed shows a broken tile</td>\n<td>filename mismatch</td>\n<td>compare character by character, including case</td>\n</tr>\n<tr>\n<td>Note content is generic</td>\n<td>text file is abstract-only or OCR is poor</td>\n<td>re-extract with more pages, or check the source PDF</td>\n</tr>\n</tbody>\n</table>\n","files":[{"path":"assets/example_concept_note.md","sizeBytes":2202,"isText":true},{"path":"assets/example_paper_note.md","sizeBytes":2059,"isText":true},{"path":"assets/example_queue.md","sizeBytes":1755,"isText":true},{"path":"references/concept-extraction.md","sizeBytes":3656,"isText":true},{"path":"references/locale/ko/note_templates.md","sizeBytes":3979,"isText":true},{"path":"references/subagent-prompt.md","sizeBytes":3337,"isText":true},{"path":"references/tag-vocabulary.md","sizeBytes":2300,"isText":true},{"path":"references/templates.md","sizeBytes":4866,"isText":true},{"path":"references/workflow.md","sizeBytes":4373,"isText":true},{"path":"scripts/extract_pdfs.py","sizeBytes":2555,"isText":true},{"path":"SKILL.md","sizeBytes":8013,"isText":true},{"path":"skill.yml","sizeBytes":2694,"isText":true}],"reviewScore":null,"reviewSummary":null,"trust":{"provenance":"trusted-source-unreviewed","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow.","bodySource":null},"bodyLocked":false,"purchaseUrl":null,"sourceUrl":null,"report":{"provenance":"trusted-source-unreviewed","screen":{"ran":true,"outcome":"clean","suspicious":0,"notes":0,"hiddenCharacters":false},"virusScan":{"engine":"clamav","status":"clean","scannedAt":"2026-09-14T20:48:27.626722Z","sha256":"54F7BBD0689FBEDCA7A91CF09BFC216745733B32A8F6702CA8517FC08479F6D0","sizeBytes":21733},"review":null,"source":{"repositoryUrl":"https://github.com/Aperivue/medsci-skills","path":"skills/obsidian-paper-vault","license":"MIT","commit":"5599b724675a1d788e03cd58dabd3db7c68ca86b","subtreeSha":"A634C386AE527DD350C9C3C1D2F8B747B2248F861D65C8A44CC64E768B4850BC","lastSyncedAt":"2026-09-27T19:46:33.449845Z"},"reviewedAt":"2026-09-14T20:53:14.634752Z","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow."},"install":[{"target":"skills-cli","command":"npx skills add https://github.com/Aperivue/medsci-skills/tree/main/skills/obsidian-paper-vault"},{"target":"claude-code","command":"claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install aperivue-medsci-skills@llmmart"},{"target":"git","command":"git clone https://github.com/Aperivue/medsci-skills.git"}]}