{"slug":"paper-autoraters","title":"paper-autoraters","summary":"Run the four paper-quality autoraters from PaperOrchestra (arXiv:2604.05018, App. F.3) — Citation F1 (P0/P1 partition + Precision/Recall/F1), Literature Review Quality (6-axis 0-100 with anti-inflation rules), SxS Overall Paper Quality (side-by-side), and SxS Literature Review Qu","platform":"Claude","tags":[],"authorName":"LLM Mart","authorSlug":"llm-mart","score":0,"source":"github","price":null,"verified":false,"createdAt":"2026-09-24T15:42:14.35474Z","repo":{"url":"https://github.com/Ar9av/PaperOrchestra","stars":663,"forks":92,"license":null,"updatedAt":"2026-09-21T17:10:41Z"},"bodyHtml":"<hr>\n<h2>name: paper-autoraters\ndescription: Run the four paper-quality autoraters from PaperOrchestra (arXiv:2604.05018, App. F.3) — Citation F1 (P0/P1 partition + Precision/Recall/F1), Literature Review Quality (6-axis 0-100 with anti-inflation rules), SxS Overall Paper Quality (side-by-side), and SxS Literature Review Quality (side-by-side). TRIGGER when the user asks to \"score this paper draft\", \"evaluate against the benchmark\", \"compare two papers\", or \"run the autoraters\".</h2>\n<h1>Paper Autoraters (App. F.3)</h1>\n<p>Faithful implementation of the four LLM-as-judge autoraters used in\nPaperOrchestra (Song et al., 2026, arXiv:2604.05018, §5 and App. F.3).</p>\n<p>These are the metrics the paper uses to demonstrate that PaperOrchestra\nbeats single-agent and AI-Scientist-v2 baselines. Use them to:</p>\n<ol>\n<li>Score a generated paper against a ground-truth paper.</li>\n<li>Compare two paper-writing pipelines side-by-side.</li>\n<li>Validate your own host-agent execution of the paper-orchestra pipeline.</li>\n</ol>\n<h2>The four autoraters</h2>\n<table>\n<thead>\n<tr>\n<th>Autorater</th>\n<th>What it does</th>\n<th>Inputs</th>\n<th>Output</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Citation F1 — P0/P1 partition</strong></td>\n<td>Partitions reference list into P0 (must-cite) and P1 (good-to-cite) given the paper text</td>\n<td>one paper text + its references list</td>\n<td>JSON <code>{ref_num: \"P0\"\\|\"P1\"}</code></td>\n</tr>\n<tr>\n<td><strong>Literature Review Quality</strong></td>\n<td>6-axis 0-100 score for Intro+Related Work, with anti-inflation hard caps</td>\n<td>one paper PDF/text + reference avg citation count</td>\n<td>JSON with <code>axis_scores</code>, <code>penalties</code>, <code>summary</code>, <code>overall_score</code></td>\n</tr>\n<tr>\n<td><strong>SxS Overall Paper Quality</strong></td>\n<td>Holistic side-by-side preference judgment</td>\n<td>two papers (PDF or text)</td>\n<td>JSON with <code>winner</code> ∈ {paper_1, paper_2, tie}</td>\n</tr>\n<tr>\n<td><strong>SxS Literature Review Quality</strong></td>\n<td>Side-by-side preference, Intro+Related Work only</td>\n<td>two papers</td>\n<td>JSON with <code>winner</code> ∈ {paper_1, paper_2, tie}</td>\n</tr>\n</tbody>\n</table>\n<p>The paper uses Gemini-3.1-Pro and GPT-5 as judges, set to temperature 0.0\n(Gemini) or default 1.0 (GPT-5, which doesn't allow temperature\nadjustment). Use whatever your host LLM is.</p>\n<h2>Workflow</h2>\n<h3>Citation F1 (compute Precision / Recall / F1 vs ground truth)</h3>\n<p>This is a two-step procedure:</p>\n<h4>Step 1: Partition the reference lists into P0 / P1</h4>\n<p>For both the ground-truth paper AND the generated paper, run the LLM with\n<code>references/citation-f1-prompt.md</code>:</p>\n<pre><code>inputs:\n  paper_text:    full paper LaTeX or markdown\n  references_str: numbered reference list (e.g., \"1. Vaswani et al. (2017)\n                  Attention Is All You Need. NeurIPS. 2. He et al. (2016)\n                  Deep Residual Learning for Image Recognition. CVPR. ...\")\n\noutput: JSON {\"1\": \"P0\", \"2\": \"P1\", \"3\": \"P0\", ...}\n</code></pre>\n<p>Save both partitions:</p>\n<ul>\n<li><code>bench/&lt;paper_id&gt;/gt_partition.json</code></li>\n<li><code>bench/&lt;paper_id&gt;/gen_partition.json</code></li>\n</ul>\n<h4>Step 2: Resolve references to entity IDs and compute F1</h4>\n<p>The paper uses Semantic Scholar paper IDs to match references between the\ntwo lists. The <code>compute_f1.py</code> script does this deterministically given\ntwo input lists:</p>\n<pre><code>python skills/paper-autoraters/scripts/compute_f1.py \\\n    --gt-partition gt_partition.json \\\n    --gt-refs gt_refs.json \\\n    --gen-partition gen_partition.json \\\n    --gen-refs gen_refs.json \\\n    --out f1_report.json\n</code></pre>\n<p>Where <code>gt_refs.json</code> and <code>gen_refs.json</code> are lists of <code>{ref_num, paper_id, title}</code> produced by your host's S2-resolution pass (the same\nfuzzy match + S2 verification used by <code>literature-review-agent/scripts/</code>).</p>\n<p>Output JSON contains P0 / P1 / overall Precision, Recall, F1.</p>\n<h3>Literature Review Quality (single paper, 6 axes)</h3>\n<p>Load <code>references/litreview-quality-prompt.md</code>. Inputs:</p>\n<ul>\n<li>The full paper PDF (or LaTeX/markdown if your host lacks PDF input)</li>\n<li><code>avg_citation_count</code> for the venue/field (used as the baseline for\ncitation count anchoring, e.g., 58.52 for CVPR 2025, 59.18 for ICLR 2025\nper the paper)</li>\n</ul>\n<p>The prompt instructs the model to evaluate ONLY the literature-review\nfunction of the paper (Introduction + Related Work / Background sections).\nIt produces a strict JSON output with per-axis scores and justifications.</p>\n<p>Critical anti-inflation rules baked into the prompt:</p>\n<table>\n<thead>\n<tr>\n<th>Rule</th>\n<th>Cap</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Default expectation</td>\n<td>overall 45-70</td>\n</tr>\n<tr>\n<td>&gt; 85 requires strong evidence on ALL axes</td>\n<td>—</td>\n</tr>\n<tr>\n<td>&gt; 90 extremely rare (near-survey-level mastery)</td>\n<td>—</td>\n</tr>\n<tr>\n<td>Any axis &lt; 50 → overall rarely &gt; 75</td>\n<td>—</td>\n</tr>\n<tr>\n<td>Mostly descriptive review</td>\n<td>Critical Analysis ≤ 60</td>\n</tr>\n<tr>\n<td>Novelty asserted without comparison</td>\n<td>Positioning ≤ 60</td>\n</tr>\n<tr>\n<td>Sparse/inconsistent citations</td>\n<td>Citation Rigor ≤ 60</td>\n</tr>\n<tr>\n<td>Citation count &lt; 50% of avg</td>\n<td>Coverage ≤ 55</td>\n</tr>\n<tr>\n<td>Citation count &gt; 120% of avg</td>\n<td>Coverage = \"strong\"</td>\n</tr>\n</tbody>\n</table>\n<p>Plus penalty table:</p>\n<table>\n<thead>\n<tr>\n<th>Penalty</th>\n<th>Range</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Overclaiming novelty</td>\n<td>-5 to -15</td>\n</tr>\n<tr>\n<td>Missing key recent work</td>\n<td>-5 to -15</td>\n</tr>\n<tr>\n<td>Mostly descriptive review</td>\n<td>-5 to -10</td>\n</tr>\n<tr>\n<td>Weak gap statements</td>\n<td>-5 to -10</td>\n</tr>\n<tr>\n<td>Citation dumping</td>\n<td>-5 to -10</td>\n</tr>\n</tbody>\n</table>\n<p>Save the output to <code>litreview_quality_score.json</code>. The score JSON is the\nsame shape used by <code>content-refinement-agent/scripts/score_delta.py</code>, so\nyou can re-use the halt-rule logic to compare iterations.</p>\n<h3>SxS Overall Paper Quality (side-by-side, full paper)</h3>\n<p>Load <code>references/sxs-paper-quality-prompt.md</code>. Inputs:</p>\n<ul>\n<li>Two paper PDFs or LaTeX files (call them <code>paper_1</code> and <code>paper_2</code>)</li>\n</ul>\n<p>The prompt produces a JSON with <code>paper_1_holistic_analysis</code>,\n<code>paper_2_holistic_analysis</code>, <code>comparison_justification</code>, and\n<code>winner ∈ {paper_1, paper_2, tie}</code>.</p>\n<p>To mitigate LLM positional bias (the paper notes this in §5.4), run the\ncomparison <strong>twice</strong> with the order swapped:</p>\n<pre><code>call_1: paper_A → paper_1, paper_B → paper_2  → winner1\ncall_2: paper_B → paper_1, paper_A → paper_2  → winner2\n</code></pre>\n<p>Final outcome: a <code>win</code> (both calls agree on paper A), <code>tie</code> (one win + one\ntie, or two ties), or <code>loss</code> (both agree on paper B). The paper uses this\nexact ordering protocol.</p>\n<h3>SxS Literature Review Quality (side-by-side, Intro+RW only)</h3>\n<p>Load <code>references/sxs-litreview-prompt.md</code>. Same input/output shape as the\nSxS paper quality autorater, but the model is instructed to evaluate\n<strong>only</strong> the Introduction and Related Work / Background sections of each\npaper. Same positional-bias mitigation: run twice, swap order.</p>\n<h2>Resources</h2>\n<ul>\n<li><code>references/citation-f1-prompt.md</code>        — verbatim P0/P1 partition prompt from App. F.3</li>\n<li><code>references/litreview-quality-prompt.md</code>  — verbatim 6-axis litreview rubric from App. F.3</li>\n<li><code>references/sxs-paper-quality-prompt.md</code>  — verbatim SxS paper-quality prompt from App. F.3</li>\n<li><code>references/sxs-litreview-prompt.md</code>      — verbatim SxS litreview prompt from App. F.3</li>\n<li><code>scripts/compute_f1.py</code> — Precision / Recall / F1 from two partition JSONs</li>\n</ul>\n","files":[{"path":"references/citation-f1-prompt.md","sizeBytes":2563,"isText":true},{"path":"references/litreview-quality-prompt.md","sizeBytes":8060,"isText":true},{"path":"references/sxs-litreview-prompt.md","sizeBytes":2406,"isText":true},{"path":"references/sxs-paper-quality-prompt.md","sizeBytes":3433,"isText":true},{"path":"scripts/compute_f1.py","sizeBytes":3690,"isText":true},{"path":"SKILL.md","sizeBytes":6595,"isText":true}],"reviewScore":null,"reviewSummary":null,"trust":{"provenance":"trusted-source-unreviewed","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow.","bodySource":null},"bodyLocked":false,"purchaseUrl":null,"sourceUrl":null,"report":{"provenance":"trusted-source-unreviewed","screen":{"ran":true,"outcome":"clean","suspicious":0,"notes":0,"hiddenCharacters":false},"virusScan":{"engine":"clamav","status":"clean","scannedAt":"2026-09-24T15:42:26.828063Z","sha256":"770890514060284B92F9BDA5A1A62B378D202EF3FA6C4A805F63650A51851F82","sizeBytes":12450},"review":null,"source":{"repositoryUrl":"https://github.com/Ar9av/PaperOrchestra","path":"skills/paper-autoraters","license":null,"commit":"36c3cc4b10370b1f905adcd4e5601e8b624c2dc3","subtreeSha":"69923D46D3AD727890567A4AAC16AB0459F33491D99E7F57E4949149260291C1","lastSyncedAt":"2026-09-24T15:42:09.852277Z"},"reviewedAt":"2026-09-24T15:42:38.700617Z","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow."},"install":[{"target":"skills-cli","command":"npx skills add https://github.com/Ar9av/PaperOrchestra/tree/main/skills/paper-autoraters"},{"target":"claude-code","command":"claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install ar9av-paperorchestra@llmmart"},{"target":"git","command":"git clone https://github.com/Ar9av/PaperOrchestra.git"}]}