{"slug":"human-eval-handoff-repair","title":"human-eval-handoff-repair","summary":"Use when validating, repairing, or mapping human-evaluation handoff packages, filled annotation CSVs, rebuttal annotation UIs, or reviewer annotation returns across package versions. Applies to checking row alignment, detecting cross-snapshot contamination, converting old labels ","platform":"Claude","tags":[],"authorName":"LLM Mart","authorSlug":"llm-mart","score":0,"source":"github","price":null,"verified":false,"createdAt":"2026-09-30T16:42:51.197302Z","repo":{"url":"https://github.com/yha9806/academic-writing-toolkit","stars":41,"forks":7,"license":"MIT","updatedAt":"2026-09-30T13:19:54Z"},"bodyHtml":"<hr>\n<h2>name: human-eval-handoff-repair\ndescription: Use when validating, repairing, or mapping human-evaluation handoff packages, filled annotation CSVs, rebuttal annotation UIs, or reviewer annotation returns across package versions. Applies to checking row alignment, detecting cross-snapshot contamination, converting old labels to current schemas, generating refill/import CSVs, and deciding whether filled annotations are safe for formal aggregation.\nallowed-tools: Read, Glob, Grep, Bash, Write</h2>\n<h1>Human Eval Handoff Repair</h1>\n<p>Use this skill when the user asks to inspect, repair, migrate, or validate human-evaluation packages or filled annotation CSVs, especially when multiple package versions exist.</p>\n<h2>First Principle</h2>\n<p>Never trust row number, filename, or visible ID alone. Treat a filled CSV as usable only after it is aligned to the target public package by stable task-specific keys and its editable labels pass schema checks.</p>\n<p>Do not fix labels by guessing. If a row cannot be matched safely, leave its annotation fields blank and produce a refill/import template plus an unmatched reference file.</p>\n<h2>Inputs To Locate</h2>\n<ul>\n<li>Target public handoff package folder or ZIP.</li>\n<li>Filled CSVs from annotators.</li>\n<li>Any older package claimed as the source version.</li>\n<li>Current task schemas from the target UI data or target task CSVs.</li>\n<li>If relevant, QC reports from the target package.</li>\n</ul>\n<p>Prefer a user-specified output directory. If none is specified, write generated reports and repaired files under <code>codex_outputs/</code> in the current workspace. Avoid synced personal document folders unless the user explicitly asks for them.</p>\n<h2>Task Types And Stable Keys</h2>\n<p>Use these stable keys before copying any labels:</p>\n<ul>\n<li>Claim Warrant Audit: <code>item_id</code>, <code>tradition</code>, <code>medium</code>, <code>period</code>, <code>layer</code>, <code>dimension_id</code>, <code>claim_text</code>.</li>\n<li>Release Suitability Audit: <code>item_id</code>, <code>source_bucket</code>, <code>pipeline_mode</code>, <code>option_a_source</code>, <code>option_b_source</code>, <code>governed_option</code>, <code>option_a_text</code>, <code>option_b_text</code>.</li>\n<li>Card Faithfulness Spot Audit: <code>audit_field_id</code>, <code>tradition</code>, <code>card_type</code>, <code>dimension_id</code>, <code>field_name</code>, <code>field_value</code>, <code>source_reference_excerpt</code>.</li>\n</ul>\n<p>If the exact stable key does not match, do not formally aggregate the row. Text-only fuzzy candidates may be generated for manual review, but must not be treated as validated mappings.</p>\n<h2>Editable Label Columns</h2>\n<p>Claim Warrant Audit editable columns:</p>\n<ul>\n<li><code>support_source</code></li>\n<li><code>overreach</code></li>\n<li><code>certainty_calibration</code></li>\n<li><code>recommended_action</code></li>\n<li><code>annotator_confidence</code></li>\n<li><code>notes</code></li>\n</ul>\n<p>Release Suitability Audit editable columns:</p>\n<ul>\n<li><code>suitability_decision</code></li>\n<li><code>reason_primary</code></li>\n<li><code>risk_in_worse_option</code></li>\n<li><code>annotator_confidence</code></li>\n<li><code>notes</code></li>\n</ul>\n<p>Card Faithfulness Spot Audit editable columns:</p>\n<ul>\n<li><code>field_status</code></li>\n<li><code>field_provenance</code></li>\n<li><code>notes</code></li>\n</ul>\n<p>Only copy editable label columns. Never copy old evidence text, image paths, metadata fields, or UI-only columns into a new package.</p>\n<h2>Label Schema Checks</h2>\n<p>Claim legal values:</p>\n<ul>\n<li><code>support_source</code>: <code>visual_observation</code>, <code>metadata_context</code>, <code>card_or_expert_context</code>, <code>b0_excerpt</code>, <code>mixed</code>, <code>none</code></li>\n<li><code>overreach</code>: <code>no_overreach</code>, <code>minor_overreach</code>, <code>major_overreach</code>, <code>unsupported</code></li>\n<li><code>certainty_calibration</code>: <code>well_calibrated</code>, <code>slightly_overstated</code>, <code>clearly_overstated</code>, <code>uncertain_or_not_checkable</code></li>\n<li><code>recommended_action</code>: <code>accept</code>, <code>revise</code>, <code>reject</code>, <code>filter_or_abstain</code></li>\n<li><code>annotator_confidence</code>: usually <code>1</code> to <code>5</code></li>\n</ul>\n<p>Release legal values:</p>\n<ul>\n<li><code>suitability_decision</code>: <code>A_better</code>, <code>B_better</code>, <code>tie_both_ok</code>, <code>tie_both_bad</code>, <code>abstain_required</code></li>\n<li><code>reason_primary</code>: target package enum only; validate from the target CSV or UI schema.</li>\n<li><code>risk_in_worse_option</code>: target package enum only; validate from the target CSV or UI schema.</li>\n<li><code>annotator_confidence</code>: usually <code>1</code> to <code>5</code></li>\n</ul>\n<p>Card legal values:</p>\n<ul>\n<li><code>field_status</code>: <code>faithful</code>, <code>over_specific</code>, <code>unsupported</code>, <code>wrong</code>, <code>unclear</code></li>\n<li><code>field_provenance</code>: <code>expert_critique</code>, <code>metadata</code>, <code>visual_observation</code>, <code>accepted_domain_context</code>, <code>none</code>, <code>unclear</code></li>\n</ul>\n<p>Normalize harmless spelling and casing only when the mapping is deterministic. Otherwise flag the value and leave it blank.</p>\n<h2>Required QC Procedure</h2>\n<p>For every filled CSV:</p>\n<ol>\n<li>Count rows against the target task row count.</li>\n<li>Check required label columns are present.</li>\n<li>Check all non-empty labels are legal enum values.</li>\n<li>Compare non-label/source columns with the target blank CSV.</li>\n<li>Compute exact stable-key overlap with the target package.</li>\n<li>Identify whether the filled CSV matches the current package, an older package, or neither.</li>\n<li>Run simple logic checks.</li>\n<li>Produce a QC JSON and Markdown summary.</li>\n</ol>\n<p>Important logic flags:</p>\n<ul>\n<li>Claim: <code>support_source=none</code> with <code>recommended_action=accept</code>.</li>\n<li>Claim: <code>overreach=unsupported</code> with <code>recommended_action=accept</code>.</li>\n<li>Claim: <code>recommended_action=filter_or_abstain</code> with <code>overreach=no_overreach</code>.</li>\n<li>Claim: <code>recommended_action=reject</code> while <code>support_source</code> is not <code>none</code> and <code>overreach=no_overreach</code>.</li>\n<li>Release: <code>abstain_required</code> with a confident preference reason.</li>\n<li>Release: <code>A_better</code> or <code>B_better</code> with notes saying both options are unusable.</li>\n<li>Card: <code>field_status=faithful</code> with <code>field_provenance=none</code>.</li>\n<li>Card: <code>field_status=wrong</code> with strong provenance unless notes justify it.</li>\n</ul>\n<p>Logic flags are review warnings, not automatic invalidation.</p>\n<h2>Decision Categories</h2>\n<ul>\n<li><code>PASS</code>: exact target alignment, legal labels, no serious logic issues.</li>\n<li><code>PASS_WITH_MINOR_QC</code>: exact alignment and legal labels, but has small note or logic warnings requiring spot check.</li>\n<li><code>PARTIAL_PASS</code>: some rows safely map by stable key; use only mapped rows, refill unmatched rows.</li>\n<li><code>REFERENCE_ONLY</code>: old package or fuzzy/text-only mapping; useful for guidance but not formal aggregation.</li>\n<li><code>FAIL_REDO</code>: no safe stable alignment, major contamination, missing required labels, or wrong task schema.</li>\n</ul>\n<h2>Mapping And Repair Rules</h2>\n<p>When migrating old filled data to a current package:</p>\n<ol>\n<li>Start from the current target blank task CSV.</li>\n<li>Build exact stable-key lookup for current rows.</li>\n<li>For each old filled row, copy only editable label columns if exactly one current row matches the stable key.</li>\n<li>Leave unmatched current rows blank.</li>\n<li>Save unmatched old rows separately for manual review.</li>\n<li>Save a lightweight import CSV with only ID/key columns plus editable label columns when the UI import is sensitive to long evidence text.</li>\n<li>Never overwrite current evidence fields with old package evidence.</li>\n</ol>\n<p>If stable overlap is low, explain that the old annotations may reflect a different snapshot and should not be imported as formal evidence.</p>\n<h2>Package-Level Public Handoff Checks</h2>\n<p>Before saying a package is safe for annotators, scan public UI/data/CSV files for forbidden internal or old-version markers.</p>\n<p>Forbidden public markers commonly include:</p>\n<ul>\n<li><code>claim_master</code></li>\n<li><code>release_master</code></li>\n<li><code>Claim Warrant Master QC</code></li>\n<li><code>Release Suitability Master QC</code></li>\n<li><code>claim_warrant_audit_master.csv</code></li>\n<li><code>release_suitability_master.csv</code></li>\n<li><code>card_context_snapshot</code></li>\n<li><code>source_result_item_id</code></li>\n<li><code>metadata_item_id</code></li>\n<li><code>mapping_status</code></li>\n<li><code>snapshot_row_hash</code></li>\n<li>old package names such as <code>mapped_v5_current_20260531</code>, <code>handoff_with_images_20260529</code>, <code>filled3_normalized</code>, <code>fallback_card_faithfulness</code></li>\n</ul>\n<p>Expected public tasks for this rebuttal package pattern:</p>\n<ul>\n<li>Claim Warrant Annotator A</li>\n<li>Claim Warrant Annotator B</li>\n<li>Release Suitability Annotator A</li>\n<li>Release Suitability Annotator B</li>\n<li>Card Faithfulness Spot Audit</li>\n</ul>\n<p>Master/coordinator files must be absent from public UI loading. If retained, place them in <code>coordinator_only/</code> or a separate coordinator-only archive.</p>\n<h2>Evidence Text And Image Checks</h2>\n<p>Check that evidence text is not legacy-short truncated. Watch for old uniform limits:</p>\n<ul>\n<li><code>visual_observations</code> around 900 characters ending in <code>...</code></li>\n<li><code>b0_excerpt</code> around 700 characters ending in <code>...</code></li>\n<li><code>source_reference_excerpt</code> around 900 characters ending in <code>...</code></li>\n<li><code>field_value</code> around 1000 characters ending in <code>...</code></li>\n</ul>\n<p>A few long fields may still end in <code>...</code> only if they hit a new high limit such as 8000 to 12000 characters. English source text is authoritative; translated text is auxiliary reading support unless the package states otherwise.</p>\n<p>Verify images exist and that the UI can load representative rows from all task types.</p>\n<h2>Card Faithfulness Caveat</h2>\n<p>For Card Faithfulness, a field/source mismatch can be the thing being audited, not necessarily a package bug. Example: if <code>field_value</code> discusses one iconographic context but <code>source_reference_excerpt</code> is about a different work, the correct human action may be <code>field_status=wrong</code> or <code>unsupported</code> and <code>field_provenance=none</code>.</p>\n<p>Do not confuse this with Claim/Release cross-snapshot contamination. Claim/Release rows evaluate model outputs for a specific image; if claim/release text comes from another snapshot, the metric is invalid.</p>\n<h2>Outputs To Produce</h2>\n<p>For validation or repair tasks, return:</p>\n<ul>\n<li>Clean or mapped CSVs.</li>\n<li>Lightweight UI import CSVs when useful.</li>\n<li>Unmatched rows reference CSV.</li>\n<li>Text-only candidate mapping CSV, clearly marked manual review only.</li>\n<li>QC JSON and QC Markdown.</li>\n<li>ZIP path and SHA256 if packaging outputs.</li>\n<li>Short final verdict stating exactly what can be formally used and what must be refilled.</li>\n</ul>\n<h2>User-Facing Summary Pattern</h2>\n<p>Keep the final answer direct:</p>\n<ul>\n<li>Current verdict: usable, partial, or redo.</li>\n<li>Exact counts, e.g. <code>164/200 safe mapped, 36 need refill</code>.</li>\n<li>Paths to outputs.</li>\n<li>Any row IDs needing spot check.</li>\n<li>Clear warning if a file is reference only, not formal human evidence.</li>\n</ul>\n","files":[{"path":"SKILL.md","sizeBytes":9441,"isText":true}],"reviewScore":null,"reviewSummary":null,"trust":{"provenance":"trusted-source-unreviewed","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow.","bodySource":null},"bodyLocked":false,"purchaseUrl":null,"sourceUrl":null,"report":{"provenance":"trusted-source-unreviewed","screen":{"ran":true,"outcome":"clean","suspicious":0,"notes":0,"hiddenCharacters":false},"virusScan":{"engine":"clamav","status":"clean","scannedAt":"2026-09-30T16:43:58.375689Z","sha256":"56B6BEF3D5516A26AE1A9B3A067539D43C3C93D5A37264A42893FCCD7D81F269","sizeBytes":3829},"review":null,"source":{"repositoryUrl":"https://github.com/yha9806/academic-writing-toolkit","path":"archive/skills/human-eval-handoff-repair","license":"MIT","commit":"184e48294d7ff735186c1b41f22313c3bcabc6bb","subtreeSha":"46DE888131EC2618F12637AEEDFAA273733A42F3F9772471C58BF6224911AD81","lastSyncedAt":"2026-09-30T16:42:49.032948Z"},"reviewedAt":"2026-09-30T16:46:17.651943Z","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow."},"install":[{"target":"skills-cli","command":"npx skills add https://github.com/yha9806/academic-writing-toolkit/tree/main/archive/skills/human-eval-handoff-repair"},{"target":"claude-code","command":"claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install yha9806-academic-writing-toolkit@llmmart"},{"target":"git","command":"git clone https://github.com/yha9806/academic-writing-toolkit.git"}]}