{"slug":"agami-reconcile","title":"agami-reconcile","summary":"Reconciles known (label, expected_value) numbers from an existing dashboard against agami's answers. Input can be a SCREENSHOT of a Metabase / Power BI / Tableau / Looker dashboard (Claude's vision extracts the pairs), a CSV, or numbers pasted inline — the user doesn't need to kn","platform":"Claude","tags":[],"authorName":"LLM Mart","authorSlug":"llm-mart","score":0,"source":"github","price":null,"verified":false,"createdAt":"2026-08-27T17:01:14.660994Z","repo":{"url":"https://github.com/AgamiAI/agami-core","stars":28,"forks":1,"license":null,"updatedAt":"2026-10-03T05:40:28Z"},"bodyHtml":"<hr>\n<h2>name: agami-reconcile\ndescription: \"Reconciles known (label, expected_value) numbers from an existing dashboard against agami's answers. Input can be a SCREENSHOT of a Metabase / Power BI / Tableau / Looker dashboard (Claude's vision extracts the pairs), a CSV, or numbers pasted inline — the user doesn't need to know which; they can just ask. For each pair, the skill generates a matching NL question, runs it through the active profile's semantic model, diffs actual vs expected, and surfaces matches in green and mismatches in red with drill-down receipts. The strongest onboarding demo for a skeptical data engineer — either we agree with their numbers (trust earned via evidence) or we surface a real definitional disagreement (trust earned via transparency).\"\nwhen_to_use: \"Use when the user says 'reconcile against this dashboard', 'do these numbers match?', 'validate against my Tableau export', '/agami-reconcile </h2>\n<h1>agami reconcile</h1>\n<p>You are running the reconciliation harness. Goal: take labeled numbers from an existing dashboard (Tableau / Looker / Mode / Metabase / Power BI / spreadsheet) — most often a <strong>screenshot</strong>, sometimes a CSV or pasted list — and prove agami can reproduce each number. When numbers match, that's evidence the semantic model is right. When they don't, the receipt drill-down explains why — typically a definitional disagreement (gross vs net, refunds in vs out, FX rate at booking vs reporting date) — which is exactly the trust signal that makes a DE relax.</p>\n<p>This skill orchestrates:</p>\n<ol>\n<li><strong>Extract</strong> the <code>(label, expected_value)</code> pairs from the input — a dashboard screenshot (via vision, confirmed with the user), a CSV, or a pasted list. Number parsing is always deterministic (<code>reconcile.py</code>).</li>\n<li><strong>Generate a matching NL question</strong> for each label.</li>\n<li><strong>Run</strong> each question through the same NL→SQL→execute pipeline as agami-query.</li>\n<li><strong>Diff</strong> actual vs expected with a tolerance.</li>\n<li><strong>Present</strong> a markdown table with per-row status; for mismatches, render the full receipt as a drill-down so the user can find the definitional disagreement.</li>\n</ol>\n<p>Spec for the deterministic helpers: <a href=\"../../scripts/reconcile.py\"><code>scripts/reconcile.py</code></a> (CSV parser + number normalization + diff with tolerance).</p>\n<h2>Conversation style</h2>\n<ul>\n<li><strong>Tight loops.</strong> This skill is a tool, not a tutorial. One question per turn, max two sentences of prose between phases.</li>\n<li><strong>Surface mismatches loud.</strong> A reconcile run with 9/12 matches and 3 mismatches is a SUCCESSFUL run — the mismatches are the value. Lead with what didn't match.</li>\n<li><strong>Don't paste raw SQL in chat.</strong> The receipt has it. Same hard rule as agami-query.</li>\n</ul>\n<hr>\n<h2>Phase 0: Preflight</h2>\n<p>Same checks as agami-query / agami-connect:</p>\n<ol>\n<li><strong>Plan-mode check</strong> per <a href=\"../../shared/plan-mode-check.md\"><code>shared/plan-mode-check.md</code></a>. This skill needs Bash + Read + Write — refuse if locked in plan mode. <strong>DO NOT write a plan file. DO NOT call <code>ExitPlanMode</code>.</strong> Refusal text: <em>\"I can't reconcile in plan mode — each row runs a live query and writes a receipt. Switch to <strong>Auto</strong> or <strong>Edit Automatically</strong> mode (Shift+Tab to cycle) and re-invoke me with the CSV path.\"</em></li>\n<li><strong>Credentials present</strong> — read <code>&lt;artifacts_dir&gt;/local/credentials</code> for the active profile. If missing, invoke <code>/agami-connect</code> to set up first; this skill needs a working DB connection.</li>\n<li><strong>Model present</strong> — <code>&lt;artifacts_dir&gt;/&lt;profile&gt;/datasource.yaml</code> must exist. If not, invoke <code>/agami-connect</code>. This skill needs an introspected model to generate questions against.</li>\n<li><strong>Input — accept any of three shapes; the user needn't know which.</strong> Detect what they gave:\n<ul>\n<li><strong>A screenshot / image</strong> of a dashboard (Metabase, Power BI, Tableau, Looker, a spreadsheet) — the common case. Go to Phase 1's <strong>vision branch</strong>.</li>\n<li><strong>A CSV</strong> — a path in <code>$ARGUMENTS</code>, or pasted inline (write inline CSV to <code>/tmp/agami-reconcile-&lt;ts&gt;.csv</code>). Go to Phase 1's <strong>CSV branch</strong>.</li>\n<li><strong>Numbers pasted inline</strong> as a list/table — treat as inline CSV.\nIf they gave nothing (or just asked \"can you check my dashboard?\"), ask once, welcoming all three: <em>\"Show me the numbers you want to check against — easiest is a <strong>screenshot of your dashboard</strong> (Metabase, Power BI, Tableau, a spreadsheet — whatever you have), but a CSV or a pasted list of <code>label: value</code> works too.\"</em> Don't make them figure out an export format.</li>\n</ul>\n</li>\n<li><strong>If they gave a file path, validate it exists.</strong> If not, surface the error and stop.</li>\n</ol>\n<hr>\n<h2>Phase 1: Extract the (label, value) pairs</h2>\n<p>Whatever the input shape, the goal is the same normalized rows JSON. <strong>Number parsing is always deterministic — it goes through <code>reconcile.py</code>, never the LLM eyeballing a value</strong> (a misread expected number would manufacture a false mismatch on a verification surface).</p>\n<h3>Vision branch — a dashboard screenshot</h3>\n<ol>\n<li><strong>Read the image</strong> and extract every labeled number you can see — KPI tiles, table cells, chart value labels — as <code>(label, raw_value)</code> pairs. Keep the label the user would recognize (\"Total Revenue\", \"Active Users — Apr\"), and the value <strong>exactly as shown, verbatim</strong> (<code>$4.2M</code>, <code>₹2.16Cr</code>, <code>42%</code>, <code>1,234</code>) — don't convert it; the normalizer does that.</li>\n<li><strong>Write the pairs as a 2-column CSV</strong> (Write tool — never a heredoc/<code>python3 -c</code>) to <code>/tmp/agami-reconcile-&lt;ts&gt;.csv</code>, then run the SAME normalizer as the CSV branch (below) so value parsing stays deterministic.</li>\n<li><strong>Confirm before reconciling — vision can misread.</strong> Show the extracted pairs as a small table and ask the user to fix any misread label/number: <em>\"I read these N numbers off your screenshot — correct anything I got wrong, then I'll reconcile.\"</em> This confirm step is <strong>mandatory</strong>: a wrong expected-value isn't a model bug but it reads like one. If a tile is ambiguous or partly cut off, say so and skip it rather than guess.</li>\n</ol>\n<h3>CSV branch — a CSV path or inline-pasted CSV</h3>\n<pre><code>python3 \"$AGAMI_PLUGIN_ROOT/scripts/reconcile.py\" parse --csv \"&lt;csv_path&gt;\" \\\n  &gt; /tmp/agami-reconcile-rows-&lt;ts&gt;.json\n</code></pre>\n<p>The helper (used by <strong>both</strong> branches) handles:</p>\n<ul>\n<li>Header detection (with-or-without first-row column names)</li>\n<li>2-column or 3+ column inputs (3rd onward are appended to the label as context)</li>\n<li>Currency symbols / magnitude suffixes / accounting parens / percent (<code>$4.2M</code>, <code>₹2.16Cr</code>, <code>(123.45)</code>, <code>42%</code>)</li>\n<li>Null sentinels (<code>n/a</code>, <code>—</code>, blank)</li>\n</ul>\n<p>Read the JSON. Each row is <code>{label, expected_value, raw_value}</code>. Discard rows where <code>expected_value</code> is null (unparseable) — surface a one-liner: <em>\"Skipped 2 rows where the value couldn't be parsed: 'X', 'Y'.\"</em></p>\n<p>Surface to the user:</p>\n<blockquote>\n<p>Parsed <code>&lt;N&gt;</code> rows from <code>&lt;the screenshot / csv_path&gt;</code>. Reconciling now — typically <code>&lt;N&gt; × 5–15s</code> per row depending on query latency.</p>\n</blockquote>\n<hr>\n<h2>Phase 2: Generate questions + execute</h2>\n<p>For each row in the parsed list:</p>\n<h3>2a — Generate the NL question</h3>\n<p>Use the LLM to translate <code>label</code> (+ context if present) into the most natural English question whose answer should be <code>expected_value</code>. Examples:</p>\n<table>\n<thead>\n<tr>\n<th>label</th>\n<th>question</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>Q3 2025 Revenue</code></td>\n<td>\"What was total revenue in Q3 2025?\"</td>\n</tr>\n<tr>\n<td><code>Active customers (Apr 2026)</code></td>\n<td>\"How many active customers did we have in April 2026?\"</td>\n</tr>\n<tr>\n<td><code>Pipeline value (open opps)</code></td>\n<td>\"What's the total pipeline value across open opportunities?\"</td>\n</tr>\n<tr>\n<td><code>Mean order size last 30 days</code></td>\n<td>\"What's the average order size over the last 30 days?\"</td>\n</tr>\n</tbody>\n</table>\n<p>The semantic model + examples library are loaded; let the LLM pick the right subject areas / entities / metrics that resolve the labeled term to a concrete query.</p>\n<h3>2b — Run via the agami-query pipeline</h3>\n<p>Invoke the same SQL-generation + execution path agami-query uses (Phases 2 + 3 of that skill — see <a href=\"../agami-query/SKILL.md\"><code>agami-query/SKILL.md</code></a>). Capture:</p>\n<ul>\n<li>The generated SQL</li>\n<li>The result (should be a single scalar, or a single row)</li>\n<li>The full chart-template HTML report (so the user can drill in for mismatches)</li>\n<li>The trust receipt (with confidence, signed-off-by, etc.)</li>\n</ul>\n<p>If the SQL fails OR the result isn't a single scalar (e.g., the LLM-generated question returned a multi-row table), capture an error: <code>Could not extract a single scalar from the result.</code> These rows show up as <code>error</code> status in the report.</p>\n<h3>2c — Diff</h3>\n<pre><code>python3 \"$AGAMI_PLUGIN_ROOT/scripts/reconcile.py\" diff \\\n  --expected \"&lt;expected_value&gt;\" \\\n  --actual \"&lt;actual_value_from_query&gt;\" \\\n  --tolerance 0.01\n</code></pre>\n<p>Default tolerance: ±1%. The user can override with <code>tolerance=N%</code> in their original ask (e.g., \"reconcile with 5% tolerance\"). Tolerance applies to numeric comparisons; for text values (rare), use exact match.</p>\n<p>Capture: <code>match</code> (bool), <code>delta</code>, <code>delta_pct</code>.</p>\n<h3>2d — Build the row record</h3>\n<p>Per row:</p>\n<pre><code>{\n  \"label\":        \"&lt;from CSV&gt;\",\n  \"question\":     \"&lt;LLM-generated NL question&gt;\",\n  \"expected\":     &lt;number&gt;,\n  \"actual\":       &lt;number or null if errored&gt;,\n  \"delta_pct\":    &lt;signed fraction or null&gt;,\n  \"match\":        true | false,\n  \"status\":       \"match\" | \"mismatch\" | \"error\",\n  \"report_path\":  \"&lt;artifacts_dir&gt;/local/charts/&lt;profile&gt;/&lt;ts&gt;.html\",  // the full chart report for this query\n  \"error\":        \"&lt;message if status=error, else null&gt;\"\n}\n</code></pre>\n<p>Append all records to <code>/tmp/agami-reconcile-results-&lt;ts&gt;.jsonl</code> so the user can inspect later.</p>\n<hr>\n<h2>Phase 3: Present</h2>\n<h3>3a — Summary line first</h3>\n<pre><code>Reconciled &lt;N&gt; numbers: &lt;M&gt; match (within ±1%), &lt;K&gt; mismatch, &lt;E&gt; error.\n</code></pre>\n<h3>3b — Mismatches table (lead with what didn't match)</h3>\n<p>Render the mismatches as a markdown table BEFORE the matches:</p>\n<pre><code>### Mismatches\n\n| Label | Expected | Got | Δ | Drill-down |\n|---|---:|---:|---:|---|\n| Q3 2025 Revenue          | $4,200,000 | $3,890,000 | -7.4% | &lt;artifacts_dir&gt;/local/charts/&amp;lt;profile&amp;gt;/...html |\n| Active customers (Apr)   | 12,450     | 11,920     | -4.3% | &lt;artifacts_dir&gt;/local/charts/&amp;lt;profile&amp;gt;/...html |\n</code></pre>\n<p>Cell formatting:</p>\n<ul>\n<li>Numbers carry the same currency / magnitude suffix as the input where unambiguous (echo the user's <code>raw_value</code> for <code>Expected</code>, format <code>Got</code> with the same shape).</li>\n<li><code>Δ</code> is the signed percent (red wins emphasized in chat by ✗ prefix if Markdown rendering allows; otherwise plain text).</li>\n<li><code>Drill-down</code> links to the chart-template HTML for that query. Open these to see the full receipt — that's where the definitional disagreement lives.</li>\n</ul>\n<p>For each mismatch row, surface a one-line interpretation under the table:</p>\n<blockquote>\n<p><strong>Q3 2025 Revenue</strong> — agami reports $3.89M, your dashboard says $4.2M (-7.4%). Open the receipt; the metric <code>revenue</code> here is <em>gross of refunds in USD at invoice date</em>. If your dashboard nets refunds, that's the gap.</p>\n</blockquote>\n<p>This is where the trust win lands. The DE doesn't have to chase the disagreement — the receipt + your interpretation does it for them.</p>\n<h3>3c — Errors block (if any)</h3>\n<pre><code>### Errors\n\n| Label | What went wrong |\n|---|---|\n| Pipeline value (open opps) | Could not extract a single scalar — the question returned 47 rows. Try rephrasing or check that the metric exists in the model. |\n</code></pre>\n<h3>3d — Matches summary (last, compact)</h3>\n<pre><code>### Matches (within ±1%)\n\n7 numbers reproduced cleanly: Q3 2025 Orders, MoM growth, Avg order size, Top customer, Customer count by region, Refund rate, Pipeline count.\n</code></pre>\n<p>Don't dump every match's drill-down — they're not interesting. The matches build the case; the mismatches drive the conversation.</p>\n<h3>3e — Closing prompt</h3>\n<pre><code>Re-run with `tolerance=5%` to see softer matches, or open any drill-down to find the definitional gap.\n</code></pre>\n<p>End the turn. The user typically:</p>\n<ul>\n<li>Opens a mismatch's drill-down, finds the definitional gap, says <em>\"the dashboard is gross-of-refunds; can we update the metric?\"</em> — chain into <code>/agami-save-correction</code> to update the metric definition.</li>\n<li>Asks <code>tolerance=5%</code> to widen the matches.</li>\n<li>Asks for a different CSV.</li>\n</ul>\n<hr>\n<h2>Hard rules</h2>\n<ol>\n<li><strong>No automatic question generation for ambiguous labels.</strong> If the label is too short or too vague (e.g., <code>Total</code>, <code>Number</code>, <code>Value</code>), surface to the user: <em>\"Row 5's label is just 'Total' — too ambiguous to translate to a question. Skipping. Add more context to the CSV (e.g., <code>Total Revenue Q3</code> instead of <code>Total</code>) and re-run.\"</em> Don't guess.</li>\n<li><strong>Receipt is non-optional.</strong> Every per-row run MUST produce a chart-template HTML report with the trust receipt — that's what the drill-down link points at, and it's what makes mismatches actionable. If the underlying query path can't produce a receipt (legacy pre-trust-layer model), refuse with: <em>\"This profile pre-dates the trust-layer launch. Re-run <code>/agami-connect</code> to enable receipts, then retry.\"</em></li>\n<li><strong>Don't write to the semantic model from this skill.</strong> Reconcile reads + diffs; it never mutates. If a definitional disagreement surfaces and the user wants to update the metric, route them through <code>/agami-save-correction</code>.</li>\n<li><strong>CSV stays local.</strong> Don't upload, don't summarize-and-send. The reconcile run produces local artifacts (<code>/tmp/agami-reconcile-results-*.jsonl</code> + the per-query chart HTML) and nothing leaves the machine.</li>\n</ol>\n<hr>\n<h2>Error handling cheat sheet</h2>\n<table>\n<thead>\n<tr>\n<th>Symptom</th>\n<th>Action</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>&lt;csv&gt;</code> doesn't exist</td>\n<td>Refuse with one line: \"File not found: <code>&lt;path&gt;</code>.\"</td>\n</tr>\n<tr>\n<td>CSV has 0 parseable rows</td>\n<td>Refuse: \"No rows with parseable numeric values. Common cause: the value column has formatting like <code>$1,234.56 (USD)</code> — try simplifying to <code>1234.56</code>.\"</td>\n</tr>\n<tr>\n<td>Every row errors out</td>\n<td>Surface a meta-error: \"All </td>\n</tr>\n<tr>\n<td>Single mismatch but huge delta (&gt; 100%)</td>\n<td>Note in the interpretation: \"The delta is large enough to suggest a unit mismatch (cents vs dollars, count vs percentage) rather than a definition gap. Check <code>agami.unit</code> on the relevant field.\"</td>\n</tr>\n<tr>\n<td>User pastes inline CSV instead of a path</td>\n<td>Accept it. Write to <code>/tmp/agami-reconcile-pasted-&lt;ts&gt;.csv</code> and proceed.</td>\n</tr>\n<tr>\n<td>Screenshot is blurry / a value is cut off / can't read a tile</td>\n<td>Don't guess the number. Extract what's legible, and tell the user which tiles you skipped: \"Couldn't read 'Pipeline value' clearly — re-snip it or type that one in.\"</td>\n</tr>\n<tr>\n<td>User says \"reconcile my dashboard\" but attaches nothing</td>\n<td>Ask for the screenshot (or CSV / pasted numbers) per Phase 0.4 — don't proceed without the expected numbers.</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h2>Hard rule for screenshots</h2>\n<p>The screenshot is an <strong>image of numbers</strong>, and a misread expected value reads exactly like a model bug. So: (1) the value is parsed by <code>reconcile.py</code>, never by eyeballing; (2) the extracted <code>(label, value)</code> table is <strong>always confirmed with the user before any query runs</strong> (Phase 1 vision branch). The image stays local — same as the CSV (Hard rule #4); it's never uploaded or summarized off-machine.</p>\n<hr>\n<h2>Roadmap (not in v1)</h2>\n<ul>\n<li><strong>Tableau / Looker / Mode export parsing</strong> — parse <code>.twb</code> / <code>.twbx</code> / JSON exports directly (today a screenshot of any of them already works via the vision branch).</li>\n<li><strong>Recurring reconcile runs</strong> — wire into <code>agami test</code> so the golden-test suite includes reconciliation against a pinned dashboard.</li>\n</ul>\n","files":[{"path":"SKILL.md","sizeBytes":85047,"isText":true}],"reviewScore":null,"reviewSummary":null,"trust":{"provenance":"trusted-source-unreviewed","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow.","bodySource":null},"bodyLocked":false,"purchaseUrl":null,"sourceUrl":null,"report":{"provenance":"trusted-source-unreviewed","screen":{"ran":true,"outcome":"clean","suspicious":0,"notes":0,"hiddenCharacters":false},"virusScan":{"engine":"clamav","status":"clean","scannedAt":"2026-09-25T07:37:40.427709Z","sha256":"CAF1E26F3609FA6AAEE667E63128198E63300A8F866EBCB7B36842FD324D1BB7","sizeBytes":29854},"review":null,"source":{"repositoryUrl":"https://github.com/AgamiAI/agami-core","path":"plugins/agami/skills/agami-reconcile","license":null,"commit":"e6eb2304e9817300d2436d108318780c21e9ebd6","subtreeSha":"8B023970F17461BC443A15EDA99FCAEBBFBC8FB8ACCBBEA5B4FD8304E5E40066","lastSyncedAt":"2026-10-04T15:24:10.096958Z"},"reviewedAt":"2026-09-25T07:37:44.2609Z","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow."},"install":[{"target":"skills-cli","command":"npx skills add https://github.com/AgamiAI/agami-core/tree/main/plugins/agami/skills/agami-reconcile"},{"target":"claude-code","command":"claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install agamiai-agami-core@llmmart"},{"target":"git","command":"git clone https://github.com/AgamiAI/agami-core.git"}]}