talkthrough
Analyze narrated screen recordings and audio files through the talkthrough MCP server — triage feedback into findings, extract specs/backlogs/action items from recordings, and correlate spoken remarks with logs via wall-clock timestamps. Use when the user mentions a screen record
Install
npx skills add https://github.com/korovin-aa97/talkthrough-mcp/tree/main/integrations/claude-code/skills/talkthrough
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install korovin-aa97-talkthrough-mcp@llmmart
git clone https://github.com/korovin-aa97/talkthrough-mcp.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole korovin-aa97/talkthrough-mcp collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Analyzing narrated recordings with talkthrough
The talkthrough MCP server turns a local video/audio file into queryable structured data: timestamped transcript segments, scene keyframes, OCR'd on-screen text, and wall-clock anchoring. No LLM inside — you bring the reasoning; it brings the evidence. Everything is lazy and token-budgeted: never ask for more than the moment you are analyzing.
Prerequisite
The talkthrough MCP server must be connected (tools like
process_media / get_transcript are visible). If not, tell the user to
install it: claude mcp add -s user talkthrough -- uvx --python ">=3.11,<3.14" "talkthrough-mcp[diarization,url]"
(see the repository README for other clients).
Core workflow
- Ingest once:
process_media(path)— idempotent by content hash; re-calls on the same file return instantly. Given a public video/audio URL instead of a file, callprocess_url(url): the source is downloaded once (the only network step; YouTube needs the[url]extra) and kept inside the job, then everything below is identical and local — never download twice; a repeat call serves the stored job. Long videos take minutes and stream progress. The summary gives youjob_id, counts, wall-clock, and a transcript preview — do NOT dump anything else eagerly. Multi-person recording (meeting/interview)? Adddiarize=true— even when the ask is just "summarize", speaker structure is part of meeting analysis — and — whenever the headcount is known —num_speakers=N(the main accuracy lever): segments getS1/S2/… labels and the summary a talk-time roster. On an already-processed job the amend re-runs ONLY diarization (no re-transcription) — still minutes on long recordings. - Orient:
get_transcript(job_id)(paginate vianext_start_mswhentruncated) orsearch(job_id, "<distinctive word>")to jump straight to the relevant moments (searches speech AND on-screen OCR text). Multi-word search defaults tomatch_mode="all_words"; use"any_word"for broader lexical recall. - Evidence per remark:
get_moment(job_id, t0-2000, t1+2000)— one call returns the transcript slice + up to 3 unique frames + their OCR text + the wall-clock range. This is the workhorse; describeobservedfrom the returned pixels, never from imagination. - Precision when needed:
get_frames(at_ms=...)for nearby keyframes;extract_frame(job_id, at_ms, crop={x,y,w,h})for an exact instant at native resolution (keyframes capture scene changes + a 1 fps floor, so sub-second moments can fall between them). - Keep verified names: after proving an anonymous label's identity,
call
label_speakers(job_id, labels={"S1":"Name"}, evidence={"S1":"intro or frame proof"}). Saved names appear in later transcript, moment, and search calls while rawS1/S2labels remain. If a diarization amend changes labels, those names move tospeaker_names_pending_reviewand stop being identities. Use the stored old-roster context anchors to re-check them. A pending label still in the roster can be confirmed, replaced, or removed; a stale pending label can only be removed withlabels={"Sx":null}. Never use a pending name in minutes or search as though it were active. A fullforce=truerebuild of a job with active or pending names must also usediarize=true; it rebuilds safely and moves every old identity to pending review, while omitting diarization is refused without changing the stored job. - Recall across sessions:
list_jobs()— the store persists; a file processed yesterday (even via CLI) is queryable byjob_idtoday.
Timestamps
Every timestamped result carries t_ms (video-relative) and, when the
recording start is known, t_wall (ISO 8601 real time). Copy t_wall
VERBATIM from the payload — never compute it from t_ms yourself
(hand-derived wall-clocks drift by whole hours). Use t_wall to
correlate remarks with server/app logs (±30 s grep window). If
wall_clock is null or low-confidence, ask the user when the recording
started and re-anchor: process_media(path, recorded_at="<ISO 8601>", force=true); when the job already has speaker identities, include
diarize=true as required by the safe-rebuild contract.
Packaged workflows (server prompts)
Prefer the server prompts when the task matches — they encode the full
method: bug (one recording → evidence-backed GitHub issue draft; silent,
narration-free recordings welcome), triage-recording (screencast →
findings JSON per the contract in examples/output-contract.schema.json),
spec-from-workshop, backlog-from-demo, meeting-actions (audio-only
friendly), correlate-with-logs.
Rules of thumb
- Audio-only jobs (.m4a/.mp3/…): transcript tools work; frame tools error by design — that error is expected, not a failure.
- Speaker labels are anonymous (
S1/S2, ordered by first voice). Mapping them to names is YOUR job: self-introductions, vocatives, the attendees list — and on video jobs the screen check is MANDATORY: for every label you map,get_frames(at_ms=<that label's longest_turn_at_ms from the roster>)and read the meeting-app name plates, the recording's title card, the active-speaker highlight BEFORE asserting the mapping. STT homophones lie about name spellings (spoken "profit" vs on-screen "Prophet") — trust OCR/frames over the transcript for names. State the mapping explicitly and mark unmapped labels "unidentified". Rostername_candidatesare raw OCR hints, not identities: they may be UI text, a job title, or somebody else's name. Inspect the cited frame and persist only defensible mappings withlabel_speakers; never auto-save a candidate. Aname_candidates_noteon a pre-0.3.1 video job explains that its legacy flat OCR may not yield hints. The job remains readable; regenerate only when useful, withforce=true, diarize=true, so old identities become pending review instead of being lost.diarize=trueneeds the[diarization]extra — its absence produces an actionable install-hint error. - Findings/quotes must cite the narrator's exact words +
t_ms(+t_wallwhen known) + the frame files you actually inspected. - Low STT/vision confidence → surface a question; never silently guess.
- Any narration language works (Whisper auto-detects; the summary reports
language+language_probability). Garbled transcript or low/wrong detection → re-callprocess_media(path, model="large-v3-turbo", force=true)(best multilingual quality) or pinlanguage="…"; domain jargon → passvocabulary="Term1, Term2". - Write digests/summaries for the recording author in the narrator's language; keep quotes verbatim in the original — translate in your own prose only, never inside a quote.
Files (talkthrough-mcp)
-
SKILL.md 7.3 KB
--- name: talkthrough description: Analyze narrated screen recordings and audio files through the talkthrough MCP server — triage feedback into findings, extract specs/backlogs/action items from recordings, and correlate spoken remarks with logs via wall-clock timestamps. Use when the user mentions a screen recording, screencast, narrated video/audio file, or asks to "watch" a recording and act on it. license: MIT metadata: author: korovin-aa97 repository: https://github.com/korovin-aa97/talkthrough-mcp --- # Analyzing narrated recordings with talkthrough The talkthrough MCP server turns a local video/audio file into queryable structured data: timestamped transcript segments, scene keyframes, OCR'd on-screen text, and wall-clock anchoring. No LLM inside — you bring the reasoning; it brings the evidence. Everything is lazy and token-budgeted: never ask for more than the moment you are analyzing. ## Prerequisite The `talkthrough` MCP server must be connected (tools like `process_media` / `get_transcript` are visible). If not, tell the user to install it: `claude mcp add -s user talkthrough -- uvx --python ">=3.11,<3.14" "talkthrough-mcp[diarization,url]"` (see the repository README for other clients). ## Core workflow 1. **Ingest once**: `process_media(path)` — idempotent by content hash; re-calls on the same file return instantly. Given a public video/audio URL instead of a file, call `process_url(url)`: the source is downloaded once (the only network step; YouTube needs the `[url]` extra) and kept inside the job, then everything below is identical and local — never download twice; a repeat call serves the stored job. Long videos take minutes and stream progress. The summary gives you `job_id`, counts, wall-clock, and a transcript preview — do NOT dump anything else eagerly. Multi-person recording (meeting/interview)? Add `diarize=true` — even when the ask is just "summarize", speaker structure is part of meeting analysis — and — whenever the headcount is known — `num_speakers=N` (the main accuracy lever): segments get `S1`/`S2`/… labels and the summary a talk-time roster. On an already-processed job the amend re-runs ONLY diarization (no re-transcription) — still minutes on long recordings. 2. **Orient**: `get_transcript(job_id)` (paginate via `next_start_ms` when `truncated`) or `search(job_id, "<distinctive word>")` to jump straight to the relevant moments (searches speech AND on-screen OCR text). Multi-word search defaults to `match_mode="all_words"`; use `"any_word"` for broader lexical recall. 3. **Evidence per remark**: `get_moment(job_id, t0-2000, t1+2000)` — one call returns the transcript slice + up to 3 unique frames + their OCR text + the wall-clock range. This is the workhorse; describe `observed` from the returned pixels, never from imagination. 4. **Precision when needed**: `get_frames(at_ms=...)` for nearby keyframes; `extract_frame(job_id, at_ms, crop={x,y,w,h})` for an exact instant at native resolution (keyframes capture scene changes + a 1 fps floor, so sub-second moments can fall between them). 5. **Keep verified names**: after proving an anonymous label's identity, call `label_speakers(job_id, labels={"S1":"Name"}, evidence={"S1":"intro or frame proof"})`. Saved names appear in later transcript, moment, and search calls while raw `S1`/`S2` labels remain. If a diarization amend changes labels, those names move to `speaker_names_pending_review` and stop being identities. Use the stored old-roster context anchors to re-check them. A pending label still in the roster can be confirmed, replaced, or removed; a stale pending label can only be removed with `labels={"Sx":null}`. Never use a pending name in minutes or search as though it were active. A full `force=true` rebuild of a job with active or pending names must also use `diarize=true`; it rebuilds safely and moves every old identity to pending review, while omitting diarization is refused without changing the stored job. 6. **Recall across sessions**: `list_jobs()` — the store persists; a file processed yesterday (even via CLI) is queryable by `job_id` today. ## Timestamps Every timestamped result carries `t_ms` (video-relative) and, when the recording start is known, `t_wall` (ISO 8601 real time). Copy `t_wall` VERBATIM from the payload — never compute it from `t_ms` yourself (hand-derived wall-clocks drift by whole hours). Use `t_wall` to correlate remarks with server/app logs (±30 s grep window). If `wall_clock` is null or low-confidence, ask the user when the recording started and re-anchor: `process_media(path, recorded_at="<ISO 8601>", force=true)`; when the job already has speaker identities, include `diarize=true` as required by the safe-rebuild contract. ## Packaged workflows (server prompts) Prefer the server prompts when the task matches — they encode the full method: `bug` (one recording → evidence-backed GitHub issue draft; silent, narration-free recordings welcome), `triage-recording` (screencast → findings JSON per the contract in `examples/output-contract.schema.json`), `spec-from-workshop`, `backlog-from-demo`, `meeting-actions` (audio-only friendly), `correlate-with-logs`. ## Rules of thumb - Audio-only jobs (.m4a/.mp3/…): transcript tools work; frame tools error by design — that error is expected, not a failure. - Speaker labels are anonymous (`S1`/`S2`, ordered by first voice). Mapping them to names is YOUR job: self-introductions, vocatives, the attendees list — and on video jobs the screen check is MANDATORY: for every label you map, `get_frames(at_ms=<that label's longest_turn_at_ms from the roster>)` and read the meeting-app name plates, the recording's title card, the active-speaker highlight BEFORE asserting the mapping. STT homophones lie about name spellings (spoken "profit" vs on-screen "Prophet") — trust OCR/frames over the transcript for names. State the mapping explicitly and mark unmapped labels "unidentified". Roster `name_candidates` are raw OCR hints, not identities: they may be UI text, a job title, or somebody else's name. Inspect the cited frame and persist only defensible mappings with `label_speakers`; never auto-save a candidate. A `name_candidates_note` on a pre-0.3.1 video job explains that its legacy flat OCR may not yield hints. The job remains readable; regenerate only when useful, with `force=true, diarize=true`, so old identities become pending review instead of being lost. `diarize=true` needs the `[diarization]` extra — its absence produces an actionable install-hint error. - Findings/quotes must cite the narrator's exact words + `t_ms` (+ `t_wall` when known) + the frame files you actually inspected. - Low STT/vision confidence → surface a question; never silently guess. - Any narration language works (Whisper auto-detects; the summary reports `language` + `language_probability`). Garbled transcript or low/wrong detection → re-call `process_media(path, model="large-v3-turbo", force=true)` (best multilingual quality) or pin `language="…"`; domain jargon → pass `vocabulary="Term1, Term2"`. - Write digests/summaries for the recording author in the narrator's language; keep quotes verbatim in the original — translate in your own prose only, never inside a quote.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.