ytdlp-ops
yt-dlp media acquisition layer feeding ffmpeg-ops: format selection avoiding post-download transcodes, clip-at-download, cookies/auth, channel archive sync, SponsorBlock, subtitles, failure triage (403s, nsig). Triggers on: yt-dlp, download video/playlist/channel, youtube to mp3.
Install
npx skills add https://github.com/0xDarkMatter/claude-mods/tree/main/skills/ytdlp-ops
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install 0xdarkmatter-claude-mods@llmmart
git clone https://github.com/0xDarkMatter/claude-mods.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole 0xdarkmatter/claude-mods collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
yt-dlp Operations
Operational expertise for yt-dlp as the acquisition layer: get the right bytes
onto disk in the right codec, politely, resumably — then hand off. Anything that
re-encodes, cuts precisely, grades, or packages after download is
ffmpeg-ops territory; AI-driven editing of what you
acquired (transcript → EDL → final cut) is cutcraft — a separate tool, not shipped
by this repo. The full chain is acquire → process → edit.
Doctrine: version first, formats second
yt-dlp vs the platforms is an arms race. Releases land near-monthly and extractors break between them — the majority of "yt-dlp is broken" reports are a stale binary. Before debugging anything, check staleness:
bash skills/ytdlp-ops/scripts/check-ytdlp-version.sh --live # vs latest GitHub release
bash skills/ytdlp-ops/scripts/check-ytdlp-version.sh --live --json | jq '.data.days_behind'
Exit 10 = installed build is >60 days behind latest, a documented flag
vanished from yt-dlp --help, or the smoke extraction hit an extractor-internal
error → update before any other triage. Exit 7 is advisory: the check ran but
could not verify extraction (YouTube bot-gated our IP, no JS runtime, or an
unrecognised failure) — common from datacenter IPs such as CI runners:
uv tool upgrade yt-dlp # pip/uv-managed install (preferred)
yt-dlp -U # standalone binary self-update only
Second rule: pick codecs at download time. The default "best" on YouTube is
VP9/AV1 + Opus in WebM/MKV. If the destination needs H.264 MP4, stating that in
-S costs nothing — discovering it after download costs a full transcode.
Cookbook
Format selection (-S over -f)
# Declarative sort (-S) — PREFER this. States preferences in priority order and
# always degrades gracefully to the nearest available. h264 + m4a merges
# natively into mp4: zero post-download transcode.
yt-dlp -S "res:1080,vcodec:h264,acodec:m4a" --merge-output-format mp4 URL
# Hard filter (-f) — exact control, but FAILS ("Requested format is not
# available") when nothing matches. Use only for genuine hard requirements,
# always with a / fallback chain:
yt-dlp -f "bv*[height<=1080][vcodec^=avc1]+ba[ext=m4a]/b[height<=1080]/b" URL
# Survey what the extractor actually offers before arguing with selectors:
yt-dlp -F URL
# Smallest acceptable file (bandwidth/storage constrained; + prefix = ascending):
yt-dlp -S "res:480,+size,+br" URL
# Best quality regardless of codec (archival source for later ffmpeg-ops work):
yt-dlp -S "res,fps,hdr:12,vcodec,acodec" --merge-output-format mkv URL
Sort-field reference, filter grammar, per-destination presets: references/format-selection.md + assets/format-presets.json.
Clip at download (--download-sections)
# Download ONLY 10:00-12:30 — ranged requests, not a full download + trim:
yt-dlp --download-sections "*10:00-12:30" -S "res:1080,vcodec:h264" URL
# Frame-accurate cut points (re-encodes around the cuts only):
yt-dlp --download-sections "*10:00-12:30" --force-keyframes-at-cuts URL
# Last 5 minutes / by chapter-title regex / multiple sections:
yt-dlp --download-sections "*-5:00-inf" URL
yt-dlp --download-sections "Intro" --download-sections "Outro" URL
Same physics as ffmpeg copy-cuts: without --force-keyframes-at-cuts the section
boundaries snap to keyframes (can be seconds off). Need many precise cuts from
one source? Download once, then use the ffmpeg-ops EDL workflow.
Audio-only extraction (STT pipelines)
# THE STT acquisition command. YouTube's best audio IS Opus — asking for opus
# means -x COPIES the stream out (no transcode, no quality loss):
yt-dlp -x --audio-format opus -o "%(id)s.%(ext)s" URL
# Zero-processing alternative — native container, no ffmpeg step at all:
yt-dlp -f "ba" -o "%(id)s.%(ext)s" URL
# Whole channel's audio for a transcription pipeline (archive = resumable):
yt-dlp -x --audio-format opus --download-archive stt-archive.txt \
-o "%(channel)s/%(id)s.%(ext)s" CHANNEL_URL
Do NOT --audio-format mp3 for STT — that's a lossy→lossy transcode that helps
nothing. Whisper-prep (16 kHz mono PCM) is the next stage:
ffmpeg-ops stt-whisper.
Playlists, channels, incremental sync
# Playlist with ID-correlated filenames + archive file (resumable, dedup-safe):
yt-dlp --download-archive archive.txt \
-o "%(playlist)s/%(playlist_index)03d - %(title).100B [%(id)s].%(ext)s" PLAYLIST_URL
# Incremental channel sync (cron-friendly): stop at the first already-archived
# video instead of re-walking the entire channel every run:
yt-dlp --download-archive archive.txt --break-on-existing --lazy-playlist \
-S "res:1080,vcodec:h264,acodec:m4a" CHANNEL_URL
# Subset selection / list without downloading:
yt-dlp -I 1:10 PLAYLIST_URL
yt-dlp --flat-playlist --print "%(id)s %(title)s" PLAYLIST_URL
# DRY-RUN any batch before committing to it — preview every output filename
# (--print implies --simulate; nothing downloads):
yt-dlp --print filename -o "%(playlist_index)03d - %(title).100B [%(id)s].%(ext)s" PLAYLIST_URL
Archive format, sync-job patterns, when --break-on-existing misfires
(non-chronological playlists):
references/playlists-archives.md.
Livestreams and premieres
# Capture a livestream from its BEGINNING, not from "now" (YouTube keeps a
# rolling live buffer; without this you get the moment you pressed enter):
yt-dlp --live-from-start URL
# Scheduled premiere/stream: poll (1-10 min between retries) and start when live:
yt-dlp --wait-for-video 60-600 URL
Live capture caveats: a crashed live download is not resumable like a VOD
(fragments expire) — write to fast local disk (-P temp:), not a network share.
For archival quality, prefer re-downloading the VOD after the stream ends; the
live manifest often caps below the post-processed VOD.
Subtitles
# Manual subs, English variants, skip live-chat pseudo-subs, as SRT:
yt-dlp --write-subs --sub-langs "en.*,-live_chat" --convert-subs srt --skip-download URL
# Auto-generated (ASR) captions — exist for most videos when manual subs don't:
yt-dlp --write-auto-subs --sub-langs en --convert-subs srt --skip-download URL
# Embed into the media file instead of a sidecar:
yt-dlp --embed-subs --sub-langs en URL
Sub formats, language matching, transcript-only workflows (subs as cheap STT): references/subtitles-metadata.md.
SponsorBlock
# Mark segments as chapters — LOSSLESS and reversible. Prefer this:
yt-dlp --sponsorblock-mark all URL
# Cut segments out of the media — modifies the file, re-encodes at boundaries:
yt-dlp --sponsorblock-remove sponsor,selfpromo URL
Category list, mark-vs-remove trade-offs, interaction with --download-sections:
references/sponsorblock.md.
Cookies and auth
# Pull cookies from a browser profile (private/members/age-gated content):
yt-dlp --cookies-from-browser firefox URL
# Chrome 127+ on Windows uses app-bound cookie encryption — extraction usually
# FAILS. Use Firefox, or export a Netscape cookies.txt and pass it directly:
yt-dlp --cookies cookies.txt URL
Account-ban warning: authenticated bulk downloading is the fastest way to get an account flagged. Use a throwaway account, always pair cookies with the politeness flags below. Details + browser matrix: references/auth-cookies.md.
Rate limiting and politeness
# The polite-bulk baseline — cap bandwidth, space out requests, retry patiently:
yt-dlp --limit-rate 4M --sleep-requests 1 \
--sleep-interval 5 --max-sleep-interval 15 \
--retries 10 --fragment-retries 10 URL
# Speed (single video, host not throttling you): parallel fragment download:
yt-dlp --concurrent-fragments 4 URL
Politeness is self-interest: 429s and IP flags cost more time than sleeps do.
Remux vs recode
# Remux: container change only — lossless, near-instant. yt-dlp's job:
yt-dlp -S "vcodec:h264,acodec:m4a" --remux-video mp4 URL
# Recode: a FULL TRANSCODE. Almost never yt-dlp's job — you give up ffmpeg-ops'
# CRF/preset/pix_fmt control for a blind default encode. If codecs must change:
yt-dlp -S "res,vcodec,acodec" URL # 1. acquire best-native
# 2. then transcode with the ffmpeg-ops web-compatible H.264 recipe.
Rule: --remux-video whenever the codecs already fit the target container;
--recode-video only for throwaway one-offs where quality control doesn't matter.
Output templates
# ID-in-brackets convention — survives renames, correlates with archive files:
yt-dlp -o "%(uploader)s/%(upload_date)s - %(title).100B [%(id)s].%(ext)s" URL
# Cross-filesystem safety (strips spaces/unicode to ASCII-safe names):
yt-dlp --restrict-filenames -o "%(title)s [%(id)s].%(ext)s" URL
# Split destination and scratch space (-P): fragments go to temp, final to home:
yt-dlp -P "D:/media" -P "temp:C:/tmp/ytdlp" URL
%(title).100B truncates at 100 bytes (UTF-8 safe — CJK titles break
char-based truncation). Full field catalog and per-type templates:
references/output-templates.md.
Metadata embedding
# Self-describing files — metadata, thumbnail and chapters travel with the media:
yt-dlp --embed-metadata --embed-thumbnail --embed-chapters URL
Beyond YouTube
yt-dlp ships ~1,800 extractors (yt-dlp --list-extractors); everything in this
skill except the YouTube-specific parts (nsig, player clients) applies unchanged
to Twitch, Vimeo, SoundCloud, TikTok, and the rest. yt-dlp -v URL names the
extractor in use. For sites with no dedicated extractor, the generic extractor
sniffs direct media/HLS URLs out of the page. When a non-YouTube site returns
403 to yt-dlp but plays fine in a browser, it's usually TLS-fingerprint
blocking — --impersonate fixes it (see
failure-triage).
Footguns
| Footgun | The trap | The rule |
|---|---|---|
| Default format selection | YouTube "best" = VP9/AV1+Opus in WebM/MKV; downstream tooling expecting MP4 forces a transcode you could have avoided | State codecs at download: -S "vcodec:h264,acodec:m4a" --merge-output-format mp4 |
-f best |
Selects best single pre-merged file — caps at ~720p on YouTube; modern high-res is always video+audio merged | Drop the -f entirely or use -S; b only as the tail of a / fallback chain |
-f hard filters |
"Requested format is not available" the moment an extractor stops offering that exact combo | Prefer -S (degrades gracefully); always end -f chains with /b |
--recode-video casually |
Full blind transcode — no CRF/preset/pix_fmt control, big quality/time cost | --remux-video when codecs fit; real transcodes via ffmpeg-ops |
--download-sections w/o --force-keyframes-at-cuts |
Clip boundaries snap to keyframes — seconds of slop | Add the flag when cuts must be exact (re-encodes at cuts only) |
Channel sync w/o --break-on-existing |
Every cron run re-walks the entire channel (thousands of metadata requests) | --download-archive + --break-on-existing --lazy-playlist |
No %(id)s in filename |
Title changes/dupes make files impossible to correlate with the archive | Always [%(id)s] in the template |
--cookies-from-browser chrome on Windows |
Chrome 127+ app-bound encryption — extraction fails | Use firefox, or export cookies.txt |
| Authenticated bulk runs, no sleeps | Account flagged/banned; IP rate-limited | Throwaway account + --sleep-requests/--sleep-interval always |
| Throttled to ~50-100 KB/s | Looks like a network problem; it's the nsig arms race | Update yt-dlp FIRST (check-ytdlp-version.sh --live) |
| "nsig extraction failed" / "unable to extract" | Debugging the command/network when the binary is stale | Same — update first; these errors mean outdated, not broken usage |
Raw %(title)s filenames |
Emoji/colons/slashes break on Windows and some CI filesystems | --restrict-filenames or .100B-truncated fields + [%(id)s] |
| Thin format list on a fresh machine | No JS runtime — YouTube player JS now needs one (EJS); runtime-less extraction is deprecated and may offer only low-res premuxed | Install deno, or --js-runtimes node; see failure-triage |
| git-bash (MSYS) path mangling | /tmp/...-style args convert per-arg — templates containing %(...)s skip conversion while plain paths convert, scattering outputs |
Use Windows-style paths (X:/dir/...) for -o/-P/--download-archive under git-bash |
pip-installed yt-dlp -U |
Self-update doesn't work for pip/uv installs (silently a no-op with a warning) | uv tool upgrade yt-dlp; -U is for the standalone binary only |
Failure triage
The ladder — run in order, stop at the first fix:
- Stale binary?
check-ytdlp-version.sh --live→ exit 10 → update. This closes most "nsig extraction failed", missing-format, and throttling cases. - Reproduce verbosely:
yt-dlp -v URL— read the actual extractor error, don't guess from the summary line. - 403 / "Sign in to confirm you're not a bot" → identity problem:
--cookies-from-browser firefox, or a different network/IP. - 429 / sudden slowdowns mid-run → rate limited: add the politeness flags,
reduce
--concurrent-fragments, back off and resume later (archive files make every run resumable). - Geo block ("not available in your country") →
--proxy URLthrough an allowed region; the old--geo-bypassheader tricks rarely work anymore.
Full decision tree with error-message → cause mapping, --extractor-args
escape hatches, and when to file upstream:
references/failure-triage.md.
Scripts
Follows the Skill Resource Protocol:
--help with examples, stdout = data only, --json envelope
(claude-mods.ytdlp-ops.version-check/v1), semantic exit codes (0 clean,
2 usage, 7 network/yt-dlp unavailable — advisory, 10 drift finding).
| Script | Job | Worked invocation |
|---|---|---|
check-ytdlp-version.sh |
Staleness verifier: --offline structural (CI gate), --live = installed-version age vs latest GitHub release + documented-flag existence in yt-dlp --help + metadata-only smoke extraction |
check-ytdlp-version.sh --live --json \| jq '.data.days_behind' — exit 10 = >60 days behind, a documented flag vanished, or an extractor-internal smoke failure; 7 = advisory (network/API unreachable, or the smoke test was blocked/unverifiable — normal on a bot-gated datacenter IP) |
References
Load on demand — one concept per file:
| Reference | Load when |
|---|---|
| format-selection.md | Any -f/-S decision, codec targeting, filter grammar, avoiding transcodes |
| playlists-archives.md | Playlists, channels, --download-archive, incremental sync jobs |
| auth-cookies.md | Private/members/age-gated content, browser cookie matrix, ban avoidance |
| output-templates.md | -o field catalog, paths, sanitization, per-type routing |
| subtitles-metadata.md | Sub download/convert/embed, transcript workflows, metadata/thumbnail embedding |
| sponsorblock.md | SponsorBlock categories, mark vs remove, chapter workflows |
| failure-triage.md | Any download failure — 403/429/geo/nsig/throttling decision tree |
Assets: format-presets.json — canonical, date-stamped
-S/flag presets per destination (web MP4, STT audio, archival, clip, mobile-small).
Self-test
bash skills/ytdlp-ops/tests/run.sh # fully offline; no network, no yt-dlp needed
Structural assertions plus the verifier's 60-day age logic exercised through its
CM_YTDLP_INSTALLED/CM_YTDLP_LATEST test seams. Real --live runs happen only
in the scheduled freshness workflow — a network blip must never fail a PR.
Files (claude-mods)
-
assets
-
format-presets.json 2.4 KB
{ "_meta": { "schema": "claude-mods.ytdlp-ops.format-presets/v1", "updated": "2026-06-12", "notes": "Canonical yt-dlp argument presets per destination. Args are list-form (paste-safe, no shell quoting surprises). Presets prefer -S sort over -f filters so they degrade gracefully when an extractor stops offering an exact combo. Verify against a current yt-dlp with scripts/check-ytdlp-version.sh --live." }, "presets": { "web-mp4-1080": { "goal": "H.264/AAC MP4 ready for browsers and editors - zero post-download transcode", "args": ["-S", "res:1080,vcodec:h264,acodec:m4a", "--merge-output-format", "mp4"], "handoff": "none - file is delivery-ready; ffmpeg-ops only if further editing is needed" }, "archival-best": { "goal": "Best available quality regardless of codec, as a master for later processing", "args": ["-S", "res,fps,hdr:12,vcodec,acodec", "--merge-output-format", "mkv", "--embed-metadata", "--embed-chapters"], "handoff": "ffmpeg-ops encoding.md for any delivery transcode from this master" }, "stt-audio": { "goal": "Audio-only acquisition for transcription pipelines - Opus copied out, no transcode", "args": ["-x", "--audio-format", "opus", "-o", "%(id)s.%(ext)s"], "handoff": "ffmpeg-ops stt-whisper.md for 16 kHz mono PCM prep" }, "clip-precise": { "goal": "Frame-accurate section download (re-encodes at the cut points only)", "args": ["--download-sections", "*START-END", "--force-keyframes-at-cuts", "-S", "res:1080,vcodec:h264,acodec:m4a"], "handoff": "replace *START-END, e.g. *10:00-12:30; multi-cut edits -> download whole + ffmpeg-ops EDL" }, "mobile-small": { "goal": "Smallest acceptable file at a resolution floor (bandwidth/storage constrained)", "args": ["-S", "res:480,+size,+br", "--merge-output-format", "mp4"], "handoff": "none" }, "channel-sync": { "goal": "Incremental channel/playlist sync - resumable, stops at first already-archived item", "args": ["--download-archive", "archive.txt", "--break-on-existing", "--lazy-playlist", "-S", "res:1080,vcodec:h264,acodec:m4a", "--sleep-requests", "1", "--sleep-interval", "5", "--max-sleep-interval", "15", "-o", "%(channel)s/%(upload_date)s - %(title).100B [%(id)s].%(ext)s"], "handoff": "see references/playlists-archives.md for the cron pattern and non-chronological caveat" } } }
-
-
references
-
auth-cookies.md 3.3 KB
# Cookies and Authentication For private, members-only, age-gated, or "Sign in to confirm you're not a bot" content. Authentication is also the highest-risk feature in yt-dlp: it ties bulk download behaviour to an identifiable account. ## `--cookies-from-browser` (first choice) ```bash yt-dlp --cookies-from-browser firefox URL yt-dlp --cookies-from-browser "firefox:profile-name" URL # specific profile yt-dlp --cookies-from-browser "brave+gnomekeyring" URL # explicit keyring (Linux) ``` Reads cookies straight from a browser profile on disk. Full syntax: `BROWSER[+KEYRING][:PROFILE][::CONTAINER]`. Supported browsers include `firefox`, `chrome`, `chromium`, `edge`, `brave`, `opera`, `vivaldi`, `safari`, `whale`. ### The browser matrix (what actually works) | Browser | Status | |---|---| | **Firefox** | Most reliable everywhere — plain SQLite cookie store. **Default choice.** | | Chrome/Edge/Brave on **Windows** | Chrome 127+ **app-bound encryption** ties cookie decryption to the browser binary — extraction usually fails. Don't fight it; use Firefox or `cookies.txt`. | | Chrome on macOS/Linux | Generally works (keychain/keyring prompt possible); close the browser first — a running Chrome locks the cookie DB | | Safari | Works on macOS; needs Full Disk Access for the terminal | ## `--cookies cookies.txt` (the fallback that always works) ```bash yt-dlp --cookies cookies.txt URL ``` A Netscape-format cookie export (browser extensions like "Get cookies.txt LOCALLY" produce it, or `yt-dlp --cookies-from-browser firefox --cookies out.txt --skip-download URL` converts browser → file once on a machine where extraction works, for use on servers). Treat the file as a **credential**: it grants full account access. Never commit it; `chmod 600`; rotate by re-exporting. Note YouTube rotates session cookies aggressively — exported cookie files go stale in days-to-weeks, so headless boxes need a refresh procedure, not a one-time export. ## Account-ban avoidance (read before any authenticated bulk run) Authenticated + high-volume + fast is the exact signature platforms ban for. The account in the cookies is the blast radius. 1. **Use a throwaway account** for anything bulk. Never a personal/work account. 2. **Always pair auth with politeness flags**: `--sleep-requests 1 --sleep-interval 5 --max-sleep-interval 15 --limit-rate 4M`. 3. **Don't parallelize across the same account/IP** (multiple yt-dlp processes sharing cookies multiplies the signature). 4. Prefer unauthenticated access whenever the content allows it — most public content needs no cookies at all; only add them when an error demands it. ## Username/password (`-u`/`-p`) — mostly dead Direct login triggers 2FA/anti-bot challenges on major platforms and is unsupported for YouTube. Cookies are the auth mechanism; treat `-u/-p` as legacy for the few small sites where it still works. ## "Sign in to confirm you're not a bot" Not strictly an auth wall — an IP-reputation challenge (heavy on datacenter/VPS IPs). Options, in order: 1. `--cookies-from-browser firefox` — a logged-in session usually passes. 2. Run from a residential IP (or proxy through one: `--proxy`). 3. Update yt-dlp — client-impersonation fixes for this challenge ship regularly. See [failure-triage.md](failure-triage.md) for the full error → cause ladder. -
failure-triage.md 7.4 KB
# Failure Triage The arms-race reality: platforms change player code and access rules continuously; yt-dlp ships countermeasures near-monthly. **Most failures are version failures.** Triage in this order — each step is cheaper than the one below it. ## Step 0 — version check (always first) ```bash bash skills/ytdlp-ops/scripts/check-ytdlp-version.sh --live ``` Exit `10` → update (`uv tool upgrade yt-dlp`, or `yt-dlp -U` for the standalone binary) and **re-run the original command before any further debugging**. Errors in the "outdated" class below are *expected* on a stale build; debugging them is wasted time. ## Step 1 — reproduce verbosely ```bash yt-dlp -v URL 2>&1 | tail -40 ``` Read the actual extractor error, not the summary line. `-v` also prints the version, install type (pip/binary), and whether ffmpeg was found — three triage answers for free. ## Error → cause map ### The "outdated yt-dlp" class (update fixes it) | Symptom | What's happening | |---|---| | `nsig extraction failed: Some formats may be missing` | YouTube changed its player JS; the throttling-token solver broke. Formats vanish AND remaining ones may crawl at ~50-100 KB/s | | `Signature extraction failed` | Same family, older mechanism | | `ERROR: Unable to extract <anything>` on a major site | Extractor broke against a site change | | Downloads suddenly throttled to dial-up speeds | Broken nsig solve — the platform serves, but slowly | | Formats that existed last week are gone | Player-client behaviour changed; newer yt-dlp rotates clients | These are **not** network problems, **not** your command, **not** rate limits. Update first. ### Missing formats / "No supported JavaScript runtime could be found" The 2026 evolution of the nsig arms race: yt-dlp now solves YouTube's player JS through an **external JS runtime** (the EJS system); runtime-less extraction is deprecated and silently degrades the format list — often to a single low-res premuxed file. Only **deno** is auto-enabled; node and bun need opt-in: ```bash yt-dlp --js-runtimes node URL # use an installed node (verify: -v shows "JS runtimes: node-…") # or install deno (auto-detected, zero config): https://deno.com ``` Measured effect on the same video: no runtime → premuxed format 18 (360p); with a runtime → the full 395+251 (AV1+Opus) ladder. If formats look thin on a fresh machine, this — not the extractor — is usually why. #### Choosing a runtime (a security decision, not a convenience one) Whatever runtime you pick will execute obfuscated JavaScript fetched from the network on every invocation. Two risks trade off: *install risk* (a new binary on the machine) vs *execution risk* (what privileges that code runs with). | | deno | node opt-in | |---|---|---| | New dependency | yes — one static signed binary, no install scripts, no dep tree | no (if already installed) | | Sandbox | default-deny: no fs/net/env unless granted — why yt-dlp trusts it by default | **none** — full user privileges | | Exposure shape | one-time, auditable at install | standing, re-occurs every invocation, compounds with automation | **Decision rule:** anything unattended (the `--break-on-existing` channel-sync cron, scheduled STT pipelines) → **deno**, no exceptions — recurring unattended execution of network-fetched code must be sandboxed. Occasional interactive use → a *per-invocation* `--js-runtimes node` grant is defensible; do NOT persist it in a config file, because persisted defaults silently become the unattended path when automation arrives later. Install deno with supply-chain discipline — cooldown-checked and version-pinned: ```bash # 1. pick the newest release >=7 days old (skip day-zero releases): curl -fsSL "https://api.github.com/repos/denoland/deno/releases?per_page=5" \ | jq -r '.[] | select(.prerelease|not) | "\(.tag_name) \(.published_at)"' # 2. install that exact version and pin it (Windows; see deno.com for others): winget install --id DenoLand.Deno --version <X.Y.Z> --exact winget pin add --id DenoLand.Deno --version <X.Y.Z> # 3. verify yt-dlp picked it up: bash skills/ytdlp-ops/scripts/check-ytdlp-version.sh --live --json | jq -r '.data.js_runtime' ``` ### HTTP 403 Forbidden 1. Stale build (above) — update first. 2. **Non-YouTube site that plays fine in a browser** — TLS-fingerprint blocking (Cloudflare and friends sniff the client hello). Fix: ```bash yt-dlp --impersonate chrome URL yt-dlp --list-impersonate-targets # what this install can mimic ``` Needs the curl_cffi extra — standalone binaries include it; pip/uv installs need `uv tool install "yt-dlp[default,curl-cffi]"`. An empty target list means the extra is missing. 3. IP reputation — datacenter/VPS IPs are heavily challenged. Try a residential IP or `--proxy socks5://...`. 4. Stale cookies — re-export / re-extract (`--cookies-from-browser firefox`). 5. Mid-download 403 on fragments — URLs expired (very slow download or paused run); just re-run, `.part` files resume. ### "Sign in to confirm you're not a bot" IP-reputation challenge. In order: logged-in cookies (`--cookies-from-browser firefox`), residential IP/proxy, update (client impersonation fixes ship regularly). See [auth-cookies.md](auth-cookies.md). ### HTTP 429 / rate limited You're sending too much, too fast, from one address: ```bash --sleep-requests 1 --sleep-interval 5 --max-sleep-interval 15 --limit-rate 4M ``` Reduce `--concurrent-fragments` to 1, stop parallel processes against the same host, and back off for hours, not seconds. Archives make resumption free. ### Geo blocks ("not available in your country") `--proxy URL` through an allowed region is the real fix. The legacy `--geo-bypass`/`--xff` header spoofing rarely works on major platforms anymore — don't burn time on it. ### Private / deleted / members-only `Private video`, `Video unavailable`, `Join this channel` — access problems, not bugs. Cookies from an account *with that access* (member, accepted viewer) or nothing. In batch runs, `--ignore-errors` keeps one dead video from killing the job. ### "Requested format is not available" Your `-f` hard filter matched nothing (catalog changed, or per-client format availability shifted). `yt-dlp -F URL` to see today's offerings; switch to `-S` sorting which cannot fail this way ([format-selection.md](format-selection.md)). ### ffmpeg-related: `merging of multiple formats` / `ffmpeg not found` yt-dlp needs ffmpeg on PATH for merge/remux/extract-audio. Point at a specific build with `--ffmpeg-location PATH`. Verify what the build can do with ffmpeg-ops `capability-scan.sh`. ## Escape hatch: `--extractor-args` Per-extractor overrides, e.g. forcing alternative player clients: ```bash yt-dlp --extractor-args "youtube:player_client=default,web_safari" URL ``` **Staleness warning:** valid client names and their behaviour churn faster than any doc — treat specific values found in forum posts (including this file's example) as expired until verified against the current [yt-dlp wiki/extractor docs](https://github.com/yt-dlp/yt-dlp/wiki). Reach for this only after an update didn't fix it. ## When it's genuinely upstream Current version + verbose log showing an extractor exception + reproducible on a clean network → check the [issue tracker](https://github.com/yt-dlp/yt-dlp/issues) (it's almost certainly already filed; platform-wide breakages get hundreds of duplicates within hours). Pin your pipeline to "wait for the next release", not to workarounds scraped from the thread. -
format-selection.md 4.6 KB
# Format Selection — `-S` sort vs `-f` filters The single highest-leverage decision in any yt-dlp invocation. Get it right and the file lands in the codec the destination needs; get it wrong and you pay a full transcode (or a hard "Requested format is not available" failure) after the fact. ## The mental model Platforms serve **separate video and audio streams** at high quality. "Downloading a video" is really: pick a video stream, pick an audio stream, merge them (yt-dlp shells out to ffmpeg for the merge). Two ways to steer the pick: | Mechanism | Style | Failure mode | |---|---|---| | `-S` (`--format-sort`) | *Preferences* — "closest to these, in this priority order" | None — always degrades to nearest available | | `-f` (`--format`) | *Filters* — "exactly this, or the next `/` alternative" | Hard error when nothing matches | **Default to `-S`.** Reach for `-f` only when a hard requirement genuinely exists (e.g. a pipeline that breaks on anything but `ext=m4a`), and even then end the chain with `/b` so a catalog change degrades instead of failing. ## `-S` sort fields (the useful subset) Comma-separated, priority order, first field dominates: | Field | Meaning | Example | |---|---|---| | `res:1080` | Resolution closest to but not exceeding 1080p | `res:720` for 720p caps | | `vcodec:h264` | Prefer this video codec family | `h264`, `h265`, `vp9`, `av01` | | `acodec:m4a` | Prefer this audio codec/container family | `m4a` (AAC), `opus` | | `ext` / `ext:mp4` | Prefer this container family | biases toward mp4/m4a | | `fps` | Higher frame rate wins | `fps:30` to cap | | `hdr:12` | Allow up to 12-bit HDR (default sort excludes some HDR) | archival masters | | `+size`, `+br` | `+` prefix inverts: prefer SMALLER size/bitrate | bandwidth-constrained | | `proto` | Prefer better download protocols (https over m3u8) | rarely needed manually | Worked examples: ```bash # Delivery-ready MP4, no transcode (h264 video + AAC audio merge natively): yt-dlp -S "res:1080,vcodec:h264,acodec:m4a" --merge-output-format mp4 URL # Absolute best quality (codec-agnostic master; expect VP9/AV1+Opus in MKV): yt-dlp -S "res,fps,hdr:12,vcodec,acodec" --merge-output-format mkv URL # Smallest file at >=480p-ish (floor the res, then ascend by size and bitrate): yt-dlp -S "res:480,+size,+br" URL ``` ## `-f` selector grammar (when you must) | Token | Meaning | |---|---| | `bv`, `bv*` | best video-only / best video (may include audio) | | `ba`, `ba*` | best audio-only / best audio | | `b` / `best` | best single PRE-MERGED file — on YouTube caps ~720p | | `wv`, `wa`, `w` | worst (testing) | | `+` | merge: `bv+ba` | | `/` | fallback chain, left wins: `bv*+ba/b` | | `[...]` | filter: `[height<=1080]`, `[vcodec^=avc1]`, `[ext=m4a]`, `[filesize<500M]` | Comparison operators: `=`, `!=`, `^=` (starts with), `$=` (ends with), `*=` (contains), and numeric `<`, `<=`, `>`, `>=`. Combine inside one bracket with implicit AND: `[height<=1080][fps<=30]`. ```bash # Exact: H.264 video at <=1080p + AAC audio, fall back to best pre-merged, then anything: yt-dlp -f "bv*[height<=1080][vcodec^=avc1]+ba[ext=m4a]/b[height<=1080]/b" URL ``` Codec string gotcha: YouTube reports H.264 as `avc1.xxxx` — match with `[vcodec^=avc1]`, not `[vcodec=h264]`. AV1 is `av01`, H.265 is `hev1`/`hvc1`. ## Avoiding the post-download transcode (the whole point) | Destination needs | Ask for at download | Why it works | |---|---|---| | MP4 for web/editors | `-S "vcodec:h264,acodec:m4a" --merge-output-format mp4` | h264+aac are mp4-native; merge is a remux | | Audio for STT | `-x --audio-format opus` | platform audio IS Opus; `-x` copies, no transcode | | MKV archive | `-S "res,fps,hdr:12,vcodec,acodec" --merge-output-format mkv` | mkv holds anything; never forces re-encode | | Anything else | download best-native, then ffmpeg-ops | yt-dlp's `--recode-video` is a blind transcode — no CRF/preset/pix_fmt control | If the needed codec genuinely isn't offered (some platforms are VP9-only at high res), that's a real transcode — do it deliberately with the ffmpeg-ops web-compatible H.264 recipe, not `--recode-video`. ## Survey before arguing ```bash yt-dlp -F URL # table of every offered format (ID, ext, res, codecs, size) yt-dlp -J URL | jq '.formats[] | {format_id, ext, vcodec, acodec, height}' ``` When a selector misbehaves, `-F` output is ground truth — extractors change what they offer (per-client, per-region, A/B tests), so yesterday's format ID list is not evidence. ## Canonical presets Machine-readable versions of these recipes (per-destination args, handoff notes): [../assets/format-presets.json](../assets/format-presets.json). -
output-templates.md 3.4 KB
# Output Templates (`-o`) and Paths (`-P`) Filenames are an API: archive correlation, sort order, and cross-filesystem safety are all decided by the template. Get the convention right once. ## The house convention ```bash -o "%(uploader)s/%(upload_date)s - %(title).100B [%(id)s].%(ext)s" ``` - `[%(id)s]` **always** — titles change, get truncated, and collide; the ID is the only stable join key to archive files and metadata. - `%(upload_date)s` prefix (YYYYMMDD) — lexicographic = chronological sort. - `%(title).100B` — truncate at 100 **bytes**, UTF-8-safe. `%(title).100s` truncates *characters* and can blow filesystem byte limits on CJK/emoji titles. - Never end a directory component with `%(title)s` alone — a title of `.` or emoji-only produces garbage paths. ## Field catalog (the useful subset) | Field | Value | |---|---| | `%(id)s` | platform video ID | | `%(title)s` | video title (sanitized for the local OS by default) | | `%(ext)s` | final extension — **always end the template with this**; yt-dlp picks it post-merge | | `%(uploader)s`, `%(channel)s` | display name / channel name | | `%(channel_id)s` | stable channel ID (display names get renamed) | | `%(upload_date)s` | YYYYMMDD | | `%(duration)s` | seconds | | `%(playlist)s`, `%(playlist_index)s` | playlist name / position (`%(playlist_index)03d` to zero-pad) | | `%(resolution)s`, `%(fps)s`, `%(vcodec)s`, `%(acodec)s` | stream properties | | `%(epoch)s` | download time (unix) — for run-stamping | Numeric fields accept printf formatting (`%(playlist_index)03d`); all fields accept the `.NB` byte-truncation suffix. Missing fields render as `NA` — provide defaults with `%(uploader|unknown)s` pipe syntax. ## Sanitization ```bash --restrict-filenames # ASCII-only, no spaces (shell/CI-safe; ugly) --windows-filenames # Windows-illegal chars stripped even on Linux (NAS/SMB) --trim-filenames 200 # hard cap on total filename length ``` Default sanitization already strips the local OS's illegal characters; `--restrict-filenames` is for files that must survive *any* downstream system (URLs, docker volumes, old CI). Pick per destination, not reflexively. ## Per-type routing Different artifact types can take different templates in one run: ```bash yt-dlp --write-subs --write-thumbnail \ -o "%(title).100B [%(id)s].%(ext)s" \ -o "subtitle:subs/%(id)s.%(ext)s" \ -o "thumbnail:thumbs/%(id)s.%(ext)s" URL ``` Types: `subtitle`, `thumbnail`, `description`, `infojson`, `chapter`, `pl_thumbnail`, `pl_description`, `pl_infojson`. ## Paths (`-P`): destination vs scratch ```bash yt-dlp -P "D:/media" -P "temp:C:/tmp/ytdlp" URL ``` - `-P home:` (or bare `-P`) — final destination. - `-P temp:` — fragments, `.part` files, and pre-merge intermediates. Putting temp on fast local disk while home is a NAS/slow volume avoids double-writing large files over the network. The final file is *moved* (not re-downloaded) on completion. - Type-specific paths compose with type-specific templates: `-P "subtitle:subs"`. ## Sidecar metadata for pipelines ```bash yt-dlp --write-info-json -o "%(id)s.%(ext)s" URL # full metadata as <id>.info.json yt-dlp --load-info-json X.info.json # re-download later without re-extracting ``` `--write-info-json` is the pipeline-friendly pattern: every downstream step (transcription, indexing, dedup) reads structured metadata from the sidecar instead of re-querying the platform. -
playlists-archives.md 4 KB
# Playlists, Channels, and Archive Files Batch acquisition done right: resumable, deduplicated, polite, and cheap to re-run. ## The archive file (`--download-archive`) ```bash yt-dlp --download-archive archive.txt -o "%(title).100B [%(id)s].%(ext)s" PLAYLIST_URL ``` The archive is a plain text file, one `extractor video_id` line per completed download (e.g. `youtube dQw4w9WgXcQ`). On every run, anything already listed is skipped *before* download. Properties worth knowing: - **Append-only and trivially repairable** — delete a line to force a re-download; concatenate archives to merge collections. - **It records IDs, not files** — moving/renaming downloaded files doesn't break it, but *deleting* a file doesn't trigger a re-download either. Keep `%(id)s` in the filename template so files and archive lines stay correlatable. - **Scope it per collection** (one archive per channel/playlist/job), not one global file — global archives make "did job X get video Y?" unanswerable. ## Incremental channel sync (the cron pattern) The naive cron job re-walks the whole channel every run — thousands of metadata requests to discover nothing is new. The right shape: ```bash yt-dlp --download-archive archive.txt --break-on-existing --lazy-playlist \ -S "res:1080,vcodec:h264,acodec:m4a" \ --sleep-requests 1 --sleep-interval 5 --max-sleep-interval 15 \ -o "%(channel)s/%(upload_date)s - %(title).100B [%(id)s].%(ext)s" \ CHANNEL_URL ``` - `--break-on-existing` — stop the entire run at the first already-archived video. - `--lazy-playlist` — process entries as they stream in, instead of fetching the full playlist metadata first. The pair turns "walk 2,000 entries" into "check the 3 newest, stop". **Caveat — only valid when new items appear at the front.** Channel upload feeds are newest-first, so this is safe. A curated playlist that gets items inserted anywhere (or sorted oldest-first) will *miss* additions behind the first archived hit: drop `--break-on-existing` for those and eat the full walk. `--break-on-reject` is the sibling for filter-based stops (e.g. with `--dateafter`); same front-loaded-ordering caveat. ## Selecting subsets ```bash yt-dlp -I 1:10 PLAYLIST_URL # items 1-10 yt-dlp -I -3: PLAYLIST_URL # last three yt-dlp -I ::2 PLAYLIST_URL # every second item yt-dlp --dateafter 20260101 CHANNEL_URL # uploaded on/after a date (YYYYMMDD) yt-dlp --match-filters "duration<600 & !is_live" CHANNEL_URL ``` `-I`/`--playlist-items` takes Python-slice-like `start:stop:step` with negative indexing. `--match-filters` runs against metadata fields — combine with `--break-on-reject` carefully (ordering caveat above). ## Enumerate without downloading ```bash # Fast listing (no per-video page fetches): yt-dlp --flat-playlist --print "%(id)s %(title)s" PLAYLIST_URL # Full playlist metadata as one JSON document: yt-dlp --flat-playlist -J PLAYLIST_URL | jq '.entries | length' ``` `--flat-playlist` skips per-entry extraction — fields like exact duration, formats, and descriptions may be missing or approximate; it's for inventory, not for metadata-accurate pipelines. ## Playlist-aware output templates ```bash -o "%(playlist)s/%(playlist_index)03d - %(title).100B [%(id)s].%(ext)s" ``` `%(playlist_index)s` is the position *in this playlist* (zero-pad it — `03d` — so shells sort correctly). For channel scrapes prefer `%(upload_date)s` prefixes: playlist indexes shift as videos are added/removed; upload dates don't. ## Robustness flags for long batch runs ```bash --retries 10 --fragment-retries 10 # transient network errors --ignore-errors # one private/deleted video doesn't kill the run --no-overwrites # never clobber an existing file --windows-filenames # force Windows-safe names even on Linux (NAS/SMB targets) ``` A failed run is rerunnable for free: the archive skips everything completed, and partially-downloaded `.part` files resume automatically. -
sponsorblock.md 3 KB
# SponsorBlock Integration yt-dlp has native SponsorBlock support: crowd-sourced segment data (sponsor reads, intros, outros, self-promo) fetched at download time and either **marked** as chapters or **removed** from the media. ## Mark vs remove (the decision) | | `--sponsorblock-mark` | `--sponsorblock-remove` | |---|---|---| | Media bytes | Untouched — adds chapter markers only | Cut out — file is modified | | Lossless | Yes | No — re-encodes around cut boundaries | | Reversible | Yes (chapters are metadata) | No | | Player behaviour | Players with auto-skip honor the chapters; others just show them | Segments simply don't exist | | Composability | Works with everything | **Incompatible with `--download-sections`**; complicates archives (file ≠ platform timeline) | **Default to `--sponsorblock-mark`.** It preserves the original media, keeps timestamps aligned with the platform (comments, transcripts, and chapter URLs still match), and the decision to skip stays with the player. Remove only for final-consumption files where the segments must be gone (e.g. media-server libraries watched on dumb clients). ## Usage ```bash # Mark everything SponsorBlock knows about as chapters (lossless): yt-dlp --sponsorblock-mark all URL # Mark only the high-confidence ad categories: yt-dlp --sponsorblock-mark sponsor,selfpromo URL # Remove sponsor reads and self-promo from the file (re-encodes at boundaries): yt-dlp --sponsorblock-remove sponsor,selfpromo URL # Custom chapter title for marked segments: yt-dlp --sponsorblock-mark all --sponsorblock-chapter-title "[SB]: %(category_names)l" URL ``` ## Categories | Category | Content | |---|---| | `sponsor` | Paid sponsor reads | | `selfpromo` | Unpaid self-promotion (merch, Patreon, other videos) | | `interaction` | "Like and subscribe" reminders | | `intro` / `outro` | Intro animations / endcards and credits | | `preview` | Recap/preview of the video itself | | `filler` | Tangents and filler (aggressive — community-tagged loosely) | | `music_offtopic` | Non-music sections in music videos | | `poi_highlight` | Point-of-interest marker (mark-only) | | `chapter` | Community-submitted chapters (mark-only) | | `all` / `default` | Everything / the default mark set | `-remove` accepts the cuttable subset (not `poi_highlight`/`chapter`). Start conservative: `sponsor,selfpromo` has high community accuracy; `filler` is noisy. ## Operational notes - **Data is crowd-sourced** — new uploads may have no segments yet (the flags then do nothing, silently). Niche channels may never be tagged. - The SponsorBlock API is an extra network dependency: `--sponsorblock-api URL` points at a mirror if the default is unreachable; downloads proceed without segment data on API failure. - Remove + `--download-archive` interact philosophically: the archive says "have video X", but the file is a *modified* X. If fidelity matters to the collection, mark instead. - Chapter-mark output composes with `--embed-chapters` (platform chapters and SponsorBlock chapters merge into one track). -
subtitles-metadata.md 3.2 KB
# Subtitles, Transcripts, and Metadata Embedding Subtitle download is also the cheapest transcription source that exists — auto-subs for a 1-hour video cost one HTTP request vs minutes of Whisper compute. ## Downloading subtitles ```bash # Survey what exists (manual and auto-generated, per language): yt-dlp --list-subs URL # Manual subs, all English variants, skip the live-chat pseudo-track, as SRT: yt-dlp --write-subs --sub-langs "en.*,-live_chat" --convert-subs srt --skip-download URL # Auto-generated (ASR) captions — present on most videos even without manual subs: yt-dlp --write-auto-subs --sub-langs en --convert-subs srt --skip-download URL # Both, preferring manual when it exists: yt-dlp --write-subs --write-auto-subs --sub-langs "en.*,-live_chat" --convert-subs srt URL ``` - `--sub-langs` takes comma-separated language regexes with `-` exclusions. `"en.*"` catches `en`, `en-US`, `en-GB`, `en-orig`. `"all,-live_chat"` = everything except the live-chat JSON track (which otherwise downloads as a huge `.json`). - `--convert-subs srt|vtt|ass|lrc` — platforms serve VTT/JSON3; most downstream tooling wants SRT. Conversion is lossless for timing/text (styling is dropped). - `--skip-download` makes it a subs-only run. ## Subs as cheap transcripts (STT shortcut) Before reaching for Whisper, check whether auto-subs are good enough: ```bash yt-dlp --write-auto-subs --sub-langs en --convert-subs srt --skip-download -o "%(id)s" URL ``` Auto-subs quality is "ASR with platform-scale models" — usually fine for search, summarisation, and topic extraction; weak on names, jargon, and punctuation. When word-level timing precision or quality matters, do real STT: acquire audio with `-x --audio-format opus`, then the ffmpeg-ops [stt-whisper](../../ffmpeg-ops/references/stt-whisper.md) pipeline. Note: auto-sub SRT contains rolling-caption duplication (each line appears twice as the window scrolls). Dedupe before feeding to an LLM — naive concatenation roughly doubles token cost. ## Embedding (subs travel with the file) ```bash yt-dlp --embed-subs --sub-langs en URL ``` Embeds as *soft* subtitles (toggleable track — `mov_text` in MP4, SRT/ASS in MKV). Burn-in (hard subs) is a re-encode and belongs to ffmpeg-ops [subtitles.md](../../ffmpeg-ops/references/subtitles.md). ## Metadata, thumbnails, chapters ```bash # The self-describing-file trio: yt-dlp --embed-metadata --embed-thumbnail --embed-chapters URL ``` - `--embed-metadata` — title/uploader/date/description into container tags. - `--embed-thumbnail` — cover art (mp4/m4a/mkv/mp3/opus targets; needs the thumbnail to be convertible — yt-dlp handles webp→jpg via ffmpeg automatically). - `--embed-chapters` — platform chapter markers as container chapters; players and editors (and ffmpeg-ops EDL tooling) can navigate them. SponsorBlock chapter marking composes with this — see [sponsorblock.md](sponsorblock.md). Sidecar alternative when files must stay pristine: ```bash yt-dlp --write-thumbnail --write-description --write-info-json URL ``` `--write-info-json` is the machine-readable everything (formats, chapters, tags, counts) — the right input for indexing pipelines; see [output-templates.md](output-templates.md) for routing sidecars into subdirs.
-
-
scripts
-
check-ytdlp-version.sh 15.9 KB
#!/usr/bin/env bash # Staleness verifier for ytdlp-ops — offline structural + live version/extractor drift. # # --offline (default): structural integrity, NO network and no yt-dlp needed. # Assets parse as JSON, every shipped reference/script/asset is cited from # SKILL.md, and every relative resource link in SKILL.md resolves. # Runs in PR CI; may block. # --live: is the INSTALLED yt-dlp still trustworthy, and do our docs still # match it? Three checks: (1) version age — fetches the latest release tag # from the GitHub API (yt-dlp versions are dates); >60 days behind = drift # (exit 10). (2) flag drift — every core flag the skill documents must still # exist in `yt-dlp --help`; renamed/removed = drift. (3) a metadata-only # smoke extraction (--simulate; nothing downloaded). # # Smoke classification is DRIFT-ON-POSITIVE-EVIDENCE, not drift-by-default: # pass -> extraction worked (exit 0) # blocked -> YouTube refused us (bot-check / sign-in / 429 / # "this video is unavailable") — advisory, exit 7 # skipped-no-jsruntime-> no deno/node to solve the nsig challenge — advisory # fail -> yt-dlp's OWN extractor internals errored # ("unable to extract", "please report this issue", # nsig/signature) — the only smoke DRIFT (exit 10) # fail-unclassified -> anything else — advisory, exit 7, and the raw first # stderr line is printed so the next reader # classifies from real output rather than guessing # # WHY the inversion: this check runs on a GitHub-hosted runner, whose source # IP is an Azure datacenter range that YouTube bot-gates. From there it is # IMPOSSIBLE to distinguish "the extractor broke" from "our IP is blocked", # and a verifier that cries wolf every Monday gets ignored — which is strictly # worse than one that under-reports (SKILL-RESOURCE-PROTOCOL §7: a live check # must never block on environmental noise). The smoke test therefore has real # teeth only from a residential IP; in CI treat a "blocked" result as "not # tested", not as "verified healthy". # # Network/API/yt-dlp unavailable = exit 7 (advisory — the scheduled workflow # warns and moves on; live checks never gate a PR). # # Usage: check-ytdlp-version.sh [--offline | --live] [--no-smoke] [--json] [-q] # Input: none (inspects the skill's own files; --live also yt-dlp + GitHub API). # Test seams: CM_YTDLP_INSTALLED / CM_YTDLP_LATEST (version strings, # e.g. 2026.05.31) bypass yt-dlp/network and imply --no-smoke. # Output: stdout = findings ("DRIFT:"/"STRUCT:" lines), or with --json one # envelope (schema claude-mods.ytdlp-ops.version-check/v1) # Stderr: progress, warnings # Exit: 0 clean, 2 usage, 7 advisory (--live: network/API/yt-dlp unavailable, # or the smoke test could not verify extraction — blocked / no JS # runtime / unclassified failure), 10 drift (>60 days behind latest, # a documented flag vanished, an extractor-internal smoke failure, or # a structural finding) # # Examples: # check-ytdlp-version.sh --offline # check-ytdlp-version.sh --live # check-ytdlp-version.sh --live --json | jq '.data.days_behind' set -uo pipefail EXIT_OK=0; EXIT_USAGE=2; EXIT_UNAVAILABLE=7; EXIT_DRIFT=10 MAX_AGE_DAYS=60 RELEASES_API="https://api.github.com/repos/yt-dlp/yt-dlp/releases/latest" # "Me at the zoo" — the first video ever uploaded to YouTube; the most # deletion-proof target that exists (metadata-only probe, nothing downloaded). SMOKE_URL="https://www.youtube.com/watch?v=jNQXAC9IVRw" MODE="offline"; QUIET=0; JSON=0; SMOKE=1 while [[ $# -gt 0 ]]; do case "$1" in --offline) MODE="offline" ;; --live) MODE="live" ;; --no-smoke) SMOKE=0 ;; --json) JSON=1 ;; -q|--quiet) QUIET=1 ;; -h|--help) awk 'NR>1 && !/^#/{exit} NR>1{sub(/^# ?/,""); print}' "$0"; exit "$EXIT_OK" ;; *) echo "ERROR: unknown argument: $1 (try --help)" >&2; exit "$EXIT_USAGE" ;; esac shift done SKILL_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" SKILL_MD="$SKILL_DIR/SKILL.md" FINDINGS=() INSTALLED=""; LATEST=""; DAYS_BEHIND=""; SMOKE_RESULT="skipped"; JS_RUNTIME="unknown" ADVISORY=0 # Smoke-failure classification. Order matters: BLOCK is tested before DRIFT. # # BLOCK_RE — YouTube refusing US, not yt-dlp failing. Every alternative below was # copied from real GitHub-hosted-runner output (run 32614850020), never guessed. # Deliberately apostrophe-free: the message is "Sign in to confirm you<U+2019>re # not a bot" with a CURLY apostrophe (U+2019), and the old pattern # "confirm you'?re not a bot" only matched the ASCII one — that single codepoint # is what red-failed this job weekly. Matching "not a bot" survives whichever # quote character YouTube ships next. BLOCK_RE="not a bot|sign in to confirm|HTTP Error 429|too many requests|this video is unavailable" # DRIFT_RE — positive evidence the EXTRACTOR broke. These are yt-dlp's own # internal error strings (not YouTube's user-facing prose), so they are stable # to match on, which is why drift is a whitelist and everything else is advisory. DRIFT_RE="unable to extract|failed to extract|please report this issue|nsig|signature extraction" emit() { [[ "$QUIET" -eq 1 ]] || printf '%s\n' "$1" >&2; } finding() { FINDINGS+=("$1"); } # Pick a working python for JSON/date work (Windows Store stub exits non-zero). PYTHON="" for c in python3 python py; do if command -v "$c" >/dev/null 2>&1 && "$c" -c "" >/dev/null 2>&1; then PYTHON="$c"; break; fi done # 2026.5.31 or 2026.05.31.232914 (nightly) -> 2026-05-31; empty on parse failure. norm_date() { local v y m d v="$(printf '%s' "$1" | cut -d. -f1-3)" IFS=. read -r y m d <<<"$v" [[ "${y:-}" =~ ^[0-9]{4}$ && "${m:-}" =~ ^[0-9]{1,2}$ && "${d:-}" =~ ^[0-9]{1,2}$ ]] || return 1 printf '%04d-%02d-%02d' "$((10#$y))" "$((10#$m))" "$((10#$d))" } # days from $1 (older, ISO) to $2 (newer, ISO); empty if no date backend exists. days_between() { if date -d "2020-01-01" +%s >/dev/null 2>&1; then echo $(( ( $(date -d "$2" +%s) - $(date -d "$1" +%s) ) / 86400 )) elif [[ -n "$PYTHON" ]]; then "$PYTHON" -c "import sys,datetime as dt; a,b=sys.argv[1:3]; print((dt.date.fromisoformat(b)-dt.date.fromisoformat(a)).days)" "$1" "$2" 2>/dev/null fi } # ── offline: structural ────────────────────────────────────────────────────── offline_checks() { emit "== check-ytdlp-version --offline (structural)" [[ -f "$SKILL_MD" ]] || { finding "STRUCT: SKILL.md missing"; return; } # 1. assets parse as JSON for a in "$SKILL_DIR"/assets/*.json; do [[ -e "$a" ]] || continue if [[ -n "$PYTHON" ]]; then "$PYTHON" -c "import json,sys; json.load(open(sys.argv[1], encoding='utf-8'))" "$a" \ >/dev/null 2>&1 || finding "STRUCT: asset not valid JSON: $(basename "$a")" elif command -v jq >/dev/null 2>&1; then jq empty "$a" >/dev/null 2>&1 || finding "STRUCT: asset not valid JSON: $(basename "$a")" fi done # 2. every shipped resource is cited from SKILL.md (dead-weight check) for d in references scripts assets; do for f in "$SKILL_DIR/$d"/*; do [[ -f "$f" ]] || continue base="$(basename "$f")" [[ "$base" == ".gitkeep" ]] && continue grep -q "$base" "$SKILL_MD" \ || finding "STRUCT: $d/$base exists on disk but is never cited from SKILL.md" done done # 3. every relative resource link in SKILL.md resolves while IFS= read -r path; do [[ -e "$SKILL_DIR/$path" ]] \ || finding "STRUCT: SKILL.md links to missing file: $path" done < <(grep -oE '\]\((references|assets|scripts|tests)/[^)#]+\)' "$SKILL_MD" \ | sed -E 's/^\]\(//; s/\)$//' | sort -u) } # ── live: installed yt-dlp vs latest release + smoke extraction ───────────── unavailable() { # $1 = human reason echo "$1" >&2 if [[ "$JSON" -eq 1 ]]; then printf '{ "error": { "code": "UNAVAILABLE", "message": "%s", "details": {} } }\n' "$1" fi exit "$EXIT_UNAVAILABLE" } live_checks() { emit "== check-ytdlp-version --live (version age + extractor smoke)" local seamed=0 [[ -n "${CM_YTDLP_INSTALLED:-}" && -n "${CM_YTDLP_LATEST:-}" ]] && { seamed=1; SMOKE=0; } # Test seam: CM_YTDLP_SMOKE_ERR injects a smoke failure (empty = success) # without yt-dlp/network, so the classifier below is exercised in CI. [[ -n "${CM_YTDLP_SMOKE_ERR+x}" ]] && SMOKE=1 # installed version if [[ "$seamed" -eq 1 ]]; then INSTALLED="$CM_YTDLP_INSTALLED" elif command -v yt-dlp >/dev/null 2>&1; then INSTALLED="$(yt-dlp --version 2>/dev/null | head -1)" fi [[ -n "$INSTALLED" ]] \ || unavailable "yt-dlp not on PATH — install: uv tool install yt-dlp (advisory, not a failure)" # latest release tag from the GitHub API if [[ "$seamed" -eq 1 ]]; then LATEST="$CM_YTDLP_LATEST" else command -v curl >/dev/null 2>&1 \ || unavailable "curl not available — cannot query the GitHub releases API" local body body="$(curl -fsSL --max-time 20 -H "Accept: application/vnd.github+json" \ ${GITHUB_TOKEN:+-H "Authorization: Bearer $GITHUB_TOKEN"} \ "$RELEASES_API" 2>/dev/null)" \ || unavailable "GitHub releases API unreachable/rate-limited (advisory, not a failure)" if command -v jq >/dev/null 2>&1; then LATEST="$(jq -r '.tag_name // empty' <<<"$body" 2>/dev/null)" else LATEST="$(grep -oE '"tag_name"[[:space:]]*:[[:space:]]*"[^"]+"' <<<"$body" \ | head -1 | sed -E 's/.*"([^"]+)"$/\1/')" fi [[ -n "$LATEST" ]] || unavailable "could not parse tag_name from the GitHub API response" fi emit " installed: $INSTALLED latest: $LATEST" # version age (yt-dlp versions ARE dates) local inst_d latest_d if inst_d="$(norm_date "$INSTALLED")" && latest_d="$(norm_date "$LATEST")"; then DAYS_BEHIND="$(days_between "$inst_d" "$latest_d")" if [[ -z "$DAYS_BEHIND" ]]; then emit " warn: no GNU date or python available — age check skipped" elif [[ "$DAYS_BEHIND" -gt "$MAX_AGE_DAYS" ]]; then finding "DRIFT: installed yt-dlp $INSTALLED is $DAYS_BEHIND days behind latest $LATEST (>$MAX_AGE_DAYS) — extractors likely broken; update" fi else emit " warn: unparseable version string(s) — age check skipped" fi # flag drift: every core flag the skill documents must still exist in this # yt-dlp's --help — a rename/removal upstream means our docs rotted. # (Skipped in seamed mode: no real yt-dlp to interrogate.) if [[ "$seamed" -eq 0 ]]; then local help_text help_text="$(yt-dlp --help 2>/dev/null)" local core_flags=(--download-sections --force-keyframes-at-cuts --download-archive --break-on-existing --lazy-playlist --cookies-from-browser --sponsorblock-mark --sponsorblock-remove --remux-video --recode-video --write-subs --write-auto-subs --sub-langs --convert-subs --embed-subs --embed-metadata --embed-thumbnail --embed-chapters --restrict-filenames --concurrent-fragments --limit-rate --sleep-requests --sleep-interval --merge-output-format --extract-audio --audio-format --flat-playlist --playlist-items --match-filters --simulate --impersonate --live-from-start --wait-for-video --print --paths) local fl for fl in "${core_flags[@]}"; do grep -q -- "$fl" <<<"$help_text" \ || finding "DRIFT: documented flag '$fl' unknown to installed yt-dlp (renamed/removed upstream?)" done fi # metadata-only smoke extraction (only when the API was reachable, so a # failure here is the extractor, not the network) if [[ "$SMOKE" -eq 1 ]]; then # NOTE: no --no-warnings here — the JS-runtime detection below greps the # "No supported JavaScript runtime" WARNING from captured stderr. local smoke_err smoke_rc if [[ -n "${CM_YTDLP_SMOKE_ERR+x}" ]]; then smoke_err="$CM_YTDLP_SMOKE_ERR" [[ -z "$smoke_err" ]] && smoke_rc=0 || smoke_rc=1 else local smoke_cmd=(yt-dlp --simulate --no-playlist --socket-timeout 15 "$SMOKE_URL") command -v timeout >/dev/null 2>&1 && smoke_cmd=(timeout 90 "${smoke_cmd[@]}") smoke_err="$("${smoke_cmd[@]}" 2>&1 >/dev/null)"; smoke_rc=$? fi if [[ "$smoke_rc" -eq 0 ]]; then SMOKE_RESULT="pass" JS_RUNTIME="present" elif grep -qiE "$BLOCK_RE" <<<"$smoke_err"; then # IP-reputation challenge (datacenter IPs hit this), NOT extractor drift — # treating it as drift would make the scheduled job flaky-red (§7). SMOKE_RESULT="blocked" ADVISORY=1 emit " warn: smoke extraction blocked by YouTube (bot-check/sign-in/429) — not drift; skipped" elif grep -q "No supported JavaScript runtime" <<<"$smoke_err"; then # YouTube now requires a JS runtime (deno/node) to extract formats. On a # runner without one the extraction fails for an ENVIRONMENT reason, not # extractor drift — treat it like the IP challenge: advisory skip, never # exit-10 (§7). Install deno on the runner (or pass --js-runtimes node) to # give the smoke test teeth. This was the false-positive that red-failed # the weekly freshness job. SMOKE_RESULT="skipped-no-jsruntime" JS_RUNTIME="missing" ADVISORY=1 emit " warn: no JS runtime (deno/node) — smoke extraction skipped, not drift; install deno or pass --js-runtimes node" elif grep -qiE "$DRIFT_RE" <<<"$smoke_err"; then # POSITIVE evidence of extractor breakage: these are yt-dlp's OWN internal # failure strings, not YouTube's prose, so they are stable to match on. SMOKE_RESULT="fail" finding "DRIFT: smoke extraction failed ($SMOKE_URL) with an extractor-internal error — extractor broken; update yt-dlp" else # Unrecognised failure. We CANNOT tell "extractor broke" from "our IP is # blocked" here, and a weekly false red gets the whole job ignored — so # this is advisory (§7), never exit 10. The first stderr line is echoed so # the next reader classifies from real output instead of guessing. SMOKE_RESULT="fail-unclassified" ADVISORY=1 emit " warn: smoke extraction failed with an unrecognised error — advisory, not drift." emit " raw: $(printf '%s' "$smoke_err" | grep -m1 -i "^ERROR" || printf '%s' "$smoke_err" | head -1)" fi # Smoke PASSED but yt-dlp still warned about a missing JS runtime: formats # were thinned. Record the environment state without treating it as drift. if [[ "$SMOKE_RESULT" == "pass" ]] && grep -q "No supported JavaScript runtime" <<<"$smoke_err"; then JS_RUNTIME="missing" emit " warn: no JS runtime (deno/node) — YouTube formats reduced; install deno or use --js-runtimes node" fi fi } case "$MODE" in offline) offline_checks ;; live) live_checks ;; esac if [[ "$JSON" -eq 1 ]]; then flist="" for f in ${FINDINGS[@]+"${FINDINGS[@]}"}; do esc="${f//\\/\\\\}"; esc="${esc//\"/\\\"}" flist="${flist:+$flist, }\"$esc\"" done printf '{ "data": { "mode": "%s", "installed": %s, "latest": %s, "days_behind": %s, "smoke": "%s", "js_runtime": "%s", "findings": [%s] }, "meta": { "count": %d, "schema": "claude-mods.ytdlp-ops.version-check/v1" } }\n' \ "$MODE" \ "$([[ -n "$INSTALLED" ]] && printf '"%s"' "$INSTALLED" || printf 'null')" \ "$([[ -n "$LATEST" ]] && printf '"%s"' "$LATEST" || printf 'null')" \ "${DAYS_BEHIND:-null}" \ "$SMOKE_RESULT" "$JS_RUNTIME" "$flist" "${#FINDINGS[@]}" else for f in ${FINDINGS[@]+"${FINDINGS[@]}"}; do printf '%s\n' "$f"; done fi if [[ "${#FINDINGS[@]}" -eq 0 ]]; then if [[ "$ADVISORY" -eq 1 ]]; then emit "check-ytdlp-version ($MODE): clean, but the smoke test could not verify extraction (advisory)" exit "$EXIT_UNAVAILABLE" fi emit "check-ytdlp-version ($MODE): clean" exit "$EXIT_OK" fi emit "check-ytdlp-version ($MODE): ${#FINDINGS[@]} finding(s)" exit "$EXIT_DRIFT"
-
-
tests
-
run.sh 10.6 KB
#!/usr/bin/env bash # Self-test for ytdlp-ops — fully offline: no network, no yt-dlp needed. # # Structural assertions (--help contract, bash -n, exit codes, offline verifier, # asset JSON) plus the verifier's 60-day age logic exercised through its # CM_YTDLP_INSTALLED / CM_YTDLP_LATEST test seams (which bypass yt-dlp and the # GitHub API and disable the smoke extraction). Real --live runs belong to the # scheduled freshness workflow only — a network blip must never fail a PR. # # Usage: bash tests/run.sh # Exit: 0 all pass, 1 one or more failures set -uo pipefail HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" SKILL="$(dirname "$HERE")" V="$SKILL/scripts/check-ytdlp-version.sh" # Pick a python that actually executes (Windows Store python3 stub exits non-zero). PYTHON="" for c in python python3 py; do if command -v "$c" >/dev/null 2>&1 && "$c" -c "" >/dev/null 2>&1; then PYTHON="$c"; break; fi done SB="$(mktemp -d)"; trap 'rm -rf "$SB"' EXIT PASS=0; FAIL=0 ok() { PASS=$((PASS+1)); printf ' PASS %s\n' "$1"; } no() { FAIL=$((FAIL+1)); printf ' FAIL %s\n' "$1"; } expect_exit() { [[ "$2" == "$3" ]] && ok "$1 (exit $3)" || no "$1 (want $2 got $3)"; } expect_has() { case "$3" in *"$2"*) ok "$1";; *) no "$1 (missing '$2')";; esac; } echo "=== ytdlp-ops self-test ===" # ── contract ────────────────────────────────────────────────────────────────── echo "-- contract --" bash -n "$V" 2>/dev/null && ok "bash -n check-ytdlp-version.sh" || no "bash -n check-ytdlp-version.sh" bash "$V" --help >/dev/null 2>&1; expect_exit "--help exits 0" 0 $? out="$(bash "$V" --help 2>/dev/null)" expect_has "--help has Examples" "xamples" "$out" expect_has "--help documents exit 7" "7" "$out" expect_has "--help documents exit 10" "10" "$out" bash "$V" --bogus >/dev/null 2>&1; expect_exit "unknown flag -> 2" 2 $? # ── offline structural mode ────────────────────────────────────────────────── echo "-- offline structural --" bash "$V" --offline >/dev/null 2>&1; expect_exit "--offline clean on shipped skill" 0 $? out="$(bash "$V" --offline --json 2>/dev/null)" expect_has "--offline --json envelope" '"schema": "claude-mods.ytdlp-ops.version-check/v1"' "$out" expect_has "--offline --json zero findings" '"count": 0' "$out" # an uncited resource must be flagged (run the verifier from a doctored copy) cp -r "$SKILL" "$SB/copy" printf '# orphan\n' > "$SB/copy/references/orphan.md" bash "$SB/copy/scripts/check-ytdlp-version.sh" --offline >"$SB/orphan.out" 2>/dev/null expect_exit "--offline flags uncited resource -> 10" 10 $? expect_has "finding names the orphan" "orphan.md" "$(cat "$SB/orphan.out")" # a ghost link must be flagged cp -r "$SKILL" "$SB/ghost" printf '\nsee [gone](references/does-not-exist.md)\n' >> "$SB/ghost/SKILL.md" bash "$SB/ghost/scripts/check-ytdlp-version.sh" --offline >"$SB/ghost.out" 2>/dev/null expect_exit "--offline flags ghost link -> 10" 10 $? expect_has "finding names the missing file" "does-not-exist.md" "$(cat "$SB/ghost.out")" # ── assets ─────────────────────────────────────────────────────────────────── echo "-- assets --" for a in "$SKILL"/assets/*.json; do if [[ -n "$PYTHON" ]]; then "$PYTHON" -c "import json,sys; json.load(open(sys.argv[1], encoding='utf-8'))" "$a" \ >/dev/null 2>&1 && ok "asset parses: $(basename "$a")" || no "asset parses: $(basename "$a")" elif command -v jq >/dev/null 2>&1; then jq empty "$a" >/dev/null 2>&1 && ok "asset parses: $(basename "$a")" || no "asset parses: $(basename "$a")" else echo " SKIP asset JSON parse (no python or jq)" fi done grep -q '"schema": "claude-mods.ytdlp-ops.format-presets/v1"' "$SKILL/assets/format-presets.json" \ && ok "format-presets schema id" || no "format-presets schema id" # ── live mode via test seams (no network, no yt-dlp) ───────────────────────── echo "-- live age logic (seamed) --" CM_YTDLP_INSTALLED=2026.01.01 CM_YTDLP_LATEST=2026.06.01 bash "$V" --live >/dev/null 2>&1 expect_exit "151 days behind -> 10" 10 $? CM_YTDLP_INSTALLED=2026.06.01 CM_YTDLP_LATEST=2026.06.01 bash "$V" --live >/dev/null 2>&1 expect_exit "in sync -> 0" 0 $? CM_YTDLP_INSTALLED=2026.05.01 CM_YTDLP_LATEST=2026.06.01 bash "$V" --live >/dev/null 2>&1 expect_exit "31 days behind (<=60) -> 0" 0 $? CM_YTDLP_INSTALLED=2026.06.01.232900 CM_YTDLP_LATEST=2026.06.01 bash "$V" --live >/dev/null 2>&1 expect_exit "nightly 4-part version parses -> 0" 0 $? out="$(CM_YTDLP_INSTALLED=2026.01.01 CM_YTDLP_LATEST=2026.06.01 bash "$V" --live --json 2>/dev/null)" expect_has "seamed --json envelope" '"schema": "claude-mods.ytdlp-ops.version-check/v1"' "$out" expect_has "seamed --json days_behind" '"days_behind": 151' "$out" expect_has "seamed --json smoke skipped" '"smoke": "skipped"' "$out" expect_has "seamed --json js_runtime unknown" '"js_runtime": "unknown"' "$out" expect_has "seamed --json DRIFT finding" "DRIFT" "$out" if [[ -n "$PYTHON" ]]; then printf '%s' "$out" | "$PYTHON" -c "import json,sys; json.load(sys.stdin)" >/dev/null 2>&1 \ && ok "seamed --json is valid JSON" || no "seamed --json is valid JSON" fi # stdout/stderr separation: data on stdout only err="$(CM_YTDLP_INSTALLED=2026.06.01 CM_YTDLP_LATEST=2026.06.01 bash "$V" --live --json 2>/dev/null >"$SB/stdout.txt"; cat "$SB/stdout.txt")" case "$err" in "{ \"data\""*) ok "stdout carries only the JSON envelope";; *) no "stdout carries only the JSON envelope";; esac # ── smoke classification (seamed; fixes the freshness false-positive) ──────── echo "-- smoke classification (seamed) --" # (a) no JS runtime on the runner = ENVIRONMENT gap, advisory skip — NEVER # exit 10. This is the exact false-positive that red-failed the weekly job. out="$(CM_YTDLP_INSTALLED=2026.06.01 CM_YTDLP_LATEST=2026.06.01 \ CM_YTDLP_SMOKE_ERR='ERROR: [youtube] jNQXAC9IVRw: No supported JavaScript runtime found, please install deno' \ bash "$V" --live --json 2>/dev/null)"; rc=$? expect_exit "no-JS-runtime smoke -> 7 (advisory, not drift)" 7 "$rc" expect_has "smoke marked skipped-no-jsruntime" '"smoke": "skipped-no-jsruntime"' "$out" expect_has "js_runtime reported missing" '"js_runtime": "missing"' "$out" case "$out" in *DRIFT*) no "no-JS-runtime must not record DRIFT";; *) ok "no-JS-runtime records no DRIFT";; esac # (b) a genuine extractor break (JS runtime present) IS drift -> 10 CM_YTDLP_INSTALLED=2026.06.01 CM_YTDLP_LATEST=2026.06.01 \ CM_YTDLP_SMOKE_ERR='ERROR: unable to extract player response; please report this issue' \ bash "$V" --live >/dev/null 2>&1 expect_exit "real extractor break -> 10 (drift)" 10 $? # (c) IP bot-challenge / 429 = advisory skip, not drift CM_YTDLP_INSTALLED=2026.06.01 CM_YTDLP_LATEST=2026.06.01 \ CM_YTDLP_SMOKE_ERR='ERROR: HTTP Error 429: Too Many Requests' \ bash "$V" --live >/dev/null 2>&1 expect_exit "bot-challenge/429 -> 7 (advisory)" 7 $? # (c2) THE REGRESSION GUARD. Verbatim stderr from a GitHub-hosted runner (run # 32614850020) - note the CURLY apostrophe U+2019 in "you’re". The old # pattern "confirm you'?re not a bot" matched only the ASCII apostrophe, so # this exact text fell through to DRIFT and red-failed the weekly freshness # job. Do not "tidy" this string - the codepoint IS the test. out="$(env CM_YTDLP_INSTALLED=2026.06.01 CM_YTDLP_LATEST=2026.06.01 CM_YTDLP_SMOKE_ERR="ERROR: [youtube] jNQXAC9IVRw: Sign in to confirm you’re not a bot. Use --cookies-from-browser or --cookies for the authentication." bash "$V" --live --json 2>/dev/null)"; rc=$? expect_exit "curly-apostrophe bot-check -> 7 (advisory)" 7 "$rc" expect_has "curly-apostrophe bot-check marked blocked" '"smoke": "blocked"' "$out" case "$out" in *DRIFT*) no "curly-apostrophe bot-check must not record DRIFT";; *) ok "curly-apostrophe bot-check records no DRIFT";; esac # (c3) the ASCII-apostrophe form of the same message must still classify blocked out="$(env CM_YTDLP_INSTALLED=2026.06.01 CM_YTDLP_LATEST=2026.06.01 "CM_YTDLP_SMOKE_ERR=ERROR: [youtube] jNQXAC9IVRw: Sign in to confirm you're not a bot." bash "$V" --live --json 2>/dev/null)"; rc=$? expect_exit "ASCII-apostrophe bot-check -> 7 (advisory)" 7 "$rc" expect_has "ASCII-apostrophe bot-check marked blocked" '"smoke": "blocked"' "$out" # (c4) "This video is unavailable" - also observed from the runner's datacenter # IP (the second probe target in that same debug run); not extractor drift. out="$(env CM_YTDLP_INSTALLED=2026.06.01 CM_YTDLP_LATEST=2026.06.01 "CM_YTDLP_SMOKE_ERR=ERROR: [youtube] BaW_jenozKc: This video is unavailable" bash "$V" --live --json 2>/dev/null)"; rc=$? expect_exit "video-unavailable -> 7 (advisory)" 7 "$rc" expect_has "video-unavailable marked blocked" '"smoke": "blocked"' "$out" # (c5) an UNRECOGNISED failure degrades to advisory, never DRIFT. This is the # inversion that stops the next reworded YouTube message re-opening this # bug: drift needs positive evidence; an unreadable failure is not evidence. out="$(env CM_YTDLP_INSTALLED=2026.06.01 CM_YTDLP_LATEST=2026.06.01 "CM_YTDLP_SMOKE_ERR=ERROR: [youtube] jNQXAC9IVRw: Some brand new message nobody has seen before" bash "$V" --live --json 2>/dev/null)"; rc=$? expect_exit "unclassified smoke failure -> 7 (advisory, not drift)" 7 "$rc" expect_has "unclassified marked fail-unclassified" '"smoke": "fail-unclassified"' "$out" case "$out" in *DRIFT*) no "unclassified failure must not record DRIFT";; *) ok "unclassified failure records no DRIFT";; esac # (d) smoke success (empty seam) -> pass, exit 0 out="$(CM_YTDLP_INSTALLED=2026.06.01 CM_YTDLP_LATEST=2026.06.01 CM_YTDLP_SMOKE_ERR='' \ bash "$V" --live --json 2>/dev/null)"; rc=$? expect_exit "smoke success -> 0" 0 "$rc" expect_has "smoke marked pass" '"smoke": "pass"' "$out" # ── SKILL.md sanity ────────────────────────────────────────────────────────── echo "-- SKILL.md --" grep -q '^name: ytdlp-ops$' "$SKILL/SKILL.md" && ok "frontmatter name" || no "frontmatter name" grep -q 'related-skills: ffmpeg-ops' "$SKILL/SKILL.md" && ok "ffmpeg-ops cross-link" || no "ffmpeg-ops cross-link" grep -q 'check-ytdlp-version.sh' "$SKILL/SKILL.md" && ok "verifier cited from SKILL.md" || no "verifier cited from SKILL.md" echo "" echo "=== $PASS passed, $FAIL failed ===" [[ "$FAIL" -eq 0 ]] || exit 1 exit 0
-
-
SKILL.md 16.8 KB
--- name: ytdlp-ops description: "yt-dlp media acquisition layer feeding ffmpeg-ops: format selection avoiding post-download transcodes, clip-at-download, cookies/auth, channel archive sync, SponsorBlock, subtitles, failure triage (403s, nsig). Triggers on: yt-dlp, download video/playlist/channel, youtube to mp3." license: MIT compatibility: "yt-dlp 2025.x+ (releases near-monthly; run the verifier FIRST when anything fails). ffmpeg on PATH required for merge/remux/extract-audio. A JS runtime (deno auto-enabled; node via --js-runtimes node) required for full YouTube format extraction. Scripts: bash." allowed-tools: "Read Write Edit Bash Glob Grep" metadata: author: claude-mods related-skills: ffmpeg-ops, debug-ops --- # yt-dlp Operations Operational expertise for yt-dlp as the **acquisition layer**: get the right bytes onto disk in the right codec, politely, resumably — then hand off. Anything that re-encodes, cuts precisely, grades, or packages after download is [ffmpeg-ops](../ffmpeg-ops/SKILL.md) territory; AI-driven editing of what you acquired (transcript → EDL → final cut) is `cutcraft` — a separate tool, not shipped by this repo. The full chain is acquire → process → edit. ## Doctrine: version first, formats second **yt-dlp vs the platforms is an arms race.** Releases land near-monthly and extractors break between them — the majority of "yt-dlp is broken" reports are a stale binary. Before debugging *anything*, check staleness: ```bash bash skills/ytdlp-ops/scripts/check-ytdlp-version.sh --live # vs latest GitHub release bash skills/ytdlp-ops/scripts/check-ytdlp-version.sh --live --json | jq '.data.days_behind' ``` Exit `10` = installed build is >60 days behind latest, a documented flag vanished from `yt-dlp --help`, or the smoke extraction hit an extractor-internal error → update before any other triage. Exit `7` is advisory: the check ran but could not verify extraction (YouTube bot-gated our IP, no JS runtime, or an unrecognised failure) — common from datacenter IPs such as CI runners: ```bash uv tool upgrade yt-dlp # pip/uv-managed install (preferred) yt-dlp -U # standalone binary self-update only ``` **Second rule: pick codecs at download time.** The default "best" on YouTube is VP9/AV1 + Opus in WebM/MKV. If the destination needs H.264 MP4, stating that in `-S` costs nothing — discovering it after download costs a full transcode. ## Cookbook ### Format selection (`-S` over `-f`) ```bash # Declarative sort (-S) — PREFER this. States preferences in priority order and # always degrades gracefully to the nearest available. h264 + m4a merges # natively into mp4: zero post-download transcode. yt-dlp -S "res:1080,vcodec:h264,acodec:m4a" --merge-output-format mp4 URL # Hard filter (-f) — exact control, but FAILS ("Requested format is not # available") when nothing matches. Use only for genuine hard requirements, # always with a / fallback chain: yt-dlp -f "bv*[height<=1080][vcodec^=avc1]+ba[ext=m4a]/b[height<=1080]/b" URL # Survey what the extractor actually offers before arguing with selectors: yt-dlp -F URL # Smallest acceptable file (bandwidth/storage constrained; + prefix = ascending): yt-dlp -S "res:480,+size,+br" URL # Best quality regardless of codec (archival source for later ffmpeg-ops work): yt-dlp -S "res,fps,hdr:12,vcodec,acodec" --merge-output-format mkv URL ``` Sort-field reference, filter grammar, per-destination presets: [references/format-selection.md](references/format-selection.md) + [assets/format-presets.json](assets/format-presets.json). ### Clip at download (`--download-sections`) ```bash # Download ONLY 10:00-12:30 — ranged requests, not a full download + trim: yt-dlp --download-sections "*10:00-12:30" -S "res:1080,vcodec:h264" URL # Frame-accurate cut points (re-encodes around the cuts only): yt-dlp --download-sections "*10:00-12:30" --force-keyframes-at-cuts URL # Last 5 minutes / by chapter-title regex / multiple sections: yt-dlp --download-sections "*-5:00-inf" URL yt-dlp --download-sections "Intro" --download-sections "Outro" URL ``` Same physics as ffmpeg copy-cuts: without `--force-keyframes-at-cuts` the section boundaries **snap to keyframes** (can be seconds off). Need many precise cuts from one source? Download once, then use the ffmpeg-ops EDL workflow. ### Audio-only extraction (STT pipelines) ```bash # THE STT acquisition command. YouTube's best audio IS Opus — asking for opus # means -x COPIES the stream out (no transcode, no quality loss): yt-dlp -x --audio-format opus -o "%(id)s.%(ext)s" URL # Zero-processing alternative — native container, no ffmpeg step at all: yt-dlp -f "ba" -o "%(id)s.%(ext)s" URL # Whole channel's audio for a transcription pipeline (archive = resumable): yt-dlp -x --audio-format opus --download-archive stt-archive.txt \ -o "%(channel)s/%(id)s.%(ext)s" CHANNEL_URL ``` Do NOT `--audio-format mp3` for STT — that's a lossy→lossy transcode that helps nothing. Whisper-prep (16 kHz mono PCM) is the next stage: ffmpeg-ops [stt-whisper](../ffmpeg-ops/references/stt-whisper.md). ### Playlists, channels, incremental sync ```bash # Playlist with ID-correlated filenames + archive file (resumable, dedup-safe): yt-dlp --download-archive archive.txt \ -o "%(playlist)s/%(playlist_index)03d - %(title).100B [%(id)s].%(ext)s" PLAYLIST_URL # Incremental channel sync (cron-friendly): stop at the first already-archived # video instead of re-walking the entire channel every run: yt-dlp --download-archive archive.txt --break-on-existing --lazy-playlist \ -S "res:1080,vcodec:h264,acodec:m4a" CHANNEL_URL # Subset selection / list without downloading: yt-dlp -I 1:10 PLAYLIST_URL yt-dlp --flat-playlist --print "%(id)s %(title)s" PLAYLIST_URL # DRY-RUN any batch before committing to it — preview every output filename # (--print implies --simulate; nothing downloads): yt-dlp --print filename -o "%(playlist_index)03d - %(title).100B [%(id)s].%(ext)s" PLAYLIST_URL ``` Archive format, sync-job patterns, when `--break-on-existing` misfires (non-chronological playlists): [references/playlists-archives.md](references/playlists-archives.md). ### Livestreams and premieres ```bash # Capture a livestream from its BEGINNING, not from "now" (YouTube keeps a # rolling live buffer; without this you get the moment you pressed enter): yt-dlp --live-from-start URL # Scheduled premiere/stream: poll (1-10 min between retries) and start when live: yt-dlp --wait-for-video 60-600 URL ``` Live capture caveats: a crashed live download is **not resumable** like a VOD (fragments expire) — write to fast local disk (`-P temp:`), not a network share. For archival quality, prefer re-downloading the VOD after the stream ends; the live manifest often caps below the post-processed VOD. ### Subtitles ```bash # Manual subs, English variants, skip live-chat pseudo-subs, as SRT: yt-dlp --write-subs --sub-langs "en.*,-live_chat" --convert-subs srt --skip-download URL # Auto-generated (ASR) captions — exist for most videos when manual subs don't: yt-dlp --write-auto-subs --sub-langs en --convert-subs srt --skip-download URL # Embed into the media file instead of a sidecar: yt-dlp --embed-subs --sub-langs en URL ``` Sub formats, language matching, transcript-only workflows (subs as cheap STT): [references/subtitles-metadata.md](references/subtitles-metadata.md). ### SponsorBlock ```bash # Mark segments as chapters — LOSSLESS and reversible. Prefer this: yt-dlp --sponsorblock-mark all URL # Cut segments out of the media — modifies the file, re-encodes at boundaries: yt-dlp --sponsorblock-remove sponsor,selfpromo URL ``` Category list, mark-vs-remove trade-offs, interaction with `--download-sections`: [references/sponsorblock.md](references/sponsorblock.md). ### Cookies and auth ```bash # Pull cookies from a browser profile (private/members/age-gated content): yt-dlp --cookies-from-browser firefox URL # Chrome 127+ on Windows uses app-bound cookie encryption — extraction usually # FAILS. Use Firefox, or export a Netscape cookies.txt and pass it directly: yt-dlp --cookies cookies.txt URL ``` **Account-ban warning:** authenticated bulk downloading is the fastest way to get an account flagged. Use a throwaway account, always pair cookies with the politeness flags below. Details + browser matrix: [references/auth-cookies.md](references/auth-cookies.md). ### Rate limiting and politeness ```bash # The polite-bulk baseline — cap bandwidth, space out requests, retry patiently: yt-dlp --limit-rate 4M --sleep-requests 1 \ --sleep-interval 5 --max-sleep-interval 15 \ --retries 10 --fragment-retries 10 URL # Speed (single video, host not throttling you): parallel fragment download: yt-dlp --concurrent-fragments 4 URL ``` Politeness is self-interest: 429s and IP flags cost more time than sleeps do. ### Remux vs recode ```bash # Remux: container change only — lossless, near-instant. yt-dlp's job: yt-dlp -S "vcodec:h264,acodec:m4a" --remux-video mp4 URL # Recode: a FULL TRANSCODE. Almost never yt-dlp's job — you give up ffmpeg-ops' # CRF/preset/pix_fmt control for a blind default encode. If codecs must change: yt-dlp -S "res,vcodec,acodec" URL # 1. acquire best-native # 2. then transcode with the ffmpeg-ops web-compatible H.264 recipe. ``` Rule: `--remux-video` whenever the codecs already fit the target container; `--recode-video` only for throwaway one-offs where quality control doesn't matter. ### Output templates ```bash # ID-in-brackets convention — survives renames, correlates with archive files: yt-dlp -o "%(uploader)s/%(upload_date)s - %(title).100B [%(id)s].%(ext)s" URL # Cross-filesystem safety (strips spaces/unicode to ASCII-safe names): yt-dlp --restrict-filenames -o "%(title)s [%(id)s].%(ext)s" URL # Split destination and scratch space (-P): fragments go to temp, final to home: yt-dlp -P "D:/media" -P "temp:C:/tmp/ytdlp" URL ``` `%(title).100B` truncates at 100 **bytes** (UTF-8 safe — CJK titles break char-based truncation). Full field catalog and per-type templates: [references/output-templates.md](references/output-templates.md). ### Metadata embedding ```bash # Self-describing files — metadata, thumbnail and chapters travel with the media: yt-dlp --embed-metadata --embed-thumbnail --embed-chapters URL ``` ## Beyond YouTube yt-dlp ships ~1,800 extractors (`yt-dlp --list-extractors`); everything in this skill except the YouTube-specific parts (nsig, player clients) applies unchanged to Twitch, Vimeo, SoundCloud, TikTok, and the rest. `yt-dlp -v URL` names the extractor in use. For sites with no dedicated extractor, the generic extractor sniffs direct media/HLS URLs out of the page. When a non-YouTube site returns 403 to yt-dlp but plays fine in a browser, it's usually TLS-fingerprint blocking — `--impersonate` fixes it (see [failure-triage](references/failure-triage.md)). ## Footguns | Footgun | The trap | The rule | |---|---|---| | Default format selection | YouTube "best" = VP9/AV1+Opus in WebM/MKV; downstream tooling expecting MP4 forces a transcode you could have avoided | State codecs at download: `-S "vcodec:h264,acodec:m4a" --merge-output-format mp4` | | `-f best` | Selects best *single pre-merged file* — caps at ~720p on YouTube; modern high-res is always video+audio merged | Drop the `-f` entirely or use `-S`; `b` only as the tail of a `/` fallback chain | | `-f` hard filters | "Requested format is not available" the moment an extractor stops offering that exact combo | Prefer `-S` (degrades gracefully); always end `-f` chains with `/b` | | `--recode-video` casually | Full blind transcode — no CRF/preset/pix_fmt control, big quality/time cost | `--remux-video` when codecs fit; real transcodes via ffmpeg-ops | | `--download-sections` w/o `--force-keyframes-at-cuts` | Clip boundaries snap to keyframes — seconds of slop | Add the flag when cuts must be exact (re-encodes at cuts only) | | Channel sync w/o `--break-on-existing` | Every cron run re-walks the entire channel (thousands of metadata requests) | `--download-archive` + `--break-on-existing --lazy-playlist` | | No `%(id)s` in filename | Title changes/dupes make files impossible to correlate with the archive | Always `[%(id)s]` in the template | | `--cookies-from-browser chrome` on Windows | Chrome 127+ app-bound encryption — extraction fails | Use `firefox`, or export `cookies.txt` | | Authenticated bulk runs, no sleeps | Account flagged/banned; IP rate-limited | Throwaway account + `--sleep-requests`/`--sleep-interval` always | | Throttled to ~50-100 KB/s | Looks like a network problem; it's the nsig arms race | Update yt-dlp FIRST (`check-ytdlp-version.sh --live`) | | "nsig extraction failed" / "unable to extract" | Debugging the command/network when the binary is stale | Same — update first; these errors mean *outdated*, not *broken usage* | | Raw `%(title)s` filenames | Emoji/colons/slashes break on Windows and some CI filesystems | `--restrict-filenames` or `.100B`-truncated fields + `[%(id)s]` | | Thin format list on a fresh machine | No JS runtime — YouTube player JS now needs one (EJS); runtime-less extraction is deprecated and may offer only low-res premuxed | Install deno, or `--js-runtimes node`; see [failure-triage](references/failure-triage.md) | | git-bash (MSYS) path mangling | `/tmp/...`-style args convert per-arg — templates containing `%(...)s` skip conversion while plain paths convert, scattering outputs | Use Windows-style paths (`X:/dir/...`) for `-o`/`-P`/`--download-archive` under git-bash | | pip-installed `yt-dlp -U` | Self-update doesn't work for pip/uv installs (silently a no-op with a warning) | `uv tool upgrade yt-dlp`; `-U` is for the standalone binary only | ## Failure triage The ladder — run in order, stop at the first fix: 1. **Stale binary?** `check-ytdlp-version.sh --live` → exit 10 → update. This closes most "nsig extraction failed", missing-format, and throttling cases. 2. **Reproduce verbosely:** `yt-dlp -v URL` — read the actual extractor error, don't guess from the summary line. 3. **403 / "Sign in to confirm you're not a bot"** → identity problem: `--cookies-from-browser firefox`, or a different network/IP. 4. **429 / sudden slowdowns mid-run** → rate limited: add the politeness flags, reduce `--concurrent-fragments`, back off and resume later (archive files make every run resumable). 5. **Geo block** ("not available in your country") → `--proxy URL` through an allowed region; the old `--geo-bypass` header tricks rarely work anymore. Full decision tree with error-message → cause mapping, `--extractor-args` escape hatches, and when to file upstream: [references/failure-triage.md](references/failure-triage.md). ## Scripts Follows the [Skill Resource Protocol](../../docs/SKILL-RESOURCE-PROTOCOL.md): `--help` with examples, stdout = data only, `--json` envelope (`claude-mods.ytdlp-ops.version-check/v1`), semantic exit codes (`0` clean, `2` usage, `7` network/yt-dlp unavailable — advisory, `10` drift finding). | Script | Job | Worked invocation | |---|---|---| | `check-ytdlp-version.sh` | Staleness verifier: `--offline` structural (CI gate), `--live` = installed-version age vs latest GitHub release + documented-flag existence in `yt-dlp --help` + metadata-only smoke extraction | `check-ytdlp-version.sh --live --json \| jq '.data.days_behind'` — exit 10 = >60 days behind, a documented flag vanished, or an extractor-internal smoke failure; 7 = advisory (network/API unreachable, or the smoke test was blocked/unverifiable — normal on a bot-gated datacenter IP) | ## References Load on demand — one concept per file: | Reference | Load when | |---|---| | [format-selection.md](references/format-selection.md) | Any `-f`/`-S` decision, codec targeting, filter grammar, avoiding transcodes | | [playlists-archives.md](references/playlists-archives.md) | Playlists, channels, `--download-archive`, incremental sync jobs | | [auth-cookies.md](references/auth-cookies.md) | Private/members/age-gated content, browser cookie matrix, ban avoidance | | [output-templates.md](references/output-templates.md) | `-o` field catalog, paths, sanitization, per-type routing | | [subtitles-metadata.md](references/subtitles-metadata.md) | Sub download/convert/embed, transcript workflows, metadata/thumbnail embedding | | [sponsorblock.md](references/sponsorblock.md) | SponsorBlock categories, mark vs remove, chapter workflows | | [failure-triage.md](references/failure-triage.md) | Any download failure — 403/429/geo/nsig/throttling decision tree | Assets: [format-presets.json](assets/format-presets.json) — canonical, date-stamped `-S`/flag presets per destination (web MP4, STT audio, archival, clip, mobile-small). ## Self-test ```bash bash skills/ytdlp-ops/tests/run.sh # fully offline; no network, no yt-dlp needed ``` Structural assertions plus the verifier's 60-day age logic exercised through its `CM_YTDLP_INSTALLED`/`CM_YTDLP_LATEST` test seams. Real `--live` runs happen only in the scheduled freshness workflow — a network blip must never fail a PR.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.