querying-markdown
Query, filter, and transform Markdown structurally with mq — a jq-like CLI for Markdown. Use to extract headings/sections/code-blocks/links from .md files, build a table of contents, pull code blocks of a given language, slice or reshape LLM prompt/output Markdown, or batch-trans
Install
npx skills add https://github.com/oaustegard/claude-skills/tree/main/plugins/development-tools/skills/querying-markdown
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install oaustegard-claude-skills@llmmart
git clone https://github.com/oaustegard/claude-skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole oaustegard/claude-skills collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
querying-markdown
mq is "jq for Markdown" — it parses a .md file into a node stream and lets
you select, filter, and transform by structure (.h2, .code("rust"),
.link) instead of by line-matching. Reach for it when the task is structural:
"every H2 title", "all bash code blocks", "a table of contents", "strip the
frontmatter". For plain substring search, grep is still the right tool; for
code (not prose) structure, use tree-sitting.
Before you use mq: is this actually a structural task?
mq parses the whole document into a node tree before it answers, and that parse cost is real (see Empirical findings). Most "query a markdown file" tasks don't need it. Decide first, using the target — not the file type:
| Your target | Use | Why |
|---|---|---|
Lines with a fixed prefix — #/## headings, > quotes, - bullets, a leading line/verse number |
grep / awk | Line-matching, not structure. grep is faster and already installed. |
| A substring anywhere | grep | mq adds nothing. |
| Code structure inside fences (ASTs, symbols, call sites) | tree-sitting | mq sees the fence, not the code inside it. |
Language-filtered code blocks (.code("bash")); links as structured (text, url) (-F json '.link') |
mq | grep can't filter a fenced block by language without a brittle hand-rolled fence state machine. |
| Markdown→Markdown transforms that must emit valid Markdown — demote/promote headings, rebuild a TOC with anchors, in-place edit | mq | sed doesn't know structure and will corrupt nesting/fences. |
If your task lands in a grep/awk row, do not install mq — close this skill and use the line tool. Diagnosed 2026-06-04: a full-KJV smoke test queried books/chapters/verses (all line-prefix structure) with mq — ~3.3 s per query where grep is milliseconds, the same answers, and a grep post-filter still needed on top. Wrong-shape corpus; mq's selectors earn their parse cost only on the structural rows.
The judgment call is whether the case is actually line-prefix or only looks it. A heading is a prefix; a heading you want demoted with its subtree, or a match you must re-emit as valid Markdown, is structure — mq's row even when the match looks like a prefix.
Setup
mq is a single static binary, not preinstalled. Install on first use (idempotent — exits early if already present, ~1s, no build step):
bash /mnt/skills/user/querying-markdown/scripts/install-mq.sh
This drops the pinned mq release into /usr/local/bin. Override the version
with MQ_VERSION=vX.Y.Z.
Usage
mq 'QUERY' file.md # query a file
cat file.md | mq 'QUERY' # query stdin
mq repl # interactive REPL — use to test syntax fast
A node stream flows left→right through |. Selectors (.h, .code, .link)
pick nodes; functions (to_text, slugify, map, len) transform them.
self is the current node.
mq '.h2 | to_text()' README.md # every H2 as plain text
mq '.code("python") | to_text()' file.md # all python code blocks
mq '.h.level' file.md # heading depth per heading
mq -F json '.h2 | to_text()' file.md # results as JSON
mq '.h2 | to_text()' file.md | wc -l # count matches (reliable idiom)
Empirical findings
Measured 2026-06-04 against a full public-domain KJV Bible (66 files, 4.28 MB).
Parse-bound, not query-bound. mq reparses the whole document on every
invocation; latency tracks document size, not selector or match count. On the
4.28 MB file every query — whether it returned 66 matches or 32,418 — ran
~3.2–3.3 s (~1.3 MB/s); on a normal-sized doc it is single-digit ms. Never loop
mq per query over a large corpus: extract once with -F json and work on the
result, or accept a constant per-call parse tax.
Selectors return nodes, not your domain concepts. .h2 over the KJV
returned 1,250 nodes — 1,184 chapter headings plus 66 eof markers the source
appended per file, while single-chapter books emitted no chapter heading at all.
.text also pulled heading text into the paragraph stream. A raw selector count
is a node count; map it to your concept with an explicit predicate
(e.g. grep -E '^[0-9]+ ' for verses) and check it against a known total before
trusting the number.
An empty result is ambiguous. Zero output means either the selector
matched nothing or mq never ran — a wrapper like time/env failed in dash,
or a malformed heredoc swallowed the command. Re-run the bare
mq 'QUERY' file.md before concluding a selector or function is broken.
(Self-inflicted 2026-06-04: a time: not found shell error read as a
to_text() defect; to_text() on code blocks works.)
Reference
Selector aliases, the built-in function library, table-of-contents and
transform recipes, in-place-edit caveats, and CLI flags live in
references/cheatsheet.md. Read it before writing a
non-trivial query — the dialect is jq-like, not jq, so the function names
differ. When unsure of syntax, mq repl gives instant feedback.
Files (claude-skills)
-
references
-
cheatsheet.md 4.3 KB
# mq cheatsheet jq-like query language for Markdown. A node stream flows left→right through `|`; selectors pick nodes, functions transform them. All examples below were verified against mq 0.5.31. ## Invocation ```bash mq 'QUERY' file.md # query a file cat file.md | mq 'QUERY' # query stdin mq -f query.mq file.md # load query from a file mq repl # interactive REPL — use this to test syntax ``` Useful flags: | Flag | Effect | | ---- | ------ | | `-F, --output-format` | `markdown` (default), `html`, `text`, `json`, `table`, `grep`, `raw`, `none` | | `-I, --input-format` | force input parse: `markdown` (default), `html`, `text`, `json`, `csv`, `yaml`, `toml`, `xml`, … | | `-o FILE` | write output to FILE | | `-A, --aggregate` | aggregate all input into a single array before the query | | `-U, --in-place` | write result back to the input file — **whole-document transforms only** (see Transforming) | | `-S QUERY` | separator query inserted between multiple input files | ## Selectors `.` prefix selects Markdown nodes: ``` .h # all headings .code # fenced code blocks .text # paragraphs (.p) .code_inline # inline code (.inline_code) .link # links .list # list items (.li) .table # tables .hr # horizontal rules ``` Selector **calls** filter by property: ``` .h(1) # only h1 .h(2, 3) # h2 and h3 .h(1..3) # h1 through h3 (range) .h2 # shorthand for .h(2) (also .h1 … .h6) .code("rust") # only rust-tagged code blocks ``` Attribute access via dot notation: ``` .h.level # heading depth 1–6 (alias .h.depth) .h.value # heading text .code.value ``` ## Verified examples ```bash # Section titles — every H2 as plain text mq '.h2 | to_text()' file.md # All headings, any level mq '.h | to_text()' file.md # Heading levels (one integer per heading) mq '.h.level' file.md # Extract only the bash code blocks mq '.code("bash") | to_text()' file.md # All links (rendered as markdown) mq '.link' file.md # Slugify the H1 (e.g. for an anchor/filename) mq '.h1 | to_text() | slugify(self)' file.md # Headings as structured JSON mq -F json '.h2 | to_text()' file.md # Count matches — pipe to wc, the reliable idiom mq '.h2 | to_text()' file.md | wc -l # Table-of-contents: nested markdown list of headings with anchor links mq '.h | let link = to_link("#" + to_text(self), to_text(self), "") | let level = .h.depth | if (not(is_none(level))): to_md_list(link, to_number(level))' file.md ``` ## Transforming `self` refers to the current node. Build transforms with the function library: ```bash # Uppercase H1 text (emitted to stdout) mq 'if (is_h1(self)): upcase(to_text(self)) else: self' file.md # Increase every heading depth by one (demote) mq 'nodes | map(increase_header_level)' file.md ``` `-U`/`--in-place` writes back **only when the query yields a complete document** (e.g. a `nodes | map(...)` transform). A filtering selector like `.h1 | …` produces a partial stream, so mq emits it to stdout and leaves the file untouched — verify the write-back in `mq repl` before relying on `-U` for a destructive edit, or just use `-o out.md`. ## Functions (built-in) Selection / iteration: `select` `filter` `reject` `map` `compact_map` `flat_map` `each` `find_index` `first` `second` `last` `take` `skip` `take_while` `skip_while` `nodes` Text: `to_text` `upcase` `downcase` `slugify` `lpad` `rpad` `ltrimstr` `rtrimstr` `lines` `unlines` `matches_url` `test` `ngram` Markdown builders: `to_link` `to_md_list` `to_number` `increase_header_level` `decrease_header_level` `promote_heading` `demote_heading` `frontmatter` `load_markdown` Aggregation: `len` `count_by` `group_by` `sort_by` `unique_by` `sum` `sum_by` `mean` `max_by` `min_by` `partition` `chunks` `zip` `transpose` `fold` Predicates: `is_h` `is_h1`…`is_h6` `is_code` `is_list` `is_text` `is_link` `is_table_cell` `is_em` `is_none` `is_empty` `is_array` `is_string` `is_number` Run `mq --list` for the full subcommand/function surface, or open `mq repl` and experiment. Full reference: https://mqlang.org/book/ ## Defining functions ``` def double(x): mul(x, 2); # named, ';' or 'end' terminates nodes | map(fn(x): upcase(x);) # anonymous (lambda) nodes | map(->(x): upcase(x);) # '->' is shorthand for 'fn' ```
-
-
scripts
-
install-mq.sh 1.1 KB
#!/usr/bin/env bash # Idempotent installer for `mq`, the jq-for-Markdown CLI (https://github.com/harehare/mq). # Fetches the static x86_64-linux binary into /usr/local/bin. Safe to re-run — exits # early if mq is already on PATH. No build step; mq is a single self-contained binary. set -euo pipefail MQ_VERSION="${MQ_VERSION:-v0.5.31}" DEST="${MQ_DEST:-/usr/local/bin/mq}" ASSET="mq-x86_64-unknown-linux-gnu" if command -v mq >/dev/null 2>&1; then echo "mq already installed: $(mq --version 2>/dev/null) ($(command -v mq))" exit 0 fi echo "Installing mq ${MQ_VERSION} -> ${DEST}" tmp="$(mktemp)" if command -v gh >/dev/null 2>&1; then # gh carries GH_TOKEN auth, sidestepping shared-container-IP rate limits. gh release download "${MQ_VERSION}" --repo harehare/mq \ --pattern "${ASSET}" --output "${tmp}" --clobber else curl -fsSL "https://github.com/harehare/mq/releases/download/${MQ_VERSION}/${ASSET}" -o "${tmp}" fi chmod +x "${tmp}" # /usr/local/bin is typically root-owned; fall back to sudo if a plain mv is denied. mv "${tmp}" "${DEST}" 2>/dev/null || sudo mv "${tmp}" "${DEST}" echo "Installed: $(mq --version)"
-
-
CHANGELOG.md 1.1 KB
# querying-markdown - Changelog All notable changes to the `querying-markdown` skill are documented in this file. The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/). ## [0.2.0] - 2026-06-04 ### Other - querying-markdown v0.2.0: use-gate + empirical findings (#687) ## [0.2.0] - 2026-06-04 ### Added - Use-gate decision table ("Before you use mq: is this actually a structural task?") routing line-prefix/substring work to grep/awk and code-structure to tree-sitting, reserving mq for language-filtered code blocks, structured links, and valid-Markdown transforms. - "Empirical findings" section: parse-bound performance (latency tracks document size, not selector/match count; ~1.3 MB/s), selectors-return-nodes-not-concepts, and the ambiguous-empty-result trap. ### Notes - Findings measured against a full KJV smoke test. The cheatsheet's `.code("lang") | to_text()` idiom is correct (an earlier "bug" report was a `dash` `time: not found` error misread as empty output). ## [0.1.0] - 2026-06-04 ### Other - Add querying-markdown skill (mq, jq-for-Markdown) (#686) -
SKILL.md 5.6 KB
--- name: querying-markdown description: Query, filter, and transform Markdown structurally with mq — a jq-like CLI for Markdown. Use to extract headings/sections/code-blocks/links from .md files, build a table of contents, pull code blocks of a given language, slice or reshape LLM prompt/output Markdown, or batch-transform docs. Triggers on "extract sections from this markdown", "get all the code blocks", "jq for markdown", "mq", or any structural query over Markdown that grep/Read can't do cleanly. metadata: version: 0.2.0 --- # querying-markdown `mq` is "jq for Markdown" — it parses a `.md` file into a node stream and lets you select, filter, and transform by structure (`.h2`, `.code("rust")`, `.link`) instead of by line-matching. Reach for it when the task is structural: "every H2 title", "all bash code blocks", "a table of contents", "strip the frontmatter". For plain substring search, `grep` is still the right tool; for code (not prose) structure, use `tree-sitting`. ## Before you use mq: is this actually a structural task? mq parses the whole document into a node tree before it answers, and that parse cost is real (see [Empirical findings](#empirical-findings)). Most "query a markdown file" tasks don't need it. Decide first, using the target — not the file type: | Your target | Use | Why | | --- | --- | --- | | Lines with a fixed prefix — `#`/`##` headings, `>` quotes, `-` bullets, a leading line/verse number | **grep / awk** | Line-matching, not structure. grep is faster and already installed. | | A substring anywhere | **grep** | mq adds nothing. | | Code *structure* inside fences (ASTs, symbols, call sites) | **tree-sitting** | mq sees the fence, not the code inside it. | | Language-filtered code blocks (`.code("bash")`); links as structured `(text, url)` (`-F json '.link'`) | **mq** | grep can't filter a fenced block by language without a brittle hand-rolled fence state machine. | | Markdown→Markdown transforms that must emit valid Markdown — demote/promote headings, rebuild a TOC with anchors, in-place edit | **mq** | `sed` doesn't know structure and will corrupt nesting/fences. | If your task lands in a grep/awk row, **do not install mq** — close this skill and use the line tool. Diagnosed 2026-06-04: a full-KJV smoke test queried books/chapters/verses (all line-prefix structure) with mq — ~3.3 s per query where grep is milliseconds, the same answers, and a grep post-filter still needed on top. Wrong-shape corpus; mq's selectors earn their parse cost only on the structural rows. The judgment call is whether the case is *actually* line-prefix or only looks it. A heading is a prefix; a heading you want demoted **with its subtree**, or a match you must re-emit as valid Markdown, is structure — mq's row even when the match looks like a prefix. ## Setup mq is a single static binary, not preinstalled. Install on first use (idempotent — exits early if already present, ~1s, no build step): ```bash bash /mnt/skills/user/querying-markdown/scripts/install-mq.sh ``` This drops the pinned `mq` release into `/usr/local/bin`. Override the version with `MQ_VERSION=vX.Y.Z`. ## Usage ```bash mq 'QUERY' file.md # query a file cat file.md | mq 'QUERY' # query stdin mq repl # interactive REPL — use to test syntax fast ``` A node stream flows left→right through `|`. Selectors (`.h`, `.code`, `.link`) pick nodes; functions (`to_text`, `slugify`, `map`, `len`) transform them. `self` is the current node. ```bash mq '.h2 | to_text()' README.md # every H2 as plain text mq '.code("python") | to_text()' file.md # all python code blocks mq '.h.level' file.md # heading depth per heading mq -F json '.h2 | to_text()' file.md # results as JSON mq '.h2 | to_text()' file.md | wc -l # count matches (reliable idiom) ``` ## Empirical findings Measured 2026-06-04 against a full public-domain KJV Bible (66 files, 4.28 MB). **Parse-bound, not query-bound.** mq reparses the whole document on every invocation; latency tracks document *size*, not selector or match count. On the 4.28 MB file every query — whether it returned 66 matches or 32,418 — ran ~3.2–3.3 s (~1.3 MB/s); on a normal-sized doc it is single-digit ms. Never loop mq per query over a large corpus: extract once with `-F json` and work on the result, or accept a constant per-call parse tax. **Selectors return nodes, not your domain concepts.** `.h2` over the KJV returned 1,250 nodes — 1,184 chapter headings plus 66 `eof` markers the source appended per file, while single-chapter books emitted no chapter heading at all. `.text` also pulled heading text into the paragraph stream. A raw selector count is a *node* count; map it to your concept with an explicit predicate (e.g. `grep -E '^[0-9]+ '` for verses) and check it against a known total before trusting the number. **An empty result is ambiguous.** Zero output means *either* the selector matched nothing *or* mq never ran — a wrapper like `time`/`env` failed in `dash`, or a malformed heredoc swallowed the command. Re-run the bare `mq 'QUERY' file.md` before concluding a selector or function is broken. (Self-inflicted 2026-06-04: a `time: not found` shell error read as a `to_text()` defect; `to_text()` on code blocks works.) ## Reference Selector aliases, the built-in function library, table-of-contents and transform recipes, in-place-edit caveats, and CLI flags live in [references/cheatsheet.md](references/cheatsheet.md). Read it before writing a non-trivial query — the dialect is jq-*like*, not jq, so the function names differ. When unsure of syntax, `mq repl` gives instant feedback.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.