Claude Skill

querying-markdown

Query, filter, and transform Markdown structurally with mq — a jq-like CLI for Markdown. Use to extract headings/sections/code-blocks/links from .md files, build a table of contents, pull code blocks of a given language, slice or reshape LLM prompt/output Markdown, or batch-trans

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download oaustegard-claude-skills-plugins_development-tools_skills_querying-markdown-e39c726.zip · 6 KB
Part of oaustegard/claude-skills — 39 skills

Install

skills CLI npx skills add https://github.com/oaustegard/claude-skills/tree/main/plugins/development-tools/skills/querying-markdown
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install oaustegard-claude-skills@llmmart
Git git clone https://github.com/oaustegard/claude-skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole oaustegard/claude-skills collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

querying-markdown

mq is "jq for Markdown" — it parses a .md file into a node stream and lets you select, filter, and transform by structure (.h2, .code("rust"), .link) instead of by line-matching. Reach for it when the task is structural: "every H2 title", "all bash code blocks", "a table of contents", "strip the frontmatter". For plain substring search, grep is still the right tool; for code (not prose) structure, use tree-sitting.

Before you use mq: is this actually a structural task?

mq parses the whole document into a node tree before it answers, and that parse cost is real (see Empirical findings). Most "query a markdown file" tasks don't need it. Decide first, using the target — not the file type:

Your target Use Why
Lines with a fixed prefix — #/## headings, > quotes, - bullets, a leading line/verse number grep / awk Line-matching, not structure. grep is faster and already installed.
A substring anywhere grep mq adds nothing.
Code structure inside fences (ASTs, symbols, call sites) tree-sitting mq sees the fence, not the code inside it.
Language-filtered code blocks (.code("bash")); links as structured (text, url) (-F json '.link') mq grep can't filter a fenced block by language without a brittle hand-rolled fence state machine.
Markdown→Markdown transforms that must emit valid Markdown — demote/promote headings, rebuild a TOC with anchors, in-place edit mq sed doesn't know structure and will corrupt nesting/fences.

If your task lands in a grep/awk row, do not install mq — close this skill and use the line tool. Diagnosed 2026-06-04: a full-KJV smoke test queried books/chapters/verses (all line-prefix structure) with mq — ~3.3 s per query where grep is milliseconds, the same answers, and a grep post-filter still needed on top. Wrong-shape corpus; mq's selectors earn their parse cost only on the structural rows.

The judgment call is whether the case is actually line-prefix or only looks it. A heading is a prefix; a heading you want demoted with its subtree, or a match you must re-emit as valid Markdown, is structure — mq's row even when the match looks like a prefix.

Setup

mq is a single static binary, not preinstalled. Install on first use (idempotent — exits early if already present, ~1s, no build step):

bash /mnt/skills/user/querying-markdown/scripts/install-mq.sh

This drops the pinned mq release into /usr/local/bin. Override the version with MQ_VERSION=vX.Y.Z.

Usage

mq 'QUERY' file.md          # query a file
cat file.md | mq 'QUERY'    # query stdin
mq repl                     # interactive REPL — use to test syntax fast

A node stream flows left→right through |. Selectors (.h, .code, .link) pick nodes; functions (to_text, slugify, map, len) transform them. self is the current node.

mq '.h2 | to_text()' README.md            # every H2 as plain text
mq '.code("python") | to_text()' file.md  # all python code blocks
mq '.h.level' file.md                     # heading depth per heading
mq -F json '.h2 | to_text()' file.md      # results as JSON
mq '.h2 | to_text()' file.md | wc -l      # count matches (reliable idiom)

Empirical findings

Measured 2026-06-04 against a full public-domain KJV Bible (66 files, 4.28 MB).

Parse-bound, not query-bound. mq reparses the whole document on every invocation; latency tracks document size, not selector or match count. On the 4.28 MB file every query — whether it returned 66 matches or 32,418 — ran ~3.2–3.3 s (~1.3 MB/s); on a normal-sized doc it is single-digit ms. Never loop mq per query over a large corpus: extract once with -F json and work on the result, or accept a constant per-call parse tax.

Selectors return nodes, not your domain concepts. .h2 over the KJV returned 1,250 nodes — 1,184 chapter headings plus 66 eof markers the source appended per file, while single-chapter books emitted no chapter heading at all. .text also pulled heading text into the paragraph stream. A raw selector count is a node count; map it to your concept with an explicit predicate (e.g. grep -E '^[0-9]+ ' for verses) and check it against a known total before trusting the number.

An empty result is ambiguous. Zero output means either the selector matched nothing or mq never ran — a wrapper like time/env failed in dash, or a malformed heredoc swallowed the command. Re-run the bare mq 'QUERY' file.md before concluding a selector or function is broken. (Self-inflicted 2026-06-04: a time: not found shell error read as a to_text() defect; to_text() on code blocks works.)

Reference

Selector aliases, the built-in function library, table-of-contents and transform recipes, in-place-edit caveats, and CLI flags live in references/cheatsheet.md. Read it before writing a non-trivial query — the dialect is jq-like, not jq, so the function names differ. When unsure of syntax, mq repl gives instant feedback.

Files (claude-skills)
  • references
    • cheatsheet.md 4.3 KB
      # mq cheatsheet
      
      jq-like query language for Markdown. A node stream flows left→right through `|`;
      selectors pick nodes, functions transform them. All examples below were verified
      against mq 0.5.31.
      
      ## Invocation
      
      ```bash
      mq 'QUERY' file.md            # query a file
      cat file.md | mq 'QUERY'      # query stdin
      mq -f query.mq file.md        # load query from a file
      mq repl                       # interactive REPL — use this to test syntax
      ```
      
      Useful flags:
      
      | Flag | Effect |
      | ---- | ------ |
      | `-F, --output-format` | `markdown` (default), `html`, `text`, `json`, `table`, `grep`, `raw`, `none` |
      | `-I, --input-format` | force input parse: `markdown` (default), `html`, `text`, `json`, `csv`, `yaml`, `toml`, `xml`, … |
      | `-o FILE` | write output to FILE |
      | `-A, --aggregate` | aggregate all input into a single array before the query |
      | `-U, --in-place` | write result back to the input file — **whole-document transforms only** (see Transforming) |
      | `-S QUERY` | separator query inserted between multiple input files |
      
      ## Selectors
      
      `.` prefix selects Markdown nodes:
      
      ```
      .h        # all headings        .code        # fenced code blocks
      .text     # paragraphs (.p)     .code_inline # inline code (.inline_code)
      .link     # links               .list        # list items (.li)
      .table    # tables              .hr          # horizontal rules
      ```
      
      Selector **calls** filter by property:
      
      ```
      .h(1)          # only h1
      .h(2, 3)       # h2 and h3
      .h(1..3)       # h1 through h3 (range)
      .h2            # shorthand for .h(2)  (also .h1 … .h6)
      .code("rust")  # only rust-tagged code blocks
      ```
      
      Attribute access via dot notation:
      
      ```
      .h.level   # heading depth 1–6 (alias .h.depth)
      .h.value   # heading text
      .code.value
      ```
      
      ## Verified examples
      
      ```bash
      # Section titles — every H2 as plain text
      mq '.h2 | to_text()' file.md
      
      # All headings, any level
      mq '.h | to_text()' file.md
      
      # Heading levels (one integer per heading)
      mq '.h.level' file.md
      
      # Extract only the bash code blocks
      mq '.code("bash") | to_text()' file.md
      
      # All links (rendered as markdown)
      mq '.link' file.md
      
      # Slugify the H1 (e.g. for an anchor/filename)
      mq '.h1 | to_text() | slugify(self)' file.md
      
      # Headings as structured JSON
      mq -F json '.h2 | to_text()' file.md
      
      # Count matches — pipe to wc, the reliable idiom
      mq '.h2 | to_text()' file.md | wc -l
      
      # Table-of-contents: nested markdown list of headings with anchor links
      mq '.h
          | let link = to_link("#" + to_text(self), to_text(self), "")
          | let level = .h.depth
          | if (not(is_none(level))): to_md_list(link, to_number(level))' file.md
      ```
      
      ## Transforming
      
      `self` refers to the current node. Build transforms with the function library:
      
      ```bash
      # Uppercase H1 text (emitted to stdout)
      mq 'if (is_h1(self)): upcase(to_text(self)) else: self' file.md
      
      # Increase every heading depth by one (demote)
      mq 'nodes | map(increase_header_level)' file.md
      ```
      
      `-U`/`--in-place` writes back **only when the query yields a complete document**
      (e.g. a `nodes | map(...)` transform). A filtering selector like `.h1 | …`
      produces a partial stream, so mq emits it to stdout and leaves the file
      untouched — verify the write-back in `mq repl` before relying on `-U` for a
      destructive edit, or just use `-o out.md`.
      
      ## Functions (built-in)
      
      Selection / iteration: `select` `filter` `reject` `map` `compact_map`
      `flat_map` `each` `find_index` `first` `second` `last` `take` `skip`
      `take_while` `skip_while` `nodes`
      
      Text: `to_text` `upcase` `downcase` `slugify` `lpad` `rpad` `ltrimstr`
      `rtrimstr` `lines` `unlines` `matches_url` `test` `ngram`
      
      Markdown builders: `to_link` `to_md_list` `to_number`
      `increase_header_level` `decrease_header_level` `promote_heading`
      `demote_heading` `frontmatter` `load_markdown`
      
      Aggregation: `len` `count_by` `group_by` `sort_by` `unique_by` `sum` `sum_by`
      `mean` `max_by` `min_by` `partition` `chunks` `zip` `transpose` `fold`
      
      Predicates: `is_h` `is_h1`…`is_h6` `is_code` `is_list` `is_text` `is_link`
      `is_table_cell` `is_em` `is_none` `is_empty` `is_array` `is_string` `is_number`
      
      Run `mq --list` for the full subcommand/function surface, or open `mq repl`
      and experiment. Full reference: https://mqlang.org/book/
      
      ## Defining functions
      
      ```
      def double(x): mul(x, 2);          # named, ';' or 'end' terminates
      nodes | map(fn(x): upcase(x);)     # anonymous (lambda)
      nodes | map(->(x): upcase(x);)     # '->' is shorthand for 'fn'
      ```
      
  • scripts
    • install-mq.sh 1.1 KB
      #!/usr/bin/env bash
      # Idempotent installer for `mq`, the jq-for-Markdown CLI (https://github.com/harehare/mq).
      # Fetches the static x86_64-linux binary into /usr/local/bin. Safe to re-run — exits
      # early if mq is already on PATH. No build step; mq is a single self-contained binary.
      set -euo pipefail
      
      MQ_VERSION="${MQ_VERSION:-v0.5.31}"
      DEST="${MQ_DEST:-/usr/local/bin/mq}"
      ASSET="mq-x86_64-unknown-linux-gnu"
      
      if command -v mq >/dev/null 2>&1; then
        echo "mq already installed: $(mq --version 2>/dev/null) ($(command -v mq))"
        exit 0
      fi
      
      echo "Installing mq ${MQ_VERSION} -> ${DEST}"
      tmp="$(mktemp)"
      
      if command -v gh >/dev/null 2>&1; then
        # gh carries GH_TOKEN auth, sidestepping shared-container-IP rate limits.
        gh release download "${MQ_VERSION}" --repo harehare/mq \
           --pattern "${ASSET}" --output "${tmp}" --clobber
      else
        curl -fsSL "https://github.com/harehare/mq/releases/download/${MQ_VERSION}/${ASSET}" -o "${tmp}"
      fi
      
      chmod +x "${tmp}"
      # /usr/local/bin is typically root-owned; fall back to sudo if a plain mv is denied.
      mv "${tmp}" "${DEST}" 2>/dev/null || sudo mv "${tmp}" "${DEST}"
      
      echo "Installed: $(mq --version)"
      
  • CHANGELOG.md 1.1 KB
    # querying-markdown - Changelog
    
    All notable changes to the `querying-markdown` skill are documented in this file. The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/).
    
    ## [0.2.0] - 2026-06-04
    
    ### Other
    
    - querying-markdown v0.2.0: use-gate + empirical findings (#687)
    
    ## [0.2.0] - 2026-06-04
    
    ### Added
    
    - Use-gate decision table ("Before you use mq: is this actually a structural task?") routing line-prefix/substring work to grep/awk and code-structure to tree-sitting, reserving mq for language-filtered code blocks, structured links, and valid-Markdown transforms.
    - "Empirical findings" section: parse-bound performance (latency tracks document size, not selector/match count; ~1.3 MB/s), selectors-return-nodes-not-concepts, and the ambiguous-empty-result trap.
    
    ### Notes
    
    - Findings measured against a full KJV smoke test. The cheatsheet's `.code("lang") | to_text()` idiom is correct (an earlier "bug" report was a `dash` `time: not found` error misread as empty output).
    
    ## [0.1.0] - 2026-06-04
    
    ### Other
    
    - Add querying-markdown skill (mq, jq-for-Markdown) (#686)
  • SKILL.md 5.6 KB
    ---
    name: querying-markdown
    description: Query, filter, and transform Markdown structurally with mq — a jq-like CLI for Markdown. Use to extract headings/sections/code-blocks/links from .md files, build a table of contents, pull code blocks of a given language, slice or reshape LLM prompt/output Markdown, or batch-transform docs. Triggers on "extract sections from this markdown", "get all the code blocks", "jq for markdown", "mq", or any structural query over Markdown that grep/Read can't do cleanly.
    metadata:
      version: 0.2.0
    ---
    
    # querying-markdown
    
    `mq` is "jq for Markdown" — it parses a `.md` file into a node stream and lets
    you select, filter, and transform by structure (`.h2`, `.code("rust")`,
    `.link`) instead of by line-matching. Reach for it when the task is structural:
    "every H2 title", "all bash code blocks", "a table of contents", "strip the
    frontmatter". For plain substring search, `grep` is still the right tool; for
    code (not prose) structure, use `tree-sitting`.
    
    ## Before you use mq: is this actually a structural task?
    
    mq parses the whole document into a node tree before it answers, and that parse
    cost is real (see [Empirical findings](#empirical-findings)). Most "query a
    markdown file" tasks don't need it. Decide first, using the target — not the
    file type:
    
    | Your target | Use | Why |
    | --- | --- | --- |
    | Lines with a fixed prefix — `#`/`##` headings, `>` quotes, `-` bullets, a leading line/verse number | **grep / awk** | Line-matching, not structure. grep is faster and already installed. |
    | A substring anywhere | **grep** | mq adds nothing. |
    | Code *structure* inside fences (ASTs, symbols, call sites) | **tree-sitting** | mq sees the fence, not the code inside it. |
    | Language-filtered code blocks (`.code("bash")`); links as structured `(text, url)` (`-F json '.link'`) | **mq** | grep can't filter a fenced block by language without a brittle hand-rolled fence state machine. |
    | Markdown→Markdown transforms that must emit valid Markdown — demote/promote headings, rebuild a TOC with anchors, in-place edit | **mq** | `sed` doesn't know structure and will corrupt nesting/fences. |
    
    If your task lands in a grep/awk row, **do not install mq** — close this skill
    and use the line tool. Diagnosed 2026-06-04: a full-KJV smoke test queried
    books/chapters/verses (all line-prefix structure) with mq — ~3.3 s per query
    where grep is milliseconds, the same answers, and a grep post-filter still
    needed on top. Wrong-shape corpus; mq's selectors earn their parse cost only on
    the structural rows.
    
    The judgment call is whether the case is *actually* line-prefix or only looks
    it. A heading is a prefix; a heading you want demoted **with its subtree**, or a
    match you must re-emit as valid Markdown, is structure — mq's row even when the
    match looks like a prefix.
    
    ## Setup
    
    mq is a single static binary, not preinstalled. Install on first use (idempotent
    — exits early if already present, ~1s, no build step):
    
    ```bash
    bash /mnt/skills/user/querying-markdown/scripts/install-mq.sh
    ```
    
    This drops the pinned `mq` release into `/usr/local/bin`. Override the version
    with `MQ_VERSION=vX.Y.Z`.
    
    ## Usage
    
    ```bash
    mq 'QUERY' file.md          # query a file
    cat file.md | mq 'QUERY'    # query stdin
    mq repl                     # interactive REPL — use to test syntax fast
    ```
    
    A node stream flows left→right through `|`. Selectors (`.h`, `.code`, `.link`)
    pick nodes; functions (`to_text`, `slugify`, `map`, `len`) transform them.
    `self` is the current node.
    
    ```bash
    mq '.h2 | to_text()' README.md            # every H2 as plain text
    mq '.code("python") | to_text()' file.md  # all python code blocks
    mq '.h.level' file.md                     # heading depth per heading
    mq -F json '.h2 | to_text()' file.md      # results as JSON
    mq '.h2 | to_text()' file.md | wc -l      # count matches (reliable idiom)
    ```
    
    ## Empirical findings
    
    Measured 2026-06-04 against a full public-domain KJV Bible (66 files, 4.28 MB).
    
    **Parse-bound, not query-bound.** mq reparses the whole document on every
    invocation; latency tracks document *size*, not selector or match count. On the
    4.28 MB file every query — whether it returned 66 matches or 32,418 — ran
    ~3.2–3.3 s (~1.3 MB/s); on a normal-sized doc it is single-digit ms. Never loop
    mq per query over a large corpus: extract once with `-F json` and work on the
    result, or accept a constant per-call parse tax.
    
    **Selectors return nodes, not your domain concepts.** `.h2` over the KJV
    returned 1,250 nodes — 1,184 chapter headings plus 66 `eof` markers the source
    appended per file, while single-chapter books emitted no chapter heading at all.
    `.text` also pulled heading text into the paragraph stream. A raw selector count
    is a *node* count; map it to your concept with an explicit predicate
    (e.g. `grep -E '^[0-9]+ '` for verses) and check it against a known total before
    trusting the number.
    
    **An empty result is ambiguous.** Zero output means *either* the selector
    matched nothing *or* mq never ran — a wrapper like `time`/`env` failed in `dash`,
    or a malformed heredoc swallowed the command. Re-run the bare
    `mq 'QUERY' file.md` before concluding a selector or function is broken.
    (Self-inflicted 2026-06-04: a `time: not found` shell error read as a
    `to_text()` defect; `to_text()` on code blocks works.)
    
    ## Reference
    
    Selector aliases, the built-in function library, table-of-contents and
    transform recipes, in-place-edit caveats, and CLI flags live in
    [references/cheatsheet.md](references/cheatsheet.md). Read it before writing a
    non-trivial query — the dialect is jq-*like*, not jq, so the function names
    differ. When unsure of syntax, `mq repl` gives instant feedback.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related