Claude Skill

delegating-with-context

Delegates work to a subagent or spawned session by writing only the task, while a PreToolUse hook pages the parent transcript through Jev and appends the chunks that task needs. Use before writing any Agent tool prompt or create_session prompt, and when asked to delegate this, ha

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download oaustegard-claude-skills-plugins_ai-and-reasoning_skills_delegating-with-context-e39c726.zip · 15 KB
Part of oaustegard/claude-skills — 39 skills

Install

skills CLI npx skills add https://github.com/oaustegard/claude-skills/tree/main/plugins/ai-and-reasoning/skills/delegating-with-context
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install oaustegard-claude-skills@llmmart
Git git clone https://github.com/oaustegard/claude-skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole oaustegard/claude-skills collection as a plugin from our marketplace. Git is the plain clone.

README

delegating-with-context

Delegate by writing only the task. A PreToolUse hook reads the session transcript, asks Jev (TypeSafe) which chunks the task needs, and appends them to the subagent's prompt. See SKILL.md for the procedure and the hook entry, and oaustegard/experiments/subagent-context-filter for the eval.

Needs a Jev transport: CF_ACCOUNT_ID + CF_API_TOKEN (+ CF_GATEWAY_ID) for the Cloudflare AI Gateway with a stored TypeSafe key, or TYPESAFE_API_KEY.

Skill manifest

Delegating with context

Writing a background brief for a subagent costs output tokens, generation time, and fidelity: exact paths, error strings and numbers get paraphrased or dropped, and the brief omits what the orchestrator did not know mattered. The context_hook.py PreToolUse hook removes the need. It reads the session transcript, asks Jev which chunks the task needs, and appends those chunks to the subagent's prompt verbatim. The hook fires after the prompt is written, so it cannot save a brief you already wrote. The saving happens here, before the tool input exists.

Procedure

  1. Check the hook is live. test -s ~/.claude/context-filter-log.jsonl && tail -1 ~/.claude/context-filter-log.jsonl shows the last filtered delegation. No file on a first delegation is normal; the hook's reply after your first Agent call confirms it (step 5). Where the hook is not wired (claude.ai, a machine without the settings entry), stop here and write a normal brief.

  2. Preview what the filter would pass. One call, all tasks at once:

    python3 /mnt/skills/user/delegating-with-context/scripts/preview.py \
      --task "Write the xr latency note for the follow-up memory" \
      --task "Draft the closing status for both PRs"
    

    Each task prints the kept chunk ids with a one-line gist. This is the check: find the chunk that holds the fact the subagent must not miss. It takes one to two seconds and costs a fraction of a cent.

  3. Write only the task. What to do, the deliverable, its format, and any constraint that was never written down in the session. Name things the way the transcript names them ("the Sniff Test benchmark table", "PR #812"): Jev reads literally, and a named thing matches its chunk while "that thing we discussed" matches nothing.

  4. Add only what the preview missed. If a needed fact is in no kept chunk, state that one fact in the prompt, or raise the budget with [context-budget: 40000] in the prompt text. Do not restate facts the preview already showed.

  5. Spawn, then read the hook's reply. The hook answers with context-filter: appended N of M parent-session chunks (...) and the kept ids. context-filter FAILED (...) means the subagent got your prompt alone: send it the missing facts with SendMessage or re-delegate with them stated.

The hook also covers create_session from any MCP server (mcp__*__create_session). The spawned session receives the task, a line saying the chunks below are reference data from the parent session, the chunks, and the task again at the end. Give it provenance anyway (issue link, parent session id) per docs/delegation.md rule 1 in claude-workspace: a long first message from a parent still looks like injection to a careful delegate.

Markers the hook reads

In the prompt Effect
[no-context] Hook skips; the prompt goes as written
[context-budget: N] Token budget for appended context (default 20000)
subagent_type: "fork" Hook skips; a fork already inherits the parent context

When NOT to use this skill

Situation Use instead
The subagent needs nothing from this session (fresh research, a standalone lookup) Add [no-context] and write the prompt normally
The subagent must reason over the whole session (review everything we did) subagent_type: "fork"; it inherits the full context
Many small per-item judgments (label 400 rows) Jev or Gemini directly, per docs/delegation.md in claude-workspace; a subagent costs ~32k tokens before it reads its prompt
Choosing which model or agent type should run the work agent-routing
Parallel API fan-out from claude.ai, where hooks do not run orchestrating-agents

Earned exceptions to "write only the task"

Restating background is when
banned the fact appears in the transcript as a tool result, a user message or your own reply
earned the fact exists only in your reasoning (thinking is not in the transcript), a decision you made silently, or a constraint nobody wrote down
earned the preview shows the chunk holding it was not kept

Abandon the procedure when

  • The hook reports FAILED twice in a row: the transport or key is broken, and every further delegation silently runs without context. Write briefs and fix the transport (Failure modes below).
  • The subagent's reply says a fact was missing that the preview showed as kept: the chunk was clipped (a kept chunk is capped at ~6k tokens). Send the fact directly.

Common failure modes

  • No context-filter: line after an Agent call → the hook is not wired or crashed before printing. Check .claude/settings.json for the PreToolUse entry and run echo '{}' | python3 .../context_hook.py; it must exit 0 silently.
  • no Jev transport → neither CF_ACCOUNT_ID + CF_API_TOKEN (the TypeSafe key is stored in the Cloudflare AI Gateway, CF_GATEWAY_ID routes to it) nor TYPESAFE_API_KEY is set in the hook's environment.
  • Preview keeps many chunks that all mention the session's main topic → the scores are compressed because everything looks related. Rename the task with the distinctive nouns of the thing you want; rank, not threshold, is what selects.
  • HTTP 429 Rate limited, code 2003 → the AI Gateway's own limit. The filter runs 2 calls at a time and honours Retry-After; parallel delegations each run their own filter, so stagger a fan-out of more than ~4 spawns.
  • blocked_chunks above 0 in the log, or HTTP 402 "Payment error from model using BYOK" → not billing. TypeSafe's edge WAF rejected the request bytes (shell commands and curl lines in a transcript read as attack payloads). The filter bisects the window and scores each blocked chunk 0.5, so it can still be kept on rank. If a chunk you need is blocked, state its fact in the prompt.
  • A kept chunk states something a later chunk corrected → each chunk is scored on its own merits. User messages are force-kept (clipped) because they carry corrections; if a correction lives in a tool result, state it.

Verification

After the subagent returns: its reply must not say a needed fact was missing, and tail -1 ~/.claude/context-filter-log.jsonl must show the delegation with kept_ids and no error key. A bad success looks like a fluent deliverable with a plausible but wrong number: compare one specific figure against the chunk it came from before relaying it.

Diagnosed failures

  • 2026-09-22: the first selection rule force-kept every user-role message. Skill bodies and stop-hook feedback are user-role (isMeta) turns, and on a 181-chunk session they filled the whole 20k budget before any tool result was considered; the chunk holding the answer was dropped. Fixed: isMeta turns are scored like tool results, and real user messages are clipped and capped at a quarter of the budget.

  • 2026-09-22: a single condensed window (every chunk clipped to ~1k chars) ranked the key chunk 6th because the fact sat mid-result, outside the clip. Paging at 4x view (five ~27k windows, 1.6 s) ranked it 1st. The filter defaults to the paged view.

  • 2026-09-23, eval over 32 delegations from 8 archived sessions: written briefs lost every fact on 5 tasks by telling the subagent to look the fact up ("check the pr-workflow config entry") instead of stating it; neither the filter nor the full transcript ever did. That is what step 3 prevents.

  • 2026-09-23, first real create_session probes: two Haiku cloud sessions received the appended context (their first-turn token counts match) and stalled at need-input, one saying the message looked truncated, the other that no request had arrived. The context ended the message and the task sat above ~20k tokens of transcript. Labelling the block as reference data and repeating the task after it, the next probe answered correctly first time.

Scripts

  • scripts/chunking.py: transcript JSONL to chunks (user, assistant, harness, tool call + result), with secret values and token patterns redacted before anything leaves the machine.
  • scripts/jevfilter.py: windows under Jev's limits (32k state plus longest question, 64k state plus all questions), one Noul per chunk, rank into the budget.
  • scripts/context_hook.py: the PreToolUse hook. Fails open and says so.
  • scripts/preview.py: step 2.

Install as a plugin (the hook wires itself; hooks/hooks.json):

claude plugin marketplace add oaustegard/claude-skills
claude plugin install delegating-with-context@oaustegard-claude-skills

Or wire it by hand in .claude/settings.json, guarded so a missing file exits 0 (python3 on a missing path exits 2, which blocks the tool call):

{"matcher": "Agent|Task|mcp__.*__create_session", "hooks": [{"type": "command", "timeout": 120,
  "command": "f=/mnt/skills/user/delegating-with-context/scripts/context_hook.py; test -f \"$f\" || exit 0; exec python3 \"$f\""}]}

Evidence: with ~20k tokens of selected chunks a fresh subagent found 99/105 required facts, against 102/105 with the whole ~81k-token transcript and 82/105 with a written brief; method and caveats in oaustegard/experiments/subagent-context-filter/RESULTS.md.

Files (claude-skills)
  • hooks
    • hooks.json 346 B
      {
        "hooks": {
          "PreToolUse": [
            {
              "matcher": "Agent|Task|mcp__.*__create_session",
              "hooks": [
                {
                  "type": "command",
                  "timeout": 120,
                  "command": "python3 \"${CLAUDE_PLUGIN_ROOT}/skills/delegating-with-context/scripts/context_hook.py\""
                }
              ]
            }
          ]
        }
      }
      
  • scripts
    • chunking.py 4.6 KB
      """Split a Claude Code session transcript (JSONL) into context chunks and scrub secrets.
      
      A chunk is one user message, one assistant text block, or one tool call paired
      with its result. Harness-injected user text (starting with '<') is skipped.
      """
      import glob, json, os, re
      
      SECRET_PATTERNS = [re.compile(p) for p in (
          r"ghp_[A-Za-z0-9]{36}", r"github_pat_[A-Za-z0-9_]{20,}", r"gh[osu]_[A-Za-z0-9]{36}",
          r"sk-[A-Za-z0-9_\-]{20,}", r"AKIA[0-9A-Z]{16}", r"eyJ[\w-]{10,}\.[\w-]{10,}\.[\w-]{10,}",
          r"xox[abp]-[A-Za-z0-9-]{10,}")]
      SECRET_NAME = re.compile(r"TOKEN|KEY|SECRET|PASSWORD|PASSWD|_PAT\b|CREDENTIAL", re.I)
      
      
      def _secret_values():
          vals = {v for k, v in os.environ.items() if SECRET_NAME.search(k) and len(v) >= 12}
          for f in glob.glob("/mnt/project/*.env"):
              for line in open(f, errors="ignore"):
                  k, _, v = line.strip().removeprefix("export ").partition("=")
                  v = v.strip().strip("'\"")
                  if SECRET_NAME.search(k) and len(v) >= 12:
                      vals.add(v)
          return sorted(vals, key=len, reverse=True)
      
      
      _VALUES = None
      
      
      def scrub(text):
          """Return (clean_text, n_replacements). Never prints or returns a secret."""
          global _VALUES
          if _VALUES is None:
              _VALUES = _secret_values()
          n = 0
          for v in _VALUES:
              if v in text:
                  n += text.count(v); text = text.replace(v, "[REDACTED]")
          for p in SECRET_PATTERNS:
              text, k = p.subn("[REDACTED]", text); n += k
          return text, n
      
      
      def _text(content):
          if isinstance(content, str):
              return content
          out = []
          for b in content or []:
              if not isinstance(b, dict):
                  continue
              if b.get("type") == "text":
                  out.append(b.get("text", ""))
              elif b.get("type") == "tool_result":
                  c = b.get("content")
                  out.append(c if isinstance(c, str) else " ".join(x.get("text", "") for x in c or [] if isinstance(x, dict)))
          return "\n".join(out)
      
      
      def chunk_transcript(path):
          """Return (chunks, n_scrubbed). Each chunk: id, kind, text/tool/input/result, prompt, chars."""
          chunks, pending, prompt, scrubbed = [], {}, "", 0
          for line in open(path, errors="ignore"):
              if not line.strip():
                  continue
              try:
                  r = json.loads(line)
              except json.JSONDecodeError:
                  continue
              t, c = r.get("type"), (r.get("message") or {}).get("content")
              if t == "user":
                  blocks = [{"type": "text", "text": c}] if isinstance(c, str) else (c or [])
                  for b in blocks:
                      if not isinstance(b, dict):
                          continue
                      if b.get("type") == "tool_result" and b.get("tool_use_id") in pending:
                          ch = pending.pop(b["tool_use_id"]); ch["result"] = _text([b]); chunks.append(ch)
                      elif b.get("type") == "text" and b.get("text", "").strip() and not b["text"].lstrip().startswith("<"):
                          # isMeta marks harness-injected user turns: skill bodies, stop-hook
                          # feedback, peer messages. They are scored like any chunk but are
                          # not the user's words, so they are never force-kept.
                          if r.get("isMeta"):
                              chunks.append({"kind": "harness", "text": b["text"], "prompt": prompt})
                              continue
                          prompt = b["text"][:400]
                          chunks.append({"kind": "user", "text": b["text"], "prompt": prompt})
              elif t == "assistant" and isinstance(c, list):
                  for b in c:
                      if b.get("type") == "text" and b.get("text", "").strip():
                          chunks.append({"kind": "assistant", "text": b["text"], "prompt": prompt})
                      elif b.get("type") == "tool_use":
                          pending[b["id"]] = {"kind": "tool", "tool": b.get("name", "?"),
                                              "input": json.dumps(b.get("input", {}), ensure_ascii=False), "prompt": prompt}
          for i, ch in enumerate(chunks):
              ch["id"] = i
              for k in ("text", "input", "result", "prompt"):
                  if k in ch:
                      ch[k], n = scrub(ch[k]); scrubbed += n
              ch["chars"] = sum(len(ch.get(k, "")) for k in ("text", "input", "result"))
          return chunks, scrubbed
      
      
      def render(chunk):
          """Full-fidelity text of a chunk, as a subagent would receive it."""
          if chunk["kind"] == "tool":
              return f"[{chunk['id']}] TOOL {chunk['tool']} input: {chunk['input']}\nresult: {chunk.get('result', '')}"
          who = {"user": "USER", "harness": "HARNESS"}.get(chunk["kind"], "ASSISTANT")
          return f"[{chunk['id']}] {who}: {chunk['text']}"
      
    • context_hook.py 4.6 KB
      #!/usr/bin/env python3
      """PreToolUse hook: append Jev-selected parent-session context to a subagent prompt.
      
      Wired on the Agent tool and on create_session (any MCP server's). The
      orchestrator writes only the task; this hook reads the session transcript,
      asks Jev which chunks the task needs, and appends them to the prompt via
      updatedInput.
      
      Skips: forks (they inherit the parent context already), prompts containing
      [no-context], and transcripts too short to be worth filtering.
      Budget: [context-budget: N] in the prompt overrides the default token budget.
      
      Fails open: on any error the tool call proceeds with the original prompt, and
      the failure is reported back to the model as additionalContext so it knows the
      subagent got no context. Exit code is always 0.
      """
      import json
      import os
      import re
      import sys
      import time
      
      sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
      
      DEFAULT_BUDGET = int(os.environ.get("CONTEXT_FILTER_BUDGET", "20000"))
      MIN_CHUNKS = int(os.environ.get("CONTEXT_FILTER_MIN_CHUNKS", "6"))
      LONG_PROMPT_CHARS = 4_000    # a prompt this long probably restates background
      LOG = os.environ.get("CONTEXT_FILTER_LOG", os.path.expanduser("~/.claude/context-filter-log.jsonl"))
      
      
      def emit(extra_context=None, updated_input=None):
          out = {"hookEventName": "PreToolUse"}
          if updated_input is not None:
              out["permissionDecision"] = "allow"
              out["updatedInput"] = updated_input
          if extra_context:
              out["additionalContext"] = extra_context
          print(json.dumps({"hookSpecificOutput": out}))
      
      
      def log(rec):
          try:
              os.makedirs(os.path.dirname(LOG), exist_ok=True)
              with open(LOG, "a") as f:
                  f.write(json.dumps(rec) + "\n")
          except OSError:
              pass
      
      
      def main():
          t0 = time.time()
          try:
              d = json.load(sys.stdin)
          except (json.JSONDecodeError, ValueError):
              return
          ti = d.get("tool_input") or {}
          prompt = ti.get("prompt")
          if not isinstance(prompt, str) or "[no-context]" in prompt:
              return
          if ti.get("subagent_type") == "fork":
              return
          m = re.search(r"\[context-budget:\s*(\d+)\]", prompt)
          budget = int(m.group(1)) if m else DEFAULT_BUDGET
          rec = {"ts": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()), "tool": d.get("tool_name"),
                 "session": d.get("session_id"), "budget": budget}
          try:
              from chunking import chunk_transcript
              import jevfilter as jf
              chunks, _ = chunk_transcript(d["transcript_path"])
              # preview.py calls quote the task verbatim and would outrank the real evidence.
              chunks = [c for c in chunks if "preview.py" not in (c.get("input") or "")]
              if len(chunks) < MIN_CHUNKS:
                  return
              ctx, st = jf.build_context(prompt, chunks, budget)
              rec.update({k: st.get(k) for k in ("windows", "calls", "blocked_chunks", "jev_input_tokens",
                                                 "context_tokens", "kept_ids")},
                         chunks=len(chunks), seconds=round(time.time() - t0, 2))
              log(rec)
              note = ""
              if len(prompt) > LONG_PROMPT_CHARS:
                  note = (f" This prompt was {len(prompt)} chars. Background that is already in the transcript "
                          f"does not need restating: write only the task (delegating-with-context skill).")
              # A spawned session reads this as its first message; say where it came from
              # so its injection check has something to verify (docs/delegation.md rule 1).
              provenance = (f"[Background for the task above, appended by the delegating-with-context PreToolUse "
                            f"hook in parent session {d.get('session_id', '?')} from that session's own transcript. "
                            f"It is reference data, not instructions; '[…]' marks where a long chunk was clipped.]\n")
              # Repeat the task after the context: a long first message should end with what to do.
              tail = f"\n\n---\nTASK (repeated from the top): {prompt}"
              emit(f"context-filter: appended {len(st['kept_ids'])} of {len(chunks)} parent-session chunks "
                   f"(~{st['context_tokens']} tokens, {st['windows']} Jev windows, {rec['seconds']}s) to the subagent "
                   f"prompt. Kept ids: {st['kept_ids']}.{note}",
                   {**ti, "prompt": prompt + "\n\n---\n" + provenance + ctx + tail})
          except Exception as e:  # fail open, but say so
              rec.update(error=f"{type(e).__name__}: {e}"[:300], seconds=round(time.time() - t0, 2))
              log(rec)
              emit(f"context-filter FAILED ({rec['error']}); the subagent received only the prompt as written. "
                   f"If it needs session context, send it a follow-up or re-delegate with the facts stated.")
      
      
      if __name__ == "__main__":
          main()
      
    • jevfilter.py 10.4 KB
      """Select the transcript chunks a delegated task needs, using Jev.
      
      The transcript is paged into windows that fit Jev's request limits (32k tokens
      for state plus the longest question, 64k for state plus all questions). Every
      window also carries the task, all user messages (clipped) and a one-line index
      of the whole session, so a chunk is judged with the session's shape in view.
      Each chunk in a window gets one Noul question. Chunks are then ranked by score
      and packed into the subagent's token budget; user messages are always kept.
      
      Transport: the Cloudflare AI Gateway (CF_ACCOUNT_ID, CF_GATEWAY_ID,
      CF_API_TOKEN; the TypeSafe key is stored in the gateway) or TypeSafe directly
      (TYPESAFE_API_KEY).
      
      A failed window raises FilterError. Callers decide what to fall back to; the
      hook passes the original prompt through unchanged.
      """
      from __future__ import annotations
      
      import concurrent.futures as cf
      import json
      import os
      import time
      import urllib.error
      import urllib.request
      
      CHARS_PER_TOKEN = 3.3        # conservative; Jev reported ~3.8 on transcript text
      STATE_TOKENS = 27_000        # per window, under the 32k state + longest-question limit
      GLOBAL_TOKENS = 5_000        # task + user messages + index, repeated in every window
      MAX_CHUNK_TOKENS = 6_000     # a single kept chunk is clipped to this in the output
      SCORE_FLOOR = 0.15           # never pad the budget with chunks below this
      CONCURRENCY = 2              # METHODS.md: the CF AI Gateway throttles hard; start at 2
      
      
      class FilterError(RuntimeError):
          pass
      
      
      class Blocked(FilterError):
          """TypeSafe's edge WAF rejected the request content. Deterministic: retrying the
          same bytes never helps. Through the AI Gateway it arrives as HTTP 402 'Payment
          error from model using BYOK' wrapping a Cloudflare 'Sorry, you have been blocked'
          page (diagnosed 2026-09-23 on a transcript full of shell commands)."""
      
      
      def tokens(s: str) -> int:
          return int(len(s) / CHARS_PER_TOKEN) + 1
      
      
      def clip(s: str | None, n: int) -> str:
          s = s or ""
          return s if len(s) <= n else s[: int(n * 0.75)] + " […] " + s[-int(n * 0.25):]
      
      
      def view(c: dict, scale: float = 1.0) -> dict:
          """Condensed chunk for Jev's state."""
          v = {"kind": c["kind"]}
          if c["kind"] == "tool":
              v["tool"] = c.get("tool", "?")
              v["input"] = clip(c.get("input"), int(300 * scale))
              v["result"] = clip(c.get("result"), int(900 * scale))
          else:
              v["text"] = clip(c.get("text"), int(1500 * scale))
          return v
      
      
      def render(c: dict, max_chars: int = int(MAX_CHUNK_TOKENS * CHARS_PER_TOKEN)) -> str:
          """Full-fidelity chunk text as the subagent receives it."""
          if c["kind"] == "tool":
              body = f"TOOL {c.get('tool', '?')} input: {c.get('input', '')}\nresult: {c.get('result', '')}"
          else:
              body = {"user": "USER: ", "harness": "HARNESS: "}.get(c["kind"], "ASSISTANT: ") + c.get("text", "")
          return f"[{c['id']}] " + clip(body, max_chars)
      
      
      def _global_state(task: str, chunks: list[dict]) -> dict:
          users = [clip(c["text"], 400) for c in chunks if c["kind"] == "user"]
          index = [f"c{c['id']} {c['kind']}{' ' + c.get('tool', '') if c['kind'] == 'tool' else ''}: "
                   + clip((c.get("text") or c.get("input") or "").replace("\n", " "), 70) for c in chunks]
          g = {"task": task, "user_messages": users, "session_index": index}
          while tokens(json.dumps(g)) > GLOBAL_TOKENS and len(g["session_index"]) > 0:
              g["session_index"] = [clip(x, max(20, len(x) // 2)) for x in g["session_index"]]
              if tokens(json.dumps(g)) > GLOBAL_TOKENS:
                  g["user_messages"] = [clip(u, 150) for u in g["user_messages"]]
              if all(len(x) <= 24 for x in g["session_index"]):
                  break
          return g
      
      
      def _question(cid: int) -> dict:
          return {"type": "noul",
                  "instructions": f"Does `chunks.c{cid}` contain information that an assistant given `task` "
                                  f"would need or benefit from to do that task?"}
      
      
      def windows(task: str, chunks: list[dict], scale: float = 1.0) -> list[dict]:
          """Pack condensed chunks into states that fit STATE_TOKENS."""
          g = _global_state(task, chunks)
          room = STATE_TOKENS - tokens(json.dumps(g))
          out, cur, used = [], {}, 0
          for c in chunks:
              v = view(c, scale)
              t = tokens(json.dumps(v)) + 8
              if cur and used + t > room:
                  out.append(cur); cur, used = {}, 0
              cur[f"c{c['id']}"] = v; used += t
          if cur:
              out.append(cur)
          return [{"state": {**g, "chunks": w}, "questions": {k: _question(int(k[1:])) for k in w}} for w in out]
      
      
      def _call(state: dict, questions: dict, timeout: int = 60) -> tuple[dict, int]:
          if os.environ.get("CF_API_TOKEN") and os.environ.get("CF_ACCOUNT_ID"):
              url = f"https://api.cloudflare.com/client/v4/accounts/{os.environ['CF_ACCOUNT_ID']}/ai/run"
              body = {"model": "typesafe/jev", "input": {"state": state, "questions": questions}}
              headers = {"Authorization": f"Bearer {os.environ['CF_API_TOKEN']}", "Content-Type": "application/json"}
              if os.environ.get("CF_GATEWAY_ID"):
                  headers["cf-aig-gateway-id"] = os.environ["CF_GATEWAY_ID"]
              unwrap = lambda r: r["result"]["result"]
          elif os.environ.get("TYPESAFE_API_KEY"):
              url = "https://api.typesafe.ai/v1/systemone"
              body = {"model": "jev-latest", "state": state, "questions": questions}
              headers = {"Authorization": f"Bearer {os.environ['TYPESAFE_API_KEY']}", "Content-Type": "application/json"}
              unwrap = lambda r: r
          else:
              raise FilterError("no Jev transport: set CF_ACCOUNT_ID+CF_API_TOKEN(+CF_GATEWAY_ID) or TYPESAFE_API_KEY")
          req = urllib.request.Request(url, data=json.dumps(body).encode(), headers=headers, method="POST")
          last = None
          for attempt in range(6):
              try:
                  with urllib.request.urlopen(req, timeout=timeout) as r:
                      res = unwrap(json.load(r))
                  return {k: a["noul"] for k, a in res["answers"].items()}, res.get("usage", {}).get("input_tokens", 0)
              except urllib.error.HTTPError as e:
                  body = e.read().decode(errors="replace")
                  if "you have been blocked" in body or "Attention Required" in body:
                      raise Blocked(f"HTTP {e.code}: request content blocked by the provider's WAF")
                  last = f"HTTP {e.code}: {body[:300]}"
                  # the gateway's own rate limit (code 2003) needs real backoff; honour Retry-After
                  wait = float(e.headers.get("Retry-After") or 0) or min(30, 2 * 2 ** attempt)
                  time.sleep(wait)
              except (urllib.error.URLError, KeyError, ValueError, TimeoutError) as e:
                  last = f"{type(e).__name__}: {str(e)[:200]}"
                  time.sleep(1.5 * 2 ** attempt)
          raise FilterError(f"Jev call failed after retries: {last}")
      
      
      BLOCKED_SCORE = 0.5          # a chunk the WAF will not let Jev see is ranked as a coin flip
      
      
      def _score_window(w: dict, depth: int = 0) -> tuple[dict, int, int, int]:
          """Score one window; on a WAF block, bisect it until the offending chunks are isolated.
      
          Returns (scores, jev_tokens, calls, blocked_chunks)."""
          try:
              ans, u = _call(w["state"], w["questions"])
              return {int(k[1:]): v for k, v in ans.items()}, u, 1, 0
          except Blocked:
              keys = list(w["state"]["chunks"])
              if len(keys) == 1 or depth >= 6:
                  return {int(k[1:]): BLOCKED_SCORE for k in keys}, 0, 1, len(keys)
              out, used, calls, blocked = {}, 0, 1, 0
              for half in (keys[: len(keys) // 2], keys[len(keys) // 2:]):
                  sub = {"state": {**w["state"], "chunks": {k: w["state"]["chunks"][k] for k in half}},
                         "questions": {k: w["questions"][k] for k in half}}
                  s, u, c, b = _score_window(sub, depth + 1)
                  out.update(s); used += u; calls += c; blocked += b
              return out, used, calls, blocked
      
      
      def score(task: str, chunks: list[dict], scale: float = 1.0) -> tuple[dict[int, float], dict]:
          """Return ({chunk_id: P(needed)}, stats)."""
          t0 = time.time()
          ws = windows(task, chunks, scale)
          scores, used, calls, blocked = {}, 0, 0, 0
          with cf.ThreadPoolExecutor(CONCURRENCY) as ex:
              for s, u, c, b in ex.map(_score_window, ws):
                  scores.update(s); used += u; calls += c; blocked += b
          return scores, {"windows": len(ws), "calls": calls, "blocked_chunks": blocked,
                          "jev_input_tokens": used, "seconds": round(time.time() - t0, 2)}
      
      
      USER_CLIP_CHARS = 1_500      # force-kept user messages are clipped to this
      USER_SHARE = 0.25            # and never take more than this share of the budget
      
      
      def select(chunks: list[dict], scores: dict[int, float], budget_tokens: int) -> dict[int, int]:
          """Return {chunk_id: max_chars} to render.
      
          User messages are kept first, clipped, newest first, up to USER_SHARE of the
          budget: they carry intent and corrections. Everything else, including a user
          message whose full text scores high, competes on score for the rest.
          """
          keep, spent = {}, 0
          for c in sorted((c for c in chunks if c["kind"] == "user"), key=lambda c: -c["id"]):
              t = tokens(render(c, USER_CLIP_CHARS))
              if spent + t > budget_tokens * USER_SHARE:
                  break
              keep[c["id"]] = USER_CLIP_CHARS; spent += t
          full = int(MAX_CHUNK_TOKENS * CHARS_PER_TOKEN)
          for c in sorted(chunks, key=lambda c: -scores.get(c["id"], 0)):
              s = scores.get(c["id"], 0)
              if s < SCORE_FLOOR:
                  break
              if keep.get(c["id"]) == full:
                  continue
              had = tokens(render(c, keep[c["id"]])) if c["id"] in keep else 0
              t = tokens(render(c, full)) - had
              if spent + t <= budget_tokens:
                  keep[c["id"]] = full; spent += t
          return keep
      
      
      def build_context(task: str, chunks: list[dict], budget_tokens: int = 20_000,
                        scale: float = 4.0) -> tuple[str, dict]:
          """scale multiplies how much of each chunk Jev sees; larger scale means more windows."""
          scores, stats = score(task, chunks, scale)
          keep = select(chunks, scores, budget_tokens)
          kept = [c for c in chunks if c["id"] in keep]
          body = "\n\n".join(render(c, keep[c["id"]]) for c in kept)
          header = (f"Context from the parent session, selected for this task: {len(kept)} of {len(chunks)} "
                    f"transcript chunks (~{tokens(body)} tokens; ids in brackets). Chunks not shown were judged "
                    f"irrelevant; if something you need is missing, say so in your reply.")
          stats.update({"kept_ids": [c["id"] for c in kept], "context_tokens": tokens(body), "scores": scores})
          return header + "\n\n" + body, stats
      
    • preview.py 2.5 KB
      #!/usr/bin/env python3
      """Show which parent-session chunks the context filter would pass for a task.
      
          python3 preview.py --task "Write the xr latency note for the follow-up memory"
          python3 preview.py --task "..." --task "..." [--budget 20000] [--transcript PATH]
      
      Reads the current session transcript (newest .jsonl under ~/.claude/projects
      for this cwd unless --transcript is given), runs the same Jev filter the
      PreToolUse hook runs, and prints one line per kept chunk. Exit 1 on a filter
      failure, with the reason on stderr: the hook would then pass the prompt through
      with no context.
      """
      import argparse
      import glob
      import json
      import os
      import sys
      
      sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
      from chunking import chunk_transcript  # noqa: E402
      import jevfilter as jf  # noqa: E402
      
      
      def current_transcript(cwd: str) -> str | None:
          slug = cwd.replace("/", "-")
          files = glob.glob(os.path.expanduser(f"~/.claude/projects/{slug}/*.jsonl"))
          return max(files, key=os.path.getmtime) if files else None
      
      
      def line(c: dict, score: float | None) -> str:
          what = c.get("tool") if c["kind"] == "tool" else c["kind"]
          text = (c.get("text") or c.get("input") or "").replace("\n", " ")
          s = "  --" if score is None else f"{score:4.2f}"
          return f"  [{c['id']:>3}] {s} {what:<14} {jf.clip(text, 90)}"
      
      
      def main() -> int:
          ap = argparse.ArgumentParser()
          ap.add_argument("--task", action="append", required=True)
          ap.add_argument("--budget", type=int, default=20_000)
          ap.add_argument("--transcript")
          a = ap.parse_args()
          path = a.transcript or current_transcript(os.getcwd())
          if not path:
              print("no transcript found for this cwd; pass --transcript", file=sys.stderr)
              return 1
          chunks, _ = chunk_transcript(path)
          # Earlier preview calls quote the task verbatim and would rank first.
          chunks = [c for c in chunks if "preview.py" not in (c.get("input") or "")]
          rc = 0
          for task in a.task:
              try:
                  ctx, st = jf.build_context(task, chunks, a.budget)
              except jf.FilterError as e:
                  print(f"FILTER FAILED for {task[:60]!r}: {e}", file=sys.stderr)
                  rc = 1
                  continue
              print(f"TASK: {task}\n  keeps {len(st['kept_ids'])} of {len(chunks)} chunks, ~{st['context_tokens']} tokens, "
                    f"{st['windows']} window(s), {st['seconds']}s")
              by_id = {c["id"]: c for c in chunks}
              for i in st["kept_ids"]:
                  print(line(by_id[i], st["scores"].get(i)))
          return rc
      
      
      if __name__ == "__main__":
          sys.exit(main())
      
  • CHANGELOG.md 1009 B
    # delegating-with-context - Changelog
    
    ## 0.2.0 — 2026-09-23
    - Covers `create_session`: the hook matches `mcp__*__create_session` and
      labels the appended block as reference data from the parent session and
      repeats the task after it (two probes stalled without that; the third
      answered). `updatedInput` on an MCP
      tool verified with a stub MCP server and on the real remote server.
    - Ships as its own plugin (`hooks/hooks.json`, `${CLAUDE_PLUGIN_ROOT}` paths);
      `registry/generate.py` now builds a standalone plugin for any skill with hooks.
    
    ## 0.1.0 — 2026-09-23
    - New skill. Write only the task when delegating; `scripts/context_hook.py`
      (PreToolUse on Agent) pages the parent transcript through Jev and appends the
      chunks the task needs. `scripts/preview.py` shows the selection before the
      prompt is written. Measured in `oaustegard/experiments/subagent-context-filter`.
    
    ## [0.1.0] - 2026-09-23
    
    ### Other
    
    - delegating-with-context 0.1.0: write only the task, a hook passes the context
    
  • README.md 480 B
    # delegating-with-context
    
    Delegate by writing only the task. A PreToolUse hook reads the session
    transcript, asks Jev (TypeSafe) which chunks the task needs, and appends them
    to the subagent's prompt. See `SKILL.md` for the procedure and the hook entry,
    and `oaustegard/experiments/subagent-context-filter` for the eval.
    
    Needs a Jev transport: `CF_ACCOUNT_ID` + `CF_API_TOKEN` (+ `CF_GATEWAY_ID`) for
    the Cloudflare AI Gateway with a stored TypeSafe key, or `TYPESAFE_API_KEY`.
    
  • SKILL.md 9.6 KB
    ---
    name: delegating-with-context
    description: Delegates work to a subagent or spawned session by writing only the task, while a PreToolUse hook pages the parent transcript through Jev and appends the chunks that task needs. Use before writing any Agent tool prompt or create_session prompt, and when asked to delegate this, hand this off, use a subagent, spawn an agent, run agents in parallel, or brief a subagent on what we have done so far.
    metadata:
      version: 0.2.0
    ---
    
    # Delegating with context
    
    Writing a background brief for a subagent costs output tokens, generation
    time, and fidelity: exact paths, error strings and numbers get paraphrased or
    dropped, and the brief omits what the orchestrator did not know mattered. The
    `context_hook.py` PreToolUse hook removes the need. It reads the session
    transcript, asks Jev which chunks the task needs, and appends those chunks to
    the subagent's prompt verbatim. The hook fires after the prompt is written, so
    it cannot save a brief you already wrote. The saving happens here, before the
    tool input exists.
    
    ## Procedure
    
    1. **Check the hook is live.** `test -s ~/.claude/context-filter-log.jsonl && tail -1 ~/.claude/context-filter-log.jsonl`
       shows the last filtered delegation. No file on a first delegation is normal;
       the hook's reply after your first Agent call confirms it (step 5). Where the
       hook is not wired (claude.ai, a machine without the settings entry), stop
       here and write a normal brief.
    
    2. **Preview what the filter would pass.** One call, all tasks at once:
    
       ```bash
       python3 /mnt/skills/user/delegating-with-context/scripts/preview.py \
         --task "Write the xr latency note for the follow-up memory" \
         --task "Draft the closing status for both PRs"
       ```
    
       Each task prints the kept chunk ids with a one-line gist. This is the
       check: find the chunk that holds the fact the subagent must not miss. It
       takes one to two seconds and costs a fraction of a cent.
    
    3. **Write only the task.** What to do, the deliverable, its format, and any
       constraint that was never written down in the session. Name things the way
       the transcript names them ("the Sniff Test benchmark table", "PR #812"):
       Jev reads literally, and a named thing matches its chunk while "that thing
       we discussed" matches nothing.
    
    4. **Add only what the preview missed.** If a needed fact is in no kept chunk,
       state that one fact in the prompt, or raise the budget with
       `[context-budget: 40000]` in the prompt text. Do not restate facts the
       preview already showed.
    
    5. **Spawn, then read the hook's reply.** The hook answers with
       `context-filter: appended N of M parent-session chunks (...)` and the kept
       ids. `context-filter FAILED (...)` means the subagent got your prompt alone:
       send it the missing facts with SendMessage or re-delegate with them stated.
    
    The hook also covers `create_session` from any MCP server
    (`mcp__*__create_session`). The spawned session receives the task, a line
    saying the chunks below are reference data from the parent session, the
    chunks, and the task again at the end. Give it
    provenance anyway (issue link, parent session id) per `docs/delegation.md`
    rule 1 in claude-workspace: a long first message from a parent still looks like
    injection to a careful delegate.
    
    ## Markers the hook reads
    
    | In the prompt | Effect |
    |---|---|
    | `[no-context]` | Hook skips; the prompt goes as written |
    | `[context-budget: N]` | Token budget for appended context (default 20000) |
    | `subagent_type: "fork"` | Hook skips; a fork already inherits the parent context |
    
    ## When NOT to use this skill
    
    | Situation | Use instead |
    |---|---|
    | The subagent needs nothing from this session (fresh research, a standalone lookup) | Add `[no-context]` and write the prompt normally |
    | The subagent must reason over the whole session (review everything we did) | `subagent_type: "fork"`; it inherits the full context |
    | Many small per-item judgments (label 400 rows) | Jev or Gemini directly, per `docs/delegation.md` in claude-workspace; a subagent costs ~32k tokens before it reads its prompt |
    | Choosing which model or agent type should run the work | `agent-routing` |
    | Parallel API fan-out from claude.ai, where hooks do not run | `orchestrating-agents` |
    
    ## Earned exceptions to "write only the task"
    
    | Restating background is | when |
    |---|---|
    | banned | the fact appears in the transcript as a tool result, a user message or your own reply |
    | earned | the fact exists only in your reasoning (thinking is not in the transcript), a decision you made silently, or a constraint nobody wrote down |
    | earned | the preview shows the chunk holding it was not kept |
    
    ## Abandon the procedure when
    
    - The hook reports FAILED twice in a row: the transport or key is broken, and
      every further delegation silently runs without context. Write briefs and fix
      the transport (Failure modes below).
    - The subagent's reply says a fact was missing that the preview showed as
      kept: the chunk was clipped (a kept chunk is capped at ~6k tokens). Send the
      fact directly.
    
    ## Common failure modes
    
    - **No `context-filter:` line after an Agent call** → the hook is not wired or
      crashed before printing. Check `.claude/settings.json` for the PreToolUse
      entry and run `echo '{}' | python3 .../context_hook.py`; it must exit 0 silently.
    - **`no Jev transport`** → neither `CF_ACCOUNT_ID` + `CF_API_TOKEN` (the
      TypeSafe key is stored in the Cloudflare AI Gateway, `CF_GATEWAY_ID` routes
      to it) nor `TYPESAFE_API_KEY` is set in the hook's environment.
    - **Preview keeps many chunks that all mention the session's main topic** →
      the scores are compressed because everything looks related. Rename the task
      with the distinctive nouns of the thing you want; rank, not threshold, is
      what selects.
    - **HTTP 429 `Rate limited`, code 2003** → the AI Gateway's own limit. The
      filter runs 2 calls at a time and honours Retry-After; parallel delegations
      each run their own filter, so stagger a fan-out of more than ~4 spawns.
    - **`blocked_chunks` above 0 in the log, or HTTP 402 "Payment error from model
      using BYOK"** → not billing. TypeSafe's edge WAF rejected the request bytes
      (shell commands and curl lines in a transcript read as attack payloads). The
      filter bisects the window and scores each blocked chunk 0.5, so it can still
      be kept on rank. If a chunk you need is blocked, state its fact in the prompt.
    - **A kept chunk states something a later chunk corrected** → each chunk is
      scored on its own merits. User messages are force-kept (clipped) because
      they carry corrections; if a correction lives in a tool result, state it.
    
    ## Verification
    
    After the subagent returns: its reply must not say a needed fact was missing,
    and `tail -1 ~/.claude/context-filter-log.jsonl` must show the delegation with
    `kept_ids` and no `error` key. A bad success looks like a fluent deliverable
    with a plausible but wrong number: compare one specific figure against the
    chunk it came from before relaying it.
    
    ## Diagnosed failures
    
    - 2026-09-22: the first selection rule force-kept every user-role message.
      Skill bodies and stop-hook feedback are user-role (`isMeta`) turns, and on a
      181-chunk session they filled the whole 20k budget before any tool result
      was considered; the chunk holding the answer was dropped. Fixed: `isMeta`
      turns are scored like tool results, and real user messages are clipped and
      capped at a quarter of the budget.
    - 2026-09-22: a single condensed window (every chunk clipped to ~1k chars)
      ranked the key chunk 6th because the fact sat mid-result, outside the clip.
      Paging at 4x view (five ~27k windows, 1.6 s) ranked it 1st. The filter
      defaults to the paged view.
    - 2026-09-23, eval over 32 delegations from 8 archived sessions: written
      briefs lost every fact on 5 tasks by telling the subagent to look the fact up
      ("check the pr-workflow config entry") instead of stating it; neither the
      filter nor the full transcript ever did. That is what step 3 prevents.
    
    - 2026-09-23, first real `create_session` probes: two Haiku cloud sessions
      received the appended context (their first-turn token counts match) and
      stalled at need-input, one saying the message looked truncated, the other
      that no request had arrived. The context ended the message and the task sat
      above ~20k tokens of transcript. Labelling the block as reference data and
      repeating the task after it, the next probe answered correctly first time.
    
    ## Scripts
    
    - `scripts/chunking.py`: transcript JSONL to chunks (user, assistant, harness,
      tool call + result), with secret values and token patterns redacted before
      anything leaves the machine.
    - `scripts/jevfilter.py`: windows under Jev's limits (32k state plus longest
      question, 64k state plus all questions), one Noul per chunk, rank into the
      budget.
    - `scripts/context_hook.py`: the PreToolUse hook. Fails open and says so.
    - `scripts/preview.py`: step 2.
    
    Install as a plugin (the hook wires itself; `hooks/hooks.json`):
    
    ```bash
    claude plugin marketplace add oaustegard/claude-skills
    claude plugin install delegating-with-context@oaustegard-claude-skills
    ```
    
    Or wire it by hand in `.claude/settings.json`, guarded so a missing file exits
    0 (python3 on a missing path exits 2, which blocks the tool call):
    
    ```json
    {"matcher": "Agent|Task|mcp__.*__create_session", "hooks": [{"type": "command", "timeout": 120,
      "command": "f=/mnt/skills/user/delegating-with-context/scripts/context_hook.py; test -f \"$f\" || exit 0; exec python3 \"$f\""}]}
    ```
    
    Evidence: with ~20k tokens of selected chunks a fresh subagent found 99/105
    required facts, against 102/105 with the whole ~81k-token transcript and 82/105
    with a written brief; method and caveats in
    `oaustegard/experiments/subagent-context-filter/RESULTS.md`.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related