container-layer
Authors and caches a personalized container environment from a Dockerfile- like spec, as a single layer or as a composition of independently cached layers. Use when the user mentions "container layer", "Containerfile", "custom container", "cache my installs", "composable layers",
Install
npx skills add https://github.com/oaustegard/claude-skills/tree/main/plugins/environment-and-config/skills/container-layer
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install oaustegard-claude-skills@llmmart
git clone https://github.com/oaustegard/claude-skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole oaustegard/claude-skills collection as a plugin from our marketplace. Git is the plain clone.
README
container-layer
Custom, cached environment overlays for Claude's ephemeral containers.
What This Does
Claude.ai and Claude Code on the Web run in ephemeral containers — every session starts from a blank slate. This skill lets you declare your environment in a Containerfile (a Dockerfile subset), build it once, cache the result as a tarball in GitHub Releases, and restore it in seconds on subsequent sessions.
First session: parse Containerfile → execute instructions → snapshot filesystem delta → push ~3 MB tarball to GitHub Releases.
Every session after: download tarball → extract → done. One fetch replaces N installs.
Components
| File | Purpose |
|---|---|
SKILL.md |
Skill metadata and documentation |
Containerfile |
Default environment spec (edit this) |
boot.sh |
Boot script for Claude.ai project instructions |
boot-ccotw.sh |
Boot script for Claude Code (SessionStart hook) |
scripts/containerfile.py |
Parser + executor with baseline diffing |
scripts/layer_cache.py |
GitHub Releases tarball cache |
scripts/cli.py |
CLI: build, restore, hash, inspect |
scripts/uv_shim.sh |
Captures ad-hoc uv pip install to Containerfile |
Supported Instructions
FETCH github:user/repo /dest # GitHub repo tarball
FETCH github:user/repo@ref /dest # Specific ref
RUN uv pip install --system pandas # Shell commands
ENV KEY=value # Environment variables
WORKDIR /path # Working directory
SNAPSHOT /path # Include in cached layer
FROM, EXPOSE, CMD, ENTRYPOINT, etc. are silently ignored (Dockerfile compatibility).
Smart Snapshotting
The executor captures a filesystem baseline before building, then diffs against it — only new files from pip install / uv pip install are included in the tarball, not the entire dist-packages directory. FETCH destinations are captured in full. This keeps tarballs small (~3 MB for a full skills repo + Python packages).
Cache Invalidation
The cache key is a SHA-256 of the Containerfile contents. Pass --invalidate-on user/repo to include a GitHub repo's HEAD SHA in the key — when that repo gets a new commit, the cache auto-invalidates and triggers a rebuild.
python3 -m scripts.cli --invalidate-on oaustegard/claude-skills restore ./Containerfile
Ad-Hoc Install Capture
Source the uv shim to automatically append new installs to your Containerfile:
source ./scripts/uv_shim.sh ./Containerfile
uv pip install --system pandas # installs AND appends RUN line
Test Repo
See container-layer-test for a working example with Claude Code on the Web SessionStart hooks.
Skill manifest
Container Layer
Build a reproducible, cached environment overlay for ephemeral containers using a Dockerfile-like spec.
When NOT to use this skill
This authors and caches a layer spec. It is not a Docker troubleshooting tool.
| Situation | Use |
|---|---|
| A build is slow or failing | read the build log; this skill will not help |
| Managing a running container | docker/podman directly |
| Session boot sequence and hooks | the workspace's own boot docs |
The tell is tense: this skill is for the environment you want next session, not the container you are fighting now.
Concept
The container resets every session, but your environment shouldn't. This skill:
- Parses a
Containerfile(Dockerfile subset) that declares your environment - Caches the built result as a tarball in GitHub Releases
- Restores from cache on subsequent boots (single fetch vs. N installs)
- Provides a
uvshim that captures ad-hoc installs back into the Containerfile
Supported Containerfile Instructions
# Environment variables
ENV KEY=value
# Shell commands (including package installs)
RUN apt-get install -y foo # system packages
RUN uv pip install pandas numpy # Python packages (preferred)
RUN pip install requests # also works
# Fetch files from URLs or GitHub
FETCH https://example.com/file.tar.gz /dest/path
FETCH github:user/repo /dest/path # latest tarball
FETCH github:user/repo@ref /dest/path # specific ref
# Set working directory for subsequent RUN commands
WORKDIR /some/path
# Declare paths to include in the cached layer snapshot
# (auto-detected for FETCH destinations and pip/uv installs)
SNAPSHOT /additional/path/to/capture
# Ignored (Dockerfile compat, no-op here):
# FROM, EXPOSE, CMD, ENTRYPOINT, LABEL, ARG, VOLUME, USER, SHELL
Usage
Single layer — build / restore
from scripts.containerfile import ContainerLayer
layer = ContainerLayer(
containerfile_path="/path/to/Containerfile",
cache_repo="oaustegard/claude-container-layers", # GitHub repo for release assets
gh_token="...",
)
# Try cache first, fall back to full build
layer.restore_or_build()
Or via CLI:
python -m scripts.cli restore /path/to/Containerfile --repo user/cache-repo
Multi-layer composition (v0.2.0+)
Decompose a heavy environment into named layers, each cached independently. Compose them in order on session start so most-changed bits don't invalidate stable bits.
from scripts.containerfile import compose
compose(
containerfile_paths=[
"layers/Containerfile", # name='base' (always-on)
"layers/Containerfile.scientific", # name='scientific'
"layers/Containerfile.mojo", # name='mojo'
],
cache_repo="user/cache-repo",
)
Each layer gets its own cache release tag layer-<name>-<hash> so retention policies (keep last N) and cache invalidation operate per-name.
Default layer names are derived from the Containerfile path:
Containerfile→baseContainerfile.scientific→scientificlayers/Containerfile.X→X
CLI equivalent:
python -m scripts.cli compose \
layers/Containerfile \
layers/Containerfile.scientific \
layers/Containerfile.mojo \
--repo user/cache-repo
Per-layer name override
If filename doesn't derive cleanly, pass --name NAME:PATH per layer:
python -m scripts.cli compose \
--name base:weird-named-file.txt \
--name mojo:other-file.txt \
weird-named-file.txt other-file.txt
Single-layer naming (back-compat)
build / restore / hash / inspect accept --name:
python -m scripts.cli restore Containerfile.mojo --name mojo
# Cache tag becomes 'layer-mojo-<hash>' instead of 'layer-<hash>'.
# Omit --name to keep the old back-compat tag for existing callers.
The uv Shim
After building, install the shim to capture future installs:
source /path/to/container-layer/scripts/uv_shim.sh /path/to/Containerfile
Now uv pip install foo both installs the package AND appends RUN uv pip install foo to your Containerfile.
Rebuilding the Cache
After modifying the Containerfile:
layer.build_and_push() # Execute, snapshot, upload
Architecture
Read scripts/containerfile.py for the parser/executor and scripts/layer_cache.py for the GitHub Releases caching logic. The cache key is a SHA-256 of the Containerfile contents — any change triggers a rebuild.
Configuration
The skill expects these environment variables (or pass as constructor args):
GH_TOKEN— GitHub token withreposcope (for releases)- Cache repo can be any repo the token has write access to
Workflow Integration
This skill is designed to be invoked from a boot script. Example Containerfile:
# Skills
FETCH github:oaustegard/claude-skills /mnt/skills/user
# Python environment
RUN uv pip install --system pandas numpy requests
# Path config
RUN echo '/mnt/skills/user/remembering' > /usr/local/lib/python3.12/dist-packages/muninn-remembering.pth
# Custom setup
ENV MY_VAR=hello
WORKDIR /home/claude
Files (claude-skills)
-
scripts
-
cli.py 8.1 KB
#!/usr/bin/env python3 """ CLI for container-layer: build, restore, or snapshot container layers. Single-layer mode (back-compat): python -m scripts.cli build /path/to/Containerfile [--repo user/repo] [--no-cache] [--name N] python -m scripts.cli restore /path/to/Containerfile [--repo user/repo] [--name N] python -m scripts.cli hash /path/to/Containerfile [--name N] python -m scripts.cli inspect /path/to/Containerfile Multi-layer composition (new in v0.2.0): python -m scripts.cli compose <containerfile1> [<containerfile2> ...] [--repo user/repo] Restores each layer in order. Each layer gets its own cache release tag `layer-<name>-<hash>` (name derived from filename: `Containerfile.mojo` -> 'mojo', `Containerfile` -> 'base'). Later layers can overwrite earlier ones — additive Docker-like semantics. Per-layer name override (rare; uncommon path/filename): python -m scripts.cli compose --name base:Containerfile.foo --name mojo:Containerfile.bar Cache invalidation (single or composed): --invalidate-on user/repo Include repo HEAD SHA in cache key --invalidate-on user/repo@branch Specific branch Multiple repos: --invalidate-on repo1 --invalidate-on repo2 """ import argparse import os import sys from .containerfile import ( ContainerLayer, compose, content_hash, default_layer_name, github_head_sha, parse_containerfile, ) def _compute_salt(invalidate_on: list[str], token: str) -> str: """Compute salt from GitHub repo HEAD SHAs.""" if not invalidate_on: return "" parts = [] for spec in invalidate_on: if "@" in spec: repo, ref = spec.rsplit("@", 1) else: repo, ref = spec, "main" sha = github_head_sha(repo, ref, token) if sha: parts.append(f"{repo}@{sha}") print(f" Salt: {repo} @ {sha}") else: print(f" WARNING: couldn't fetch HEAD for {repo}, skipping from salt") return "|".join(parts) def cmd_build(args): """Execute the Containerfile and optionally push to cache.""" salt = _compute_salt(args.invalidate_on or [], args.token) layer = ContainerLayer( containerfile_path=args.containerfile, cache_repo=args.repo, gh_token=args.token, salt=salt, layer_name=args.name, ) if args.no_cache: result = layer.build_only() else: result = layer.build_and_push() if result.success: print(f"\n✓ Build complete (tag: {layer.tag})") if result.snapshot_paths: print(f" Snapshot paths: {len(result.snapshot_paths)} entries") if result.env_vars: print(f" Environment: {len(result.env_vars)} vars set") else: print("\n✗ Build failed") for err in result.errors: print(f" {err}") sys.exit(1) def cmd_restore(args): """Try to restore from cache, fall back to build.""" salt = _compute_salt(args.invalidate_on or [], args.token) layer = ContainerLayer( containerfile_path=args.containerfile, cache_repo=args.repo, gh_token=args.token, salt=salt, layer_name=args.name, ) result = layer.restore_or_build() if result.success: print(f"\n✓ Environment ready (tag: {layer.tag})") else: print("\n✗ Restore failed") for err in result.errors: print(f" {err}") sys.exit(1) def cmd_compose(args): """Restore (or build+push on miss) a sequence of named layers.""" salt = _compute_salt(args.invalidate_on or [], args.token) # Parse --name overrides: list of "name:path" strings to (name, path) pairs. name_overrides: dict[str, str] = {} for spec in args.name or []: if ":" not in spec: print(f"✗ --name expects 'name:path', got: {spec}") sys.exit(2) name, path = spec.split(":", 1) name_overrides[os.path.abspath(path)] = name paths = args.containerfiles if not paths: print("✗ compose requires at least one Containerfile path") sys.exit(2) # Apply overrides where path matches; fall back to default derivation names = [name_overrides.get(os.path.abspath(p)) for p in paths] results = compose( containerfile_paths=paths, cache_repo=args.repo, gh_token=args.token, salt=salt, names=names, ) succeeded = [r for r in results if r.success] print( f"\n=== Compose summary: {len(succeeded)}/{len(paths)} layers ready ===" ) if len(succeeded) != len(paths): sys.exit(1) def cmd_hash(args): """Print the cache key hash (or full tag, if --name) of a Containerfile.""" salt = _compute_salt(args.invalidate_on or [], args.token) h = content_hash(args.containerfile, extra_salt=salt) if args.name: print(f"layer-{args.name}-{h}") else: print(h) def cmd_inspect(args): """Parse and display the instructions in a Containerfile.""" instructions = parse_containerfile(args.containerfile) salt = _compute_salt(args.invalidate_on or [], args.token) layer_name = args.name or default_layer_name(args.containerfile) h = content_hash(args.containerfile, extra_salt=salt) print(f"Containerfile: {args.containerfile}") print(f"Derived name: {layer_name}") print(f"Cache key: {h}") print(f"Full tag: layer-{layer_name}-{h}") print(f"Instructions: {len(instructions)}") print() for inst in instructions: print(f" [{inst.line_num:3d}] {inst.directive:10s} {inst.args}") def main(): parser = argparse.ArgumentParser(description="Container layer manager") parser.add_argument( "--token", default=os.environ.get("GH_TOKEN", ""), help="GitHub token (default: $GH_TOKEN)", ) parser.add_argument( "--repo", default="oaustegard/claude-container-layers", help="GitHub repo for cache storage", ) parser.add_argument( "--invalidate-on", action="append", help="GitHub repo whose HEAD SHA is included in cache key " "(e.g. user/repo or user/repo@branch). Repeatable.", ) sub = parser.add_subparsers(dest="command", required=True) p_build = sub.add_parser("build", help="Execute Containerfile and cache result") p_build.add_argument("containerfile") p_build.add_argument( "--name", help="Layer name for cache release tag (default: derived from filename, " "e.g. 'Containerfile.mojo' -> 'mojo'). Pass empty to use old " "back-compat tag 'layer-<hash>'.", ) p_build.add_argument("--no-cache", action="store_true", help="Skip cache push") p_build.set_defaults(func=cmd_build) p_restore = sub.add_parser("restore", help="Restore from cache or build") p_restore.add_argument("containerfile") p_restore.add_argument( "--name", help="Layer name for cache release tag (see `build --name`).", ) p_restore.set_defaults(func=cmd_restore) p_compose = sub.add_parser( "compose", help="Restore a sequence of named layers in order (cache miss = build+push)", ) p_compose.add_argument( "containerfiles", nargs="+", help="Ordered list of Containerfile paths to restore", ) p_compose.add_argument( "--name", action="append", help="Per-layer name override, formatted 'name:path'. Repeatable. " "Paths not listed get a default name derived from filename.", ) p_compose.set_defaults(func=cmd_compose) p_hash = sub.add_parser("hash", help="Print Containerfile cache key") p_hash.add_argument("containerfile") p_hash.add_argument( "--name", help="If set, prints full tag `layer-<name>-<hash>` instead of bare hash." ) p_hash.set_defaults(func=cmd_hash) p_inspect = sub.add_parser("inspect", help="Show parsed instructions") p_inspect.add_argument("containerfile") p_inspect.add_argument( "--name", help="Override the layer name displayed in inspection output." ) p_inspect.set_defaults(func=cmd_inspect) args = parser.parse_args() args.func(args) if __name__ == "__main__": main() -
containerfile.py 20 KB
""" Containerfile parser and executor. Parses a Dockerfile-like spec and executes the supported subset of instructions, tracking which filesystem paths are modified for layer snapshot/caching. """ import hashlib import json import os import re import shlex import subprocess import urllib.request from dataclasses import dataclass, field from pathlib import Path # Instructions we execute EXECUTABLE_INSTRUCTIONS = {"ENV", "RUN", "FETCH", "WORKDIR", "SNAPSHOT"} # Instructions we silently ignore (Dockerfile compat) IGNORED_INSTRUCTIONS = { "FROM", "EXPOSE", "CMD", "ENTRYPOINT", "LABEL", "ARG", "VOLUME", "USER", "SHELL", "HEALTHCHECK", "STOPSIGNAL", "ONBUILD", } # Well-known paths that pip/uv/apt install into. # `_dedup_paths` uses these as baseline keys so that auto-tracked snapshot # paths (from RUN commands) get the diff-vs-baseline treatment instead of # whole-tree capture. # # Lists all common Python versions so the diff logic works regardless of # which interpreter is active. The single-entry list pre-0.2.1 hardcoded # python3.12, so containers on python3.11 captured all of dist-packages # instead of just newly-installed files. # # /usr/{bin,lib,share} are here so `apt-get install` layers diff correctly # — `_exec_run` auto-adds these to snapshot_paths on any apt invocation # (one .deb pulls in shared libs / headers / docs sprawled across all # three). Pre-0.2.2 they fell into the whole-tree path: adding `zstd` # to a layer captured ~3GB raw → ~600MB compressed. # # Despite the name, this list is no longer Python-only. Kept the name for # backwards-compat with any external importer; consider renaming to # BASELINED_PATHS in a future major bump. PYTHON_INSTALL_PATHS = [ "/usr/local/lib/python3.10/dist-packages", "/usr/local/lib/python3.11/dist-packages", "/usr/local/lib/python3.12/dist-packages", "/usr/local/lib/python3.13/dist-packages", "/usr/local/bin", "/usr/bin", "/usr/lib", "/usr/share", "/home/claude/.local/lib", "/home/claude/.local/bin", ] @dataclass class Instruction: """A parsed Containerfile instruction.""" line_num: int directive: str args: str raw: str @dataclass class BuildResult: """Result of executing a Containerfile.""" success: bool snapshot_paths: list[str] content_hash: str errors: list[str] = field(default_factory=list) env_vars: dict[str, str] = field(default_factory=dict) def parse_containerfile(path: str) -> list[Instruction]: """Parse a Containerfile into a list of Instructions.""" instructions = [] content = Path(path).read_text() # Handle line continuations content = re.sub(r'\\\n\s*', ' ', content) for line_num, line in enumerate(content.splitlines(), 1): line = line.strip() # Skip comments and blank lines if not line or line.startswith('#'): continue # Extract directive and args match = re.match(r'^([A-Z]+)\s+(.*)', line) if not match: continue directive, args = match.group(1), match.group(2).strip() if directive in EXECUTABLE_INSTRUCTIONS: instructions.append(Instruction(line_num, directive, args, line)) elif directive in IGNORED_INSTRUCTIONS: continue # silently skip else: print(f" WARNING line {line_num}: unknown instruction '{directive}', skipping") return instructions def content_hash(path: str, extra_salt: str = "") -> str: """SHA-256 hash of a Containerfile's contents (the cache key). Optionally include extra_salt (e.g. a git SHA) so the cache invalidates when external dependencies change. """ content = Path(path).read_text().strip() if extra_salt: content += f"\n# salt: {extra_salt}" return hashlib.sha256(content.encode()).hexdigest()[:16] def default_layer_name(containerfile_path: str) -> str: """Derive a default layer name from a Containerfile path. `Containerfile` -> 'base' `Containerfile.mojo` -> 'mojo' `layers/Containerfile.scientific` -> 'scientific' `foo/bar.txt` -> 'bar' (fallback: file stem) Names are lowercased and stripped of non-[a-z0-9-_] characters so they're safe to embed in GitHub Release tags (`layer-<name>-<hash>`). """ basename = os.path.basename(containerfile_path) if basename == "Containerfile": raw = "base" elif basename.startswith("Containerfile."): raw = basename[len("Containerfile."):] else: # Fallback: stem of whatever path was passed raw = os.path.splitext(basename)[0] # Sanitize: keep alphanumerics, hyphen, underscore, period -> hyphen safe = re.sub(r"[^a-z0-9_-]", "-", raw.lower()).strip("-") return safe or "layer" def github_head_sha(repo: str, ref: str = "main", token: str = "") -> str: """Fetch the HEAD SHA of a GitHub repo ref. Returns empty string on failure.""" try: url = f"https://api.github.com/repos/{repo}/commits/{ref}" headers = {"Accept": "application/vnd.github+json"} if token: headers["Authorization"] = f"token {token}" req = urllib.request.Request(url, headers=headers) with urllib.request.urlopen(req, timeout=10) as resp: data = json.loads(resp.read()) return data.get("sha", "")[:12] except Exception: return "" def snapshot_baseline(paths: list[str]) -> dict[str, set[str]]: """Capture the set of files currently in each path (for diffing later).""" baseline = {} for p in paths: p = os.path.normpath(p) if os.path.isdir(p): files = set() for root, dirs, fnames in os.walk(p): for f in fnames: files.add(os.path.join(root, f)) baseline[p] = files elif os.path.isfile(p): baseline[p] = {p} else: baseline[p] = set() return baseline def diff_paths(baseline: dict[str, set[str]], paths: list[str]) -> list[str]: """ Given a baseline snapshot and current paths, return a list of new/modified files to include in the layer tarball. For FETCH destinations (not in baseline), include everything. """ new_files = [] for p in paths: p = os.path.normpath(p) if p not in baseline: # New path (e.g. FETCH destination) — include whole tree if os.path.exists(p): new_files.append(p) continue if os.path.isdir(p): current = set() for root, dirs, fnames in os.walk(p): for f in fnames: current.add(os.path.join(root, f)) added = current - baseline[p] if added: new_files.extend(sorted(added)) elif os.path.isfile(p): if p not in baseline[p]: new_files.append(p) return new_files class ContainerfileExecutor: """Executes a parsed Containerfile, tracking modified paths.""" def __init__(self, gh_token: str | None = None): self.snapshot_paths: list[str] = [] self.env: dict[str, str] = dict(os.environ) self.workdir: str = "/home/claude" self.gh_token = gh_token or os.environ.get("GH_TOKEN", "") self.errors: list[str] = [] self._baseline: dict[str, set[str]] = {} def execute(self, instructions: list[Instruction]) -> BuildResult: """Execute all instructions, return result with snapshot paths.""" file_hash = "" # Caller should set this # Capture baseline of well-known install paths before building self._baseline = snapshot_baseline(PYTHON_INSTALL_PATHS) for inst in instructions: try: handler = getattr(self, f"_exec_{inst.directive.lower()}", None) if handler: print(f" [{inst.line_num}] {inst.raw}") handler(inst) else: self.errors.append(f"Line {inst.line_num}: no handler for {inst.directive}") except Exception as e: msg = f"Line {inst.line_num}: {inst.directive} failed: {e}" self.errors.append(msg) print(f" ERROR: {msg}") return BuildResult( success=False, snapshot_paths=self._dedup_paths(), content_hash=file_hash, errors=self.errors, env_vars={k: v for k, v in self.env.items() if k not in os.environ or os.environ[k] != v}, ) return BuildResult( success=True, snapshot_paths=self._dedup_paths(), content_hash=file_hash, errors=self.errors, env_vars={k: v for k, v in self.env.items() if k not in os.environ or os.environ[k] != v}, ) def _exec_env(self, inst: Instruction): """ENV KEY=value or ENV KEY value""" if '=' in inst.args: key, _, value = inst.args.partition('=') value = value.strip('"').strip("'") else: parts = inst.args.split(None, 1) key = parts[0] value = parts[1] if len(parts) > 1 else "" self.env[key.strip()] = value os.environ[key.strip()] = value def _exec_workdir(self, inst: Instruction): """WORKDIR /path""" path = inst.args.strip() os.makedirs(path, exist_ok=True) self.workdir = path def _exec_run(self, inst: Instruction): """RUN command — execute shell command, detect package installs.""" cmd = inst.args # Detect pip/uv installs to track snapshot paths if re.search(r'\b(pip|uv pip)\s+install\b', cmd): for p in PYTHON_INSTALL_PATHS: if p not in self.snapshot_paths: self.snapshot_paths.append(p) # Auto-add --break-system-packages if not present (externally managed envs) if '--break-system-packages' not in cmd: cmd = cmd.replace('install', 'install --break-system-packages', 1) # Detect apt installs if re.search(r'\bapt(-get)?\s+install\b', cmd): self.snapshot_paths.extend([ "/usr/lib", "/usr/bin", "/usr/share", ]) result = subprocess.run( cmd, shell=True, cwd=self.workdir, env=self.env, capture_output=True, text=True, timeout=300, ) if result.stdout.strip(): # Print last 5 lines of stdout to avoid noise lines = result.stdout.strip().splitlines() for line in lines[-5:]: print(f" {line}") if len(lines) > 5: print(f" ... ({len(lines) - 5} lines omitted)") if result.returncode != 0: raise RuntimeError(f"Command failed (exit {result.returncode}): {result.stderr.strip()}") def _exec_fetch(self, inst: Instruction): """FETCH source dest — fetch from URL or GitHub.""" parts = shlex.split(inst.args) if len(parts) < 2: raise ValueError("FETCH requires <source> <dest>") source, dest = parts[0], parts[1] os.makedirs(dest, exist_ok=True) self.snapshot_paths.append(dest) if source.startswith("github:"): self._fetch_github(source[7:], dest) elif source.startswith("http://") or source.startswith("https://"): self._fetch_url(source, dest) else: raise ValueError(f"Unknown FETCH source: {source}") def _exec_snapshot(self, inst: Instruction): """SNAPSHOT /path — explicitly add a path to the snapshot.""" path = inst.args.strip() if path and path not in self.snapshot_paths: self.snapshot_paths.append(path) def _fetch_github(self, spec: str, dest: str): """Fetch a GitHub repo tarball. Spec: user/repo or user/repo@ref""" if '@' in spec: repo, ref = spec.rsplit('@', 1) else: repo, ref = spec, "main" url = f"https://codeload.github.com/{repo}/tar.gz/{ref}" tarball = f"/tmp/_fetch_{repo.replace('/', '_')}.tar.gz" headers = "" if self.gh_token: headers = f'-H "Authorization: token {self.gh_token}"' cmd = f'curl -sL {headers} "{url}" -o "{tarball}" && tar -xzf "{tarball}" -C "{dest}" --strip-components=1 && rm -f "{tarball}"' result = subprocess.run(cmd, shell=True, capture_output=True, text=True, timeout=120) if result.returncode != 0: raise RuntimeError(f"GitHub fetch failed: {result.stderr.strip()}") print(f" Fetched {repo}@{ref} → {dest}") def _fetch_url(self, url: str, dest: str): """Fetch a URL to a destination.""" filename = url.rsplit('/', 1)[-1] or "download" dest_file = os.path.join(dest, filename) cmd = f'curl -sL "{url}" -o "{dest_file}"' result = subprocess.run(cmd, shell=True, capture_output=True, text=True, timeout=120) if result.returncode != 0: raise RuntimeError(f"URL fetch failed: {result.stderr.strip()}") # Auto-extract tarballs if filename.endswith(('.tar.gz', '.tgz')): subprocess.run(f'tar -xzf "{dest_file}" -C "{dest}" && rm -f "{dest_file}"', shell=True, capture_output=True, timeout=60) print(f" Fetched and extracted {filename} → {dest}") else: print(f" Fetched {filename} → {dest}") def _dedup_paths(self) -> list[str]: """Deduplicate, filter, and diff snapshot paths. For paths that existed before build (pip install targets), only capture new files. For FETCH destinations and explicit SNAPSHOTs, capture everything. """ seen = set() result = [] # Separate paths into baselined (pip/apt targets) and new (FETCH/SNAPSHOT) baselined = set(self._baseline.keys()) for p in self.snapshot_paths: p = os.path.normpath(p) if p in seen or not os.path.exists(p): continue seen.add(p) if p in baselined: # For baselined paths, compute diff — individual new files new_files = diff_paths({p: self._baseline[p]}, [p]) result.extend(new_files) else: # For FETCH destinations and explicit SNAPSHOTs, take the whole tree result.append(p) return result class ContainerLayer: """ High-level interface: parse a Containerfile, execute or restore from cache. A layer can optionally have a `layer_name`. When set, the cache release tag becomes `layer-<name>-<hash>` instead of `layer-<hash>`. This enables composing multiple named layers (each cached independently) into a single container, and lets cache-retention policies operate per-name. Back-compat: if `layer_name` is None, the tag stays `layer-<hash>` so existing single-Containerfile callers don't see their cache invalidate. """ def __init__( self, containerfile_path: str, cache_repo: str = "oaustegard/claude-container-layers", gh_token: str | None = None, salt: str = "", layer_name: str | None = None, ): self.containerfile_path = containerfile_path self.cache_repo = cache_repo self.gh_token = gh_token or os.environ.get("GH_TOKEN", "") self.layer_name = layer_name self._hash = content_hash(containerfile_path, extra_salt=salt) @property def tag(self) -> str: if self.layer_name: return f"layer-{self.layer_name}-{self._hash}" return f"layer-{self._hash}" def restore_or_build(self) -> BuildResult: """Try cache restore first, fall back to full build.""" from . import layer_cache print(f"Container layer hash: {self._hash}") if layer_cache.try_restore(self.cache_repo, self.tag, self.gh_token): print("✓ Restored from cache") # Still need to replay ENV instructions instructions = parse_containerfile(self.containerfile_path) env_instructions = [i for i in instructions if i.directive == "ENV"] executor = ContainerfileExecutor(self.gh_token) for inst in env_instructions: executor._exec_env(inst) return BuildResult( success=True, snapshot_paths=[], content_hash=self._hash, env_vars=executor.env, ) print("Cache miss — building from Containerfile...") return self.build_and_push() def build_and_push(self) -> BuildResult: """Execute Containerfile, snapshot, and push to cache.""" from . import layer_cache instructions = parse_containerfile(self.containerfile_path) executor = ContainerfileExecutor(self.gh_token) result = executor.execute(instructions) result.content_hash = self._hash if result.success and result.snapshot_paths: print(f"\nSnapshotting {len(result.snapshot_paths)} paths...") layer_cache.build_and_push( result.snapshot_paths, self.cache_repo, self.tag, self.gh_token ) return result def build_only(self) -> BuildResult: """Execute Containerfile without caching (for testing).""" instructions = parse_containerfile(self.containerfile_path) executor = ContainerfileExecutor(self.gh_token) result = executor.execute(instructions) result.content_hash = self._hash return result def compose( containerfile_paths: list[str], cache_repo: str = "oaustegard/claude-container-layers", gh_token: str | None = None, salt: str = "", names: list[str | None] | None = None, ) -> list[BuildResult]: """Restore (or build+push, on miss) a sequence of named layers in order. Each Containerfile becomes a named ContainerLayer with its own cache key and GitHub Release. Layers are restored sequentially — later layers' file modifications can overwrite earlier ones, mirroring Docker's additive-overlay semantics. Args: containerfile_paths: Ordered list of Containerfile paths. cache_repo: Single cache repo for all layers (each gets its own release within it, tagged `layer-<name>-<hash>`). gh_token: GitHub token; falls back to $GH_TOKEN. salt: Optional salt applied to every layer's hash (typically a git HEAD SHA so the whole composition invalidates together when the source repo advances). names: Optional per-layer name overrides. None entries fall back to `default_layer_name()`. List length must match `containerfile_paths` if provided. Returns: List of BuildResult, one per layer, in the same order as input. Stops on first failure (later layers in the list aren't attempted). """ if names is not None and len(names) != len(containerfile_paths): raise ValueError( f"names length ({len(names)}) must match containerfile_paths " f"length ({len(containerfile_paths)})" ) results: list[BuildResult] = [] for i, cf_path in enumerate(containerfile_paths): explicit_name = names[i] if names else None name = explicit_name or default_layer_name(cf_path) print(f"\n=== Composing layer [{i + 1}/{len(containerfile_paths)}]: {name} ({cf_path}) ===") layer = ContainerLayer( containerfile_path=cf_path, cache_repo=cache_repo, gh_token=gh_token, salt=salt, layer_name=name, ) result = layer.restore_or_build() results.append(result) if not result.success: print(f"\n✗ Compose halted at layer '{name}' (errors above)") break return results -
layer_cache.py 8.4 KB
""" Layer cache: snapshot filesystem paths into a tarball and store/retrieve via GitHub Releases on a designated repo. Cache key = SHA-256 of Containerfile contents, used as the release tag. """ import json import os import subprocess import urllib.error import urllib.request from datetime import UTC, datetime TARBALL_NAME = "layer.tar.gz" def _gh_api( endpoint: str, token: str, method: str = "GET", data: bytes | None = None, content_type: str = "application/json", timeout: int = 30, ) -> dict | None: """Make a GitHub API request.""" url = f"https://api.github.com{endpoint}" if endpoint.startswith("/") else endpoint headers = { "Authorization": f"token {token}", "Accept": "application/vnd.github+json", } if data and content_type: headers["Content-Type"] = content_type req = urllib.request.Request(url, data=data, headers=headers, method=method) try: with urllib.request.urlopen(req, timeout=timeout) as resp: body = resp.read() return json.loads(body) if body else {} except urllib.error.HTTPError as e: if e.code == 404: return None raise def _find_release(repo: str, tag: str, token: str) -> dict | None: """Find a release by tag.""" return _gh_api(f"/repos/{repo}/releases/tags/{tag}", token) def _create_release(repo: str, tag: str, token: str) -> dict: """Create a release (or return existing one).""" existing = _find_release(repo, tag, token) if existing: return existing payload = json.dumps({ "tag_name": tag, "name": f"Container Layer {datetime.now(UTC).strftime('%Y-%m-%dT%H%M%SZ')} {tag}", "body": "Auto-generated container layer cache. Safe to delete.", "draft": False, "prerelease": True, }).encode() result = _gh_api(f"/repos/{repo}/releases", token, method="POST", data=payload) if not result: raise RuntimeError(f"Failed to create release {tag} on {repo}") return result def _upload_asset(upload_url: str, filepath: str, token: str): """Upload a release asset.""" # upload_url has {?name,label} template suffix — strip it upload_url = upload_url.split("{")[0] upload_url += f"?name={TARBALL_NAME}" with open(filepath, "rb") as f: data = f.read() size_mb = len(data) / (1024 * 1024) print(f" Uploading {size_mb:.1f} MB...") headers = { "Authorization": f"token {token}", "Content-Type": "application/gzip", "Content-Length": str(len(data)), } req = urllib.request.Request(upload_url, data=data, headers=headers, method="POST") with urllib.request.urlopen(req, timeout=120) as resp: result = json.loads(resp.read()) print(f" Uploaded: {result.get('browser_download_url', 'ok')}") def _find_asset_url(release: dict) -> str | None: """Find the layer tarball asset URL in a release.""" for asset in release.get("assets", []): if asset["name"] == TARBALL_NAME: return asset["url"] # API URL (needs Accept header for download) return None def try_restore(repo: str, tag: str, token: str) -> bool: """ Try to restore a cached layer from GitHub Releases. Returns True if successfully restored, False if cache miss. """ print(f" Checking cache: {repo} @ {tag}") release = _find_release(repo, tag, token) if not release: print(" Cache miss: no release found") return False asset_url = _find_asset_url(release) if not asset_url: print(" Cache miss: release exists but no tarball asset") return False # Download the asset print(" Cache hit — downloading layer...") tarball = "/tmp/_layer_restore.tar.gz" # Stream via urllib so the token stays in request headers, never in argv. # The previous implementation shelled out to curl with the token interpolated # into the command string; on TimeoutExpired, Python's exception __str__ # echoed the full cmd (including the token) into stderr/logs/transcripts. # urllib keeps the secret in the Request object and out of any error message. headers = { "Authorization": f"token {token}", "Accept": "application/octet-stream", } req = urllib.request.Request(asset_url, headers=headers) try: # Per-read timeout, not wall-clock — multi-hundred-MB layers on slow # links won't trip the old hardcoded 120s ceiling. with urllib.request.urlopen(req, timeout=60) as resp, open(tarball, "wb") as out: while True: chunk = resp.read(1024 * 1024) if not chunk: break out.write(chunk) except (urllib.error.URLError, TimeoutError, OSError) as e: print(f" Download failed: {type(e).__name__}: {e}") if os.path.exists(tarball): os.remove(tarball) return False if not os.path.exists(tarball): print(" Download failed: no file written") return False size_mb = os.path.getsize(tarball) / (1024 * 1024) print(f" Downloaded {size_mb:.1f} MB — extracting...") # Extract from root to restore absolute paths result = subprocess.run( f'tar -xzf "{tarball}" -C / 2>&1', shell=True, capture_output=True, text=True, timeout=120, ) os.remove(tarball) if result.returncode != 0: # Some permission errors are expected and harmless errors = [l for l in result.stderr.splitlines() if "Cannot" not in l] if errors: print(f" Extraction warnings: {'; '.join(errors[:3])}") print(" Layer restored") return True def build_and_push( snapshot_paths: list[str], repo: str, tag: str, token: str, ): """ Create a tarball from snapshot_paths and upload as a GitHub Release asset. """ if not snapshot_paths: print(" No paths to snapshot") return # Filter to existing paths existing = [p for p in snapshot_paths if os.path.exists(p)] if not existing: print(" No existing paths to snapshot") return print(" Paths to snapshot:") for p in existing: # Get size size = subprocess.run( f'du -sh "{p}" 2>/dev/null | cut -f1', shell=True, capture_output=True, text=True, ).stdout.strip() print(f" {p} ({size})") tarball = "/tmp/_layer_build.tar.gz" # Build tarball with absolute paths (rooted at /). Feed the path list to # tar via a NUL-delimited -T file rather than the command line: a large # snapshot (thousands of individual files) overflows the single argv string # passed to `/bin/sh -c`, raising OSError [Errno 7] "Argument list too long" # at exec — which aborts the layer build entirely. -T reads names from a # file, so argv stays tiny regardless of how many paths are snapshotted. # --null pairs with NUL separators so paths with spaces need no quoting. filelist = "/tmp/_layer_build_files.txt" with open(filelist, "wb") as fh: fh.write(b"\0".join(os.fsencode(p) for p in existing)) cmd = f'tar -czf "{tarball}" --null -T "{filelist}" 2>&1' result = subprocess.run(cmd, shell=True, capture_output=True, text=True, timeout=300) try: os.remove(filelist) except OSError: pass if not os.path.exists(tarball): print(f" Tarball creation failed: {result.stderr.strip()}") return size_mb = os.path.getsize(tarball) / (1024 * 1024) print(f" Layer tarball: {size_mb:.1f} MB") if size_mb > 2000: print(f" WARNING: tarball is {size_mb:.0f} MB — GitHub release assets max at 2GB") os.remove(tarball) return # Create release and upload try: # Delete existing release if present (to replace the asset) existing_release = _find_release(repo, tag, token) if existing_release: _gh_api( f"/repos/{repo}/releases/{existing_release['id']}", token, method="DELETE", ) print(" Replaced existing cache entry") release = _create_release(repo, tag, token) _upload_asset(release["upload_url"], tarball, token) print(f" ✓ Layer cached as {repo} release: {tag}") except Exception as e: print(f" Cache push failed: {e}") finally: if os.path.exists(tarball): os.remove(tarball) -
test_baseline_paths.py 4.9 KB
""" Unit tests for the PYTHON_INSTALL_PATHS baseline fix (v0.2.1). Locks in the behavior that dist-packages for python3.10/3.11/3.12/3.13 are ALL recognized as baselined paths — so a SNAPSHOT directive that references any of them gets the diff-vs-baseline treatment instead of full-tree capture. Run: cd container-layer python3 -m scripts.test_baseline_paths """ import os import shutil import tempfile import unittest from pathlib import Path from . import containerfile from .containerfile import ( PYTHON_INSTALL_PATHS, ContainerfileExecutor, snapshot_baseline, ) class TestPythonInstallPathsCoverage(unittest.TestCase): """PYTHON_INSTALL_PATHS must include every dist-packages version we run on.""" def test_python_3_10_through_3_13_dist_packages_present(self): for v in ("3.10", "3.11", "3.12", "3.13"): self.assertIn( f"/usr/local/lib/python{v}/dist-packages", PYTHON_INSTALL_PATHS, f"missing python{v}/dist-packages — diff-vs-baseline won't fire on " f"containers using this version", ) def test_local_bin_still_listed(self): self.assertIn("/usr/local/bin", PYTHON_INSTALL_PATHS) def test_system_apt_target_dirs_present(self): """`_exec_run` auto-snapshots /usr/{bin,lib,share} on apt installs. These must be in PYTHON_INSTALL_PATHS too, else the apt-install path falls back to whole-tree capture (regression: 678MB layer ballooned to 1625MB when a single libfuse2 apt-install fired the auto-snapshot before these entries were added).""" for p in ("/usr/bin", "/usr/lib", "/usr/share"): self.assertIn( p, PYTHON_INSTALL_PATHS, f"missing {p} — apt-install layers will whole-tree-capture this dir", ) def test_user_local_bin_present(self): """Sanity: dot-local paths for non-root installs still listed.""" self.assertIn("/home/claude/.local/lib", PYTHON_INSTALL_PATHS) self.assertIn("/home/claude/.local/bin", PYTHON_INSTALL_PATHS) class TestBaselineCapturesNonExistentPaths(unittest.TestCase): """snapshot_baseline returns an entry for every requested path, even if the path doesn't exist — so _dedup_paths can later treat it as baselined regardless of build-time presence.""" def test_missing_path_in_baseline_keys(self): baseline = snapshot_baseline(["/nonexistent/path/here"]) self.assertIn("/nonexistent/path/here", baseline) self.assertEqual(baseline["/nonexistent/path/here"], set()) class TestDedupPathsDiffsBaselinedDirEvenWhenItGrows(unittest.TestCase): """Regression test for the bug: with python3.11 in PYTHON_INSTALL_PATHS, a SNAPSHOT directive on /usr/local/lib/python3.11/dist-packages should trigger the diff path (only NEW files captured), not the whole-tree path.""" def setUp(self): self.tmpdir = Path(tempfile.mkdtemp()) self.addCleanup(lambda: shutil.rmtree(self.tmpdir, ignore_errors=True)) def _fake_baselined_dir_in_install_paths(self, simulate_path: str): """Patch PYTHON_INSTALL_PATHS to include a tempdir we control, so we can test the executor's baseline + diff path without touching real /usr/local.""" self._patch_paths = simulate_path self._orig = containerfile.PYTHON_INSTALL_PATHS containerfile.PYTHON_INSTALL_PATHS = [simulate_path] + [ p for p in self._orig if not p.endswith("dist-packages") ] self.addCleanup(setattr, containerfile, "PYTHON_INSTALL_PATHS", self._orig) def test_only_new_files_captured_in_baselined_dir(self): simulated = str(self.tmpdir / "fake-dist-packages") os.makedirs(simulated, exist_ok=True) # Pre-existing "baseline" file (Path(simulated) / "existing.py").write_text("# already there\n") self._fake_baselined_dir_in_install_paths(simulated) executor = ContainerfileExecutor() # Capture baseline as the real executor would, before any RUN executor._baseline = snapshot_baseline(containerfile.PYTHON_INSTALL_PATHS) # Simulate a SNAPSHOT directive having added this path executor.snapshot_paths.append(simulated) # Simulate a RUN command adding new files into the same dir (Path(simulated) / "added_by_run.py").write_text("# new\n") (Path(simulated) / "another.py").write_text("# also new\n") snapshotted = executor._dedup_paths() # The diff path should pick up ONLY the new files, not the existing one names = {os.path.basename(p) for p in snapshotted} self.assertIn("added_by_run.py", names) self.assertIn("another.py", names) self.assertNotIn("existing.py", names) # And it should be a list of individual files (diff), not the whole dir self.assertNotIn(simulated, snapshotted) if __name__ == "__main__": unittest.main(verbosity=2) -
test_named_layers.py 6.6 KB
""" Unit tests for named-layer + compose support (v0.2.0). Runnable standalone: cd container-layer python3 -m scripts.test_named_layers """ import os import tempfile import unittest from unittest import mock from .containerfile import ( BuildResult, ContainerLayer, compose, default_layer_name, ) class TestDefaultLayerName(unittest.TestCase): """Path -> layer name derivation.""" def test_bare_containerfile_becomes_base(self): self.assertEqual(default_layer_name("Containerfile"), "base") self.assertEqual(default_layer_name("/abs/path/Containerfile"), "base") self.assertEqual(default_layer_name("./rel/Containerfile"), "base") def test_dot_suffix_becomes_name(self): self.assertEqual(default_layer_name("Containerfile.mojo"), "mojo") self.assertEqual(default_layer_name("Containerfile.torch-cpu"), "torch-cpu") self.assertEqual(default_layer_name("layers/Containerfile.scientific"), "scientific") def test_lowercase_and_sanitize(self): self.assertEqual(default_layer_name("Containerfile.JuliaSR"), "juliasr") self.assertEqual(default_layer_name("Containerfile.foo bar"), "foo-bar") # Note: os.path.basename strips directory components first, so # "Containerfile.weird/path" becomes "path" before suffix-stripping. def test_fallback_to_file_stem(self): self.assertEqual(default_layer_name("foo/bar.txt"), "bar") self.assertEqual(default_layer_name("/tmp/my_layer"), "my_layer") def test_empty_or_all_special_chars_falls_back(self): self.assertEqual(default_layer_name("Containerfile.@@@"), "layer") self.assertEqual(default_layer_name("Containerfile.---"), "layer") class TestContainerLayerTag(unittest.TestCase): """tag property: with vs without layer_name.""" def _layer(self, path: str, name: str = None) -> ContainerLayer: # ContainerLayer reads the file for hashing; create a temp file. with tempfile.NamedTemporaryFile(mode="w", suffix=".Containerfile", delete=False) as f: f.write("RUN echo hello\n") cf_path = f.name self.addCleanup(os.unlink, cf_path) return ContainerLayer( containerfile_path=cf_path, cache_repo="testorg/test-cache", gh_token="test-token", layer_name=name, ) def test_unnamed_layer_uses_back_compat_tag(self): layer = self._layer("Containerfile") self.assertRegex(layer.tag, r"^layer-[a-f0-9]{16}$") def test_named_layer_uses_named_tag(self): layer = self._layer("Containerfile", name="scientific") self.assertRegex(layer.tag, r"^layer-scientific-[a-f0-9]{16}$") def test_same_content_same_hash_regardless_of_name(self): layer_a = self._layer("Containerfile", name="alpha") layer_b = self._layer("Containerfile", name="beta") # Different names -> different tags, but same underlying content hash self.assertNotEqual(layer_a.tag, layer_b.tag) self.assertEqual(layer_a._hash, layer_b._hash) class TestCompose(unittest.TestCase): """compose() orchestrates per-layer restore in order.""" def setUp(self): # Three fake containerfiles in a clean tempdir so basename == filename self.tmpdir = tempfile.mkdtemp() self.addCleanup(self._cleanup_tmpdir) self.cf_paths = [] for name, body in [ ("Containerfile", "RUN echo base\n"), ("Containerfile.scientific", "RUN echo sci\n"), ("Containerfile.mojo", "RUN echo mojo\n"), ]: path = os.path.join(self.tmpdir, name) with open(path, "w") as f: f.write(body) self.cf_paths.append(path) def _cleanup_tmpdir(self): import shutil shutil.rmtree(self.tmpdir, ignore_errors=True) def test_compose_invokes_restore_or_build_on_each_layer_in_order(self): calls = [] def fake_restore(self): calls.append((self.layer_name, self.containerfile_path)) return BuildResult(success=True, snapshot_paths=[], content_hash="abc") with mock.patch.object(ContainerLayer, "restore_or_build", fake_restore): results = compose( containerfile_paths=self.cf_paths, cache_repo="testorg/cache", gh_token="t", ) self.assertEqual(len(results), 3) self.assertTrue(all(r.success for r in results)) # Order preserved + names derived from filename suffix: # Containerfile -> 'base', Containerfile.scientific -> 'scientific', etc. self.assertEqual([n for n, _ in calls], ["base", "scientific", "mojo"]) def test_compose_halts_on_first_failure(self): attempt_log = [] def flaky_restore(self): attempt_log.append(self.layer_name) success = self.layer_name != "scientific" return BuildResult( success=success, snapshot_paths=[], content_hash="abc", errors=[] if success else ["simulated failure"], ) with mock.patch.object(ContainerLayer, "restore_or_build", flaky_restore): results = compose( containerfile_paths=self.cf_paths, cache_repo="testorg/cache", gh_token="t", ) # Stopped at scientific (the 2nd layer), didn't try mojo self.assertEqual(attempt_log, ["base", "scientific"]) self.assertEqual(len(results), 2) self.assertFalse(results[1].success) def test_compose_respects_explicit_name_overrides(self): seen_names = [] def capture(self): seen_names.append(self.layer_name) return BuildResult(success=True, snapshot_paths=[], content_hash="abc") with mock.patch.object(ContainerLayer, "restore_or_build", capture): compose( containerfile_paths=self.cf_paths[:2], names=["custom-1", None], # second falls back to default cache_repo="testorg/cache", gh_token="t", ) # First is overridden, second derived from filename self.assertEqual(seen_names[0], "custom-1") self.assertEqual(seen_names[1], "scientific") def test_compose_rejects_mismatched_names_length(self): with self.assertRaises(ValueError): compose( containerfile_paths=self.cf_paths, # 3 paths names=["only-one"], # 1 name cache_repo="testorg/cache", gh_token="t", ) if __name__ == "__main__": unittest.main(verbosity=2) -
uv_shim.sh 2.5 KB
#!/bin/bash # uv shim — wraps the real uv binary, captures pip install commands # to a Containerfile for reproducibility. # # Usage: # source /path/to/uv_shim.sh /path/to/Containerfile # # After sourcing, `uv pip install foo` will: # 1. Run the real uv pip install foo # 2. Append `RUN uv pip install foo` to the Containerfile # # To bypass the shim: use the full path to uv directly. _CONTAINERFILE="${1:-}" _REAL_UV="$(which uv 2>/dev/null)" if [ -z "$_REAL_UV" ]; then echo "uv_shim: ERROR — uv not found in PATH" >&2 return 1 2>/dev/null || exit 1 fi if [ -z "$_CONTAINERFILE" ]; then echo "uv_shim: ERROR — must specify Containerfile path" >&2 echo " Usage: source uv_shim.sh /path/to/Containerfile" >&2 return 1 2>/dev/null || exit 1 fi uv() { # Pass through to real uv "$_REAL_UV" "$@" local exit_code=$? # If it was a successful pip install, capture it if [ $exit_code -eq 0 ]; then # Check if this is a pip install command local is_pip_install=0 local args_after_install="" local seen_pip=0 local seen_install=0 for arg in "$@"; do if [ "$seen_install" -eq 1 ]; then # Skip flags like --system, --break-system-packages case "$arg" in --system|--break-system-packages|--quiet|-q) ;; *) args_after_install="$args_after_install $arg" ;; esac fi [ "$arg" = "pip" ] && seen_pip=1 [ "$seen_pip" -eq 1 ] && [ "$arg" = "install" ] && seen_install=1 && is_pip_install=1 done if [ "$is_pip_install" -eq 1 ] && [ -n "$args_after_install" ]; then local install_line="RUN uv pip install --system$args_after_install" # Check if this line already exists in the Containerfile if [ -f "$_CONTAINERFILE" ]; then if ! grep -qF "$install_line" "$_CONTAINERFILE"; then echo "$install_line" >> "$_CONTAINERFILE" echo "uv_shim: captured → $install_line" >&2 fi else echo "$install_line" >> "$_CONTAINERFILE" echo "uv_shim: created Containerfile with → $install_line" >&2 fi fi fi return $exit_code } echo "uv_shim: active — installs will be captured to $_CONTAINERFILE" -
__init__.py 26 B
# container-layer scripts -
__main__.py 30 B
from .cli import main main()
-
-
boot-ccotw.sh 2.7 KB
#!/bin/bash # Container-layer boot for Claude Code (Web + CLI) # Called by SessionStart hook. stdout goes into Claude's context. # # Ephemeral containers: bootstraps skill, restores cached layer or builds fresh. # Idempotent within a session via marker file. set -e MARKER="/tmp/.container-layer-booted" PROJECT_DIR="${CLAUDE_PROJECT_DIR:-.}" CONTAINERFILE="$(cd "$PROJECT_DIR" && pwd)/Containerfile" # ── Functions ── _output_skills() { local skills_dir="${1:-/mnt/skills/user}" [ -d "$skills_dir" ] || return 0 echo "" echo "Available skills:" for skill_dir in "$skills_dir"/*/; do local skill_file="${skill_dir}SKILL.md" if [ -f "$skill_file" ]; then local name=$(grep -m1 "^name:" "$skill_file" | sed 's/name: *//') [ -n "$name" ] && echo " - $name" fi done } _source_env() { # Source from project dir for envfile in "$PROJECT_DIR"/*.env "$PROJECT_DIR"/.env; do [ -f "$envfile" ] && { set -a; . "$envfile" 2>/dev/null; set +a; } done # Also check Claude.ai project files for envfile in /mnt/project/*.env; do [ -f "$envfile" ] && { set -a; . "$envfile" 2>/dev/null; set +a; } done } # ── Idempotency ── if [ -f "$MARKER" ]; then echo "Environment ready (cached)." _output_skills exit 0 fi # ── Main ── _source_env # Bootstrap the container-layer skill from GitHub SKILL_DIR="/tmp/_container_layer" if [ ! -f "$SKILL_DIR/scripts/containerfile.py" ]; then echo "Bootstrapping container-layer..." mkdir -p "$SKILL_DIR" if curl -sL "https://codeload.github.com/oaustegard/claude-skills/tar.gz/main" \ | tar -xz --strip-components=2 -C "$SKILL_DIR" "claude-skills-main/container-layer/" 2>/dev/null; then echo " ✓ container-layer skill loaded" else echo " ✗ bootstrap failed (check network/token)" exit 1 fi fi # Apply the Containerfile if [ -f "$CONTAINERFILE" ] && [ -f "$SKILL_DIR/scripts/containerfile.py" ]; then echo "Applying container layer: $CONTAINERFILE" INVALIDATE_ARGS="" if [ -n "$INVALIDATE_ON" ]; then for repo in $INVALIDATE_ON; do INVALIDATE_ARGS="$INVALIDATE_ARGS --invalidate-on $repo" done fi cd "$SKILL_DIR" python3 -m scripts.cli \ --token "${GH_TOKEN:-}" \ --repo "${LAYER_CACHE_REPO:-oaustegard/claude-container-layers}" \ $INVALIDATE_ARGS \ restore "$CONTAINERFILE" 2>&1 cd - > /dev/null echo "✓ Container layer applied" else echo "No Containerfile found at $CONTAINERFILE — skipping." fi touch "$MARKER" _output_skills # Custom post-boot hook [ -f "$PROJECT_DIR/post-boot.sh" ] && bash "$PROJECT_DIR/post-boot.sh" 2>&1 exit 0 -
boot.sh 2.9 KB
#!/bin/bash # Muninn boot via container-layer # Replaces the manual curl+tar+setup with a cacheable Containerfile. # # Flow: # 1. Bootstrap: fetch container-layer skill (tiny, ~30KB) # 2. Restore/build: container-layer handles skills fetch + pip installs, # using cached tarball from GitHub Releases when available # 3. Muninn init: source env vars, run Python boot (Turso, always live) # # Cache invalidation: includes claude-skills repo HEAD SHA in cache key, # so a skills push auto-triggers a full rebuild on next boot. set -e # ── Step 1: Bootstrap the container-layer skill ────────── # We need the skill itself before we can use it. This is the one # unavoidable fetch — everything else can be cached. LAYER_SKILL="/mnt/skills/user/container-layer" if [ ! -f "$LAYER_SKILL/scripts/containerfile.py" ]; then echo "Bootstrapping container-layer skill..." mkdir -p "$LAYER_SKILL" curl -sL "https://codeload.github.com/oaustegard/claude-skills/tar.gz/main" \ | tar -xz --strip-components=2 -C "$LAYER_SKILL" "claude-skills-main/container-layer/" 2>/dev/null \ || { echo "Bootstrap fetch failed"; exit 1; } fi # ── Step 2: Locate Containerfile ───────────────────────── # Look in project files first, then fall back to the skill's bundled default CONTAINERFILE="" for candidate in \ "/mnt/project/Containerfile" \ "$LAYER_SKILL/Containerfile"; do [ -f "$candidate" ] && CONTAINERFILE="$candidate" && break done if [ -z "$CONTAINERFILE" ]; then echo "ERROR: No Containerfile found" exit 1 fi # ── Step 3: Source credentials ─────────────────────────── set -a for envfile in /mnt/project/*.env; do [ -f "$envfile" ] && . "$envfile" 2>/dev/null done set +a # ── Step 4: Restore or build the layer ─────────────────── echo "Container layer: $CONTAINERFILE" cd "$LAYER_SKILL" python3 -m scripts.cli \ --token "$GH_TOKEN" \ --invalidate-on oaustegard/claude-skills \ restore "$CONTAINERFILE" # ── Step 5: Output available skills ────────────────────── echo "" echo "<available_skills source=\"/mnt/skills/user\">" for skill_dir in /mnt/skills/user/*/; do skill_file="${skill_dir}SKILL.md" if [ -f "$skill_file" ]; then name=$(grep -m1 "^name:" "$skill_file" | sed 's/name: *//') desc=$(grep -m1 "^description:" "$skill_file" | sed 's/description: *//') [ -n "$name" ] && echo "<skill><n>$name</n><description>$desc</description><location>${skill_file}</location></skill>" fi done echo "</available_skills>" # ── Step 6: Muninn boot (always live — queries Turso) ──── echo "" echo "Booting Muninn..." python3 << 'PYEOF' import os from scripts import boot print(boot(mode=os.environ.get('BOOT_MODE'))) PYEOF -
CHANGELOG.md 4.3 KB
# container-layer - Changelog All notable changes to the `container-layer` skill are documented in this file. The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/). ## [0.4.0] - 2026-08-25 ### Other - top skills: separate by omission, and correct the guidance that said otherwise (#777) ## [0.2.3] - 2026-08-25 ### Other - creating-skill: use Anthropic's quick_validate.py instead of a hand-rolled check (#775) - Deprecate mapping-codebases; adopt ruff 0.16.0 baseline (#747) - container-layer: fix ARG_MAX crash building large snapshot tarballs (#723) ## [0.2.2] - 2026-05-17 ### Other - container-layer v0.2.2: baseline /usr/{bin,lib,share} for apt-install layers (#653) ## [0.2.2] - 2026-05-17 ### Fixed - `PYTHON_INSTALL_PATHS` now includes `/usr/bin`, `/usr/lib`, `/usr/share` so apt-install layers get the diff-vs-baseline treatment instead of whole-tree capture. `_exec_run` already auto-adds these three paths to `snapshot_paths` when it sees `apt-get install`, but before this fix they weren't baselined, so any layer with `apt-get install <one tiny package>` ballooned by ~600MB compressed (the full `/usr/{bin,lib,share}` trees). Verified empirically against the libfuse2 install in oaustegard/claude-workspace-fuse: layer jumped from 940MB → 1625MB compressed. Closes oaustegard/claude-skills#652. ### Notes - Despite the name, `PYTHON_INSTALL_PATHS` is no longer python-only. Kept the name for backwards compatibility with any external importer; consider renaming to `BASELINED_PATHS` in a future major bump. - `snapshot_baseline()` walks these dirs at build-time. `/usr/lib` is the largest (~2GB raw) — adds a few seconds to baseline capture. Acceptable for build-time; negligible vs. the layer-restore time it saves. ## [0.2.1] - 2026-05-17 ### Other - container-layer v0.2.1: cover python3.10-3.13 dist-packages in baseline (#651) ## [0.2.1] - 2026-05-17 ### Fixed - `PYTHON_INSTALL_PATHS` now lists python3.10 through python3.13 dist-packages instead of only python3.12. The previous single entry missed containers using 3.11 (which is most of them), causing the diff-vs-baseline logic in `_dedup_paths` to silently fall through to whole-tree snapshot — every layer captured all of dist-packages instead of just newly-installed files. Empirically verified against oaustegard/claude-workspace-fuse: cached "slim" layer was 678MB compressed with 712M torch + 944M modular still embedded. ### Added - Regression tests in `scripts/test_baseline_paths.py` lock in the new coverage and assert the diff path is taken when a SNAPSHOT directive references a baselined dir. ## [0.2.0] - 2026-05-17 ### Added - add ISO date prefix to layer release names for sortability (#535) ### Other - container-layer v0.2.0: named layers + compose (#650) - container-layer: stop leaking GH_TOKEN via subprocess TimeoutExpired (#596) - Remove _MAP.md files, direct agents to tree-sitting for code navigation (#545) - marketplace: restructure as category-based plugins for Claude Code discovery (#530) - container-layer: add README - container-layer: update boot-ccotw.sh (fix function ordering, abs paths) - container-layer: update scripts/cli.py - container-layer: update scripts/containerfile.py - container-layer: add boot.sh - container-layer: add Containerfile ## [0.2.0] - 2026-05-17 ### Added - Named layers: `ContainerLayer(..., layer_name="X")` produces cache release tag `layer-X-<hash>` instead of `layer-<hash>`. Enables per-name cache retention policies and multi-layer composition without collisions. - `default_layer_name(path)`: derives layer name from filename (`Containerfile` → `base`, `Containerfile.mojo` → `mojo`, etc.). - `compose(containerfile_paths, ...)`: orchestrates sequential restore (or build+push on miss) of multiple named layers. Mirrors Docker's additive-overlay semantics — later layers can overwrite earlier ones' files. - CLI `compose` subcommand for the same. - `--name` flag on `build` / `restore` / `hash` / `inspect` subcommands for single-layer naming. ### Backwards compatibility - Unnamed layers (no `layer_name`, no `--name`) keep the old `layer-<hash>` tag. Existing single-Containerfile callers see no cache invalidation. - Existing `build_and_push()` / `restore_or_build()` / `build_only()` methods unchanged. ## [0.1.0] - 2026-04-04 ### Other - Add container-layer skill -
Containerfile 1.4 KB · in bundle
-
README.md 2.8 KB
# container-layer Custom, cached environment overlays for Claude's ephemeral containers. ## What This Does Claude.ai and Claude Code on the Web run in ephemeral containers — every session starts from a blank slate. This skill lets you declare your environment in a `Containerfile` (a Dockerfile subset), build it once, cache the result as a tarball in GitHub Releases, and restore it in seconds on subsequent sessions. **First session:** parse Containerfile → execute instructions → snapshot filesystem delta → push ~3 MB tarball to GitHub Releases. **Every session after:** download tarball → extract → done. One fetch replaces N installs. ## Components | File | Purpose | |------|---------| | `SKILL.md` | Skill metadata and documentation | | `Containerfile` | Default environment spec (edit this) | | `boot.sh` | Boot script for Claude.ai project instructions | | `boot-ccotw.sh` | Boot script for Claude Code (SessionStart hook) | | `scripts/containerfile.py` | Parser + executor with baseline diffing | | `scripts/layer_cache.py` | GitHub Releases tarball cache | | `scripts/cli.py` | CLI: `build`, `restore`, `hash`, `inspect` | | `scripts/uv_shim.sh` | Captures ad-hoc `uv pip install` to Containerfile | ## Supported Instructions ```dockerfile FETCH github:user/repo /dest # GitHub repo tarball FETCH github:user/repo@ref /dest # Specific ref RUN uv pip install --system pandas # Shell commands ENV KEY=value # Environment variables WORKDIR /path # Working directory SNAPSHOT /path # Include in cached layer ``` `FROM`, `EXPOSE`, `CMD`, `ENTRYPOINT`, etc. are silently ignored (Dockerfile compatibility). ## Smart Snapshotting The executor captures a filesystem baseline before building, then diffs against it — only new files from `pip install` / `uv pip install` are included in the tarball, not the entire `dist-packages` directory. `FETCH` destinations are captured in full. This keeps tarballs small (~3 MB for a full skills repo + Python packages). ## Cache Invalidation The cache key is a SHA-256 of the Containerfile contents. Pass `--invalidate-on user/repo` to include a GitHub repo's HEAD SHA in the key — when that repo gets a new commit, the cache auto-invalidates and triggers a rebuild. ```bash python3 -m scripts.cli --invalidate-on oaustegard/claude-skills restore ./Containerfile ``` ## Ad-Hoc Install Capture Source the uv shim to automatically append new installs to your Containerfile: ```bash source ./scripts/uv_shim.sh ./Containerfile uv pip install --system pandas # installs AND appends RUN line ``` ## Test Repo See [container-layer-test](https://github.com/oaustegard/container-layer-test) for a working example with Claude Code on the Web SessionStart hooks. -
SKILL.md 5.7 KB
--- name: container-layer description: >- Authors and caches a personalized container environment from a Dockerfile- like spec, as a single layer or as a composition of independently cached layers. Use when the user mentions "container layer", "Containerfile", "custom container", "cache my installs", "composable layers", "uv shim", or wants package installations, skills and environment config to survive an ephemeral session. Also for snapshotting, restoring or rebuilding that environment, and for capturing ad-hoc installs into a reproducible spec. metadata: version: 0.4.0 --- # Container Layer Build a reproducible, cached environment overlay for ephemeral containers using a Dockerfile-like spec. ## When NOT to use this skill This authors and caches a layer spec. It is not a Docker troubleshooting tool. | Situation | Use | |---|---| | A build is slow or failing | read the build log; this skill will not help | | Managing a running container | docker/podman directly | | Session boot sequence and hooks | the workspace's own boot docs | The tell is tense: this skill is for the environment you want next session, not the container you are fighting now. ## Concept The container resets every session, but your environment shouldn't. This skill: 1. Parses a `Containerfile` (Dockerfile subset) that declares your environment 2. Caches the built result as a tarball in GitHub Releases 3. Restores from cache on subsequent boots (single fetch vs. N installs) 4. Provides a `uv` shim that captures ad-hoc installs back into the Containerfile ## Supported Containerfile Instructions ```dockerfile # Environment variables ENV KEY=value # Shell commands (including package installs) RUN apt-get install -y foo # system packages RUN uv pip install pandas numpy # Python packages (preferred) RUN pip install requests # also works # Fetch files from URLs or GitHub FETCH https://example.com/file.tar.gz /dest/path FETCH github:user/repo /dest/path # latest tarball FETCH github:user/repo@ref /dest/path # specific ref # Set working directory for subsequent RUN commands WORKDIR /some/path # Declare paths to include in the cached layer snapshot # (auto-detected for FETCH destinations and pip/uv installs) SNAPSHOT /additional/path/to/capture # Ignored (Dockerfile compat, no-op here): # FROM, EXPOSE, CMD, ENTRYPOINT, LABEL, ARG, VOLUME, USER, SHELL ``` ## Usage ### Single layer — build / restore ```python from scripts.containerfile import ContainerLayer layer = ContainerLayer( containerfile_path="/path/to/Containerfile", cache_repo="oaustegard/claude-container-layers", # GitHub repo for release assets gh_token="...", ) # Try cache first, fall back to full build layer.restore_or_build() ``` Or via CLI: ```bash python -m scripts.cli restore /path/to/Containerfile --repo user/cache-repo ``` ### Multi-layer composition (v0.2.0+) Decompose a heavy environment into named layers, each cached independently. Compose them in order on session start so most-changed bits don't invalidate stable bits. ```python from scripts.containerfile import compose compose( containerfile_paths=[ "layers/Containerfile", # name='base' (always-on) "layers/Containerfile.scientific", # name='scientific' "layers/Containerfile.mojo", # name='mojo' ], cache_repo="user/cache-repo", ) ``` Each layer gets its own cache release tag `layer-<name>-<hash>` so retention policies (keep last N) and cache invalidation operate per-name. Default layer names are derived from the Containerfile path: - `Containerfile` → `base` - `Containerfile.scientific` → `scientific` - `layers/Containerfile.X` → `X` CLI equivalent: ```bash python -m scripts.cli compose \ layers/Containerfile \ layers/Containerfile.scientific \ layers/Containerfile.mojo \ --repo user/cache-repo ``` ### Per-layer name override If filename doesn't derive cleanly, pass `--name NAME:PATH` per layer: ```bash python -m scripts.cli compose \ --name base:weird-named-file.txt \ --name mojo:other-file.txt \ weird-named-file.txt other-file.txt ``` ### Single-layer naming (back-compat) `build` / `restore` / `hash` / `inspect` accept `--name`: ```bash python -m scripts.cli restore Containerfile.mojo --name mojo # Cache tag becomes 'layer-mojo-<hash>' instead of 'layer-<hash>'. # Omit --name to keep the old back-compat tag for existing callers. ``` ### The uv Shim After building, install the shim to capture future installs: ```bash source /path/to/container-layer/scripts/uv_shim.sh /path/to/Containerfile ``` Now `uv pip install foo` both installs the package AND appends `RUN uv pip install foo` to your Containerfile. ### Rebuilding the Cache After modifying the Containerfile: ```python layer.build_and_push() # Execute, snapshot, upload ``` ## Architecture Read `scripts/containerfile.py` for the parser/executor and `scripts/layer_cache.py` for the GitHub Releases caching logic. The cache key is a SHA-256 of the Containerfile contents — any change triggers a rebuild. ## Configuration The skill expects these environment variables (or pass as constructor args): - `GH_TOKEN` — GitHub token with `repo` scope (for releases) - Cache repo can be any repo the token has write access to ## Workflow Integration This skill is designed to be invoked from a boot script. Example Containerfile: ```dockerfile # Skills FETCH github:oaustegard/claude-skills /mnt/skills/user # Python environment RUN uv pip install --system pandas numpy requests # Path config RUN echo '/mnt/skills/user/remembering' > /usr/local/lib/python3.12/dist-packages/muninn-remembering.pth # Custom setup ENV MY_VAR=hello WORKDIR /home/claude ```
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.