docker-agent-deploy
Use this skill when exposing a Docker Agent as a server (MCP, HTTP API, A2A, ACP, or OpenAI-compatible chat), distributing an agent via an OCI registry with `docker agent share`, or measuring agent quality with `docker agent eval`. Even if the user just says they want to "turn my
Install
npx skills add https://github.com/docker/skills/tree/main/skills/docker-agent-deploy
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install docker-skills@llmmart
git clone https://github.com/docker/skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole docker/skills collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Docker Agent: Serving, Sharing, and Evaluating
Overview
This skill owns the integration surface of Docker Agent: making an agent
reachable by other software (docker agent serve), distributing it through
an OCI registry the way container images are distributed (docker agent share), and proving it still behaves after a change (docker agent eval).
It assumes the agent config already exists — see docker-agent-config for
authoring it, and docker-agent-run for interactive/local invocation.
When to use this skill
Activate this skill when:
- The user wants an agent reachable over MCP, an OpenAI-compatible chat endpoint, a plain HTTP API, or A2A/ACP.
- The user wants to publish an agent to Docker Hub (or any OCI registry) or pull one someone else published.
- The user wants automated evaluations (regression tests) for an agent, or wants to gate CI on eval results.
Do not use this skill when
Do not use this skill when:
- The task is authoring the agent.yaml itself (models, toolsets, sub_agents) — use
docker-agent-config. - The task is running the agent interactively on a developer's machine, choosing
--safety/--sandbox, or aliases — usedocker-agent-run.
Core guidance
Serving an agent
Five server modes, each with its own default loopback listen address — never expose any of them beyond loopback without authentication:
Mode Default listen Auth flag Has --safety?serve mcp127.0.0.1:8081--auth-token(only with--http)Yes (only with --http)serve api127.0.0.1:8080--auth-tokenNo serve chat127.0.0.1:8083--api-key/--api-key-envYes serve a2a127.0.0.1:8082--auth-tokenYes serve acp(stdio only) n/a No docker agent serve mcp ./agent.yaml --http --listen 127.0.0.1:9090 --auth-token "$TOKEN"serve mcpdefaults to stdio transport (for local clients like Claude Desktop); pass--httponly when you need a network-reachable MCP endpoint, and set--auth-tokenwhenever you do.Binding any server flag to a non-loopback address without an auth token/key is refused;
--insecure-no-authexists to force it and must be treated as a deliberate, documented exception, never a default.serve mcp(with--http),serve chat, andserve a2aexpose--safety(strict/balanced/restricted/autonomous); Docker's docs state it defaults torestrictedfor these modes when unset.serve apiandserve acpexpose no--safetyflag at all. Never raise--safetytoautonomouson a network-reachable listener; if a served agent must approve more, preferbalancedand keep auth enabled.serve apiaccepts a directory instead of a single file: every.yaml/.yml/.hclin it is exposed under/api/agents. Use--session-workingdir-rootto confine session working directories when the server is reachable by more than one user.
Sharing agents via OCI registries
- Push and pull agent configs the same way you push and pull images — same
registry, same
docker loginauth:docker agent share push ./agent.yaml docker.io/username/my-agent:latest docker agent share pull docker.io/username/my-agent:latest instruction_filecontents are inlined into the pushed artifact automatically, so a published agent stays self-contained — you do not need to bundle the referenced files separately.- Pin
sub_agentsthat reference the pushed artifact to a digest (name@sha256:...) once published, to avoid a per-run registry lookup and to guarantee the exact config a consumer gets. - Use
--forceonshare pullonly when you intend to overwrite a local copy that already exists; without it, an existing local config is left untouched.
Evaluating agents
- Evals live in an
evals/directory next to the agent config by default; each eval is one JSON session file capturing a user message, the recorded tool calls, and anevalsobject with the scoring criteria. - Create eval sessions from real conversations rather than hand-writing
JSON: run the agent interactively, then use the
/evalslash command in the TUI to save the session, and edit inrelevance/size/assertionscriteria afterward. - Four scoring dimensions: Tool Calls (F1 against the recorded sequence),
Relevance (LLM-judge,
--judge-model, defaultanthropic/claude-opus-5), Size (S/M/L/XL response-length bucket), and Assertions (deterministic checks; see the complete assertion-type list inreferences/eval-format.md). Prefer assertions overrelevancewhen a check can be exact: they need no judge model and are deterministic, not approximation-prone. - Evaluations run inside containers for isolation; a Docker-compatible
runtime is required. Dedicated provider API keys
(
ANTHROPIC_API_KEY/OPENAI_API_KEY) are forwarded automatically.GITHUB_TOKEN/GH_TOKENare not forwarded automatically (they're broad host credentials, not model keys) — pass them explicitly with-e GITHUB_TOKENwhen an agent's provider needs one (e.g.github-copilot). - Gate CI on regressions, not on absolute scores, with
--baseline:
A previously-passing eval that now fails always gates regardless of tolerance; cost changes are reported but never gate. A baseline or run with zero evaluations (e.g. andocker agent eval ./agent.yaml --baseline results/2026-08-01-run.json --regression-tolerance 0.05--onlypattern matching nothing) is rejected rather than reported as passing. - Use
--keep-containersplus your runtime'sexecto inspect a failed eval's container; the eval's.dbsession file holds the full conversation for offline debugging.
Verify
- After changing a served agent's config, re-run its evals with the same
explicit
--safetyvalue used in the deployment before restarting the listener — this catches an approval-policy regression before it reaches traffic. If a rollout must be rolled back, restore the prior config and safety flag; never restore an unauthenticated listener as a rollback shortcut.
Related skills
- For writing or changing the underlying
agent.yaml, usedocker-agent-config. - For local/interactive runs, safety-mode choice, and sandboxing, use
docker-agent-run.
References
references/eval-format.md— full eval session JSON schema and CLI flag table.references/sources.md— provenance of every rule in this skill.
Assets
assets/eval-session-example.json— a minimal eval session file to copy and adapt.
Checks
checks/verification.md— Verification runbook for serving, sharing, and evaluating an agent.
Files (skills)
-
agents
-
openai.yaml 402 B
interface: display_name: 'Docker Agent: Serving, Sharing, and Evaluating' short_description: Rules for exposing Docker Agents as servers, distributing them via OCI registries, and evaluating them for regressions in CI. default_prompt: Use this skill when serving a Docker Agent as an MCP/API/A2A/chat server, sharing it via a registry, or evaluating it. policy: allow_implicit_invocation: true
-
-
assets
-
eval-session-example.json 1.3 KB
{ "id": "41b179a2-ed19-4ae2-a45d-95775aaa90f7", "title": "Counting Files in Local Folder", "messages": [ { "message": { "message": { "role": "user", "content": "How many files in the local folder?" } } }, { "message": { "agent_name": "root", "message": { "role": "assistant", "tool_calls": [ { "id": "call_abc123", "type": "function", "function": { "name": "list_directory", "arguments": "{\"path\":\"./\"}" } } ] } } }, { "message": { "agent_name": "root", "message": { "role": "assistant", "content": "There are 2 files in the local folder: README.md and agent.yaml." } } } ], "evals": { "relevance": [ "The response mentions exactly 2 files", "The response lists README.md and agent.yaml" ], "assertions": [ { "name": "used list_directory", "type": "tool_called", "value": "list_directory" }, { "name": "no error message", "type": "not_contains", "value": "error" } ], "size": "S", "working_dir": "my-project" } }
-
-
checks
-
verification.md 2.3 KB
# Verification Runbook for serving, sharing, and evaluating ## 1. A server binds where and how you intend ```bash docker agent serve mcp ./agent.yaml --http --listen 127.0.0.1:9090 --auth-token "$TOKEN" & sleep 1 curl -s -o /dev/null -w "%{http_code}\n" http://127.0.0.1:9090/ # no token curl -s -o /dev/null -w "%{http_code}\n" -H "Authorization: Bearer $TOKEN" http://127.0.0.1:9090/ ``` Pass: the process reports it is listening on the given address; the request without the bearer token is rejected (non-2xx, typically 401), and the request with the correct token succeeds. Fail: it starts on a non-loopback address without `--auth-token`/`--insecure-no-auth` set explicitly, or the unauthenticated request succeeds — stop it, add/fix authentication, and restart. ## 2. A pushed agent round-trips ```bash docker agent share push ./agent.yaml docker.io/<user>/<name>:<tag> docker agent share pull docker.io/<user>/<name>:<tag> --force docker agent debug config docker.io/<user>/<name>:<tag> ``` Pass: the pulled config's resolved form matches the original (including any `instruction_file` contents, now inlined). Fail: a resolution error or missing instruction content — check for `instruction_file` paths outside the config directory, which are rejected on push. ## 3. Evals run and produce a report ```bash docker agent eval ./agent.yaml ./evals ``` Pass: console summary shows per-eval pass/fail plus Tool Calls / Relevance / Size / Assertions metrics, and a `results/` directory with JSON, `.db`, and log files. Fail: "No model is currently available" inside the eval container — use `docker-agent-run` for local credential/model troubleshooting, then check the eval-specific boundary: credentials must be set in the invoking shell. Dedicated model keys are forwarded automatically; `GITHUB_TOKEN`/`GH_TOKEN` are not — pass the one your provider needs explicitly (e.g. `-e GITHUB_TOKEN`). ## 4. A regression gate actually gates ```bash docker agent eval ./agent.yaml --baseline results/<prior-run>.json --regression-tolerance 0.05 ``` Pass: exit code 0 when quality is within tolerance of the baseline, non-zero when a previously-passing eval now fails or an aggregate rate drops beyond tolerance. Fail (gate never fails): confirm the baseline file actually contains evaluations — a baseline or run with none is rejected, not treated as passing.
-
-
references
-
eval-format.md 2.6 KB
# Eval session format and CLI flags ## Eval directory layout ``` my-agent/ ├── agent.yaml └── evals/ ├── <uuid>.json # one eval session └── results/ # auto-created output ├── <run-name>.json ├── <run-name>.log ├── <run-name>.db └── <run-name>-sessions.json ``` ## Eval session JSON schema Use `assets/eval-session-example.json` (relative to the skill root) for the complete example. The shape below is only a structural skeleton, not an eval to run: ```json { "id": "<session UUID>", "title": "<scenario title>", "messages": [], "evals": {} } ``` - `messages` holds the recorded conversation. A user entry nests the message as `message.message`; assistant entries also set `message.agent_name`. Recorded tool calls live in `message.message.tool_calls`. - `evals` holds scoring criteria and working-directory setup; see the fields below. Keep the concrete conversation and assertions in the asset rather than maintaining another copy here. ## `evals` object fields | Field | Type | Description | | --- | --- | --- | | `relevance` | string[] | Statements that must be true about the response; scored by the LLM judge | | `assertions` | object[] | Deterministic checks: `{name, type, value}` | | `size` | string | Expected response size: `S`, `M`, `L`, `XL` | | `working_dir` | string | Subdirectory under `evals/working_dirs/` mounted as the container's working dir | | `setup` | string | Shell script run in the container before the agent executes | ## Assertion types `contains`, `not_contains`, `equals`, `starts_with`, `ends_with`, `regex`, `cost_threshold` (dollar amount ceiling), `tool_called`. ## CLI flags | Flag | Default | Description | | --- | --- | --- | | `-c, --concurrency` | `16` | Concurrent evaluation runs | | `--judge-model` | `anthropic/claude-opus-5` | Model for relevance scoring | | `--output` | `<eval-dir>/results` | Results/logs/session-db directory | | `--only` | (all) | Only run evals matching these filename patterns | | `--base-image` | (default) | Custom base image for eval containers | | `--container-runtime` | `docker` | Runtime executable (e.g. `podman`) | | `--keep-containers` | `false` | Keep containers after evaluation for inspection | | `-e, --env` | (none) | Extra env vars to forward into the container | | `--repeat` | `1` | Repeat each eval k times; also unlocks `pass@k`/`pass^k` | | `--baseline` | (none) | Prior run JSON to regress-test against | | `--regression-tolerance` | `0` | Aggregate-rate drop allowed before `--baseline` fails | Source: https://docs.docker.com/ai/docker-agent/features/evaluation/. -
sources.md 1.6 KB
# Sources - `docker agent serve --help`, `docker agent serve mcp --help`, `docker agent serve api --help`, `docker agent serve chat --help`, `docker agent serve a2a --help`, `docker agent serve acp --help`, `docker agent share --help`, `docker agent share push --help`, `docker agent share pull --help`, `docker agent eval --help` — verified locally against docker-agent as shipped with Docker CLI 29.7.2. - https://docs.docker.com/ai/docker-agent/features/cli/ — full serve mode flag tables (default listen addresses, `--safety` default `restricted`, `--auth-token`/`--api-key`, `--session-workingdir-root`), and `share push`/`share pull` reference including signing (`--key`, `--encrypt`). - https://docs.docker.com/ai/docker-agent/features/mcp-mode/ — MCP server mode setup and stdio vs `--http` transport. - https://docs.docker.com/ai/docker-agent/features/evaluation/ — eval directory structure, eval session JSON schema, scoring metrics (Tool Calls F1, Relevance, Size, Assertions), `--baseline`/`--regression-tolerance` gate rules, provider-credential forwarding into eval containers. - https://docs.docker.com/ai/docker-agent/concepts/distribution/ — agent distribution over OCI registries, `instruction_file` inlining on push, signing/encrypting agents. > **Version note:** the docs above describe optional `--key`/`--encrypt` signing flags for `share push`/`share pull`. The locally installed `docker agent share push --help` / `docker agent share pull --help` (Docker CLI 29.7.2) show no such flags — do not tell a user to pass `--key`/`--encrypt` against this version; verify with `docker agent share push --help` before relying on signing.
-
-
SKILL.md 7.6 KB
--- name: docker-agent-deploy description: Use this skill when exposing a Docker Agent as a server (MCP, HTTP API, A2A, ACP, or OpenAI-compatible chat), distributing an agent via an OCI registry with `docker agent share`, or measuring agent quality with `docker agent eval`. Even if the user just says they want to "turn my agent into an MCP server", "let Claude Desktop use my agent", "publish my agent to Docker Hub", "push my agent like an image", or "test my agent in CI", this skill applies. Covers `serve mcp/api/a2a/acp/chat` listen addresses and auth flags, `share push/pull`, eval session JSON format, scoring metrics, and the `--baseline` regression gate. license: Apache-2.0 compatibility: Requires the docker-agent CLI plugin (Docker Desktop 4.63+, or standalone). `docker agent eval` additionally requires a Docker-compatible container runtime (Docker Desktop/Engine, or Podman via `--container-runtime`). Verified against docker-agent as shipped with Docker CLI 29.7.2. --- # Docker Agent: Serving, Sharing, and Evaluating ## Overview This skill owns the integration surface of Docker Agent: making an agent reachable by other software (`docker agent serve`), distributing it through an OCI registry the way container images are distributed (`docker agent share`), and proving it still behaves after a change (`docker agent eval`). It assumes the agent config already exists — see `docker-agent-config` for authoring it, and `docker-agent-run` for interactive/local invocation. ## When to use this skill Activate this skill when: - The user wants an agent reachable over MCP, an OpenAI-compatible chat endpoint, a plain HTTP API, or A2A/ACP. - The user wants to publish an agent to Docker Hub (or any OCI registry) or pull one someone else published. - The user wants automated evaluations (regression tests) for an agent, or wants to gate CI on eval results. ## Do not use this skill when Do not use this skill when: - The task is authoring the agent.yaml itself (models, toolsets, sub_agents) — use `docker-agent-config`. - The task is running the agent interactively on a developer's machine, choosing `--safety`/`--sandbox`, or aliases — use `docker-agent-run`. ## Core guidance ### Serving an agent - Five server modes, each with its own default loopback listen address — never expose any of them beyond loopback without authentication: | Mode | Default listen | Auth flag | Has `--safety`? | | --- | --- | --- | --- | | `serve mcp` | `127.0.0.1:8081` | `--auth-token` (only with `--http`) | Yes (only with `--http`) | | `serve api` | `127.0.0.1:8080` | `--auth-token` | No | | `serve chat` | `127.0.0.1:8083` | `--api-key` / `--api-key-env` | Yes | | `serve a2a` | `127.0.0.1:8082` | `--auth-token` | Yes | | `serve acp` | (stdio only) | n/a | No | ```bash docker agent serve mcp ./agent.yaml --http --listen 127.0.0.1:9090 --auth-token "$TOKEN" ``` - `serve mcp` defaults to stdio transport (for local clients like Claude Desktop); pass `--http` only when you need a network-reachable MCP endpoint, and set `--auth-token` whenever you do. - Binding any server flag to a non-loopback address without an auth token/key is refused; `--insecure-no-auth` exists to force it and must be treated as a deliberate, documented exception, never a default. - `serve mcp` (with `--http`), `serve chat`, and `serve a2a` expose `--safety` (`strict`/`balanced`/`restricted`/`autonomous`); Docker's docs state it defaults to `restricted` for these modes when unset. `serve api` and `serve acp` expose no `--safety` flag at all. Never raise `--safety` to `autonomous` on a network-reachable listener; if a served agent must approve more, prefer `balanced` and keep auth enabled. - `serve api` accepts a directory instead of a single file: every `.yaml`/`.yml`/`.hcl` in it is exposed under `/api/agents`. Use `--session-workingdir-root` to confine session working directories when the server is reachable by more than one user. ### Sharing agents via OCI registries - Push and pull agent configs the same way you push and pull images — same registry, same `docker login` auth: ```bash docker agent share push ./agent.yaml docker.io/username/my-agent:latest docker agent share pull docker.io/username/my-agent:latest ``` - `instruction_file` contents are inlined into the pushed artifact automatically, so a published agent stays self-contained — you do not need to bundle the referenced files separately. - Pin `sub_agents` that reference the pushed artifact to a digest (`name@sha256:...`) once published, to avoid a per-run registry lookup and to guarantee the exact config a consumer gets. - Use `--force` on `share pull` only when you intend to overwrite a local copy that already exists; without it, an existing local config is left untouched. ### Evaluating agents - Evals live in an `evals/` directory next to the agent config by default; each eval is one JSON session file capturing a user message, the recorded tool calls, and an `evals` object with the scoring criteria. - Create eval sessions from real conversations rather than hand-writing JSON: run the agent interactively, then use the `/eval` slash command in the TUI to save the session, and edit in `relevance`/`size`/`assertions` criteria afterward. - Four scoring dimensions: Tool Calls (F1 against the recorded sequence), Relevance (LLM-judge, `--judge-model`, default `anthropic/claude-opus-5`), Size (S/M/L/XL response-length bucket), and Assertions (deterministic checks; see the complete assertion-type list in `references/eval-format.md`). Prefer assertions over `relevance` when a check can be exact: they need no judge model and are deterministic, not approximation-prone. - Evaluations run inside containers for isolation; a Docker-compatible runtime is required. Dedicated provider API keys (`ANTHROPIC_API_KEY`/`OPENAI_API_KEY`) are forwarded automatically. `GITHUB_TOKEN`/`GH_TOKEN` are **not** forwarded automatically (they're broad host credentials, not model keys) — pass them explicitly with `-e GITHUB_TOKEN` when an agent's provider needs one (e.g. `github-copilot`). - Gate CI on regressions, not on absolute scores, with `--baseline`: ```bash docker agent eval ./agent.yaml --baseline results/2026-08-01-run.json --regression-tolerance 0.05 ``` A previously-passing eval that now fails always gates regardless of tolerance; cost changes are reported but never gate. A baseline or run with zero evaluations (e.g. an `--only` pattern matching nothing) is rejected rather than reported as passing. - Use `--keep-containers` plus your runtime's `exec` to inspect a failed eval's container; the eval's `.db` session file holds the full conversation for offline debugging. ### Verify - After changing a served agent's config, re-run its evals with the same explicit `--safety` value used in the deployment before restarting the listener — this catches an approval-policy regression before it reaches traffic. If a rollout must be rolled back, restore the prior config and safety flag; never restore an unauthenticated listener as a rollback shortcut. ## Related skills - For writing or changing the underlying `agent.yaml`, use `docker-agent-config`. - For local/interactive runs, safety-mode choice, and sandboxing, use `docker-agent-run`. ## References - `references/eval-format.md` — full eval session JSON schema and CLI flag table. - `references/sources.md` — provenance of every rule in this skill. ## Assets - `assets/eval-session-example.json` — a minimal eval session file to copy and adapt. ## Checks - `checks/verification.md` — Verification runbook for serving, sharing, and evaluating an agent. -
skill.yaml 923 B
schema: v1 id: docker-agent-deploy version: 0.1.1 title: 'Docker Agent: Serving, Sharing, and Evaluating' description: Rules for exposing Docker Agents as servers, distributing them via OCI registries, and evaluating them for regressions in CI. owns: - docker-agent-serve - docker-agent-share - docker-agent-eval - evals/*.json use_when: - The task is exposing an agent over MCP, an HTTP API, an OpenAI-compatible chat endpoint, or A2A/ACP. - The task is publishing an agent to or pulling one from an OCI registry with docker agent share. - The task is writing, running, or gating CI on docker agent eval sessions. do_not_use_when: - The main task is authoring or editing the agent.yaml itself (models, toolsets, sub_agents). - The main task is running an agent interactively on a developer machine, or choosing --safety/--sandbox for local use. delegates_to: - docker-agent-config - docker-agent-run
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.