agent-safety
Use when bounding an LLM agent that already runs — scoping its task domain, gating tools to least privilege, defending against prompt injection in untrusted web/email/RAG text, requiring human approval on irreversible actions, capping runtime and cost, or triaging what it already
Install
npx skills add https://github.com/ericrisco/rsc-harness/tree/main/skills/agent-safety
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install ericrisco-rsc-harness@llmmart
git clone https://github.com/ericrisco/rsc-harness.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole ericrisco/rsc-harness collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Agent safety
You are the security review for an agent's agency, not for its code. The loop works,
tools are wired, memory persists — your job is to make that autonomy bounded. If you
want to review ordinary endpoints, auth, or secrets handling, that is
../secure-coding/SKILL.md — this skill is the Agentic Top 10, the risks that exist only
because a model has tools and autonomy. If the loop or tools do not exist yet, that is
../building-agents/SKILL.md. You arrive after both.
references/threat-model.md carries the OWASP Agentic Top 10 2026 risks mapped to the
controls below, the pre-ship guardrail checklist, and the incident-triage flow for "the
agent did X" — open it when you are reviewing before ship or reconstructing an incident.
The ownership split
Agent security splits into four layers — Model · Harness · Tools · Environment. The model provider owns only the Model layer (alignment, refusals). Everything else is yours: the Harness (loop, memory, context assembly), the Tools (what the agent can do), and the Environment (creds, network, blast radius). Do not outsource a layer you own to "the model is aligned."
Three excesses cause almost every agentic incident. Cut all three:
- Excessive functionality — tools the task never needs.
- Excessive permissions — broader scopes/creds than the tool needs.
- Excessive autonomy — acting without checking back when it should.
The operating principle is least agency: autonomy is earned per task, not defaulted.
Scope limits
- Declare the allowed task domain as a hard boundary in the system prompt. Why: an undeclared scope is an infinite scope; "you are a refund assistant; you do not touch payroll" is a constraint a reviewer can check.
- Deny by default — the agent starts with zero tools. Each tool earns its place by a task justification. Why: an opt-out tool list grows; an opt-in list stays minimal.
- Segregate the instruction channel from the data channel. System/developer prompt = trusted instructions. Everything the agent reads at runtime = data, never instructions. Why: this single boundary is what stops indirect injection (LLM01).
Tool gating / least agency
Give every tool a profile: read / write / exec / send, the exact resources it may touch, and an allowlist (never a wildcard). Block destructive flags and secret paths at the tool boundary, not in the prompt — the prompt is advisory, the boundary is enforced.
# Bad: one wildcard tool = unbounded blast radius, runs anything the loop emits
def run_shell(cmd: str) -> str:
return subprocess.run(cmd, shell=True, capture_output=True, text=True).stdout
# Good: narrow tool, allowlisted root, denied patterns, no shell
ALLOWED_ROOT = pathlib.Path("/srv/agent/workspace").resolve()
DENY = ("*.key", "*.pem", "*secret*", "*.env", "id_rsa*")
def read_file(path: str) -> str:
p = (ALLOWED_ROOT / path).resolve()
if not p.is_relative_to(ALLOWED_ROOT): # no traversal out of scope
raise PermissionError("path outside workspace")
if any(p.match(g) for g in DENY): # never read secrets
raise PermissionError("denied pattern")
return p.read_text()
- Issue task-scoped, short-lived tokens — not the session's broad creds. A credential should be valid only for the specific tool and the duration of one task. Why: a hijacked loop cannot reuse a session-wide token it never held.
- Prefer read-only by default; writes/sends/exec are separate, gated tools. Why: most steps only need to read, so most steps should be unable to mutate anything.
Injection defense
Treat all external data as untrusted: user messages, retrieved documents, API responses, emails, web pages, other agents' output. Sanitize and delimit before it enters context, and never let external text reach a privileged tool unmediated.
| Source | Trust level | Required mediation before it can act |
|---|---|---|
| System / developer prompt | Trusted | none (this is the only instruction channel) |
| End-user chat message | Untrusted | delimit; treat as data, not commands |
| Retrieved RAG / KB document | Untrusted | delimit; strip instruction-like spans |
| Fetched web page / API JSON | Untrusted | parse to schema; no raw text → tool args |
| Inbound email / ticket body | Untrusted | delimit; HITL on any action it requests |
| Another agent's message | Untrusted | same as external user input |
# Bad: retrieved chunk flows straight into a privileged action
chunk = retriever.search(q)[0].text # attacker-controlled doc
agent.call_tool("send_email", to=extract_to(chunk), body=chunk)
# Good: external content is quarantined data; the action is schema-validated + gated
chunk = retriever.search(q)[0].text
ctx = f"<retrieved untrusted>\n{chunk}\n</retrieved untrusted>" # delimited, labeled
proposal = agent.draft("send_email", context=ctx) # model proposes
args = SendEmail.model_validate(proposal.args) # schema or reject
if args.to_domain not in ALLOWED_DOMAINS: # exfil guard
raise PermissionError("recipient outside allowlist")
require_human_approval("send_email", args) # irreversible → HITL
- Validate every tool-call argument against a strict schema before execution. Why: a schema rejects the surprise field, encoded payload, or off-allowlist recipient injection produces.
- Watch for exfiltration shapes — unexpected outbound URLs, base64 blobs, recipients outside the allowlist. Why: data theft is the common payload of a successful injection.
Human-in-the-loop by risk class
Do not approve every action — reported ~93% of permission prompts get approved without being read, so blanket prompting trains a rubber stamp. Gate by risk class, keyed on reversibility × blast radius. Bind each approval to the exact parameters with a short-lived token so the approved action cannot be swapped after the click.
| Action type (examples) | Reversible? | Blast radius | Control |
|---|---|---|---|
| Read file, search KB, fetch page | n/a | none | auto |
| Write to scratch workspace, internal draft | yes | local | log-only |
| Mutate prod DB, deploy, change config | hard | system | approve (HITL) |
| Send email/payment to external party, post live | no | external | approve (HITL) |
| Delete backups, rotate prod creds, mass-delete | no | catastrophic | block (or step-up auth) |
- Step up for the top row — high-value irreversible actions deserve fresh auth, not the ambient session. Why: a hijacked session should not also hold the keys to the worst action.
Memory hygiene
- Validate and sanitize content before it is stored. Why: memory poisoning persists across sessions (OWASP Agentic T1) — unlike session-scoped injection, a poisoned memory re-attacks every future run until purged.
- Isolate memory per user and per session; do not let one user's writes color another's reads. Why: shared memory is a cross-tenant injection channel.
- Expire entries and cap memory size; redact PII (SSN, cards, API keys) before persist. Why: stale instructions and leaked secrets both age into liabilities.
Runtime kill-switches
A looping or hijacked agent must hit a wall on its own. Set hard caps, fail closed:
- Tool-call rate cap (e.g. ~30 calls/min) — runaway loops trip it before they do damage.
- Cost cap per session (e.g. ~$10) — a wallet attack stops at a known ceiling.
- Loop / step cap — a fixed max iterations kills the infinite plan.
- Wall-clock timeout — a stuck agent is terminated, not left running.
Log every tool call (arguments redacted) and alert on repeated approval-bypass attempts.
Anti-patterns
| Anti-pattern | Why it bites | Do instead |
|---|---|---|
| Approve every action | ~93% rubber-stamped; the real risky one slips through | Gate by risk class; HITL only on irreversible/external |
| One broad session token shared by all tools | Hijacked loop reuses it everywhere | Task-scoped, short-lived per-tool tokens |
| Trust RAG / fetched / email content | Indirect injection (LLM01) becomes direct tool execution | Delimit as untrusted data; mediate before any tool |
Wildcard run_shell(cmd) tool |
Unbounded blast radius | Narrow tools, allowlisted resources, denied patterns |
| Raw external text piped into tool args | Attacker controls the action's parameters | Schema-validate args; allowlist recipients/domains |
| Scope/limits stated only in the prompt | Prompt is advisory; the model can be talked out of it | Enforce at the tool/harness boundary |
| No loop / cost / rate cap | A hijacked or looping agent runs until it runs out of money | Hard fail-closed kill-switches |
| Redact PII only in the UI | The secret was already written to memory/logs | Redact before persistence, at the source |
Files (rsc-harness)
-
evals
-
cases.yaml 3.2 KB
skill: agent-safety should_trigger: - prompt: "Our support agent can call any shell command — lock it down." why: Wildcard exec tool with unbounded blast radius; core tool-gating / least-agency case. - prompt: "An email in the inbox told the agent to forward all invoices to an external address and it did." why: Indirect prompt injection (LLM01). Non-obvious — reads like a bug report, but the cause is untrusted email content reaching a privileged send tool. - prompt: "Before we ship this autonomous agent to prod, review what it's allowed to do." why: Pre-ship review of scope, permissions, and HITL — exactly the bounded-autonomy review this skill is. - prompt: "Posa límits a l'agent perquè no pugui esborrar res sense aprovació." why: Catalan. HITL-by-risk-class on an irreversible delete plus scope limiting. - prompt: "The agent's memory keeps repeating a wrong instruction across sessions." why: Memory poisoning (OWASP Agentic T1). Non-obvious — sounds like a quality bug, is actually contaminated persistent memory. - prompt: "Review our MCP server's tool registry — the scopes look way too broad." why: Over-broad tool scopes / excessive permissions; deny-by-default and per-tool allowlists. should_not_trigger: - prompt: "Help me build the agent's tool-calling loop." route_to: building-agents why: Construction of the agent (loop/tools), not constraining one that already runs. - prompt: "Is this Express auth endpoint vulnerable to broken access control?" route_to: secure-coding why: Plain web/app OWASP on an endpoint; no model, tools, or autonomy involved. - prompt: "Measure whether my agent answers correctly with a golden set." route_to: agent-eval why: Quality / correctness scoring, not safety constraint. - prompt: "Write a sharper system prompt so the agent is more helpful." route_to: prompt-engineering why: Capability tuning, not a safety boundary. - prompt: "Show me how much each model call costs this month." route_to: cost-tracking why: Reporting spend; agent-safety sets the cost *cap*, it does not report the analytics. capability: - scenario: > A research agent has web-fetch, send-email, and file-write tools plus persistent memory shared across sessions. Design the guardrails that make its autonomy bounded. must_include: - Deny-by-default tools with a per-tool profile (read web vs send email vs write file); no wildcard tool. - Treats web-fetched content as untrusted; delimits it as data; never pipes external text straight into send-email args. - HITL approval for send-email bound to exact parameters (irreversible, external blast radius); recipient/domain allowlist as exfil guard. - File-write scoped to an allowlisted workspace root; blocks secret patterns (*.key/*.env/*secret*) and path traversal. - Memory validated before store, isolated per session, expiring, PII-redacted before persist. - Runtime kill-switches: tool-call rate cap, per-session cost cap, loop/step cap, wall-clock timeout, all failing closed. - Task-scoped short-lived tokens rather than one broad session credential. - Maps each control to OWASP Agentic Top 10 (indirect injection / tool misuse / memory poisoning / excessive agency). -
README.md 909 B
# Evals: agent-safety These cases are the trigger contract and capability rubric for the `agent-safety` skill. `should_trigger` and `should_not_trigger` assert that the description routes correctly — each negative names the real sibling it should route to instead (building-agents, secure-coding, agent-eval, prompt-engineering, cost-tracking). Run them through the repo's eval harness if one is wired up, or read them directly as the routing spec when reviewing the description. The `capability` block is not a pass/fail script: it is a rubric a human or an LLM-as-judge scores a sample guardrail design against — the design must cover every `must_include` item (deny-by-default tools, untrusted-content mediation, HITL on irreversible actions, memory hygiene, runtime caps) and map them to the OWASP Agentic Top 10. This is a process/review skill, so there is no `verify.sh`; rigor lives in this eval.
-
-
references
-
threat-model.md 4 KB
# Threat model: OWASP Agentic Top 10 (2026) → controls Mapping of the named agentic risks to the control in `../SKILL.md` that mitigates each. Sources: OWASP Gen AI Security Project, "OWASP Top 10 for Agentic Applications for 2026" (genai.owasp.org); OWASP Top 10 for LLM Applications v2025; OWASP AI Agent Security Cheat Sheet (cheatsheetseries.owasp.org). All accessed 2026-06-02. | Agentic risk | What it is | Control in this skill | | ---------------------------------- | ----------------------------------------------------------------------- | ------------------------------------------------------- | | Memory poisoning (T1) | Contamination persists across sessions, re-attacking every future run | Memory hygiene: validate before store, isolate, expire | | Tool misuse & exploitation | Abuse via unsafe composition / recursion / excessive calls *with* valid perms | Tool gating + runtime kill-switches (rate/loop caps) | | Privilege compromise | Broad or stolen creds let the loop act beyond its task | Task-scoped short-lived tokens; least agency | | Indirect prompt injection (LLM01) | Untrusted content carries instructions the model executes | Trust-boundary table; segregate data from instructions | | Excessive agency | Acting without check-back on irreversible/external actions | HITL by risk class; deny-by-default scope | | Data exfiltration | Stolen data leaves via tool output / outbound calls | Output schema validation; exfil-shape detection | The OWASP framing of root causes — excessive **functionality**, **permissions**, **autonomy** — maps one-to-one onto the three excesses in the SKILL body. Cutting all three is the deny-by-default + least-agency posture. ## Pre-ship guardrail checklist Before an autonomous loop reaches production, confirm: - [ ] Allowed task domain is declared in the system prompt as a hard boundary. - [ ] Tools are deny-by-default; each enabled tool has a written task justification. - [ ] Every tool has a profile (read / write / exec / send) and an allowlist, not a wildcard. - [ ] No tool can read `*.key` / `*.pem` / `*secret*` / `*.env` or run with destructive flags. - [ ] Creds are task-scoped and short-lived, not one broad session token. - [ ] All external content (web, email, RAG, API, other agents) is delimited as untrusted. - [ ] Tool-call args are schema-validated; recipients/domains are allowlisted (exfil guard). - [ ] HITL is gated by risk class; irreversible/external actions require approval bound to exact params. - [ ] Memory is validated before store, isolated per user/session, expiring, PII-redacted. - [ ] Runtime caps exist and fail closed: rate, cost, loop/step, wall-clock timeout. - [ ] Every tool call is logged (redacted); repeated approval-bypass alerts. ## Incident triage: "the agent did X" When an agent has already done something it should not have, work the loop in order: 1. **Contain.** Revoke the task tokens / creds the loop is holding; trip the kill-switch (pause the loop). Stop the bleeding before diagnosing. 2. **Find the trust boundary it crossed.** Which untrusted source reached a privileged tool? Trace the action's parameters back to their origin — usually an email body, RAG chunk, fetched page, or a poisoned memory entry. 3. **Purge poisoned memory.** If the bad behavior repeats across sessions, the cause is stored, not in-context. Delete the contaminated entries; do not just restart the session. 4. **Add the missing gate.** Reclassify that action's risk, move it behind HITL or block, tighten the tool allowlist, or add the schema/exfil check that would have caught it. 5. **Confirm with a replay.** Re-run the triggering input against the new guardrail and verify the action is now refused or escalated.
-
-
SKILL.md 10.3 KB
--- name: agent-safety description: "Use when bounding an LLM agent that already runs — scoping its task domain, gating tools to least privilege, defending against prompt injection in untrusted web/email/RAG text, requiring human approval on irreversible actions, capping runtime and cost, or triaging what it already did. NOT building the loop, tools, or RAG (that is `building-agents`)." tags: [agent-security, guardrails, prompt-injection, least-privilege, owasp-agentic] recommends: [building-agents, secure-coding, agent-eval] origin: risco --- # Agent safety You are the security review for an agent's **agency**, not for its code. The loop works, tools are wired, memory persists — your job is to make that autonomy *bounded*. If you want to review ordinary endpoints, auth, or secrets handling, that is `../secure-coding/SKILL.md` — this skill is the *Agentic* Top 10, the risks that exist only because a model has tools and autonomy. If the loop or tools do not exist yet, that is `../building-agents/SKILL.md`. You arrive *after* both. `references/threat-model.md` carries the OWASP Agentic Top 10 2026 risks mapped to the controls below, the pre-ship guardrail checklist, and the incident-triage flow for "the agent did X" — open it when you are reviewing before ship or reconstructing an incident. ## The ownership split Agent security splits into four layers — **Model · Harness · Tools · Environment**. The model provider owns only the Model layer (alignment, refusals). Everything else is yours: the Harness (loop, memory, context assembly), the Tools (what the agent can *do*), and the Environment (creds, network, blast radius). Do not outsource a layer you own to "the model is aligned." Three excesses cause almost every agentic incident. Cut all three: - **Excessive functionality** — tools the task never needs. - **Excessive permissions** — broader scopes/creds than the tool needs. - **Excessive autonomy** — acting without checking back when it should. The operating principle is **least agency**: autonomy is earned per task, not defaulted. ## Scope limits - **Declare the allowed task domain as a hard boundary in the system prompt.** Why: an undeclared scope is an infinite scope; "you are a refund assistant; you do not touch payroll" is a constraint a reviewer can check. - **Deny by default — the agent starts with zero tools.** Each tool earns its place by a task justification. Why: an opt-out tool list grows; an opt-in list stays minimal. - **Segregate the instruction channel from the data channel.** System/developer prompt = trusted instructions. Everything the agent reads at runtime = data, never instructions. Why: this single boundary is what stops indirect injection (LLM01). ## Tool gating / least agency Give every tool a profile: **read / write / exec / send**, the exact resources it may touch, and an **allowlist** (never a wildcard). Block destructive flags and secret paths at the tool boundary, not in the prompt — the prompt is advisory, the boundary is enforced. ```python # Bad: one wildcard tool = unbounded blast radius, runs anything the loop emits def run_shell(cmd: str) -> str: return subprocess.run(cmd, shell=True, capture_output=True, text=True).stdout ``` ```python # Good: narrow tool, allowlisted root, denied patterns, no shell ALLOWED_ROOT = pathlib.Path("/srv/agent/workspace").resolve() DENY = ("*.key", "*.pem", "*secret*", "*.env", "id_rsa*") def read_file(path: str) -> str: p = (ALLOWED_ROOT / path).resolve() if not p.is_relative_to(ALLOWED_ROOT): # no traversal out of scope raise PermissionError("path outside workspace") if any(p.match(g) for g in DENY): # never read secrets raise PermissionError("denied pattern") return p.read_text() ``` - **Issue task-scoped, short-lived tokens — not the session's broad creds.** A credential should be valid only for the specific tool and the duration of one task. Why: a hijacked loop cannot reuse a session-wide token it never held. - **Prefer read-only by default; writes/sends/exec are separate, gated tools.** Why: most steps only need to read, so most steps should be unable to mutate anything. ## Injection defense Treat **all** external data as untrusted: user messages, retrieved documents, API responses, emails, web pages, other agents' output. Sanitize and delimit before it enters context, and **never let external text reach a privileged tool unmediated.** | Source | Trust level | Required mediation before it can act | | ------------------------------ | ----------- | ------------------------------------------- | | System / developer prompt | Trusted | none (this is the only instruction channel) | | End-user chat message | Untrusted | delimit; treat as data, not commands | | Retrieved RAG / KB document | Untrusted | delimit; strip instruction-like spans | | Fetched web page / API JSON | Untrusted | parse to schema; no raw text → tool args | | Inbound email / ticket body | Untrusted | delimit; HITL on any action it requests | | Another agent's message | Untrusted | same as external user input | ```python # Bad: retrieved chunk flows straight into a privileged action chunk = retriever.search(q)[0].text # attacker-controlled doc agent.call_tool("send_email", to=extract_to(chunk), body=chunk) ``` ```python # Good: external content is quarantined data; the action is schema-validated + gated chunk = retriever.search(q)[0].text ctx = f"<retrieved untrusted>\n{chunk}\n</retrieved untrusted>" # delimited, labeled proposal = agent.draft("send_email", context=ctx) # model proposes args = SendEmail.model_validate(proposal.args) # schema or reject if args.to_domain not in ALLOWED_DOMAINS: # exfil guard raise PermissionError("recipient outside allowlist") require_human_approval("send_email", args) # irreversible → HITL ``` - **Validate every tool-call argument against a strict schema before execution.** Why: a schema rejects the surprise field, encoded payload, or off-allowlist recipient injection produces. - **Watch for exfiltration shapes** — unexpected outbound URLs, base64 blobs, recipients outside the allowlist. Why: data theft is the common payload of a successful injection. ## Human-in-the-loop by risk class Do **not** approve every action — reported ~93% of permission prompts get approved without being read, so blanket prompting trains a rubber stamp. Gate by **risk class**, keyed on reversibility × blast radius. Bind each approval to the **exact parameters** with a short-lived token so the approved action cannot be swapped after the click. | Action type (examples) | Reversible? | Blast radius | Control | | ----------------------------------------------- | ----------- | ------------ | ------------------ | | Read file, search KB, fetch page | n/a | none | **auto** | | Write to scratch workspace, internal draft | yes | local | **log-only** | | Mutate prod DB, deploy, change config | hard | system | **approve (HITL)** | | Send email/payment to external party, post live | no | external | **approve (HITL)** | | Delete backups, rotate prod creds, mass-delete | no | catastrophic | **block** (or step-up auth) | - **Step up for the top row** — high-value irreversible actions deserve fresh auth, not the ambient session. Why: a hijacked session should not also hold the keys to the worst action. ## Memory hygiene - **Validate and sanitize content before it is stored.** Why: memory poisoning persists across sessions (OWASP Agentic T1) — unlike session-scoped injection, a poisoned memory re-attacks every future run until purged. - **Isolate memory per user and per session; do not let one user's writes color another's reads.** Why: shared memory is a cross-tenant injection channel. - **Expire entries and cap memory size; redact PII (SSN, cards, API keys) before persist.** Why: stale instructions and leaked secrets both age into liabilities. ## Runtime kill-switches A looping or hijacked agent must hit a wall on its own. Set hard caps, fail closed: - **Tool-call rate cap** (e.g. ~30 calls/min) — runaway loops trip it before they do damage. - **Cost cap per session** (e.g. ~$10) — a wallet attack stops at a known ceiling. - **Loop / step cap** — a fixed max iterations kills the infinite plan. - **Wall-clock timeout** — a stuck agent is terminated, not left running. Log every tool call (arguments redacted) and alert on repeated approval-bypass attempts. ## Anti-patterns | Anti-pattern | Why it bites | Do instead | | ---------------------------------------------- | ------------------------------------------------------------- | --------------------------------------------------- | | Approve every action | ~93% rubber-stamped; the real risky one slips through | Gate by risk class; HITL only on irreversible/external | | One broad session token shared by all tools | Hijacked loop reuses it everywhere | Task-scoped, short-lived per-tool tokens | | Trust RAG / fetched / email content | Indirect injection (LLM01) becomes direct tool execution | Delimit as untrusted data; mediate before any tool | | Wildcard `run_shell(cmd)` tool | Unbounded blast radius | Narrow tools, allowlisted resources, denied patterns | | Raw external text piped into tool args | Attacker controls the action's parameters | Schema-validate args; allowlist recipients/domains | | Scope/limits stated only in the prompt | Prompt is advisory; the model can be talked out of it | Enforce at the tool/harness boundary | | No loop / cost / rate cap | A hijacked or looping agent runs until it runs out of money | Hard fail-closed kill-switches | | Redact PII only in the UI | The secret was already written to memory/logs | Redact before persistence, at the source |
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.