Claude Skill

agent-safety

Use when bounding an LLM agent that already runs — scoping its task domain, gating tools to least privilege, defending against prompt injection in untrusted web/email/RAG text, requiring human approval on irreversible actions, capping runtime and cost, or triaging what it already

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download ericrisco-rsc-harness-skills_agent-safety-953fef5.zip · 8 KB
Part of ericrisco/rsc-harness — 46 skills

Install

skills CLI npx skills add https://github.com/ericrisco/rsc-harness/tree/main/skills/agent-safety
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install ericrisco-rsc-harness@llmmart
Git git clone https://github.com/ericrisco/rsc-harness.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole ericrisco/rsc-harness collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Agent safety

You are the security review for an agent's agency, not for its code. The loop works, tools are wired, memory persists — your job is to make that autonomy bounded. If you want to review ordinary endpoints, auth, or secrets handling, that is ../secure-coding/SKILL.md — this skill is the Agentic Top 10, the risks that exist only because a model has tools and autonomy. If the loop or tools do not exist yet, that is ../building-agents/SKILL.md. You arrive after both.

references/threat-model.md carries the OWASP Agentic Top 10 2026 risks mapped to the controls below, the pre-ship guardrail checklist, and the incident-triage flow for "the agent did X" — open it when you are reviewing before ship or reconstructing an incident.

The ownership split

Agent security splits into four layers — Model · Harness · Tools · Environment. The model provider owns only the Model layer (alignment, refusals). Everything else is yours: the Harness (loop, memory, context assembly), the Tools (what the agent can do), and the Environment (creds, network, blast radius). Do not outsource a layer you own to "the model is aligned."

Three excesses cause almost every agentic incident. Cut all three:

  • Excessive functionality — tools the task never needs.
  • Excessive permissions — broader scopes/creds than the tool needs.
  • Excessive autonomy — acting without checking back when it should.

The operating principle is least agency: autonomy is earned per task, not defaulted.

Scope limits

  • Declare the allowed task domain as a hard boundary in the system prompt. Why: an undeclared scope is an infinite scope; "you are a refund assistant; you do not touch payroll" is a constraint a reviewer can check.
  • Deny by default — the agent starts with zero tools. Each tool earns its place by a task justification. Why: an opt-out tool list grows; an opt-in list stays minimal.
  • Segregate the instruction channel from the data channel. System/developer prompt = trusted instructions. Everything the agent reads at runtime = data, never instructions. Why: this single boundary is what stops indirect injection (LLM01).

Tool gating / least agency

Give every tool a profile: read / write / exec / send, the exact resources it may touch, and an allowlist (never a wildcard). Block destructive flags and secret paths at the tool boundary, not in the prompt — the prompt is advisory, the boundary is enforced.

# Bad: one wildcard tool = unbounded blast radius, runs anything the loop emits
def run_shell(cmd: str) -> str:
    return subprocess.run(cmd, shell=True, capture_output=True, text=True).stdout
# Good: narrow tool, allowlisted root, denied patterns, no shell
ALLOWED_ROOT = pathlib.Path("/srv/agent/workspace").resolve()
DENY = ("*.key", "*.pem", "*secret*", "*.env", "id_rsa*")

def read_file(path: str) -> str:
    p = (ALLOWED_ROOT / path).resolve()
    if not p.is_relative_to(ALLOWED_ROOT):           # no traversal out of scope
        raise PermissionError("path outside workspace")
    if any(p.match(g) for g in DENY):                # never read secrets
        raise PermissionError("denied pattern")
    return p.read_text()
  • Issue task-scoped, short-lived tokens — not the session's broad creds. A credential should be valid only for the specific tool and the duration of one task. Why: a hijacked loop cannot reuse a session-wide token it never held.
  • Prefer read-only by default; writes/sends/exec are separate, gated tools. Why: most steps only need to read, so most steps should be unable to mutate anything.

Injection defense

Treat all external data as untrusted: user messages, retrieved documents, API responses, emails, web pages, other agents' output. Sanitize and delimit before it enters context, and never let external text reach a privileged tool unmediated.

Source Trust level Required mediation before it can act
System / developer prompt Trusted none (this is the only instruction channel)
End-user chat message Untrusted delimit; treat as data, not commands
Retrieved RAG / KB document Untrusted delimit; strip instruction-like spans
Fetched web page / API JSON Untrusted parse to schema; no raw text → tool args
Inbound email / ticket body Untrusted delimit; HITL on any action it requests
Another agent's message Untrusted same as external user input
# Bad: retrieved chunk flows straight into a privileged action
chunk = retriever.search(q)[0].text          # attacker-controlled doc
agent.call_tool("send_email", to=extract_to(chunk), body=chunk)
# Good: external content is quarantined data; the action is schema-validated + gated
chunk = retriever.search(q)[0].text
ctx = f"<retrieved untrusted>\n{chunk}\n</retrieved untrusted>"   # delimited, labeled
proposal = agent.draft("send_email", context=ctx)                # model proposes
args = SendEmail.model_validate(proposal.args)                   # schema or reject
if args.to_domain not in ALLOWED_DOMAINS:                        # exfil guard
    raise PermissionError("recipient outside allowlist")
require_human_approval("send_email", args)                       # irreversible → HITL
  • Validate every tool-call argument against a strict schema before execution. Why: a schema rejects the surprise field, encoded payload, or off-allowlist recipient injection produces.
  • Watch for exfiltration shapes — unexpected outbound URLs, base64 blobs, recipients outside the allowlist. Why: data theft is the common payload of a successful injection.

Human-in-the-loop by risk class

Do not approve every action — reported ~93% of permission prompts get approved without being read, so blanket prompting trains a rubber stamp. Gate by risk class, keyed on reversibility × blast radius. Bind each approval to the exact parameters with a short-lived token so the approved action cannot be swapped after the click.

Action type (examples) Reversible? Blast radius Control
Read file, search KB, fetch page n/a none auto
Write to scratch workspace, internal draft yes local log-only
Mutate prod DB, deploy, change config hard system approve (HITL)
Send email/payment to external party, post live no external approve (HITL)
Delete backups, rotate prod creds, mass-delete no catastrophic block (or step-up auth)
  • Step up for the top row — high-value irreversible actions deserve fresh auth, not the ambient session. Why: a hijacked session should not also hold the keys to the worst action.

Memory hygiene

  • Validate and sanitize content before it is stored. Why: memory poisoning persists across sessions (OWASP Agentic T1) — unlike session-scoped injection, a poisoned memory re-attacks every future run until purged.
  • Isolate memory per user and per session; do not let one user's writes color another's reads. Why: shared memory is a cross-tenant injection channel.
  • Expire entries and cap memory size; redact PII (SSN, cards, API keys) before persist. Why: stale instructions and leaked secrets both age into liabilities.

Runtime kill-switches

A looping or hijacked agent must hit a wall on its own. Set hard caps, fail closed:

  • Tool-call rate cap (e.g. ~30 calls/min) — runaway loops trip it before they do damage.
  • Cost cap per session (e.g. ~$10) — a wallet attack stops at a known ceiling.
  • Loop / step cap — a fixed max iterations kills the infinite plan.
  • Wall-clock timeout — a stuck agent is terminated, not left running.

Log every tool call (arguments redacted) and alert on repeated approval-bypass attempts.

Anti-patterns

Anti-pattern Why it bites Do instead
Approve every action ~93% rubber-stamped; the real risky one slips through Gate by risk class; HITL only on irreversible/external
One broad session token shared by all tools Hijacked loop reuses it everywhere Task-scoped, short-lived per-tool tokens
Trust RAG / fetched / email content Indirect injection (LLM01) becomes direct tool execution Delimit as untrusted data; mediate before any tool
Wildcard run_shell(cmd) tool Unbounded blast radius Narrow tools, allowlisted resources, denied patterns
Raw external text piped into tool args Attacker controls the action's parameters Schema-validate args; allowlist recipients/domains
Scope/limits stated only in the prompt Prompt is advisory; the model can be talked out of it Enforce at the tool/harness boundary
No loop / cost / rate cap A hijacked or looping agent runs until it runs out of money Hard fail-closed kill-switches
Redact PII only in the UI The secret was already written to memory/logs Redact before persistence, at the source
Files (rsc-harness)
  • evals
    • cases.yaml 3.2 KB
      skill: agent-safety
      
      should_trigger:
        - prompt: "Our support agent can call any shell command — lock it down."
          why: Wildcard exec tool with unbounded blast radius; core tool-gating / least-agency case.
        - prompt: "An email in the inbox told the agent to forward all invoices to an external address and it did."
          why: Indirect prompt injection (LLM01). Non-obvious — reads like a bug report, but the cause is untrusted email content reaching a privileged send tool.
        - prompt: "Before we ship this autonomous agent to prod, review what it's allowed to do."
          why: Pre-ship review of scope, permissions, and HITL — exactly the bounded-autonomy review this skill is.
        - prompt: "Posa límits a l'agent perquè no pugui esborrar res sense aprovació."
          why: Catalan. HITL-by-risk-class on an irreversible delete plus scope limiting.
        - prompt: "The agent's memory keeps repeating a wrong instruction across sessions."
          why: Memory poisoning (OWASP Agentic T1). Non-obvious — sounds like a quality bug, is actually contaminated persistent memory.
        - prompt: "Review our MCP server's tool registry — the scopes look way too broad."
          why: Over-broad tool scopes / excessive permissions; deny-by-default and per-tool allowlists.
      
      should_not_trigger:
        - prompt: "Help me build the agent's tool-calling loop."
          route_to: building-agents
          why: Construction of the agent (loop/tools), not constraining one that already runs.
        - prompt: "Is this Express auth endpoint vulnerable to broken access control?"
          route_to: secure-coding
          why: Plain web/app OWASP on an endpoint; no model, tools, or autonomy involved.
        - prompt: "Measure whether my agent answers correctly with a golden set."
          route_to: agent-eval
          why: Quality / correctness scoring, not safety constraint.
        - prompt: "Write a sharper system prompt so the agent is more helpful."
          route_to: prompt-engineering
          why: Capability tuning, not a safety boundary.
        - prompt: "Show me how much each model call costs this month."
          route_to: cost-tracking
          why: Reporting spend; agent-safety sets the cost *cap*, it does not report the analytics.
      
      capability:
        - scenario: >
            A research agent has web-fetch, send-email, and file-write tools plus persistent
            memory shared across sessions. Design the guardrails that make its autonomy bounded.
          must_include:
            - Deny-by-default tools with a per-tool profile (read web vs send email vs write file); no wildcard tool.
            - Treats web-fetched content as untrusted; delimits it as data; never pipes external text straight into send-email args.
            - HITL approval for send-email bound to exact parameters (irreversible, external blast radius); recipient/domain allowlist as exfil guard.
            - File-write scoped to an allowlisted workspace root; blocks secret patterns (*.key/*.env/*secret*) and path traversal.
            - Memory validated before store, isolated per session, expiring, PII-redacted before persist.
            - Runtime kill-switches: tool-call rate cap, per-session cost cap, loop/step cap, wall-clock timeout, all failing closed.
            - Task-scoped short-lived tokens rather than one broad session credential.
            - Maps each control to OWASP Agentic Top 10 (indirect injection / tool misuse / memory poisoning / excessive agency).
      
    • README.md 909 B
      # Evals: agent-safety
      
      These cases are the trigger contract and capability rubric for the `agent-safety` skill.
      `should_trigger` and `should_not_trigger` assert that the description routes correctly —
      each negative names the real sibling it should route to instead (building-agents,
      secure-coding, agent-eval, prompt-engineering, cost-tracking). Run them through the repo's
      eval harness if one is wired up, or read them directly as the routing spec when reviewing
      the description. The `capability` block is not a pass/fail script: it is a rubric a human or
      an LLM-as-judge scores a sample guardrail design against — the design must cover every
      `must_include` item (deny-by-default tools, untrusted-content mediation, HITL on
      irreversible actions, memory hygiene, runtime caps) and map them to the OWASP Agentic Top
      10. This is a process/review skill, so there is no `verify.sh`; rigor lives in this eval.
      
  • references
    • threat-model.md 4 KB
      # Threat model: OWASP Agentic Top 10 (2026) → controls
      
      Mapping of the named agentic risks to the control in `../SKILL.md` that mitigates each.
      Sources: OWASP Gen AI Security Project, "OWASP Top 10 for Agentic Applications for 2026"
      (genai.owasp.org); OWASP Top 10 for LLM Applications v2025; OWASP AI Agent Security Cheat
      Sheet (cheatsheetseries.owasp.org). All accessed 2026-06-02.
      
      | Agentic risk                       | What it is                                                              | Control in this skill                                   |
      | ---------------------------------- | ----------------------------------------------------------------------- | ------------------------------------------------------- |
      | Memory poisoning (T1)              | Contamination persists across sessions, re-attacking every future run   | Memory hygiene: validate before store, isolate, expire  |
      | Tool misuse & exploitation         | Abuse via unsafe composition / recursion / excessive calls *with* valid perms | Tool gating + runtime kill-switches (rate/loop caps)    |
      | Privilege compromise               | Broad or stolen creds let the loop act beyond its task                  | Task-scoped short-lived tokens; least agency            |
      | Indirect prompt injection (LLM01)  | Untrusted content carries instructions the model executes               | Trust-boundary table; segregate data from instructions  |
      | Excessive agency                   | Acting without check-back on irreversible/external actions              | HITL by risk class; deny-by-default scope               |
      | Data exfiltration                  | Stolen data leaves via tool output / outbound calls                     | Output schema validation; exfil-shape detection         |
      
      The OWASP framing of root causes — excessive **functionality**, **permissions**,
      **autonomy** — maps one-to-one onto the three excesses in the SKILL body. Cutting all three
      is the deny-by-default + least-agency posture.
      
      ## Pre-ship guardrail checklist
      
      Before an autonomous loop reaches production, confirm:
      
      - [ ] Allowed task domain is declared in the system prompt as a hard boundary.
      - [ ] Tools are deny-by-default; each enabled tool has a written task justification.
      - [ ] Every tool has a profile (read / write / exec / send) and an allowlist, not a wildcard.
      - [ ] No tool can read `*.key` / `*.pem` / `*secret*` / `*.env` or run with destructive flags.
      - [ ] Creds are task-scoped and short-lived, not one broad session token.
      - [ ] All external content (web, email, RAG, API, other agents) is delimited as untrusted.
      - [ ] Tool-call args are schema-validated; recipients/domains are allowlisted (exfil guard).
      - [ ] HITL is gated by risk class; irreversible/external actions require approval bound to exact params.
      - [ ] Memory is validated before store, isolated per user/session, expiring, PII-redacted.
      - [ ] Runtime caps exist and fail closed: rate, cost, loop/step, wall-clock timeout.
      - [ ] Every tool call is logged (redacted); repeated approval-bypass alerts.
      
      ## Incident triage: "the agent did X"
      
      When an agent has already done something it should not have, work the loop in order:
      
      1. **Contain.** Revoke the task tokens / creds the loop is holding; trip the kill-switch
         (pause the loop). Stop the bleeding before diagnosing.
      2. **Find the trust boundary it crossed.** Which untrusted source reached a privileged
         tool? Trace the action's parameters back to their origin — usually an email body, RAG
         chunk, fetched page, or a poisoned memory entry.
      3. **Purge poisoned memory.** If the bad behavior repeats across sessions, the cause is
         stored, not in-context. Delete the contaminated entries; do not just restart the session.
      4. **Add the missing gate.** Reclassify that action's risk, move it behind HITL or block,
         tighten the tool allowlist, or add the schema/exfil check that would have caught it.
      5. **Confirm with a replay.** Re-run the triggering input against the new guardrail and
         verify the action is now refused or escalated.
      
  • SKILL.md 10.3 KB
    ---
    name: agent-safety
    description: "Use when bounding an LLM agent that already runs — scoping its task domain, gating tools to least privilege, defending against prompt injection in untrusted web/email/RAG text, requiring human approval on irreversible actions, capping runtime and cost, or triaging what it already did. NOT building the loop, tools, or RAG (that is `building-agents`)."
    tags: [agent-security, guardrails, prompt-injection, least-privilege, owasp-agentic]
    recommends: [building-agents, secure-coding, agent-eval]
    origin: risco
    ---
    
    # Agent safety
    
    You are the security review for an agent's **agency**, not for its code. The loop works,
    tools are wired, memory persists — your job is to make that autonomy *bounded*. If you
    want to review ordinary endpoints, auth, or secrets handling, that is
    `../secure-coding/SKILL.md` — this skill is the *Agentic* Top 10, the risks that exist only
    because a model has tools and autonomy. If the loop or tools do not exist yet, that is
    `../building-agents/SKILL.md`. You arrive *after* both.
    
    `references/threat-model.md` carries the OWASP Agentic Top 10 2026 risks mapped to the
    controls below, the pre-ship guardrail checklist, and the incident-triage flow for "the
    agent did X" — open it when you are reviewing before ship or reconstructing an incident.
    
    ## The ownership split
    
    Agent security splits into four layers — **Model · Harness · Tools · Environment**. The
    model provider owns only the Model layer (alignment, refusals). Everything else is yours:
    the Harness (loop, memory, context assembly), the Tools (what the agent can *do*), and the
    Environment (creds, network, blast radius). Do not outsource a layer you own to "the model
    is aligned."
    
    Three excesses cause almost every agentic incident. Cut all three:
    
    - **Excessive functionality** — tools the task never needs.
    - **Excessive permissions** — broader scopes/creds than the tool needs.
    - **Excessive autonomy** — acting without checking back when it should.
    
    The operating principle is **least agency**: autonomy is earned per task, not defaulted.
    
    ## Scope limits
    
    - **Declare the allowed task domain as a hard boundary in the system prompt.** Why: an
      undeclared scope is an infinite scope; "you are a refund assistant; you do not touch
      payroll" is a constraint a reviewer can check.
    - **Deny by default — the agent starts with zero tools.** Each tool earns its place by a
      task justification. Why: an opt-out tool list grows; an opt-in list stays minimal.
    - **Segregate the instruction channel from the data channel.** System/developer prompt =
      trusted instructions. Everything the agent reads at runtime = data, never instructions.
      Why: this single boundary is what stops indirect injection (LLM01).
    
    ## Tool gating / least agency
    
    Give every tool a profile: **read / write / exec / send**, the exact resources it may
    touch, and an **allowlist** (never a wildcard). Block destructive flags and secret paths
    at the tool boundary, not in the prompt — the prompt is advisory, the boundary is enforced.
    
    ```python
    # Bad: one wildcard tool = unbounded blast radius, runs anything the loop emits
    def run_shell(cmd: str) -> str:
        return subprocess.run(cmd, shell=True, capture_output=True, text=True).stdout
    ```
    
    ```python
    # Good: narrow tool, allowlisted root, denied patterns, no shell
    ALLOWED_ROOT = pathlib.Path("/srv/agent/workspace").resolve()
    DENY = ("*.key", "*.pem", "*secret*", "*.env", "id_rsa*")
    
    def read_file(path: str) -> str:
        p = (ALLOWED_ROOT / path).resolve()
        if not p.is_relative_to(ALLOWED_ROOT):           # no traversal out of scope
            raise PermissionError("path outside workspace")
        if any(p.match(g) for g in DENY):                # never read secrets
            raise PermissionError("denied pattern")
        return p.read_text()
    ```
    
    - **Issue task-scoped, short-lived tokens — not the session's broad creds.** A credential
      should be valid only for the specific tool and the duration of one task. Why: a hijacked
      loop cannot reuse a session-wide token it never held.
    - **Prefer read-only by default; writes/sends/exec are separate, gated tools.** Why: most
      steps only need to read, so most steps should be unable to mutate anything.
    
    ## Injection defense
    
    Treat **all** external data as untrusted: user messages, retrieved documents, API
    responses, emails, web pages, other agents' output. Sanitize and delimit before it enters
    context, and **never let external text reach a privileged tool unmediated.**
    
    | Source                         | Trust level | Required mediation before it can act        |
    | ------------------------------ | ----------- | ------------------------------------------- |
    | System / developer prompt      | Trusted     | none (this is the only instruction channel) |
    | End-user chat message          | Untrusted   | delimit; treat as data, not commands        |
    | Retrieved RAG / KB document    | Untrusted   | delimit; strip instruction-like spans       |
    | Fetched web page / API JSON    | Untrusted   | parse to schema; no raw text → tool args    |
    | Inbound email / ticket body    | Untrusted   | delimit; HITL on any action it requests     |
    | Another agent's message        | Untrusted   | same as external user input                 |
    
    ```python
    # Bad: retrieved chunk flows straight into a privileged action
    chunk = retriever.search(q)[0].text          # attacker-controlled doc
    agent.call_tool("send_email", to=extract_to(chunk), body=chunk)
    ```
    
    ```python
    # Good: external content is quarantined data; the action is schema-validated + gated
    chunk = retriever.search(q)[0].text
    ctx = f"<retrieved untrusted>\n{chunk}\n</retrieved untrusted>"   # delimited, labeled
    proposal = agent.draft("send_email", context=ctx)                # model proposes
    args = SendEmail.model_validate(proposal.args)                   # schema or reject
    if args.to_domain not in ALLOWED_DOMAINS:                        # exfil guard
        raise PermissionError("recipient outside allowlist")
    require_human_approval("send_email", args)                       # irreversible → HITL
    ```
    
    - **Validate every tool-call argument against a strict schema before execution.** Why: a
      schema rejects the surprise field, encoded payload, or off-allowlist recipient injection
      produces.
    - **Watch for exfiltration shapes** — unexpected outbound URLs, base64 blobs, recipients
      outside the allowlist. Why: data theft is the common payload of a successful injection.
    
    ## Human-in-the-loop by risk class
    
    Do **not** approve every action — reported ~93% of permission prompts get approved without
    being read, so blanket prompting trains a rubber stamp. Gate by **risk class**, keyed on
    reversibility × blast radius. Bind each approval to the **exact parameters** with a
    short-lived token so the approved action cannot be swapped after the click.
    
    | Action type (examples)                          | Reversible? | Blast radius | Control            |
    | ----------------------------------------------- | ----------- | ------------ | ------------------ |
    | Read file, search KB, fetch page                | n/a         | none         | **auto**           |
    | Write to scratch workspace, internal draft      | yes         | local        | **log-only**       |
    | Mutate prod DB, deploy, change config            | hard        | system       | **approve (HITL)** |
    | Send email/payment to external party, post live | no          | external     | **approve (HITL)** |
    | Delete backups, rotate prod creds, mass-delete  | no          | catastrophic | **block** (or step-up auth) |
    
    - **Step up for the top row** — high-value irreversible actions deserve fresh auth, not the
      ambient session. Why: a hijacked session should not also hold the keys to the worst action.
    
    ## Memory hygiene
    
    - **Validate and sanitize content before it is stored.** Why: memory poisoning persists
      across sessions (OWASP Agentic T1) — unlike session-scoped injection, a poisoned memory
      re-attacks every future run until purged.
    - **Isolate memory per user and per session; do not let one user's writes color another's
      reads.** Why: shared memory is a cross-tenant injection channel.
    - **Expire entries and cap memory size; redact PII (SSN, cards, API keys) before persist.**
      Why: stale instructions and leaked secrets both age into liabilities.
    
    ## Runtime kill-switches
    
    A looping or hijacked agent must hit a wall on its own. Set hard caps, fail closed:
    
    - **Tool-call rate cap** (e.g. ~30 calls/min) — runaway loops trip it before they do damage.
    - **Cost cap per session** (e.g. ~$10) — a wallet attack stops at a known ceiling.
    - **Loop / step cap** — a fixed max iterations kills the infinite plan.
    - **Wall-clock timeout** — a stuck agent is terminated, not left running.
    
    Log every tool call (arguments redacted) and alert on repeated approval-bypass attempts.
    
    ## Anti-patterns
    
    | Anti-pattern                                   | Why it bites                                                   | Do instead                                          |
    | ---------------------------------------------- | ------------------------------------------------------------- | --------------------------------------------------- |
    | Approve every action                           | ~93% rubber-stamped; the real risky one slips through         | Gate by risk class; HITL only on irreversible/external |
    | One broad session token shared by all tools    | Hijacked loop reuses it everywhere                            | Task-scoped, short-lived per-tool tokens            |
    | Trust RAG / fetched / email content            | Indirect injection (LLM01) becomes direct tool execution      | Delimit as untrusted data; mediate before any tool  |
    | Wildcard `run_shell(cmd)` tool                 | Unbounded blast radius                                        | Narrow tools, allowlisted resources, denied patterns |
    | Raw external text piped into tool args         | Attacker controls the action's parameters                     | Schema-validate args; allowlist recipients/domains  |
    | Scope/limits stated only in the prompt         | Prompt is advisory; the model can be talked out of it         | Enforce at the tool/harness boundary                |
    | No loop / cost / rate cap                       | A hijacked or looping agent runs until it runs out of money   | Hard fail-closed kill-switches                      |
    | Redact PII only in the UI                       | The secret was already written to memory/logs                 | Redact before persistence, at the source            |
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related