{"slug":"agent-loop-testing","title":"agent-loop-testing","summary":"Use this skill when you need evidence-bounded loop state, plan/action/observation cycles, stop conditions, budgets, repetition, and trace evidence; triggers include Agent 循环 and Agent loop.","platform":"Claude","tags":[],"authorName":"LLM Mart","authorSlug":"llm-mart","score":0,"source":"github","price":null,"verified":false,"createdAt":"2026-09-24T14:54:52.569036Z","repo":{"url":"https://github.com/naodeng/awesome-qa-skills","stars":230,"forks":31,"license":null,"updatedAt":"2026-09-22T13:51:34Z"},"bodyHtml":"<hr>\n<h2>name: agent-loop-testing\ndescription: Use this skill when you need evidence-bounded loop state, plan/action/observation cycles, stop conditions, budgets, repetition, and trace evidence; triggers include Agent 循环 and Agent loop.</h2>\n<h1>Agent Loop Testing</h1>\n<h2>When to Use</h2>\n<ul>\n<li>Use this skill when you need evidence-bounded analysis, design, or validation preparation for loop state, plan/action/observation cycles, stop conditions, budgets, repetition, and trace evidence.</li>\n<li>Use it to review an Agent, RAG, or LLM plan, result, or evidence set and produce actionable improvements.</li>\n<li>Use it when context is incomplete but a bounded first pass with assumptions, gaps, and human-decision boundaries is still useful.</li>\n</ul>\n<h2>Output Format Options</h2>\n<ul>\n<li>Default to Markdown organized by domain risk, evidence state, priority, and boundary.</li>\n<li>When the user requests tables, CSV, JSON, or ticket fields, preserve the same finding fields, evidence, and decision boundaries.</li>\n<li>Before machine consumption, confirm the schema, enums, required fields, and evidence sources.</li>\n</ul>\n<h2>How to Use</h2>\n<ol>\n<li>Read and follow <code>prompts/agent-loop-testing.md</code>, including its input audit, domain coverage, and output order.</li>\n<li>Extract scope, environment, version, time window, constraints, success criteria, and available evidence, with attention to loop state, plan/action/observation, stop condition, budget, repetition.</li>\n<li>Separate confirmed facts, evidence-backed inferences, candidate recommendations, and Human decisions before ranking by risk and evidence strength.</li>\n<li>Turn high-risk items into preconditions, steps, expected behavior or decision criteria, required evidence, and a validation method.</li>\n<li>When input is incomplete, deliver a bounded first pass, state unsupported conclusions, and never present static material as execution evidence.</li>\n</ol>\n<h2>Reference Files</h2>\n<ul>\n<li>Always read <code>prompts/agent-loop-testing.md</code>; it is the complete execution specification for this skill.</li>\n<li>For evaluation, read <code>evals/eval.yaml</code> and the matching cases under <code>evals/cases/</code>.</li>\n<li>Load <code>references/</code>, <code>examples/</code>, <code>scripts/</code>, or <code>output-formats.md</code> only when those directories exist and the task needs them.</li>\n</ul>\n<h2>Core Constraints</h2>\n<ul>\n<li>Keep the analysis focused on loop state, plan/action/observation cycles, stop conditions, budgets, repetition, and trace evidence; do not replace business owners or Human risk acceptance, exception approval, or safety decisions.</li>\n<li>Never invent system behavior, fields, model outputs, data, thresholds, root causes, execution records, or pass claims.</li>\n<li>Static design, plans, file presence, or a dry run retain their evidence state and cannot become proof of real execution.</li>\n<li>When evidence is insufficient, use pending confirmation, blocked, unassessed, or NOT_SCORED and give the smallest validation method.</li>\n<li>For user data, production, or safety work, use least privilege, masked data, mocks, dry runs, or isolation.</li>\n</ul>\n<h2>Delivery Checklist</h2>\n<ul>\n<li><input disabled=\"disabled\" type=\"checkbox\"> Covered loop state, plan/action/observation, stop condition, budget, repetition, with source, evidence state, and validation method for each.</li>\n<li><input disabled=\"disabled\" type=\"checkbox\"> Separated facts, inferences, candidate recommendations, gaps, and Human decisions.</li>\n<li><input disabled=\"disabled\" type=\"checkbox\"> Gave high-risk items P0/P1/P2/P3 or an equivalent priority, owner role, and close condition.</li>\n<li><input disabled=\"disabled\" type=\"checkbox\"> Did not turn plans, static checks, or dry runs into test execution, all-passed, or safety-approved claims.</li>\n<li><input disabled=\"disabled\" type=\"checkbox\"> Stated residual risk, stop/escalation conditions, and next actions.</li>\n</ul>\n<h2>Common Pitfalls</h2>\n<ul>\n<li>Listing checks without triggers, expected concerns, owner roles, close conditions, and evidence.</li>\n<li>Treating adjacent tests or model tools as a complete Agent loop judgment.</li>\n<li>Using unexplained numbers for false precision or writing correlation as causation.</li>\n<li>Refusing incomplete input, or pretending that incomplete evidence is conclusive.</li>\n</ul>\n<h2>Best Practices</h2>\n<ul>\n<li>Start with paths most likely to cause user harm, business loss, or decision blockage.</li>\n<li>Use the smallest verifiable experiment to reduce uncertainty and record conditions, versions, sources, and evidence.</li>\n<li>Make the Skill independently installable, executable, and reviewable by another engineer.</li>\n</ul>\n","files":[{"path":"agents/openai.yaml","sizeBytes":290,"isText":true},{"path":"evals/cases/basic-success.yaml","sizeBytes":1411,"isText":true},{"path":"evals/cases/edge-incomplete-input.yaml","sizeBytes":917,"isText":true},{"path":"evals/cases/edge-scope-boundary.yaml","sizeBytes":953,"isText":true},{"path":"evals/eval.yaml","sizeBytes":441,"isText":true},{"path":"evals/local-rules.json","sizeBytes":137,"isText":true},{"path":"evals/trigger-prompts.csv","sizeBytes":582,"isText":false},{"path":"prompts/agent-loop-testing.md","sizeBytes":4680,"isText":true},{"path":"SKILL.md","sizeBytes":4134,"isText":true}],"reviewScore":null,"reviewSummary":null,"trust":{"provenance":"trusted-source-unreviewed","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow.","bodySource":null},"bodyLocked":false,"purchaseUrl":null,"sourceUrl":null,"report":{"provenance":"trusted-source-unreviewed","screen":{"ran":true,"outcome":"clean","suspicious":0,"notes":0,"hiddenCharacters":false},"virusScan":{"engine":"clamav","status":"clean","scannedAt":"2026-09-24T14:56:39.508241Z","sha256":"934B0DDBC716183277CE9C01503E0D36A1D2CBC1569DF6F19F3CA9ADFBBAEB41","sizeBytes":7035},"review":null,"source":{"repositoryUrl":"https://github.com/naodeng/awesome-qa-skills","path":"skills/en/testing-types/agent-loop-testing","license":null,"commit":"c44b8922085e01bafc804d1ffa3f21d4cec1d1c5","subtreeSha":"376B10C7307D337D485338D481E1C28655BFABC9F0DD25637886BA1BC7402CB7","lastSyncedAt":"2026-09-24T14:54:50.849933Z"},"reviewedAt":"2026-09-24T15:03:38.386844Z","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow."},"install":[{"target":"skills-cli","command":"npx skills add https://github.com/naodeng/awesome-qa-skills/tree/main/skills/en/testing-types/agent-loop-testing"},{"target":"claude-code","command":"claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install naodeng-awesome-qa-skills@llmmart"},{"target":"git","command":"git clone https://github.com/naodeng/awesome-qa-skills.git"}]}