Claude Skill

Benchmark virtual agents with scripted multi-turn conversations using Agent Evaluation

Run concurrent scripted conversations against a target agent to measure whether it stays on task, responds correctly, and holds up in repeatable test cases.

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download agentskillexchange-skills-skills_benchmark-virtual-agents-with-scripted-multi-turn-conversations-using-agent-evaluation-07beb56.zip · 0 KB
Part of agentskillexchange/skills — 249 skills

Install

skills CLI npx skills add https://github.com/agentskillexchange/skills/tree/main/skills/benchmark-virtual-agents-with-scripted-multi-turn-conversations-using-agent-evaluation
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install agentskillexchange-skills@llmmart
Git git clone https://github.com/agentskillexchange/skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole agentskillexchange/skills collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Benchmark virtual agents with scripted multi-turn conversations using Agent Evaluation

Run concurrent scripted conversations against a target agent to measure whether it stays on task, responds correctly, and holds up in repeatable test cases.

Prerequisites

Python environment, target agent endpoint or integration, optional AWS services such as Bedrock or SageMaker

Installation

No source-backed install or usage instructions could be extracted automatically. Review the upstream project before running this skill in a sensitive workflow.

Documentation

Source

Files (skills)
  • SKILL.md 1.5 KB
    ---
    name: "Benchmark virtual agents with scripted multi-turn conversations using Agent Evaluation"
    slug: "benchmark-virtual-agents-with-scripted-multi-turn-conversations-using-agent-evaluation"
    description: "Run concurrent scripted conversations against a target agent to measure whether it stays on task, responds correctly, and holds up in repeatable test cases."
    github_stars: 358
    verification: "listed"
    source: "https://github.com/awslabs/agent-evaluation"
    author: "AWS Labs"
    publisher_type: "open_source_project"
    category: "Runbooks & Diagnostics"
    framework: "Custom Agents"
    tool_ecosystem:
      github_repo: "awslabs/agent-evaluation"
      github_stars: 358
    ---
    
    # Benchmark virtual agents with scripted multi-turn conversations using Agent Evaluation
    
    Run concurrent scripted conversations against a target agent to measure whether it stays on task, responds correctly, and holds up in repeatable test cases.
    
    ## Prerequisites
    
    Python environment, target agent endpoint or integration, optional AWS services such as Bedrock or SageMaker
    
    ## Installation
    
    No source-backed install or usage instructions could be extracted automatically. Review the upstream project before running this skill in a sensitive workflow.
    
    - Source: https://github.com/awslabs/agent-evaluation
    
    ## Documentation
    
    - https://awslabs.github.io/agent-evaluation/
    
    ## Source
    
    - [Agent Skill Exchange](https://agentskillexchange.com/skills/benchmark-virtual-agents-with-scripted-multi-turn-conversations-using-agent-evaluation/)
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related