Claude Skill

Benchmark deep research agents across factual, quality, and process dimensions with MiroEval

Score deep research agents on benchmark tasks using factual verification, report-quality scoring, and process evaluation before model or workflow changes ship.

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download agentskillexchange-skills-skills_benchmark-deep-research-agents-across-factual-quality-and-process-dimensions-with-miroeval-07beb56.zip · 0 KB
Part of agentskillexchange/skills — 249 skills

Install

skills CLI npx skills add https://github.com/agentskillexchange/skills/tree/main/skills/benchmark-deep-research-agents-across-factual-quality-and-process-dimensions-with-miroeval
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install agentskillexchange-skills@llmmart
Git git clone https://github.com/agentskillexchange/skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole agentskillexchange/skills collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Benchmark deep research agents across factual, quality, and process dimensions with MiroEval

Score deep research agents on benchmark tasks using factual verification, report-quality scoring, and process evaluation before model or workflow changes ship.

Prerequisites

Python, uv, model result JSON, required API keys for judge and retrieval services

Installation

No source-backed install or usage instructions could be extracted automatically. Review the upstream project before running this skill in a sensitive workflow.

Documentation

Source

Files (skills)
  • SKILL.md 1.4 KB
    ---
    name: "Benchmark deep research agents across factual, quality, and process dimensions with MiroEval"
    slug: "benchmark-deep-research-agents-across-factual-quality-and-process-dimensions-with-miroeval"
    description: "Score deep research agents on benchmark tasks using factual verification, report-quality scoring, and process evaluation before model or workflow changes ship."
    github_stars: 34
    verification: "listed"
    source: "https://github.com/MiroMindAI/MiroEval"
    author: "MiroMindAI"
    publisher_type: "organization"
    category: "Code Quality & Review"
    framework: "Multi-Framework"
    tool_ecosystem:
      github_repo: "MiroMindAI/MiroEval"
      github_stars: 34
    ---
    
    # Benchmark deep research agents across factual, quality, and process dimensions with MiroEval
    
    Score deep research agents on benchmark tasks using factual verification, report-quality scoring, and process evaluation before model or workflow changes ship.
    
    ## Prerequisites
    
    Python, uv, model result JSON, required API keys for judge and retrieval services
    
    ## Installation
    
    No source-backed install or usage instructions could be extracted automatically. Review the upstream project before running this skill in a sensitive workflow.
    
    - Source: https://github.com/MiroMindAI/MiroEval
    
    ## Documentation
    
    - https://github.com/MiroMindAI/MiroEval
    
    ## Source
    
    - [Agent Skill Exchange](https://agentskillexchange.com/skills/benchmark-deep-research-agents-across-factual-quality-and-process-dimensions-with-miroeval/)
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related