Claude Skill

Benchmark OpenClaw coding agents against repeatable real tasks before rollout with PinchBench

Run a real-task benchmark suite against OpenClaw agents so model or harness changes can be compared before they hit production workflows.

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download agentskillexchange-skills-skills_benchmark-openclaw-coding-agents-against-repeatable-real-tasks-before-rollout-with-pinchbench-07beb56.zip · 0 KB
Part of agentskillexchange/skills — 249 skills

Install

skills CLI npx skills add https://github.com/agentskillexchange/skills/tree/main/skills/benchmark-openclaw-coding-agents-against-repeatable-real-tasks-before-rollout-with-pinchbench
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install agentskillexchange-skills@llmmart
Git git clone https://github.com/agentskillexchange/skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole agentskillexchange/skills collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Benchmark OpenClaw coding agents against repeatable real tasks before rollout with PinchBench

Run a real-task benchmark suite against OpenClaw agents so model or harness changes can be compared before they hit production workflows.

Prerequisites

Running OpenClaw instance, Python 3.10+, uv, PinchBench repository checkout, model provider credentials as documented upstream

Installation

Use the upstream install or setup path that matches your environment:

Requirements and caveats from upstream:

  • Note: Model IDs must include their provider prefix (e.g. openrouter/, anthropic/). OpenRouter is the default provider used for routing.
  • Python 3.10+

Basic usage or getting-started notes:

Documentation

Source

Files (skills)
  • SKILL.md 1.9 KB
    ---
    name: "Benchmark OpenClaw coding agents against repeatable real tasks before rollout with PinchBench"
    slug: "benchmark-openclaw-coding-agents-against-repeatable-real-tasks-before-rollout-with-pinchbench"
    description: "Run a real-task benchmark suite against OpenClaw agents so model or harness changes can be compared before they hit production workflows."
    github_stars: 1003
    verification: "security_reviewed"
    source: "https://github.com/pinchbench/skill"
    author: "pinchbench"
    publisher_type: "organization"
    category: "Code Quality & Review"
    framework: "OpenClaw"
    tool_ecosystem:
      github_repo: "pinchbench/skill"
      github_stars: 1003
    ---
    
    # Benchmark OpenClaw coding agents against repeatable real tasks before rollout with PinchBench
    
    Run a real-task benchmark suite against OpenClaw agents so model or harness changes can be compared before they hit production workflows.
    
    ## Prerequisites
    
    Running OpenClaw instance, Python 3.10+, uv, PinchBench repository checkout, model provider credentials as documented upstream
    
    ## Installation
    
    Use the upstream install or setup path that matches your environment:
    - git clone https://github.com/pinchbench/skill.git
    
    Requirements and caveats from upstream:
    - **Note:** Model IDs must include their provider prefix (e.g. openrouter/, anthropic/). [OpenRouter](https://openrouter.ai) is the default provider used for routing.
    - Python 3.10+
    
    Basic usage or getting-started notes:
    - **Tool usage** — Can the model call the right tools with the right parameters?
    - bash
    - # Clone the skill
    
    - Source: https://github.com/pinchbench/skill
    - Extracted from upstream docs: https://raw.githubusercontent.com/pinchbench/skill/HEAD/README.md
    
    ## Documentation
    
    - https://pinchbench.com
    
    ## Source
    
    - [Agent Skill Exchange](https://agentskillexchange.com/skills/benchmark-openclaw-coding-agents-against-repeatable-real-tasks-before-rollout-with-pinchbench/)
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related