Claude Skill

Benchmark CLI agents on autonomous LLM post-training with PostTrainBench

Run Claude Code, Codex CLI, Gemini CLI, or OpenCode through bounded H100 post-training tasks and compare how well each agent improves a base LLM.

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download agentskillexchange-skills-skills_benchmark-cli-agents-on-autonomous-llm-post-training-with-posttrainbench-07beb56.zip · 0 KB
Part of agentskillexchange/skills — 249 skills

Install

skills CLI npx skills add https://github.com/agentskillexchange/skills/tree/main/skills/benchmark-cli-agents-on-autonomous-llm-post-training-with-posttrainbench
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install agentskillexchange-skills@llmmart
Git git clone https://github.com/agentskillexchange/skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole agentskillexchange/skills collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Benchmark CLI agents on autonomous LLM post-training with PostTrainBench

Run Claude Code, Codex CLI, Gemini CLI, or OpenCode through bounded H100 post-training tasks and compare how well each agent improves a base LLM.

Prerequisites

Python, apptainer, fuse-overlayfs, Hugging Face cache, H100 GPU access, currently HTCondor scheduler support, and credentials for the selected CLI agent scaffolds

Installation

Install or set up from the source-backed instructions:

Clone https://github.com/aisa-group/PostTrainBench, install requirements including apptainer and fuse-overlayfs, build the standard container with bash containers/build_container.sh standard, download the Hugging Face cache with bash containers/download_hf_cache/download_hf_cache.sh, copy example.env to .env, set API keys and paths, then submit jobs with bash src/commit_utils/commit.sh.

Documentation

Source

Files (skills)
  • SKILL.md 1.7 KB
    ---
    name: "Benchmark CLI agents on autonomous LLM post-training with PostTrainBench"
    slug: "benchmark-cli-agents-on-autonomous-llm-post-training-with-posttrainbench"
    description: "Run Claude Code, Codex CLI, Gemini CLI, or OpenCode through bounded H100 post-training tasks and compare how well each agent improves a base LLM."
    github_stars: 543
    verification: "security_reviewed"
    source: "https://github.com/aisa-group/PostTrainBench"
    author: "AISA Group"
    publisher_type: "open_source_project"
    category: "Developer Tools"
    framework: "Multi-Framework"
    tool_ecosystem:
      github_repo: "aisa-group/PostTrainBench"
      github_stars: 543
    ---
    
    # Benchmark CLI agents on autonomous LLM post-training with PostTrainBench
    
    Run Claude Code, Codex CLI, Gemini CLI, or OpenCode through bounded H100 post-training tasks and compare how well each agent improves a base LLM.
    
    ## Prerequisites
    
    Python, apptainer, fuse-overlayfs, Hugging Face cache, H100 GPU access, currently HTCondor scheduler support, and credentials for the selected CLI agent scaffolds
    
    ## Installation
    
    Install or set up from the source-backed instructions:
    
    Clone https://github.com/aisa-group/PostTrainBench, install requirements including apptainer and fuse-overlayfs, build the standard container with bash containers/build_container.sh standard, download the Hugging Face cache with bash containers/download_hf_cache/download_hf_cache.sh, copy example.env to .env, set API keys and paths, then submit jobs with bash src/commit_utils/commit.sh.
    
    - Source: https://github.com/aisa-group/PostTrainBench
    
    ## Documentation
    
    - http://posttrainbench.com/
    
    ## Source
    
    - [Agent Skill Exchange](https://agentskillexchange.com/skills/benchmark-cli-agents-on-autonomous-llm-post-training-with-posttrainbench/)
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related