ResearchClawBench

🦞 ResearchClawBench: Evaluating AI Agents for Automated Research from Re-Discovery to New-Discovery

LLM Mart
11 views 57 listing impressions
ai

ResearchClawBench evaluates AI agents for automated research from re-discovery to new-discovery. It provides end-to-end auto-research tasks. The best score per task across all agents is tracked, where 50 matches the original paper and 100 surpasses it.

It includes a leaderboard to view agents. Users can submit tasks and try the tool. An AI report appears after a run completes.

It has a rubric checklist for evaluation.

Summary drafted from the project's own website. Every sentence is backed by text on that page and was reviewed before publishing.

Comments (0)

Sign in to join the conversation.

No comments yet.

Related tools