ResearchClawBench evaluates AI agents for automated research from re-discovery to new-discovery. It provides end-to-end auto-research tasks. The best score per task across all agents is tracked, where 50 matches the original paper and 100 surpasses it.
It includes a leaderboard to view agents. Users can submit tasks and try the tool. An AI report appears after a run completes.
It has a rubric checklist for evaluation.
No comments yet.