You can now compare different coding agents side by side. Agent Eval is a simple tool that runs your own tasks on agents like Claude Code, Aider, and Codex. It measures pass rate, cost, time, and consistency. No more guessing which agent works best for your project.
Each task is defined in a plain YAML file. The tool creates a fresh copy of your code for each agent run. This keeps results clean and repeatable. You get a table that shows exactly how each agent performs.
Anyone on a team can use this to make data-backed choices. It helps you pick the right agent without relying on gut feelings or online reviews.
Global
mkdir -p ~/.claude/skills/agent-evalProject
mkdir -p .claude/skills/agent-evalSource Repository
Find Skillsvercel-labs/skills
Find and install the perfect skill to extend your AI agent
Microsoft Foundrymicrosoft/azure-skills
Build, deploy, and improve AI agents on Microsoft Foundry from start to finish
Azure Aimicrosoft/azure-skills
Search, transcribe, and analyze with Azure AI tools for smarter apps
Azure Hosted Copilot Sdkmicrosoft/azure-skills
Build, deploy, and manage your Copilot SDK apps on Azure with ease
Skill Creatoranthropics/skills
Create, test, and improve AI agent skills with easy step-by-step guidance
Image Editagentspace-so/runcomfy-agent-skills
Smart router picks the best AI model for your image editing needs
Agentspaceagentspace-so/skills
See your AI agent's live folder from any browser instantly
Openclaw Secure Linux Cloudxixu-me/skills
Secure self-hosting of OpenClaw with private control and SSH tunneling