AI AGENT ADDONS

When you run multiple AI agents on the same task, you need a fair way to pick the best one. This tool scores and ranks every agent result. You can use a metric like execution time or accuracy. Or you can use an LLM judge that reads the outputs and decides the winner.

If your top agents are very close, a hybrid mode first checks the numbers then uses the LLM to break ties. After ranking, you can merge the winning agent into your project with one command.

Add Eval skill to your workflow

Global

mkdir -p ~/.claude/skills/eval

Project

mkdir -p .claude/skills/eval

Source Repository

Stars
19,285
Forks
2,656
Watchers
19,285
License
MIT
Last Push
3 months ago
Created
11 months ago