Eval-driven development helps you measure how well an AI agent does its job. You write down what success looks like before the agent starts. Then you check if it passes or fails each test.
This framework tracks results with pass@k metrics. For example pass@3 means at least one success in three tries. You can catch regressions when changes break something that used to work.
Use it to improve agent reliability over time. The same idea works for testing new features or checking older ones still work.
Global
mkdir -p ~/.claude/skills/eval-harnessProject
mkdir -p .claude/skills/eval-harnessSource Repository
Grill Memattpocock/skills
Stress-test your plan with relentless questions until we both understand
Tddmattpocock/skills
Write one test at a time then code to make it pass
Test Driven Developmentobra/superpowers
Write a failing test first then code just enough to pass
Webapp Testinganthropics/skills
Test your local web apps quickly with Playwright automation and screenshots
Qamattpocock/skills
Turn bug reports into GitHub issues through natural conversation without technical fuss
Migrate To Shoehornmattpocock/skills
Replace unsafe as assertions with type-safe partial test data easily
Playwright Best Practicescurrents-dev/playwright-best-practices-skill
Master Playwright testing with best practices for reliable and fast tests
Google Agents Cli Evalgoogle/agents-cli
Run evaluations on your AI agent, find failures, and improve its quality step by step