The ADK Evaluation Guide helps you test how well your AI agent works. It explains how to run evaluations and read the scores. You can check things like whether the agent used the right tools and gave correct answers.
This guide also shows you how to fix problems when scores are low. You will learn to adjust your agent's instructions and test again. This eval-fix loop makes your agent better over time.
Use this guide when you run adk eval or debug your agent's results. It covers metrics, evalset schemas, and common failure causes.
Global
mkdir -p ~/.claude/skills/adk-eval-guideProject
mkdir -p .claude/skills/adk-eval-guideSource Repository
Grill Memattpocock/skills
Stress-test your plan with relentless questions until we both understand
Tddmattpocock/skills
Write one test at a time then code to make it pass
Test Driven Developmentobra/superpowers
Write a failing test first then code just enough to pass
Webapp Testinganthropics/skills
Test your local web apps quickly with Playwright automation and screenshots
Qamattpocock/skills
Turn bug reports into GitHub issues through natural conversation without technical fuss
Migrate To Shoehornmattpocock/skills
Replace unsafe as assertions with type-safe partial test data easily
Playwright Best Practicescurrents-dev/playwright-best-practices-skill
Master Playwright testing with best practices for reliable and fast tests
Google Agents Cli Evalgoogle/agents-cli
Run evaluations on your AI agent, find failures, and improve its quality step by step