Evals are like unit tests for AI agents. This skill helps you set up a formal evaluation framework for Claude Code sessions. You define what success looks like before you start building. Then you track how well the agent performs using pass@k metrics.
You can write different types of evals. Capability evals test new features. Regression evals make sure old features still work. Each eval has clear pass or fail criteria.
The framework uses three kinds of graders. Code graders run automated checks. Model graders use AI to judge open-ended results. Human graders flag items for manual review. This way you catch regressions and improve reliability over time.
Global
mkdir -p ~/.claude/skills/eval-harnessProject
mkdir -p .claude/skills/eval-harnessSource Repository
Find Skillsvercel-labs/skills
Find and install the perfect skill to extend your AI agent
Microsoft Foundrymicrosoft/azure-skills
Build, deploy, and improve AI agents on Microsoft Foundry from start to finish
Azure Aimicrosoft/azure-skills
Search, transcribe, and analyze with Azure AI tools for smarter apps
Azure Hosted Copilot Sdkmicrosoft/azure-skills
Build, deploy, and manage your Copilot SDK apps on Azure with ease
Triagemattpocock/skills
Triage issues with a state machine driven by clear roles and agent briefs
Handoffmattpocock/skills
Hand off your work to another AI agent with a clear summary
Image Editagentspace-so/runcomfy-agent-skills
Smart router picks the best AI model for your image editing needs
Agentspaceagentspace-so/skills
See your AI agent's live folder from any browser instantly