Evaluating AI agents is different from testing regular software. Agents make their own decisions and can take many paths to a goal. This skill explains how to judge their performance using outcome-focused methods.
Key ideas include multi-dimensional rubrics that score accuracy, completeness, and tool use. The 95% finding shows token usage drives most performance, followed by tool calls and model choice.
You will learn to handle non-determinism and build evaluation that catches regressions. Use LLM-as-judge and human evaluation for best results.
Global
mkdir -p ~/.claude/skills/agent-evaluationProject
mkdir -p .claude/skills/agent-evaluationSource Repository
Find Skillsvercel-labs/skills
Find and install the perfect skill to extend your AI agent
Microsoft Foundrymicrosoft/azure-skills
Build, deploy, and improve AI agents on Microsoft Foundry from start to finish
Azure Aimicrosoft/azure-skills
Search, transcribe, and analyze with Azure AI tools for smarter apps
Azure Hosted Copilot Sdkmicrosoft/azure-skills
Build, deploy, and manage your Copilot SDK apps on Azure with ease
Triagemattpocock/skills
Triage issues with a state machine driven by clear roles and agent briefs
Handoffmattpocock/skills
Hand off your work to another AI agent with a clear summary
Image Editagentspace-so/runcomfy-agent-skills
Smart router picks the best AI model for your image editing needs
Agentspaceagentspace-so/skills
See your AI agent's live folder from any browser instantly