An evaluator acts like a judge for your AI models. It checks if the output is accurate or full of mistakes. You can set up an LLM-as-judge that uses a prompt to score answers. Or you can use a code evaluator that follows fixed rules without calling an AI.
This skill helps you create, update, and run these judges on the Arize platform. You can score individual spans, whole traces, or entire experiments. The results help you know if your AI is working as expected.
If something goes wrong, the skill tells you exactly what failed. It never makes up fake results. You fix the issue and try again.
Global
mkdir -p ~/.claude/skills/arize-evaluatorProject
mkdir -p .claude/skills/arize-evaluatorSource Repository
Find Skillsvercel-labs/skills
Find and install the perfect skill to extend your AI agent
Microsoft Foundrymicrosoft/azure-skills
Build, deploy, and improve AI agents on Microsoft Foundry from start to finish
Azure Aimicrosoft/azure-skills
Search, transcribe, and analyze with Azure AI tools for smarter apps
Azure Hosted Copilot Sdkmicrosoft/azure-skills
Build, deploy, and manage your Copilot SDK apps on Azure with ease
Triagemattpocock/skills
Triage issues with a state machine driven by clear roles and agent briefs
Handoffmattpocock/skills
Hand off your work to another AI agent with a clear summary
Image Editagentspace-so/runcomfy-agent-skills
Smart router picks the best AI model for your image editing needs
Agentspaceagentspace-so/skills
See your AI agent's live folder from any browser instantly