Phoenix Evals helps you build tests for your AI applications. These tests check if your app responds correctly. You can use your own code or a language model to judge answers. Then you can compare results with human ratings to see how well your app works.
You can start with built in test types or create your own. The tool works with Python and TypeScript. It also lets you run experiments on large sets of data. This makes it easy to find and fix problems in your AI app.
Many teams use Phoenix Evals to improve their apps before going live. It helps catch mistakes and ensures your app is reliable and accurate.
Global
mkdir -p ~/.claude/skills/phoenix-evalsProject
mkdir -p .claude/skills/phoenix-evalsSource Repository
Find Skillsvercel-labs/skills
Find and install the perfect skill to extend your AI agent
Microsoft Foundrymicrosoft/azure-skills
Build, deploy, and improve AI agents on Microsoft Foundry from start to finish
Azure Aimicrosoft/azure-skills
Search, transcribe, and analyze with Azure AI tools for smarter apps
Azure Hosted Copilot Sdkmicrosoft/azure-skills
Build, deploy, and manage your Copilot SDK apps on Azure with ease
Triagemattpocock/skills
Triage issues with a state machine driven by clear roles and agent briefs
Handoffmattpocock/skills
Hand off your work to another AI agent with a clear summary
Image Editagentspace-so/runcomfy-agent-skills
Smart router picks the best AI model for your image editing needs
Agentspaceagentspace-so/skills
See your AI agent's live folder from any browser instantly