Online evaluations help you check how well your AI responds to users. You attach judges to your AI configs to automatically score each answer. Judges use an LLM to give a score between 0.0 and 1.0.
You can use built-in judges for accuracy, relevance, and safety. You can also create your own judges for specific needs. This helps you monitor quality and improve your AI over time.
Anyone managing AI responses can set up evaluations with just a few steps. No complicated setup needed. Just a LaunchDarkly account and API token.
Global
mkdir -p ~/.claude/skills/online-evalsProject
mkdir -p .claude/skills/online-evalsSource Repository
Find Skillsvercel-labs/skills
Find and install the perfect skill to extend your AI agent
Microsoft Foundrymicrosoft/azure-skills
Build, deploy, and improve AI agents on Microsoft Foundry from start to finish
Azure Aimicrosoft/azure-skills
Search, transcribe, and analyze with Azure AI tools for smarter apps
Azure Hosted Copilot Sdkmicrosoft/azure-skills
Build, deploy, and manage your Copilot SDK apps on Azure with ease
Triagemattpocock/skills
Triage issues with a state machine driven by clear roles and agent briefs
Handoffmattpocock/skills
Hand off your work to another AI agent with a clear summary
Image Editagentspace-so/runcomfy-agent-skills
Smart router picks the best AI model for your image editing needs
Agentspaceagentspace-so/skills
See your AI agent's live folder from any browser instantly