Evaluation & benchmarking

Judge agent output, benchmark performance, and verify correctness.

Editor's picks

muratcankoylan/Agent-Skills-for-Context-Engineering

16.9K

This skill should be used when the user asks to "implement LLM-as-judge", "compare model outputs", "create evaluation rubrics", "mitigate evaluation bias", or mentions direct scoring, pairwise comparison, position bias, evaluation pipelines, or automated quality assessment.

Rigorous evaluation of agent output

affaan-m/everything-claude-code

225.6K

Evaluates agent output against 5-axis quality rubric (accuracy, completeness, clarity, actionability, conciseness). Use after any non-trivial task when the user wants a quality assessment, or when the agent-self-evaluation skill is active. Produces structured scorecard with evidence and improvement suggestions.

Judge agent performance

automagik-dev/genie

322

Performance-obsessed, benchmark-driven analysis demanding measured evidence (Matteo Collina inspiration)

Multi-perspective benchmarking

anthropics/claude-plugins-official

Solve competition math (IMO, Putnam, USAMO) with adversarial verification that catches what self-verification misses. Fresh-context verifiers attack proofs with specific failure patterns. Calibrated abstention over bluffing.

Adversarial verification for hard problems

Browse all in Agent & Harness Development

The picks above are curated. Search the full directory for everything matching “Evaluation & benchmarking” or related keywords.