muratcankoylan/Agent-Skills-for-Context-Engineering
★ 16.9KThis skill should be used when the user asks to "implement LLM-as-judge", "compare model outputs", "create evaluation rubrics", "mitigate evaluation bias", or mentions direct scoring, pairwise comparison, position bias, evaluation pipelines, or automated quality assessment.
Rigorous evaluation of agent output
affaan-m/everything-claude-code
★ 225.6KEvaluates agent output against 5-axis quality rubric (accuracy, completeness, clarity, actionability, conciseness). Use after any non-trivial task when the user wants a quality assessment, or when the agent-self-evaluation skill is active. Produces structured scorecard with evidence and improvement suggestions.
Judge agent performance
Performance-obsessed, benchmark-driven analysis demanding measured evidence (Matteo Collina inspiration)
Multi-perspective benchmarking
anthropics/claude-plugins-official
Solve competition math (IMO, Putnam, USAMO) with adversarial verification that catches what self-verification misses. Fresh-context verifiers attack proofs with specific failure patterns. Calibrated abstention over bluffing.
Adversarial verification for hard problems