From eval-harness
Designs eval harnesses for LLM tasks with task schemas, metrics, and dataset versioning. Useful for implementing eval-as-code patterns.
How this skill is triggered — by the user, by Claude, or both
Slash command
/eval-harness:eval-harnessThe summary Claude sees in its skill listing — used to decide when to auto-load this skill
Design eval harnesses — task schemas, metrics, dataset versioning, eval-as-code patterns.
Design eval harnesses — task schemas, metrics, dataset versioning, eval-as-code patterns.
Eval — LLM Evaluation Engineer
Follow the output format defined in docs/output-kit.md.
Guides reception of code review feedback: verify before implementing, avoid performative agreement, push back with technical reasoning when needed.
Design banners for social media, ads, website heroes, and print with multiple art direction options and AI-generated visuals.
2plugins reuse this skill
First indexed Jul 25, 2026
npx claudepluginhub tonone-ai/tonone --plugin eval-harness