From evals
Designs eval harnesses for LLM tasks with task schemas, metrics, and dataset versioning. Useful for implementing eval-as-code patterns.
How this skill is triggered — by the user, by Claude, or both
Slash command
/evals:eval-harnessThe summary Claude sees in its skill listing — used to decide when to auto-load this skill
Design eval harnesses — task schemas, metrics, dataset versioning, eval-as-code patterns.
Design eval harnesses — task schemas, metrics, dataset versioning, eval-as-code patterns.
Eval — LLM Evaluation Engineer
Follow the output format defined in docs/output-kit.md.
npx claudepluginhub tonone-ai/tonone --plugin evalsGuides collaborative design exploration before implementation: explores context, asks clarifying questions, proposes approaches, and writes a design doc for user approval.
Creates structured, bite-sized implementation plans from specs or requirements before writing code. Useful for breaking down multi-step tasks into testable steps with file structure and task boundaries.
Resolves in-progress git merge or rebase conflicts by analyzing history, understanding intent, and preserving both changes where possible. Runs automated checks after resolution.
2plugins reuse this skill
First indexed Jul 25, 2026