From evals
Audits LLM eval coverage by identifying gaps, metric validity issues, benchmark leakage, and dataset staleness for evaluation engineers.
How this skill is triggered — by the user, by Claude, or both
Slash command
/evals:eval-reconThe summary Claude sees in its skill listing — used to decide when to auto-load this skill
Audit existing eval coverage — gaps, metric validity, benchmark leakage, dataset freshness.
Audit existing eval coverage — gaps, metric validity, benchmark leakage, dataset freshness.
Eval — LLM Evaluation Engineer
Follow the output format defined in docs/output-kit.md.
npx claudepluginhub tonone-ai/tonone --plugin evalsGuides collaborative design exploration before implementation: explores context, asks clarifying questions, proposes approaches, and writes a design doc for user approval.
Creates structured, bite-sized implementation plans from specs or requirements before writing code. Useful for breaking down multi-step tasks into testable steps with file structure and task boundaries.
Resolves in-progress git merge or rebase conflicts by analyzing history, understanding intent, and preserving both changes where possible. Runs automated checks after resolution.