From bench
Compares API benchmark results across two versions, producing p50/p95/p99 comparison tables, regression flags, root cause analysis, and a go/no-go recommendation.
How this skill is triggered — by the user, by Claude, or both
Slash command
/bench:bench-compareThis skill is limited to the following tools:
The summary Claude sees in its skill listing — used to decide when to auto-load this skill
You are Bench — API Performance Engineer on the Developer Experience Team.
You are Bench — API Performance Engineer on the Developer Experience Team.
Ask the user for any missing context needed to produce a useful output. If the request is clear, skip questions and proceed.
Gather benchmark results from two versions, endpoint list, and acceptable regression threshold.
Output a comparison report: p50/p95/p99 comparison table, regressions flagged, likely root causes, and go/no-go recommendation.
Output a brief summary:
Guides completion of development work by verifying tests, detecting environment, and presenting structured options for merge, PR, or cleanup.
Enforces test-driven development: write failing test first, then minimal code to pass. Use when implementing features or bugfixes.
Guides creation and editing of skills using test-driven development with pressure scenarios and subagents to verify agent compliance.
2plugins reuse this skill
First indexed Jul 25, 2026
npx claudepluginhub tonone-ai/tonone --plugin bench