From BitRouter
Use when planning, launching, resuming, auditing, or reproducing a BitRouter Terminal-Bench 2.1 benchmark with Harbor, Terminus 2, and AWS EC2, including short policy iterations, one-time controls, model comparisons, or public full runs.
How this skill is triggered — by the user, by Claude, or both
Slash command
/bitrouter:run-bitrouter-benchmarkThe summary Claude sees in its skill listing — used to decide when to auto-load this skill
Treat the benchmark as a fail-closed experiment, not a best-effort load test. Keep the validated method fixed, make environment-specific inputs explicit, and accept no quality or cost claim until trial, routing, settlement, attribution, and cleanup evidence all agree.
Treat the benchmark as a fail-closed experiment, not a best-effort load test. Keep the validated method fixed, make environment-specific inputs explicit, and accept no quality or cost claim until trial, routing, settlement, attribution, and cleanup evidence all agree.
This skill is an operational document. It does not supply a runner, infrastructure module, or historical configuration. Inspect the selected BitRouter and Harbor revisions, adapt commands to their current interfaces, and freeze the resulting inputs before spending.
| Class | Typical scale | Trials | Valid use |
|---|---|---|---|
| Mechanism iteration | short 13 or a predeclared small slice | 1 per case | Debug routing, learning, metering, and runtime behavior |
| Replicated experiment | predeclared slice or full set | At least 3 per case | Estimate paired variation across clean policy lineages |
| Public reproduction | full 89-task set | 5 per case | Publish a reproducible Terminal-Bench result with provenance |
Never describe a one-trial tuning run as a public model score. Never describe r3 on tasks that taught r1-r2 as held-out generalization.
Keep these invariant unless deliberately starting a different benchmark method:
Configure AWS identity, account, region, network, instance shapes, source revisions, models, providers, secret sources, prices, task manifest, trial count, concurrency, timeouts, and central-host provision/reuse mode for each environment.
control+r1+r2 still records accepted r2 feedback before completion and must not invent r3.Stop the current lineage, preserve partial evidence, and clean up when any of these is true:
Do not apply feedback after a rejected round. Do not relaunch a started identity, rerun a control to improve its score, estimate unknown cost, delete failed attempts, or change concurrency inside a lineage.
| Need | Read |
|---|---|
| Experimental meaning, controls, trials, held-out design | methodology.md |
| Required inputs, AWS/IAM, source/model/price checklist | configuration.md |
| EC2/Harbor/Terminus lifecycle, resume, settlement, cleanup | operations.md |
| Strict gates, calculations, report and registry fields | acceptance-and-reporting.md |
| Diagnosing known benchmark failures | qna.md |
npx claudepluginhub bitrouter/bitrouter --plugin bitrouterRuns a canary suite of tasks to measure harness performance against ground truth, recording scores in trace-log.jsonl. Use before/after harness changes or via /mk:benchmark.
Run the pre-registered solo-vs-coordinated A/B benchmark: three arms (solo session, current orchestration, delegation-only sweep) over the same task matrix at matched verification rigor, measuring dollar cost, token band shift, quality, rework, and wall-clock. Use when the user asks "is orchestration worth it", "benchmark the pipeline against a solo run", "measure delegation value", "orchestration benchmark", or wants the crossover threshold below which a solo session beats delegation.
Drives the effortmining benchmark harness to measure pass-rate and token cost per effort tier, validate the instrument, run the matrix, grade, analyze, report, or refit the calibration table.