By tonone-ai
Design, run, and audit LLM evaluation pipelines — from experiment design and A/B analysis to eval harnesses, automated regression suites, and coverage gap detection — with CI integration and human eval orchestration.
Analyze A/B test results — statistical significance, practical significance, and segmentation.
Design an A/B test — power analysis, randomization, and success metrics.
Design eval harnesses — task schemas, metrics, dataset versioning, eval-as-code patterns.
Audit existing eval coverage — gaps, metric validity, benchmark leakage, dataset freshness.
Build automated regression suites — golden sets, threshold alerting, CI integration for model changes.
Uses power tools
Uses Bash, Write, or Edit tools
Own this plugin?
Verify ownership to unlock analytics, metadata editing, and a verified badge. GitHub access is read-only (username + org membership).
Sign in to claimOwn this plugin?
Verify ownership to unlock analytics, metadata editing, and a verified badge. GitHub access is read-only (username + org membership).
Sign in to claimBased on adoption, maintenance, documentation, and repository signals. Not a security audit or endorsement.
npx claudepluginhub tonone-ai/tonone --plugin evalsML/AI engineer — model training, MLOps, feature engineering, LLM integration
Editorial "AAS AI Product & Evaluation Ops" bundle for Claude Code from Agentic Awesome Skills.
AI Agent Team Operating System for Claude Code with 155 MCP tools (incl. 40+ ecosystem research tools), 25 agent templates, 14 hooks / 12 lifecycle events. Persistent team management, structured meetings, task wall with pipeline workflows, company loop engine, real-time React dashboard, and Ecosystem Research Platform v2 (progressive 4-stage deep-review funnel: shallow auto-summary → on-demand architecture → debate-based finalist evaluation → reference/integrate marking, with project-customizable thresholds, append-only history snapshots, and Failed self-learning).
Implement comprehensive evaluation strategies for LLM applications using automated metrics, human feedback, and benchmarking. Use when testing LLM performance, measuring AI application quality, or establishing evaluation frameworks.
v9.54.1 — Reliability wave: tangle contextual review correction loop with hard round ceiling, progress-supervised review rounds (per-agent stall watch, descendant-tree kills), council diversity and agy pin fixes, marketplace generator source-of-truth fix, provider troubleshooting runbook and cost-expectations docs. Run /octo:setup.
(forwward) Lean agent skills for building, shipping, strategy, and growth — no context bloat.
Design a fine-tuning pipeline
Diagnose runtime infrastructure issues — cold starts, timeouts, scaling problems, network failures. Use when asked about "infra is slow", "cold starts", "network issues", "why is this timing out", "scaling problem", "latency spikes", or "service is down".
Write UX copy for a feature, flow, or component
Audit existing model evaluation code
Identify open legal questions and research gaps