By sickn33
Run structured AI product evaluation and operations workflows including A/B testing, LLM agent benchmarking, product analytics, and observability.
Structured guide for setting up A/B tests with mandatory gates for hypothesis, metrics, and execution readiness.
Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on real-world benchmarks
Expert in building products that wrap AI APIs (OpenAI, Anthropic, etc. ) into focused tools people will pay for. Not just "ChatGPT but different" - products that solve specific problems with AI.
Analytics de produto — PostHog, Mixpanel, eventos, funnels, cohorts, retencao, north star metric, OKRs e dashboards de produto.
Strategies for managing LLM context windows including summarization, trimming, routing, and avoiding context rot
Own this plugin?
Verify ownership to unlock analytics, metadata editing, and a verified badge. GitHub access is read-only (username + org membership).
Sign in to claimOwn this plugin?
Verify ownership to unlock analytics, metadata editing, and a verified badge. GitHub access is read-only (username + org membership).
Sign in to claimBased on adoption, maintenance, documentation, and repository signals. Not a security audit or endorsement.
A complete local skill catalog for coding agents—from project inspection and agent-owned selection to a reproducible, reviewable plan.
Current release: V15.1.0. This release includes AAS Core for complete local catalog search, agent-owned selection, manifest validation, planning, and diagnosis. Apply and recovery remain experimental and outside the supported preview path.
Codex or Claude inspects your project, enumerates its primary capabilities, searches and compares candidates across the complete local AAS catalog, and chooses the exact skills. Core imposes no semantic policy that favors a small stack; the manifest format has an explicit technical maximum of 128 skills. All 1,968 skills in the current catalog remain individually searchable, readable, and selectable. AAS Core does not rank or recommend skills. Its read-only compose_stack tool validates and returns the agent-owned manifest in memory; a client or the aas CLI persists the reviewed stack and its optional selection-evidence sidecar.
Read the AAS Core preview guide →
Project
-> inspected by Codex or Claude (not by AAS)
-> agent searches and reads the complete local catalog
-> AAS MCP (local stdio, read-only)
-> Codex or Claude chooses exact skill IDs
-> compose_stack validates the selection in memory (read-only)
-> client or AAS CLI persists aas-stack.json and optional evidence
-> AAS CLI validate + immutable plan preview
-> human review (optionally in Workbench)
The 1,967+ reusable SKILL.md playbooks, specialized plugins, bundles, workflows, and direct installers remain important. They are the content, curation, distribution, and compatibility layers around AAS Core—not competing primary products.
This is an independent community project. It is not affiliated with, sponsored by, endorsed by, or authorized by Google. Google, Antigravity, Gemini, and related product names are referenced only to describe compatibility and install targets. The GitHub repository is canonical; the hosted catalog and browser-local Workbench are companion discovery and review surfaces, not a hosted control plane.
The agent composes. You control. AAS keeps the stack reproducible.
AAS Core gives the repository one product model:
Plugin-safe Claude Code distribution of Agentic Awesome Skills with 1,933 supported skills.
Editorial "AAS Security Engineer" bundle for Claude Code from Agentic Awesome Skills.
Plugin-safe Claude Code distribution of Agentic Awesome Skills with 1,916 supported skills.
Editorial "Web Designer" bundle for Claude Code from Agentic Awesome Skills.
Editorial "AAS QA & Test Automation" bundle for Claude Code from Agentic Awesome Skills.
npx claudepluginhub sickn33/agentic-awesome-skills --plugin agentic-bundle-aas-ai-product-evaluation-opsEditorial "LLM Application Developer" bundle for Claude Code from Agentic Awesome Skills.
Implement comprehensive evaluation strategies for LLM applications using automated metrics, human feedback, and benchmarking. Use when testing LLM performance, measuring AI application quality, or establishing evaluation frameworks.
26 Agent Skills (several with runnable, unit-tested scripts) for building, evaluating, securing, and monitoring reliable LLM & AI-agent apps.
Benchmark, evaluate, and optimize skills to ensure reliable performance across all LLMs
🤖 AI Engineer — AI Engineer + LLM Systems Specialist
Multi-agent collaboration plugin for Claude Code. Spawn N parallel subagents that compete on code optimization, content drafts, research approaches, or any problem that benefits from diverse solutions. Evaluate by metric or LLM judge, merge the winner. 7 slash commands, agent templates, git DAG orchestration, message board coordination.