From ramco-brain
Carefully evaluate and grade the Ramco ERP "Brain" chatbot's answers against the ground-truth question bank (BrainEvals-Layer1). Use this skill WHENEVER you are scoring, grading, judging, benchmarking, or running an evaluation of Brain answers — or when the user says "eval the brain", "score these answers", "run the evals", "grade against ground truth", "judge as correct/hallucination", or asks how well the Brain answers. Two rules define careful grading here and this skill enforces them: (1) drop two-word and otherwise low-quality/malformed questions from the eval set — they are not real questions; (2) an answer that is MORE comprehensive than the ground truth (a superset) is CORRECT, not a hallucination — a hallucination is ONLY inventing facts that are not there. Judge on informational equivalence of the important points, semantically, in English.
How this skill is triggered — by the user, by Claude, or both
Slash command
/ramco-brain:eval-brainThe summary Claude sees in its skill listing — used to decide when to auto-load this skill
The Brain answers via the `brain_chatbot` agentic API; you grade each answer against a
The Brain answers via the brain_chatbot agentic API; you grade each answer against a
ground-truth expected_response in the question bank. Grading is LLM-as-judge, and the
whole value of this skill is doing it carefully and fairly — neither generous nor harsh, and
never mistaking thoroughness for hallucination.
A two-word "question" is not a good question and must not be scored. Before grading, filter the
bank and set aside (don't delete — mark excluded with a reason) any item that is:
EXPECTED_UNVERIFIABLE — the expected_response itself is wrong or unanswerable (a
test-bank quality issue, not a Brain failure).Report the excluded set separately with counts by reason, and score only the clean set. A rough heuristic for "too short": fewer than ~4 meaningful words and no clear ask — but use judgment, not just a word count; a short but well-formed question ("What does GRN stand for?") stays in.
The Brain often answers more comprehensively than the terse ground truth. That is good, not a failure. Grade on informational equivalence of the important points, not on length or format:
expected_response and those
points are correct, the answer is correct — even if the ground truth is only a subset
of what the answer says.Put simply: correct = (all important expected points present and right) AND (nothing fabricated). Extra truth is a bonus; missing an important point is a miss; invented facts are the only hallucination.
brain_chatbot (two-tier: unlimited
tool-rounds first, retry capped at 25 on error/limit). Capture brain_answer + brain_sources.expected_response on the 0–5 scale (see
references/rubric.md):
5 fully correct & complete · 4 complete, trivial omission · 3 core present, some
specifics missing · 2 topic present, major specifics missing · 1 barely any expected
content · 0 absent or incorrect. Apply Rule 2 while scoring: a superset answer whose
important points match earns a 5 (or 4 for a trivial omission) — the ground truth being a
subset does not cap the score.FULL = score ≥ 4, PARTIAL = 2–3, MISSING = 0–1;
answerable = score ≥ 2.hallucination (bool — only true for fabricated facts, per Rule 2),
gap_present (a genuine Brain knowledge gap, not a mere retrieval/surfacing miss), and the
gap_category (see rubric).source_reference cites an
idealized kb/<COMP>/layer*.json scheme that does not physically exist in the delivered
Brain (which holds the same content as markdown pages + JSONL registers + journey JSON). A
"wrong file cited" is not wrong if the fact is present. Use the GT→Brain layer mapping.brain/, read the cited pages) — many "gaps"
are the chatbot picking the wrong page (RETRIEVAL_ONLY), not absent knowledge.gap_description so a human reviewer can calibrate.id, score(0-5), verdict, answerable, brain_answer, brain_sources, hallucination(bool), gap_present(bool), gap_category, gap_severity, correct_facts, missing_or_wrong, suggested_resolution. Plus, for excluded items: excluded(true), exclusion_reason.
Full rubric, gap-category enum, worked examples, and the bank layout are in the references:
ALWAYS produce:
# Layer-1 Eval — <run id / date>
## Curation
- kept: N excluded: M (by reason: too-short K, out-of-scope K, unverifiable K, malformed K)
## Headline (clean set)
- mean score X.XX/5 · FULL P% · answerable P% · hallucinations H (fabricated facts only)
## Breakdown
- by difficulty, by category/BPC
## Gaps (genuine knowledge gaps only)
- by gap_category, top items with one-sentence descriptions + suggested_resolution
## Regressions vs prior run
- any previously-FULL question now PARTIAL/MISSING (investigate — see reindex-brain Regression Gate)
npx claudepluginhub rushyop/rushy-claude-plugins --plugin ramco-brainGuides reception of code review feedback: verify before implementing, avoid performative agreement, push back with technical reasoning when needed.
Design banners for social media, ads, website heroes, and print with multiple art direction options and AI-generated visuals.