Ryker_Feng 41658c5ff4
feat(skills): add skill review quality gate (#4037)
* feat(skills): add skill review quality gate

* fix(skills): skip review eval fixtures in CI

* fix(skills): ignore review eval fixtures in bundled scans

* fix(skill-review): harden review gate boundaries

* fix(skills): address skill review gate feedback
2026-07-11 15:58:07 +08:00

40 lines
1.0 KiB
Markdown

# Eval Design For Skills
Use evals to prove routing and behavior claims. Static review can recommend evals, but it must not claim runtime verification without retained runs.
## Trigger Evals
A trigger eval set should include:
- positive cases that should invoke the skill;
- negative cases that should route elsewhere;
- sibling collision cases when another skill has a similar boundary;
- short rationales for each case.
The existing trigger-eval list shape is supported:
```json
[
{
"query": "Review this skill for publication readiness",
"should_trigger": true,
"rationale": "Explicit skill review request"
}
]
```
## Behavior Evals
Behavior evals should retain:
- reviewed package digest;
- model and runtime identity;
- prompt and expected behavior;
- tool trace;
- output artifact;
- assertion or grading result.
## Baseline Comparisons
To claim improvement, retain both the baseline and candidate package digests. The report can only use `regression_verified` when comparison artifacts and grading evidence are present.