Ryker_Feng 41658c5ff4
feat(skills): add skill review quality gate (#4037)
* feat(skills): add skill review quality gate

* fix(skills): skip review eval fixtures in CI

* fix(skills): ignore review eval fixtures in bundled scans

* fix(skill-review): harden review gate boundaries

* fix(skills): address skill review gate feedback
2026-07-11 15:58:07 +08:00

1.0 KiB

Eval Design For Skills

Use evals to prove routing and behavior claims. Static review can recommend evals, but it must not claim runtime verification without retained runs.

Trigger Evals

A trigger eval set should include:

  • positive cases that should invoke the skill;
  • negative cases that should route elsewhere;
  • sibling collision cases when another skill has a similar boundary;
  • short rationales for each case.

The existing trigger-eval list shape is supported:

[
  {
    "query": "Review this skill for publication readiness",
    "should_trigger": true,
    "rationale": "Explicit skill review request"
  }
]

Behavior Evals

Behavior evals should retain:

  • reviewed package digest;
  • model and runtime identity;
  • prompt and expected behavior;
  • tool trace;
  • output artifact;
  • assertion or grading result.

Baseline Comparisons

To claim improvement, retain both the baseline and candidate package digests. The report can only use regression_verified when comparison artifacts and grading evidence are present.