mirror of
https://github.com/bytedance/deer-flow.git
synced 2026-08-28 23:58:39 +00:00
* feat(skills): add skill review quality gate * fix(skills): skip review eval fixtures in CI * fix(skills): ignore review eval fixtures in bundled scans * fix(skill-review): harden review gate boundaries * fix(skills): address skill review gate feedback
1.0 KiB
1.0 KiB
Eval Design For Skills
Use evals to prove routing and behavior claims. Static review can recommend evals, but it must not claim runtime verification without retained runs.
Trigger Evals
A trigger eval set should include:
- positive cases that should invoke the skill;
- negative cases that should route elsewhere;
- sibling collision cases when another skill has a similar boundary;
- short rationales for each case.
The existing trigger-eval list shape is supported:
[
{
"query": "Review this skill for publication readiness",
"should_trigger": true,
"rationale": "Explicit skill review request"
}
]
Behavior Evals
Behavior evals should retain:
- reviewed package digest;
- model and runtime identity;
- prompt and expected behavior;
- tool trace;
- output artifact;
- assertion or grading result.
Baseline Comparisons
To claim improvement, retain both the baseline and candidate package digests. The report can only use regression_verified when comparison artifacts and grading evidence are present.