mirror of
https://github.com/bytedance/deer-flow.git
synced 2026-09-28 23:46:21 +00:00
* fix(skills): accept portable frontmatter forms * fix(skills): normalize portable tool names * Safely preserve parenthesized portable skill tool patterns Portable Agent Skills declarations such as Bash(tvly *) contain spaces inside a command pattern. Keep those patterns as single literal entries while preserving exact names from the existing YAML-list form, so skill loading no longer fragments valid metadata or rewrites mixed-case MCP tools. Constraint: DeerFlow's current skill policy matches exact tool names and does not inspect Bash arguments Constraint: Agent Skills scalar syntax uses whitespace-separated entries with parenthesized command patterns Rejected: raw.split() | fragments Bash(tvly *) into unrelated tool names Rejected: normalize YAML-list entries | breaks case-sensitive MCP/runtime tool names Rejected: map Bash(...) to bash | broadens command-scoped declarations into unrestricted shell access Confidence: high Scope-risk: narrow Reversibility: clean Directive: Keep Bash(...) entries literal and inactive until DeerFlow has an explicit command-pattern authorization model Tested: 175 focused parser, validation, installer, review, loader, and tool-policy tests; Ruff check and format; compileall; git diff --check Not-tested: Full backend suite stopped at pre-existing Windows mode assertion test_runtime_config_store_file_is_owner_only Related: #4912 * Preserve exact custom tool names in portable skill parsing Portable scalar frontmatter needs alias normalization for known DeerFlow-compatible names, but generic case conversion corrupts MCP and custom tool identifiers. The tokenizer also treated quoted or escaped parentheses as structural delimiters, rejecting valid command patterns. Preserve unknown names and parse quoted or escaped patterns without broadening Bash(...) into bash. Constraint: Runtime skill policy uses exact tool-name matching Constraint: Parenthesized patterns remain literal because argument-level authorization is not implemented Rejected: Generic CamelCase-to-snake_case for every scalar | rewrites custom/MCP names Rejected: Map Bash(...) to bash | broadens command-scoped declarations into unrestricted shell access Confidence: high Scope-risk: narrow Reversibility: clean Directive: Add an explicit alias before supporting another portable tool name; keep command-pattern authorization separate Tested: 225 skills tests passed, 1 skipped; Ruff check; Ruff format --check; compileall; git diff --check Not-tested: Full backend suite remains affected by unrelated Windows permissions/path and missing Lark CLI tests Related: #4984; #4912 * Preserve case-sensitive exact tool authorities Case-folding a scalar declaration before alias lookup can turn literal write into write_file, substituting a different runtime authority. Keep exact portable spellings as aliases and preserve lowercase, custom, and MCP names; strengthen activation coverage for spaced Bash patterns and command fragments. Constraint: Runtime skill policy uses exact tool-name matching Constraint: Bash(...) remains literal and inactive because command-pattern authorization is not implemented Rejected: Case-insensitive alias lookup | maps lowercase runtime tools onto built-in authorities Rejected: Broaden the parser into command-pattern authorization | outside this PR's scope Confidence: high Scope-risk: narrow Reversibility: clean Directive: Add aliases only for documented portable spellings; preserve all other scalar names verbatim Tested: 226 skills tests passed, 1 skipped; Ruff check; Ruff format --check; compileall; git diff --check Not-tested: Full backend suite remains affected by unrelated Windows permissions/path and missing Lark CLI tests; GitNexus index refresh remains stale Related: #4984; #5016297602 * Support portable Glob and Grep skill aliases Portable Agent Skills commonly declare Glob and Grep, but DeerFlow exposes the runtime tools as glob and grep. Add explicit exact-spelling aliases and activation coverage so imported skills retain search-tool access without broad normalization. Constraint: Runtime skill policy uses exact tool-name matching Constraint: Alias conversion is limited to documented portable spellings Rejected: Case-fold all scalar names | can substitute custom or MCP authorities Rejected: Map arbitrary names by convention | breaks exact runtime compatibility Confidence: high Scope-risk: narrow Reversibility: clean Directive: Keep the alias table explicit and preserve unknown scalar names verbatim Tested: 228 skills tests passed, 1 skipped; Ruff check; Ruff format --check; compileall; git diff --check Not-tested: Full backend suite has unrelated environment failures on Windows; GitNexus index reports stale line mappings Related: #4984; #5026257899 --------- Co-authored-by: kriptoburak <kriptoburak@users.noreply.github.com>
23 KiB
23 KiB
Skills System (packages/harness/deerflow/skills/)
- Location: global public skills live under
deer-flow/skills/public/; user-authored custom skills live under{DEER_FLOW_HOME}/users/{user_id}/skills/custom/; globally managed integration skills live under{DEER_FLOW_HOME}/integrations/skills/{provider}/; per-user integration credentials remain under{DEER_FLOW_HOME}/users/{user_id}/integrations/{provider}/{config,data} - Format: Directory with
SKILL.md(YAML frontmatter: name, description, license, allowed-tools as a spec-compatible string or YAML list, argument-hint, required-secrets). Exact portable spellings such asBash,WebFetch,WebSearch,Glob,Grep,Read,Write, andEditmap tobash,web_fetch,web_search,glob,grep,read_file,write_file, andstr_replace; lowercase or otherwise unknown scalar names and YAML-list entries preserve their exact runtime spelling. Argument-scoped entries remain literal and inactive because the tool policy does not inspect arguments; the scalar tokenizer keeps spaces, quotes, and escaped parentheses inside patterns intact. - Loading:
load_skills()recursively scans public, per-user custom, global integration, and legacy custom locations forSKILL.md, parses metadata, and reads enabled state from extensions_config.json plus per-user skill state for non-public categories; that directory is a package boundary, so no nestedSKILL.mdis registered as a runtime skill. A custom skill directory may be a one-level symlink to an external directory for compatibility with operator-managed skill trees; activation still rejects a symlinkedSKILL.mdor deeper path escape. SkillScan has a deliberately narrower packaging rule: known eval fixtures are permitted as support data, while other nestedSKILL.mdfiles are reported as package defects. It parses runtime metadata and reads enabled state from extensions_config.json. - External reload:
POST /api/skills/reloadis an admin-only, process-local invalidation hook for trusted MinIO/NFS/CSI writes.SkillStorageinstances do not cache a catalog —load_skills()scans on every call — so the route clears all(app_config, user_id)entries and the rendered prompt-section LRU, then waits up to the shared refresh timeout for the existing off-loop single-flight refresh. Each invalidation receives a generation-bound result handle; a successful scan atomically replaces the global enabled-skills cache, while a loader-level failure propagates to the HTTP waiter and preserves the last-known-good global cache. Per-user/config scans capture the refresh version and cannot repopulate shared caches if invalidation occurs while they are loading. A timed-out HTTP wait fails generically while the daemon refresh worker continues. Subsequent runs rescan after a successful reload; active runs keep their existing snapshot. Each Uvicorn worker/Kubernetes Pod must be targeted separately. Direct mount writes bypass install/edit validation, SkillScan, and history, so mounted roots are an operator-controlled trust boundary. - Tool policy: Agent
allowed-toolsdeclarations apply dynamically only to slash-activated skills and skills captured inThreadState.skill_contextthrough configuredread_fileloads; passive enabled skills and custom-agent/subagent skill allowlists remain discoverable without clamping the baseline toolset. Subagents render only skill discovery metadata at startup and reuse the same adjacentSkillActivationMiddleware+SkillToolPolicyMiddlewarepair as the lead; their configuredskillsfield limits discovery and activation instead of eagerly loading bodies or unioning policies. Slash policy is dominant for its run, preventing subsequently read skills from widening explicit authority; autonomous captured skills use the existing union only when no slash source exists.tool_searchanddescribe_skillstay available as framework discovery infrastructure, while every discovered or promoted business tool still requires active-policy permission for schema visibility and execution;task,list_background_tasks, andcancel_background_tasklikewise require explicit declarations. Each active model call intentionally reloads the full live registry so enable/disable changes, frontmatter edits, and custom/public name-shadow winners take effect without a stale TTL or unsafe direct-path cache; all tool calls produced by that model step reuse the resulting source-and-path-signed decision. Registry failures and all-invalid active sets fail closed, while stale individual paths are skipped when another valid skill remains. This is best-effort behavioral scoping, not a hard security boundary: alternate loading paths are not captured and bounded autonomous context may evict entries. - Sandbox projection:
skills/projection.pymaterializes enabled-only trees at{base_dir}/skills_view/publicand{base_dir}/users/{user_id}/skills_view/{custom,legacy,integrations}. It copies files into the view (_copy_into_view) so a sandbox write cannot mutate the canonical skill inode; the operational trade-off is an O(total bytes) I/O and per-user storage multiplier across rebuilds, prioritized for write isolation over zero-copy hardlinks. Steady-state freshness checks combine source and view metadata tree digests, so in-sandbox view tampering is detected and repaired on the next acquire. Storage writes, archive installs, deletes, and toggles rebuild under a cross-process lock; Gateway boot ensures only the shared public view, while each user view is repaired lazily on first sandbox acquire. Managed integration packages are global, but their projected category is per-user because enabled state is isolated. Rebuilds stage a complete tree and reconcile it with per-file atomic replacement, so unrelated enabled skills remain continuously visible; disable/delete paths remove only the affected package before mutating to preserve fail-closed behavior. User projection rebuilds re-read global enable state from disk instead of the process singleton, so a toggle handled by another Gateway worker is reflected on the next acquire. Gateway public-skill toggles take the public projection lock before the sharedextensions_config_write_lock, re-read an existing config from disk, persist the full model shape, and rebuild before responding; keep this as one worker-owned critical section so MCP writes cannot interleave and request cancellation cannot release either lock while the worker still runs. The shared public steady-state signature check runs without the global projection lock; stale/error paths take the lock and re-check before rebuilding or clearing. User-scope checks remain serialized per user. Category root inodes remain stable so live bind mounts observe content changes without sandbox recreation. Projection failures clear the affected view before raising. - Injection (legacy / default): Enabled skills are listed in the agent system prompt with full metadata and container paths (
<available_skills>block). Controlled byskills.deferred_discovery: false(default). - Deferred discovery (
skills.deferred_discovery: true): Skills are listed by name only in a compact<skill_index>block, keeping the system prompt prefix-cache friendly. The agent calls thedescribe_skilltool at runtime to fetch full metadata for skills it wants to use, then loads the SKILL.md viaread_file. Two new modules support this path:skills/catalog.py—SkillCatalog(immutable, searchable; query forms:select:a,b,+prefix, free-text regex);select:returns all requested skills without a result cap; other modes cap atMAX_RESULTS=5.skills/describe.py—build_describe_skill_tool(catalog)builds thedescribe_skilltool as a closure;build_skill_search_setup(skills, enabled, ...)produces aSkillSearchSetup(describe_skill_tool, skill_names)that is wired into both the LangGraph agent factory (agent.py) and the embedded client (client.py).
- Slash activation:
/skill-name taskloads that enabled skill'sSKILL.mdfor the current model call only. The resolver rejects leading whitespace, missing separators, reserved channel commands (/new,/help,/bootstrap,/status,/models,/memory,/goal), disabled skills, and skills outside a custom agent's whitelist. - Installation:
POST /api/skills/installextracts .skill ZIP archive to custom/ directory - Managed integrations: Lark/Feishu CLI support installs one global official
lark-*pack as read-onlySkillCategory.INTEGRATIONentries under/mnt/skills/integrations/lark-cli/...; enabled flags, app configuration, and OAuth data remain per-user. Install resolves the newestlarksuite/clirelease from GitHub (releases/latest) at install time (falling back to a bottom-line pinned version if the lookup fails) rather than hard-coding the pack version; integrity relies on the official host + structural archive guards + a recorded hash of the effective installed tree after shared guidance injection (not a pinned archive-byte SHA, which GitHub does not keep stable). The Gateway image still installs a pinned@larksuite/clibinary, soget_lark_integration_statussurfaceslatest_available_versionandruntime_version_mismatchfor the UI. AIO installs additionally verify and publish official Linux amd64/arm64 binaries under{DEER_FLOW_HOME}/integrations/lark-cli/sandbox-cli, mounted read-only at/mnt/integrations/lark-cli/runtime;/mnt/integrations/lark-cli/config(app credentials, incl. the long-livedappSecret) is mounted read-only into the sandbox, its emptyconfig/lockssubdirectory is over-mounted writable forlark-clicoordination files, and/mnt/integrations/lark-cli/data(refreshable OAuth tokens) stays writable, all mapping to owner-only per-user directories. Sandbox trust boundary: the credential-bearing config and data dirs are still readable by arbitrary sandbox processes (the agent'sbashtool, or code reached via prompt-injection in a tool result), so the app secret and tokens are exposed to sandbox-side code even though they never reach the browser — the read-only config mount only prevents in-sandbox tampering, not read/exfiltration. The sidecar credential-broker (Pattern B, issue #4338) is the fix that removes these plaintext mounts from sandbox execution: setLARK_CLI_BROKER_IMAGEon the provisioner (seedocker/lark-cli-broker/) and the Gateway sendsprovision_lark_cli_brokeron sandbox create. The provisioner then runs alark-cli-brokersidecar that owns the per-userconfig/config/locks/datamounts (mounted into the sidecar only, at/var/lark/{config,config/locks,data}with only the nested locks mount writable) and serves thelark-clicommand surface on Pod loopback (http://127.0.0.1:8788); a shim init container (install-shim) writes a forwardinglark-cliinto the shared runtimeemptyDir, so the sandbox getsDEERFLOW_LARK_BROKER_URL+ a shim on PATH but no credential files. The on-PATHbin/lark-cliis a/bin/shlauncher that resolves a Python 3 interpreter and execs the Python shim body (bin/lark-cli-shim.py) beside it, so broker mode does not silently ENOEXEC on a sandbox image without a#!/usr/bin/env python3-resolvable interpreter — it fails loudly (exit 127, actionable message) and can be pinned withDEERFLOW_LARK_BROKER_PYTHON. The broker runslark-cliin the sidecar's cwd and cannot see the sandbox filesystem, so cwd is intentionally not forwarded and file-I/O subcommands relative to the sandbox cwd are unsupported (command surface only). An optionalDEERFLOW_LARK_BROKER_DENY_SUBCOMMANDSdenylist (comma-separated command prefixes, forwarded from the provisioner) lets the broker refuse secret-dumping subcommands before spawning the binary.lark_cli_env_overlay(broker=True)therefore omitsLARKSUITE_CLI_CONFIG_DIR/DATA_DIR;sandbox_lark_broker_active()(TTL-cached provisioner/api/capabilitiesprobe, tight timeout + longer negative caching on the bash hot path) selects broker vs. binary mode for both the bash env overlay and status.DEER_FLOW_LARK_CLI_SANDBOX_RUNTIME_DIRsupplies a validated, symlink-free pre-staged runtime for air-gapped deployments. For the remote provisioner (K8s), the runtime binary is otherwise provisioned by an optional init container + sharedemptyDir(Pattern A): setLARK_CLI_INIT_IMAGEon the provisioner (seedocker/lark-cli-init/) and the Gateway sendsprovision_lark_cli_runtimeon sandbox create once the pack is installed, so remote installs skip the Gateway-side GitHub download entirely. Broker (Pattern B) supersedes the init-container binary (Pattern A) when both images are configured.get_lark_integration_status(check_runtime=True)surfacessandbox_runtime_mode(none/gateway-download/init-container/broker) andsandbox_runtime_ready(remote modes read the provisionerGET /api/capabilities:lark_cli_init_image/lark_cli_broker_image) so a green UI can't hide a chat-timelark-cli: command not found. Cheap status probes are explicitly not live-verified; users authorize or reconnect through the browser device-flow endpoints instead of running terminal commands. - SkillScan:
packages/harness/deerflow/skills/skillscan/is the native deterministic scanner for.skillarchives and agent-managed skill writes. It runs offline before the LLM scanner, emits structured findings (rule_id,severity,file,line,message,remediation, redactedevidence— category/analyzer are encoded in therule_idprefix), blocksCRITICAL, and passes warning findings intoscan_skill_content(). The moderation adapter must normalize both plain-text responses and LangChain Responses API text blocks before parsing the required JSON decision.scan_archive_preflight()/scan_skill_dir()are pure sync functions (dispatch off the event loop);enforce_static_scan()applies the blocking policy and theskill_scan.enabledkill switch. The Python instance-client signal deliberately follows only a one-level, same-scope evidence chain (PR #4265 review): a proven imported constructor bound to a simple name, optional name-to-name alias propagation, rebinding invalidation, and a constructor-supported outbound method or context-manager use; bare canonical-looking names never fall back to module identity. Nested scopes never inherit client handles and inherit only constructor aliases proven stable by a binding-only enclosing-scope prepass. Comprehensions, walrus-bearing statements, annotations, executable expressions inside complex binding targets, unsupported operations, and ambiguous flows produce no finding from this signal; skipped constructs invalidate all names they may bind, while representative false negatives are pinned bytest_python_declared_false_negatives_stay_unreported. Compound bodies are walked from isolated copies so wrapping code inif True:is not a bypass, while copied scope entries, binding-only prepasses, and AST visits consume a deterministic work budget and the walk stops after its first sink. Budget or recursion exhaustion skips only this best-effort signal and retains deterministic findings already collected for the file. Do not add Semgrep/OpenGrep or YAML rule-engine dependencies to the core path; Phase 1 rule specs live in Python constants next to their analyzers inskillscan/orchestrator.py. - Skill Review Core:
packages/harness/deerflow/skills/review/provides read-only package snapshots, deterministic facts, resource/eval analysis, report rendering, and the CLI (python -m deerflow.skills.review.cli). It reuses the shared frontmatter helper and SkillScan; it must not importapp.*, execute target scripts, install dependencies, or call networks. JSON contracts live incontracts/skill_review/. Thereview_skill_packagebuilt-in tool labels results withreview_subject_entryand neverskill_context_entry, so reviewing a target does not activate it, bind itsrequired-secrets, or apply itsallowed-tools. Its model-visibleToolMessage.contentis a compact JSON payload with untrusted control tags neutralized; the full raw review payload, including Markdown renders, stays inToolMessage.artifact. CI should run the CLI with--fail-on error --fail-on-incompleteso blocker/error findings and truncated/not-assessed packages fail the gate. The publicskills/public/skill-reviewerskill owns semantic readiness review and suggestions only; mutation and runtime experiments remain owned byskill-creator.
Request-Scoped Secrets (required-secrets)
Lets a caller pass per-request, short-lived end-user credentials (e.g. an ERP token) to a skill's sandbox scripts without the value entering the prompt, tool arguments, the executed command string, or traces (issue #3861).
- Declare: a skill lists the secrets it needs in
SKILL.mdfrontmatter —required-secrets:as a string list or{name, optional}mappings.nameis both the lookup key and the env var name exposed to scripts. Parsed byskills/parser.py::parse_required_secretsintoSkill.required_secrets(SecretRequirement); malformed entries are dropped with a warning. - Carry: the caller sends values out-of-band in the run request's
context.secretsmapping (never a message).runtime/secret_context.pyowns the contract (SECRETS_CONTEXT_KEY,extract_request_secrets). The existingcontextpassthrough carries it toruntime.contextwithout mirroring intoconfigurable.build_run_configstill setsconfigurable.thread_idon the context path — the checkpointer requires it. MCP servers can read the same live carrier declaratively throughmcpServers.<server>.headers_from_context, and custom MCP interceptors throughextract_request_secrets(request.runtime.context)— notlanggraph.config.get_config()["context"], which isNoneinside a tool call because the run context rides the LangGraph runtime rather than the propagatedRunnableConfig; seedocs/MCP_SERVER.md. - Admission and redaction ownership:
services.py::start_run()validates both legacy request mappings,metadata.auth_tokenandconfig.metadata.auth_token, before any run or thread persistence.runtime/secret_context.py::redact_config_secrets()also removes nested config metadata secrets from observable and persisted config copies; historicalRunResponse.kwargsapplies the same redaction non-mutatively, leaving storedRunRecorddata unchanged. Keep callers onconfig.context.secretsrather than adding another credential carrier. Scheduled task definitions have no durable credential carrier:ScheduledTaskServicesupplies onlyscheduled_task_id,scheduled_task_run_id, andscheduled_triggeras run metadata. - Bind (point A+):
SkillActivationMiddleware._resolve_secret_bindingsrecomputes the injection set (runtime.context[__active_skill_secrets]) on every model call from two unioned sources, then REPLACES the key. (1) Slash: the run's most recent/skillactivation, persisted as a source on the run context (only the activated skill's canonical container path, never its declared secrets) so the whole tool loop after the activation call keeps the binding; a new activation replaces it. Slash reads the genuine user text viaget_original_user_content_text;InputSanitizationMiddlewarepreserves it (ORIGINAL_USER_CONTENT_KEY), so activation fires even after sanitization. (2) In-context (autonomous invocation): skills the model actually loaded in this thread —ThreadState.skill_contextentries. Both sources resolve the live registry skill by normalized container path on every call (_resolve_registry_skill) and bind only that skill's own declared secrets — enabled + allowlist checked for both; thesecrets-autonomous: falseopt-out (malformed values fail closed tofalse) additionally gates the in-context path but exempts explicit slash. Resolving by registry — not by trusting the source's stored data — is what makes a caller-forged__slash_skill_secret_sourceharmless (runtime.contextis caller-mergeable; the gateway also strips caller__-keys inbuild_run_config), #3938. Authorization is three-gated regardless of activation style: skill enabled by the operator × values supplied per-request by the caller (context.secrets) × names declared in frontmatter (∩ semantics). Because the set is recomputed per call, a skill evicted fromskill_context(capacity) or a caller that stops supplying a value loses injection on the next call. The injected value always comes from the caller's request, never the host environment (scrubbed first — see below), so a declared name that also exists in the host env is safe: the caller's value wins and the host value is dropped (the #3861 per-user-key-overrides-shared-key case). Missing required secrets are logged once per binding change, not injected; binding changes are recorded as amiddleware:skill_secretsjournal event (skill and secret names only, never values). - Inject:
bash_toolreads the injection set and passes it asexecute_command(env=...). Scope is the activation turn/run only — a run without/skillactivation injects nothing. - AIO image requirement: on
AioSandboxthe env path uses thebash.execAPI (POST /v1/bash/exec), which upstream all-in-one-sandbox only ships since1.9.3— older images (including alatesttag frozen on the1.0.0.xline) 404 the whole/v1/bash/*namespace.AioSandboxdetects the 404, remembers the capability gap on the instance, and fails fast with an actionable upgrade error instead of letting the model retry raw 404s; there is deliberately no fallback through the legacy shell path because none keeps the secret values out of the command string (#3921). Regression tests:tests/test_aio_sandbox.py::TestBashExecUnsupportedFailFast. - Inherited-env scrub:
execute_commandno longer leaks the Gateway'sos.environto skill subprocesses —env_policy.build_sandbox_envdrops secret-looking names (*KEY*/*SECRET*/*TOKEN*/*PASS*/*CREDENTIAL*/*DSN*+ a connection-string denylist likeDATABASE_URL/REDIS_URL/GH_PAT, plus no-flag credential sources likeMYSQL_PWD/REDISCLI_AUTH/PGPASSFILE/PGSERVICEFILE) so platform credentials never reach a skill; a skill that needs one must declare it. - Leak surfaces sealed (verified by a real-gateway e2e run — secret reaches the sandbox but none of these): prompt (value never in a message), trace (
tracing/metadata.pynever copiescontext), checkpoint (secrets live onruntime.context, not graph state), audit (journal records names only), stdout (tools.py::mask_secret_valuesredacts injected values from bash output), and run-record persistence + run API (services.py::start_runstoresredact_config_secrets(body.config)soruns.kwargs_jsonandRunResponse.kwargsnever carry the secret). - Historical retention: API response hiding prevents legacy
metadata.auth_tokenandconfig.metadata.auth_tokenfrom being returned now; it does not delete values already retained in databases, run events, logs, snapshots, exports, or backups. Deployments that ever used either legacy carrier must rotate the credential and clean every retained copy under their retention policy. Restarting or upgrading DeerFlow performs neither action. - Scope / non-goals: no persistence/vaulting — values are request-scoped and never stored server-side, so long-lived use means the caller re-supplies
context.secretson each request while the skill stays inskill_context; subagents do not inherit the skill injection set. MCP interceptors may independently consume the same supported request-scoped carrier. Tests:tests/test_skill_request_scoped_secrets.py,tests/test_mcp_session_pool.py.