JasonH e1352bcdc0
fix(agents): distinguish undeclared and empty skill allowed-tools (#5669)
* fix(agents): distinguish undeclared and empty skill allowed-tools

* fix(ci): reduce guidance and isolate PostgreSQL test mocks

---------

Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
2026-09-22 11:36:28 +08:00

5.4 KiB

Agent System

Lead Agent (packages/harness/deerflow/agents/lead_agent/agent.py):

  • make_lead_agent(config: RunnableConfig) is the published langgraph.json entry point; preserve its signature and bare-graph return type.
  • Gateway calls assemble_lead_agent(config, *, app_config=None) for a LeadAgentAssembly(graph, descriptor); make_lead_agent returns .graph. assembly_descriptor.py::build_assembly_descriptor() records the model after runtime overrides, rendered prompt hash, authorized tools and middleware order. Consumers must unwrap via runtime/runs/worker.py::_agent_graph to also accept third-party factories returning bare graphs. The hashed skill catalog preserves allowed_tools=None (legacy allow-all), () (explicit allow-none for business tools), and non-empty allowlists as distinct policy states; allowlist ordering does not affect the fingerprint.
  • Dynamic model selection via create_chat_model() with thinking/vision support
  • Tools loaded via get_available_tools() - combines sandbox, built-in, MCP, community, and subagent tools
  • System prompt generated by apply_prompt_template() with skills, memory, and subagent instructions
  • Custom Agent memory_enabled: false disables memory reads, writes, tools, and compaction flushes while retaining date context; manual compaction trusts the state-producing checkpoint's agent binding, and global disable stays authoritative. Embedded clients cache the named policy by agent/user until reset_agent(); unreadable configs keep the legacy enabled default and log a warning.
  • Prompt trust: framework authority uses the system channel; user/model-influenced text uses the sanitized HumanMessage data channel. Never interpolate untrusted values into system text—even escaped tags do not neutralize natural-language injection (PR #5090).
  • Pass the same rendered prompt and middleware objects to the graph and assembly descriptor so observers describe the live graph, including Custom Agent allowed_subagents scope.

ThreadState (packages/harness/deerflow/agents/thread_state.py):

  • Extends AgentState with: sandbox, thread_data, title, artifacts, todos, uploaded_files, viewed_images, goal, promoted, delegations, skill_context, summary_text
  • Uses custom reducers: merge_artifacts (deduplicate), merge_viewed_images (merge/clear), merge_goal (preserve the active goal across ordinary state updates unless the goal writer replaces it), merge_promoted (catalog-hash-scoped deferred tool promotions), merge_delegations (append task delegation entries, same id latest wins, terminal status never downgraded, capped to the most recent entries), and merge_skill_context (dedupe active-skill references by path, keep the most recently read entries; entries store a name/path/description reference, not the SKILL.md body). summary_text is a LastValue channel updated by summarization and projected into model requests as durable context data instead of being stored as a messages item.
  • Delta-mode merge_message_writes normalizes the current message state once, then folds normalized writes in order with message-ID position indexes and deferred tombstone compaction. It preserves public add_messages behavior, including duplicate IDs, replacement position, removal errors, REMOVE_ALL_MESSAGES, null-write errors, and missing-ID allocation order, without rescanning the accumulated state for every write. Keep this full-parity contract covered by differential tests: LangGraph's private _messages_delta_reducer is also linear, but intentionally omits some of those public add_messages semantics and cannot be substituted directly.

Runtime Configuration (via config.configurable):

  • thinking_enabled - Enable model's extended thinking
  • model_name - Select specific LLM model
  • is_plan_mode - Enable TodoList middleware
  • subagent_enabled - Enable task delegation tool
  • max_concurrent_subagents - Per-response task call concurrency limit (clamped by SubagentLimitMiddleware)
  • max_total_subagents - Optional per-run total delegation cap override (falls back to subagents.max_total_per_run, clamped to 1-50) Gateway and DeerFlowClient.stream() always provide the runtime run_id; custom graph integrations must do the same. If it is absent, enforcement deliberately counts the thread's full delegation ledger (fail-restrictive) and emits a warning.

Direct subagent runtime: create_deerflow_agent(..., subagent_runtime=runtime) is the explicit dependency-injection path for direct graph callers. Reuse one deerflow.subagents.SubagentRuntime across every graph that belongs to the same application capacity boundary. With the default subagent feature it binds middleware concurrency/total limits, the ordinary task tool, one real execution controller, and any active durable-batch submitter to the same snapshot. A caller-owned batch repository requires await runtime.start() (or async with runtime) before graph construction and stop() at shutdown; the factory fails closed while that worker is stopped, and already-built bound batch tools must fail unavailable after it stops rather than falling through to another process-global submitter. The factory never creates SQL infrastructure, renders the caller-owned system_prompt, or mounts Gateway API/UI routes. Full middleware takeover cannot be combined with this runtime; direct callers and custom subagent middleware remain responsible for model-visible call-policy wording.