JasonH e1352bcdc0
fix(agents): distinguish undeclared and empty skill allowed-tools (#5669)
* fix(agents): distinguish undeclared and empty skill allowed-tools

* fix(ci): reduce guidance and isolate PostgreSQL test mocks

---------

Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
2026-09-22 11:36:28 +08:00

47 lines
5.4 KiB
Markdown

### Agent System
**Lead Agent** (`packages/harness/deerflow/agents/lead_agent/agent.py`):
- `make_lead_agent(config: RunnableConfig)` is the published `langgraph.json`
entry point; preserve its signature and bare-graph return type.
- Gateway calls `assemble_lead_agent(config, *, app_config=None)` for a
`LeadAgentAssembly(graph, descriptor)`; `make_lead_agent` returns `.graph`.
`assembly_descriptor.py::build_assembly_descriptor()` records the model after
runtime overrides, rendered prompt hash, authorized tools and middleware order.
Consumers must unwrap via `runtime/runs/worker.py::_agent_graph` to also accept
third-party factories returning bare graphs.
The hashed skill catalog preserves `allowed_tools=None` (legacy allow-all),
`()` (explicit allow-none for business tools), and non-empty allowlists as
distinct policy states; allowlist ordering does not affect the fingerprint.
- Dynamic model selection via `create_chat_model()` with thinking/vision support
- Tools loaded via `get_available_tools()` - combines sandbox, built-in, MCP, community, and subagent tools
- System prompt generated by `apply_prompt_template()` with skills, memory, and subagent instructions
- Custom Agent `memory_enabled: false` disables memory reads, writes, tools, and compaction flushes while retaining date context; manual compaction trusts the state-producing checkpoint's agent binding, and global disable stays authoritative. Embedded clients cache the named policy by agent/user until `reset_agent()`; unreadable configs keep the legacy enabled default and log a warning.
- **Prompt trust**: framework authority uses the system channel; user/model-influenced text uses the sanitized `HumanMessage` data channel. Never interpolate untrusted values into system text—even escaped tags do not neutralize natural-language injection (PR #5090).
- Pass the same rendered prompt and middleware objects to the graph and assembly descriptor so observers describe the live graph, including Custom Agent `allowed_subagents` scope.
**ThreadState** (`packages/harness/deerflow/agents/thread_state.py`):
- Extends `AgentState` with: `sandbox`, `thread_data`, `title`, `artifacts`, `todos`, `uploaded_files`, `viewed_images`, `goal`, `promoted`, `delegations`, `skill_context`, `summary_text`
- Uses custom reducers: `merge_artifacts` (deduplicate), `merge_viewed_images` (merge/clear), `merge_goal` (preserve the active goal across ordinary state updates unless the goal writer replaces it), `merge_promoted` (catalog-hash-scoped deferred tool promotions), `merge_delegations` (append task delegation entries, same id latest wins, terminal status never downgraded, capped to the most recent entries), and `merge_skill_context` (dedupe active-skill references by path, keep the most recently read entries; entries store a name/path/description reference, not the SKILL.md body). `summary_text` is a LastValue channel updated by summarization and projected into model requests as durable context data instead of being stored as a `messages` item.
- Delta-mode `merge_message_writes` normalizes the current message state once,
then folds normalized writes in order with message-ID position indexes and
deferred tombstone compaction. It preserves public `add_messages` behavior,
including duplicate IDs, replacement position, removal errors,
`REMOVE_ALL_MESSAGES`, null-write errors, and missing-ID allocation order,
without rescanning the accumulated state for every write. Keep this
full-parity contract covered by differential tests: LangGraph's private
`_messages_delta_reducer` is also linear, but intentionally omits some of
those public `add_messages` semantics and cannot be substituted directly.
**Runtime Configuration** (via `config.configurable`):
- `thinking_enabled` - Enable model's extended thinking
- `model_name` - Select specific LLM model
- `is_plan_mode` - Enable TodoList middleware
- `subagent_enabled` - Enable task delegation tool
- `max_concurrent_subagents` - Per-response `task` call concurrency limit (clamped by `SubagentLimitMiddleware`)
- `max_total_subagents` - Optional per-run total delegation cap override (falls back to `subagents.max_total_per_run`, clamped to 1-50)
Gateway and `DeerFlowClient.stream()` always provide the runtime `run_id`; custom
graph integrations must do the same. If it is absent, enforcement deliberately
counts the thread's full delegation ledger (fail-restrictive) and emits a warning.
**Direct subagent runtime**: `create_deerflow_agent(..., subagent_runtime=runtime)` is the explicit dependency-injection path for direct graph callers. Reuse one `deerflow.subagents.SubagentRuntime` across every graph that belongs to the same application capacity boundary. With the default subagent feature it binds middleware concurrency/total limits, the ordinary `task` tool, one real execution controller, and any active durable-batch submitter to the same snapshot. A caller-owned batch repository requires `await runtime.start()` (or `async with runtime`) before graph construction and `stop()` at shutdown; the factory fails closed while that worker is stopped, and already-built bound batch tools must fail unavailable after it stops rather than falling through to another process-global submitter. The factory never creates SQL infrastructure, renders the caller-owned `system_prompt`, or mounts Gateway API/UI routes. Full middleware takeover cannot be combined with this runtime; direct callers and custom subagent middleware remain responsible for model-visible call-policy wording.