Aari ff0a6768c2
feat(subagents): add unified capacity and durable batch execution (#4998)
* feat(subagents): add capacity controls and durable batches

* fix(helm): sync subagent config schema version

* fix(subagents): preserve batch history without worker

* fix(subagents): support explicit factory runtimes

* fix: address durable batch review findings
2026-08-25 07:49:38 +08:00

5.0 KiB

Agent System

Lead Agent (packages/harness/deerflow/agents/lead_agent/agent.py):

  • Entry point: make_lead_agent(config: RunnableConfig) registered in langgraph.json. Its signature and bare-graph return type are a published ABI: LangGraph Server calls it directly, so neither may change.
  • assemble_lead_agent(config, *, app_config=None) -> LeadAgentAssembly(graph, descriptor) is the richer entry point the Gateway uses; make_lead_agent is a thin wrapper returning .graph. The descriptor is built by deerflow/agents/assembly_descriptor.py::build_assembly_descriptor() and captures what only the factory knows — the model resolved after runtime overrides, the rendered prompt hash, the tool list left by authorization, and the composed middleware stack in order. Consumers of a factory result must unwrap .graph defensively (see runtime/runs/worker.py::_agent_graph), because a third-party factory still returns a bare graph.
  • Dynamic model selection via create_chat_model() with thinking/vision support
  • Tools loaded via get_available_tools() - combines sandbox, built-in, MCP, community, and subagent tools
  • System prompt generated by apply_prompt_template() with skills, memory, and subagent instructions
  • Each assembly renders the system prompt and composes middleware exactly once; the same prompt and middleware objects must be passed to both create_agent() and the assembly descriptor so extension observations match the running graph, including Custom Agent allowed_subagents scope.

ThreadState (packages/harness/deerflow/agents/thread_state.py):

  • Extends AgentState with: sandbox, thread_data, title, artifacts, todos, uploaded_files, viewed_images, goal, promoted, delegations, skill_context, summary_text
  • Uses custom reducers: merge_artifacts (deduplicate), merge_viewed_images (merge/clear), merge_goal (preserve the active goal across ordinary state updates unless the goal writer replaces it), merge_promoted (catalog-hash-scoped deferred tool promotions), merge_delegations (append task delegation entries, same id latest wins, terminal status never downgraded, capped to the most recent entries), and merge_skill_context (dedupe active-skill references by path, keep the most recently read entries; entries store a name/path/description reference, not the SKILL.md body). summary_text is a LastValue channel updated by summarization and projected into model requests as durable context data instead of being stored as a messages item.
  • Delta-mode merge_message_writes normalizes the current message state once, then folds normalized writes in order with message-ID position indexes and deferred tombstone compaction. It preserves public add_messages behavior, including duplicate IDs, replacement position, removal errors, REMOVE_ALL_MESSAGES, null-write errors, and missing-ID allocation order, without rescanning the accumulated state for every write. Keep this full-parity contract covered by differential tests: LangGraph's private _messages_delta_reducer is also linear, but intentionally omits some of those public add_messages semantics and cannot be substituted directly.

Runtime Configuration (via config.configurable):

  • thinking_enabled - Enable model's extended thinking
  • model_name - Select specific LLM model
  • is_plan_mode - Enable TodoList middleware
  • subagent_enabled - Enable task delegation tool
  • max_concurrent_subagents - Per-response task call concurrency limit (clamped by SubagentLimitMiddleware)
  • max_total_subagents - Optional per-run total delegation cap override (falls back to subagents.max_total_per_run, clamped to 1-50) Gateway and DeerFlowClient.stream() always provide the runtime run_id; custom graph integrations must do the same. If it is absent, enforcement deliberately counts the thread's full delegation ledger (fail-restrictive) and emits a warning.

Direct subagent runtime: create_deerflow_agent(..., subagent_runtime=runtime) is the explicit dependency-injection path for direct graph callers. Reuse one deerflow.subagents.SubagentRuntime across every graph that belongs to the same application capacity boundary. With the default subagent feature it binds middleware concurrency/total limits, the ordinary task tool, one real execution controller, and any active durable-batch submitter to the same snapshot. A caller-owned batch repository requires await runtime.start() (or async with runtime) before graph construction and stop() at shutdown; the factory fails closed while that worker is stopped, and already-built bound batch tools must fail unavailable after it stops rather than falling through to another process-global submitter. The factory never creates SQL infrastructure, renders the caller-owned system_prompt, or mounts Gateway API/UI routes. Full middleware takeover cannot be combined with this runtime; direct callers and custom subagent middleware remain responsible for model-visible call-policy wording.