### Agent System **Lead Agent** (`packages/harness/deerflow/agents/lead_agent/agent.py`): - `make_lead_agent(config: RunnableConfig)` is the published `langgraph.json` entry point; preserve its signature and bare-graph return type. - Gateway calls `assemble_lead_agent(config, *, app_config=None)` for a `LeadAgentAssembly(graph, descriptor)`; `make_lead_agent` returns `.graph`. `assembly_descriptor.py::build_assembly_descriptor()` records the model after runtime overrides, rendered prompt hash, authorized tools and middleware order. Consumers must unwrap via `runtime/runs/worker.py::_agent_graph` to also accept third-party factories returning bare graphs. The hashed skill catalog preserves `allowed_tools=None` (legacy allow-all), `()` (explicit allow-none for business tools), and non-empty allowlists as distinct policy states; allowlist ordering does not affect the fingerprint. - Dynamic model selection via `create_chat_model()` with thinking/vision support - Tools loaded via `get_available_tools()` - combines sandbox, built-in, MCP, community, and subagent tools - System prompt generated by `apply_prompt_template()` with skills, memory, and subagent instructions - Custom Agent `memory_enabled: false` disables memory reads, writes, tools, and compaction flushes while retaining date context; manual compaction trusts the state-producing checkpoint's agent binding, and global disable stays authoritative. Embedded clients cache the named policy by agent/user until `reset_agent()`; unreadable configs keep the legacy enabled default and log a warning. - **Prompt trust**: framework authority uses the system channel; user/model-influenced text uses the sanitized `HumanMessage` data channel. Never interpolate untrusted values into system text—even escaped tags do not neutralize natural-language injection (PR #5090). - Pass the same rendered prompt and middleware objects to the graph and assembly descriptor so observers describe the live graph, including Custom Agent `allowed_subagents` scope. **ThreadState** (`packages/harness/deerflow/agents/thread_state.py`): - Extends `AgentState` with: `sandbox`, `thread_data`, `title`, `artifacts`, `todos`, `uploaded_files`, `viewed_images`, `goal`, `promoted`, `delegations`, `skill_context`, `summary_text` - Uses custom reducers: `merge_artifacts` (deduplicate), `merge_viewed_images` (merge/clear), `merge_goal` (preserve the active goal across ordinary state updates unless the goal writer replaces it), `merge_promoted` (catalog-hash-scoped deferred tool promotions), `merge_delegations` (append task delegation entries, same id latest wins, terminal status never downgraded, capped to the most recent entries), and `merge_skill_context` (dedupe active-skill references by path, keep the most recently read entries; entries store a name/path/description reference, not the SKILL.md body). `summary_text` is a LastValue channel updated by summarization and projected into model requests as durable context data instead of being stored as a `messages` item. - Delta-mode `merge_message_writes` normalizes the current message state once, then folds normalized writes in order with message-ID position indexes and deferred tombstone compaction. It preserves public `add_messages` behavior, including duplicate IDs, replacement position, removal errors, `REMOVE_ALL_MESSAGES`, null-write errors, and missing-ID allocation order, without rescanning the accumulated state for every write. Keep this full-parity contract covered by differential tests: LangGraph's private `_messages_delta_reducer` is also linear, but intentionally omits some of those public `add_messages` semantics and cannot be substituted directly. **Runtime Configuration** (via `config.configurable`): - `thinking_enabled` - Enable model's extended thinking - `model_name` - Select specific LLM model - `is_plan_mode` - Enable TodoList middleware - `subagent_enabled` - Enable task delegation tool - `max_concurrent_subagents` - Per-response `task` call concurrency limit (clamped by `SubagentLimitMiddleware`) - `max_total_subagents` - Optional per-run total delegation cap override (falls back to `subagents.max_total_per_run`, clamped to 1-50) Gateway and `DeerFlowClient.stream()` always provide the runtime `run_id`; custom graph integrations must do the same. If it is absent, enforcement deliberately counts the thread's full delegation ledger (fail-restrictive) and emits a warning. **Direct subagent runtime**: `create_deerflow_agent(..., subagent_runtime=runtime)` is the explicit dependency-injection path for direct graph callers. Reuse one `deerflow.subagents.SubagentRuntime` across every graph that belongs to the same application capacity boundary. With the default subagent feature it binds middleware concurrency/total limits, the ordinary `task` tool, one real execution controller, and any active durable-batch submitter to the same snapshot. A caller-owned batch repository requires `await runtime.start()` (or `async with runtime`) before graph construction and `stop()` at shutdown; the factory fails closed while that worker is stopped, and already-built bound batch tools must fail unavailable after it stops rather than falling through to another process-global submitter. The factory never creates SQL infrastructure, renders the caller-owned `system_prompt`, or mounts Gateway API/UI routes. Full middleware takeover cannot be combined with this runtime; direct callers and custom subagent middleware remain responsible for model-visible call-policy wording.