Hyeonsang Cho f7f4a022e6
fix(agents): remove provider tool-call blocks when guards strip calls (#5447)
* fix(agents): remove provider tool-call blocks when guards strip calls

Token-budget and loop-detection hard stops, subagent-limit truncation,
and safety-finish-reason suppression removed calls from tool_calls and
the raw additional_kwargs payload, but left the provider's own
tool-call blocks in AIMessage.content. Provider adapters re-serialize
those blocks: langchain_anthropic sends a tool_use block whose id is
not in tool_calls, and the OpenAI Responses input builder sends every
function_call block. ChatAnthropic stores any tool-calling response as
a block list, so a guard firing on a Claude tool call always left a
tool_use without a tool_result. A truncated subagent call failed the
next model request of the same run; a hard stop was checkpointed under
the same message id and failed every later turn of the thread.

clone_ai_message_with_tool_calls now trims content tool-call blocks to
the calls that remain on the message: tool_use and LangChain v1
tool_call/tool_call_chunk by id, Responses function_call and
custom_tool_call by call_id (their id is the fc_ item id), Google GenAI
function_call by id, and id-less blocks by name in order. Blocks for
calls still on invalid_tool_calls stay, because
DanglingToolCallMiddleware answers those calls with placeholder
results. The token-budget and loop-detection hard stops now build their
messages through the helper instead of their own copies, and
ClarificationMiddleware drops its private filter, which matched
Responses blocks by item id.

* docs(changelog): reference #5447 in the orphaned tool-call block entry

* fix(agents): skip id-matched calls in the id-less block budget

The name budget for id-less content tool-call blocks counted every
retained call, including calls whose own id-bearing block had already
matched. In mixed-shape content, a retained call "a" with a
function_call block carrying id "a" also let a same-named id-less block
survive, leaving the unpaired block this helper exists to remove.

Collect the retained ids that id-bearing blocks matched first, and build
the name budget only from retained calls outside that set. Content with
no id-bearing blocks keeps the full budget, so the Gemini path is
unchanged.
2026-09-15 22:22:12 +08:00

36 KiB
Raw Blame History

Middleware Chain

Compaction preserves all state-level SystemMessages as framework instructions, including untagged legacy reminders. Transient instructions belong in request wrappers. A fully rescued partition skips compaction.

After latest-user rescue, if the inherited trimmer empties an AI/Tool-only window, format it and use _build_summary_input_text(strategy="last"). Keep normal human-anchored trimming and the final-message fallback for mixed windows whose human anchor falls outside the token-limited tail; head-first restoration can lose recent tool results. Tail truncation prefixes \n...\n only when marker and content fit. Budget raw sections before HTML escaping, wrappers, and prompt (not the final request); escape after trimming to preserve entities. Pass trim_tokens_to_summarize=None explicitly through the factory; omission restores LangChain's 4000-token default.

Persisted delegation verdicts are untrusted durable context; ledger rendering revalidates them and ignores malformed values. Completed is not accepted; retain useful work and address acceptance gaps.

Assembly order: tool_error_handling_middleware.py::_build_runtime_middlewares (exposed as build_lead_runtime_middlewares), then ../lead_agent/agent.py::build_middlewares appends lead-only entries. Optional entries require their config/runtime condition.

Message provenance. At injection/rewrite, always stamp additional_kwargs via deerflow_extension_api.provenance.provenance_kwargs(): deerflow_content_kind, deerflow_producer_kind, optional deerflow_producer_entity_id. All are server-owned inbound metadata; stamp even without observers, since downstream cannot recover producers. Producers: DynamicContext (reminder/memory), DurableContext (contract/data), SystemMessageCoalescing, ViewImage, SkillActivation. Summarization/Title use SystemOperationKind.SUMMARIZATION/.TITLE model-call attribution; summaries enter via DurableContext's stamped durable_context_data, not separate messages. Memory only queues extraction; recall uses DynamicContext's dynamic_context_memory stamp.

Middleware self-description. Behaviour-configurable middleware implements release_policy_parameters() -> dict[str, object] (duck-typed deerflow_extension_api.release.ReleasePolicyProvider, no base class). Use JSON-serialisable values and canonical_hash for long text, not prompt copies. collect_release_policies() gathers stack declarations; update them alongside every behaviour-affecting field. Summarization declares enabled task_continuity retention settings (otherwise None). DurableContext declares its normalized skills root, sorted read-tool names and continuity switch, so each capture/injection policy affects assembly identity without depending on private-field probing. Continuity history readers share shape validation, including the capture failure path and DurableContext rendering, so malformed persisted metadata cannot abort ordinary compaction or a model call.

Removing tool calls. Use clone_ai_message_with_tool_calls, not a bare tool_calls update: adapters resend stale content tool-call blocks, which strict providers reject.

Shared runtime base (build_lead_runtime_middlewares; subagents reuse most of this via build_subagent_runtime_middlewares):

  1. InputSanitizationMiddleware - First, so it is the outermost wrap_model_call wrapper; every inner middleware (including LLM retries) sees sanitized messages. additional_kwargs.original_user_content is server-owned provenance: Gateway strips caller-supplied values for non-internal run requests, trusted IM calls may carry the string they captured before adding transport/file context, and the middleware replaces any non-string value before wrapping. Uploads and sanitization retain first-writer-wins only for validated strings. Caller markers are marked untrusted_input, never stripped; scope is every turn.

  2. ToolOutputBudgetMiddleware - Caps model-bound tool output per app config. Externalizes oversized results to tool_output.storage_subdir (default .tool-results, constant TOOL_RESULTS_DIRNAME) under thread outputs, leaving a typed synopsis + read_file reference. These process-feedback files are excluded from workspace-change scans and delivery verification. wrap_model_call elides successful write_file content only in model-bound requests (#5328) after a later successful same-path read_file/write_file/str_replace; the on-disk file becomes the reference. Preserves the newest keep_recent_writes writes. Pairs call occurrences via tool_call_args.pair_tool_call_results and rewrites through shared tool_call_args helpers; controls: elide_superseded_writes, superseded_write_min_chars.

  3. ToolResultSanitizationMiddleware - Neutralizes framework/injection tags (e.g. <system-reminder>) and boundary markers in remote-content tool results (web_fetch/web_search/image_search/web_capture) so attacker-controlled fetched pages cannot forge trusted framework context. Mirrors InputSanitizationMiddleware's user-input guardrail for the other untrusted-content entry point; sits inner of ToolOutputBudgetMiddleware (neutralizes the raw output, then the budget truncates). Local tool output (bash/read_file) is left untouched. Scope is a name-based allowlist for the first-party web tools, plus every MCP-sourced tool via its deerflow_mcp metadata tag, so an MCP server naming its fetcher fetch_url is still covered

    Result-rewriting middlewares between the raw callable boundary and the model-visible result append a declared entry to additional_kwargs["deerflow_tool_transforms"] via agents/middlewares/tool_transform_meta.py::append_tool_transform. The trail is ordered by application — the last entry produced the final visible bytes — so an observer classifies raw→visible transforms from facts rather than by sniffing output wording.

  4. ThreadDataMiddleware - Creates per-thread directories under the user's isolation scope (backend/.deer-flow/users/{user_id}/threads/{thread_id}/user-data/{workspace,uploads,outputs}); resolves identity via resolve_runtime_user_id(runtime), including Gateway runtime context and standalone LangGraph Server auth, then falls back to the request ContextVar / "default"

  5. UploadsMiddleware - Tracks and injects newly uploaded files into conversation (lead agent only); upload existence checks use the same runtime-resolved user bucket as thread-data creation

  6. SandboxMiddleware - Acquires sandbox, stores sandbox_id in state. The lead runtime normally owns the thread's physical Agent-skill projection; delegated subagents and the prompt-only bootstrap agent are non-owners, so their narrower discovery allowlists never rebuild the shared thread view or force eager sandbox acquisition.

  7. DanglingToolCallMiddleware - Injects placeholder ToolMessages for AIMessage tool_calls that lack responses (e.g., user interruption), preserving raw provider tool-call payloads in additional_kwargs["tool_calls"]; malformed tool-call names and arguments are sanitized in the model-bound request so strict OpenAI-compatible providers do not reject the next request

  8. LLMErrorHandlingMiddleware - Converts provider/model failures to recoverable assistant errors. Async cancellation at admission, provider execution, retry events, or backoff releases only the call's own half-open probe (ownership assigned under the circuit lock), then propagates unchanged, without retry or failure accounting.

  9. Authorization / GuardrailMiddleware - Up to two independent pre-tool-call gates run here. When authorization.enabled, the AuthorizationProvider instance already used for Layer 1 capability filtering is wrapped by GuardrailAuthorizationAdapter and reused for Layer 2 execution checks. A generated tool_search bypasses the adapter's second provider call only when the current build has a concrete deferred setup; its catalog was already filtered by Layer 1, and an ordinary same-named tool without that deferred setup receives no exemption. When guardrails.enabled, the explicitly configured GuardrailProvider is appended after authorization and still evaluates every call, including tool_search. Authorization therefore runs outermost and can deny before an external guardrail call; both use the existing middleware's fail-closed, audit, sync/async, and error-ToolMessage behavior. See the authorization RFC and docs/GUARDRAILS.md.

    Every guardrail decision path publishes a neutral deerflow.authz.outcome.AuthorizationOutcome into the per-run runtime context, keyed by tool_call_id under the __-prefixed __authorization_outcome key (so build_run_config strips caller-supplied forgeries). Consumers pop it; the publisher and the consumer share only that contract module.

  10. SandboxAuditMiddleware - Audits sandboxed shell/file operations before tool execution; command classification is defense-in-depth and audit, not a security boundary (the sandbox is the isolation boundary). Command substitution is judged by position, not the presence of $(: command position ($(curl url), `curl url`, the word after |/&&/;, an eval/source argument) executes fetched content and is blocked; value position (x=$(curl url), echo $(curl url), an argument, a for word list) only captures output and passes (#4611). So _HIGH_RISK_COMMAND_POSITION_PATTERNS is matched anchored against each sub-command from _split_compound_command(split_pipes=True), never the whole string; pipe-spanning rules (| sh, base64 -d | ...) still use _classify_command's whole-command Pass 1. _COMMAND_POSITION_PREFIX extends the anchor over leading assignments and exec wrappers (FOO=1 $(curl url), env/command/builtin/exec/nohup/time/sudo/doas); its assignment branch requires whitespace before the substitution, which keeps x=$(curl url) in value position. Two contexts are deliberately position-blind (matched whole-command in Pass 1, since they execute their input anywhere, e.g. xargs sh -c "$(curl url)"): an eval/source argument, and an interpreter code-string flag — -c (shells, python), -e (perl/ruby/node), -p (perl/node), -r (php) — plus the here-string (<<<) reaching the same place via stdin. All three substitution spellings ($(, <(, `) share one _RISKY_SUBSTITUTION opener. An unquoted newline splits like ; (else echo hi\n$(curl url) evades the anchored rules). Heredoc bodies are data: _split_compound_command records headers (<<EOF, <<-EOF, <<'EOF') and consumes their bodies verbatim, so a body line starting $(curl url) isn't promoted to a command position; <<< (here-string, needs look-ahead + look-behind) and a << inside $(( ))/(( )) (bit shift, arithmetic depth tracked with the quote flags) must not open one. This is a heuristic, not shell parsing — an unterminated body consumes the rest of the string, an unclosed (( only disables heredoc detection, and the failure direction is always toward more command positions, not fewer. Known gaps: process substitution outside eval/source (. <(curl u)) is undetected, and two-step forms (x=$(curl u); eval "$x") need dataflow analysis. No config gate — appended unconditionally in _build_runtime_middlewares, for both lead and subagents.

  11. ReadBeforeWriteMiddleware - (optional, read_before_write.enabled, default on) Outermost write gate (#3857): read_file stamps a content hash on its ToolMessage; write_file (existing file, incl. append) and str_replace are blocked unless the newest mark for the path matches its current hash. Sits outside ToolProgress/ToolErrorHandling (a block consumes no ToolProgress slot); blocked results self-stamp deerflow_tool_meta and carry deerflow_write_block ({path, tool}). Marks live on messages, so summarization dropping the read invalidates the gate; writes never refresh marks. Gate check + execution are serialized per (thread, path); "Error: ..."-string sandboxes (AIO/E2B) fail open. It owns the composed call's sandbox authorization scope; SandboxAuthorizationError becomes an error ToolMessage. Its wrap_model_call swaps blocked calls' dead payload (content, old_str/new_str) for a deterministic placeholder in the model-bound request only (elide_blocked_payloads, elide_min_chars); state, receipts, and the journal keep the originals. Policy stays in the gate; the shared tool_call_args helper rewrites every arg surface together and every model-bound arg rewrite must use it.

  12. ToolProgressMiddleware - (optional, if tool_progress.enabled) State-machine-based stagnation guard (RFC #3177). Outer wrapper around ToolErrorHandlingMiddleware so its tool wrapper receives results already stamped with deerflow_tool_meta. Recoverable problems stay WARNED with hints; retryable non-recoverable problems escalate to BLOCKED; stop-category failures block immediately. It is independent of LoopDetection's turn-level call-pattern guard. Effective phase changes emit middleware:tool_progress, including reset when a later agent invocation (for example, goal continuation) clears WARNED/BLOCKED state. The state machine supplies the action and count threshold (null for category rules, recovery, and reset); recorder calls happen after releasing the state lock. Events include the actual sync/async hook plus bounded status, error, and recovery fields. Server-owned recorder keys determine subagent attribution. Arguments, content, prompts, and hashes are never persisted; recorder failures are fail-open. See event semantics.

  13. ToolReceiptMiddleware + ToolErrorHandlingMiddleware - ToolReceiptMiddleware is (optional, if verification.receipts_enabled, default on). It is the outermost wrap_tool_call layer — registered ahead of entries 9-12 — because Guardrail/SandboxAudit/ReadBeforeWrite/ToolProgress can short-circuit a call with their own ToolMessage (and SandboxAudit rebuilds medium-risk results); an inner receipt layer would silently gap the ledger on those results (ordering constraints in deerflow.extensions.ordering). Normal results still carry the deerflow_tool_meta status ToolErrorHandlingMiddleware stamps on the inner return path; short-circuit messages self-stamp meta or fall back to message.status. It stamps deterministic provenance (tool name, status, args/output hashes, byte count, timestamp) onto direct ToolMessage results and every matching ToolMessage carried in Command.update.messages, including delegated task, present_file, view_image, and tool_search results; before model calls it derives a hidden receipt ledger (display ids r1..rN) from message state, and when the 2,000-character budget is exceeded the newest receipts are retained in chronological order with their original ids plus an older-receipts omission marker. Rendering returns both the text and its retained receipt subset; every response that received a ledger carries only that exact server-owned subset, never omitted receipts. Snapshot validation accepts a strictly consecutive positive original-id range (for example r24–r30) rather than requiring r1, so subagent terminal citation verification resolves ids against evidence present in the citing turn even when later summarization drops and renumbers tool messages. Model-generated citation IDs are digit-bounded before integer conversion; oversized IDs are ignored as malformed input rather than raising through task write-back. Citation parsing deduplicates exact (id, anchor) pairs, not IDs alone, so repeated identical references stay compact while every distinct anchor claim is verified. Gateway strips delegated receipts/verdicts from external messages. ToolErrorHandlingMiddleware receives AppConfig, converts tool exceptions into error ToolMessages so the run can continue instead of aborting, stamps every result with deerflow_tool_meta (status / error_type / recoverable_by_model / recommended_next_action / source) via tool_result_meta.normalize_tool_result, stamps structured metadata for task exception wrappers, and stamps skill-read metadata for downstream durable-context capture. Task tool result text is generated from the same status/result/error inputs as the structured metadata so callers do not hand-write a second protocol string.

Authorization identity is independent of enforcement. Gateway strips client identity overrides: only the server auth source sets is_internal, and only authenticated IM body.context supplies channel_user_id (never body.config). build_principal_from_context applies role defaults, strict provenance, and copied attributes; RBAC rejects unknown defaults. Delegation and GuardrailMiddleware share this identity. Layer 1 precedes deferred assembly across agent paths and its provider is reused for Layer 2; framework skill/memory ordering stays stable. Trusted DeerFlowClient.stream() accepts identity overrides. Its graph key always includes effective storage user_id and, when enforced, the full Principal; nested attributes are copied so mutation cannot hide stale cache state.

Gateway route authorization uses authz.py::resolve_route_permissions() as the single provider integration point for both AuthMiddleware and decorator-only authentication. When enabled, it evaluates the six registered threads:* / runs:* permissions as resource="route" requests whose targets are the full resource:action strings. Decisions use the async provider API and are cached for the request in AuthContext; decorators do not call the provider again. Provider resolution or decision errors follow authorization.fail_closed, scoped per permission for decision errors. When authorization is disabled, the legacy complete permission set is returned without resolving a provider. Existing owner_check enforcement and require_admin_user() management gates remain independent and unchanged. Tests: tests/test_authorization_route_permissions.py, tests/test_auth.py, and tests/test_auth_middleware.py.

Model authorization uses authz.py::resolve_model_authorization() (same cached-provider, internal-role, and principal-building path as route authorization) as the Gateway integration point for the models router: list_models filters names through filter_resources(principal, "model", ...), and get_model enforces authorize(resource="model", action="use") with a deny surfacing as 403; provider errors follow authorization.fail_closed (fail-open returns the unfiltered list / proceeds). At runtime, lead_agent/agent.py::_authorize_model_name — called from _make_lead_agent and from DeerFlowClient._ensure_agent — applies the same model:use check to the resolved model name. On deny it scans the filter_resources-visible names (excluding the denied model), re-verifying each candidate with authorize("model", "use") before falling back, because a custom provider may allow list while denying use; no usable fallback raises under fail_closed and keeps the original model under fail-open. The built-in RBAC provider maps this to the per-role models policy key. Tests: tests/test_models_authorization.py.

Sandbox authorization (sandbox:execute) gates every sandbox acquisition before provider.acquire — including this middleware's eager path (before_agent / abefore_agent skip acquisition on deny instead of raising, deferring to the lazy per-tool gate). See the sandbox module guide authorization-gate paragraph and tests/test_sandbox_authorization.py.

Before changing a later authorization phase, read the authorization RFC and its implementation notes. The notes are the cumulative handoff record for merged PR behavior, reviewer feedback, trust-boundary decisions, deferred scope, and required regression coverage.

Lead-only middlewares (build_middlewares, appended after the base):

  1. DynamicContextMiddleware - Injects date and optional memory outside the static prompt. Opt-out removes its frozen server memory but retains date/user messages. Date follows server-local time unless DEER_FLOW_DATE_TIMEZONE names an IANA zone (invalid values fall back locally).
  2. SkillActivationMiddleware - Detects strict /skill-name task syntax on the latest real user message, resolves only enabled and runtime-allowed skills, injects the SKILL.md body as hidden current-turn context, and records a middleware:skill_activation audit event
  3. SkillToolPolicyMiddleware - Applies allowed-tools only after real activation; passive enabled skills and a custom agent's configured skill allowlist do not clamp the lead toolset. A run-scoped slash activation is authoritative and suppresses skill_context as a policy source, so reading another skill cannot widen the explicit skill's tools; without slash activation, skills captured after configured read_file loads retain the existing union semantics. The middleware filters model-visible schemas and blocks unauthorized execution, resolving canonical paths against the live enabled/agent-allowed registry on every model call, then stores a versioned, JSON-safe, middleware-token-bound decision signed by policy source plus active paths in run context for the resulting tool calls to reuse. The next model call always refreshes it, and malformed, foreign, stale, or unmatched decisions fall back to live resolution. tool_search and describe_skill remain framework-safe discovery tools under a restrictive policy; they may reveal or promote metadata, but a deferred business tool must still be declared by the active policy before its schema or execution can survive the policy middleware. The decision's owner token is authorization-sensitive, so its reserved context key is owned by runtime.secret_context and included in REDACTED_CONTEXT_KEYS for observable and persisted context copies. Registry load failures and a non-empty active set with no authorized skill fail closed to framework-safe tools; an individual stale path is skipped only when at least one valid active skill remains. This is best-effort behavioral scoping rather than a hard security boundary: alternate loads such as bash cat are not captured, and bounded autonomous skill_context can evict old entries. task is not framework-exempt, so a restricted skill cannot delegate around its policy. The middleware must remain immediately after SkillActivationMiddleware (which publishes the slash source through runtime.secret_context's public path helpers authenticated by a required token shared only within the assembled middleware chain) and immediately before DurableContextMiddleware; assembly and compiled-graph tests pin ordering, token sharing, schema filtering, and execution blocking.
  4. DurableContextMiddleware - Captures task delegations into ThreadState.delegations (including in-progress dispatches and terminal result summaries) and loaded skill-file references (name/path/description, parsed in-memory - not the body) into ThreadState.skill_context before summarization can compact the paired tool-call/result messages, then projects durable context into each model request. Static authority rules are injected as a SystemMessage; untrusted field values (summary_text, delegation results, skill descriptions) are injected separately as a hidden HumanMessage data block so compressed history, delegated work, and which skills are active stay visible without being stored as messages or promoted to system-role instructions. build_subagent_runtime_middlewares also attaches this middleware immediately before subagent summarization so a compacted summary_text is projected ahead of a preserved assistant/tool tail instead of leaving strict providers with an assistant-first request.
  5. SummarizationMiddleware - (optional, if enabled) Compacts near token limits; memory flush follows runtime policy, while manual compaction trusts the checkpoint-bound agent, so opt-out never captures removed turns. It preserves the latest real user request by ID and tagged DynamicContext reminders while allowing stale ID-swap peers into the summary. Moving the cutoff backward can retain old AI/tool turns and make first-turn compaction a no-op.
  6. TodoListMiddleware - (optional, if is_plan_mode) Task tracking with the write_todos tool
  7. TokenUsageMiddleware - (optional, if token_usage.enabled) Records token usage metrics; subagent usage is read from terminal ToolMessage.additional_kwargs in the current run and merged back into the dispatching AIMessage by message position. The same state update marks the ToolMessage with subagent_token_usage_attributed=true, so checkpoint replay or middleware re-entry cannot add the cumulative snapshot twice; missing/malformed usage or a result with no matching dispatch remains unmarked and retryable.
  8. TitleMiddleware - Auto-generates the thread title after the first complete exchange and normalizes structured message content before prompting the title model. If a first-turn run is interrupted before this middleware can write a title, runtime/runs/worker.py keeps the run in a finalizing state, persists a local fallback title from the latest checkpoint or original run input, and then syncs it to threads_meta.display_name. Replacement runs admitted by multitask_strategy="interrupt" / "rollback" wait for older same-thread finalization before entering the graph; the interrupted run only skips the fallback title write once a later run has started and may have advanced the checkpoint.
  9. MemoryMiddleware - Queues conversations for async memory update (filters to user + final AI responses); captures the runtime-resolved user so standalone LangGraph Server reads and writes stay in the same bucket
  10. ViewImageMiddleware - (optional, if the model supports vision) Appends a hidden HumanMessage with base64 image data, identified by a reserved ID prefix plus a server-owned metadata marker, to ModelRequest.messages in wrap_model_call / awrap_model_call. The payload lives only in that request and is never returned as a state update, so no checkpoint carries it and an interrupted run cannot strand it in history; state keeps only the lightweight viewed_images metadata. It owns that context and rebuilds it per call: its own message is swept out of the request first — a thread checkpointed by the earlier before_model/after_model pair (which wrote the payload into state and took it back out with RemoveMessage) can carry one that reached state but was never removed, and leaving it in would resend that base64 in every later request for the life of the thread — then a freshly built one is appended when warranted. The sweep requires both the reserved ID prefix and the server-owned marker, so a client cannot get its own message dropped; unmarked leftovers predating the marker are left in place and merely not duplicated
  11. McpRoutingMiddleware - Auto-promotes deferred schemas matching the latest real user message before DeferredToolFilterMiddleware, without executing tools. New names emit middleware:tool_promotion with source=routing_hint; repeat passes emit nothing
  12. DeferredToolPromotionAuditMiddleware - Observes final tool_search Commands; keep it outer of SkillToolPolicyMiddleware so denied names are excluded. It atomically claims new names per lead run or subagent execution to dedupe parallel searches, derives subagent attribution from the server-installed recorder, returns the original Command, and omits private payloads
  13. DeferredToolFilterMiddleware - (optional, if tool_search.enabled) Hides deferred (MCP) tool schemas from the bound model until tool_search or McpRoutingMiddleware promotes them (reads per-thread promotions from ThreadState.promoted, hash-scoped)
  14. SystemMessageCoalescingMiddleware - Merges every SystemMessage into a single leading SystemMessage per request; provider-agnostic fix for strict backends (vLLM/SGLang/Qwen/Anthropic) that reject non-leading system messages. Touches the per-request payload only (checkpoint state unchanged); on midnight crossings only the latest dynamic_context_reminder SystemMessage survives. The subagent builder places its date-only context middleware immediately before this coalescer, so the built-in subagent prompt and hidden date reminder still reach providers as one leading system block
  15. SubagentLimitMiddleware - (optional, if subagent_enabled) Truncates excess ordinary task tool calls to enforce both the per-response concurrency limit (max_concurrent_subagents, resolved against startup subagent_runtime.max_running and the 1-64 safety range before construction) and the per-run total delegation cap (max_total_subagents runtime override or subagents.max_total_per_run, default 6, clamped to 1-50). The total cap counts current-run entries in the durable delegation ledger (entries are tagged with run_id when captured), so repeated planning checkpoints in one run cannot keep launching legal-sized batches indefinitely, while later user turns in the same thread get a fresh run budget. Explicit durable batch_task calls are a separate mode with persisted total/live/running limits and are not rewritten into ordinary ledger entries. If the ordinary cap is exhausted, the middleware strips remaining task calls, forces finish_reason="stop", and appends a visible limit note so the run can synthesize existing results instead of ending with an empty tool-call response.
  16. LoopDetectionMiddleware - (optional, if loop_detection.enabled) Detects repeated tool-call loops; hard-stop clears structured, raw, and content-block tool calls before forcing a final text answer; stamps loop_capped via consume_stop_reason (#3875 Phase 2), symmetric to TokenBudgetMiddleware; persists warned-state transitions (first per call hash or per tool-frequency burst) and hard stops as middleware:loop_detection, attributed with is_subagent and the optional agent_id, without tool arguments, message content, tool results, or argument-derived hashes. Ordinary task subagents get dedicated recorder keys through a parent-loop proxy; never pass RunJournal into their isolated loop. Durable batch subagents have no parent run journal and do not persist these transitions State is run-scoped: new user runs get fresh budgets; same-run goal continuations share history. Keep sibling warnings isolated and lifecycle hooks topology-stable. Before changing this guard, read Loop detection lifecycle for fallback identity, cleanup/LRU/reset, severity ordering, and test invariants.
  17. TokenBudgetMiddleware - token_budget.enabled: shares run-ID budgets across continuations; missing/invalid IDs clear invocation state.
  18. Custom middlewares - (optional) Any custom_middlewares passed to build_middlewares are injected here, before config-declared extensions and the terminal-response/safety/clarification tail
  19. Configured extension middlewares - extensions.middlewares in config.yaml or extensions_config.json optionally accepts module.path:ClassName strings or {class, kwargs} objects. deerflow.reflection.resolve_class loads AgentMiddleware classes; import, class, and constructor errors fail agent creation. kwargs must be JSON-compatible; YAML dates/timestamps become ISO strings. Order: built-ins/custom and loop/token guards → extensions → terminal-response/safety/clarification tail. Subagents share the list before their safety tail; separate lead/subagent lists are unsupported. Trusted operator config only: paths instantiate arbitrary code. Gateway skill/MCP toggles preserve it in raw JSON; adding an API write path requires explicit trust-boundary review.
  20. TerminalResponseMiddleware - When a provider returns an empty terminal AIMessage after tool execution, injects a hidden recovery prompt and retries the model once; a second empty response is replaced in checkpoint state by a visible error fallback marked for the run worker, so the run finishes as an error instead of a silent success
  21. ModelLengthFinishReasonMiddleware - Records stop_reason=model_length_capped when provider-specific length detectors match a terminal AIMessage without tool-call intent (finish_reason=length / MAX_TOKENS, or stop_reason=max_tokens), preserving the original assistant content and never reparsing textual tool-call-like envelopes
  22. SafetyFinishReasonMiddleware - (optional, if safety_finish_reason.enabled) Suppresses tool execution when the provider safety-terminated the response (e.g. finish_reason=content_filter); registered after terminal-response/custom/configured middlewares so LangChain's reverse-order after_model dispatch runs it first
  23. ClarificationMiddleware - Intercepts ask_clarification, writes a readable ToolMessage.content fallback plus a structured ToolMessage.artifact.human_input payload, and interrupts via Command(goto=END) (must be last). after_model drops same-turn sibling tool calls so they cannot run before the user answers; a malformed ask_clarification parked on invalid_tool_calls is the same stop signal. disable_clarification runs keep the siblings. Payloads are versioned — legacy free_text/choice_with_other stay version: 1; the v2 form mode (from fields) is version: 2 so older frontends reject it and fall back to plain text. Field normalization is deterministic and lives in the middleware (it short-circuits before tool execution, so tool-arg typing gives no runtime validation), and it is atomic: any structurally broken entry — non-dict, bad/duplicate name, a name colliding with a JS Object.prototype member (__proto__/constructor), or exceeding the caps (16 fields / 24 options per field / 200 chars per text / MAX_FORM_SERIALIZED_BYTES = 16KB UTF-8, the per-item caps alone admitting forms whose IM text fallback overruns channel limits) — degrades the whole form to the legacy option/free-text modes, so a card never renders "complete" while missing a field. Benign issues degrade locally (unknown types — incl. unhashable JSON like type: [], which must not raise from the membership probe — and option-less selects become text); options are trimmed/deduped with blanks dropped (form- and top-level) since the frontend rejects blank labels. XML-to-dict option payloads are recursively flattened from dict/list containers in source order, scalar leaves kept, residual XML tags stripped before that trimming. Checkboxes are booleans defaulting to "no"; required on one means consent semantics. The response protocol is unchanged (v1 text/option): form cards submit a text summary as response_kind: "text", so journal persistence needs no new allowlist entries. Because this middleware can short-circuit before on_tool_end, RunJournal does a root-run reconciliation for ToolMessages whose tool_call_id came from the current run, so cards survive checkpoint compaction. That reconciliation is not ask_clarification-only — any middleware that answers a tool call has the same gap, and a result the user saw must not vanish on reload (#4666 — ReadBeforeWriteMiddleware blocked-write errors reached the UI but not the event store). It is bounded by three conditions, not a name allowlist: the message is user-visible, the call belongs to this run's lead agent (_remember_current_run_tool_calls records lead-agent calls only; subagent results stay in subagent.step), and it is not already persisted. Human Input Card replies are hide_from_ui HumanMessages with additional_kwargs.human_input_response; RunJournal persists only allowlisted hidden sources (currently ask_clarification) as llm.human.input.