mirror of
https://github.com/bytedance/deer-flow.git
synced 2026-09-14 16:08:41 +00:00
* fix(sandbox): report truncated remote glob and grep results
BoxLite, Tenki, E2B, and OpenSandbox run find/grep in the sandbox, cap
the raw output with `| head`, and then filter those lines in Python:
ignored directories such as node_modules are dropped and grep's glob
scope is applied. They reported truncated only when max_results matches
survived the filter. When the capped lines were mostly filtered out, a
search with real matches past the cap came back short or empty with
truncated=False, and glob_tool/grep_tool rendered it as "No files
matched" / "No matches found". With the default max_results=200 and
1,200 files under node_modules, glob("**/*.py") reported no matches for
a workspace that has src/app.py.
remote_search_command now lets one line past its limit through, and
parse_remote_search_output(..., limit=) returns RemoteSearchOutput(text,
truncated): the first `limit` lines and whether the extra line arrived.
Exactly `limit` lines stays a complete result. Each provider passes the
cap it already computed to both calls and returns that truncated from
glob and grep when fewer than max_results results survive filtering.
The glob and grep tools now describe an empty truncated result as
incomplete instead of reporting no matches, which also covers AIO grep's
forwarded truncated flag. Sandbox.glob/grep document truncated as "the
matches may be incomplete".
* docs(changelog): reference #5427 in the remote search truncation entry
---------
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
168 KiB
168 KiB
Changelog
All notable changes to DeerFlow are documented in this file.
The format is based on Keep a Changelog and this project adheres to Semantic Versioning.
[Unreleased]
This section accumulates work toward the 2.1.0 milestone (milestone 2).
⚠ Breaking changes
- gateway: Request trace ids are now issued unconditionally, and every
Gateway HTTP response carries an
X-Trace-Idheader. Previously both were gated behindlogging.enhance.enabled, which now controls log output only — whether records carry atrace_idfield, and in which format. The header cannot be turned off; installations running the defaultenabled: falsewill start seeing it after upgrading. Scheduled tasks, MCP task notification runs, IM channel messages, and the embeddedDeerFlowClientbind an id per unit of work, so the id also reaches the run record, the checkpoint metadata, and Langfuse traces that previously had none. Adeerflow_trace_idsupplied in a run request'smetadataorconfig.contextis now ignored and overwritten so the response header, the logs, and the persisted run cannot disagree — send theX-Trace-Idrequest header to pin a correlation id across services.loggingremains restart-required. No config keys were added or removed. (#5119) - skills: Sandboxes now reserve
/mnt/skillsfor managed enabled-only projections.DEER_FLOW_HOST_SKILLS_PATHandSKILLS_HOST_PATHare no longer used; Docker/AIO and hostPath deployments derive projection paths fromDEER_FLOW_HOST_BASE_DIR. E2B operator mounts targeting/mnt/skillsor any child path are skipped with a warning so they cannot shadow the managed projection; move extra E2B content to a different container path. User projections re-read global enable state from disk so toggles propagate across Gateway workers on the next sandbox acquire. Existing E2B sandboxes retain their creation-time snapshot until they are recreated. PVC-backed provisioner deployments still mount the operator-supplied PVC snapshot directly, so disabled-skill filesystem isolation does not apply in PVC mode until dynamic PVC materialization is implemented. (#4178) - sandbox: E2B now enforces
sandbox.replicasas a process-local capacity limit. The defaultwaitpolicy waits foracquire_timeout, then fails the agent turn. DeerFlow does not retry the turn automatically. Useburstwithburst_limitto permit bounded extra VMs. Therejectpolicy can remove one warm VM before it returns a capacity error. (#4391) - skills: A directory containing
SKILL.mdis now a runtime package boundary. NestedSKILL.mdfiles inside that package are supporting data and are no longer registered as independent skills; unusual custom layouts must move independently loadable skills under a namespace directory without its ownSKILL.md. (#4098) - memory: The memory system is now pluggable (
memory.manager_classselects a backend; defaultdeermemis self-contained). DeerMem-private settings moved from the top level ofmemory:intomemory.backend_config, and the/memory/configresponse (andclient.get_memory_config()) changed shape. (#4122) - memory:
/memory/configandclient.get_memory_config()no longer return flat DeerMem fields (storage_path,max_facts,debounce_seconds,token_counting,guaranteed_*,staleness_*, ...). They return{enabled, mode, injection_enabled, manager_class, backend_config}wherebackend_configis an opaque dict the active backend self-interprets. Memory data responses (/memory,/memory/statusdata) are unchanged. External API/SDK clients reading the old flat fields must readbackend_configinstead. (#4122) - memory: Custom
memory.storage_classmoved: the old default pathdeerflow.agents.memory.storage.FileMemoryStorageno longer exists (nowdeerflow.agents.memory.backends.deermem.deermem.core.storage.FileMemoryStorage). CustomMemoryStoragesubclasses must acceptconfigin__init__(was no-arg). A broken/oldstorage_classlogs an error and falls back toFileMemoryStorage(won't crash) -- update the path + signature to restore it. (#4122) - memory:
storage_pathsemantics changed from a FILE path to a root DIRECTORY. Pre-abstraction, an absolutestorage_pathwas the shared memory file (opting out of per-user isolation) and a relative value was the global file under the data base_dir. Nowstorage_path(absolute or relative) is the root directory; per-user memory lives at{storage_path}/users/{uid}/memory.json. An upgrade keeping the old defaultstorage_path: memory.json(a relative file name) would orphan per-user memory or hitNotADirectoryErroron save, so the legacy migration drops file-stylestorage_pathvalues (ending in.json) with a warning and the factory raises ifstorage_pathresolves to an existing file. Setmemory.backend_config.storage_pathto a directory for a custom root. (#4122) - memory:
memory.mode: toolwith a backend that does not implementsearch()now fails fast at Gateway startup with aValueErrorfrom theMemoryManagerinvariant, instead of starting successfully and silently returning empty results on everymemory_searchcall. Both shipping backends implementsearch()(DeerMem retrieves;noopreturns[]), so this only affects a custom backend that onboards without overridingsearch(). It is intentional -- silent empties are worse than a loud startup error. Fix: switch tomode: middlewareor overridesearch()(and setsupports_search=True). (#4324) - config:
database.checkpoint_delta_snapshot_frequencymoved todatabase.checkpoint_delta.snapshot_frequencyand its default changed from1000to10. A legacy top-level value is still honored with a deprecation warning and mapped onto the nested key (an explicitly set nested key wins). Deployments that relied on the old default now snapshot 100x more often in delta mode -- setdatabase.checkpoint_delta.snapshot_frequency: 1000explicitly to keep the previous cadence. (#4516) - docker: The published entry port now binds to loopback (
127.0.0.1) by default in both compose files, matching the documented local-trust deployment model. Deployments that relied on the old0.0.0.0binding must setBIND_HOSTto expose the stack on other interfaces. (#4618)
Added
Scheduler
- scheduler: Scheduled tasks accept
interval(schedule_spec.every_seconds) in addition toonceandcron. Cadence is UTCnow + Nwith no missed-beat catch-up. N is at leastscheduler.min_once_delay_seconds(default 60s) and at most 30 days.
Authentication
- auth: Personal access tokens (PAT) for programmatic API access:
POST/GET/DELETE /api/v1/auth/patsmanage tokens (shown once, stored as SHA-256 digests); a default-deny route policy admits only the thread/run lifecycle routes, narrowed further by the token'sthreads/runsscopes, and any request dimension that carries cancel capability (?action=,multitask_strategy) additionally requiresruns:cancel. (#5041) - auth: Login throttling parameters are configurable via
auth.local.max_login_attempts(default 5, min 2) andauth.local.lockout_seconds(default 300), resolved live so a config reload applies on the next login without a Gateway restart — the unblock path for deployments behind corporate proxies/NAT where many users share one egress IP. Defaults are unchanged. (#5110)
Agents & runtime
- scheduler: Scheduled tasks can pin
assistant_idtolead_agent(the default) or a custom agent the owner already has. Unknown or malformed names return 422. The workspace create/edit form exposes the same choice. ([#5286]) - gateway:
GET /api/threads/{thread_id}/runs/pagewalks thread run history with a(created_at, run_id)keyset cursor ({data, has_more, next_before_created_at, next_before_run_id}).GET /api/threads/{thread_id}/runsstill returns a bare array of the newest 100 runs so LangGraph SDK clients keep working. (#5282) - middleware: New
TokenBudgetMiddlewareenforces a per-run token budget, shared additively across the lead agent and subagents. (#3412) - middleware: Structured tool-result metadata and a tool-progress state machine give the runtime first-class visibility into multi-step tool flows. (#3601)
- context: Record the effective memory identity per run and persist durable context (system messages, memory, and tool state) across summarization, emitting it as structured runtime metadata so compaction no longer drops it. (#3556, #3887, #3906)
- runtime: Goal continuations let a run resume toward a goal across multiple
agent turns, with
continuation_counttracked and capped. (#3858) - subagents: A system-maintained delegation ledger prevents redundant re-delegation of an in-flight task, and a total delegation cap bounds fan-out per run. (#3877, #4115)
- subagents: Persist and display subagent step history in the thread. (#3845)
- tools: Structured synopses replace raw oversized tool output in previews. (#3377)
- files: Deterministic read-before-write version gate for file tools prevents clobbering concurrent edits. (#3912)
- gateway: Cache-aware cost accounting attributes token costs to cached vs. uncached paths; a Redis stream bridge enables distributed event streaming; and manual context compaction is exposed to the user. (#3920, #3191, #3969)
- gateway: The stream-bridge heartbeat interval is configurable via
stream_bridge.heartbeat_interval_seconds(default 15s), so deployments behind aggressive proxy idle timeouts can tune SSE,/wait, and internal subscribers together. (#5017) - runtime: Dual-mode checkpoint storage with LangGraph
DeltaChannelcuts thread storage from O(N²) to near-linear for long research/coding runs. (#4292) - runtime: Delta-mode checkpoint history cache (memory/redis) with O(1)
incremental composition, configured via
database.checkpoint_cache. (#4638) - agent: Config-declared lead-agent middlewares let deployments add custom
AgentMiddlewareclasses without patching the runtime chain. (#3964) - agents: Per-agent model and generation settings (
temperature,max_tokens,thinking_enabled,reasoning_effort) override the shared model profile. (#4347) - runtime: Record terminal artifact-delivery receipts so runs expected to
present_filesno longer report success when delivery fails. (#4365) - uploads: Lazy-load historical files via a
list_uploaded_filestool instead of injecting the full manifest. (#4174) - scheduler:
scheduler.recursion_limitinconfig.yamlsets the LangGraph super-step cap for scheduled runs (default 1000, matching the web UI's interactive budget, clamped bymax_recursion_limit). (#4848) - runtime: Every tool call now carries a runtime-stamped, tamper-evident
tool receipt, and a bounded receipt ledger is injected into the model
context so agents can cite execution evidence in their reports. Enabled
by default via the new
verificationconfig section. (#4659) - subagents: Subagent delegations are now verifiable, layering RFC #4651:
every subagent's report contract requires citing tool receipts (e.g.
[r3 write_file]) and attaching a verifiable handle to each deliverable, the lead agent cross-checks those citations against the subagent's actual execution record, andacceptance_criteriaon ataskdelegation are checked deterministically parent-side (file existence/non-emptiness, recorded test-command exit status) with anything undecidable reported UNVERIFIED instead of silently passed. (#5076, #5090, #5109) - clarification: Human-input (clarification) cards support structured form fields, so an agent can request exactly the input it needs instead of free text only. (#4406)
- subagents: Built-in subagents now receive the current-date context anchor, so delegated tasks involving relative dates behave like tasks the lead agent handles directly. (#4797)
- subagents: A Settings page manages a deployment-level Subagent catalog
(admin-managed worker definitions alongside built-in and
config.yamlones), and Custom Agents can restrict delegation to an explicit worker allowlist enforced at both prompt and execution time. (#4887) - subagents: Subagent concurrency is now governed by one process-wide
capacity controller, and an opt-in
batch_tasktool runs large collections of independent items as durable, resumable SQL-backed batches with leases, bounded retries, pause/resume/cancel, and a chat panel for tracking progress. (#4998) - agents: The current-date context injected into lead- and subagent
prompts honors the optional
DEER_FLOW_DATE_TIMEZONEenv var (IANA name, e.g.Asia/Shanghai), so users in non-UTC deployments are no longer told the wrong "today" around midnight; unset keeps server-local behavior. (#5154) - subagents: Delegated subagents can discover files uploaded in earlier
turns: the parent run's validated
uploaded_filesboundary seeds the subagent's graph state, makinglist_uploaded_fileseligible for normal tool-policy filtering (durablebatch_taskworkers keep it disabled). (#5170) - agents: The read-before-write gate now elides the dead payload of a
blocked
write_file/str_replacecall (content,old_str,new_str) from model-bound requests. A blocked call never ran and must be re-issued after a re-read, so the original arguments only cost context; stored history, receipts, and the run journal keep them. Blocked results are paired with call occurrences (tool-call ids may repeat across turns), and a request whose history was rewritten drops OpenAIresp_response ids souse_previous_response_idchaining cannot resume the original server-side history. Controlled byread_before_write.elide_blocked_payloads(default on) andread_before_write.elide_min_chars(default 2000). - agents:
ToolOutputBudgetMiddlewarenow also elides thecontentof a successfulwrite_filecall from model-bound requests once the same path was read or modified again later in the conversation. After a successful write the file on disk is the source of truth, and the read-before-write gate forces aread_filebefore the next modification, so the historical copy was redundant with that read and long report-writing runs carried every section twice. The newesttool_output.keep_recent_writessuccessful writes (default 1) always stay visible,str_replacepayloads are never touched, and stored history, receipts, and the run journal keep the original arguments. Controlled bytool_output.elide_superseded_writes(default on) andtool_output.superseded_write_min_chars(default 2000).
Memory
- memory: Memory consolidation synthesizes fragmented facts, and a staleness
review prunes silently-outdated facts using LLM-assigned per-fact
expected_valid_days/staleFactsToExtend. (#3996, #3860, #4143) - memory: Guaranteed injection of correction facts (with graceful fallback) so user corrections always reach the model. (#3592)
- memory: Slim the pluggable
MemoryManagerinterface for backend onboarding - new backends no longer implement unused abstract methods, and DeerMem-specific hook injection moves out of the shared factory. (#4326) - memory: Incremental agent-scoped Markdown fact storage isolates per-agent facts and updates a single fact without rewriting or reindexing the whole collection. (#4279)
- memory: Memory message processing adds a conversation watermark, trivial-turn filtering, and a durable queue so extraction no longer re-feeds the full conversation every turn. (#4447)
- memory: A built-in FTS5/BM25 retrieval adapter provides full-text search over stored memories without an external retrieval service. (#4360)
- memory: New pluggable memory backends: OpenViking and mem0 over HTTP, plus Honcho as a user-model memory provider. (#4509, #4528, #4730)
- memory: A hybrid fact eviction policy blends multiple signals when deciding which stored facts to drop as memory fills. (#4789)
Skills
- skills: The built-in image-generation skill can use OpenAI-compatible Images APIs for generation and reference-image editing, with configurable endpoint, model, size, and output format.
- skills: Native SkillScan (phase 1) statically analyzes skill packages at
load, and
describe_skillenables deferred discovery so the model fetches a skill's schema on demand instead of loading all skills up front. (#3033, #3775) - skills: Per-user custom skill isolation with sandbox mounting. (#3889)
- skills: The skill list reopens after a skill is selected, so several skills can be attached in a row. (#4639)
- skills: Install local
.skillarchives directly from the Skills settings page, reusing the existing per-user installer and security scan. (#5039) - skills: Podcast-generation Volcengine voices are configurable per speaker gender, with trimmed blank-safe defaults. (#5156)
Models & integrations
- community: New web search/fetch engines - GroundRoute, Crawl4AI
(
web_fetch), and a fastCRW provider - plus a Browserlessweb_capturescreenshot tool and Braveimage_search. (#3675, #3821, #3585, #3881, #3866) - mcp: Per-server
tool_call_timeoutfor MCP tool calls, and routing hints that guide the model to the right server. (#3843, #4004) - mcp: Add an official OpenViking
/mcpexample that exposes the native tool set through DeerFlow's generic MCP client. (#4745) - community: Agentic browser control as a first-class thread capability - Playwright-backed browser sessions the agent operates while the user observes or takes over from the workspace. (#4187)
- community: Lark/Feishu CLI integration bundles the runtime install, the
official
lark-*skill pack, and an interactive auth flow so the integration is no longer environment-dependent. (#3971) - integrations: Lark/Feishu app credentials can be switched per user from Settings > Integrations: new App ID/Secret values are validated before anything is committed, and the previous OAuth token is revoked after a successful switch. (#4703)
- acp: MiniMax Code (
mcode acp) is supported and documented as a native external coding agent, and ACP thought chunks are no longer concatenated into tool results. (#4846) - models: A Z.AI GLM-5.3-Flash profile keeps thinking permanently enabled and stops generic reasoning-effort forwarding, since the model rejects disabled thinking and only accepts its own effort values. (#5074)
- community: New web search providers - Serply (with news and scholar verticals) and Tencent Cloud WSA - plus native recency filters (day/week/month/year) shared across DDGS, Brave, Tavily, and SearXNG. (#5023, #5057, #5099)
- community: New Sofya
web_searchandweb_fetchprovider - search results carry the content of each page, capped per result so a default search stays inline. (#5239) - knowledge: Opt-in read-only RAGFlow retrieval exposes a
knowledge_search(query)agent tool over configured RAGFlow datasets, with a dataset-ID allowlist and credential/dataset-id redaction on error paths. (#4955) - knowledge: Opt-in read-only LightRAG retrieval is an alternative
knowledge_search(query)provider for the sameknowledgegroup — operators pick RAGFlow or LightRAG by which entry is configured. It queries LightRAG's structured/query/dataendpoint (no LLM generation), formats the ranked chunks as citation-numbered text, redacts the optionalX-API-Keyon every model-visible path, and maps server error messages to actionable tool errors. (#5209)
MCP
- mcp: A durable task runtime for MCP: long-running tool tasks survive Gateway restarts through a durable driver, and their progress and completion notifications surface in the chat UI. (#4665, #4690, #4833)
- mcp: Shared MCP servers can inject per-user credentials: a single server entry authenticates each DeerFlow user with their own header value, unmapped users are denied by default, and stored credentials are masked in Gateway API responses. (#4868)
- mcp: Per-server
tool_name_prefixoption lets servers that already namespace their own tools keep their original tool names; the default behavior is unchanged. (#4624) - mcp: Settings > Tools can add, edit, and delete MCP servers through targeted Gateway endpoints, with a copy-paste JSON workflow that preserves advanced fields and masked secret placeholders. (#5022)
- mcp: Shared HTTP/SSE servers can map request-scoped secrets to headers
via
headers_from_context: callers supply per-request values inconfig.context.secrets, the config stores only key names, and missing values deny by default. (#5010) - mcp: An optional
parallel-searchserver entry (https://search.parallel.ai/mcp, HTTP, no auth by default) ships inextensions_config.example.json, disabled by default; enabling it exposesparallel-search_web_searchandparallel-search_web_fetch, with optional Bearer authentication documented. (#5028)
Channels
- channels: Expose the IM
channel_user_idto sandbox commands asDEERFLOW_CHANNEL_USER_ID. (#3926) - channels: Queue rapid same-thread messages and preserve topic-card previews across batches. (#3988)
- channels: Inbound webhook deduplication moves to Postgres, so several Gateway pods can serve the same IM channel without double-processing events. (#4210)
- channels: DingTalk inbound messages support file and image attachments. (#4423)
- channels: New Buzz (Nostr) channel connector, including the frontend experience for the channel. (#4649, #4727)
- channels:
/agent listand/agent use <name>IM commands let a conversation switch to the owner's Custom Agents: the selection is persisted in thread metadata (restored after restart) and wins over stale channel defaults, IM-created threads route through the same agent when opened in the web UI, and/agentis reserved across the slash-skill parser, frontend, and TUI so no skill can shadow the command. (#5168)
Auth & guardrails
- auth: Generic OIDC/SSO authentication with Keycloak support. (#3506)
- guardrails: Authenticated runtime context is exposed in
GuardrailRequest, and security interventions are persisted as run events. (#3665, #3837) - auth: "Keep me signed in" login option with a centralized session-cookie
policy (persistent
Securecookies on HTTPS, session cookies on public HTTP). (#4255) - auth: Deployments can close local self-registration to restrict new accounts to SSO/OIDC provisioning. (#4311)
- authz: Built-in RBAC authorization provider with a unified factory, plus tool-authorization enforcement at both assembly (tools removed before the model sees them) and runtime (denied calls blocked). (#4260, #4370)
- authz: Gateway route permissions are derived from the configured AuthorizationProvider rather than a fixed table. (#4439)
- authz: Model authorization is enforced at Gateway routes and again in
the agent runtime, and
sandbox:executeis checked when a sandbox is acquired - users can no longer reach models or sandboxes they are not authorized for. (#4540, #4911) - authz:
GET /auth/menow surfaces the caller's effective route permissions (RFC #4063 Phase 4), read from theAuthContextthe auth middleware already stamps on every authenticated request — no extra provider evaluations — so the frontend can hide actions the caller's role cannot perform. (#5228)
Sandbox & provisioner
- sandbox: New E2B and BoxLite (micro-VM) sandbox providers; BoxLite ships with a warm pool. (#3883, #3940, #3951)
- provisioner: ClusterIP Services and scoped per-skill PVC mounts, plus a configurable sandbox container port. (#4016, #3928)
- sandbox: New cloud sandbox providers: Tenki and OpenSandbox. (#4382, #4877)
- sandbox: An optional lark-cli credential broker sidecar (K8s provisioner mode) keeps Lark app secrets and OAuth tokens out of the sandbox filesystem entirely - the sandbox sees only a shim that forwards commands to a loopback broker in the pod. Off by default. (#4501)
- sandbox: The E2B mount-upload wall-clock deadline is configurable via
mount_upload_deadline_seconds(default 120s). (#4876) - sandbox: An E2B sandbox carries a structured
MountUploadResult(truncated,reason, upload totals) after creation — preserved across warm-pool reclaim within the process — so mount truncation by a resource limit is observable in code instead of only in Gateway logs. (#4884) - sandbox: Opt-in controlled egress for local Docker AIO sandboxes:
sandbox.network.modesupportsisolated(per-sandbox internal bridge with no outbound route) andallowlist(the same bridge plus a domain-allowlist HTTP(S) policy sidecar that resolves destinations itself and rejects IP literals and ECH); denied public domains can be approved through the Human Input card (temporary grant or allow-for-this-sandbox), non-interactive runs fail closed, and the sandbox API is no longer published directly in restricted modes. Requires Docker Engine 28+;openremains the default. (#5152)
Extensions & plugins
- extensions: An out-of-tree Python extension system: extensions can
contribute middleware, task-lifecycle and system-model observers, Gateway
services, and HTTP routers, and are managed with
deerflow extensionsinstall/enable/disable/remove. (#4636, #4684, #4780) - extensions: Extensions can observe what the agent did - message
provenance, middleware policy declarations, agent-assembly fingerprints,
context-compaction records, guardrail decisions, and the MCP origin of a
tool.
deerflow-extension-apimoves to 0.2.0; extensions written against 0.1 are refused at startup with an install hint. (#4863)
Persistence
- persistence: A custom PostgreSQL schema can be selected via
postgres_schema; ORM, LangGraph checkpointer, and store tables are all created there, and the schema is created automatically at startup. (#3442)
Frontend
- frontend: Branching support for assistant turns and side conversations for quoted follow-ups. (#3950, #3934)
- frontend: Regenerate the latest answer. (#3637)
- frontend: Citation-sources evidence panel, workspace change review for
agent runs, and a visualized
ask_clarificationcard. (#3907, #3945, #3956) - frontend: Voice dictation, prompt-history recall with arrow keys, composer input polishing, and a "(thought for N seconds)" thinking-duration chip. (#4036, #3718, #3986, #3627)
- frontend: Feature-gate the agents UI behind the
agents_apiflag, and persist AI turn duration in backend and UI. (#3769, #3663) - frontend: Render slash-skill activations as inline chips. (#3981)
- frontend: Localized AI-assistance disclaimer. (#4374)
- frontend: Pin recent chats. (#4442)
- frontend: Validate
/goalobjective length in the composer. (#4337) - frontend: Real-time context window usage is shown as a conversation grows. (#3183)
- frontend: The latest user turn can be edited and rerun in place. (#4377)
- frontend: Replies can be typed and sent while a clarification card is pending. (#4530)
- suggestions: The number of follow-up suggestions is configurable via
suggestions.max_suggestions(default 3). (#4533) - artifacts: Text artifacts can be edited inline in the artifact panel. (#4596)
- artifacts: Markdown artifacts open rendered in a new-window reader (with "View source" and "Download" fallbacks), and all files presented in a run can be downloaded as one zip archive derived from the run's delivery receipt. (#5056, #5117)
- frontend: Browser Live is available in Custom Agent chats. (#4719)
- frontend: A conversation outline navigates long chats: past 5 user turns, a compact side menu lists the conversation's questions and jumps between them, tracking the current section. (#5025)
- frontend: Scheduled tasks can be duplicated into an editable draft that carries over the configuration but not the run history. (#5064)
- threads: Branched conversations get distinguishing titles
(automatic
Title (2),Title (3)sibling numbering) and the recent-chats list shows parent-child lineage with tree connectors. (#4983) - frontend: Chats can be archived and restored: an Archive sidebar action with an Undo toast, Recent chats / Archived tabs above search, and per-chat restore controls; SQL and Memory stores apply the archive filter before pagination while messages, files, links, pin state, and the current URL are preserved. (#5236)
- projects: Project workspaces organize chats (Projects MVP Phase 1): a sidebar Projects section with flat/grouped list modes, a project detail page with a paginated thread list, project-scoped new chats whose threads are pre-created with their project so a run can never land outside it, branch membership inheritance, move-between-projects, and project create/rename/archive/restore/delete. (#5265)
- artifacts: Completed CSV/TSV artifacts preview as bounded tables (up to 200 rows × 50 columns, 50 rows per page, sticky headers, optional first-row header) parsed off the main thread over the existing 1 MiB range loader — literal strings, leading zeros, and multiline quoted cells survive, long cells open in a copyable dialog, and source view remains one toggle away. (#5284)
Observability & tooling
- observability: Trace-id correlation with enhanced logging and agent observability via Monocle. (#3902, #4024)
- tooling: A Hermes-like terminal workbench (
deerflowCLI) backed byDeerFlowClient, plus a redacted community support-bundle generator. (#3760, #3886) - setup: The setup wizard now asks whether OpenAI-compatible gateway models support thinking, and a Volcengine Coding Plan quick-setup path was added. (#3428, #4141)
- tui:
clearcommand. (#4306) - tui: The TUI supports a transparent terminal background. (#4631)
- gateway: New
GET /health/readyreadiness probe runs a bounded databaseSELECT 1and returns 503 while the database is unreachable (200not_configuredfor the memory backend);/healthstays pure liveness, and the production compose healthcheck now gates on readiness. (#5166) - observability: Deferred tool promotions — both routing-hint
auto-promotion and explicit
tool_search— persist as privacy-minimalmiddleware:tool_promotionrun events (tool names, source, count, agent attribution; no queries, schemas, or results), observed only after skill-policy filtering so denied schemas are never reported as effective promotions. (#5183)
Changed
- frontend performance: Keep the public root and localized docs static; lazy-load closed workspace panels and editor/highlighter dependencies; incrementally derive streamed message state; bound streaming Markdown work; virtualize long message and chat lists; pause offscreen decorative effects; and enforce representative route JS/CSS budgets.
- browser: Negotiate binary Browser Live JPEG frames, retain the legacy JSON/base64 protocol for older clients, coalesce presentation to the latest frame per refresh, and revoke replaced object URLs.
- artifacts: Stream regular text artifacts with HTTP byte-range support and limit the initial Web UI preview to 1 MiB until the user explicitly loads the complete file.
- sandbox: The Helm chart now defaults per-sandbox Services to
ClusterIPinstead ofNodePort, so the code-execution sandbox is reachable only inside the cluster via Service DNS (http://sandbox-<id>-svc.<ns>.svc.cluster.local) and is no longer bound on every node's interfaces - including the externally-reachable ones on GKE/EKS/AKS. Existing chart installs flip NodePort -> ClusterIP on upgrade. To preserve the old reachability (an external probe hitting the 30xxx port, or the Docker-Compose/hybrid path where the gateway is not in K8s), setprovisioner.sandboxServiceType: NodePort(withprovisioner.nodeHostif needed). The provisioner itself is unchanged (mode-aware since #4016). (#4190) - skills: An active restrictive skill must explicitly list
taskinallowed-toolsto delegate to a subagent. Read-only discovery infrastructure (tool_searchanddescribe_skill) remains available, but cannot grant schema visibility or execution for a denied business tool. (#4098) - memory: Pre-abstraction top-level
memory.*DeerMem fields (storage_path,max_facts,debounce_seconds,model_name,token_counting,staleness_*,consolidation_*, ...) are auto-migrated intobackend_configon load with a warning, so an upgrade does NOT silently revert customized settings to defaults (model_name->backend_config.model.model). Move them undermemory.backend_configinconfig.yamlto silence the warning. (#4122) - memory: Added
memory.mode(middleware|tool);toolmode registers memory tools (memory_search/add/update/delete) the model calls directly instead of passive per-turn summarization.manager_classresolution is now fail-fast (raisesValueErroron an unknown backend instead of silently falling back). (#4023) - middleware: Declarative layered middleware builder;
ThreadDatanow runs beforeUploads. (#3809) - sandbox: The host->virtual output-masking regex now has a single owner, eliminating duplicated pattern compilation. (#4108)
- docs:
AGENTS.mdis now the source of truth for agent guidance, imported byCLAUDE.mdvia@AGENTS.md; module guides refreshed. (#3770) - memory: The OpenViking memory backend now uses the official OpenViking
adapter; the old trusted-mode
auth_mode/accountfields are rejected in favor of a credential-bound USER API key. (#4707) - gateway: Threads created before the run-event journal have their checkpoint history backfilled as seed events before the first new run, so legacy conversations stay visible and correctly ordered after an upgrade. (#4590)
- agents: Subagent delegation is now routed by net benefit: the lead agent defaults to direct execution unless parallel latency, specialist capability, or context isolation clearly pays off. (#4384)
Fixed
- sandbox: Stop remote
globandgrepfrom reporting "no matches" when their output was cut off. BoxLite, Tenki, E2B, and OpenSandbox cap the search's raw output and then filter it in Python (ignored directories such asnode_modules, the pattern orglobscope), but they reportedtruncatedonly whenmax_resultswas reached. When the capped lines were all filtered out, a search with real matches past the cap came back empty and complete. The search now passes one line beyond its cap so a cut-off result is reported as truncated, and theglobandgreptools say an empty truncated result is incomplete instead of "No matches found". (#5427) - sandbox: Stop host paths reaching the model when output joins them with
:, as$PATHand$PYTHONPATHdo. The matched path ran on through the rest of the list, so every later entry under the same root was left unmasked; extra masking passes recovered one entry each, which hid the leak for short lists. Masking now ends a matched path at:. A symlink inside a mount whose target lies outside every mount is now shown by its mount path instead of the target's host path in command output andglobresults. (#5418) - sandbox: Stop BoxLite
grepfrom ignoring the directory part ofglob. It compared only file names, sosrc/*.jsmatched every.jsfile in the tree. The glob now applies to the path relative to the search root, the same scope asglob()and the other providers. (#5419) - models: Stop every Claude model after the first from losing its
credential when the Claude Code OAuth token is handed off through
CLAUDE_CODE_OAUTH_TOKEN_FILE_DESCRIPTOR. EveryClaudeChatModelinstance loaded credentials again, but a descriptor can be drained only once, so the title, summarization, and subagent models — and every later run — had no credential and failed withTypeError: Could not resolve authentication method. The token is now read once per process and reused. (#5411) - models: Stop the lead agent from failing to build whenever a model with
supports_reasoning_effort: truealso gets areasoning_effortfrom its profile — at the top level, inwhen_thinking_enabledorwhen_thinking_disabled, or from theextra_body.thinkingdisable path. The regular lead-agent build forwards the requested effort even when unset, so the key reached the provider constructor twice and raisedTypeError: got multiple values for keyword argument 'reasoning_effort'. The requested value now layers like per-agentmodel_settings: it replaces a top-level profile value, an unset request keeps that value, and the thinking-mode settings still decide the final one. Codex keeps its own level check. (#5403) - runtime: Stop a keyed run retry from failing with 500 on the SQL run
store. HTTP admissions do not pass a
user_id; the SQL store stamps the request user on the row, but the process-local run record keptNone. A retry with the sameIdempotency-Keythat reached another Gateway worker, or the same worker after the finished run was cleaned up, compared the two owners, took its own run for another user's, and raised. The same mismatch dropped HTTP runs from owner-scoped history reads and skipped the MCPbackground_tasksprojection for them, sovaluesevents for these runs now includebackground_tasks.RunManagernow resolves an omitted owner from the request user the way the SQL store does, so every store records the same owner. (#5401) - runtime: Stop a cross-worker idempotent run reuse from permanently
blocking the thread on the reusing worker. The reuse registered the hydrated
store row as a local run record, but only the owning worker finalizes and
cleans up its records, so the copy kept its admission-time
pending/runningstatus forever: every laterrejectadmission for that thread on the worker returned 409 until a restart, its run reads kept reporting the stale status, and orphan reconciliation skipped the run if the owner crashed. A cancel sent to that worker also took the local-owner path and marked the owner's still-running rowinterrupted. The reusing worker now returns a detached store-only handle instead, so cancel follows the non-owner contract. (#5393) - skills: Stop writing resolved secrets into
extensions_config.jsonwhen a skill is toggled. The Gateway skill toggle andDeerFlowClient.update_skillloaded the file throughExtensionsConfig.from_file(), which replaces every$VARvalue with the environment value, and wrote that model back — so a"$GITHUB_TOKEN"reference was persisted as the plaintext token and an unset variable was permanently replaced with"".DeerFlowClient.update_mcp_configdid the same for every key other thanmcpServers. These writers now edit the raw on-disk JSON and validate the candidate the way the runtime loads it, so placeholders and hand-written structure survive; the MCP router shares the same raw loader. Files rewritten by an earlier toggle keep their plaintext values: restore the$VARreferences and rotate the exposed credentials. (#5357) - gateway: Honor
disable_clarificationandgithub_tokenonly for internally-authenticated callers, the waynon_interactivealready was. Both keys were forwarded frombody.contextregardless of the caller and were not scrubbed from the free-formbody.configthat the run config copies verbatim, so any session or PAT caller could set them.disable_clarificationis the stronger of the two:ClarificationMiddlewareanswers every clarification —risk_confirmationincluded — with "proceed without asking", andSandboxMiddlewarereads it as the same non-interactive signal asnon_interactive.github_tokenreachedruntime.context, where the bash tool exports it asGH_TOKEN/GITHUB_TOKEN, and a copy smuggled throughbody.config['configurable']was persisted in the checkpoint store. The scheduler, IM channels, and the GitHub webhook channel authenticate over the internal request channel and are unaffected. (#5338) - artifacts: Keep
PUT /api/threads/{id}/artifacts/{path}confined to/mnt/user-data/outputs. The outputs-only guard was a string-prefix check on the raw path, so a percent-encoded..(outputs/%2e%2e/uploads/x.txt) — which nginx forwards untouched and Starlette decodes — passed it, and the resolver only confines touser-data/, letting a caller overwrite a sibling upload or workspace file in their own thread. Dot segments are now collapsed before the prefix check, and the resolved host path is re-checked against the resolved outputs root so a symlink planted insideoutputs/cannot redirect the write either. The rule now lives in one shared helper that IM-channel attachment delivery uses as well, so the two copies cannot drift. (#5321) - gateway: Stop persisting a caller-supplied
deerflow_trace_idon the run record.body.metadatareaches both the live run config, which the run worker restamps, and the run record echoed verbatim by the runs API; only the first was covered, so a client could make the most durable surface of a run disagree with theX-Trace-Idand the log lines from the same request. The id is now stamped once at the trust boundary,config.contextis closed off the same way, and a thread's own metadata is no longer seeded with the run-scoped id of whichever run created it. (#5119) - gateway: Expose
X-Trace-IdinAccess-Control-Expose-Headers. It is not CORS-safelisted, so split-origin browser clients — the ones that cannot read the Gateway's logs either — could not read the correlation id they are meant to quote in a bug report. (#5119) - gateway: Keep
X-Trace-Idon unhandled-exception 500s. Starlette'sServerErrorMiddlewareemits those through the raw send outside every user middleware, so the 500 for a server bug — the response most in need of correlation — was the only one shipped without the id.TraceMiddlewarenow sends its own 500 carrying the header before re-raising; the server's exception logging is untouched and mid-stream failures propagate unchanged. This fallback is emitted outsideCORSMiddlewareand stays CORS-opaque, so split-origin browser clients cannot read the id on this one response — same as theServerErrorMiddleware500 it replaces. (#5119) - gateway: Strip a forged
deerflow_trace_idfrom the persisted request echo.body.configis stored verbatim asruns.kwargs_jsonand served back by the runs API, so a forged id inconfig.metadataorconfig.contextsurvived on that one surface while every other carried the real id.redact_config_secretsnow drops the key from both containers, andbuild_run_configmerges run metadata onto a copy so the server-stamped id can no longer be written through into the caller's request body. (#5119) - artifacts: Keep explicit full-file loading scoped to the source thread, so a same-path artifact in another conversation keeps its 1 MiB preview. (#4634)
- sandbox:
SandboxAuditMiddlewareno longer blocks ordinary command substitution that only captures output. The rule now judges position instead of matching any$(:x=$(curl url),echo $(curl url), an argument, and aforword list all run normally, while a substitution in command position ($(curl url), after a|/&&/;, behind leading assignments or anenv/nohup/timestyle wrapper, or as aneval/sourceargument) still blocks because it executes fetched content. An interpreter's code-string flag (bash -c,python -c,perl -e,node -p,php -r, and the<<<here-string) is treated as an execution context wherever it appears, sobash -c "$(curl url)"blocks;source <(curl url)and the backtick spelling ofeval/sourcenow block too, neither of which was detected before. An unquoted newline separates statements like;, soecho hifollowed by a new line starting$(curl url)blocks as well, while heredoc bodies are consumed as data — writing a file whose content happens to start a line with$(curl url)is not a command. Variable expansions whose name merely starts with a risky executable ($shell,$bashrc,$python_version) and lookalike binaries (shellcheck,shasum) are no longer false positives. (#4611, #4623) - mcp: Isolate Settings > Tools enable/disable updates to one MCP server, so
an unrelated disallowed stdio command no longer blocks every switch; allow
disabling a disallowed target while still rejecting its re-enable, preserve
the raw extensions config, honor the MCP-spec
transportalias when enabling SSE/HTTP servers, surface backend validation details in the UI, and atomically replace the shared config for MCP, skill, and embedded-client updates so interrupted writes cannot leave it truncated. (#4574, #4577) - runtime: Thread metadata now switches to
runningonly after the run passes the startup barrier, so pending-cancelled runs no longer briefly projectrunning; clients may observe the prior thread status during worker startup. (#4450) - runtime: Re-check orphan candidates through an atomic, lease-aware takeover claim so a successful heartbeat after the scan keeps the run active and only one reconciler reports recovery. (#4424, #4434)
- skills: Apply
allowed-toolsonly to slash-activated or actually loaded lead-agent skills, preventing passive enabled skills and evaluation fixtures from removing MCP, web, file, and delegation tools from every run. (#4095, #4098, #4192) - models: Honor
api_baseon everyBaseChatOpenAIsubclass (VllmChatModel,MindIEChatModel,PatchedChatMiMo,PatchedChatStepFun,PatchedChatMiniMax), not justChatOpenAI/PatchedChatOpenAI. Those five previously dropped the configured endpoint silently and then failed every request with an opaqueunexpected keyword argument 'api_base'; the unknown-config-key warning was disabled for them as well. Both now gate onissubclass(BaseChatOpenAI). (#4146) - agents: Coalesce
SystemMessages before the LLM request; ensure a visible response after tool runs; avoid a default LLM title call before stream end; reserve ellipsis room so the local title respectsmax_chars; and snap the tool-output tail forward so fallback truncation respectsmax_chars. (#3711, #4033, #3885, #4052, #4017) - agents: Skip dateless reminders in the dynamic-context date scan; load
SOUL.mdfrom agent dirs withoutconfig.yaml; requireconfig.yamlinupdate_agent's legacy-agent guard; and refuse emptySOUL.mdupdates. (#3685, #4136, #4166, #4219) - middleware: Window the loop-detection tool-frequency counter so long runs
no longer false-trip; prevent the title middleware from streaming tokens;
fix positional fallback consuming an unrelated todo when the same-content list
is exhausted; acquire the token-budget lock across
_apply,before_agent,_clear_run_state, and_drain_pending_warnings; drop orphanToolMessages so strict providers don't 400; sanitize invalid tool-call arguments; and recover from empty tool-call names and malformed tool-call ids in dangling repair. (#4072, #3566, #3709, #3714, #4080, #4193, #4008, #4246) - subagents: Inherit
LoopDetectionMiddlewareand summarization middleware so tool loops break and steps are captured; surface the turn-budget cap asMAX_TURNS_REACHEDwith a partial result; unify guardrail caps on the additivestop_reason+token_budget; inject durable context before compaction; preserve the parent checkpoint namespace; prohibit thetasktool in the general-purpose system prompt; re-buffer subagent events on flush failure to avoid losing steps; and fix the lostloop_cappedstop reason when a subagent'srun_idisNone. (#3931, #4009, #3949, #3980, #4040, #4215, #4161, #4082, #4059) - memory: Harden against null/empty edge cases - skip whitespace-only facts;
coerce null
confidence/source.confidencein updates, searches, and the three remaining raw reads; treat explicitnullbackend_configvalues as omitted; fixKeyError/UnboundLocalErrorwhen a fact has no id or the facts list is empty; stop the busy-spin in the debounced update queue; and flush the memory queue on graceful shutdown to prevent loss. (#3719, #4074, #4076, #4034, #4217, #3993, #3992, #4073, #4181) - runs: Close multi-worker ownership gaps in run atomicity; fail-stop local
execution when lease renewal cannot be confirmed before its deadline and
fence late completion writes after peer takeover; degrade cancel to lease
takeover for multi-worker; keep
create_threadidempotent when the insert loses a race; readstop_reasonfrom runtime context; and persist run duration in checkpoints for history reads. (#4003, #4064, #4414, #3800, #4188, #4118, #4431) - runtime: Serialize SQLite event-store writes to prevent per-thread
sequence collisions; skip hidden human messages in the journal; and drop the
silent delta-discard in
_merge_stream_text. (#4077, #3698, #4085) - gateway: Attach thread-message feedback by real
event_type; offload blocking filesystem IO in artifact serving, gateway uploads, and the Discord channel; limit the uploaded-file context manifest; and live-tail malformed Redis reconnect ids. (#3651, #3551, #3935, #3927, #3917, #4012) - uploads: Claim the converted-Markdown companion filename before writing
it, so two convertible uploads sharing a stem (or a convertible plus a
same-stem
.mdupload) no longer silently clobber each other within one request. Whenuploads.auto_convert_documentsis on, the companion.mdnow gets a unique name (e.g.a_1.md);POST /threads/{id}/uploadsandDeerFlowClient.upload_filesboth report the actual name inmarkdown_file. (#4288) - config: Coerce null object config sections to their defaults; honor the
unified database configuration in the store and sync checkpointer; and have
legacy DB backfill create missing
Indexobjects on existing tables. (#3573, #3904, #3994, #4090) - models: Apply the
stream_chunk_timeoutdefault to allBaseChatOpenAIsubclasses; and normalizeapi_base->base_urlforChatOpenAIwith a warning on unknown config keys. (#4102, #3790) - mcp: Isolate tool-discovery failures per server; synchronize the session-pool singleton lifecycle; invalidate the tools cache on config content
- skills: Activate a slash skill once per run, not per model call; close the
skill-install security-scan coverage gap; recognize fully deleted skill
packages in review CI and remaining
requests/httpxmethods as network sinks in SkillScan; reuse the resolved app config in the no-arg skills prompt section; and reload mounted skills without restarting the Gateway. (#4103, #3924, #4169, #4130, #4160, #4264) - sandbox: Guard the reverse path-translation and output-masking regexes
with segment boundaries; handle one-sided line ranges and empty files in
read_file/str_replace; align the AIO bash working directory; useos.sepin the reverse-resolve containment check on Windows; normalize Windows backslash paths in bash commands; stopglob/grep/lsfrom surfacing disabled skills' files; and allow valid heredoc commands in the sandbox audit. (#4035, #4053, #4078, #4079, #4051, #4058, #3869, #4096, #3786) - sandbox: Synchronize the sandbox provider singleton lifecycle (with concurrency regression tests) and keep k8s calls off the event loop in the provisioner. (#3730, #3941)
- sandbox: Align sandbox artifact mounts with the channel user; fix
local-dev (
make dev) on non-root / NFS hosts; reap macOS nginx processes on stop; and fix production Postgres UV-extras detection in Docker. (#3729, #3590, #3828, #3897) - channels: Validate the channel provider before resolving its config;
dedupe GitHub webhook redeliveries and drop redundant GitHub review-comment
webhook fan-out; scope the slash-skill whitelist check to the run's owner;
batch Feishu file messages into one thread and dispatch Feishu group commands
prefixed with a bot @mention; accept leading @mentions before
/connectbind codes and don't treat a bare "connect" as a bind command; stop Feishu from creating thread topics and throttle card updates; let the UI runtime channel config win overconfig.yaml; fixrequire_mentiongating on whitespace-onlybot_login/mention_login; guard null quote fields in WeCom; and key inbound dedupe on chat-scoped workspaces so Telegram, Feishu, WeChat and DingTalk redeliveries stop re-running the agent on a default (unbound) configuration, releasing the dedupe key on transient failures so a redelivery can still recover. (#4100, #4104, #4131, #4129, #3753, #4229, #4222, #4251, #3810, #3674, #4055, #4069, #4287) - frontend: Preserve messages and durable context across summarization;
preserve artifacts and stabilize artifact paths during streaming; resolve
relative artifact image paths; retain presented artifacts in the header
dropdown; keep orphan tool messages visible; show assistant text during tool
steps; reset new chat on client-side navigation; prevent stream cancellation
on concurrent submit; fix stale-run reconnect and cancel handling; fix chat
math rendering, single-tilde markdown, double reasoning rendering, UTF-16
markdown binary classification, and
<memory>tags in Streamdown; make recent-chat rows fully clickable; validate attachment limits before upload and fix uploaded-file metadata in message copy; fix mobile workspace and accessibility blockers, the card tool-message bug, and side-chat toolbar / panel-button behavior; block unresolved suggestion-template placeholders; refresh notification permissions; show the branch action only for completed turns; enable regenerate in custom agent chats; and generate a fallback title for interrupted first-turn runs. (#3826, #3791, #4094, #4038, #3854, #3880, #4114, #3673, #3878, #3908, #3557, #4245, #3870, #3966, #4209, #3733, #3900, #3944, #3740, #3976, #3959, #3961, #3764, #3768, #4147, #3967, #3874, #3644) - tui: Interrupt an active run before
/quitexits. (#4235) - harness: Don't flag the outline as truncated at exactly
MAX_OUTLINE_ENTRIESheadings. (#3856) - tracing: Attach Langfuse trace metadata to the goal evaluator. (#4202)
- context: Resolve the context-compress bug. (#4065)
- threaddata: Fix
AttributeErrorwhenruntime.contextisNone. (#3989) - goal: Stop
continuation_countdouble-bump during stand-down. (#4199) - circuit-breaker: Stop wedging after a non-retriable half-open probe. (#3991)
- github: Match
allow_authorslogins case-insensitively. (#4218) - community:
image_searchnow returns the full-resolution image URL. (#3990) - skills: Offload blocking filesystem IO in the skill-history endpoint. (#3563)
- skills: Don't treat a lazily evaluated PEP 695 type alias as a network sink in SkillScan. (#4315)
- tracing: Resolve the Langfuse trace user from runtime context. (#3794)
- guardrails: Propagate internal owner attribution into the guardrail context. (#3839)
- subagents: Clamp the subagent limit consistently with
MIN_SUBAGENT_LIMIT. (#4081) - subagents: Load user-scoped skills. (#4356)
- mcp: Per-server fail-soft OAuth priming, and persist rotated refresh tokens. (#4084)
- mcp: Ignore malformed path-like text. (#4456)
- auth: Resolve email accounts case-insensitively. (#4101)
- auth: Recover from setup-status timeouts. (#4371)
- scheduler: Close a dispatch race that could launch two runs for one scheduled task. (#4105)
- channels: Buffer and drain GitHub comments queued during a busy run. (#4133)
- channels: Escape Slack reserved characters before mrkdwn conversion. (#4197)
- channels: Check
response.success()on Feishu card/reaction SDK calls. (#4234) - channels: Drop inbound DingTalk messages that carry no conversation identity. (#4316)
- channels: Receive inbound Telegram attachments. (#4392)
- memory: Consolidated facts inherit
expected_valid_daysfrom their sources. (#4225) - config: Sync
_memory_configwith AppConfig auto-reload. (#4208) - postgres: Harden the async engine with
pool_recycleandcommand_timeoutto stop stale-connection 504s. (#4230) - harness: Add a timeout to
invoke_acp_agentto prevent indefinite hangs. (#4238) - community: Surface the target-page error status in
web_fetch(Browserless). (#4239) - sandbox: Widen the BoxLite/AIO tenant hash and verify identity on reclaim. (#4171)
- sandbox: Make an empty
old_stra no-op instr_replaceon any file. (#4256) - sandbox: Serialize E2B release transitions. (#4355)
- sandbox: Bound E2B output-synchronization resources. (#4364)
- sandbox: Unwrap
Overwrite-wrapped sandbox state inafter_agent. (#4381) - sandbox: Bypass proxies for local AIO traffic. (#4444)
- models: Surface length-capped model responses instead of dropping them. (#4309)
- streaming: Keep large file generation responsive. (#4354)
- streaming: Expose custom events to
astream_events. (#4403) - streaming: Signal replay history gaps. (#4426)
- summarization: Summarize with the run model and fall back on summary-provider failure. (#4361)
- runtime: Remove transient image context after model calls. (#4267)
- runtime: Stop subgraph stream frames from impersonating root frames. (#4407)
- runtime: Reject unsupported run options and stream modes. (#4430)
- runtime: Serialize checkpoint writes with active runs, linearize
delta-mode checkpoint resume, and accept the SDK's default
stream_resumable=falseto avoid resume races. (#4437, #4460, #4468) - checkpoint: Unwrap
Overwritefirst writes into empty channels. (#4383) - nginx: Allow long chat prompts through
/api/langgraph/without a raw 500. (#4277) - gateway: Prefer
X-Trace-Idovermetadata.deerflow_trace_idwhen the header is set. (#4283) - gateway: Seed branch run-events so inherited history survives forking. (#4385)
- gateway: Scope branch-history seed run ids per inherited turn. (#4459)
- frontend: Harden artifact and markdown rendering. (#4117)
- frontend: Classify a symlink replacing a file distinctly from deleted in workspace-change review. (#4170)
- frontend: Offload blocking filesystem IO in the workspace-change text-cache lifecycle. (#4268)
- frontend: Encode artifact URL path segments. (#4278)
- frontend: Clarify run-duration display. (#4348)
- frontend: Preserve regenerate state in branched threads. (#4358)
- frontend: Default the reasoning-effort label to Medium when unset. (#4373)
- frontend: Strip and parse the
<current_uploads>upload-context tag. (#4402) - frontend: Keep leading orphan tool messages visible. (#4408)
- frontend: Keep completed subtask cards stable after reload. (#4432)
- frontend: Apply message-image
maxWidthvia inline style. (#4446) - frontend: Restore resizing for the artifacts and sidecar panels. (#4469)
- frontend: Allow dev-server access from non-localhost hosts. (#4471)
- safety: Backfill empty content-filter responses so they don't poison the thread. (#4394)
- tools: Exclude injected runtime from the
list_uploaded_filesschema. (#4376) - mcp: Bound MCP server bring-up — tool discovery (subprocess spawn +
initialize+tools/list) and persistent stdio session initialization — with a new per-serversession_init_timeout(default 60s,nulldisables), so a hung stdio server can no longer block agent construction, or the whole Gateway event loop, indefinitely.tool_call_timeoutstill bounds individual stdio tool calls. (#4657) - runtime: Tool-output budget externalization no longer trips run delivery
verification. The default
.tool-resultsstorage dir (and any customtool_output.storage_subdir) is excluded from workspace-change snapshots and produced-artifact detection, so a run that only externalized oversized tool outputs succeeds instead of failing as an error. (#4657) - frontend: Hide stale follow-up suggestion chips while a turn is still streaming. (#3396)
- frontend: Fix streaming render glitches: stop the word animation from replaying, keep step text stable, preserve message order during long runs, and keep reasoning above the answer. (#4266, #4510, #4513, #4578)
- frontend: Encode thread IDs in chat routes so IDs with special characters no longer break navigation. (#4302)
- frontend: Render citation links from React children. (#4486)
- frontend: Localize conversation export failure messages. (#4493)
- frontend: Sync side panel state when a drag collapses the panel. (#4556)
- frontend: Render one workspace-change card per run instead of duplicates. (#4559)
- frontend: Refresh the active artifact's content when it changes. (#4584)
- gateway: Reject non-positive read limits in API requests. (#4284)
- gateway: Handle a null
config.configurablewhen resolving the thread id instead of failing. (#4301) - gateway: Unify thread id validation across API routes. (#4589)
- gateway: Merge concurrent thread metadata updates instead of letting them silently overwrite each other's changes. (#4489)
- gateway: Expose the run metadata response header to cross-origin clients, so a split-origin frontend learns new run ids instead of staying stuck on the new-thread placeholder route until reload. (#4535)
- gateway: Replay edit and rerun from a settled checkpoint so the edited prompt actually runs (previously a first turn's edit replayed the original prompt and vanished after reload), and keep a manual rename through the rerun. (#4534, #4539)
- runtime: Cancel a run from any live gateway worker, not only the one that owns it, so the stop button no longer depends on request routing. (#4500)
- runtime: Close a replacement run when interrupt or rollback admission is cancelled mid-flight, instead of stranding an unseen active run on the thread. (#4472)
- runtime: Regenerating a response now preserves the thread's current title and supports the latest interrupted response whose partial message never reached a checkpoint. (#4480, #4524)
- agents: Classify web_fetch error pages such as 404s as errors rather than successful evidence, so retries and stagnation guards can react. (#4314)
- agents: Handle XML-to-dict option shapes when normalizing clarification choices. (#4527)
- subagents: Run delegated subagents with isolated callbacks and lazy
skill activation, fixing cross-event-loop failures and passive skills
stripping baseline tools like
write_file. (#4497) - sandbox: Handle overwrite-wrapped state when ensuring the sandbox is initialized. (#4429)
- sandbox: Reconcile E2B sandboxes safely: pick the first healthy candidate, adopt the canonical instance per user and thread, defer a peer's live duplicates, and reap orphans after a grace window. (#4443)
- sandbox: Claim ownership before destroying a sandbox that failed its readiness check, so a peer gateway can no longer adopt the not-yet-ready sandbox and kill a live turn. (#4505)
- sandbox: Allow grep to search a single file. (#4512)
- sandbox: Enforce the E2B capacity limit deployment-wide when sandbox ownership uses Redis, so multiple gateways cannot create past it. (#4575)
- skills: Activate managed integration skills from the managed integrations root on slash invocation. (#4570)
- skills: Offload blocking filesystem IO when updating a skill and serialize concurrent writes. (#3565)
- mcp: Ignore oversized path-like text. (#4582)
- memory: Harden long-term memory: reject duplicate facts inside the create critical section, truncate injected mem0 context on entry boundaries, and keep task-scoped instructions such as "inspect only" out of long-term memory. (#4599, #4600, #4604)
- scheduler: Keep a successfully launched scheduled run's slot and run id when post-launch bookkeeping fails, preventing a later dispatch from launching a duplicate run. (#4504)
- config: Treat a deleted extensions config file as absent instead of raising, so tool and skill config resolution keeps working. (#4275)
- config: Normalize the
postgres://short scheme for the async ORM engine. (#4293) - console: Disable cost reporting when model pricing mixes currencies instead of reporting a meaningless cross-currency total. (#4564)
- browserless: Accept the
timeoutconfig key and harden its coercion. (#4519) - docker: Send
Connection: upgradeonly when the browser requests it, fixing login-page refresh loops when the Docker dev stack is accessed via a remote host. (#4250) - runtime: Group JSONL batch event writes by run, so a batch covering several runs no longer lands all events in the first run's file and makes later runs unreadable through per-run APIs. (#4938)
- runtime: Restore standalone LangGraph Studio compatibility: the graph
entrypoint and file-based app load again, the Studio identity can discover
system assistants, and the documented
langgraph devworkflow works. (#4760, #4838) - gateway: Stamp
turn_durationon a run's last AI message only in/messages/page, so multi-step turns no longer repeat the same run lifetime as thinking latency on every intermediate message. (#4755) - gateway: Preserve exact history attribution beyond the event page limit, so older AI messages on long-lived threads are no longer credited to a later turn's run and duration. (#4953)
- gateway: Reject MCP task cancellation with HTTP 503 when the task worker is stopped, instead of acknowledging a cancellation that would never run. (#4963)
- middleware: Correct four context-handling defects: fallback dynamic-context injection targets the latest user message instead of resurrecting an old prompt as the current turn; bare string blocks in list-form user content are sanitized like all other user text; duplicate placeholders are no longer emitted for the same invalid tool call; and summarization no longer compresses away the current request's user message while leaving the previous turn's behind. (#4667, #4668, #4693, #4882)
- middleware: Restore the system-prompt injection that teaches the model
about the
write_todostool, which the todo middleware's model-call override had silently dropped. (#4735) - agents: Make SQL agent-store signatures content-sensitive, so an agent update that reuses its previous timestamp no longer leaves the GitHub agent registry serving stale webhook routing. (#4709)
- tools: Resolve presented files with the runtime user, so
present_filesno longer rejects valid artifacts as outside the outputs directory when the request user context is unavailable. (#4677) - tools: Retain a strong reference to deferred subagent cleanup tasks, so garbage collection can no longer destroy a pending cleanup and leak cancelled subagent records, locks, and memory. (#4928)
- subagents: Give every background subagent run a server-side execution ID, so concurrent runs that reuse a provider tool-call ID can no longer overwrite, poll, or cancel each other's state. (#4758)
- harness: Offload ACP workspace creation and MCP config loading from the event loop, so invoking an ACP agent no longer raises blocking-IO errors or stalls other async work. (#4965)
- mcp: Reject non-finite
poll_after_secondsvalues on task snapshots when they arrive, so a bad polling interval no longer crashes scheduling and persistence after a successful poll. (#4750) - mcp: Keep the configured
grant_typeauthoritative overextra_token_paramsduring OAuth token exchange, so extra parameters can no longer silently switch the configured flow and be rejected by the token endpoint. (#4860) - mcp: Exclude the internal stdio MCP temp directory (
.mcp/tmp) from workspace changes, so MCP temporary and debug files no longer appear alongside user deliverables or crowd real changes out of the file budget. (#4898) - mcp: Cancel the remote task when a durable task submission is cancelled mid-flight, so an interrupted submission no longer leaves a remote task running with no record to poll or stop. (#4933)
- sandbox: Accept the documented E2B reconciliation config fields, so valid E2B configuration no longer produces misleading startup warnings. (#4772)
- sandbox: Bound E2B mount upload resource use per file, per mount, and across the whole upload pass (shared size and file budgets plus a wall-clock deadline), so large mounts can no longer spike Gateway memory or hold sandbox capacity indefinitely. (#4812, #4842)
- sandbox: Preserve trailing whitespace in E2B-synced filenames and tolerate out-of-range remote mtimes, so output sync no longer re-downloads files repeatedly or aborts mid-sync. (#4861)
- sandbox: Reject non-finite Redis lease-timing values in sandbox
ownership config at parse time instead of crashing with an
OverflowErrorduring startup. (#4960) - sandbox: Resolve structured skill reads through the sandbox provider's
path mappings, so
read_fileopens legacy and per-user custom skills under the same enabled-state projection aslsand shell execution. (#4792) - skills: Parse Responses API content blocks in the moderation scanner, so valid skill-management decisions returned as content blocks are no longer rejected as unparseable. (#4936)
- memory: Reject non-positive and non-finite timeout and character-limit settings in the Honcho and Mem0 backends at config parse time, so a bad value fails fast instead of silently truncating stored text or crashing on the first HTTP call. (#4783, #4823)
- memory: Scope custom-agent bootstrap facts to the selected agent's bucket, so facts learned during setup no longer leak into the default bucket and influence ordinary lead-agent conversations. (#4804)
- artifacts: Support atomic saves on Windows, and serve a SHA-256 ETag on
artifact reads so inline preview and editing work on non-secure contexts
such as plain-HTTP LAN origins where
crypto.subtleis unavailable. (#4629, #4865) - frontend: Keep conversation order stable around long runs: the submitted user message no longer renders twice or sinks below its own processing steps, and after a mid-run page reload a turn's steps can no longer appear above the user message that started the run. (#4620, #4660, #4834)
- frontend: Stop matching
<header>as<head>when injecting the base href into HTML artifact previews, so relative assets in report fragments that begin with<header>now load in the sandboxed preview iframe. (#4625) - frontend: Open landing-page case studies on a public read-only
/showcase/route so anonymous visitors are no longer redirected to login. (#4635) - frontend: Sort the chats page by pinned state, so pinned threads no longer render below unpinned ones. (#4643)
- frontend: Keep
<think>pairs written inside markdown inline code in the rendered content instead of hollowing them out into the Reasoning panel, and restore the copy button for turns that contain only reasoning. (#4647) - frontend: Surface model-loading failures with a workspace error banner and retry action instead of a silently empty model list. (#4840, #5021)
- frontend: Preserve copy and other actions on completed assistant messages while a later turn is still streaming. (#4844)
- frontend: Keep the browser live stream connected after a successful reconnect instead of tearing down the new socket and immediately creating another. (#4951)
- frontend: Reuse the shared clipboard fallback when copying the Lark authorization link, so the copy action works in browsers without the Clipboard API. (#4767)
- frontend: Use consistent "DeerFlow" casing in the composer disclaimer and fix the "What's New" heading on the landing page. (#4970)
- channels: Bound inbound intake with a fixed worker pool and bounded admission queues, and await real cross-thread tasks on shutdown, so message floods are rejected promptly instead of accumulating and channel shutdown no longer tears down transports with work still in flight. (#4800, #4816)
- channels: Offload outbound attachment file IO for Feishu, Telegram, and WeCom to worker threads, so sending a large artifact no longer stalls the Gateway event loop. (#4633)
- channels: Run Telegram connection-identity lookups on the Gateway event loop, so inbound messages and commands no longer crash with a cross-loop error when channel connections are enabled. (#4815)
- feishu: Keep file receiving off the event loop and preserve every inbound attachment: duplicate provider filenames no longer overwrite each other, writes can no longer be redirected outside the thread bucket, and a failed attachment no longer blocks the rest of the message. (#4627, #4903)
- dingtalk: Strip leading
@botmentions before command classification, so slash commands like/newsent in group chats are recognized instead of treated as plain chat. (#4724) - discord: Refuse to start typing-indicator loops after the channel stops, so shutdown no longer leaves an infinite typing task sending events in the background. (#4752)
- wecom: Serialize WebSocket start/stop transitions and await the SDK's real receive-task shutdown, so stopping the WeCom channel can no longer return before the socket closes or clear a newer connection's state. (#4762)
- buzz: Drop replayed events across reconnects using a persistent seen-id store, so the agent no longer re-answers the last message in a channel after a relay or Gateway restart. (#4888)
- lark: Keep the CLI lock directory writable inside sandboxes while the credential-bearing config root stays read-only, restoring Lark API commands that previously failed with a read-only filesystem error. (#4701)
- scheduler: Enforce the global
max_concurrent_runsbudget for manual triggers too, returning HTTP 409 when the cap is reached instead of letting manual launches exceed it. (#4769) - scheduler: Coerce serialized task timestamps on read, so scheduled-task operations no longer fail when string-form timestamp values reach the database layer. (#4785)
- scheduler: Support safe multi-instance scheduler recovery: startup no
longer treats live runs owned by peer Gateway instances as local
leftovers, so a restarting instance cannot interrupt a live run or trigger
a duplicate execution; multi-instance mode is opt-in via
scheduler.multi_instance. (#4713) - scheduler: Enqueue busy scheduled task runs instead of skipping them:
occurrences that hit a busy reused thread now wait in a durable queue
(bounded by
scheduler.queue_timeout_seconds) and survive Gateway restarts, and the UI explains the queueing behavior. (#4918) - cli: Add
--recursion-limitto headless--print,--json, and--cliruns, so long-running agent loops are no longer stuck at the default recursion limit of 100. (#4615) - dev: Exclude backend runtime state from the Uvicorn reload watcher in
the backend
make devlauncher, so an agent task writing files under the runtime tree can no longer restart the Gateway and reset concurrent users' requests. (#4759) - dev: Resolve diagnostic script paths from the script's own location, so root diagnostic commands work when invoked from any working directory. (#4736)
- docker: Harden local and container startup:
make upwaits for the Gateway health probe before declaring the stack ready, Docker startup no longer aborts when.envis missing, the Gateway can writeextensions_config.jsonin production, runtime data stays out of the image build context, log commands resolve the checkout root correctly, and the default loopback origins are allowed so the dev setup page can hydrate. (#4658, #4806, #4852, #4853, #4956, #4959) - gateway: Stamp the server-authoritative feed position onto persisted messages, so an early user message no longer vanishes or jumps into the middle of the step stream once history exceeds one page and context compaction has fired. (#4696)
- lark: Preserve the new app secret during managed credential switches by
clearing the previous app's OAuth data before the replacement is written,
so the subsequent browser authorization no longer resolves an empty
client_secret. (#4820) - messages: Drop legacy
<uploaded_files>tag handling: the backend treats the pre-#4174 spelling as ordinary content and strips only<current_uploads>, while the frontend keeps stripping the legacy tag so old threads still render cleanly. (#4826) - skills: Reject a blank
SKILL.mddescription at the write gate, matching what the loader already requires, so editing a custom skill with an empty description no longer writes a file the loader then rejects - which destroyed the skill on disk. (#4867) - sandbox: Make the model-facing
descriptionargument optional (empty by default) acrossbash,ls,glob,grep,read_file,write_file,str_replace, andtask, so providers that omit it are no longer rejected before execution. (#4878) - sandbox: Bound Windows command execution: host commands run in a new
process group killed via
taskkill /T /Fon timeout so a descendant cannot hold the call open, and output flows through the existing bounded 10 MiB capture. (#4946) - sandbox: Scope the Windows MSYS path-conversion exclusion to safe virtual path prefixes instead of disabling conversion globally, so host-native CLI launchers that need normal conversion work again. (#5003)
- skills: Rebuild per-user skill storage after an app-config hot reload, so it no longer stays bound to paths from the previous config instance. (#4972)
- skills: Tokenize portable
allowed-toolsscalars with parenthesis awareness, soBash(tvly *)-style entries stay intact, unmatched parentheses are rejected instead of silently fragmenting, and argument-scoped entries remain literal rather than broadening access. (#4984) - agents: Normalize
ToolMessages returned insideCommandresults, so error payloads no longer earn a default success receipt and tool-progress tracking sees them. (#4977) - mcp: Tear down the in-flight session owner when
get_sessionis cancelled during eviction, so a cancelled caller no longer leaks the owner task or parks past its timeout. (#5008) - mcp: Reconnect ordinary stdio tools after a transport disconnect: the failed pooled session is evicted (only if still registered), the original error surfaces without automatic replay, and a later retry starts a fresh subprocess. (#5018)
- mcp: Preserve pooled stdio sessions after protocol timeouts during
durable MCP task polling - a 408 is not a disconnect - so task state
survives and the next poll no longer reports
task_not_found. (#5027) - mcp: Reject credentials that cannot travel as HTTP header values (trailing newline or whitespace, non-ASCII) at the config boundary, so the transport's exception - which echoes the full value - can no longer leak a secret into model context, checkpoints, and traces. (#5066)
- subagents: Clean up the background-task entry when the poller exits unexpectedly and drop a PENDING registry entry when submission fails, so a failed or crashed poll no longer leaks the entry or leaves the subagent running unattended. (#5069)
- subagents: Stop the zombie PENDING registry entry on the submit-failure path, and derive the capacity snapshot's queued count from the waiters' length instead of iterating a deque other threads mutate. (#5086)
- channels: Synchronize
ChannelStorereads with mutations, soget_thread_id()/list_entries()can no longer raisedictionary changed size during iteration. (#5083) - discord: Retain strong references to ack-reaction tasks and drain them on shutdown, so a GC pass can no longer silently drop a reaction or pin the channel across restart cycles. (#5049)
- buzz: Move seen-event persistence off the event loop with coalesced atomic writes, preserving dirty generations when events arrive mid-write and awaiting the final flush on shutdown. (#5103)
- streaming: Stop an
IndexErrorinMemoryStreamBridge._make_gapwhen a subscriber reconnects to an empty or drained stream with an expired cursor. (#5047) - uploads: Keep deduplicated filenames within the 255-byte limit by truncating the stem on a UTF-8 code-point boundary, so two max-length files that differ only by a dedupe suffix upload successfully instead of failing the whole batch. (#5059)
- frontend: Format structured upload error details (FastAPI validation
issues, objects, arrays) instead of showing
[object Object]. (#5071) - frontend: Keep a renamed thread's title in sync across the active chat header, document title, search results, and metadata caches without a reload. (#5045)
- frontend: Truncate selected model names to the selector button width, so long model names ellipsize in the composer and sidecar instead of overflowing. (#5050)
- frontend: Truncate long subtask card titles to one line with a tooltip,
so a delegation whose model omitted
description(falling back to the full prompt) no longer overflows the chat layout. (#5136) - dev: Default the frontend dev server to Webpack on all platforms
(
DEER_FLOW_DEV_BUNDLER=turboopts back into Turbopack), avoiding Turbopack's macOS PostCSS worker leak and its Windows runtime panics. (#5036, #5133) - scripts: Run repo shell scripts through an explicit interpreter
(
bash scripts/...), so a lost executable bit - zip/tarball downloads,core.fileMode=false, non-POSIX filesystems - no longer breaksmake docker-startand friends withPermission denied. (#5031) - deps: Depend on the renamed
tenkipackage instead of the PyPI-removedtenki-sandbox(sametenki_sandboximport), so clean checkouts can resolve dependencies again onmake dev/uv sync. (#5087) - memory: Memory reads configured to stop the turn now raise a
backend-neutral
MemoryReadErrorthat prompt assembly preserves instead of swallowing: strict OpenViking (read: raise), Mem0, and Honcho reads propagate, OpenViking scope-resolution failures follow the configured read policy, and the 5-second injection deadline honors the same policy (fail-open continues without the context; strict raises with the timeout as its cause). (#4726) - memory: Buffered memory extraction is cancelled when a custom agent is deleted or cleared, so pending debounce timers can no longer resurrect the deleted per-agent memory scope or overwrite a fresh clear with a stale pending update. (#5123)
- agents: Conversation titles are generated from the user's original
message content when upload (or other) context wrappers have been injected
into the text, so titles no longer quote server-injected
<current_uploads>context; attachment-only messages keep theNew Conversationfallback. (#4729) - agents: Fraction summarization triggers resolve against the model's
declared
context_window(now translated into the LangChain profile), and an unresolvable fraction clause degrades to a never-firing trigger with a warning instead of crashing the whole agent build; percent-style and non-finite trigger values are rejected at config load. (#4901) - agents: Custom-agent storage falls back to file-backed storage only when search mode cannot resolve the main application config — invalid configuration and missing config paths surface instead of silently switching storage backends — and store IO moves off the event loop in async agent routes. (#4952)
- middleware: Loop-detection hard stops win across a whole tool-call batch: selecting a soft warning no longer ends inspection, so a later call in the same response that crosses an operator-configured hard limit is rejected instead of riding along with the earlier warning. (#5245)
- sandbox:
list_dirandglobin the five remote sandbox providers (E2B, OpenSandbox, AIO, Tenki, BoxLite) return filenames verbatim instead of stripping whitespace, so files whose names begin or end with spaces are no longer listed under paths that then do not exist. (#4980) - sandbox: Concurrent subagents sharing one thread sandbox run under process-local execution leases with task-scoped AIO shell sessions, so one sibling finishing can no longer release the shared sandbox underneath the others or corrupt the implicit persistent session; a healthy replacement session is promoted after corruption instead of returning to it. (#5134)
- sandbox: The Bash tool guides the agent to detect its execution
environment with evidence (
uname -s,sw_vers,uname -a) instead of model assumptions, and rejected host paths direct it to command-only probes or allowed virtual paths rather than repeating the blocked command. (#5111) - sandbox: The Docker AIO compatibility capability allowlist gains
FOWNER, so AIO images whose startupchmods/run/user/1000(e.g. 1.11.0) start again under the hardened default capabilities whileno-new-privilegesstays on. (#5163) - sandbox: File appends no longer destroy existing content when the pre-read fails: E2B append treats only missing-file errors as an empty file and re-raises anything else instead of overwriting the file with just the appended tail, and AIO appends use the server's native append mode without a pre-read at all. (#5261, #5278)
- skills: Skill markdown is read explicitly as UTF-8, so localized skills
validate on Windows hosts whose default code page is not UTF-8 instead of
failing with
UnicodeDecodeError. (#4995) - mcp: The MCP tools cache re-initializes after a runtime config change:
a module-level
asyncio.Lockbound to a closed event loop and an unsynchronized initialized flag had left every post-update call failing withLock is bound to a different event loopor racing across worker threads. (#5062) - mcp: Sync-wrapped MCP tools keep LangGraph
ToolRuntimeinjection (the annotation-less sync wrapper is nowfunctools.wraps-transparent), so per-user scope resolution and durable task submission no longer run withruntime=None— which had routed completion-notification runs under the default lead agent instead of the thread's custom agent. (#5164) - auth: A duplicate OAuth identity no longer reports "Email already
registered" — the two integrity violations are distinguished — and the
partial OAuth-identity index now declares
postgresql_whereso Postgres builds the intended partial index instead of a full one. (#5026) - frontend: The mobile sidebar trigger stays clickable on regular and
custom-agent welcome pages, where multi-line (longer localized) welcome
text in a same-
z-indexoverlay could cover the header's tappable area. (#5149) - subagents:
SubagentResultlifecycle timestamps are UTC-aware on every writer, matching the repo-wide convention instead of stamping local wall-clock time on non-UTC hosts. (#5153) - browser: Background live-frame scheduler tasks retain strong
references, so garbage collection can no longer silently stop the Browser
Live view from refreshing by stranding a pending guard that was cleared
only in a collected task's
finally. (#5155) - runtime: Duplicate
on_llm_endcallbacks for the same LangChain run id persist one durablellm.ai.responseevent (replayed usage merged by generation position, first callback canonical), so append-only message APIs no longer return duplicate responses from providers that re-fire the callback with usage populated. (#5187) - runtime: Terminal finalization completes after a cancellation raised inside the completion hook or task-stop fan-out, so extension observers run and the stream END marker is published before the interruption is re-raised; the finalization tail stays interruptible. (#5191)
- models: A cancelled LLM call releases its owned circuit-breaker
recovery probe across provider execution, concurrency admission, and
backoff, so subsequent calls no longer see
CircuitBreakerOpenafter a cancellation; probe ownership is fenced by a per-call token so cancelling an older call cannot release another call's probe. (#5197) - runtime: The embedded
DeerFlowClientkeys its graph cache by effective user in every authorization mode and materializes the same user in runtime context, so sequential reuse for different users can no longer serve a graph assembled with another user's prompt and workspace state. (#5206) - runtime: Cancelled workspace-change snapshot captures drain their
already-running scan and clean up the per-run text cache instead of
leaking
deerflow-workspace-changes-*directories, while metadata-only captures propagate cancellation promptly without waiting on the scan. (#5232, #5234) - persistence: Gateway startup tolerates a database already migrated to
the explicitly-reviewed newer revision (
0019_thread_incarnations), so rolling back to this image after a newer deployment stays possible; other unknown revisions, an empty version table, and multiple version rows still fail closed. (#5219) - community: Tavily Extract results without a
titlefall back to the result or requested URL as the display heading instead of raisingKeyErrorand discarding usable page content. (#5280) - uploads: Fenced code blocks are excluded from uploaded-document outlines, so code comments and bold examples inside fences no longer crowd real sections out of the 50-entry heading budget. (#5281)
- subagents: After compaction, the delegation ledger distinguishes execution completion from task acceptance and preserves bounded examples of a completed subagent's unmet and unverified acceptance criteria, so the lead agent repairs the remaining gaps instead of treating the completed result as fully done. (#5287)
- dev:
_pick_python()validates interpreter candidates through/usr/bin/env, mirroring how the frontend is launched, somake devstarts the frontend on Windows where Microsoft Store Python stubs satisfy Bash-side probes but notenv. (#5181)
Performance
- runtime: Index
MemoryRunStorebythread_idandMemoryRunEventStoreevents byrun_idto avoid O(n) scans. (#3562, #3686) - subagents: Deduplicate streamed AI messages via a seen-id set (O(n²) -> O(n)). (#3687)
- sandbox: Cache
LocalSandboxpath-rewrite regexes and local-path masking patterns per instance instead of recompiling per search match. (#3648, #3713) - messages: Index tool-call results per group. (#4411)
- frontend: Coalesce streaming renders to a frame budget instead of per chunk. (#4425)
- frontend: Stop re-deriving message content on every stream chunk. (#4441)
- sandbox:
read_filereads only the requested line range from the sandbox instead of fetching the whole file first. (#3824) - browser: Encode Browser Live progress frames as JPEG to cut progress payload size. (#4836)
- middleware: Inject
view_imagecontent viawrap_model_callinstead of a checkpointed hidden message, so up to 20 MB of base64 no longer sits in two checkpoints per viewed image and an interrupted run can no longer leave the payload behind. (#5014) - frontend: Cache settled copy-data derivation across streaming chunks, so each chunk no longer re-derives toolbar/copy text for every settled message. (#5095)
- runtime: Bound gateway memory after terminal runs, stopping the post-GC low-water mark from creeping upward across completed sessions. (#5112)
- frontend: Chat streams request
messages-tuple+updates+custominstead of fullvaluesstate snapshots — the retransmitted historicalvaluesmessages were ~75% of SSE payload — folding reducer events into rendered state locally while keeping full snapshots only for replay-gap recovery. (#5159)
Security
- prompt-injection: New input-sanitization middleware defends against
prompt-injection, forged framework tags in the input guardrail are blocked,
and system context is injected as a
SystemMessagefor role isolation. (#3662, #4155, #3661) - prompt-injection: HTML-escape untrusted content rendered into model prompts
- prompt-injection: Close two input-sanitization bypasses.
hide_from_uiand a humanname="summary"tellis_genuine_user_messagethat the framework authored a message, which skips sanitization entirely. Untrusted run input and thread-state writes carrying either marker are now marked server-side and sanitized regardless, so a caller can no longer land a raw<system-reminder>outside the user-input boundary markers that the lead-agent prompt declares trusted framework data. The markers themselves are preserved, so messages that usehide_from_uionly to stay out of the transcript — quoted conversation context, sidecar context, the agent save command, HumanInputCard replies — keep doing that, and trusted internal launchers are unaffected. Sanitization also covers every genuine user message instead of only the newest: the transformation is request-scoped, so a last-turn-only scan neutralized a payload for exactly one model call and then replayed it verbatim from the next turn on. (#5375) - secrets: Scrub inherited secret environment variables (
MYSQL_PWD,REDISCLI_AUTH, abbreviated*_PASS, and PostgresPGPASSFILE) from the skill environment; request-scoped secrets are bound for both slash-activated and autonomously-invoked skills. (#4018, #4026, #3871, #3938) - web_fetch: SSRF guard for self-hosted providers. (#3942)
- guardrails: An empty allowlist now denies all tools instead of failing open. (#4067)
- authz: Global skills-management endpoints now require admin; the legacy
skills mount is gated by user visibility; artifacts honor a trusted
owner-user-idheader; and the trusted authorization principal is propagated through the runtime. (#3855, #3985, #3982, #4203) - auth: Persist the
csrf_tokencookie for the access-token lifetime. (#3872) - storage: Stop persisting base64 image data in checkpoint state. (#4140)
- mcp: Reject legacy MCP credentials in run metadata. (#4448)
- mcp: Constrain stdio launcher arguments and environment variables at
the config API, rejecting launcher flags and env names that could turn an
allowlisted
npx/uvxserver registration into arbitrary code execution. (#4617) - auth: Harden validation of the post-login
nextpath. (#4587) - runtime: Honor the LangGraph Server's authenticated user identity across agents, uploads, thread data, memory, and skills, and reject client-supplied auth identity fields. (#4538)
- frontend: Send the session cookie on model, workspace-change, and ranged artifact reads in split-origin deployments. (#4827)
- frontend: Restore sanitization in custom streamdown rehype chains, so
artifact markdown previews and the memory settings summary can no longer
render hostile HTML such as
javascript:links oron*event handlers. (#4987) - skills: Copy projected skill files instead of hardlinking them, so a sandboxed write can no longer mutate the canonical skill source, and fail closed on a drifted projection namespace on every platform, including Windows. (#4825, #4830)
- scripts: Redact secret-shaped keys (
db_pass,signing_key, ...) wherever they appear in bundled config, not only under well-known key names. (#4242) - sandbox: Sanitize MCP-sourced tool results through the same trust
boundary as the built-in web tools, so a hostile or compromised MCP server
can no longer hand the model forged
<system-reminder>or user-input boundary tags. (#4839) - sandbox: Harden local Docker sandbox containers: published ports bind
the Docker bridge gateway instead of
0.0.0.0when the sandbox host is non-loopback (DEER_FLOW_SANDBOX_BIND_HOST=0.0.0.0restores the broad bind), Docker's default seccomp profile replaces unconditionalseccomp=unconfined(opt back in withDEER_FLOW_SANDBOX_SECCOMP_UNCONFINED=1), and containers drop all capabilities, getno-new-privileges, and run with bounded resources. (#4986) - authz: Enforce run-create authorization on stateless stream/wait
endpoints (
runs:create), and require boththreads:writeandruns:createfor scheduled-task create, update, resume, and manual-trigger mutations. (#5030) - authz: Re-check the authorization policy before reusing a persisted
sandbox, so a revoked
sandbox:executegrant takes effect on the next sandbox-backed turn instead of outliving the policy in the cached sandbox. (#5006) - skills: Enforce custom-Agent skill allowlists at the sandbox filesystem
level: an explicit
skillspolicy materializes a signed per-user/thread skills view, so a custom agent with shell or file tools can no longer read skills its policy excludes. (#5077) - runs: Reject cancel/rollback actions on GET stream joins with
405 Method Not Allowed- cancel-then-stream is a POST operation - closing a state change that CSRF middleware deliberately exempted on safe methods; action-less GET joins are unchanged. (#5092) - lark: Lark CLI credential trees on Windows enforce private ACLs
(gateway-SID-only protected DACL, reparse-point rejection, handle-relative
traversal), extending the POSIX
0700/0600confidentiality contract against inherited grants, junction redirection, hard-link aliasing, and TOCTOU replacement. (#5141) - sandbox:
SSH_AUTH_SOCKis scrubbed from the sandbox subprocess environment — inheriting the host ssh-agent socket lets sandboxed code sign and authenticate with every key the agent holds — unless a skill explicitly declares it via required-secrets. (#5145) - artifacts: Serve XML artifacts as download attachments like HTML and
SVG.
GET /api/threads/{id}/artifacts/{path}rendered.xml,.xsl, and.rdffiles — and+xmltypes such as.rsswherever the host MIME database maps them — inline in the application origin, so an XML document with an XHTML-namespaced<script>, written by a prompt-injected agent and opened from a chat link, could call the API with the viewer's session. Every XML MIME type (text/xml,application/xml,text/xsl, any+xmlsubtype) is now treated as active content, including.skillarchive members; the artifacts panel keeps previewing XML through its ranged fetch. (#5353)
Documentation
- docs: Clarify how
LocalSandboxProviderresolvessandbox.mounts[].host_pathunder production Docker, with gateway bind-mount and config examples. (#3833) - docs: Document that Crawl4AI >= 0.9 requires a bearer token. (#4518)
- docs: Document the GitHub inbound-dedupe TTL semantics, including what redeliveries are not deduped, and tighten the redelivery tests. (#4274)
- docs: Update the agent AGENTS.md and ARCHITECTURE.md guides. (#4817)
- docs: Document the Honcho memory backend with a dedicated guide and a long-term memory section entry in the README. (#4822)
- docs: Align custom-agent documentation with the API across the English
and Chinese agents/threads/lead-agent pages: the required ASCII
namerequest field, lowercase storage,/api/agents/checkname-availability behavior, and no auto-derived slug fromdisplay_name. (#4944)
Internal
- tests: Migrate frontend unit tests to rstest and run hook-level tests in a DOM environment. (#3703, #4453)
- tests: Require explicit opt-in for live client tests. (#4482)
- tests: Rename the LLM-error test stand-in instead of the shared FakeError. (#4744)
- tests: Replace the magic unwritable absolute path in tool-output tests with a self-constructed failure condition. (#4722)
- tests: Add multi-turn message-stream invariants as graph integration tests. (#3708)
- tests: Add trace-based behavioral tests with Monocle Test Tools, asserting agent routing, tool calls, and token/duration cost. (#4025)
- tests: Cover passive skill tool visibility in the MCP layer. (#4247)
- tests: Add SQL and concurrent-reconciler coverage for lease-aware orphan recovery. (#4427)
- tests: Restore memory updater regression coverage. (#4490)
- tests: Lock in POST logout from the gateway-offline banner. (#4506)
- tests: Document known instance-client false negatives in the SkillScan tests. (#4644)
- refactor: Extract frontend placeholder detection into a tested utility. (#3783)
- refactor: Consolidate E2B client lifecycle helpers and reuse the kill helper during warm-pool eviction. (#4262, #4298)
- refactor: Name the E2B capacity-ledger meta-field count so the admission offset is explicit. (#4764)
- dev: Trace self/cls attribute chains and local aliases in the blocking-IO detector's call graph, closing false negatives. (#4200)
- ci: Publish the lark-cli-init and lark-broker images. (#4558)
- dev: Route host-side pnpm consumers through a shared runner with a Corepack fallback so local workflows work without a pnpm shim. (#4405)
- bench: Add an isolated checkpoint channel-mode benchmark comparing
fullanddeltaacross latency, storage, and replay metrics. (#4395) - deps: Bump
cryptography49.0.0 -> 50.0.0,postcss8.4.31 -> 8.5.25,h24.3.0 -> 4.4.1,langgraph-checkpoint-sqliteandlanggraph-checkpoint-postgres3.1.0 -> 3.1.1, andnanoid5.1.6 -> 5.1.16. (#4681, #4683, #4737, #4738, #4747, #4748) - bench: Add a reproducible hybrid memory-eviction evaluation under
backend/scripts/benchmark/deermem_eviction/with a deterministic, blind-by-construction grader for the #4789 policy. (#4810) - bench: Measure Postgres checkpoint/blob/write storage growth in the checkpoint benchmark alongside memory and SQLite. (#5051)
- tests: Exclude
tests/blocking_io/frommake test; the dedicatedmake test-blocking-iosuite (and its CI workflow) remains the owner. (#5105) - refactor: Share sandbox identity derivation and acquire serialization across the five remote sandbox providers (RFC #4741), replacing five per-provider lock tables that grew unboundedly with process lifetime; derived ids are pinned byte-identical by per-provider golden vectors. (#5089)
- ci: Split the backend unit-test workflow into four parallel shards —
each on its own runner with isolated Postgres/Redis — using
pytest-splitwith a committed duration baseline and failing closed when the baseline is missing;make testremains the canonical full offline suite. (#5137) - dev: Launch the Playwright
webServer's Next.js throughpnpm exec, so Windows E2E runs resolve the platform package binary instead of failing on the extensionless POSIX shim. (#5185) - tests: Skip the POSIX mode-bit skill-permission assertions on Windows,
where the
chmodcontract is unobservable, so Windows contributors can reach a green backend-suite baseline. (#5244)
2.0.0 — 2026-06-15
DeerFlow 2.0 is a ground-up rewrite around a "super agent" harness with
sub-agents, persistent memory, sandbox execution, and an extensible
skills/tools system. It shares no code with the 1.x line, which now lives on
the main-1.x branch.
This release closes milestone 2.0.0 with 180 merged pull requests since the first 2.0 milestone tag.
⚠ Breaking changes
- harness: Hydrate runs from
RunStoreand persist interrupted status. Run cancellation/multitask semantics now require a working RunStore on the worker that owns the run; cross-worker cancels return 409 instead of silently appearing successful. (#2932)
Added
Agents & runtime
- agent: Custom-agent self-updates with user isolation — agents can persist
edits to their own
SOUL.md/config.yamlfrom inside a normal chat. (#2713) - loop-detection: Make loop detection configurable with per-tool frequency overrides; keep configurable on/off switch. (#2586, #2711)
- loop-detection: Defer warning injection so detector pairs cleanly with tool-call lifecycle. (#2752)
- run: Propagate
model_namefrom the gateway request through the runtime and persistence stack into the SQLite-backed store. (#2775) - subagents: Stream subagent token usage to the header via terminal task events. (#2882)
- memory: Add
memory.token_countingconfig to opt out of tiktoken for network-restricted deployments. (#3465) - suggest: Make AI follow-up question suggestions optional. (#3591)
Models & integrations
- models: Add StepFun reasoning model adapter. (#3461)
- community: Add Brave Search web search tool. (#3528)
- channels: Enhance Discord with mention-only mode, thread routing, and typing indicators. (#2842)
- im: Add user-owned IM channel connections — users can bind their own Slack/Telegram/Discord/Feishu/DingTalk/WeChat/WeCom accounts on top of the operator-configured bots. (#3487)
- models: Add patched MiMo reasoning content support. (#3298)
- models: Add MiniMax provider for image/video/podcast skills plus a new music-generation skill. (#3437)
- community: Add SearXNG and Browserless web search/fetch tools. (#3451)
- community: Add Serper Google Images provider for
image_search. (#3575) - channels: Stream Telegram agent replies by editing the placeholder message in place. (#3534)
Observability
- trace: Set the LangGraph trace name to
lead_agent(or the custom agent'sagent_name) for cleaner Langfuse/LangSmith traces. (#3101) - frontend: Refine token usage display modes. (#2329)
- defaults: Enable token usage tracking by default. (#2841)
- defaults: Raise default summarization trigger threshold. (#3174)
- trace: Attribute subagent spans to the parent thread's Langfuse trace. (#3611)
Skills
- skill: Add
blocking-io-guardskill for blocking-IO triage and runtime anchors. (#3503) - skill: Add maintainer issue and PR workflow skill. (#3554)
- skill: Strengthen the maintainer orchestrator review workflow. (#3606)
Performance
- harness: Push thread metadata filters into SQL instead of post-filtering in Python. (#2865)
- runtime: Index runs by
thread_idto avoid O(n) scans inRunManager. (#3499) - runtime: Index messages in
MemoryRunEventStoreto avoid O(n) scans. (#3531) - persistence: Cache
Base.to_dictcolumn reflection per class. (#3654) - sandbox: Speed up
should_ignore_namein glob/grep walks. (#3657)
Security
- upload: Reject symlinked upload destinations. (#2623)
- uploads: Add Windows support for safe symlink-protected uploads. (#2794)
- mcp: Mask sensitive values in MCP config API responses. (#2667)
- mcp: Harden the MCP config endpoint against malformed input. (#3425)
- auth: Reject cross-site auth POSTs. (#2740)
- gateway: Cap skill artifact preview decompression to prevent zip-bomb-style abuse. (#2963)
- sandbox: Mount the host Docker socket only in aio (DooD) sandbox mode. (#3517)
- sandbox: Do not bind-mount host CLI auth dirs by default. (#3521)
Fixed
Runtime, gateway & persistence
- runtime: Rollback restore checkpoint now supersedes newer checkpoints. (#2582)
- runtime: Persist run message summaries. (#2850)
- runtime: Bound
write_fileexecution-failure observations to keep failure traces from blowing out the context. (#3133) - runtime: Protect the sync singleton's init and reset paths. (#3413)
- runtime: Avoid PostgreSQL aggregate
FOR UPDATEon run events. (#2962) - runs: Restore historical runs from persistent store after a gateway restart. (#2989)
- gateway: Return ISO 8601 timestamps from threads endpoints. (#2599)
- gateway: Make cancel idempotent for already-interrupted runs. (#3058)
- gateway: Split
stream_existing_runinto per-method routes for unique OpenAPIoperationIds. (#3228) - events: Serialize structured DB event content. (#2762)
- persistence: Emit timezone-aware timestamps from SQLite-backed stores. (#3130)
- persistence: Reuse token usage model grouping expression. (#2910)
- runs: Ignore stale run reconnect conflicts. (#3284)
- nginx: Defer CORS to the gateway allowlist instead of double-applying it. (#2861)
- persistence: Fix runtime journal run lifecycle events. (#3470)
- gateway: Enforce thread ownership on stateless run endpoints. (#3473)
- runtime: Propagate interrupt through SSE values events for the LangGraph SDK. (#3605)
- serialization: Strip base64 image data from streamed values events. (#3631)
- history: Strip base64 image data from REST endpoint responses. (#3535)
- gateway: Attribute token usage to the actual models. (#3658)
Agents, subagents & middleware
- subagents: Make subagent timeout terminal state atomic. (#2583)
- subagents: Use model override for tools and middleware. (#2641)
- subagents: Consolidate
system_promptand skills into a singleSystemMessage. (#2701) - subagent: Isolate subagents from the parent run's checkpointer. (#3559)
- agents: Make
update_agenthonorruntime.contextuser_idlikesetup_agentdoes. (#2867) - agents: Resolve duplicate
todoschannel type conflict inTodoMiddleware. (#3200) - agents: Offload blocking filesystem IO in the custom-agent router off the event loop. (#3457)
- agents: Keep new agent bootstrap in user scope. (#2784)
- loop-detection: Keep tool-call pairing on warn injection. (#2725)
- middleware: Sync raw tool-call metadata. (#2757)
- middleware: Handle invalid tool calls in dangling pairing middleware. (#2891)
- middleware: Prevent todo completion reminder IM-message leak. (#2907)
- middleware: Normalize tool result adjacency before model calls. (#2939)
- agents: Require
config.yamlinresolve_agent_dirto skip memory-only directories. (#3481) - agents: Sync
agent_nameacross context/configurable and reject empty soul. (#3553) - middleware: Offload the uploads scan in
UploadsMiddlewareoff the event loop. (#3311) - middleware: Offload memory injection off the event loop to prevent tiktoken blocking. (#3411)
- middleware: Externalize oversized tool output into the sandbox for non-mounted sandboxes. (#3417)
- middleware: Preserve the sandbox reducer in middleware state. (#3629)
- subagents: Raise general-purpose
max_turnsto 150 and default timeout to 30 min. (#3610)
Memory & tracing
- memory: Replace short-lived
asyncio.run()with a persistent event loop. (#2627) - memory: Isolate queued memory updates by agent. (#2941)
- memory: Parse wrapped memory-update JSON responses. (#3252)
- tracing: Propagate
session_idanduser_idinto Langfuse traces. (#2944) - trace: Decode unicode escape sequences in non-ASCII memory trace info. (#3104)
Tools, sandbox & MCP
- mcp: Fix env resolution in MCP config lists. (#2556)
- models: Record Codex token usage in
usage_metadata. (#2585) - sandbox: Supplement
list_runninginRemoteSandboxBackend. (#2716) - sandbox: Disable MSYS path conversion for Git Bash on Windows. (#2766)
- sandbox: Avoid blocking sandbox readiness polling. (#2822)
- sandbox: Uphold the
/mnt/user-datacontract at theSandboxAPI boundary. (#2881) - sandbox: Scope provisioner PVC data by user. (#2973)
- sandbox: Merge idempotent sandbox state updates. (#3518)
- tools: Introduce
Runtimetype alias to eliminate Pydantic serialization warnings. (#2774) - tools: Preserve
tool_searchpromotions across re-entrantget_available_tools. (#2885) - harness: Wrap async-only config tools for sync client execution. (#2878)
- harness: Wrap all async-only tools for sync clients. (#2935)
- tool-search: Reliably hide deferred MCP schemas by removing the ContextVar. (#3342)
- search: Fix DDGS Wikipedia region handling. (#3423)
- web_fetch: Support a proxy for the Jina reader in restricted networks. (#3430)
- sandbox: Persist lazily-acquired sandbox state via
Command. (#3464) - sandbox: Fix stale AIO sandbox cache reuse. (#3494)
- sandbox: Create a shell session before retrying on a fresh id. (#3577)
- sandbox: Stop flagging string-literal path fragments as unsafe absolute paths. (#3623)
- sandbox: Return an actionable hint when
read_filehits a binary file. (#3624) - mcp: Make stdio MCP-produced files resolvable via virtual sandbox paths. (#3600)
- mcp: Surface admin-required state on the settings tools page. (#3533)
- mcp: Add a tools cache reset endpoint. (#3602)
- uploads: Fix the upload file size contract. (#3408)
Skills & channels
- skills: Enforce
allowed-toolsmetadata. (#2626) - skills: Harden slash skill activation across chat channels. (#3466)
- skills: Fix custom skill install permissions. (#3241)
- channels: Authenticate gateway command requests. (#2742)
- skills: Surface the offending line and a quoting hint on SKILL.md YAML errors. (#3335)
- skills: Keep skill archive installation off the event loop. (#3505)
- channels: Ignore hidden control messages when extracting replies. (#3270)
- channels: Reload config on channel restart. (#3514)
- channels: Surface WeCom WebSocket connection failures. (#3526)
- channels: Close the Discord file handle after upload. (#3561)
- channels: Require a bound identity for user-owned IM messages. (#3578)
- channels: Scope IM files and helper commands to the owner. (#3579)
- channels: Make runtime provider state authoritative. (#3580)
- channels: Harden runtime credential management APIs. (#3581)
- channels: Make the channel connect flow deterministic. (#3582)
- channels: Centralize shared channel retry helpers. (#3583)
- channels: Add operational guardrails. (#3584)
- channels: Unsubscribe channel listeners by equality. (#3608)
Auth
- auth: Replace setup-status 429 rate limit with a cached response. (#2915)
- auth: Persist auto-generated JWT secret so it survives restarts. (#2933)
- auth: Align auth-disabled mode with mock history loading. (#3471)
Frontend
- frontend: Restore
localhostfallback forgetGatewayConfigin prod mode. (#2718) - chat: Prevent the first user message from being swallowed in new conversations. (#2731)
- frontend: Use backend thread token usage for the header total. (#2800)
- frontend: Wait for async chat submit before clearing the input. (#2940)
- frontend: Resolve login page flickering and the resize-observer loop. (#2954)
- frontend: Deduplicate restored thread messages. (#2958)
- frontend: Avoid duplicate optimistic user message. (#3002)
- frontend: Hide the copy button for streaming assistant messages. (#3176)
- frontend: Show a new thread in the sidebar immediately on creation. (#3283)
- frontend: Isolate new chat thread messages. (#3508)
- frontend: Cap deeply nested list indentation to prevent render crashes. (#3393, #3570)
- token-usage: Dedupe token usage aggregation by message id. (#2770)
- frontend: Fall back to Streamdown clipboard copy. (#3397)
- frontend: Remove the Backspace shortcut for deleting prompt attachments. (#3410)
- frontend: Restructure the Memory settings toolbar into two rows. (#3433)
- suggestions: Strip inline
<think>reasoning before parsing follow-up questions. (#3435) - frontend: Stop fetching follow-up suggestions when they are disabled. (#3599)
- frontend: Paginate the workspace chat list beyond 50 threads. (#3485)
- frontend: Prevent user message bubble overflow with long unbreakable strings. (#3488)
- frontend: Keep the workspace interactive when the SSR auth probe cannot reach the gateway. (#3495)
- frontend: Render user messages as plain text and cap blockquote nesting. (#3502)
- frontend: Reset the active chat after deletion. (#3519)
- frontend: Improve the mobile workspace layout. (#3646)
- frontend: Render full content for multi-part AI messages. (#3649)
Build, deploy, scripts & config
- packaging: Add
postgresextra for store/checkpointer support; clarify install guidance. (#2584) - harness: Resolve runtime paths from the project root. (#2642)
- docker: Force nginx to resolve upstream names at request time. (#2717)
- docker: Default Gateway to a single worker to prevent multi-worker breakage. (#3475)
- scripts: Preserve
uvextras acrossmake devrestarts. (#2767, #2754) - scripts: Clean up local nginx on stop. (#3005)
- deploy: Fall back to
python/opensslwhenpython3is absent for secret generation. (#3074) - config: Make the reload boundary discoverable from code. (#3144, #3153)
- replay-e2e: Key replay fixtures by caller and conversation. (#3453)
- setup: Refresh LLM provider wizard defaults. (#3421)
- config: Coerce null
config.yamllist sections to an empty list. (#3434) - scripts: Exclude runtime state from gateway reload. (#3426)
- scripts: Create the backend/sandbox dir before the uvicorn reload-exclude. (#3460)
- scripts: Stop next-server correctly after
make start-daemon. (#3498) - makefile: Fix per-commit hooks installation. (#3569)
- replay-e2e: Match replay by conversation, not the living system prompt. (#3436)
Changed
- provider (refactor): Share assistant payload replay matching across providers. (#3307)
- lead-agent (refactor): Make
build_middlewarespublic to drop the last cross-module private import. (#3458) - todo (refactor): Remove the unused completion reminder counter. (#3530)
Documentation
- Document blocking-IO detection usage and maintenance. (#3233)
- Clean standalone LangGraph server remnants from docs. (#3301)
- Add AI assistance disclosure to the PR template and CONTRIBUTING. (#3398)
- Document custom AIO sandbox images. (#3548)
Internal
- dev: Add async/thread boundary detector. (#2936)
- runtime: Add lifecycle end-to-end coverage. (#2946)
- windows: Add
PYTHONIOENCODINGandPYTHONUTF8to backend Makefile targets. (#3069) - blocking-io: Fail-loud repo-root resolution and shared detector CLI shim. (#3512)
- runtime: Add a Blockbuster runtime anchor for
JsonlRunEventStoreasync IO. (#3313) - ci: Consolidate PR/issue labeling and fix the reviewing-job crash and label thrash. (#3455)