* fix(gateway): confine artifact PUT to /mnt/user-data/outputs after path resolution
The outputs-only guard on PUT /api/threads/{id}/artifacts/{path} was a
string-prefix check on the raw path. A percent-encoded `..`
(`outputs/%2e%2e/uploads/x.txt`) survives nginx's variable proxy_pass
untouched, is decoded by Starlette, passes the prefix check, and the
resolver only confines the result to `user-data/` -- so an owner could
overwrite a sibling upload or workspace file in their own thread.
Collapse dot segments before the prefix check, and re-check the resolved
host path against the resolved outputs root so a symlink planted inside
`outputs/` cannot redirect the write either. The normalized virtual path
is what the response echoes and what non-mounted sandboxes receive.
* refactor(gateway): share the outputs-confinement rule with channel attachments
Review follow-up on #5321: the "only under /mnt/user-data/outputs" rule was
implemented independently by the artifact editor and by IM-channel
attachment delivery, and the two copies had already drifted.
Move it into app/gateway/path_utils.py as normalize_outputs_virtual_path
(collapse `..` before the prefix check) and resolve_outputs_confined_path
(re-check the resolved host path against the resolved outputs root, which
also catches a symlink planted inside outputs/). PUT /artifacts and
ChannelManager._resolve_attachments both call the helper; artifact_archive
keeps its stricter ZIP-member rules layered on top.
Tests that previously stubbed resolve_thread_virtual_path for the editor now
stub resolve_outputs_confined_path, and the channel attachment tests patch
path_utils.get_paths, which the helper binds at import like the other
consumers. The confinement itself is pinned by tests/test_gateway_path_utils.py.
36 KiB
IM Channels System (app/channels/)
Bridges external messaging platforms (Feishu, Slack, Telegram, Discord, DingTalk, GitHub) to the DeerFlow agent via Gateway's LangGraph-compatible API.
Architecture: Channels communicate with Gateway through the langgraph-sdk HTTP client (same as the frontend), ensuring threads are created and managed server-side. The internal SDK client injects process-local internal auth plus a matching CSRF cookie/header pair so Gateway accepts state-changing thread/run requests from channel workers without relying on browser session cookies.
Components:
message_bus.py- Async pub/sub hub (InboundMessage→ queue → dispatcher;OutboundMessage→ callbacks → channels)store.py- JSON-file persistence mappingchannel_name:chat_id[:topic_id]→thread_id(keys arechannel:chatfor root conversations andchannel:chat:topicfor threaded conversations). Every access to_datamust be protected by_lock;list_entries()snapshots keys and copied entries under the lock, then formats the result after releasing it so concurrent channel threads cannot resize the dictionary during iteration without extending the critical section.manager.py- Core dispatcher: creates threads viaclient.threads.create(), routes commands including/goal(setting a goal persists it through Gateway and then routes the objective as a chat turn) and/agent(listis owner-scoped;usecreates a fresh thread and persists the Custom Agent selection under both the channel restart key and Web's canonical routing metadata so existing checkpoint lineage never changes agent across IM/Web), keeps Slack/Discord onclient.runs.wait(), usesclient.runs.stream(["messages-tuple", "values"])for Feishu/Telegram incremental outbound updates, serializes same-thread Feishu turns in-manager when the channel'sChannelRunPolicy.serialize_thread_runs=Trueso rapid follow-ups queue instead of tripping the runtime busy reply, and switches toclient.runs.create()(fire-and-forget, returns once the run ispending) for channels whoseChannelRunPolicy.fire_and_forget=Trueso long autonomous runs do not hit the SDK default 300shttpx.ReadTimeoutA swallowed streaming failure publishes its final outbound before releasing the inbound dedupe key, so a provider redelivery can retry without overtaking the terminal reply. What may be published from the stream is an allowlist, not a denylist (_accumulate_stream_text): only assistant message types — LangChain serializesAIMessage.typeas"ai"andAIMessageChunk.typeas"AIMessageChunk", plus the OpenAI-style"assistant"spelling for foreign runtimes — become displayable text. The previous rule rejected only payloads whosetypecontained"tool"and therefore published everything else, which leaked DeerFlow's hidden model context to every streaming IM channel:DynamicContextMiddlewareinjects recalled memory as a hiddenHumanMessage(type == "human") and rewrites the user's own turn into a newHumanMessage,DurableContextMiddlewareinjects a hidden<durable_context_data>HumanMessage, and LangGraph fans those state writes out on themessages-tuplestream. Proved live on a Buzz relay, which published a<memory>fact block and, in another run, a verbatim echo of the user's own message as the assistant's reply. Matching is by prefix (ai/assistant), never substring, because ordinary words contain"ai"(chain,domain). The message type is resolved by_stream_payload_type, which handles both themodel_dump()shape DeerFlow's own gateway emits and LangChain'sto_json()constructor shape (whose top-leveltypeis the literal"constructor", with the class name at the tail of theidpath). A barestrpayload is no longer accepted at all: it carries no type information, so it cannot be attributed to the assistant, and nothing in DeerFlow produces one (runtime/serialization.py::serialize_messages_tuplealways emits[message_dict, metadata]).base.py- AbstractChannelbase class (start/stop/send lifecycle). Provider callbacks that submit coroutines from SDK threads must use_submit_threadsafe_coroutine(): it creates and retains the realasyncio.Taskon the owner loop instead of treatingrun_coroutine_threadsafe()'s proxy Future as a completion signal. Submission is closed atomically with shutdown, andstop()must call_close_and_drain_threadsafe_futures()before tearing down SDK resources.service.py- Manages lifecycle of all configured channels fromconfig.yaml. Shutdown closes manager admission first and keeps transports alive until every manager worker/follow-up watcher has exited. A successful manager stop therefore owns no live handler; if the Gateway's outer timeout cancels shutdown, the service retains its channel objects and global singleton so unfinished resources are not detached and cleanup can be retried.slack.py/feishu.py/telegram.py/discord.py/dingtalk.py- Platform-specific implementations (feishu.pytracks the running cardmessage_idin memory and patches the same card in place;telegram.pyaccepts inbound text/photos/documents, preserves media captions, hands token-free attachment bytes to the shared upload pipeline, edits the "Working on it..." stream target in place viaeditMessageText, and can optionally send final Markdown replies as Rich Messages throughchannels.telegram.rich_messages;discord.pyregisters typing loops before inbound handling yields and_start_typing()refuses work once_runningis false; becausestop()runs on the main loop while typing tasks belong to_discord_loop, normal cross-thread cancellation, awaiting, and map cleanup are scheduled there with a bounded wait, while_run_client()drains tasks in itsfinallyblock before an exception/disconnect can make that loop unusable; an already-stopped foreign loop must never have its tasks awaited from the main loop; to that end every outbound cross-loop call (send/send_file/_get_channel_or_thread) goes through_run_on_discord_loop, which bounds the await (DISCORD_OUTBOUND_TIMEOUT_SECONDS30s;DISCORD_UPLOAD_TIMEOUT_SECONDS120s for file uploads, whose unbounded-size payload needs room for a slow uplink plus 429 retry-after) and fails fast with aRuntimeErrorwhen the client loop is missing or not running — a dead client becomes a logged send failure instead of a permanently hungChannelManagerworker — andis_runningreports client-thread aliveness (likefeishu.py) soensure_channel_readycan restart the channel after_run_client()exits on a fatal error;dingtalk.pyoptionally uses AI Card streaming for in-place updates whencard_template_idis configured, and overridesreceive_fileto download inbound images (picture/richText) and documents (file) bydownloadCodeinto the thread uploads bucket, mirroringfeishu.py)buzz.py- Buzz (Nostr relay) implementation: one NIP-42-authenticated websocket, pubkey-allowlist + mention/DM/thread-follow gating, streaming replies via in-place kind-40003 edits; requires thebuzzdependency extra. Its durable seen-event replay guard uses asyncaseen()/arecord()/aflush()boundaries: initial JSON reads and coalesced atomic writes run off the Gateway event loop, andstop()awaits the final flush before returning. Subscription model (operator-facing version in IM_CHANNEL_CONNECTIONS.md): the relay fans kind-9 chat events out only to channel-scoped subscriptions, proved against a live relay —REQ {"kinds":[9]}is accepted and answered withEOSEbut never receives an event (the connector authenticated and then sat silent forever),REQ {"kinds":[9],"#h":[uuid]}works, and a multi-value#hmatches nothing, so it is strictly one REQ per channel (the same shape as Buzz's ownbuzz-acpharness). Every connection therefore rebuilds three kinds of subscription after NIP-42 auth:buzz-discovery({"kinds":[39000]}, a historical query returning exactly the channels this identity belongs to, one stored event each thenEOSE— adding#preturns zero, do not "narrow" it),buzz-membership({"kinds":[44100,44101],"#p":[us],"since":<connect time −MEMBERSHIP_LOOKBACK_SECONDS>}, the relay-signed member-added/member-removed notifications whoseptag names the affected member andhtag the channel), and onebuzz-chat-<uuid>per discovered channel. Chat subscriptions open as each kind-39000 arrives; the discoveryEOSEis the completeness barrier that retries any that failed and warns when discovery found nothing. A kind-44100 for our pubkey subscribes to the new channel live (then re-issues discovery so its name/type reach the DM-detection cache, but only when that channel's metadata is actually missing — an unconditional refresh is how a burst of 44100s multiplied into one discovery pass each); a kind-44101 issuesbuzz_nostr.close_framefor exactly that channel's subscription and drops its metadata. The membership subscription is scoped to LIVE events and this is load-bearing: buzz-relay stores 44100/44101 and serves history newest-first (default limit 2000), so an unscoped filter replayed the whole membership history on every connect — every stored add read as live, re-running discovery once each (M+1 discovery passes × N stored kind-39000 events per connect, observed live as twochannel discovery completelines and one channel logged<unnamed>because the 44100 path subscribed before its metadata arrived), re-subscribing channels we have since been removed from, and letting a stored 44101 transiently unsubscribe a channel we are still in.sinceis anchored at the moment the socket opened (_session_started_at) minusMEMBERSHIP_LOOKBACK_SECONDS(60) of slack, which covers both relay clock skew and a membership change published during the connect/auth handshake; the slack can only cost an idempotent replay of the last minute. A relayCLOSEDframe is recovered, not merely forgotten — every subscription on the socket fails silently when dropped, so_handle_closedre-issues it, bounded byMAX_RESUBSCRIBE_ATTEMPTS(3) per subscription id per connection and per auth epoch (_resubscribe_attemptsis reset on session start, on re-auth, and bystop(), so pre-auth rejections — which the auth branch already recovers wholesale — never spend the authenticated session's budget). The first retry is immediate (the common case is a one-off hiccup); later ones back off 1s then 2s, awaited inline in the read loop rather than in a background task that could outlive its own socket.auth-required:before this socket has completed its NIP-42 handshake is the expected bootstrap sequence, not a refusal, and_handle_closedshort-circuits it ahead of every recovery path: the connector opens its control REQs immediately in case the relay serves unauthenticated reads, a closed relay answersauth-required:plus anAUTHchallenge, and the auth branch re-opens everything. That case is logged at DEBUG and consumes neither the permanent-refusal branch nor the retry budget (a chat subscription is still dropped from_chat_subscriptions, since it genuinely is not subscribed and discovery is what re-opens it). Treating it as permanent produced an operator-facing warning claiming discovery/membership tracking was DOWN in the same run where discovery then completed, every channel was subscribed, and a brand-new channel's kind-44100 was picked up one second later. The boundary is the per-socket_auth_completedflag, set once the signed AUTH event has been sent and cleared on session entry, session exit, andstop(); the same reason after that stays loud, and says the subscription is down until the relay's next AUTH challenge or the next reconnect rather than borrowing the non-auth wording. Then_is_transient_closedecides whether to retry at all: NIP-01/NIP-42auth-required:/restricted:/blocked:/mute:/invalid:/pow:prefixes and buzz-relay's own removal/revocation prose are permanent (do not fight the relay over a channel that is no longer ours; a post-authauth-required:is recovered by the AUTH branch, not by re-issuing),rate-limited:/error:/no-reason-at-all and anything unrecognized are transient — the default resolves toward "keep listening" because going silently deaf is the failure this exists to remove, and the attempt budget bounds a wrong guess. A chatCLOSEDis only ever recovered for a channel already in_chat_subscriptions: aCLOSEDis relay-supplied, so acting on an unknown one would let a relay induce a subscription just by naming a channel. Every subscription that goes unlistened is logged at WARNING, never INFO. Subscription ids are deterministic per channel precisely so one can be replaced or closed without disturbing the others on the socket, and_chat_subscriptionsis per-socket state cleared on session end, on re-auth (a pre-auth REQ may have been rejected), and bystop(). Three bounds on remote-fed state:MAX_CACHED_CHANNELS(512) caps the kind-39000 metadata cache, the watermark map, and the resubscribe-attempt map;MAX_CHANNEL_SUBSCRIPTIONS(256, well under buzz-relay's own 1024-per-connection ceiling) caps live chat subscriptions — at the cap new channels are refused and named in a warning rather than evicting a working subscription. Known bound (documented, not fixed): the relay caps historical delivery at 2000 events per subscription, newest-first, even with asince, so >2000 unread messages in a single channel across a disconnect loses the oldest — the relay never sends them and the watermark advances past them. That is the one remaining path that can skip; everything else fails toward replay. Trust model (operator-facing version in IM_CHANNEL_CONNECTIONS.md): every inboundEVENTis authenticated at the singlehandle_relay_framechoke point — the NIP-01 id is recomputed from the delivered payload and the BIP-340 Schnorr signature verified against the claimedpubkey(buzz_nostr.verify_event, pure and total: malformed input returnsFalse, never raises) — soev["pubkey"], the authorization principal for both the allowlist and the/connectbind, cannot be forged by a relay the DeerFlow operator does not run. What remains trusted is the authorship of kind-39000 channel metadata: any member can sign one, and because per-channel subscriptions are now driven by discovery, a forged kind-39000 has two effects rather than one — it can mark a channeltype: "dm"(relaxingrequire_mentionfor that channel) and it can induce a chat subscription for a channel of the forger's choosing, since the channels we listen to are exactly the channels we hold metadata for. Neither makes anything be acted on:allowed_usersand per-event signature verification are independent gates, so an induced subscription only means the relay reads its own traffic back to a subscriber that drops it, bounded byMAX_CHANNEL_SUBSCRIPTIONS(which refuses rather than evicts, so it cannot displace a real channel). Same for a forged kind-44100, except itsptag is re-checked locally so it must at least name us. Closing this needs a configured trusted relay pubkey, whichrelay_urlis not.allowed_usersis deny-by-default (empty = nobody, unlike siblings' empty = everyone), sostart()logs a WARNING when it is empty and each drop logs at DEBUG. The resubscribe cursor (since) is per channel, advances only for events that were actually processed, and never pastnow + MAX_FUTURE_SKEW_SECONDS, because it is peer-supplied (created_at) and a single future-dated event otherwise made the connector permanently deaf. Per channel rather than global is the safety-critical half: subscriptions are per channel, so one shared cursor is the newest event seen in any channel and a busy channel would drag it past a quiet channel's unread messages, skipping them on the next reconnect — measured on a live relay, three channels of one identity sat ~28h apart. Per-channel cursors can only ever cost duplicate delivery (absorbed by the manager'sevent_iddedupe), and an evicted cursor degrades to "nosince", i.e. the relay's default backlog — both fail toward replay, never toward a miss. Streaming tracks every oversize chunk index (_stream_targetsfor chunk 0,_stream_tailsfor the rest), since the manager republishes cumulative text and repostingchunks[1:]per update flooded the channel; all of it is per-connection state cleared bystop(), and the remote-fed kind-39000 cache is capped atMAX_CACHED_CHANNELS.send()refuses outright to publish text carrying a hidden model-context wrapper (<memory>,<durable_context_data>,<system-reminder>—_HIDDEN_CONTEXT_MARKERS), logging at ERROR and clearing the stream bookkeeping on a blockedis_final. This is defense in depth behind the manager's allowlist, and it lives here rather than in a sibling connector because on Buzz a leak is permanent: every streaming update is an immutable public Nostr event, so a corrective edit only changes what clients render while the original leaked event stays on the relay. Matching is on the literal opening tag, so a reply that merely talks about memory is still published Buzz seen-event shutdown stays bounded and retryable:aflush()awaits any in-flight write, attempts at most one final snapshot, leaves a still-changing generation dirty for fail-open replay, and returns without a live persistence timer. A Gateway cancellation does not cancel the underlying worker-thread write;BuzzChannel.stop()tracks cleanup completion separately from transport admission so ChannelService can retry the retained channel without racing a second snapshot against that write. The store is quiesced before stop (including the already-stopped guard), so a timed-out relay task that records after stop only marks data dirty and cannot schedule detached file work; a repeatedstop()still drains that dirty state, whileBuzzChannel.start()explicitly resumes scheduling and flushes it automatically.github.py- Webhook-driven GitHub channel. Inbound messages come fromPOST /api/webhooks/github; outbound is log-only because GitHub agents post explicitly withghfrom their sandbox when they choose to comment or create a PRapp/gateway/routers/channel_connections.py- Browser-facing user connection and disconnect APIsdeerflow.persistence.channel_connections- SQL-backed user-owned connection, optional credential, connect state, and conversation store
Message Flow:
- External platform -> Channel impl ->
MessageBus.publish_inbound()- For GitHub, the webhook router verifies the delivery then calls
fanout_event(bus, ...); matching agent bindings publish oneInboundMessageeach instead of a long-polling channel worker. - Telegram photo/document updates use the largest photo size or document metadata, preserve
message.caption, enforce the hosted Bot API's 20,000,000-byte download ceiling before and after download, and never expose the token-bearing Bot API file URL. Downloaded bytes cross the adapter/manager boundary only throughmessage_bus.INBOUND_FILE_CONTENT_KEY; the manager consumes that transient field before persisting safe upload metadata. - Feishu/Lark inbound image/file downloads read at most 20,000,001 bytes and reject anything above 20,000,000 bytes before persistence or sandbox sync. Oversize, unsafe-path, and path-resolution failures rewrite only that attachment placeholder to
Failed to obtain the [type]so later attachments in the same message can still load.
- For GitHub, the webhook router verifies the delivery then calls
ChannelManager._dispatch_loop()consumes from queue- For user-owned channel connections, incoming messages carry
connection_id,owner_user_id, andworkspace_id;owner_user_idbecomes the DeerFlow runuser_id, while the raw platform user id remainschannel_user_id. The Gateway acceptschannel_user_idonly from an internally authenticated channel caller's top-levelbody.context, clears it from both free-formbody.configsections, and writes it into runtime context only (neverconfigurable, which is checkpointed).bash_toolexposes it to sandbox commands as the fixed env varDEERFLOW_CHANNEL_USER_ID— via a shell-quoted command-string prefix, NOT theexecute_command(env=...)channel, which is reserved for request-scoped secrets and would switchAioSandboxonto thebash.execpath (image >= 1.9.3, fresh session per call). Per-call injection keeps group-chat identity correct (one thread/sandbox, many senders) without depending on the AIO shell's session semantics: every IM-channel command carries an explicitexport VAR=<id>;(valid id) orunset VAR;(empty / non-str / over the 256-char cap). The AIO no-env path reuses a persistent shell session (the reason for the class lock, #1433), so a bare command could otherwise resolve a stale id an earlier sender exported; theunsetcloses the window the length/type guard would open (a dropped id would inherit the previous sender's value). Non-IM runs (nochannel_user_idin context) are left untouched. Not injected on the Windows local sandbox (its PowerShell/cmd.exe fallback has noexport/unset). Propagates acrosstaskdelegation:task_toolcaptures the dispatching turn's id and the subagent executor forwards it into the subagent's runtime context, same as the guardrail attribution fields. The runtime-context value is authorization-grade at the Gateway/guardrail boundary, but the exported shell variable remains informational because any bash command can overwrite its own environment; skills must not treat the shell variable itself as authenticated identity. Tests:tests/test_gateway_services.py,tests/test_channel_user_id_env.py - For chat: look up/create thread through Gateway's LangGraph-compatible API
- Feishu/Telegram chat:
runs.stream()→ accumulate AI text → publish multiple outbound updates (is_final=False) → publish final outbound (is_final=True) - Slack/Discord chat:
runs.wait()→ extract final response → publish outbound 6b. GitHub chat (ChannelRunPolicy.fire_and_forget=True):runs.create()returns once the run ispending; the manager does not wait for the final state and does not publish an outbound. The agent posts its own reply mid-run viaghfrom the sandbox.ConflictErroron a busy thread still trips the standardTHREAD_BUSY_MESSAGEpath (log-only on GitHub); when the channel's policy also setsbuffer_followups_on_busy=True(GitHub's default — see "Follow-up buffering while busy" below), the triggering message is additionally captured into a per-thread buffer instead of only logged, so a concurrent comment is not silently dropped. - Feishu channel sends one running reply card up front, then patches the same card for each outbound update (card JSON sets
config.update_multi=truefor Feishu's patch API requirement). Messages already sent inside an existing Feishu topic carry a compact source-message preview in that card, and queued same-thread follow-ups patch their own source message's card from queued → running → final without falling back to the generic busy reply. - Telegram streaming: the "Working on it..." placeholder message is registered as the stream target; non-final updates
editMessageTextit in place (channel-side throttle: 1s in private chats, 3s in groups due to Telegram's 20 msg/min group cap; 4096-char truncation; rate-limited updates dropped); the final update performs the last edit and splits >4096 texts into follow-up messages - DingTalk AI Card mode (when
card_template_idconfigured):runs.stream()→ create card with initial text → stream updates viaPUT /v1.0/card/streaming→ finalize onis_final=True. Falls back tosampleMarkdownif card creation or streaming fails - For commands (
/new,/status,/models,/memory,/goal,/agent,/help): handle locally or query Gateway API./agent listreads only the effective owner's Custom Agents;/agent use <name>validates in the same owner bucket, creates a new thread, and persistschannel_agent_namein its metadata. A Custom Agent also writes canonicalmetadata.agent_name, which makes thread-search results route Web continuation through/workspace/agents/<name>/chats/<thread_id>;lead_agentdeliberately omits that canonical key and stays on the ordinary chat route. The manager cacheschannel_agent_namefor the hot path and reloads it after restart before routing a resumed turn. An explicit selection also normalizesagent_nameacross top-level run context plus the existing RunnableConfigcontextandconfigurablecarriers (or clears all three forlead_agent) before Gateway'ssetdefaultcompatibility merge, so stale channel defaults cannot silently win. - Outbound → channel callbacks → platform reply
- GitHub is the exception: the channel logs the final assistant message and does not auto-post it to GitHub. Agents use the sandbox
ghCLI (gh issue comment,gh pr comment,gh pr create, etc.) for intentional writeback, so silence is cheap when several agents fan out on the same event.
- GitHub is the exception: the channel logs the final assistant message and does not auto-post it to GitHub. Agents use the sandbox
Owner-scoped file storage: inbound files, uploads, and output artifacts are staged under the DeerFlow owner's bucket so they land where the agent run reads/writes (users/{user_id}/threads/{thread_id}/user-data/{uploads,outputs}). ChannelManager._handle_chat resolves the storage owner once via _channel_storage_user_id(msg) (sanitized owner id, falling back to safe(msg.user_id) for unbound auth-enabled channels — mirroring _resolve_run_params's run identity; None only when no identity is available) and threads it as the user_id= kwarg through the file pipeline:
Channel.receive_file(msg, thread_id, user_id=...)— owner-bound channels persist downloaded files under the owner's bucket instead of the default bucketFeishuChannel._receive_single_file(...)/DingTalkChannel._receive_single_file(...)— normalize provider filenames, claim a collision-free basename and write it throughwrite_upload_file_no_symlinkunder the same channel lock; the returned basename drives both the agent-visible virtual path and non-local sandbox syncsandbox_files.py— non-mounted Feishu/DingTalk syncs acquire unique non-releasing execution holders, drain blockingupdate_fileworkers across repeated cancellation, and release only after the last sandbox operation, so a parallel run cannot close the shared client mid-upload_ingest_inbound_files(...)and the underlyingensure_uploads_dir/get_uploads_dir— owner-scoped via the same kwarg_resolve_attachments/_prepare_artifact_delivery— resolve output artifacts from the bound owner's bucket throughapp.gateway.path_utils.resolve_outputs_confined_path, the same outputs-only rule the artifact editor uses, so a siblinguploads//workspace/path or a symlink planted inoutputs/is skipped with a warning The cached value is reused for both the blocking (runs.wait) and streaming (_handle_streaming_chat) paths, so uploads and artifact delivery always target the same bucket even if a channel returns a rewrittenInboundMessagefromreceive_file. The bucket id matches the memory bucket resolved by_resolve_memory_user_id(both normalize throughmake_safe_user_id).
Configuration (config.yaml -> channels):
langgraph_url- LangGraph-compatible Gateway API base URL (default:http://localhost:8001/api)gateway_url- Gateway API URL for auxiliary commands (default:http://localhost:8001)- In Docker Compose, IM channels run inside the
gatewaycontainer, solocalhostpoints back to that container. Usehttp://gateway:8001/apiforlanggraph_urlandhttp://gateway:8001forgateway_url, or setDEER_FLOW_CHANNELS_LANGGRAPH_URL/DEER_FLOW_CHANNELS_GATEWAY_URL. - Per-channel configs:
feishu(app_id, app_secret),slack(bot_token, app_token),telegram(bot_token, optionalrich_messagesfor final Markdown Rich Messages),dingtalk(client_id, client_secret, optionalcard_template_idfor AI Card streaming),github(operator kill-switchenabled, plusdefault_mention_loginfor mention-required GitHub triggers),buzz(relay_url, private_key)
User-owned channel connections (config.yaml -> channel_connections):
- Disabled by default. It is a user-binding layer on top of the existing
channels.*runtime config, not a replacement for provider bot credentials. - No public IP, OAuth callback URL, or provider webhook route is required by the current implementation.
- Telegram uses a deep-link
/start <code>flow over the existing long-polling worker. Slack, Discord, Feishu/Lark, DingTalk, WeChat, and WeCom use/connect <code>over their existing outbound channel workers. - WeChat timing settings (
polling_timeout,polling_retry_delay,qrcode_poll_interval,qrcode_poll_timeout) accept only positive finite seconds; invalid values fall back to their defaults so polling cannot enter a hot loop or sleep forever. - WeCom serializes
start()andstop()for each channel instance. The SDKconnect()task covers connection setup only; after the handshake, the SDK owns a separate receive task. Shutdown cancels an in-progress connection attempt and awaits the SDK's actual asynchronous receive-task/socket cleanup before releasing lifecycle state or allowing a restart. Cancellation ofstop()still propagates, but only after owned cleanup finishes and lifecycle references are cleared; real connection failures remain reported by_on_ws_task_done. - Frontend APIs:
GET /api/channels/providers,GET /api/channels/connections,POST /api/channels/{provider}/connect, andDELETE /api/channels/connections/{connection_id}. - Browser APIs remain protected by normal Gateway auth/CSRF. Provider messages arrive through the already-configured channel workers.
- Provider-level
connection_statusreflects the user's newest connection row. With no binding it isnot_connected, except in auth-disabled local mode where a configured running channel reportsconnectedbecause all channel messages already route to the default user. - Slack replies use the configured operator bot token from
channels.slackunless per-connection credentials are present; unreadable or corrupt stored credentials are treated as unavailable. - Telegram, Slack, Discord, Feishu/Lark, DingTalk, WeChat, and WeCom workers resolve incoming platform identities to connection records before reaching
ChannelManager. - Connect-code ordering vs
allowed_users: inbound workers consume a valid/connect <code>(or Telegram/start <code>) before applying theallowed_usersfilter, so a newly allowlisted-but-unbound user can bootstrap their first bind via the browser flow. Consequence:allowed_usersis not a bind-time defense — any sender who possesses a valid code can consume it (not only allowlisted users). The bind security model rests on the code's confidentiality:secrets.token_urlsafe(16), 600 s TTL, one-timeconsume_oauth_state, and codes surfaced only in the initiating browser (never echoed to chat).allowed_usersstill gates ordinary (non-bind) messages. - Single-active-owner transfer semantics: an external identity is keyed by
(provider, external_account_id, workspace_id). The latest successful bind wins —upsert_connectionrevokes other owners' active rows for the same identity (ownership transfer). This invariant is enforced at the DB layer by the partial unique indexuq_channel_connection_active_identity(WHERE status != 'revoked'), so concurrent connects from different owners cannot both endconnected; the losing writer retries against the now-visible state.find_connection_by_external_identitytherefore resolves deterministically. - See
backend/docs/IM_CHANNEL_CONNECTIONS.mdfor provider setup, operational notes, and the architecture diagrams (connect-code flow, single-active-owner transfer, sync vs streaming dispatch, owner-scoped file storage pipeline).
GitHub event-driven agents (webhook-driven IM channel):
- Custom agents declare a
github:block in theirconfig.yamlto bind to repos and event triggers; the webhook route is fail-closed by default (mounted only whenGITHUB_WEBHOOK_SECRETis set) and exempt from auth/CSRF because authenticity is enforced by HMAC. - Registry caching is keyed by the configured agent store's opaque signature: file storage uses agent config mtimes, while database storage hashes the ordered owner/name/config/soul contents so same-timestamp writes still invalidate webhook routing.
- Outbound is log-only by design: each agent posts its own reply mid-run via the
ghCLI from its sandbox, so the manager usesfire_and_forget=Trueandruns.create()returns once pending. - Follow-up buffering while busy (issue #4121): because outbound is log-only, the pre-existing
THREAD_BUSY_MESSAGEreply on aConflictErrorwas invisible to the commenter — a comment posted while a run was already active looked like it had been silently ignored. WhenChannelRunPolicy.buffer_followups_on_busy=True(GitHub's default), aConflictErroronruns.create()now also appends the triggering message to a per-thread, in-memory buffer (ChannelManager._followup_buffers) — deduped by GitHub delivery id, capped atFOLLOWUP_BUFFER_MAX_PER_THREAD(20, oldest dropped with a WARNING log on overflow). The first successfulruns.create()on a thread now captures itsrun_idand spawns a background watcher that subscribes to that run'sStreamBridgestream; once it observesEND_SENTINEL, the watcher drains up toFOLLOWUP_DRAIN_BATCH_SIZE(10) buffered entries into one<followups-while-busy>-wrapped input and fires a follow-upruns.create()— itself watched the same way, so a backlog deeper than one batch chains into further drain cycles instead of growing one unbounded input. If that follow-upruns.create()itself hitsConflictError(e.g. a manual Web UI turn or a scheduled run raced onto the same thread), the batch is requeued rather than lost, and is retried whenever this manager next successfully creates and watches a run on that thread. Reactions/acknowledgment (e.g. GitHub'seyes/confusedreaction API) on buffered comments are intentionally out of scope for this mechanism and left to a follow-up — comments are coalesced silently. Plumbing: the watcher needs the Gateway'sStreamBridgesingleton, whichChannelManagerdid not previously have access to; it is threaded fromapp.py's lifespan (whereapp.state.stream_bridgeis already set bylanggraph_runtime) throughstart_channel_service(get_stream_bridge=...)→ChannelService.__init__→ChannelManager.__init__, as a zero-arg closure mirroring the existinglaunch_run=lambda **kwargs: launch_scheduled_thread_run(app=app, **kwargs)pattern used forScheduledTaskServicein the same lifespan function. AChannelManagerconstructed without it (e.g. directly in a test) still buffers safely — it just has no watcher to auto-drain. Scope limitation: the buffer and watcher state are per-process, in-memory. UnderGATEWAY_WORKERS>1or multi-pod, a follow-up comment routed to a different worker process than the one running the busy thread's agent will not see that buffer. This is a known, deliberately deferred limitation with the same shape as the cross-pod gap described for issue #4120 (a shared buffer store or IM-leader election would be needed to close it) — single-process/single-pod deployments, the safe default, see no correctness issue from this, only the documented per-process scope. - See backend/docs/GITHUB_AGENTS.md for the architecture diagrams: webhook → fan-out →
InboundMessagedispatch,preferred_thread_id = UUID5(repo, number, agent_name)thread determinism, mention-handle precedence chain, GH token lifecycle viaGH_TOKEN/GITHUB_TOKENper-callextra_env, and the narrowConflictError(HTTP 409) thread-create race recovery.