Hyeonsang Cho 8e86729aa0
fix(gateway): confine artifact PUT to /mnt/user-data/outputs after path resolution (#5321)
* fix(gateway): confine artifact PUT to /mnt/user-data/outputs after path resolution

The outputs-only guard on PUT /api/threads/{id}/artifacts/{path} was a
string-prefix check on the raw path. A percent-encoded `..`
(`outputs/%2e%2e/uploads/x.txt`) survives nginx's variable proxy_pass
untouched, is decoded by Starlette, passes the prefix check, and the
resolver only confines the result to `user-data/` -- so an owner could
overwrite a sibling upload or workspace file in their own thread.

Collapse dot segments before the prefix check, and re-check the resolved
host path against the resolved outputs root so a symlink planted inside
`outputs/` cannot redirect the write either. The normalized virtual path
is what the response echoes and what non-mounted sandboxes receive.

* refactor(gateway): share the outputs-confinement rule with channel attachments

Review follow-up on #5321: the "only under /mnt/user-data/outputs" rule was
implemented independently by the artifact editor and by IM-channel
attachment delivery, and the two copies had already drifted.

Move it into app/gateway/path_utils.py as normalize_outputs_virtual_path
(collapse `..` before the prefix check) and resolve_outputs_confined_path
(re-check the resolved host path against the resolved outputs root, which
also catches a symlink planted inside outputs/). PUT /artifacts and
ChannelManager._resolve_attachments both call the helper; artifact_archive
keeps its stricter ZIP-member rules layered on top.

Tests that previously stubbed resolve_thread_virtual_path for the editor now
stub resolve_outputs_confined_path, and the channel attachment tests patch
path_utils.get_paths, which the helper binds at import like the other
consumers. The confinement itself is pinned by tests/test_gateway_path_utils.py.
2026-09-10 15:01:56 +08:00

36 KiB
Raw Blame History

IM Channels System (app/channels/)

Bridges external messaging platforms (Feishu, Slack, Telegram, Discord, DingTalk, GitHub) to the DeerFlow agent via Gateway's LangGraph-compatible API.

Architecture: Channels communicate with Gateway through the langgraph-sdk HTTP client (same as the frontend), ensuring threads are created and managed server-side. The internal SDK client injects process-local internal auth plus a matching CSRF cookie/header pair so Gateway accepts state-changing thread/run requests from channel workers without relying on browser session cookies.

Components:

  • message_bus.py - Async pub/sub hub (InboundMessage → queue → dispatcher; OutboundMessage → callbacks → channels)
  • store.py - JSON-file persistence mapping channel_name:chat_id[:topic_id]thread_id (keys are channel:chat for root conversations and channel:chat:topic for threaded conversations). Every access to _data must be protected by _lock; list_entries() snapshots keys and copied entries under the lock, then formats the result after releasing it so concurrent channel threads cannot resize the dictionary during iteration without extending the critical section.
  • manager.py - Core dispatcher: creates threads via client.threads.create(), routes commands including /goal (setting a goal persists it through Gateway and then routes the objective as a chat turn) and /agent (list is owner-scoped; use creates a fresh thread and persists the Custom Agent selection under both the channel restart key and Web's canonical routing metadata so existing checkpoint lineage never changes agent across IM/Web), keeps Slack/Discord on client.runs.wait(), uses client.runs.stream(["messages-tuple", "values"]) for Feishu/Telegram incremental outbound updates, serializes same-thread Feishu turns in-manager when the channel's ChannelRunPolicy.serialize_thread_runs=True so rapid follow-ups queue instead of tripping the runtime busy reply, and switches to client.runs.create() (fire-and-forget, returns once the run is pending) for channels whose ChannelRunPolicy.fire_and_forget=True so long autonomous runs do not hit the SDK default 300s httpx.ReadTimeout A swallowed streaming failure publishes its final outbound before releasing the inbound dedupe key, so a provider redelivery can retry without overtaking the terminal reply. What may be published from the stream is an allowlist, not a denylist (_accumulate_stream_text): only assistant message types — LangChain serializes AIMessage.type as "ai" and AIMessageChunk.type as "AIMessageChunk", plus the OpenAI-style "assistant" spelling for foreign runtimes — become displayable text. The previous rule rejected only payloads whose type contained "tool" and therefore published everything else, which leaked DeerFlow's hidden model context to every streaming IM channel: DynamicContextMiddleware injects recalled memory as a hidden HumanMessage (type == "human") and rewrites the user's own turn into a new HumanMessage, DurableContextMiddleware injects a hidden <durable_context_data> HumanMessage, and LangGraph fans those state writes out on the messages-tuple stream. Proved live on a Buzz relay, which published a <memory> fact block and, in another run, a verbatim echo of the user's own message as the assistant's reply. Matching is by prefix (ai / assistant), never substring, because ordinary words contain "ai" (chain, domain). The message type is resolved by _stream_payload_type, which handles both the model_dump() shape DeerFlow's own gateway emits and LangChain's to_json() constructor shape (whose top-level type is the literal "constructor", with the class name at the tail of the id path). A bare str payload is no longer accepted at all: it carries no type information, so it cannot be attributed to the assistant, and nothing in DeerFlow produces one (runtime/serialization.py::serialize_messages_tuple always emits [message_dict, metadata]).
  • base.py - Abstract Channel base class (start/stop/send lifecycle). Provider callbacks that submit coroutines from SDK threads must use _submit_threadsafe_coroutine(): it creates and retains the real asyncio.Task on the owner loop instead of treating run_coroutine_threadsafe()'s proxy Future as a completion signal. Submission is closed atomically with shutdown, and stop() must call _close_and_drain_threadsafe_futures() before tearing down SDK resources.
  • service.py - Manages lifecycle of all configured channels from config.yaml. Shutdown closes manager admission first and keeps transports alive until every manager worker/follow-up watcher has exited. A successful manager stop therefore owns no live handler; if the Gateway's outer timeout cancels shutdown, the service retains its channel objects and global singleton so unfinished resources are not detached and cleanup can be retried.
  • slack.py / feishu.py / telegram.py / discord.py / dingtalk.py - Platform-specific implementations (feishu.py tracks the running card message_id in memory and patches the same card in place; telegram.py accepts inbound text/photos/documents, preserves media captions, hands token-free attachment bytes to the shared upload pipeline, edits the "Working on it..." stream target in place via editMessageText, and can optionally send final Markdown replies as Rich Messages through channels.telegram.rich_messages; discord.py registers typing loops before inbound handling yields and _start_typing() refuses work once _running is false; because stop() runs on the main loop while typing tasks belong to _discord_loop, normal cross-thread cancellation, awaiting, and map cleanup are scheduled there with a bounded wait, while _run_client() drains tasks in its finally block before an exception/disconnect can make that loop unusable; an already-stopped foreign loop must never have its tasks awaited from the main loop; to that end every outbound cross-loop call (send / send_file / _get_channel_or_thread) goes through _run_on_discord_loop, which bounds the await (DISCORD_OUTBOUND_TIMEOUT_SECONDS 30s; DISCORD_UPLOAD_TIMEOUT_SECONDS 120s for file uploads, whose unbounded-size payload needs room for a slow uplink plus 429 retry-after) and fails fast with a RuntimeError when the client loop is missing or not running — a dead client becomes a logged send failure instead of a permanently hung ChannelManager worker — and is_running reports client-thread aliveness (like feishu.py) so ensure_channel_ready can restart the channel after _run_client() exits on a fatal error; dingtalk.py optionally uses AI Card streaming for in-place updates when card_template_id is configured, and overrides receive_file to download inbound images (picture/richText) and documents (file) by downloadCode into the thread uploads bucket, mirroring feishu.py)
  • buzz.py - Buzz (Nostr relay) implementation: one NIP-42-authenticated websocket, pubkey-allowlist + mention/DM/thread-follow gating, streaming replies via in-place kind-40003 edits; requires the buzz dependency extra. Its durable seen-event replay guard uses async aseen() / arecord() / aflush() boundaries: initial JSON reads and coalesced atomic writes run off the Gateway event loop, and stop() awaits the final flush before returning. Subscription model (operator-facing version in IM_CHANNEL_CONNECTIONS.md): the relay fans kind-9 chat events out only to channel-scoped subscriptions, proved against a live relay — REQ {"kinds":[9]} is accepted and answered with EOSE but never receives an event (the connector authenticated and then sat silent forever), REQ {"kinds":[9],"#h":[uuid]} works, and a multi-value #h matches nothing, so it is strictly one REQ per channel (the same shape as Buzz's own buzz-acp harness). Every connection therefore rebuilds three kinds of subscription after NIP-42 auth: buzz-discovery ({"kinds":[39000]}, a historical query returning exactly the channels this identity belongs to, one stored event each then EOSE — adding #p returns zero, do not "narrow" it), buzz-membership ({"kinds":[44100,44101],"#p":[us],"since":<connect time MEMBERSHIP_LOOKBACK_SECONDS>}, the relay-signed member-added/member-removed notifications whose p tag names the affected member and h tag the channel), and one buzz-chat-<uuid> per discovered channel. Chat subscriptions open as each kind-39000 arrives; the discovery EOSE is the completeness barrier that retries any that failed and warns when discovery found nothing. A kind-44100 for our pubkey subscribes to the new channel live (then re-issues discovery so its name/type reach the DM-detection cache, but only when that channel's metadata is actually missing — an unconditional refresh is how a burst of 44100s multiplied into one discovery pass each); a kind-44101 issues buzz_nostr.close_frame for exactly that channel's subscription and drops its metadata. The membership subscription is scoped to LIVE events and this is load-bearing: buzz-relay stores 44100/44101 and serves history newest-first (default limit 2000), so an unscoped filter replayed the whole membership history on every connect — every stored add read as live, re-running discovery once each (M+1 discovery passes × N stored kind-39000 events per connect, observed live as two channel discovery complete lines and one channel logged <unnamed> because the 44100 path subscribed before its metadata arrived), re-subscribing channels we have since been removed from, and letting a stored 44101 transiently unsubscribe a channel we are still in. since is anchored at the moment the socket opened (_session_started_at) minus MEMBERSHIP_LOOKBACK_SECONDS (60) of slack, which covers both relay clock skew and a membership change published during the connect/auth handshake; the slack can only cost an idempotent replay of the last minute. A relay CLOSED frame is recovered, not merely forgotten — every subscription on the socket fails silently when dropped, so _handle_closed re-issues it, bounded by MAX_RESUBSCRIBE_ATTEMPTS (3) per subscription id per connection and per auth epoch (_resubscribe_attempts is reset on session start, on re-auth, and by stop(), so pre-auth rejections — which the auth branch already recovers wholesale — never spend the authenticated session's budget). The first retry is immediate (the common case is a one-off hiccup); later ones back off 1s then 2s, awaited inline in the read loop rather than in a background task that could outlive its own socket. auth-required: before this socket has completed its NIP-42 handshake is the expected bootstrap sequence, not a refusal, and _handle_closed short-circuits it ahead of every recovery path: the connector opens its control REQs immediately in case the relay serves unauthenticated reads, a closed relay answers auth-required: plus an AUTH challenge, and the auth branch re-opens everything. That case is logged at DEBUG and consumes neither the permanent-refusal branch nor the retry budget (a chat subscription is still dropped from _chat_subscriptions, since it genuinely is not subscribed and discovery is what re-opens it). Treating it as permanent produced an operator-facing warning claiming discovery/membership tracking was DOWN in the same run where discovery then completed, every channel was subscribed, and a brand-new channel's kind-44100 was picked up one second later. The boundary is the per-socket _auth_completed flag, set once the signed AUTH event has been sent and cleared on session entry, session exit, and stop(); the same reason after that stays loud, and says the subscription is down until the relay's next AUTH challenge or the next reconnect rather than borrowing the non-auth wording. Then _is_transient_close decides whether to retry at all: NIP-01/NIP-42 auth-required:/restricted:/blocked:/mute:/invalid:/pow: prefixes and buzz-relay's own removal/revocation prose are permanent (do not fight the relay over a channel that is no longer ours; a post-auth auth-required: is recovered by the AUTH branch, not by re-issuing), rate-limited:/error:/no-reason-at-all and anything unrecognized are transient — the default resolves toward "keep listening" because going silently deaf is the failure this exists to remove, and the attempt budget bounds a wrong guess. A chat CLOSED is only ever recovered for a channel already in _chat_subscriptions: a CLOSED is relay-supplied, so acting on an unknown one would let a relay induce a subscription just by naming a channel. Every subscription that goes unlistened is logged at WARNING, never INFO. Subscription ids are deterministic per channel precisely so one can be replaced or closed without disturbing the others on the socket, and _chat_subscriptions is per-socket state cleared on session end, on re-auth (a pre-auth REQ may have been rejected), and by stop(). Three bounds on remote-fed state: MAX_CACHED_CHANNELS (512) caps the kind-39000 metadata cache, the watermark map, and the resubscribe-attempt map; MAX_CHANNEL_SUBSCRIPTIONS (256, well under buzz-relay's own 1024-per-connection ceiling) caps live chat subscriptions — at the cap new channels are refused and named in a warning rather than evicting a working subscription. Known bound (documented, not fixed): the relay caps historical delivery at 2000 events per subscription, newest-first, even with a since, so >2000 unread messages in a single channel across a disconnect loses the oldest — the relay never sends them and the watermark advances past them. That is the one remaining path that can skip; everything else fails toward replay. Trust model (operator-facing version in IM_CHANNEL_CONNECTIONS.md): every inbound EVENT is authenticated at the single handle_relay_frame choke point — the NIP-01 id is recomputed from the delivered payload and the BIP-340 Schnorr signature verified against the claimed pubkey (buzz_nostr.verify_event, pure and total: malformed input returns False, never raises) — so ev["pubkey"], the authorization principal for both the allowlist and the /connect bind, cannot be forged by a relay the DeerFlow operator does not run. What remains trusted is the authorship of kind-39000 channel metadata: any member can sign one, and because per-channel subscriptions are now driven by discovery, a forged kind-39000 has two effects rather than one — it can mark a channel type: "dm" (relaxing require_mention for that channel) and it can induce a chat subscription for a channel of the forger's choosing, since the channels we listen to are exactly the channels we hold metadata for. Neither makes anything be acted on: allowed_users and per-event signature verification are independent gates, so an induced subscription only means the relay reads its own traffic back to a subscriber that drops it, bounded by MAX_CHANNEL_SUBSCRIPTIONS (which refuses rather than evicts, so it cannot displace a real channel). Same for a forged kind-44100, except its p tag is re-checked locally so it must at least name us. Closing this needs a configured trusted relay pubkey, which relay_url is not. allowed_users is deny-by-default (empty = nobody, unlike siblings' empty = everyone), so start() logs a WARNING when it is empty and each drop logs at DEBUG. The resubscribe cursor (since) is per channel, advances only for events that were actually processed, and never past now + MAX_FUTURE_SKEW_SECONDS, because it is peer-supplied (created_at) and a single future-dated event otherwise made the connector permanently deaf. Per channel rather than global is the safety-critical half: subscriptions are per channel, so one shared cursor is the newest event seen in any channel and a busy channel would drag it past a quiet channel's unread messages, skipping them on the next reconnect — measured on a live relay, three channels of one identity sat ~28h apart. Per-channel cursors can only ever cost duplicate delivery (absorbed by the manager's event_id dedupe), and an evicted cursor degrades to "no since", i.e. the relay's default backlog — both fail toward replay, never toward a miss. Streaming tracks every oversize chunk index (_stream_targets for chunk 0, _stream_tails for the rest), since the manager republishes cumulative text and reposting chunks[1:] per update flooded the channel; all of it is per-connection state cleared by stop(), and the remote-fed kind-39000 cache is capped at MAX_CACHED_CHANNELS. send() refuses outright to publish text carrying a hidden model-context wrapper (<memory>, <durable_context_data>, <system-reminder>_HIDDEN_CONTEXT_MARKERS), logging at ERROR and clearing the stream bookkeeping on a blocked is_final. This is defense in depth behind the manager's allowlist, and it lives here rather than in a sibling connector because on Buzz a leak is permanent: every streaming update is an immutable public Nostr event, so a corrective edit only changes what clients render while the original leaked event stays on the relay. Matching is on the literal opening tag, so a reply that merely talks about memory is still published Buzz seen-event shutdown stays bounded and retryable: aflush() awaits any in-flight write, attempts at most one final snapshot, leaves a still-changing generation dirty for fail-open replay, and returns without a live persistence timer. A Gateway cancellation does not cancel the underlying worker-thread write; BuzzChannel.stop() tracks cleanup completion separately from transport admission so ChannelService can retry the retained channel without racing a second snapshot against that write. The store is quiesced before stop (including the already-stopped guard), so a timed-out relay task that records after stop only marks data dirty and cannot schedule detached file work; a repeated stop() still drains that dirty state, while BuzzChannel.start() explicitly resumes scheduling and flushes it automatically.
  • github.py - Webhook-driven GitHub channel. Inbound messages come from POST /api/webhooks/github; outbound is log-only because GitHub agents post explicitly with gh from their sandbox when they choose to comment or create a PR
  • app/gateway/routers/channel_connections.py - Browser-facing user connection and disconnect APIs
  • deerflow.persistence.channel_connections - SQL-backed user-owned connection, optional credential, connect state, and conversation store

Message Flow:

  1. External platform -> Channel impl -> MessageBus.publish_inbound()
    • For GitHub, the webhook router verifies the delivery then calls fanout_event(bus, ...); matching agent bindings publish one InboundMessage each instead of a long-polling channel worker.
    • Telegram photo/document updates use the largest photo size or document metadata, preserve message.caption, enforce the hosted Bot API's 20,000,000-byte download ceiling before and after download, and never expose the token-bearing Bot API file URL. Downloaded bytes cross the adapter/manager boundary only through message_bus.INBOUND_FILE_CONTENT_KEY; the manager consumes that transient field before persisting safe upload metadata.
    • Feishu/Lark inbound image/file downloads read at most 20,000,001 bytes and reject anything above 20,000,000 bytes before persistence or sandbox sync. Oversize, unsafe-path, and path-resolution failures rewrite only that attachment placeholder to Failed to obtain the [type] so later attachments in the same message can still load.
  2. ChannelManager._dispatch_loop() consumes from queue
  3. For user-owned channel connections, incoming messages carry connection_id, owner_user_id, and workspace_id; owner_user_id becomes the DeerFlow run user_id, while the raw platform user id remains channel_user_id. The Gateway accepts channel_user_id only from an internally authenticated channel caller's top-level body.context, clears it from both free-form body.config sections, and writes it into runtime context only (never configurable, which is checkpointed). bash_tool exposes it to sandbox commands as the fixed env var DEERFLOW_CHANNEL_USER_ID — via a shell-quoted command-string prefix, NOT the execute_command(env=...) channel, which is reserved for request-scoped secrets and would switch AioSandbox onto the bash.exec path (image >= 1.9.3, fresh session per call). Per-call injection keeps group-chat identity correct (one thread/sandbox, many senders) without depending on the AIO shell's session semantics: every IM-channel command carries an explicit export VAR=<id>; (valid id) or unset VAR; (empty / non-str / over the 256-char cap). The AIO no-env path reuses a persistent shell session (the reason for the class lock, #1433), so a bare command could otherwise resolve a stale id an earlier sender exported; the unset closes the window the length/type guard would open (a dropped id would inherit the previous sender's value). Non-IM runs (no channel_user_id in context) are left untouched. Not injected on the Windows local sandbox (its PowerShell/cmd.exe fallback has no export/unset). Propagates across task delegation: task_tool captures the dispatching turn's id and the subagent executor forwards it into the subagent's runtime context, same as the guardrail attribution fields. The runtime-context value is authorization-grade at the Gateway/guardrail boundary, but the exported shell variable remains informational because any bash command can overwrite its own environment; skills must not treat the shell variable itself as authenticated identity. Tests: tests/test_gateway_services.py, tests/test_channel_user_id_env.py
  4. For chat: look up/create thread through Gateway's LangGraph-compatible API
  5. Feishu/Telegram chat: runs.stream() → accumulate AI text → publish multiple outbound updates (is_final=False) → publish final outbound (is_final=True)
  6. Slack/Discord chat: runs.wait() → extract final response → publish outbound 6b. GitHub chat (ChannelRunPolicy.fire_and_forget=True): runs.create() returns once the run is pending; the manager does not wait for the final state and does not publish an outbound. The agent posts its own reply mid-run via gh from the sandbox. ConflictError on a busy thread still trips the standard THREAD_BUSY_MESSAGE path (log-only on GitHub); when the channel's policy also sets buffer_followups_on_busy=True (GitHub's default — see "Follow-up buffering while busy" below), the triggering message is additionally captured into a per-thread buffer instead of only logged, so a concurrent comment is not silently dropped.
  7. Feishu channel sends one running reply card up front, then patches the same card for each outbound update (card JSON sets config.update_multi=true for Feishu's patch API requirement). Messages already sent inside an existing Feishu topic carry a compact source-message preview in that card, and queued same-thread follow-ups patch their own source message's card from queued → running → final without falling back to the generic busy reply.
  8. Telegram streaming: the "Working on it..." placeholder message is registered as the stream target; non-final updates editMessageText it in place (channel-side throttle: 1s in private chats, 3s in groups due to Telegram's 20 msg/min group cap; 4096-char truncation; rate-limited updates dropped); the final update performs the last edit and splits >4096 texts into follow-up messages
  9. DingTalk AI Card mode (when card_template_id configured): runs.stream() → create card with initial text → stream updates via PUT /v1.0/card/streaming → finalize on is_final=True. Falls back to sampleMarkdown if card creation or streaming fails
  10. For commands (/new, /status, /models, /memory, /goal, /agent, /help): handle locally or query Gateway API. /agent list reads only the effective owner's Custom Agents; /agent use <name> validates in the same owner bucket, creates a new thread, and persists channel_agent_name in its metadata. A Custom Agent also writes canonical metadata.agent_name, which makes thread-search results route Web continuation through /workspace/agents/<name>/chats/<thread_id>; lead_agent deliberately omits that canonical key and stays on the ordinary chat route. The manager caches channel_agent_name for the hot path and reloads it after restart before routing a resumed turn. An explicit selection also normalizes agent_name across top-level run context plus the existing RunnableConfig context and configurable carriers (or clears all three for lead_agent) before Gateway's setdefault compatibility merge, so stale channel defaults cannot silently win.
  11. Outbound → channel callbacks → platform reply
    • GitHub is the exception: the channel logs the final assistant message and does not auto-post it to GitHub. Agents use the sandbox gh CLI (gh issue comment, gh pr comment, gh pr create, etc.) for intentional writeback, so silence is cheap when several agents fan out on the same event.

Owner-scoped file storage: inbound files, uploads, and output artifacts are staged under the DeerFlow owner's bucket so they land where the agent run reads/writes (users/{user_id}/threads/{thread_id}/user-data/{uploads,outputs}). ChannelManager._handle_chat resolves the storage owner once via _channel_storage_user_id(msg) (sanitized owner id, falling back to safe(msg.user_id) for unbound auth-enabled channels — mirroring _resolve_run_params's run identity; None only when no identity is available) and threads it as the user_id= kwarg through the file pipeline:

  • Channel.receive_file(msg, thread_id, user_id=...) — owner-bound channels persist downloaded files under the owner's bucket instead of the default bucket
  • FeishuChannel._receive_single_file(...) / DingTalkChannel._receive_single_file(...) — normalize provider filenames, claim a collision-free basename and write it through write_upload_file_no_symlink under the same channel lock; the returned basename drives both the agent-visible virtual path and non-local sandbox sync
  • sandbox_files.py — non-mounted Feishu/DingTalk syncs acquire unique non-releasing execution holders, drain blocking update_file workers across repeated cancellation, and release only after the last sandbox operation, so a parallel run cannot close the shared client mid-upload
  • _ingest_inbound_files(...) and the underlying ensure_uploads_dir / get_uploads_dir — owner-scoped via the same kwarg
  • _resolve_attachments / _prepare_artifact_delivery — resolve output artifacts from the bound owner's bucket through app.gateway.path_utils.resolve_outputs_confined_path, the same outputs-only rule the artifact editor uses, so a sibling uploads//workspace/ path or a symlink planted in outputs/ is skipped with a warning The cached value is reused for both the blocking (runs.wait) and streaming (_handle_streaming_chat) paths, so uploads and artifact delivery always target the same bucket even if a channel returns a rewritten InboundMessage from receive_file. The bucket id matches the memory bucket resolved by _resolve_memory_user_id (both normalize through make_safe_user_id).

Configuration (config.yaml -> channels):

  • langgraph_url - LangGraph-compatible Gateway API base URL (default: http://localhost:8001/api)
  • gateway_url - Gateway API URL for auxiliary commands (default: http://localhost:8001)
  • In Docker Compose, IM channels run inside the gateway container, so localhost points back to that container. Use http://gateway:8001/api for langgraph_url and http://gateway:8001 for gateway_url, or set DEER_FLOW_CHANNELS_LANGGRAPH_URL / DEER_FLOW_CHANNELS_GATEWAY_URL.
  • Per-channel configs: feishu (app_id, app_secret), slack (bot_token, app_token), telegram (bot_token, optional rich_messages for final Markdown Rich Messages), dingtalk (client_id, client_secret, optional card_template_id for AI Card streaming), github (operator kill-switch enabled, plus default_mention_login for mention-required GitHub triggers), buzz (relay_url, private_key)

User-owned channel connections (config.yaml -> channel_connections):

  • Disabled by default. It is a user-binding layer on top of the existing channels.* runtime config, not a replacement for provider bot credentials.
  • No public IP, OAuth callback URL, or provider webhook route is required by the current implementation.
  • Telegram uses a deep-link /start <code> flow over the existing long-polling worker. Slack, Discord, Feishu/Lark, DingTalk, WeChat, and WeCom use /connect <code> over their existing outbound channel workers.
  • WeChat timing settings (polling_timeout, polling_retry_delay, qrcode_poll_interval, qrcode_poll_timeout) accept only positive finite seconds; invalid values fall back to their defaults so polling cannot enter a hot loop or sleep forever.
  • WeCom serializes start() and stop() for each channel instance. The SDK connect() task covers connection setup only; after the handshake, the SDK owns a separate receive task. Shutdown cancels an in-progress connection attempt and awaits the SDK's actual asynchronous receive-task/socket cleanup before releasing lifecycle state or allowing a restart. Cancellation of stop() still propagates, but only after owned cleanup finishes and lifecycle references are cleared; real connection failures remain reported by _on_ws_task_done.
  • Frontend APIs: GET /api/channels/providers, GET /api/channels/connections, POST /api/channels/{provider}/connect, and DELETE /api/channels/connections/{connection_id}.
  • Browser APIs remain protected by normal Gateway auth/CSRF. Provider messages arrive through the already-configured channel workers.
  • Provider-level connection_status reflects the user's newest connection row. With no binding it is not_connected, except in auth-disabled local mode where a configured running channel reports connected because all channel messages already route to the default user.
  • Slack replies use the configured operator bot token from channels.slack unless per-connection credentials are present; unreadable or corrupt stored credentials are treated as unavailable.
  • Telegram, Slack, Discord, Feishu/Lark, DingTalk, WeChat, and WeCom workers resolve incoming platform identities to connection records before reaching ChannelManager.
  • Connect-code ordering vs allowed_users: inbound workers consume a valid /connect <code> (or Telegram /start <code>) before applying the allowed_users filter, so a newly allowlisted-but-unbound user can bootstrap their first bind via the browser flow. Consequence: allowed_users is not a bind-time defense — any sender who possesses a valid code can consume it (not only allowlisted users). The bind security model rests on the code's confidentiality: secrets.token_urlsafe(16), 600 s TTL, one-time consume_oauth_state, and codes surfaced only in the initiating browser (never echoed to chat). allowed_users still gates ordinary (non-bind) messages.
  • Single-active-owner transfer semantics: an external identity is keyed by (provider, external_account_id, workspace_id). The latest successful bind wins — upsert_connection revokes other owners' active rows for the same identity (ownership transfer). This invariant is enforced at the DB layer by the partial unique index uq_channel_connection_active_identity (WHERE status != 'revoked'), so concurrent connects from different owners cannot both end connected; the losing writer retries against the now-visible state. find_connection_by_external_identity therefore resolves deterministically.
  • See backend/docs/IM_CHANNEL_CONNECTIONS.md for provider setup, operational notes, and the architecture diagrams (connect-code flow, single-active-owner transfer, sync vs streaming dispatch, owner-scoped file storage pipeline).

GitHub event-driven agents (webhook-driven IM channel):

  • Custom agents declare a github: block in their config.yaml to bind to repos and event triggers; the webhook route is fail-closed by default (mounted only when GITHUB_WEBHOOK_SECRET is set) and exempt from auth/CSRF because authenticity is enforced by HMAC.
  • Registry caching is keyed by the configured agent store's opaque signature: file storage uses agent config mtimes, while database storage hashes the ordered owner/name/config/soul contents so same-timestamp writes still invalidate webhook routing.
  • Outbound is log-only by design: each agent posts its own reply mid-run via the gh CLI from its sandbox, so the manager uses fire_and_forget=True and runs.create() returns once pending.
  • Follow-up buffering while busy (issue #4121): because outbound is log-only, the pre-existing THREAD_BUSY_MESSAGE reply on a ConflictError was invisible to the commenter — a comment posted while a run was already active looked like it had been silently ignored. When ChannelRunPolicy.buffer_followups_on_busy=True (GitHub's default), a ConflictError on runs.create() now also appends the triggering message to a per-thread, in-memory buffer (ChannelManager._followup_buffers) — deduped by GitHub delivery id, capped at FOLLOWUP_BUFFER_MAX_PER_THREAD (20, oldest dropped with a WARNING log on overflow). The first successful runs.create() on a thread now captures its run_id and spawns a background watcher that subscribes to that run's StreamBridge stream; once it observes END_SENTINEL, the watcher drains up to FOLLOWUP_DRAIN_BATCH_SIZE (10) buffered entries into one <followups-while-busy>-wrapped input and fires a follow-up runs.create() — itself watched the same way, so a backlog deeper than one batch chains into further drain cycles instead of growing one unbounded input. If that follow-up runs.create() itself hits ConflictError (e.g. a manual Web UI turn or a scheduled run raced onto the same thread), the batch is requeued rather than lost, and is retried whenever this manager next successfully creates and watches a run on that thread. Reactions/acknowledgment (e.g. GitHub's eyes/confused reaction API) on buffered comments are intentionally out of scope for this mechanism and left to a follow-up — comments are coalesced silently. Plumbing: the watcher needs the Gateway's StreamBridge singleton, which ChannelManager did not previously have access to; it is threaded from app.py's lifespan (where app.state.stream_bridge is already set by langgraph_runtime) through start_channel_service(get_stream_bridge=...)ChannelService.__init__ChannelManager.__init__, as a zero-arg closure mirroring the existing launch_run=lambda **kwargs: launch_scheduled_thread_run(app=app, **kwargs) pattern used for ScheduledTaskService in the same lifespan function. A ChannelManager constructed without it (e.g. directly in a test) still buffers safely — it just has no watcher to auto-drain. Scope limitation: the buffer and watcher state are per-process, in-memory. Under GATEWAY_WORKERS>1 or multi-pod, a follow-up comment routed to a different worker process than the one running the busy thread's agent will not see that buffer. This is a known, deliberately deferred limitation with the same shape as the cross-pod gap described for issue #4120 (a shared buffer store or IM-leader election would be needed to close it) — single-process/single-pod deployments, the safe default, see no correctness issue from this, only the documented per-process scope.
  • See backend/docs/GITHUB_AGENTS.md for the architecture diagrams: webhook → fan-out → InboundMessage dispatch, preferred_thread_id = UUID5(repo, number, agent_name) thread determinism, mention-handle precedence chain, GH token lifecycle via GH_TOKEN/GITHUB_TOKEN per-call extra_env, and the narrow ConflictError (HTTP 409) thread-create race recovery.