* fix: bound MCP server bring-up timeouts and exclude externalized tool outputs from delivery verification Two related robustness fixes: 1. MCP server bring-up was unbounded. tool_call_timeout only covered session.call_tool(); tool discovery (subprocess spawn + initialize + tools/list) and persistent stdio session initialization could hang forever, blocking agent construction (and on the Gateway event loop, the whole process). Add a per-server session_init_timeout (default DEFAULT_MCP_SESSION_INIT_TIMEOUT = 60s, null disables) that bounds both discovery and pooled-session initialization. The session pool's existing cancellation handling tears down a session stuck mid-creation in its own task. 2. ToolOutputBudgetMiddleware externalizes oversized tool outputs into outputs/.tool-results/ (configurable tool_output.storage_subdir). The workspace-change scanner and run delivery verification counted those files as produced artifacts, so any run that externalized a tool output without also presenting a real artifact failed with "Artifact delivery incomplete". Exclude TOOL_RESULTS_DIRNAME via a shared constant (mirroring BROWSER_FRAMES_DIRNAME) and thread the configured storage_subdir through snapshot capture so both workspace-changes events and delivery verification stay clean. * review: enforce single-segment tool_output.storage_subdir; document discovery-timeout cleanup Address review feedback: 1. A custom tool_output.storage_subdir with a path separator (e.g. cache/tool-results) silently no-oped the workspace-scanner exclusion: os.walk yields one-segment dirnames, so a nested value never matched and its files were counted as produced artifacts again. ToolOutputConfig now validates storage_subdir as a single directory name (rejects separators, .., absolute, empty) with tests, so the exclusion is always sound. 2. The discovery-timeout path now documents why cancellation is safe, mirroring the session-init note: discovery runs inside the adapter's nested async context managers, and stdio_client's finally terminates the process tree (SIGTERM->SIGKILL on POSIX, process-tree on Windows), so a timed-out npx subprocess and its children are reaped rather than accumulating. * review: log session-init timeouts and align API response model default with runtime config Address second-round review feedback: 1. A session-init timeout raised TimeoutError without any log, unlike the discovery timeout which logs a WARNING. Wrap the bounded get_session in a try/except that logs the timeout (server name + seconds) and re-raises, so operators can diagnose tool-call failures caused by hung MCP sessions. 2. McpServerConfigResponse.session_init_timeout defaulted to None while McpServerConfig defaults to 60s: a server created via PUT /api/mcp/config without the field was persisted with null (no timeout) while the same server created in the config file got 60s. Align the response-model default to DEFAULT_MCP_SESSION_INIT_TIMEOUT so API-created and file-created servers behave the same; an explicit null still opts out. * review: narrow the discovery-timeout handler to the bounded wait_for path The except TimeoutError clause covered both the bounded wait_for branch and the bare discovery branch. With session_init_timeout opted out (None), a TimeoutError raised by discovery itself would hit the %.1f format with None: logging raises TypeError internally, the WARNING is silently dropped, and a --- Logging error --- traceback goes to stderr. Narrow the handler to wrap only the wait_for call, where the branch condition guarantees the timeout value is not None. A discovery-internal TimeoutError on the opted-out path now falls through to the generic failure handler and is reported as 'tool discovery failed' with exc_info. Covered by a regression test that asserts the skip is reported without any broken format.
9.9 KiB
MCP (Model Context Protocol) Configuration
DeerFlow supports configurable MCP servers and skills to extend its capabilities, which are loaded from a dedicated extensions_config.json file in the project root directory.
Setup
-
Copy
extensions_config.example.jsontoextensions_config.jsonin the project root directory.# Copy example configuration cp extensions_config.example.json extensions_config.json -
Enable the desired MCP servers or skills by setting
"enabled": true. -
Configure each server’s command, arguments, and environment variables as needed.
-
Restart the application to load and register MCP tools.
Routing Hints
Use routing when an MCP server should be preferred for specific requests, such
as internal database questions that should use a PostgreSQL MCP tool before web
search. Routing hints are soft model guidance: they add a
<mcp_routing_hints> prompt section, but they do not forbid other tools. Use
agent-level allow/deny policy for hard restrictions. If tool_search.enabled
defers MCP tool schemas, matching routing metadata can also auto-promote the
deferred schema before the model call. Auto-promotion is controlled by the
top-level config.yaml -> tool_search.auto_promote_top_k setting.
{
"mcpServers": {
"postgres": {
"enabled": true,
"type": "stdio",
"command": "npx",
"args": ["-y", "@modelcontextprotocol/server-postgres", "postgresql://localhost/mydb"],
"routing": {
"mode": "prefer",
"priority": 50,
"keywords": ["orders", "users", "SQL", "database", "table"]
},
"tools": {
"query": {
"routing": {
"mode": "prefer",
"priority": 100,
"keywords": ["query database", "orders table", "metrics"]
}
}
}
}
}
}
routing.mode:offdisables hints;preferemits hints.routing.priority:0to100; higher-priority hints are rendered first. Whentool_search.enabled=true, priority also orders auto-promote matches.routing.keywords: operator-authored terms that describe when to prefer the MCP tool. Empty keywords are allowed but do not emit a hint line and do not trigger auto-promotion. Auto-promote matching is a case-insensitive substring test against the latest user message (not token/word-boundary matching), so prefer distinctive keywords — a short term likeapialso matchesrapid. Over-matching only exposes an extra tool schema (soft/additive), never disables other tools.tools.<original_tool_name>.routing: overrides only the fields explicitly set for that tool. The key is the MCP server's original tool name, before the<server>_prefix added for model binding. If the server-levelrouting.modeisoff, a tool override must setmode: "prefer"; setting onlypriorityorkeywordsstill inheritsoffand emits no hint.tool_search.auto_promote_top_k: global limit for auto-promoted deferred MCP schemas per model call. Default3; valid range1..5.
Tool Name Prefixes
DeerFlow prefixes discovered MCP tool names with <server_name>_ by default.
This avoids collisions when two enabled servers expose tools with the same
name. A server that already namespaces its own tools can opt out:
{
"mcpServers": {
"semantic-scholar": {
"type": "stdio",
"command": "uvx",
"args": ["s2-mcp-server"],
"tool_name_prefix": false
}
}
}
With this setting, a server tool named semantic_scholar_search_papers keeps
that name instead of becoming
semantic-scholar_semantic_scholar_search_papers. The default is true for
backward compatibility. Disable it only when every resulting tool name remains
unique across the enabled servers. Stdio tools continue to use DeerFlow's
persistent per-thread session pool regardless of this setting.
Server Timeouts (Stdio MCP Servers)
Two independent timeouts bound stdio MCP servers. session_init_timeout covers
server bring-up — tool discovery (subprocess spawn + initialize +
tools/list) and persistent-session initialization — and defaults to 60s so a
hung server (e.g. npx blocked on a package download, or a server that never
answers initialize) cannot block agent construction indefinitely. Set it to
null to disable:
{
"mcpServers": {
"github": {
"enabled": true,
"type": "stdio",
"command": "npx",
"args": ["-y", "@modelcontextprotocol/server-github"],
"env": {
"GITHUB_TOKEN": "$GITHUB_TOKEN"
},
"session_init_timeout": 60,
"tool_call_timeout": 60
}
}
}
tool_call_timeout limits each individual tool call in seconds and applies only
to stdio servers; http and sse servers use transport-level timeouts, and
DeerFlow logs a warning if tool_call_timeout is configured for those
transports.
Filesystem MCP Servers
DeerFlow already provides built-in file tools for thread-scoped workspace access. Do not add an MCP filesystem server for the same DeerFlow workspace. The overlapping file tools use different path semantics, which can make LLM tool selection and file access behavior unstable.
DeerFlow does not currently adapt the MCP Roots mode for filesystem servers. In
particular, it does not publish per-thread MCP roots or map DeerFlow sandbox
paths such as /mnt/user-data/... to paths accepted by
@modelcontextprotocol/server-filesystem. Use DeerFlow's built-in file tools
for DeerFlow workspace files.
OAuth Support (HTTP/SSE MCP Servers)
For http and sse MCP servers, DeerFlow supports OAuth token acquisition and automatic token refresh.
- Supported grants:
client_credentials,refresh_token - Configure per-server
oauthblock inextensions_config.json - Secrets should be provided via environment variables (for example:
$MCP_OAUTH_CLIENT_SECRET)
Example:
{
"mcpServers": {
"secure-http-server": {
"enabled": true,
"type": "http",
"url": "https://api.example.com/mcp",
"oauth": {
"enabled": true,
"token_url": "https://auth.example.com/oauth/token",
"grant_type": "client_credentials",
"client_id": "$MCP_OAUTH_CLIENT_ID",
"client_secret": "$MCP_OAUTH_CLIENT_SECRET",
"scope": "mcp.read",
"refresh_skew_seconds": 60
}
}
}
}
Custom Tool Interceptors
You can register custom interceptors that run before every MCP tool call. This is useful for injecting per-request headers (e.g., user auth tokens from the LangGraph execution context), logging, or metrics.
Declare interceptors in extensions_config.json using the mcpInterceptors field:
{
"mcpInterceptors": [
"my_package.mcp.auth:build_auth_interceptor"
],
"mcpServers": { ... }
}
Each entry is a Python import path in module:variable format (resolved via resolve_variable). The variable must be a no-arg builder function that returns an async interceptor compatible with MultiServerMCPClient’s tool_interceptors interface, or None to skip.
Example interceptor that injects an authorization header from the request-scoped LangGraph secret context:
from langgraph.config import get_config
def build_auth_interceptor():
async def interceptor(request, handler):
config = get_config()
secrets = (config.get("context") or {}).get("secrets") or {}
token = secrets.get("MCP_AUTH_TOKEN")
if token:
request = request.override(
headers={**(request.headers or {}), "Authorization": f"Bearer {token}"}
)
return await handler(request)
return interceptor
Supply the credential on each run request through config.context.secrets:
{
"metadata": {"source": "my-client"},
"config": {
"context": {
"secrets": {"MCP_AUTH_TOKEN": "<request-scoped credential>"}
}
}
}
Both metadata.auth_token and config.metadata.auth_token are rejected with HTTP 422 at run admission and are never supported
interceptor paths. Do not put credentials in either metadata surface; use
config.context.secrets, whose values remain available to the live interceptor
but are removed from persisted and API-visible run configuration copies.
- A single string value is accepted and normalized to a one-element list.
- Invalid paths or builder failures are logged as warnings without blocking other interceptors.
- The builder return value must be
callable; non-callable values are skipped with a warning.
Migrating legacy MCP credentials
Deployments that previously sent metadata.auth_token or config.metadata.auth_token must:
- Update the caller and interceptor to use
config.context.secretsas shown above. - Rotate the exposed credential before resuming authenticated MCP traffic.
- Locate and remove every retained legacy copy according to the deployment's retention policy, including database rows, run events, application or proxy logs, snapshots, exports, and backups.
Current history APIs hide legacy metadata.auth_token and config.metadata.auth_token values, but hiding a response does not erase
material already retained by those systems. Restarting or upgrading DeerFlow does
not rotate credentials or perform historical cleanup; operators must complete
both actions explicitly.
How It Works
MCP servers expose tools that are automatically discovered and integrated into DeerFlow’s agent system at runtime. Once enabled, these tools become available to agents without additional code changes.
Example Capabilities
MCP servers can provide access to:
- Databases (e.g., PostgreSQL)
- External APIs (e.g., GitHub, Brave Search)
- Browser automation (e.g., Puppeteer)
- Custom MCP server implementations
Learn More
For detailed documentation about the Model Context Protocol, visit:
https://modelcontextprotocol.io