* feat(skills): export custom skill packages with revision preview * docs(gateway): keep export guidance within size budget * ci: retry checks after transient uv setup download failure * docs: focus skill export agent guidance on maintenance invariants * fix(skills): handle export disconnects and bound archive transfers * docs(gateway): remove redundant export guidance to fit merged budget * fix(skills): reset export idle deadline after transfer progress
41 KiB
API Reference
This document provides a complete reference for the DeerFlow backend APIs.
Overview
DeerFlow backend exposes two sets of APIs:
- LangGraph-compatible API - Agent interactions, threads, and streaming (
/api/langgraph/*) - Gateway API - Models, MCP, skills, uploads, and artifacts (
/api/*)
All APIs are accessed through the Nginx reverse proxy at port 2026.
For agent conversations, clients can either pre-create a thread
(POST /api/langgraph/threads) or start immediately with the stateless stream
endpoint (POST /api/langgraph/runs/stream). The latter auto-creates a thread
and returns thread_id and run_id in the response Content-Location header.
Authentication
Browser sessions authenticate with the access_token session cookie issued at
login. Programmatic clients can instead use a personal access token (PAT)
sent as a Bearer credential:
POST /api/threads/search
Authorization: Bearer dfp_...
Content-Type: application/json
{}
PATs require a configured database backend (SQLite/PostgreSQL) — on the
memory-only backend, Bearer credentials are rejected and PAT management routes
return 503.
Personal Access Tokens
Base URL: /api/v1/auth
PAT management requires an interactive session (a PAT cannot manage PATs or change passwords, so a leaked automation token cannot mint fresh credentials). The raw token is returned exactly once at creation; only its SHA-256 digest is stored server-side.
Create Token
POST /api/v1/auth/pats
Content-Type: application/json
Request Body:
{
"name": "ci-runner",
"scopes": ["threads:read", "runs:create", "runs:read"],
"expires_in_days": 90
}
scopes— subset of the route permissions:threads:read,threads:write,threads:delete,runs:create,runs:read,runs:cancel. A PAT can only narrow its owning user's permissions, never widen them.expires_in_days— optional (1–365); omitted means the token never expires.
Response (201):
{
"id": "0f0c6e6a-...",
"name": "ci-runner",
"scopes": ["runs:create", "runs:read", "threads:read"],
"expires_at": "2026-11-25T10:30:00Z",
"created_at": "2026-08-27T10:30:00Z",
"token": "dfp_..."
}
Save token immediately — it cannot be retrieved again.
List Tokens
GET /api/v1/auth/pats
Returns the caller's tokens with last_used_at / revoked_at audit fields;
never returns digests or raw tokens.
Revoke Token
DELETE /api/v1/auth/pats/{pat_id}
Revocation is immediate.
PAT Constraints
- A request carrying an
Authorizationheader that fails validation gets a hard401— it never falls back to the session cookie. - Cancel capability requires
runs:cancelon every request dimension that carries it, not just the dedicated cancel route:?action=interrupt|rollbackonPOST /api/threads/{thread_id}/runs/{run_id}/stream(action-less joins stay atruns:read), andmultitask_strategy=interrupt|rollbackon run creation (the defaultrejectstays atruns:create). Joining a run's stream is pure observation — an observer disconnecting never cancels the run. - Route-level default-deny: PAT requests are admitted only to the
thread/run lifecycle routes the v1 scopes govern —
POST /api/threads(create),POST /api/threads/search(list),GET/PATCH/DELETE /api/threads/{thread_id}, the threadgoal/state/compact/history/branchessubroutes, and exactly the implemented/runssubroutes (GET|POST /api/threads/{thread_id}/runs, the POST-onlystream,wait,regenerate/prepare, andedit-regenerate/preparecollection endpoints,GET /api/threads/{thread_id}/runs/{run_id}plus itscancel(POST),join/messages/events/workspace-changes(GET), andGET|POST .../runs/{run_id}/stream), plusPOST /api/runs/stream|waitandGET /api/runs/{run_id}/messages|feedback. A route added under/runsis denied until explicitly added to the policy. Every other authenticated route — memory, agents, models, MCP/skills config, integrations, channels, uploads — answers403to PAT callers regardless of scopes. Scope enforcement alone only constrains permission-decorated routes, so the allowlist is the outer boundary; session-cookie callers are unaffected. - PAT credentials never carry admin capability, even when the owning user is an admin. This includes extension-contributed admin routes: the extension principal projection suppresses every admin signal for PAT callers.
- Revoking or deleting the owning user invalidates their PATs on the next request.
LangGraph-compatible API
Base URL: /api/langgraph
The public LangGraph-compatible API follows LangGraph SDK conventions. In the unified nginx deployment, Gateway owns /api/langgraph/* and translates those paths to its native /api/* run, thread, and streaming routers.
Threads
Create Thread
POST /api/langgraph/threads
Content-Type: application/json
Request Body:
{
"metadata": {}
}
Response:
{
"thread_id": "abc123",
"created_at": "2024-01-15T10:30:00Z",
"metadata": {}
}
Get Thread State
GET /api/langgraph/threads/{thread_id}/state
Response:
{
"values": {
"messages": [...],
"sandbox": {...},
"artifacts": [...],
"thread_data": {...},
"title": "Conversation Title"
},
"next": [],
"config": {...}
}
Runs
Create Run
Execute the agent with input.
POST /api/langgraph/threads/{thread_id}/runs
Content-Type: application/json
Idempotency-Key: <unique key for this logical request> # optional
The thread-scoped create, stream, and wait endpoints accept an optional
Idempotency-Key header. Retrying with the same authenticated user, thread_id,
and key reuses the existing run instead of executing the input again. The key is
shared across /runs, /runs/stream, and /runs/wait for a given user and
thread, so the same key string cannot back two different calls even across those
endpoints. Reuse is bound to the original input and assistant_id; a retry
that changes either returns 409. Generate a new key for every intentional user
action; reuse a key only when retrying that same action after an uncertain HTTP
result. Keys may be at most 255 characters. Stateless /api/langgraph/runs/*
endpoints do not support this header because requests without an explicit thread
create a new temporary conversation.
Retrying a still-running run that this worker cannot stream returns 409 from
/runs/stream (Run ... is not active on this worker and cannot be streamed)
with no Retry-After. The same shape on /runs/wait returns 200
{"status": "<durable status>", "error": ...} without blocking for a final
state. Retrying a finished run through /runs/wait also returns that durable
status payload rather than the latest thread checkpoint: a later run on the
same thread may have advanced the head, and /wait does not claim that head
as this run's result. That status is the durable row after completion, not
the hydrated record from admission time. The original creating /wait still
returns this run's checkpoint even if a retry overlaps while it is waiting. Retrying a finished run whose SSE log is gone emits a gap frame
(stream_replay_gap, recovery: reload_durable_state) on the creating
/runs/stream endpoint and closes without an end frame; reload durable
thread/run state instead of treating the stream as empty. Observer joins of
that same run still end with end. Stateless /api/langgraph/runs/stream
does not accept this header and keeps the existing missing-stream close of
end; the gap signal is only on a thread-scoped creating retry.
Request Body:
{
"input": {
"messages": [
{
"role": "user",
"content": "Hello, can you help me?"
}
]
},
"config": {
"recursion_limit": 100,
"configurable": {
"model_name": "gpt-4",
"thinking_enabled": false,
"is_plan_mode": false
}
},
"stream_mode": ["values", "messages-tuple", "custom"]
}
Stream Mode Compatibility:
- Use:
values,messages-tuple,custom,updates,debug,tasks,checkpoints - Unsupported modes, including
messages,events, andtools, return422before a run is created. DeerFlow never substitutesvaluesfor an unsupported mode.
Run Option Compatibility:
- Supported concurrency strategies:
reject,rollback, andinterrupt - Compatibility default:
if_not_exists="create"; this matches DeerFlow's current behavior - Artifact delivery is enforced automatically when a run creates or modifies regular files under
/mnt/user-data/outputs.present_filesmust present at least one path produced by the current run (or a directory containing it), and the terminal receipt must be persisted; presenting only an unrelated file does not satisfy delivery. Runs without changed outputs retain ordinary conversational behavior.artifact_deliveryis not a client-settable run option. - Unsupported options return
422:webhook,stream_resumable=true,after_seconds,feedback_keys, any non-nullon_completionvalue (including the SDK values"complete"and"continue"),if_not_exists="reject", andmultitask_strategy="enqueue" stream_resumable=falseis accepted: it is the LangGraph SDK's default and requests the non-resumable stream DeerFlow already serves- Undeclared SDK options, including
checkpoint_duringanddurability, also return422instead of being silently discarded
When outputs changed during the run, run.delivery events retain the Slice 1
facts (presented, paths, and by_tool) and add produced_paths,
presented_paths, matched_paths, plus an explicit verdict: verification,
stage (presented, mismatched, or not_started), and satisfied. Receipts
for runs without changed outputs keep their existing shape.
Recursion Limit:
config.recursion_limit caps the number of graph steps LangGraph will execute
in a single run. The unified Gateway path defaults to 100 in
build_run_config (see backend/app/gateway/services.py), which is a safer
starting point for plan-mode or subagent-heavy runs. Clients can still set
recursion_limit explicitly in the request body; increase it if you run deeply
nested subagent graphs. Scheduled-task launches do not take a client body: they
use scheduler.recursion_limit from config.yaml (default 1000, matching
the web UI). For safety, the Gateway clamps any supplied
value to a configurable server ceiling (max_recursion_limit in config.yaml,
default 1000) so a single run cannot execute unbounded graph steps (runaway
LLM cost / DoS); invalid or non-positive values fall back to the 100 default.
Configurable Options:
model_name(string): Override the default modelthinking_enabled(boolean): Enable extended thinking for supported modelsis_plan_mode(boolean): Enable TodoList middleware for task tracking
Response: Server-Sent Events (SSE) stream
event: values
data: {"messages": [...], "title": "..."}
event: messages
data: {"content": "Hello! I'd be happy to help.", "role": "assistant"}
event: end
data: {}
Get Run History
GET /api/langgraph/threads/{thread_id}/runs
Response:
{
"runs": [
{
"run_id": "run123",
"status": "success",
"created_at": "2024-01-15T10:30:00Z"
}
]
}
Stream Run
Stream responses in real-time.
POST /api/langgraph/threads/{thread_id}/runs/stream
Content-Type: application/json
Idempotency-Key: <unique key for this logical request> # optional
Same request body as Create Run. Returns SSE stream.
Stateless Stream Run
Start a conversation without creating a thread first. Gateway auto-creates a
thread when config.configurable.thread_id is omitted, and returns both
identifiers in the response Content-Location header.
POST /api/langgraph/runs/stream
Content-Type: application/json
Accept: text/event-stream
Through Nginx, /api/langgraph/runs/stream is rewritten to the native Gateway
path POST /api/runs/stream.
Request Body: Same as Create Run. Omit thread_id to start a
new conversation; include it to continue an existing one:
{
"input": {
"messages": [
{
"role": "user",
"content": "Hello, can you help me?"
}
]
},
"config": {
"recursion_limit": 100,
"configurable": {
"model_name": "gpt-4",
"thinking_enabled": false,
"is_plan_mode": false
}
},
"stream_mode": ["values", "messages-tuple", "custom"]
}
Response: Server-Sent Events (SSE) stream with a Content-Location header:
Content-Location: /api/threads/{thread_id}/runs/{run_id}
Clients should parse thread_id and run_id from this header (the path ends
with /runs/{run_id}). Persist thread_id and send it back on the next turn
via config.configurable.thread_id to keep conversation history.
Continuing a conversation:
{
"input": {
"messages": [
{
"role": "user",
"content": "What did I just ask?"
}
]
},
"config": {
"configurable": {
"thread_id": "abc123",
"model_name": "gpt-4"
}
},
"stream_mode": ["values", "messages-tuple", "custom"]
}
Gateway API
Base URL: /api
Models
List Models
Get all available LLM models from configuration.
GET /api/models
Response:
{
"models": [
{
"name": "gpt-4",
"display_name": "GPT-4",
"supports_thinking": false,
"supports_vision": true
},
{
"name": "claude-3-opus",
"display_name": "Claude 3 Opus",
"supports_thinking": false,
"supports_vision": true
},
{
"name": "deepseek-v3",
"display_name": "DeepSeek V3",
"supports_thinking": true,
"supports_vision": false
}
]
}
Get Model Details
GET /api/models/{model_name}
Response:
{
"name": "gpt-4",
"display_name": "GPT-4",
"model": "gpt-4",
"max_tokens": 4096,
"supports_thinking": false,
"supports_vision": true
}
MCP Configuration
Get MCP Config
Get current MCP server configurations.
GET /api/mcp/config
Requires an authenticated admin session. Sensitive env/header/OAuth secret
values are masked in the response. Environment placeholders outside secret
containers are returned in their raw form so editing cannot expose or persist
their expanded values. Invalid operator-authored JSON/config shapes return
400 instead of being reported as a Gateway fault.
Response:
{
"mcp_servers": {
"github": {
"enabled": true,
"type": "stdio",
"command": "npx",
"args": ["-y", "@modelcontextprotocol/server-github"],
"env": {
"GITHUB_TOKEN": "***"
},
"description": "GitHub operations"
}
}
}
Update MCP Config
Update MCP server configurations.
PUT /api/mcp/config
Content-Type: application/json
Requires an authenticated admin session. API-managed stdio MCP servers may
only use allowed executable names for command (default: npx, uvx). Set
DEER_FLOW_MCP_STDIO_COMMAND_ALLOWLIST to a comma-separated list when a
deployment needs additional trusted launchers.
Request Body:
{
"mcp_servers": {
"github": {
"enabled": true,
"type": "stdio",
"command": "npx",
"args": ["-y", "@modelcontextprotocol/server-github"],
"env": {
"GITHUB_TOKEN": "$GITHUB_TOKEN"
},
"description": "GitHub operations"
}
}
}
Response:
{
"mcp_servers": {
"github": {
"enabled": true,
"type": "stdio",
"command": "npx",
"args": ["-y", "@modelcontextprotocol/server-github"],
"env": {
"GITHUB_TOKEN": "***"
},
"description": "GitHub operations"
}
}
}
Update One MCP Server State
Enable or disable one configured MCP server without replacing the full extensions configuration.
PATCH /api/mcp/config
Content-Type: application/json
Requires an authenticated admin session. Enabling a stdio server validates
that server's command against the same allowlist used by the full PUT
endpoint. Disabling a server does not require its command to be allowlisted, and
invalid commands on other servers do not block the update. The endpoint
preserves secrets, environment-variable placeholders, skills, custom server
fields, and other top-level extensions config. SSE/HTTP targets may use either
DeerFlow's type field or the MCP-spec transport field.
Request Body:
{
"server_name": "semantic-scholar",
"enabled": false
}
The response is the full masked MCP configuration, matching GET and PUT.
An unknown server_name returns 404; attempting to enable a server with a
disallowed stdio command returns 400.
Add MCP Servers
Add one or more servers without replacing existing entries. The Gateway
re-reads the file under the shared configuration lock, so concurrent sibling
changes are preserved. Existing names return 409.
POST /api/mcp/config/servers
Content-Type: application/json
The request body uses the same mcp_servers map as the full PUT endpoint.
Replace One MCP Server
Completely replace one existing server while preserving sibling entries.
Omitted ordinary fields are deleted or reset; explicit *** placeholders
restore the corresponding stored secret.
A disabled stdio replacement may keep a syntactically valid command outside
the allowlist for offline editing. Command-shape and code-injecting environment
variable checks still run when saving; the allowlist and executable-argument
policy run when the server is enabled.
PUT /api/mcp/config/server
Content-Type: application/json
{
"server_name": "github",
"server": {
"enabled": true,
"command": "npx",
"args": ["-y", "@modelcontextprotocol/server-github"],
"env": {"GITHUB_TOKEN": "***"}
}
}
Delete One MCP Server
Delete one server without replacing sibling entries. The server name is a path parameter and the DELETE request has no body. Percent-encode names before placing them in the URL; the path converter also keeps legacy empty and slash-containing names addressable.
DELETE /api/mcp/config/servers/{server_name}
All targeted mutations return the full masked MCP configuration. Before any write, the Gateway resolves environment variables in a copy and validates the same expanded document the runtime will load while persisting the original raw placeholders.
Reset MCP Tools Cache
Clear cached MCP tools and persistent MCP sessions process-wide. This affects all threads and users in the current Gateway process. Tools are loaded again from configured MCP servers on the next agent run or tool lookup.
POST /api/mcp/cache/reset
Requires an authenticated admin session.
Response:
{
"success": true,
"message": "MCP tools cache reset. Tools will reload on next use."
}
Skills
List Skills
Get all available skills.
GET /api/skills
Response:
{
"skills": [
{
"name": "pdf-processing",
"display_name": "PDF Processing",
"description": "Handle PDF documents efficiently",
"enabled": true,
"license": "MIT",
"path": "public/pdf-processing"
},
{
"name": "frontend-design",
"display_name": "Frontend Design",
"description": "Design and build frontend interfaces",
"enabled": false,
"license": "MIT",
"path": "public/frontend-design"
}
]
}
Get Skill Details
GET /api/skills/{skill_name}
Response:
{
"name": "pdf-processing",
"display_name": "PDF Processing",
"description": "Handle PDF documents efficiently",
"enabled": true,
"license": "MIT",
"path": "public/pdf-processing",
"allowed_tools": ["read_file", "write_file", "bash"],
"content": "# PDF Processing\n\nInstructions for the agent..."
}
Enable Skill
POST /api/skills/{skill_name}/enable
Response:
{
"success": true,
"message": "Skill 'pdf-processing' enabled"
}
Disable Skill
POST /api/skills/{skill_name}/disable
Response:
{
"success": true,
"message": "Skill 'pdf-processing' disabled"
}
Install Skill
Install a skill from a .skill file.
POST /api/skills/install
Content-Type: multipart/form-data
Request Body:
file: The.skillfile to install
Response:
{
"success": true,
"message": "Skill 'my-skill' installed successfully",
"skill": {
"name": "my-skill",
"display_name": "My Skill",
"path": "custom/my-skill"
}
}
Export a Custom Skill
Admin session authentication is required for both requests. PAT credentials cannot export. Only the current user's custom skill is eligible; public, legacy and integration fallback is never used. A disabled custom skill remains eligible.
GET /api/skills/custom/{skill_name}/export-manifestreturnsskill_name,revision(SHA-256 or null),can_export,file_count,directory_count,total_bytes,files(path,type,size,executable),requirements(compatibility,allowed_tools,required_secretsnames and optional flags), and structuredwarnings/blockers. Paths are relative;.is the package root, counted in directory/entry totals. Structural blockers return a non-downloadable manifest. Declarations are not credential values or dependency verification.GET /api/skills/custom/{skill_name}/export?expected_revision=<64 lowercase hex characters>recaptures content and rejects stale previews with 409 before sending ZIP headers. Successful responses carryapplication/zip, attachment<skill_name>.skill, accurateContent-Length,Cache-Control: private, no-store, andX-Content-Type-Options: nosniff.
Error detail contains a safe code, message, and optional relative path. Codes/statuses: skill_not_found 404, skill_changed 409, skill_export_limit_exceeded 413, skill_export_unsupported 422, skill_export_busy 429, skill_export_timeout 503, skill_export_failed 500; existing 401/403 auth behavior applies. Limits are 4096 entries including directories, 64 MiB/file, 100 MiB raw/ZIP, 1 MiB frontmatter, 1024 UTF-8 bytes per ZIP path and depth 32. Frontmatter preflight rejects YAML aliases and bounds structure to 32 nesting levels / 16384 parser events before constructing YAML objects. A 5-second lock wait and 60-second cooperative worker deadline bound work; blocking OS calls cannot be forcibly interrupted. Two export slots are shared across all users in each Gateway process; both previews and downloads use them, and 429 means that process-wide capacity is occupied. Slots remain held through worker drain and temporary-file cleanup. The streaming phase has a separate 120-second inactivity deadline, reset after each successful ASGI send. A continuously progressing transfer may exceed 120 seconds overall; a stalled send does not reset the deadline. Expiry aborts the incomplete download (no replacement JSON after ZIP headers); clients must retry. Client disconnect during preparation cancels and drains the worker, then exits the handler normally rather than leaking a synthetic task cancellation. No export cache, persistent job or sharing URL is created.
Raw skill files, sidecars and empty directories are preserved. No hooks/scripts run during export and no secrets are redacted from package files. Import still uses normal security scanning and conflict checks. Export requires no-follow descriptor-relative host filesystem operations; unsupported platforms receive 422 rather than following links unsafely.
Reload Skills
Invalidate the skill prompt caches for every user in the current Gateway process. Subsequent runs rescan the configured public, custom, and legacy skill directories; runs that have already started keep their existing skill snapshot.
POST /api/skills/reload
The request has no body and requires an authenticated administrator. For a cookie-authenticated request, send the CSRF cookie value in the matching header:
curl -X POST http://localhost:2026/api/skills/reload \
-b cookies.txt \
-H "X-CSRF-Token: <csrf_token-cookie-value>"
Response:
{
"success": true,
"scope": "process",
"message": "Skill caches invalidated; subsequent runs in this Gateway process will rescan the latest skills."
}
success confirms cache invalidation, not that every file on disk was valid:
malformed skills retain the existing parser behavior of being skipped and
logged. The endpoint returns 401 for unauthenticated callers, 403 for
non-admin users, and a generic 500 if the invalidation mechanism itself
fails or the process-local background scan does not finish within the cache
refresh timeout. A loader-level failure, such as an unavailable mounted root,
does not publish an empty catalog: the last successfully loaded process cache
remains available. A timed-out scan continues in its daemon worker and can
still populate the process cache when it finishes.
The scope is deliberately process-local. Each Uvicorn worker or Kubernetes Pod must be called directly; repeated requests through a load-balanced Service do not guarantee that every instance is reached. External MinIO/NFS/CSI writes bypass the validation, SkillScan, and history used by the install/edit APIs, so the mounted directory must be writable only by trusted operators.
File Uploads
Upload Files
Upload one or more files to a thread.
POST /api/threads/{thread_id}/uploads
Content-Type: multipart/form-data
Request Body:
files: One or more files to upload
Response:
{
"success": true,
"files": [
{
"filename": "document.pdf",
"size": 1234567,
"path": ".deer-flow/threads/abc123/user-data/uploads/document.pdf",
"virtual_path": "/mnt/user-data/uploads/document.pdf",
"artifact_url": "/api/threads/abc123/artifacts/mnt/user-data/uploads/document.pdf",
"markdown_file": "document.md",
"markdown_path": ".deer-flow/threads/abc123/user-data/uploads/document.md",
"markdown_virtual_path": "/mnt/user-data/uploads/document.md",
"markdown_artifact_url": "/api/threads/abc123/artifacts/mnt/user-data/uploads/document.md"
}
],
"message": "Successfully uploaded 1 file(s)"
}
Supported Document Formats (auto-converted to Markdown):
- PDF (
.pdf) - PowerPoint (
.ppt,.pptx) - Excel (
.xls,.xlsx) - Word (
.doc,.docx)
List Uploaded Files
GET /api/threads/{thread_id}/uploads/list
Response:
{
"files": [
{
"filename": "document.pdf",
"size": 1234567,
"path": ".deer-flow/threads/abc123/user-data/uploads/document.pdf",
"virtual_path": "/mnt/user-data/uploads/document.pdf",
"artifact_url": "/api/threads/abc123/artifacts/mnt/user-data/uploads/document.pdf",
"extension": ".pdf",
"modified": 1705997600.0
}
],
"count": 1
}
Delete File
DELETE /api/threads/{thread_id}/uploads/{filename}
Response:
{
"success": true,
"message": "Deleted document.pdf"
}
Thread Cleanup
Remove DeerFlow-managed local thread files under .deer-flow/threads/{thread_id} after the LangGraph thread itself has been deleted.
DELETE /api/threads/{thread_id}
Response:
{
"success": true,
"message": "Deleted local thread data for abc123"
}
Error behavior:
422for invalid thread IDs500returns a generic{"detail": "Failed to delete local thread data."}response while full exception details stay in server logs
Artifacts
Get Artifact
Download or view an artifact generated by the agent.
GET /api/threads/{thread_id}/artifacts/{path}
Path Examples:
/api/threads/abc123/artifacts/mnt/user-data/outputs/result.txt/api/threads/abc123/artifacts/mnt/user-data/uploads/document.pdf
Query Parameters:
download(boolean): Iftrue, force download with Content-Disposition header
Response: File content with appropriate Content-Type
Error Responses
All APIs return errors in a consistent format:
{
"detail": "Error message describing what went wrong"
}
HTTP Status Codes:
400- Bad Request: Invalid input404- Not Found: Resource not found422- Validation Error: Request validation failed500- Internal Server Error: Server-side error
Authentication
DeerFlow supports four HTTP identity sources. They share the same thread/run isolation rules but differ in whether a row is created in users and how external identities are mapped. See AUTH_DESIGN.md for the full design.
| Model | Entry | users table |
Isolation key |
|---|---|---|---|
| Browser session | access_token cookie after login/register |
Yes | users.id |
| OIDC / SSO | OAuth callback → cookie | Yes | users.id (see SSO.md) |
| IM channel binding | Connect code + channel_connections |
Bound to registered user | channel_connections.owner_user_id |
| Internal Auth | X-DeerFlow-Internal-Token + X-DeerFlow-Owner-User-Id |
No | Owner string on threads_meta.user_id |
IM channel binding and Internal Auth are both platform-trust integrations: DeerFlow trusts the channel/platform to authenticate end users. IM bindings persist the mapping in channel_connections / channel_conversations and require a DeerFlow users row. Internal Auth lets a platform call the Gateway API directly with a deployment-shared token and a per-request owner header—no users row, but thread/run/checkpoint isolation works the same way.
Browser session (default)
DeerFlow enforces authentication for all non-public HTTP routes. Public routes are limited to health/docs metadata and these public auth endpoints:
POST /api/v1/auth/initializecreates the first admin account when no admin exists.POST /api/v1/auth/login/locallogs in with email/password and sets an HttpOnlyaccess_tokencookie.POST /api/v1/auth/registercreates a regularuseraccount and sets the session cookie.POST /api/v1/auth/logoutclears the session cookie.GET /api/v1/auth/setup-statusreports whether the first admin still needs to be created.
The authenticated auth endpoints are:
GET /api/v1/auth/mereturns the current user.POST /api/v1/auth/change-passwordchanges password, optionally changes email during setup, incrementstoken_version, and reissues the cookie.
Protected state-changing requests also require the CSRF double-submit token: send the csrf_token cookie value as the X-CSRF-Token header. Login/register/initialize/logout are bootstrap auth endpoints: they are exempt from the double-submit token but still reject hostile browser Origin headers.
User isolation is enforced from the authenticated user context:
- Thread metadata is scoped by
threads_meta.user_id; search/read/write/delete APIs only expose the current user's threads. - Thread files live under
{base_dir}/users/{user_id}/threads/{thread_id}/user-data/and are exposed inside the sandbox as/mnt/user-data/. - Memory and custom agents are stored under
{base_dir}/users/{user_id}/....
Note: MCP outbound connections can still use OAuth for configured HTTP/SSE MCP servers; that is separate from DeerFlow API authentication.
Internal Auth (platform HTTP integration)
For server-to-server integrations (e.g. a Feishu or WeCom/Enterprise WeChat bot backend), configure:
export DEER_FLOW_INTERNAL_AUTH_TOKEN="<long-random-secret>"
| Header | Required | Description |
|---|---|---|
X-DeerFlow-Internal-Token |
Yes | Must match DEER_FLOW_INTERNAL_AUTH_TOKEN; missing/invalid → 401 |
X-DeerFlow-Owner-User-Id |
Yes for per-user isolation | Platform user id (e.g. feishu_ou_alice, wecom_user_bob); omit → default bucket |
Does not use browser cookies or CSRF tokens. Does not insert into users; sets threads_meta.user_id / runs.user_id from the owner header. DeerFlow validates only the platform token—not whether the owner id represents a real end user; user validity is entirely the platform's responsibility. See AUTH_DESIGN.md — Internal Auth for trust boundaries, persistence, and security notes.
Use the standard Gateway thread/run endpoints (POST /api/threads, POST /api/threads/{thread_id}/runs/stream, etc.) with the headers above on every request.
Rate Limiting
No rate limiting is implemented by default. For production deployments, configure rate limiting in Nginx:
limit_req_zone $binary_remote_addr zone=api:10m rate=10r/s;
location /api/ {
limit_req zone=api burst=20 nodelay;
proxy_pass http://backend;
}
Streaming Support
Gateway's LangGraph-compatible API streams run events with Server-Sent Events (SSE).
Thread-scoped streaming (thread must exist):
POST /api/langgraph/threads/{thread_id}/runs/stream
Accept: text/event-stream
Stateless streaming (no pre-created thread; Gateway auto-creates one):
POST /api/langgraph/runs/stream
Accept: text/event-stream
Both endpoints return Content-Location: /api/threads/{thread_id}/runs/{run_id}.
The DeerFlow web UI and LangGraph SDK clients rely on this header to discover the
assigned thread_id and run_id on the first message of a new chat.
SSE replay retention and gaps
Clients may reconnect to a run stream with Last-Event-ID. Replay history is
bounded by stream_bridge.queue_maxsize (default 256) and, for Redis, by the
rolling stream_ttl_seconds. A retained cursor resumes after that event with no
additional control frame.
When a syntactically valid cursor is older than the retained watermark, the
server sends exactly one gap event before any retained data and closes that
subscription without an end event:
event: gap
data: {"code":"stream_replay_gap","run_id":"run-123","requested_event_id":"1718000000000-1","earliest_available_event_id":"1718000000100-42","latest_available_event_id":"1718000000200-84","recovery":"reload_durable_state"}
The frame deliberately has no SSE id:. Both earliest_available_event_id and
latest_available_event_id are string | null (they are null when no events
are retained in the buffer). Consumers must reload durable thread state and
persisted run events/messages, then may reconnect from latest_available_event_id
to follow newer live events, or rejoin without a cursor when the buffer is empty
(latest_available_event_id is null). A gap does not cancel the active run.
The same signal applies when a no-cursor subscriber has already established an
empty-stream wait but the first Redis wake-up falls behind before delivery; in
that case requested_event_id is null. Malformed cursor handling is
backend-specific and is not the same as a valid cursor that was evicted.
SDK Usage
Python (LangGraph SDK)
from langgraph_sdk import get_client
client = get_client(url="http://localhost:2026/api/langgraph")
run_meta: dict[str, str] = {}
def on_run_created(meta) -> None:
# langgraph-sdk 0.3.x parses Content-Location only when this callback is set.
if meta.thread_id:
run_meta["thread_id"] = meta.thread_id
run_meta["run_id"] = meta.run_id
# Option A: stateless stream — no thread pre-creation
# Gateway auto-creates a thread and returns thread_id/run_id in Content-Location.
async for event in client.runs.stream(
None,
"lead_agent",
input={"messages": [{"role": "user", "content": "Hello"}]},
config={"configurable": {"model_name": "gpt-4"}},
stream_mode=["values", "messages-tuple", "custom"],
on_run_created=on_run_created,
):
print(event)
thread_id = run_meta["thread_id"] # persist before the next turn
# Option A (continued): same thread on the next turn
async for event in client.runs.stream(
None,
"lead_agent",
input={"messages": [{"role": "user", "content": "What did I just ask?"}]},
config={"configurable": {"thread_id": thread_id, "model_name": "gpt-4"}},
stream_mode=["values", "messages-tuple", "custom"],
on_run_created=on_run_created,
):
print(event)
# Option B: thread-scoped stream — create thread first, then stream
thread = await client.threads.create()
async for event in client.runs.stream(
thread["thread_id"],
"lead_agent",
input={"messages": [{"role": "user", "content": "Hello"}]},
config={"configurable": {"model_name": "gpt-4"}},
stream_mode=["values", "messages-tuple", "custom"],
on_run_created=on_run_created,
):
print(event)
JavaScript/TypeScript
// Using fetch for Gateway API
const response = await fetch('/api/models');
const data = await response.json();
console.log(data.models);
function parseRunLocation(contentLocation: string | null) {
if (!contentLocation) return null;
const match = /\/threads\/([^/]+)\/runs\/([^/]+)/.exec(contentLocation);
if (!match) return null;
return { threadId: match[1], runId: match[2] };
}
// Option A: stateless stream — no thread pre-creation
let threadId: string | undefined;
const firstResponse = await fetch("/api/langgraph/runs/stream", {
method: "POST",
headers: {
"Content-Type": "application/json",
Accept: "text/event-stream",
},
body: JSON.stringify({
input: { messages: [{ role: "user", content: "Hello" }] },
stream_mode: ["values", "messages-tuple", "custom"],
}),
});
const created = parseRunLocation(firstResponse.headers.get("Content-Location"));
threadId = created?.threadId;
console.log("thread_id:", created?.threadId, "run_id:", created?.runId);
// Option B: continue the same thread on the next turn
const followUpResponse = await fetch("/api/langgraph/runs/stream", {
method: "POST",
headers: {
"Content-Type": "application/json",
Accept: "text/event-stream",
},
body: JSON.stringify({
input: { messages: [{ role: "user", content: "What did I just ask?" }] },
config: { configurable: { thread_id: threadId } },
stream_mode: ["values", "messages-tuple", "custom"],
}),
});
// Option C: thread-scoped stream when you already have a thread_id
const streamResponse = await fetch(`/api/langgraph/threads/${threadId}/runs/stream`, {
method: "POST",
headers: {
"Content-Type": "application/json",
Accept: "text/event-stream",
},
body: JSON.stringify({
input: { messages: [{ role: "user", content: "Hello" }] },
stream_mode: ["values", "messages-tuple", "custom"],
}),
});
const reader = streamResponse.body?.getReader();
// Decode and parse SSE frames from reader in your client code.
cURL Examples
# List models
curl http://localhost:2026/api/models
# Get MCP config
curl http://localhost:2026/api/mcp/config
# Upload file
curl -X POST http://localhost:2026/api/threads/abc123/uploads \
-F "files=@document.pdf"
# Enable skill
curl -X POST http://localhost:2026/api/skills/pdf-processing/enable
# Stateless stream — no thread pre-creation
curl -s -D - -N -X POST http://localhost:2026/api/langgraph/runs/stream \
-H "Content-Type: application/json" \
-H "Accept: text/event-stream" \
-d '{
"input": {"messages": [{"role": "user", "content": "Hello"}]},
"config": {
"recursion_limit": 100,
"configurable": {"model_name": "gpt-4"}
},
"stream_mode": ["values", "messages-tuple", "custom"]
}'
# Read Content-Location: /api/threads/{thread_id}/runs/{run_id} from the headers.
# Continue the same thread on the next turn
curl -s -N -X POST http://localhost:2026/api/langgraph/runs/stream \
-H "Content-Type: application/json" \
-H "Accept: text/event-stream" \
-d '{
"input": {"messages": [{"role": "user", "content": "What did I just ask?"}]},
"config": {
"configurable": {"thread_id": "abc123", "model_name": "gpt-4"}
},
"stream_mode": ["values", "messages-tuple", "custom"]
}'
# Thread-scoped flow — create thread first, then stream
curl -X POST http://localhost:2026/api/langgraph/threads \
-H "Content-Type: application/json" \
-d '{}'
curl -X POST http://localhost:2026/api/langgraph/threads/abc123/runs/stream \
-H "Content-Type: application/json" \
-H "Accept: text/event-stream" \
-d '{
"input": {"messages": [{"role": "user", "content": "Hello"}]},
"config": {
"recursion_limit": 100,
"configurable": {"model_name": "gpt-4"}
},
"stream_mode": ["values", "messages-tuple", "custom"]
}'
The unified Gateway path defaults
config.recursion_limitto 100 for plan-mode and subagent-heavy runs. Clients may still setconfig.recursion_limitexplicitly — see the Create Run section for details. Scheduled-task launches usescheduler.recursion_limitfromconfig.yamlinstead of a client body.
Chat archive and restore
POST /api/threads/search accepts archived: true for archived chats or
archived: false for recent chats (including legacy rows without an archive flag).
Omit the field or use null to include both. Filtering applies before limit and
offset and is scoped to the authenticated user. Combine it with the existing
metadata and status filters when needed.
Archive with PATCH /api/threads/{thread_id} and body
{"metadata":{"deerflow_archived":true}}; use false to restore. The flag must be
a JSON boolean. Writes containing only boolean pin/archive flags preserve
updated_at and all other metadata. The owner-checked endpoint returns the normal
thread metadata response; original thread and artifact URLs remain available.
Archiving does not cancel runs, pause schedules, or change retention.