mirror of
https://github.com/bytedance/deer-flow.git
synced 2026-08-08 13:58:38 +00:00
* fix(sandbox): project enabled skills into sandbox views * fix(skills): keep projection mutations consistent * fix(skills): fail closed on projection errors * fix(skills): isolate per-scope failures during boot projection rebuild rebuild_all_skill_projections() propagated any exception from the public rebuild or from a single user's rebuild straight out of the gateway lifespan startup, uncaught. A single broken user directory (bad permissions, corrupted _skill_states.json, unreadable content) would therefore abort gateway boot for every user, not just that one - _rebuild_*_locked already fails closed internally (clears the view and re-raises), so the boot loop only needed to stop treating that re-raise as fatal. Each scope's rebuild now fails closed independently and boot continues; a scope left empty by a boot failure self-heals on the next sandbox acquire via ensure_skill_projections(). Also patches deerflow.skills.projection.rebuild_all_skill_projections in the memory-flush lifespan test fixture, matching the two sibling fixtures in the same file — this call is now on the lifespan startup path and the fixture's minimal SimpleNamespace config predates it. * test(skills): update authz test for the projection-aware public toggle _persist_shared_skill_state (introduced earlier in this branch) reads the shared extensions_config.json fresh from disk under the projection lock instead of through the cached get_extensions_config() singleton - that's the whole point of the fix (stale worker caches must not clobber another worker's concurrent update). The name no longer exists on the skills router module, so the test's monkeypatch of it started raising AttributeError instead of exercising the endpoint. The mock storage in this test isn't a real LocalSkillStorage instance, so _persist_shared_skill_state's projection-mutation branch is already skipped (nullcontext) and it falls back to a fresh ExtensionsConfig() for the nonexistent tmp config_path - no replacement monkeypatch needed. * fix(sandbox): make skill projection ensure best-effort in acquire acquire() called _ensure_skills_projection() directly, outside any try/except, in both LocalSandboxProvider and AioSandboxProvider. Every other skill-mount setup path in these providers has always caught exceptions and logged a warning rather than failing sandbox acquire outright (e.g. when config.yaml can't be resolved) - these two new call sites broke that contract, so any projection failure (including simply not having a config.yaml, as in CI's test environment) now failed acquire() itself instead of just leaving skill mounts off. _ensure_skills_projection now catches its own exceptions and returns None; both providers' callers already tolerate that (a None projection skips the skill-specific mounts, matching the existing degrade path) after making _append_public_skill_mapping and the custom/legacy mount block in LocalSandboxProvider explicitly None-safe. Caught by running the full suite with config.yaml removed, matching CI's environment - not caught locally because a real config.yaml was present, masking the failure. * fix(sandbox): make E2B skill projection mounts best-effort _skill_projection_mounts called ensure_skill_projections with no guard, unlike Local/AIO's _ensure_skills_projection. A raise propagated out of _apply_mounts before the configured-mounts loop ran, so a skills projection failure dropped the operator's own configured mounts too - only caught by create()'s outer warning, with nothing applied at all. Swallow here and return an empty mount list on failure, matching the Local/AIO pattern: still fail-closed for skills, but no longer widens the blast radius to unrelated configured mounts. Review feedback from PR #4178. * docs(skills): document projection trade-offs flagged in review - _update_tree_digest: note the metadata-only (not content) hashing trade-off and why runtime writes through this codebase are still covered regardless (rebuild-under-lock + rename always changes inode). - LocalSandboxProvider.acquire: note the acquire-time self-heal cost (cheap on a fresh manifest, ~400ms rebuild under lock on stale/drift). - skill_projection_mutation: drop the no-op except-Exception-then-raise; a raise from the mutation already propagates past the yield with the view left cleared, no explicit re-raise needed. - provisioner README: spell out that hostPath skills volumes require the gateway and K8s node to share DEER_FLOW_HOST_BASE_DIR (single-node or shared storage), and that the custom/legacy volumes' hostPath type Directory (not DirectoryOrCreate) makes a violation of that assumption a visible Pod-creation failure instead of a silent empty mount. Review feedback from PR #4178. * fix(skills): lazily repair user projections * fix(skills): close projection review gaps * fix(skills): refresh user projection enable state * fix(skills): close projection review follow-ups * fix(skills): preserve state across projection writes --------- Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
192 lines
8.6 KiB
YAML
192 lines
8.6 KiB
YAML
# DeerFlow Production Environment
|
|
# Usage: make up
|
|
#
|
|
# Services:
|
|
# - nginx: Reverse proxy (port 2026, configurable via PORT env var)
|
|
# - frontend: Next.js production server
|
|
# - gateway: FastAPI Gateway API + agent runtime
|
|
# - redis: Redis Streams backend for cross-worker SSE stream bridge
|
|
# - provisioner: (optional) Sandbox provisioner for Kubernetes mode
|
|
#
|
|
# Key environment variables (set via environment/.env or scripts/deploy.sh):
|
|
# DEER_FLOW_PROJECT_ROOT — project root for relative runtime paths
|
|
# DEER_FLOW_HOME — runtime data dir, default .deer-flow under $DEER_FLOW_PROJECT_ROOT (or cwd)
|
|
# DEER_FLOW_CONFIG_PATH — path to config.yaml
|
|
# DEER_FLOW_EXTENSIONS_CONFIG_PATH — path to extensions_config.json
|
|
# DEER_FLOW_SKILLS_PATH — skills dir, default $DEER_FLOW_PROJECT_ROOT/skills
|
|
# DEER_FLOW_DOCKER_SOCKET — Docker socket path for aio/DooD mode, default /var/run/docker.sock (used only by the opt-in docker-compose.dood.yaml overlay)
|
|
# DEER_FLOW_REPO_ROOT — repo root (used for skills host path in DooD)
|
|
# BETTER_AUTH_SECRET — required for frontend auth/session security
|
|
# DEER_FLOW_INTERNAL_AUTH_TOKEN — shared internal Gateway auth token for multi-worker IM channels
|
|
#
|
|
# LangSmith tracing is disabled by default (LANGSMITH_TRACING=false).
|
|
# Set LANGSMITH_TRACING=true and LANGSMITH_API_KEY in .env to enable it.
|
|
#
|
|
# Access: http://localhost:${PORT:-2026}
|
|
|
|
services:
|
|
# ── Redis Stream Bridge ────────────────────────────────────────────────────
|
|
redis:
|
|
image: redis:7-alpine
|
|
container_name: deer-flow-redis
|
|
command: ["redis-server", "--appendonly", "yes"]
|
|
volumes:
|
|
- redis-data:/data
|
|
healthcheck:
|
|
test: ["CMD", "redis-cli", "ping"]
|
|
interval: 5s
|
|
timeout: 3s
|
|
retries: 10
|
|
networks:
|
|
- deer-flow
|
|
restart: unless-stopped
|
|
|
|
# ── Reverse Proxy ──────────────────────────────────────────────────────────
|
|
nginx:
|
|
image: nginx:alpine
|
|
container_name: deer-flow-nginx
|
|
ports:
|
|
- "${PORT:-2026}:2026"
|
|
volumes:
|
|
- ./nginx/nginx.conf:/etc/nginx/nginx.conf.template:ro
|
|
command: >
|
|
sh -c "cp /etc/nginx/nginx.conf.template /etc/nginx/nginx.conf
|
|
&& nginx -g 'daemon off;'"
|
|
depends_on:
|
|
- frontend
|
|
- gateway
|
|
networks:
|
|
- deer-flow
|
|
restart: unless-stopped
|
|
|
|
# ── Frontend: Next.js Production ───────────────────────────────────────────
|
|
frontend:
|
|
build:
|
|
context: ../
|
|
dockerfile: frontend/Dockerfile
|
|
target: prod
|
|
args:
|
|
PNPM_STORE_PATH: ${PNPM_STORE_PATH:-/root/.local/share/pnpm/store}
|
|
NPM_REGISTRY: ${NPM_REGISTRY:-}
|
|
container_name: deer-flow-frontend
|
|
environment:
|
|
- BETTER_AUTH_SECRET=${BETTER_AUTH_SECRET}
|
|
- DEER_FLOW_INTERNAL_GATEWAY_BASE_URL=http://gateway:8001
|
|
env_file:
|
|
- ../frontend/.env
|
|
networks:
|
|
- deer-flow
|
|
restart: unless-stopped
|
|
|
|
# ── Gateway API ────────────────────────────────────────────────────────────
|
|
gateway:
|
|
build:
|
|
context: ../
|
|
dockerfile: backend/Dockerfile
|
|
args:
|
|
APT_MIRROR: ${APT_MIRROR:-}
|
|
UV_IMAGE: ${UV_IMAGE:-ghcr.io/astral-sh/uv:0.7.20}
|
|
UV_INDEX_URL: ${UV_INDEX_URL:-https://pypi.org/simple}
|
|
UV_EXTRAS: ${UV_EXTRAS:-}
|
|
NPM_REGISTRY: ${NPM_REGISTRY:-}
|
|
LARK_CLI_NPM_VERSION: ${LARK_CLI_NPM_VERSION:-1.0.65}
|
|
container_name: deer-flow-gateway
|
|
# Gateway hosts the agent runtime with in-process RunManager + StreamBridge
|
|
# singletons -- run state lives in this worker's memory. Default to a single
|
|
# worker: the Redis stream bridge solves cross-worker SSE delivery and
|
|
# reconnect, but run cancel, request dedup, and per-worker IM channel
|
|
# services remain worker-local. Override GATEWAY_WORKERS only when the
|
|
# Redis stream bridge is enabled and those limitations are acceptable.
|
|
command: sh -c "cd backend && PYTHONPATH=. uv run uvicorn app.gateway.app:app --host 0.0.0.0 --port 8001 --workers ${GATEWAY_WORKERS:-1}"
|
|
volumes:
|
|
- ${DEER_FLOW_CONFIG_PATH}:/app/backend/config.yaml:ro
|
|
- ${DEER_FLOW_EXTENSIONS_CONFIG_PATH}:/app/backend/extensions_config.json:ro
|
|
- ../skills:/app/skills:ro
|
|
- ${DEER_FLOW_HOME}:/app/backend/.deer-flow
|
|
# DooD: the host Docker socket is NOT mounted by default. It is added only
|
|
# for aio (pure-DooD) sandbox mode via the opt-in docker-compose.dood.yaml
|
|
# overlay (appended by scripts/deploy.sh). See backend/docs/CONFIGURATION.md
|
|
# Security Note section for details.
|
|
|
|
# CLI auth dirs (Claude Code / Codex) are NOT mounted by default: they
|
|
# expose the entire ~/.claude and ~/.codex (history, projects, global
|
|
# config, credentials) into the container. Mount them only when you use
|
|
# the Claude/Codex CLI login as a model provider or ACP agent, via the
|
|
# opt-in docker-compose.cli-auth.yaml overlay. Prefer an env token
|
|
# (CLAUDE_CODE_OAUTH_TOKEN, see .env.example / backend/docs/CONFIGURATION.md).
|
|
working_dir: /app
|
|
environment:
|
|
- CI=true
|
|
- DEER_FLOW_PROJECT_ROOT=/app
|
|
- DEER_FLOW_HOME=/app/backend/.deer-flow
|
|
- DEER_FLOW_CONFIG_PATH=/app/backend/config.yaml
|
|
- DEER_FLOW_EXTENSIONS_CONFIG_PATH=/app/backend/extensions_config.json
|
|
- DEER_FLOW_STREAM_BRIDGE_REDIS_URL=${DEER_FLOW_STREAM_BRIDGE_REDIS_URL:-redis://redis:6379/0}
|
|
- DEER_FLOW_CHANNELS_LANGGRAPH_URL=${DEER_FLOW_CHANNELS_LANGGRAPH_URL:-http://gateway:8001/api}
|
|
- DEER_FLOW_CHANNELS_GATEWAY_URL=${DEER_FLOW_CHANNELS_GATEWAY_URL:-http://gateway:8001}
|
|
- DEER_FLOW_INTERNAL_AUTH_TOKEN=${DEER_FLOW_INTERNAL_AUTH_TOKEN}
|
|
# DooD path/network translation
|
|
- DEER_FLOW_HOST_BASE_DIR=${DEER_FLOW_HOME}
|
|
- DEER_FLOW_SANDBOX_HOST=host.docker.internal
|
|
# Proxy values (HTTP_PROXY/HTTPS_PROXY/ALL_PROXY) are inherited from ../.env via env_file.
|
|
# Only NO_PROXY is declared here so internal service hostnames are always exempt from the proxy.
|
|
- NO_PROXY=${NO_PROXY:-}${NO_PROXY:+,}localhost,127.0.0.1,::1,gateway,frontend,nginx,provisioner,openviking,host.docker.internal
|
|
- no_proxy=${no_proxy:-}${no_proxy:+,}localhost,127.0.0.1,::1,gateway,frontend,nginx,provisioner,openviking,host.docker.internal
|
|
env_file:
|
|
- ../.env
|
|
extra_hosts:
|
|
- "host.docker.internal:host-gateway"
|
|
depends_on:
|
|
redis:
|
|
condition: service_healthy
|
|
networks:
|
|
- deer-flow
|
|
restart: unless-stopped
|
|
|
|
# ── Sandbox Provisioner (optional, Kubernetes mode) ────────────────────────
|
|
provisioner:
|
|
build:
|
|
context: ./provisioner
|
|
dockerfile: Dockerfile
|
|
args:
|
|
APT_MIRROR: ${APT_MIRROR:-}
|
|
PIP_INDEX_URL: ${PIP_INDEX_URL:-}
|
|
container_name: deer-flow-provisioner
|
|
volumes:
|
|
- ~/.kube/config:/root/.kube/config:ro
|
|
environment:
|
|
- K8S_NAMESPACE=deer-flow
|
|
- SANDBOX_IMAGE=enterprise-public-cn-beijing.cr.volces.com/vefaas-public/all-in-one-sandbox:latest
|
|
# Optional lark-cli init image (Pattern A). Empty ⇒ legacy runtime mount.
|
|
# Set to a published tag (e.g. deer-flow/lark-cli-init:v1.0.65) to provision
|
|
# the sandbox lark-cli runtime via an init container + emptyDir.
|
|
- LARK_CLI_INIT_IMAGE=${LARK_CLI_INIT_IMAGE:-}
|
|
# Optional lark-cli broker image (Pattern B, issue #4338). When set, the
|
|
# sandbox gets a shim + broker sidecar that holds the credentials, so the
|
|
# plaintext config/data are never mounted into the sandbox. Supersedes
|
|
# LARK_CLI_INIT_IMAGE when both are set. Empty ⇒ broker off.
|
|
- LARK_CLI_BROKER_IMAGE=${LARK_CLI_BROKER_IMAGE:-}
|
|
- THREADS_HOST_PATH=${DEER_FLOW_HOME}/threads
|
|
- DEER_FLOW_HOST_BASE_DIR=${DEER_FLOW_HOME}
|
|
- KUBECONFIG_PATH=/root/.kube/config
|
|
- NODE_HOST=host.docker.internal
|
|
- K8S_API_SERVER=https://host.docker.internal:26443
|
|
env_file:
|
|
- ../.env
|
|
extra_hosts:
|
|
- "host.docker.internal:host-gateway"
|
|
networks:
|
|
- deer-flow
|
|
restart: unless-stopped
|
|
healthcheck:
|
|
test: ["CMD", "curl", "-f", "http://localhost:8002/health"]
|
|
interval: 10s
|
|
timeout: 5s
|
|
retries: 6
|
|
volumes:
|
|
redis-data:
|
|
|
|
networks:
|
|
deer-flow:
|
|
driver: bridge
|