mirror of
https://github.com/bytedance/deer-flow.git
synced 2026-09-14 16:08:41 +00:00
* feat(gateway): add /health/ready readiness probe backed by the database
## Why
GET /health only proves the process is up: it returns 200 even when the persistence engine cannot reach the database. Orchestrators already treat it as a readiness gate (docker-compose.yaml marks the gateway service healthy and nginx depends_on service_healthy), so a DB outage or a still-migrating Postgres leaves the stack 'healthy' while every request fails.
## What changed
- New GET /health/ready endpoint: bounded SELECT 1 against the existing persistence engine (deerflow.persistence.engine.get_engine) with a 2s timeout.
- Response is 200 {'status': 'ready', 'database': 'ok'} when reachable, 503 {'status': 'degraded', 'database': 'unreachable'} when the probe fails, and 200 ready with database=not_configured for backend=memory (nothing to probe).
- GET /health is unchanged (pure liveness), and /health/ready is public through the existing /health auth whitelist.
- docker-compose.yaml gateway healthcheck now polls /health/ready so service_healthy reflects database reachability.
- Documented both endpoints in backend/app/gateway/AGENTS.md.
## Surface area
- [x] Backend API - new GET /health/ready endpoint under backend/app/gateway
- [x] Sandbox / Docker - gateway healthcheck in docker/docker-compose.yaml now gates on readiness
- [ ] Frontend UI / Agents / Skills / Dependencies
- [x] Default behavior change - existing /health unchanged; the prod compose healthcheck is stricter (503 while the database is unreachable)
## Validation
- New unit tests in backend/tests/test_gateway_health.py cover ok / unreachable / not_configured probe results and the 200/503 payload mapping (6 passed).
- app.gateway.app imports cleanly and registers both /health and /health/ready.
- ruff check + ruff format clean.
## AI assistance
**Tool(s) used:** Codex (coding agent)
**How you used it:** design, implementation, and unit tests produced with AI assistance; reviewed before commit.
- [ ] I've read and understand every line of this change and take responsibility for it — it's not unreviewed AI output.
* fix(helm): point the gateway readiness probe at /health/ready
## Why
Review on #5166 (willem-bd, P1): the chart still probed /health for readiness,
so Kubernetes marked the pod ready and routed traffic while the database was
unreachable - exactly the failure mode /health/ready was added to catch.
## What changed
- deploy/helm/deer-flow/templates/gateway-deployment.yaml: readinessProbe
httpGet.path now hits /health/ready (DB-backed, 503 while the database is
unreachable). The liveness probe stays on /health.
## Verification
- One-line path change inside the existing readinessProbe block; git diff
confirms only the readiness path changed (liveness untouched).
## AI assistance
**Tool(s) used:** Codex (coding agent)
**How you used it:** implemented the reviewer-requested probe path change; reviewed before commit.
- [ ] I've read and understand every line of this change and take responsibility for it — it's not unreviewed AI output.
* fix(gateway): readiness probe also checks the effective checkpointer/Store backend
## Why
Follow-up review on #5166 (willem-bd, P1): get_engine() only represents the
ORM backend selected by `database:`. The legacy `checkpointer:` section takes
precedence for the LangGraph checkpointer and Store, so a split configuration
(a local SQLite/memory `database:` with `checkpointer.type: postgres`) could
report 200 while the PostgreSQL backend agent runs depend on was down.
## What changed
- GET /health/ready now probes both persistence halves: the ORM engine behind
`database:` (unchanged) and the effective LangGraph checkpointer/Store
backend resolved with the runtime's own rule (legacy `checkpointer:` config,
otherwise derived from `database:`), for memory/sqlite/postgres backends.
- The payload gains a `checkpointer` field with the same
ok / not_configured / unreachable vocabulary as `database`; 503 degraded is
returned when either probe is unreachable.
- Probes are bounded by the existing 2s timeout: sqlite via aiosqlite SELECT 1
on the resolved path, postgres via a bounded psycopg AsyncConnection SELECT 1
on the DSN with the configured search_path. A missing driver for a configured
backend degrades readiness (the runtime could not run either).
- Documented the two-probe semantics in the endpoint docstring and
backend/app/gateway/AGENTS.md.
## Verification
- New tests: healthy ORM engine + unreachable legacy checkpointer backend ->
503 degraded with database: ok / checkpointer: unreachable; checkpointer
probe mapping for memory/sqlite(postgres missing-driver) backends; existing
payload tests now pin the checkpointer field.
- cd backend && python -m pytest tests/test_gateway_health.py: 11 passed.
- app.gateway.app imports cleanly; ruff check + ruff format clean.
## AI assistance
**Tool(s) used:** Codex (coding agent)
**How you used it:** design, implementation, and unit tests produced with AI assistance; reviewed before commit.
- [ ] I've read and understand every line of this change and take responsibility for it — it's not unreviewed AI output.
* fix(gateway): bound /health/ready to one deadline and probe the startup checkpointer snapshot
## Why
Second round of review on #5166 (zhfeng P1/P2, willem-bd P1/P1/P2). Three
correctness issues remained in the readiness endpoint:
- The database and checkpointer probes ran sequentially, each allowed 2s, so a
healthy response could take almost 4s - past Kubernetes' 1s default
readinessProbe timeout and inside Docker's 3s client timeout. A slow but
healthy backend could make every replica unready.
- The checkpointer probe re-resolved process-wide, hot-reloaded configuration
per request, while app.state.checkpointer/store are built once from the
startup_config snapshot in langgraph_runtime(). After a live config edit the
endpoint could probe a backend the running gateway does not use, and a
resolution failure was swallowed into None -> not_configured -> 200.
- The SQLite probe opened the path with aiosqlite.connect(), which creates the
file when missing: a deleted checkpoint database was silently resurrected as
an empty file and reported ok instead of surfacing the outage.
## What changed
- backend/app/gateway/health.py: the two probes now run concurrently beneath a
single endpoint-wide deadline (_READINESS_DEADLINE_SECONDS=3.0) so a healthy
response completes within one probe window (~2s), never the sum of both.
A probe that overruns the deadline degrades the endpoint instead of hanging.
- langgraph_runtime() now records the checkpointer/Store config resolved from
the same startup_config snapshot its checkpointer/store singletons are built
from (app.state.checkpointer_config); /health/ready probes that snapshot and
never re-resolves hot-reloaded config. resolve_checkpointer_config() returns
None on resolution failure and the endpoint fails closed (503, checkpointer:
unreachable) instead of reporting not_configured.
- The SQLite probe opens disk-backed paths with the non-creating mode=rw URI
flag, so a missing database file stays missing and yields unreachable;
in-memory forms (:memory:, file:...mode=memory) have nothing external to
probe and report not_configured like the memory backend.
- Orchestrator timeouts now sit above the endpoint bound: Helm readinessProbe
gains timeoutSeconds: 5 (Kubernetes default is 1s) and the docker-compose
gateway healthcheck client timeout moves from 3s to 5s.
## Verification
- New regression tests: concurrent probes keep total elapsed time within one
probe window; a probe ignoring its budget trips the endpoint deadline to 503;
missing SQLite file stays absent and yields unreachable; in-memory SQLite
forms map to not_configured; missing startup snapshot / config resolution
failure fail closed to 503; resolve_checkpointer_config() raising is covered.
- cd backend && python -m pytest tests/test_gateway_health.py: 21 passed;
tests/test_gateway_docs_toggle.py and lifespan/shutdown gateway suites pass.
- ruff check + ruff format clean on all changed files.
## AI assistance
**Tool(s) used:** Codex (coding agent)
**How you used it:** implemented the reviewer-requested concurrency/deadline, startup-snapshot probing, fail-closed resolution, and non-creating SQLite probe; reviewed before commit.
- [ ] I've read and understand every line of this change and take responsibility for it — it's not unreviewed AI output.
* fix(gateway): serialize connection-opening readiness probes behind a strict gate
## Why
Review on #5166 (willem-bd, P1): every request to /health/ready opened a new
PostgreSQL connection in _probe_postgres_backend, outside both the ORM pool and
the runtime checkpointer pool. The route is public through the /health auth
prefix and nginx proxies /health/*, so concurrent unauthenticated requests
could create an unbounded number of connections (each held for up to two
seconds), exhaust PostgreSQL max_connections, and take down both normal
traffic and the readiness probe itself.
## What changed
- backend/app/gateway/health.py: connection-opening checkpointer probes
(sqlite connect, postgres AsyncConnection.connect) now run inside a strict
per-process gate - an asyncio.Lock cached per running event loop - so at
most one probe connection can be in flight per worker process. Requests
that queue behind the gate are still shed by the existing endpoint-wide
deadline, so a flood cannot pile up new connections or open files.
- Memory and unknown-backend decisions stay outside the gate; payload and
probe semantics are unchanged. The serialization is documented in the
module docstring and backend/app/gateway/AGENTS.md.
## Verification
- New regression test: 8 concurrent readiness_payload() requests against an
instrumented sqlite probe assert the maximum number of in-flight probe
connections is 1 while every request still returns 200.
- cd backend && python -m pytest tests/test_gateway_health.py: 22 passed;
tests/test_gateway_docs_toggle.py and tests/test_gateway_lifespan_shutdown.py
also pass on the merged main head.
- ruff check + ruff format clean on all changed files.
## AI assistance
**Tool(s) used:** Codex (coding agent)
**How you used it:** implemented the reviewer-requested strict concurrency bound for the public readiness probe; reviewed before commit.
- [ ] I've read and understand every line of this change and take responsibility for it — it's not unreviewed AI output.
210 lines
9.7 KiB
YAML
210 lines
9.7 KiB
YAML
# DeerFlow Production Environment
|
|
# Usage: make up
|
|
#
|
|
# Services:
|
|
# - nginx: Reverse proxy (port 2026, configurable via PORT env var)
|
|
# - frontend: Next.js production server
|
|
# - gateway: FastAPI Gateway API + agent runtime
|
|
# - redis: Redis Streams backend for cross-worker SSE stream bridge
|
|
# - provisioner: (optional) Sandbox provisioner for Kubernetes mode
|
|
#
|
|
# Key environment variables (set via environment/.env or scripts/deploy.sh):
|
|
# DEER_FLOW_PROJECT_ROOT — project root for relative runtime paths
|
|
# DEER_FLOW_HOME — runtime data dir, default .deer-flow under $DEER_FLOW_PROJECT_ROOT (or cwd)
|
|
# DEER_FLOW_CONFIG_PATH — path to config.yaml
|
|
# DEER_FLOW_EXTENSIONS_CONFIG_PATH — path to extensions_config.json
|
|
# DEER_FLOW_SKILLS_PATH — skills dir, default $DEER_FLOW_PROJECT_ROOT/skills
|
|
# DEER_FLOW_DOCKER_SOCKET — Docker socket path for aio/DooD mode, default /var/run/docker.sock (used only by the opt-in docker-compose.dood.yaml overlay)
|
|
# DEER_FLOW_REPO_ROOT — repo root (used for skills host path in DooD)
|
|
# BETTER_AUTH_SECRET — required for frontend auth/session security
|
|
# DEER_FLOW_INTERNAL_AUTH_TOKEN — shared internal Gateway auth token for multi-worker IM channels
|
|
#
|
|
# LangSmith tracing is disabled by default (LANGSMITH_TRACING=false).
|
|
# Set LANGSMITH_TRACING=true and LANGSMITH_API_KEY in .env to enable it.
|
|
#
|
|
# Access: http://localhost:${PORT:-2026}
|
|
|
|
services:
|
|
# ── Redis Stream Bridge ────────────────────────────────────────────────────
|
|
redis:
|
|
image: redis:7-alpine
|
|
container_name: deer-flow-redis
|
|
command: ["redis-server", "--appendonly", "yes"]
|
|
volumes:
|
|
- redis-data:/data
|
|
healthcheck:
|
|
test: ["CMD", "redis-cli", "ping"]
|
|
interval: 5s
|
|
timeout: 3s
|
|
retries: 10
|
|
networks:
|
|
- deer-flow
|
|
restart: unless-stopped
|
|
|
|
# ── Reverse Proxy ──────────────────────────────────────────────────────────
|
|
nginx:
|
|
image: nginx:alpine
|
|
container_name: deer-flow-nginx
|
|
# Loopback-only by default: DeerFlow's agent can execute commands, so the
|
|
# documented default deployment is a local trusted environment. A bare
|
|
# "${PORT}:2026" would bind 0.0.0.0 instead, which does not match that.
|
|
# Set BIND_HOST=0.0.0.0 only behind your own TLS/auth front door, and
|
|
# complete first-run setup before the host becomes reachable.
|
|
ports:
|
|
- "${BIND_HOST:-127.0.0.1}:${PORT:-2026}:2026"
|
|
volumes:
|
|
- ./nginx/nginx.conf:/etc/nginx/nginx.conf.template:ro
|
|
command: >
|
|
sh -c "cp /etc/nginx/nginx.conf.template /etc/nginx/nginx.conf
|
|
&& nginx -g 'daemon off;'"
|
|
depends_on:
|
|
frontend:
|
|
condition: service_started
|
|
gateway:
|
|
condition: service_healthy
|
|
networks:
|
|
- deer-flow
|
|
restart: unless-stopped
|
|
|
|
# ── Frontend: Next.js Production ───────────────────────────────────────────
|
|
frontend:
|
|
build:
|
|
context: ../
|
|
dockerfile: frontend/Dockerfile
|
|
target: prod
|
|
args:
|
|
PNPM_STORE_PATH: ${PNPM_STORE_PATH:-/root/.local/share/pnpm/store}
|
|
NPM_REGISTRY: ${NPM_REGISTRY:-}
|
|
container_name: deer-flow-frontend
|
|
environment:
|
|
- BETTER_AUTH_SECRET=${BETTER_AUTH_SECRET}
|
|
- DEER_FLOW_INTERNAL_GATEWAY_BASE_URL=http://gateway:8001
|
|
env_file:
|
|
- ../frontend/.env
|
|
networks:
|
|
- deer-flow
|
|
restart: unless-stopped
|
|
|
|
# ── Gateway API ────────────────────────────────────────────────────────────
|
|
gateway:
|
|
build:
|
|
context: ../
|
|
dockerfile: backend/Dockerfile
|
|
args:
|
|
APT_MIRROR: ${APT_MIRROR:-}
|
|
UV_IMAGE: ${UV_IMAGE:-ghcr.io/astral-sh/uv:0.11.1}
|
|
UV_INDEX_URL: ${UV_INDEX_URL:-https://pypi.org/simple}
|
|
UV_EXTRAS: ${UV_EXTRAS:-}
|
|
NPM_REGISTRY: ${NPM_REGISTRY:-}
|
|
LARK_CLI_NPM_VERSION: ${LARK_CLI_NPM_VERSION:-1.0.65}
|
|
container_name: deer-flow-gateway
|
|
# Gateway hosts the agent runtime with in-process RunManager + StreamBridge
|
|
# singletons -- run state lives in this worker's memory. Default to a single
|
|
# worker: the Redis stream bridge solves cross-worker SSE delivery and
|
|
# reconnect, but run cancel, request dedup, and per-worker IM channel
|
|
# services remain worker-local. Override GATEWAY_WORKERS only when the
|
|
# Redis stream bridge is enabled and those limitations are acceptable.
|
|
command: sh -c "cd backend && PYTHONPATH=. uv run --no-sync uvicorn app.gateway.app:app --host 0.0.0.0 --port 8001 --workers ${GATEWAY_WORKERS:-1}"
|
|
volumes:
|
|
- ${DEER_FLOW_CONFIG_PATH}:/app/backend/config.yaml:ro
|
|
# Writable on purpose: the Gateway edits this file at runtime (MCP
|
|
# enable/disable, PUT/PATCH /api/mcp/config, skill updates). config.yaml
|
|
# above stays read-only because no API writes it.
|
|
- ${DEER_FLOW_EXTENSIONS_CONFIG_PATH}:/app/backend/extensions_config.json
|
|
- ../skills:/app/skills:ro
|
|
- ${DEER_FLOW_HOME}:/app/backend/.deer-flow
|
|
# DooD: the host Docker socket is NOT mounted by default. It is added only
|
|
# for aio (pure-DooD) sandbox mode via the opt-in docker-compose.dood.yaml
|
|
# overlay (appended by scripts/deploy.sh). See backend/docs/CONFIGURATION.md
|
|
# Security Note section for details.
|
|
|
|
# CLI auth dirs (Claude Code / Codex) are NOT mounted by default: they
|
|
# expose the entire ~/.claude and ~/.codex (history, projects, global
|
|
# config, credentials) into the container. Mount them only when you use
|
|
# the Claude/Codex CLI login as a model provider or ACP agent, via the
|
|
# opt-in docker-compose.cli-auth.yaml overlay. Prefer an env token
|
|
# (CLAUDE_CODE_OAUTH_TOKEN, see .env.example / backend/docs/CONFIGURATION.md).
|
|
working_dir: /app
|
|
environment:
|
|
- CI=true
|
|
- DEER_FLOW_PROJECT_ROOT=/app
|
|
- DEER_FLOW_HOME=/app/backend/.deer-flow
|
|
- DEER_FLOW_CONFIG_PATH=/app/backend/config.yaml
|
|
- DEER_FLOW_EXTENSIONS_CONFIG_PATH=/app/backend/extensions_config.json
|
|
- DEER_FLOW_STREAM_BRIDGE_REDIS_URL=${DEER_FLOW_STREAM_BRIDGE_REDIS_URL:-redis://redis:6379/0}
|
|
- DEER_FLOW_CHANNELS_LANGGRAPH_URL=${DEER_FLOW_CHANNELS_LANGGRAPH_URL:-http://gateway:8001/api}
|
|
- DEER_FLOW_CHANNELS_GATEWAY_URL=${DEER_FLOW_CHANNELS_GATEWAY_URL:-http://gateway:8001}
|
|
- DEER_FLOW_INTERNAL_AUTH_TOKEN=${DEER_FLOW_INTERNAL_AUTH_TOKEN}
|
|
# DooD path/network translation
|
|
- DEER_FLOW_HOST_BASE_DIR=${DEER_FLOW_HOME}
|
|
- DEER_FLOW_SANDBOX_HOST=host.docker.internal
|
|
# Proxy values (HTTP_PROXY/HTTPS_PROXY/ALL_PROXY) are inherited from ../.env via env_file.
|
|
# Only NO_PROXY is declared here so internal service hostnames are always exempt from the proxy.
|
|
- NO_PROXY=${NO_PROXY:-}${NO_PROXY:+,}localhost,127.0.0.1,::1,gateway,frontend,nginx,provisioner,openviking,host.docker.internal
|
|
- no_proxy=${no_proxy:-}${no_proxy:+,}localhost,127.0.0.1,::1,gateway,frontend,nginx,provisioner,openviking,host.docker.internal
|
|
env_file:
|
|
- ../.env
|
|
extra_hosts:
|
|
- "host.docker.internal:host-gateway"
|
|
depends_on:
|
|
redis:
|
|
condition: service_healthy
|
|
healthcheck:
|
|
# /health/ready runs both probes concurrently within a 3s endpoint
|
|
# deadline; the client timeout must stay above that bound.
|
|
test: ["CMD", "python", "-c", "import urllib.request; response = urllib.request.urlopen('http://127.0.0.1:8001/health/ready', timeout=5); raise SystemExit(0 if response.status == 200 else 1)"]
|
|
interval: 5s
|
|
timeout: 5s
|
|
retries: 30
|
|
start_period: 20s
|
|
networks:
|
|
- deer-flow
|
|
restart: unless-stopped
|
|
|
|
# ── Sandbox Provisioner (optional, Kubernetes mode) ────────────────────────
|
|
provisioner:
|
|
build:
|
|
context: ./provisioner
|
|
dockerfile: Dockerfile
|
|
args:
|
|
APT_MIRROR: ${APT_MIRROR:-}
|
|
PIP_INDEX_URL: ${PIP_INDEX_URL:-}
|
|
container_name: deer-flow-provisioner
|
|
volumes:
|
|
- ~/.kube/config:/root/.kube/config:ro
|
|
environment:
|
|
- K8S_NAMESPACE=deer-flow
|
|
- SANDBOX_IMAGE=enterprise-public-cn-beijing.cr.volces.com/vefaas-public/all-in-one-sandbox:latest
|
|
# Optional lark-cli init image (Pattern A). Empty ⇒ legacy runtime mount.
|
|
# Set to a published tag (e.g. deer-flow/lark-cli-init:v1.0.65) to provision
|
|
# the sandbox lark-cli runtime via an init container + emptyDir.
|
|
- LARK_CLI_INIT_IMAGE=${LARK_CLI_INIT_IMAGE:-}
|
|
# Optional lark-cli broker image (Pattern B, issue #4338). When set, the
|
|
# sandbox gets a shim + broker sidecar that holds the credentials, so the
|
|
# plaintext config/data are never mounted into the sandbox. Supersedes
|
|
# LARK_CLI_INIT_IMAGE when both are set. Empty ⇒ broker off.
|
|
- LARK_CLI_BROKER_IMAGE=${LARK_CLI_BROKER_IMAGE:-}
|
|
- THREADS_HOST_PATH=${DEER_FLOW_HOME}/threads
|
|
- DEER_FLOW_HOST_BASE_DIR=${DEER_FLOW_HOME}
|
|
- KUBECONFIG_PATH=/root/.kube/config
|
|
- NODE_HOST=host.docker.internal
|
|
- K8S_API_SERVER=https://host.docker.internal:26443
|
|
env_file:
|
|
- ../.env
|
|
extra_hosts:
|
|
- "host.docker.internal:host-gateway"
|
|
networks:
|
|
- deer-flow
|
|
restart: unless-stopped
|
|
healthcheck:
|
|
test: ["CMD", "curl", "-f", "http://localhost:8002/health"]
|
|
interval: 10s
|
|
timeout: 5s
|
|
retries: 6
|
|
volumes:
|
|
redis-data:
|
|
|
|
networks:
|
|
deer-flow:
|
|
driver: bridge
|