* ci(helm): publish chart to charts/ namespace prefix on GHCR
Push the Helm chart to ghcr.io/<owner>/charts/deer-flow (via a `charts/`
prefix on the `helm push` target) instead of the bare `deer-flow`
package. This namespaces the chart apart from the image packages
(deer-flow-{backend,frontend,provisioner}) without renaming the chart:
Chart.yaml `name` stays `deer-flow`, so dir = chart name = in-cluster
resource/selector names, and no `nameOverride` hack is needed.
The chart is new in 2.1.0 (chart infra landed in #3987, after v2.0.0),
so 2.1.0 is the first chart release. Early nightly builds remain at the
legacy non-prefixed ghcr.io/<owner>/deer-flow.
Also refresh chart release docs:
- Replace the removed scripts/build-and-push.sh in the chart README with
raw `docker build`/`push` commands (contexts/args match container.yaml).
- Point NOTES.txt's empty-registry warning at the README section instead
of the removed script.
- Retarget RELEASING.md version examples from 2.2.0 to 2.1.0.
- Bump the chart README's helm prerequisite to 3+.
* docs(helm): require helm 3.8+ and note legacy chart package cleanup
Address PR #4175 review:
- bump documented minimum to helm 3.8 (OCI registry support stabilized
there; earlier 3.x needs HELM_EXPERIMENTAL_OCI=1)
- add a post-release note to delete/revoke the legacy bare
ghcr.io/<owner>/deer-flow chart package after 2.1.0
* fix subagent total delegation cap
* fix embedded subagent run cap context
* fix subagent cap config consistency
* fix resumed subagent run cap boundary
* fix legacy resume subagent boundary
* address subagent cap review feedback
---------
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
* bench: add provider-agnostic sandbox benchmark with BoxLite warm pool results
- scripts/bench_sandbox_provider.py: CLI for measuring acquire/run/release
across providers, scenarios, workloads, and concurrency levels
- scripts/summarize_bench.py: JSONL aggregation with p50/p95/p99 tables
- bench_results.jsonl: 110 turns across 7 scenarios on real BoxLite 0.9.7
Key findings:
cold acquire: ~860ms
warm reclaim: ~14ms (60x speedup)
release: ~0ms
warm_hit_rate: 95% (warm_same_thread)
* perf(boxlite): skip health check for recently-released warm pool boxes
Boxes released within health_check_skip_seconds (default 5.0s)
are promoted directly without the ~14ms echo-ok round-trip.
A VM alive seconds ago is overwhelmingly likely to still be alive.
Add sandbox.health_check_skip_seconds config option.
Set to 0 to always health-check (old behaviour).
Benchmark (warm_same_thread, noop, 20 iters):
acquire p50: 14.9ms → 0.0ms
total p50: 29.9ms → 14.0ms
* chore: move benchmark scripts into backend/scripts/benchmark/
* fix: address BoxLite benchmark review findings
* fix(boxlite): only skip warm reclaim checks for released boxes
* fix(benchmark): keep BoxLite shim workaround off the event loop
* fix(boxlite): invalidate dead boxes from command path
* test(boxlite): cover skip window and invalidation edge cases
* fix(boxlite): treat sandbox-has-been-closed as terminal in _exec
* fix(boxlite): harden warm-pool reclaim and benchmark accounting
* fix(boxlite): validate warm-pool reclaims by default
* fix(config): expose boxlite health-check skip setting
* fix(boxlite): tighten failure classification and benchmark workaround
* Update config_version to 21 in values.yaml
---------
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
* feat(helm): add production-ready Helm chart for Kubernetes deployment
Adds deploy/helm/deer-flow, a native-Kubernetes translation of the
production docker-compose stack, plus CI to publish its images and chart.
* ci(release): gate releases on version-source consistency
Add a reusable verify-versions workflow invoked by both chart.yaml and
container.yaml on v* tags. It runs scripts/verify_versions.sh against the
tag and fails the release — skipping all image and chart publishing — when
Chart.yaml (version + appVersion), backend/pyproject.toml, or
frontend/package.json don't all match the tag.
Add scripts/verify_versions.sh (the check, also runnable locally) and
scripts/bump_version.sh (bumps all four sources in lockstep, then
self-verifies). Document the release flow in RELEASING.md and link it from
AGENTS.md.
* fix(deploy): address Helm chart review feedback (#3987)
Three review items from willem-bd:
1. nginx IPv6 listen strip never matched. The sed pattern required a `;`
immediately after `2026`, but the rendered config emits
`listen [::]:2026 default_server;` (space + `default_server` before the
`;`), so the line was never deleted and nginx crash-looped on pods
without IPv6 (`socket() :::2026 failed (97: Address family not
supported)`). Drop the trailing `;` from the pattern so it matches.
Same latent bug fixed in docker-compose-dev.yaml.
2. Passwords were spliced into DSNs verbatim, so a password containing
URL-special chars (@ : / # ? % [ ] space) produced a malformed DSN and
a confusing parse error. Add a `deer-flow.urlEscape` helper
(replace-based: Sprig lacks urlqueryescape, and regexReplaceAllLiteral
treats the replacement as a regex template so `[`/`]`/`?` break it) and
apply it to the password in the postgres and redis DSNs. The raw
`postgres-password` / `redis-password` keys stay unencoded - they back
POSTGRES_PASSWORD / REDIS_PASSWORD, not a URL segment.
3. NODE_HOST defaulted to "gateway", which can never route: the gateway
Service is ClusterIP:8001 and knows nothing of a sandbox NodePort, so a
user who skips the caveat gets unreachable sandboxes with no error at
install time. Default NODE_HOST to the provisioner pod's node IP via
the downward API (status.hostIP) - a NodePort is exposed on every node,
so <node-IP>:<NodePort> routes from the gateway on most clusters.
`provisioner.nodeHost` remains an override for CNIs/policies that block
pod->node-IP traffic. Updated NOTES.txt, values.yaml, and the chart
README. (#3929 remains the long-term fix - ClusterIP + cluster-DNS URL
removes NODE_HOST and the NodePort exposure entirely.)
Validated with helm lint, helm template (incl. a special-char password
rendering the encoded DSNs), and a sed pattern-match check.
* fix(deploy): address round-2 Helm chart review feedback (#3987)
Three "Medium" items from willem-bd:
1. No helm lint / helm template gate before publish. A template regression
ships as an immutable OCI artifact (GHCR won't overwrite --version), so
gate packaging on `helm lint` + `helm template --include-crds` in
chart.yaml before `helm package`. (ct lint / helm-unittest deferred.)
2. Action pinning inconsistent + PR body overstates it. SHA-pin
actions/checkout (v6.0.3, df4cb1c0) and actions/attest-build-provenance
(v2.4.0, e8998f94) across the publishing workflows (chart.yaml,
container.yaml, verify-versions.yml), matching the existing docker/*
SHA-pin pattern. Resolves the checkout @v4/@v6 mismatch and makes the
"SHA-pinned actions" claim accurate. Other pre-existing workflows left
untouched (out of scope for this PR).
3. Provisioner RBAC broader than needed. Dropped the unused update/patch
verbs and the pods/exec + events rules from the provisioner Role -
audited against docker/provisioner/app.py, which only calls
get/create/delete on pods and get/list/create/delete on services. Fixed
NOTES.txt to accurately describe the grant instead of understating it as
"create Pods and Services". The remaining scope concern - verbs apply to
all Pods in the namespace, not just sandbox Pods - is still deferred
(RBAC can't scope by label; needs a dedicated namespace or admission
control), now noted in NOTES.txt and README.
Validated with helm lint + helm template (narrowed Role renders with
exactly get/list/watch/create/delete).
* feat(helm): enable sandbox+web tools out of the box
The chart's default config loaded zero agent tools (config.tools empty ->
"Total tools loaded: 0"), so a fresh install gave an agent that could do
nothing useful. Add tool_groups + tools to the default config block:
- web: web_search (ddg), web_fetch (jina), image_search - no API key
- file:read: ls, read_file, glob, grep
- file:write: write_file, str_replace
- bash
The file/bash tools run inside the AIO sandbox the chart already
configures; the web tools need outbound internet from the gateway pod
(swap backends or drop entries for air-gapped clusters - see
config.example.yaml).
Also bump config_version 15 -> 19 to match config.example.yaml (the chart
had drifted behind). NOTES.txt and the README example updated to match.
* ci(helm): add chart validation + config_version drift check on PR
Extend the chart workflow with a PR-triggered validate-chart job that runs
helm lint, helm template --include-crds, and a config_version drift check:
it parses config_version from both config.example.yaml and the chart's
values.yaml and fails the build (with a ::error:: naming the files to bump)
if the chart is behind the example. This catches the kind of drift this
PR is fixing - the chart sat at v15 while the example moved to v19 - before
it can merge again.
verify-versions and publish-chart stay tag-only; publish-chart now
needs: [verify-versions, validate-chart]. validate-chart runs on both
PRs and tag pushes: the tag arm is required because a job that `needs`
a skipped job is itself skipped under the default success() check, so
validate-chart must actually run on tag pushes or publish-chart would
never fire.
* Bump config version to 20