* bench: add provider-agnostic sandbox benchmark with BoxLite warm pool results - scripts/bench_sandbox_provider.py: CLI for measuring acquire/run/release across providers, scenarios, workloads, and concurrency levels - scripts/summarize_bench.py: JSONL aggregation with p50/p95/p99 tables - bench_results.jsonl: 110 turns across 7 scenarios on real BoxLite 0.9.7 Key findings: cold acquire: ~860ms warm reclaim: ~14ms (60x speedup) release: ~0ms warm_hit_rate: 95% (warm_same_thread) * perf(boxlite): skip health check for recently-released warm pool boxes Boxes released within health_check_skip_seconds (default 5.0s) are promoted directly without the ~14ms echo-ok round-trip. A VM alive seconds ago is overwhelmingly likely to still be alive. Add sandbox.health_check_skip_seconds config option. Set to 0 to always health-check (old behaviour). Benchmark (warm_same_thread, noop, 20 iters): acquire p50: 14.9ms → 0.0ms total p50: 29.9ms → 14.0ms * chore: move benchmark scripts into backend/scripts/benchmark/ * fix: address BoxLite benchmark review findings * fix(boxlite): only skip warm reclaim checks for released boxes * fix(benchmark): keep BoxLite shim workaround off the event loop * fix(boxlite): invalidate dead boxes from command path * test(boxlite): cover skip window and invalidation edge cases * fix(boxlite): treat sandbox-has-been-closed as terminal in _exec * fix(boxlite): harden warm-pool reclaim and benchmark accounting * fix(boxlite): validate warm-pool reclaims by default * fix(config): expose boxlite health-check skip setting * fix(boxlite): tighten failure classification and benchmark workaround * Update config_version to 21 in values.yaml --------- Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
3.7 KiB
BoxLite backend
Runs each DeerFlow sandbox as a BoxLite micro-VM — a daemonless, OCI-native VM with its own kernel (libkrun/KVM on Linux, Hypervisor.framework on macOS). Motivated by the resource/cold-start pain with the default AIO Docker sandbox in #3439 and #3213; discussion in #3936.
Configuration
sandbox:
use: deerflow.community.boxlite:BoxliteProvider
image: python:3.12-slim # any OCI image (default: python:3.12-slim)
memory_mib: 1024 # per-box memory cap (optional)
cpus: 2 # per-box vCPUs (optional)
replicas: 3 # active + warm VM cap per gateway process (default: 3)
idle_timeout: 600 # warm VM idle seconds before stop; 0 disables
health_check_skip_seconds: 0.0 # optional low-latency mode: skip reclaim
# health checks for recent releases; 0 keeps
# reliability-first validation (default: 0.0)
environment: # injected into every command
PYTHONUNBUFFERED: "1"
Install the optional runtime before selecting this provider:
pip install "deerflow-harness[boxlite]"
The boxlite package is an optional DeerFlow harness extra, not part of the
default install. It is also limited to the host platforms and architectures
where BoxLite publishes wheels and can boot micro-VMs. Unsupported development
hosts, such as Windows, should use another sandbox provider or run DeerFlow from
a supported Linux/macOS environment.
Host requirement: BoxLite boots micro-VMs, so a Linux host needs KVM — i.e. nested virtualization when DeerFlow runs inside a cloud VM. macOS uses Hypervisor.framework. This is the main deployment constraint to weigh vs. the container-based providers.
Design
DeerFlow's Sandbox contract is synchronous; BoxLite's SDK is async-native and
its box handles are event-loop-affine. The provider owns one private asyncio
loop on a daemon thread and marshals every coroutine onto it via
run_coroutine_threadsafe. BoxLite boxes are named deterministically from
user_id:thread_id, released into an in-process warm pool after each agent turn,
and reclaimed by the same thread on the next acquire.
| File | Role |
|---|---|
provider.py |
SandboxProvider lifecycle + the private-loop bridge |
box.py |
Sandbox adapter; execute_command + file ops |
Contract coverage
The full Sandbox surface is implemented. File operations run as shell commands
inside the box and reuse deerflow.sandbox.search, mirroring e2b_sandbox:
execute_command—sh -lc, with per-call env and timeout.read_file/write_file/update_file—catand chunkedbase64(binary-safe, no arg-size limit).download_file— 100 MB cap, restricted to the/mnt/user-dataprefix.list_dir/glob/grep—find/grepwith busybox-portable flags; results filtered/capped in Python.
The provider creates /mnt/user-data/{workspace,uploads,outputs} and
/mnt/skills on box start so those virtual paths resolve natively.
Warm-pool capacity is governed by sandbox.replicas across active + warm VMs.
sandbox.idle_timeout controls how long released warm VMs stay running; 0
disables idle reaping. Active boxes are never evicted to satisfy the cap.
Status
Verified end-to-end against a live box (provider resolution → execute_command
→ file ops) on macOS/HVF. Linux/KVM validation and benchmarks vs. the AIO
sandbox are tracked in
#3936.