Jeremy Schoemaker 6b4f803354
fix(sandbox): preserve trailing whitespace in filenames from list_dir and glob in remote providers (#4980)
* fix(sandbox): stop stripping filenames when parsing find output in remote providers

The list_dir and glob parsers in the e2b, OpenSandbox, AIO, Tenki, and
BoxLite providers called .strip() on every line of find output. A
filename that legitimately ends (or begins) in whitespace was corrupted,
so the listed path never resolved on any follow-up file API call, and
the remote providers diverged from LocalSandbox, which preserves such
names via pathlib.

splitlines() already removes the line terminators, so filter empty lines
only and keep each entry verbatim. Same class of bug as the e2b
_sync_outputs_to_host fix (#4861), applied to the search parsers.

Adds a trailing-space regression test per provider at the seam each
suite already uses.

* fix(sandbox): split find output on \n only, and rename the tenki test

Review follow-ups from willem-bd:

- aio_sandbox.list_dir used str.splitlines(), which also breaks records on
  \v, \f, \x1c-\x1e and \x85 - all legal inside a Linux filename, and all
  contrary to this PR's own rule that the newline is the only delimiter.
  find emits \n and nothing else, so split("\n") is the correct parse.
- Renamed test_search_preserves_trailing_space_in_filename to
  test_list_dir_and_glob_preserve_trailing_space_in_filename, matching the
  sibling tests in test_opensandbox_provider.py and test_boxlite_provider.py.
  The body covers list_dir and glob; it never touches grep.
2026-09-02 09:23:04 +08:00
..

BoxLite backend

Runs each DeerFlow sandbox as a BoxLite micro-VM — a daemonless, OCI-native VM with its own kernel (libkrun/KVM on Linux, Hypervisor.framework on macOS). Motivated by the resource/cold-start pain with the default AIO Docker sandbox in #3439 and #3213; discussion in #3936.

Configuration

sandbox:
  use: deerflow.community.boxlite:BoxliteProvider
  image: python:3.12-slim         # any OCI image (default: python:3.12-slim)
  memory_mib: 1024                # per-box memory cap (optional)
  cpus: 2                         # per-box vCPUs (optional)
  replicas: 3                     # active + warm VM cap per gateway process (default: 3)
  idle_timeout: 600               # warm VM idle seconds before stop; 0 disables
  health_check_skip_seconds: 0.0  # optional low-latency mode: skip reclaim
                                  #   health checks for recent releases; 0 keeps
                                  #   reliability-first validation (default: 0.0)
  environment:                    # injected into every command
    PYTHONUNBUFFERED: "1"

Install the optional runtime before selecting this provider:

pip install "deerflow-harness[boxlite]"

The boxlite package is an optional DeerFlow harness extra, not part of the default install. It is also limited to the host platforms and architectures where BoxLite publishes wheels and can boot micro-VMs. Unsupported development hosts, such as Windows, should use another sandbox provider or run DeerFlow from a supported Linux/macOS environment.

Host requirement: BoxLite boots micro-VMs, so a Linux host needs KVM — i.e. nested virtualization when DeerFlow runs inside a cloud VM. macOS uses Hypervisor.framework. This is the main deployment constraint to weigh vs. the container-based providers.

Design

DeerFlow's Sandbox contract is synchronous; BoxLite's SDK is async-native and its box handles are event-loop-affine. The provider owns one private asyncio loop on a daemon thread and marshals every coroutine onto it via run_coroutine_threadsafe. BoxLite boxes are named deterministically from user_id:thread_id, released into an in-process warm pool after each agent turn, and reclaimed by the same thread on the next acquire.

File Role
provider.py SandboxProvider lifecycle + the private-loop bridge
box.py Sandbox adapter; execute_command + file ops

Contract coverage

The full Sandbox surface is implemented. File operations run as shell commands inside the box and reuse deerflow.sandbox.search, mirroring e2b_sandbox:

  • execute_commandsh -lc, with per-call env and timeout.
  • read_file / write_file / update_filecat and chunked base64 (binary-safe, no arg-size limit).
  • download_file — 100 MB cap, restricted to the /mnt/user-data prefix.
  • list_dir / glob / grepfind / grep with busybox-portable flags; results filtered/capped in Python.

The provider creates /mnt/user-data/{workspace,uploads,outputs} and /mnt/skills on box start so those virtual paths resolve natively.

Warm-pool capacity is governed by sandbox.replicas across active + warm VMs. sandbox.idle_timeout controls how long released warm VMs stay running; 0 disables idle reaping. Active boxes are never evicted to satisfy the cap.

Status

Verified end-to-end against a live box (provider resolution → execute_command → file ops) on macOS/HVF. Linux/KVM validation and benchmarks vs. the AIO sandbox are tracked in #3936.