deer-flow/backend/tests/test_local_sandbox_command_timeout.py
Zeren Wang a06a6fed7e
feat(harness): deterministic acceptance checklist for subagent delegations (RFC #4651, layer 2) (#5109)
* feat(harness): deterministic acceptance checklist for subagent delegations (RFC #4651, layer 2)

PR4 of RFC #4651: check lead-supplied acceptance_criteria in code when a
subagent completes, so objectively checkable requirements can never be
silently passed by a self-report.

- subagents/acceptance_checks.py: deterministic leaf families —
  file:<path> exists|non-empty and file_written:<path> read through
  read_current_file_content scoped to the shared thread workspace; the
  read uses the sandbox-native virtual path form (the local read
  validator and provider mount tables resolve /mnt/user-data/... paths,
  not host paths); the scope decision canonicalizes with realpath on the
  local sandbox so workspace symlinks cannot escape into uploads; a
  remote provider's "Error: ..." return string is normalized to a
  failed check (provider-typed via is_local_sandbox); a
  UnicodeDecodeError marks a binary deliverable as existing and
  non-empty; out-of-scope paths degrade to UNVERIFIED.
  tests_passed:<command> anchors to a matching recorded bash execution
  with status=success and a test-summary shape; matching is
  shell-structure aware with control-flow attribution (span must end at
  the last segment with provable execution), negating-option values are
  ineligible evidence and a target negated anywhere in the command
  degrades the match, extra flags must be selection-preserving, extra
  positionals widen only after a path-scoped criterion, truncated
  commands degrade via command_truncated, the summary shape is read
  only from output attributable to the matched segment (preceding
  segments provably silent by invocation form), and pass shapes require
  a nonzero passed count. Criterion text is neutralized with
  neutralize_untrusted_tags before storage/rendering. Anything else
  renders UNVERIFIED, never silently passed.
- executor: accumulate bounded bash command/output evidence per streamed
  chunk (merged by tool_call_id, newest-capped) so subagent
  summarization compacting earlier messages cannot erase a recorded
  execution; the recorded status is the actual shell exit status parsed
  from the output's exit marker (signed codes included; the remote
  Command exited with code N form is accepted only as the whole trimmed
  output), falling back to deerflow_tool_meta only when no marker
  exists.
- sandbox providers: e2b/opensandbox/tenki/boxlite append the
  LocalSandbox-style "Exit Code: N" marker on nonzero exit even with
  non-empty output; aio propagates the SDK's structured exit_code on
  both exec paths the same way; local timeouts append Exit Code: 124;
  and _truncate_bash_output always preserves a trailing exit marker
  (signed included) inside its budget, with a 32-char floor raising any
  smaller configured limit, so the actual shell outcome always survives
  in the output text.
- task_tool: run the checklist offloaded (asyncio.to_thread) on the
  completed branch, failure-isolated; stamp the verdict into result
  metadata and render the per-criterion section into the model-visible
  result text.
- status contract: additive subagent_acceptance_verdict transport with
  read-side structural validation.
- delegation ledger: entry carries the verdict and renders a compact
  acceptance segment; gateway strips caller-forged verdicts from both
  ledger entries and message metadata, like the citation verdict.
- blocking-IO anchor pins the offload (teeth proven red->green); leaf
  read errors catch only OSError/SandboxError so unexpected errors reach
  the task-tool-level isolation instead of being mislabeled.

* fix(harness): close acceptance evidence gaps from review (RFC #4651 PR4)

- negating options: overlap with a matched criterion target is now
  checked by path/nodeid prefix, not exact token equality — excluding a
  sub-path of the criterion's selection (pytest tests --deselect
  tests/unit/test_auth.py) degrades to UNVERIFIED instead of holds
- output attribution: any redirection token in the matched final segment
  makes the recorded tail non-attributable (> / >> / 2> are word
  characters to the parser, so redirection was invisible to the matcher)
- silent-source allowlist narrowed from any *activate suffix to the
  */bin/activate shape
- status_contract docstring: restore the shared-fixture sentence and
  note subagent_acceptance_verdict is deliberately outside the fixture
- executor: update_bash_executions publishes [] (stream carried no
  bash-family calls) instead of collapsing it into None, mirroring
  update_tool_receipts

* fix(harness): close acceptance residual gaps from re-review (RFC #4651 PR4)

- tests_passed: add error outcomes to the fail shapes — "4 passed, 1 error"
  and pytest's "ERROR <nodeid>" short summary no longer satisfy the pass
  shape when the exit status is swallowed (|| true) or absent; zero-error
  counts stay clean.
- file leaves: bound the deliverable read — a "wc -c" shell size probe
  answers files above 50k bytes without loading ~2x their size, honoring
  the host-bash kill switch and falling back to the full read on any
  non-integer rendering, so verdicts never get less sound.
- executor: record the exit marker text as status_marker on harvested bash
  evidence; the leaf detail now reports the marker actually seen instead of
  asserting a failure indistinguishable from the command's own trailing text.
- extend the blocking-IO anchor to drive the probe branch inside the
  offload; teeth re-verified red->green.

* fix(harness): close acceptance forgery and bound gaps from P2 re-review (RFC #4651 PR4)

- file leaves: never read unbounded — size is established first (os.stat on
  the validated local host path, so the host-bash-disabled configuration
  needs no shell; a guarded wc -c on remote providers that renders
  missing/unreadable in its own words). Above the 50k cap the leaf answers
  from the size alone, at/below it the full read runs, and an
  unestablishable size degrades to UNVERIFIED instead of an unlimited
  fallback read.
- output attribution: source/. prefixes are never provably silent — a
  crafted */bin/activate path shape says nothing about what the script
  prints, so sourced segments can no longer lend a passing summary.
- executable identity: an explicitly path-spelled criterion now requires
  the same normalized executable path; the basename rule stays only for
  deliberately bare criterion commands.

* fix(harness): run acceptance size probe outside subagent-controlled state (RFC #4651 PR4)

- remote probe no longer runs in the sandbox's persistent shell: a fresh
  env -i /bin/sh with absolute-path stat/realpath (poisoned functions,
  aliases, PATH, exported functions, IFS, locale cannot steer it), plus a
  marker env routing AIO onto a fresh per-call bash.exec session.
- metadata-only: stat never opens content, so a FIFO deliverable cannot
  block the parent for the provider's idle timeout; non-regular files
  (fifo/dir/symlink) degrade to UNVERIFIED.
- containment canonicalized against the literal mount root: a
  final-component symlink or a swapped parent directory (root included)
  cannot redirect the check outside shared storage; unprovable layouts
  degrade to UNVERIFIED.

* fix(harness): canonicalize probe containment against the canonical mount root (RFC #4651 PR4)

Literal-root equality made every remote file leaf permanently UNVERIFIED
on e2b and Tenki, which realize /mnt/user-data as a symlink to the home
dir by default (e2b bootstrap 'sudo ln -sfn', Tenki best-effort symlink).
Containment now compares the file's realpath against the mount root's
realpath — exactly what the provider's own read path resolves, so probe
and read-back stay consistent; final-component symlinks stay rejected by
the non-dereferencing stat, and an intermediate dir-link escape under a
sane root still lands ESCAPED. The inner script is a module constant and
the suite now executes the composed probe for real against on-disk
layouts (real dir, symlinked prefix, final symlink, fifo, missing,
dir-link escape), which the canned-output stub could not see.

* fix(harness): close bare-criterion negation and CDPATH summary channels (RFC #4651 PR4)

- matching: a criterion with no positional selection target (bare pytest,
  make test) stands for the runner's default selection, so ANY negating
  option (--ignore/--deselect/...) makes the recorded run a different
  selection — unprovable. The overlap guard only sees consumed criterion
  tokens, which a bare criterion does not have; scoped criteria keep the
  unrelated-exclusion behavior.
- attribution: cd is no longer blanket-silent — CDPATH makes cd print the
  resolved (subagent-chosen) destination and the pass shapes match as
  substrings, so one mkdir 'all tests passed' plus an export minted a pass
  for any quiet command. A cd argument or CDPATH= value (export or leading
  assignment) carrying any summary shape makes the segment non-silent;
  shape-free cd dir wrappers keep matching.
- docs: _truncate_bash_output states the effective 32-char floor (the
  guarantee previously read as an unconditional max_chars bound).

* fix(harness): close env-assignment and expansion channels in acceptance matching (RFC #4651 PR4)

Self-audit in the shape of the last review rounds — channels the matcher
classified as accounted-for that can change what runs, narrow the
selection, or lend the summary text:

- env assignments are no longer blanket-stripped: only an allowlist of
  inert display/CI knobs (CI, NO_COLOR, PY_COLORS, ...) may prefix a
  matched span, and a non-allowlisted assignment in any preceding segment
  (pure-assignment or export NAME=) is state pollution — PATH redirects
  the executable, LD_PRELOAD/PYTHONPATH/NODE_OPTIONS inject code,
  PYTEST_ADDOPTS/GOFLAGS/MAKEFILES inject selection-changing inputs,
  BASH_ENV runs arbitrary shell startup. All degrade to unprovable.
- runtime expansions: any span token carrying /$( )/backticks, any
  negating-option value carrying an expansion or glob (unknown excluded
  set), and any extra executed token carrying glob metacharacters
  (crafted option-looking filenames narrow invisibly) are unprovable.
  Criterion-side globs stay self-consistent (literal match).
- cd: an argument carrying a runtime expansion or glob is non-silent
  (unknown destination, unknown print); CDPATH= assignments are now
  handled as state pollution at the match layer, subsuming the
  value-shape special case.

* fix(harness): persistent-shell evidence, exact env sets, option-arity scoping (RFC #4651 PR4)

- tests_passed: on a persistent-shell provider (new
  Sandbox.persistent_shell_sessions capability, set by AioSandbox) every
  leaf degrades to UNVERIFIED — any earlier call in the shared session
  could have mutated the state the clean-looking run executed in, and
  only a fresh controlled session (RFC section 6 verifier) can prove
  otherwise. The flag is read from the provider registry without
  acquiring a sandbox.
- env assignments: the allowlist is gone — no variable is provably inert
  across repositories (CI/DEBUG are routinely read by tests). The span's
  assignment prefix must equal the criterion's exactly (values included,
  order-insensitive); any assignment or export NAME= in a preceding
  segment is state pollution.
- scoping: positional targets are now read by option arity, so a path
  embedded in an option (--basetemp=/tmp/p, --junitxml=/tmp/r.xml) never
  counts as a selection target and an extra positional after such a
  criterion narrows the default selection it denotes.

* fix(harness): stamp shell provenance at harvest, close export/unset and arity gaps (RFC #4651 PR4)

* fix(harness): split physical newlines as shell separators in acceptance matching (RFC #4651 PR4)

* fix(harness): scope cd wrappers to thread data roots, pin accepted boundaries (RFC #4651 PR4)

* fix(harness): preserve criterion connectors, prove file_written readable, fail-closed shell capability (RFC #4651 PR4)

* fix(harness): compare only the connector prefix, tolerate trailing criterion semicolons (RFC #4651 PR4)

* fix(harness): preserve continuation-line operators, keep ./-spelled executable identity (RFC #4651 PR4)

* fix(harness): render criteria single-line so a multiline criterion cannot inject a forged checklist line (RFC #4651 PR4)

* fix(harness): reject parent-traversal executable tokens in acceptance matching (RFC #4651 PR4)

* fix(harness): reject parent-traversal negated values in acceptance matching (RFC #4651 PR4)
2026-09-01 16:13:41 +08:00

206 lines
8.1 KiB
Python

"""Regression tests for blocking-command timeout handling in LocalSandbox.
These pin the fix for the "starting a server hangs the whole turn" bug:
a backgrounded long-lived process must not keep the bash tool blocked until
the timeout, and a genuinely blocking foreground command must be terminated
(process group and all) once it exceeds the timeout.
The platform-specific cases exercise real subprocess and process-tree/group
semantics, while the shared tests pin the user-facing timeout notice.
"""
import os
import shlex
import sys
import time
from pathlib import Path
import pytest
from deerflow.config.sandbox_config import SandboxConfig
from deerflow.sandbox.local import local_sandbox
from deerflow.sandbox.local.local_sandbox import LocalSandbox
posix_only = pytest.mark.skipif(os.name == "nt", reason="POSIX process-group semantics")
linux_proc_fd_only = pytest.mark.skipif(not Path("/proc/self/fd").exists(), reason="requires Linux /proc fd links")
@posix_only
def test_backgrounded_process_returns_promptly():
"""A backgrounded long-lived process (e.g. a dev server started with `&`)
must return as soon as the foreground command finishes, instead of
blocking the bash tool until the timeout because it inherited the
captured pipe."""
sandbox = LocalSandbox("t")
start = time.monotonic()
output = sandbox.execute_command("sleep 5 & echo serving", timeout=10)
elapsed = time.monotonic() - start
assert elapsed < 3, f"expected prompt return, took {elapsed:.1f}s"
assert "serving" in output
@posix_only
@linux_proc_fd_only
def test_backgrounded_process_does_not_inherit_deleted_temp_capture(tmp_path):
"""A backgrounded process that forgets to redirect output must not inherit
an anonymous deleted temp file for fd 1. That would be an invisible,
unbounded disk leak for long-lived processes that keep writing."""
marker = tmp_path / "fd1"
script = f"import os, pathlib, time; pathlib.Path({str(marker)!r}).write_text(os.readlink('/proc/self/fd/1')); time.sleep(2)"
sandbox = LocalSandbox("t")
output = sandbox.execute_command(f"{shlex.quote(sys.executable)} -c {shlex.quote(script)} & echo launched", timeout=10)
assert "launched" in output
for _ in range(50):
if marker.exists():
break
time.sleep(0.1)
assert marker.exists()
assert " (deleted)" not in marker.read_text()
@posix_only
def test_foreground_blocking_command_times_out_with_notice():
"""A foreground command that never exits is terminated at the timeout and
the agent receives an explanatory notice instead of a generic error."""
sandbox = LocalSandbox("t")
start = time.monotonic()
output = sandbox.execute_command("while true; do sleep 0.2; done", timeout=1)
elapsed = time.monotonic() - start
assert elapsed < 5, f"timeout not enforced, took {elapsed:.1f}s"
assert "timed out" in output.lower()
def test_timeout_notice_formats_fractional_and_singular_timeouts(monkeypatch):
monkeypatch.setattr(LocalSandbox, "_get_shell", lambda self: "/bin/sh")
runner = "_run_windows_command" if os.name == "nt" else "_run_posix_command"
monkeypatch.setattr(LocalSandbox, runner, staticmethod(lambda args, timeout, env=None: ("", "", 0, True)))
assert "after 1.5 seconds" in LocalSandbox("t").execute_command("wait", timeout=1.5)
assert "after 1 second" in LocalSandbox("t").execute_command("wait", timeout=1)
def test_timeout_output_carries_authoritative_failure_marker(monkeypatch):
"""A timed-out command is a failed execution: the output must carry an
exit marker so exit-status evidence (acceptance checklist) cannot read a
partial passing summary as success."""
monkeypatch.setattr(LocalSandbox, "_get_shell", lambda self: "/bin/sh")
monkeypatch.setattr(LocalSandbox, "_run_posix_command", staticmethod(lambda args, timeout, env=None: ("12 passed\n", "", 0, True)))
output = LocalSandbox("t").execute_command("make test", timeout=1)
assert "timed out" in output.lower()
assert output.endswith("Exit Code: 124")
def test_windows_timeout_returns_notice(monkeypatch):
monkeypatch.setattr(local_sandbox.os, "name", "nt")
monkeypatch.setattr(LocalSandbox, "_get_shell", lambda self: "cmd.exe")
monkeypatch.setattr(
LocalSandbox,
"_run_windows_command",
staticmethod(lambda args, timeout, env: ("partial out", "partial err", 0, True)),
)
output = LocalSandbox("t").execute_command("wait", timeout=1.5)
assert "partial out" in output
assert "Std Error:" in output
assert "partial err" in output
assert "after 1.5 seconds" in output
assert "Unexpected error" not in output
@pytest.mark.skipif(os.name != "nt", reason="Windows process-tree semantics")
def test_windows_foreground_timeout_is_wall_clock_bound(monkeypatch):
monkeypatch.setattr(LocalSandbox, "_get_shell", lambda self: r"C:\Windows\System32\cmd.exe")
sandbox = LocalSandbox("t")
start = time.monotonic()
output = sandbox.execute_command("ping -n 5 127.0.0.1", timeout=0.2)
elapsed = time.monotonic() - start
assert elapsed < 2, f"timeout not enforced for process tree, took {elapsed:.1f}s"
assert "after 0.2 seconds" in output
@pytest.mark.skipif(os.name != "nt", reason="Windows bounded-capture semantics")
def test_windows_command_output_is_bounded(monkeypatch):
monkeypatch.setattr(
LocalSandbox,
"_get_shell",
lambda self: r"C:\Windows\System32\WindowsPowerShell\v1.0\powershell.exe",
)
emitted_chars = local_sandbox._COMMAND_CAPTURE_LIMIT_BYTES + 1024
output = LocalSandbox("t").execute_command(f"[Console]::Out.Write('x' * {emitted_chars})", timeout=20)
assert len(output) < emitted_chars
assert "output truncated after 10485760 of 10486784 bytes" in output
@posix_only
def test_foreground_timeout_kills_whole_process_group(tmp_path):
"""On timeout the entire process group is killed, not just the direct
child, so child processes spawned by the command do not survive."""
marker = tmp_path / "alive"
sandbox = LocalSandbox("t")
sandbox.execute_command(f"while true; do touch {marker}; sleep 0.2; done", timeout=1)
assert marker.exists()
first_mtime = marker.stat().st_mtime
time.sleep(1.5)
assert marker.stat().st_mtime == first_mtime, "process group survived the timeout"
@posix_only
def test_command_reading_stdin_does_not_block():
"""stdin is redirected from /dev/null, so a command that reads stdin gets
immediate EOF instead of blocking until the timeout."""
sandbox = LocalSandbox("t")
start = time.monotonic()
output = sandbox.execute_command("read x; echo got", timeout=10)
elapsed = time.monotonic() - start
assert elapsed < 3, f"stdin read blocked, took {elapsed:.1f}s"
assert "got" in output
@posix_only
def test_normal_command_output_exit_code_and_stderr():
"""Ordinary commands keep their existing output contract: stdout,
appended Std Error section, and a non-zero Exit Code line."""
sandbox = LocalSandbox("t")
assert "hello" in sandbox.execute_command("echo hello")
assert "Exit Code: 3" in sandbox.execute_command("exit 3")
combined = sandbox.execute_command("echo out; echo oops >&2")
assert "out" in combined
assert "Std Error:" in combined
assert "oops" in combined
def test_sandbox_config_exposes_command_timeout_default():
cfg = SandboxConfig(use="deerflow.sandbox.local:LocalSandboxProvider")
assert cfg.bash_command_timeout == 600
def test_sandbox_config_exposes_health_check_skip_seconds_default():
cfg = SandboxConfig(use="deerflow.sandbox.local:LocalSandboxProvider")
assert cfg.health_check_skip_seconds is None
def test_bash_tool_description_guides_backgrounding_long_lived_processes():
"""The bash tool description (seen by the model) must tell it to background
long-lived processes like servers, so it doesn't block the turn in the
foreground. This is the prompt-side half of the server-hang fix."""
from deerflow.sandbox.tools import bash_tool
description = bash_tool.description.lower()
assert "background" in description
assert "server" in description