mirror of
https://github.com/bytedance/deer-flow.git
synced 2026-09-14 16:08:41 +00:00
* feat(harness): deterministic acceptance checklist for subagent delegations (RFC #4651, layer 2) PR4 of RFC #4651: check lead-supplied acceptance_criteria in code when a subagent completes, so objectively checkable requirements can never be silently passed by a self-report. - subagents/acceptance_checks.py: deterministic leaf families — file:<path> exists|non-empty and file_written:<path> read through read_current_file_content scoped to the shared thread workspace; the read uses the sandbox-native virtual path form (the local read validator and provider mount tables resolve /mnt/user-data/... paths, not host paths); the scope decision canonicalizes with realpath on the local sandbox so workspace symlinks cannot escape into uploads; a remote provider's "Error: ..." return string is normalized to a failed check (provider-typed via is_local_sandbox); a UnicodeDecodeError marks a binary deliverable as existing and non-empty; out-of-scope paths degrade to UNVERIFIED. tests_passed:<command> anchors to a matching recorded bash execution with status=success and a test-summary shape; matching is shell-structure aware with control-flow attribution (span must end at the last segment with provable execution), negating-option values are ineligible evidence and a target negated anywhere in the command degrades the match, extra flags must be selection-preserving, extra positionals widen only after a path-scoped criterion, truncated commands degrade via command_truncated, the summary shape is read only from output attributable to the matched segment (preceding segments provably silent by invocation form), and pass shapes require a nonzero passed count. Criterion text is neutralized with neutralize_untrusted_tags before storage/rendering. Anything else renders UNVERIFIED, never silently passed. - executor: accumulate bounded bash command/output evidence per streamed chunk (merged by tool_call_id, newest-capped) so subagent summarization compacting earlier messages cannot erase a recorded execution; the recorded status is the actual shell exit status parsed from the output's exit marker (signed codes included; the remote Command exited with code N form is accepted only as the whole trimmed output), falling back to deerflow_tool_meta only when no marker exists. - sandbox providers: e2b/opensandbox/tenki/boxlite append the LocalSandbox-style "Exit Code: N" marker on nonzero exit even with non-empty output; aio propagates the SDK's structured exit_code on both exec paths the same way; local timeouts append Exit Code: 124; and _truncate_bash_output always preserves a trailing exit marker (signed included) inside its budget, with a 32-char floor raising any smaller configured limit, so the actual shell outcome always survives in the output text. - task_tool: run the checklist offloaded (asyncio.to_thread) on the completed branch, failure-isolated; stamp the verdict into result metadata and render the per-criterion section into the model-visible result text. - status contract: additive subagent_acceptance_verdict transport with read-side structural validation. - delegation ledger: entry carries the verdict and renders a compact acceptance segment; gateway strips caller-forged verdicts from both ledger entries and message metadata, like the citation verdict. - blocking-IO anchor pins the offload (teeth proven red->green); leaf read errors catch only OSError/SandboxError so unexpected errors reach the task-tool-level isolation instead of being mislabeled. * fix(harness): close acceptance evidence gaps from review (RFC #4651 PR4) - negating options: overlap with a matched criterion target is now checked by path/nodeid prefix, not exact token equality — excluding a sub-path of the criterion's selection (pytest tests --deselect tests/unit/test_auth.py) degrades to UNVERIFIED instead of holds - output attribution: any redirection token in the matched final segment makes the recorded tail non-attributable (> / >> / 2> are word characters to the parser, so redirection was invisible to the matcher) - silent-source allowlist narrowed from any *activate suffix to the */bin/activate shape - status_contract docstring: restore the shared-fixture sentence and note subagent_acceptance_verdict is deliberately outside the fixture - executor: update_bash_executions publishes [] (stream carried no bash-family calls) instead of collapsing it into None, mirroring update_tool_receipts * fix(harness): close acceptance residual gaps from re-review (RFC #4651 PR4) - tests_passed: add error outcomes to the fail shapes — "4 passed, 1 error" and pytest's "ERROR <nodeid>" short summary no longer satisfy the pass shape when the exit status is swallowed (|| true) or absent; zero-error counts stay clean. - file leaves: bound the deliverable read — a "wc -c" shell size probe answers files above 50k bytes without loading ~2x their size, honoring the host-bash kill switch and falling back to the full read on any non-integer rendering, so verdicts never get less sound. - executor: record the exit marker text as status_marker on harvested bash evidence; the leaf detail now reports the marker actually seen instead of asserting a failure indistinguishable from the command's own trailing text. - extend the blocking-IO anchor to drive the probe branch inside the offload; teeth re-verified red->green. * fix(harness): close acceptance forgery and bound gaps from P2 re-review (RFC #4651 PR4) - file leaves: never read unbounded — size is established first (os.stat on the validated local host path, so the host-bash-disabled configuration needs no shell; a guarded wc -c on remote providers that renders missing/unreadable in its own words). Above the 50k cap the leaf answers from the size alone, at/below it the full read runs, and an unestablishable size degrades to UNVERIFIED instead of an unlimited fallback read. - output attribution: source/. prefixes are never provably silent — a crafted */bin/activate path shape says nothing about what the script prints, so sourced segments can no longer lend a passing summary. - executable identity: an explicitly path-spelled criterion now requires the same normalized executable path; the basename rule stays only for deliberately bare criterion commands. * fix(harness): run acceptance size probe outside subagent-controlled state (RFC #4651 PR4) - remote probe no longer runs in the sandbox's persistent shell: a fresh env -i /bin/sh with absolute-path stat/realpath (poisoned functions, aliases, PATH, exported functions, IFS, locale cannot steer it), plus a marker env routing AIO onto a fresh per-call bash.exec session. - metadata-only: stat never opens content, so a FIFO deliverable cannot block the parent for the provider's idle timeout; non-regular files (fifo/dir/symlink) degrade to UNVERIFIED. - containment canonicalized against the literal mount root: a final-component symlink or a swapped parent directory (root included) cannot redirect the check outside shared storage; unprovable layouts degrade to UNVERIFIED. * fix(harness): canonicalize probe containment against the canonical mount root (RFC #4651 PR4) Literal-root equality made every remote file leaf permanently UNVERIFIED on e2b and Tenki, which realize /mnt/user-data as a symlink to the home dir by default (e2b bootstrap 'sudo ln -sfn', Tenki best-effort symlink). Containment now compares the file's realpath against the mount root's realpath — exactly what the provider's own read path resolves, so probe and read-back stay consistent; final-component symlinks stay rejected by the non-dereferencing stat, and an intermediate dir-link escape under a sane root still lands ESCAPED. The inner script is a module constant and the suite now executes the composed probe for real against on-disk layouts (real dir, symlinked prefix, final symlink, fifo, missing, dir-link escape), which the canned-output stub could not see. * fix(harness): close bare-criterion negation and CDPATH summary channels (RFC #4651 PR4) - matching: a criterion with no positional selection target (bare pytest, make test) stands for the runner's default selection, so ANY negating option (--ignore/--deselect/...) makes the recorded run a different selection — unprovable. The overlap guard only sees consumed criterion tokens, which a bare criterion does not have; scoped criteria keep the unrelated-exclusion behavior. - attribution: cd is no longer blanket-silent — CDPATH makes cd print the resolved (subagent-chosen) destination and the pass shapes match as substrings, so one mkdir 'all tests passed' plus an export minted a pass for any quiet command. A cd argument or CDPATH= value (export or leading assignment) carrying any summary shape makes the segment non-silent; shape-free cd dir wrappers keep matching. - docs: _truncate_bash_output states the effective 32-char floor (the guarantee previously read as an unconditional max_chars bound). * fix(harness): close env-assignment and expansion channels in acceptance matching (RFC #4651 PR4) Self-audit in the shape of the last review rounds — channels the matcher classified as accounted-for that can change what runs, narrow the selection, or lend the summary text: - env assignments are no longer blanket-stripped: only an allowlist of inert display/CI knobs (CI, NO_COLOR, PY_COLORS, ...) may prefix a matched span, and a non-allowlisted assignment in any preceding segment (pure-assignment or export NAME=) is state pollution — PATH redirects the executable, LD_PRELOAD/PYTHONPATH/NODE_OPTIONS inject code, PYTEST_ADDOPTS/GOFLAGS/MAKEFILES inject selection-changing inputs, BASH_ENV runs arbitrary shell startup. All degrade to unprovable. - runtime expansions: any span token carrying /$( )/backticks, any negating-option value carrying an expansion or glob (unknown excluded set), and any extra executed token carrying glob metacharacters (crafted option-looking filenames narrow invisibly) are unprovable. Criterion-side globs stay self-consistent (literal match). - cd: an argument carrying a runtime expansion or glob is non-silent (unknown destination, unknown print); CDPATH= assignments are now handled as state pollution at the match layer, subsuming the value-shape special case. * fix(harness): persistent-shell evidence, exact env sets, option-arity scoping (RFC #4651 PR4) - tests_passed: on a persistent-shell provider (new Sandbox.persistent_shell_sessions capability, set by AioSandbox) every leaf degrades to UNVERIFIED — any earlier call in the shared session could have mutated the state the clean-looking run executed in, and only a fresh controlled session (RFC section 6 verifier) can prove otherwise. The flag is read from the provider registry without acquiring a sandbox. - env assignments: the allowlist is gone — no variable is provably inert across repositories (CI/DEBUG are routinely read by tests). The span's assignment prefix must equal the criterion's exactly (values included, order-insensitive); any assignment or export NAME= in a preceding segment is state pollution. - scoping: positional targets are now read by option arity, so a path embedded in an option (--basetemp=/tmp/p, --junitxml=/tmp/r.xml) never counts as a selection target and an extra positional after such a criterion narrows the default selection it denotes. * fix(harness): stamp shell provenance at harvest, close export/unset and arity gaps (RFC #4651 PR4) * fix(harness): split physical newlines as shell separators in acceptance matching (RFC #4651 PR4) * fix(harness): scope cd wrappers to thread data roots, pin accepted boundaries (RFC #4651 PR4) * fix(harness): preserve criterion connectors, prove file_written readable, fail-closed shell capability (RFC #4651 PR4) * fix(harness): compare only the connector prefix, tolerate trailing criterion semicolons (RFC #4651 PR4) * fix(harness): preserve continuation-line operators, keep ./-spelled executable identity (RFC #4651 PR4) * fix(harness): render criteria single-line so a multiline criterion cannot inject a forged checklist line (RFC #4651 PR4) * fix(harness): reject parent-traversal executable tokens in acceptance matching (RFC #4651 PR4) * fix(harness): reject parent-traversal negated values in acceptance matching (RFC #4651 PR4)
206 lines
8.1 KiB
Python
206 lines
8.1 KiB
Python
"""Regression tests for blocking-command timeout handling in LocalSandbox.
|
|
|
|
These pin the fix for the "starting a server hangs the whole turn" bug:
|
|
a backgrounded long-lived process must not keep the bash tool blocked until
|
|
the timeout, and a genuinely blocking foreground command must be terminated
|
|
(process group and all) once it exceeds the timeout.
|
|
|
|
The platform-specific cases exercise real subprocess and process-tree/group
|
|
semantics, while the shared tests pin the user-facing timeout notice.
|
|
"""
|
|
|
|
import os
|
|
import shlex
|
|
import sys
|
|
import time
|
|
from pathlib import Path
|
|
|
|
import pytest
|
|
|
|
from deerflow.config.sandbox_config import SandboxConfig
|
|
from deerflow.sandbox.local import local_sandbox
|
|
from deerflow.sandbox.local.local_sandbox import LocalSandbox
|
|
|
|
posix_only = pytest.mark.skipif(os.name == "nt", reason="POSIX process-group semantics")
|
|
linux_proc_fd_only = pytest.mark.skipif(not Path("/proc/self/fd").exists(), reason="requires Linux /proc fd links")
|
|
|
|
|
|
@posix_only
|
|
def test_backgrounded_process_returns_promptly():
|
|
"""A backgrounded long-lived process (e.g. a dev server started with `&`)
|
|
must return as soon as the foreground command finishes, instead of
|
|
blocking the bash tool until the timeout because it inherited the
|
|
captured pipe."""
|
|
sandbox = LocalSandbox("t")
|
|
start = time.monotonic()
|
|
output = sandbox.execute_command("sleep 5 & echo serving", timeout=10)
|
|
elapsed = time.monotonic() - start
|
|
|
|
assert elapsed < 3, f"expected prompt return, took {elapsed:.1f}s"
|
|
assert "serving" in output
|
|
|
|
|
|
@posix_only
|
|
@linux_proc_fd_only
|
|
def test_backgrounded_process_does_not_inherit_deleted_temp_capture(tmp_path):
|
|
"""A backgrounded process that forgets to redirect output must not inherit
|
|
an anonymous deleted temp file for fd 1. That would be an invisible,
|
|
unbounded disk leak for long-lived processes that keep writing."""
|
|
marker = tmp_path / "fd1"
|
|
script = f"import os, pathlib, time; pathlib.Path({str(marker)!r}).write_text(os.readlink('/proc/self/fd/1')); time.sleep(2)"
|
|
sandbox = LocalSandbox("t")
|
|
|
|
output = sandbox.execute_command(f"{shlex.quote(sys.executable)} -c {shlex.quote(script)} & echo launched", timeout=10)
|
|
|
|
assert "launched" in output
|
|
for _ in range(50):
|
|
if marker.exists():
|
|
break
|
|
time.sleep(0.1)
|
|
assert marker.exists()
|
|
assert " (deleted)" not in marker.read_text()
|
|
|
|
|
|
@posix_only
|
|
def test_foreground_blocking_command_times_out_with_notice():
|
|
"""A foreground command that never exits is terminated at the timeout and
|
|
the agent receives an explanatory notice instead of a generic error."""
|
|
sandbox = LocalSandbox("t")
|
|
start = time.monotonic()
|
|
output = sandbox.execute_command("while true; do sleep 0.2; done", timeout=1)
|
|
elapsed = time.monotonic() - start
|
|
|
|
assert elapsed < 5, f"timeout not enforced, took {elapsed:.1f}s"
|
|
assert "timed out" in output.lower()
|
|
|
|
|
|
def test_timeout_notice_formats_fractional_and_singular_timeouts(monkeypatch):
|
|
monkeypatch.setattr(LocalSandbox, "_get_shell", lambda self: "/bin/sh")
|
|
runner = "_run_windows_command" if os.name == "nt" else "_run_posix_command"
|
|
monkeypatch.setattr(LocalSandbox, runner, staticmethod(lambda args, timeout, env=None: ("", "", 0, True)))
|
|
|
|
assert "after 1.5 seconds" in LocalSandbox("t").execute_command("wait", timeout=1.5)
|
|
assert "after 1 second" in LocalSandbox("t").execute_command("wait", timeout=1)
|
|
|
|
|
|
def test_timeout_output_carries_authoritative_failure_marker(monkeypatch):
|
|
"""A timed-out command is a failed execution: the output must carry an
|
|
exit marker so exit-status evidence (acceptance checklist) cannot read a
|
|
partial passing summary as success."""
|
|
monkeypatch.setattr(LocalSandbox, "_get_shell", lambda self: "/bin/sh")
|
|
monkeypatch.setattr(LocalSandbox, "_run_posix_command", staticmethod(lambda args, timeout, env=None: ("12 passed\n", "", 0, True)))
|
|
|
|
output = LocalSandbox("t").execute_command("make test", timeout=1)
|
|
|
|
assert "timed out" in output.lower()
|
|
assert output.endswith("Exit Code: 124")
|
|
|
|
|
|
def test_windows_timeout_returns_notice(monkeypatch):
|
|
monkeypatch.setattr(local_sandbox.os, "name", "nt")
|
|
monkeypatch.setattr(LocalSandbox, "_get_shell", lambda self: "cmd.exe")
|
|
monkeypatch.setattr(
|
|
LocalSandbox,
|
|
"_run_windows_command",
|
|
staticmethod(lambda args, timeout, env: ("partial out", "partial err", 0, True)),
|
|
)
|
|
|
|
output = LocalSandbox("t").execute_command("wait", timeout=1.5)
|
|
|
|
assert "partial out" in output
|
|
assert "Std Error:" in output
|
|
assert "partial err" in output
|
|
assert "after 1.5 seconds" in output
|
|
assert "Unexpected error" not in output
|
|
|
|
|
|
@pytest.mark.skipif(os.name != "nt", reason="Windows process-tree semantics")
|
|
def test_windows_foreground_timeout_is_wall_clock_bound(monkeypatch):
|
|
monkeypatch.setattr(LocalSandbox, "_get_shell", lambda self: r"C:\Windows\System32\cmd.exe")
|
|
sandbox = LocalSandbox("t")
|
|
|
|
start = time.monotonic()
|
|
output = sandbox.execute_command("ping -n 5 127.0.0.1", timeout=0.2)
|
|
elapsed = time.monotonic() - start
|
|
|
|
assert elapsed < 2, f"timeout not enforced for process tree, took {elapsed:.1f}s"
|
|
assert "after 0.2 seconds" in output
|
|
|
|
|
|
@pytest.mark.skipif(os.name != "nt", reason="Windows bounded-capture semantics")
|
|
def test_windows_command_output_is_bounded(monkeypatch):
|
|
monkeypatch.setattr(
|
|
LocalSandbox,
|
|
"_get_shell",
|
|
lambda self: r"C:\Windows\System32\WindowsPowerShell\v1.0\powershell.exe",
|
|
)
|
|
emitted_chars = local_sandbox._COMMAND_CAPTURE_LIMIT_BYTES + 1024
|
|
|
|
output = LocalSandbox("t").execute_command(f"[Console]::Out.Write('x' * {emitted_chars})", timeout=20)
|
|
|
|
assert len(output) < emitted_chars
|
|
assert "output truncated after 10485760 of 10486784 bytes" in output
|
|
|
|
|
|
@posix_only
|
|
def test_foreground_timeout_kills_whole_process_group(tmp_path):
|
|
"""On timeout the entire process group is killed, not just the direct
|
|
child, so child processes spawned by the command do not survive."""
|
|
marker = tmp_path / "alive"
|
|
sandbox = LocalSandbox("t")
|
|
sandbox.execute_command(f"while true; do touch {marker}; sleep 0.2; done", timeout=1)
|
|
|
|
assert marker.exists()
|
|
first_mtime = marker.stat().st_mtime
|
|
time.sleep(1.5)
|
|
assert marker.stat().st_mtime == first_mtime, "process group survived the timeout"
|
|
|
|
|
|
@posix_only
|
|
def test_command_reading_stdin_does_not_block():
|
|
"""stdin is redirected from /dev/null, so a command that reads stdin gets
|
|
immediate EOF instead of blocking until the timeout."""
|
|
sandbox = LocalSandbox("t")
|
|
start = time.monotonic()
|
|
output = sandbox.execute_command("read x; echo got", timeout=10)
|
|
elapsed = time.monotonic() - start
|
|
|
|
assert elapsed < 3, f"stdin read blocked, took {elapsed:.1f}s"
|
|
assert "got" in output
|
|
|
|
|
|
@posix_only
|
|
def test_normal_command_output_exit_code_and_stderr():
|
|
"""Ordinary commands keep their existing output contract: stdout,
|
|
appended Std Error section, and a non-zero Exit Code line."""
|
|
sandbox = LocalSandbox("t")
|
|
|
|
assert "hello" in sandbox.execute_command("echo hello")
|
|
assert "Exit Code: 3" in sandbox.execute_command("exit 3")
|
|
|
|
combined = sandbox.execute_command("echo out; echo oops >&2")
|
|
assert "out" in combined
|
|
assert "Std Error:" in combined
|
|
assert "oops" in combined
|
|
|
|
|
|
def test_sandbox_config_exposes_command_timeout_default():
|
|
cfg = SandboxConfig(use="deerflow.sandbox.local:LocalSandboxProvider")
|
|
assert cfg.bash_command_timeout == 600
|
|
|
|
|
|
def test_sandbox_config_exposes_health_check_skip_seconds_default():
|
|
cfg = SandboxConfig(use="deerflow.sandbox.local:LocalSandboxProvider")
|
|
assert cfg.health_check_skip_seconds is None
|
|
|
|
|
|
def test_bash_tool_description_guides_backgrounding_long_lived_processes():
|
|
"""The bash tool description (seen by the model) must tell it to background
|
|
long-lived processes like servers, so it doesn't block the turn in the
|
|
foreground. This is the prompt-side half of the server-hang fix."""
|
|
from deerflow.sandbox.tools import bash_tool
|
|
|
|
description = bash_tool.description.lower()
|
|
assert "background" in description
|
|
assert "server" in description
|