* fix(sandbox): cut read_file output at a line boundary and name the next start_line
read_file head-truncates at a character offset and its marker told the model
to continue with start_line/end_line while reporting only character counts,
so the cut usually fell mid-line and the model had to guess which line to
continue from. The cut now lands on the last line boundary the budget allows,
and the marker reports lines shown of lines total, keeps the character
counts, and names the exact next start_line. When the line at the cut is
longer than 4,096 characters (minified sources, one-line JSON) the cut stays
at the character limit and the marker names the line it fell inside, so a
re-read of that line is the continuation. Reads under the limit are unchanged.
* fix(sandbox): make the read_file continuation hold for ranged reads and long lines
Line numbers in the truncation marker are now file line numbers: read_file_tool
passes start_line - 1 as the line offset, so a ranged read that is itself
truncated names the right next line instead of one relative to its slice. A
ranged read is a provider slice joined with newlines, so the tool also says
so and a trailing newline there counts as an empty last line.
The long-line fallback now names a continuation only when it makes progress:
a read from the cut line when the whole line fits such a read, a single-line
read (start_line = end_line) when only the line alone fits max_chars, and bash
when even that cannot return it; the single-line form names no further line
after the last line of the read. The budget reserves one extra character so a
newline sitting exactly at the limit still counts as a complete line, the
"fits a fresh read" check uses a pessimistic estimate of the follow-up read's
marker, and a budget too small for any marker still returns a marker instead
of a bare prefix.
Adds unit cases for the ranged-read offset, the newline-at-budget edge, the
single-line-read and bash forms, tiny budgets and empty last lines, plus
end-to-end tests that drive read_file_tool with a LocalSandbox and follow the
markers across reads, asserting the kept segments reproduce the file without
gap or overlap.
* fix(sandbox): keep naming the next line after a bounded read's last line
A ranged read with an end_line below the file's length is a slice that stops
mid-file, so the single-line-read continuation must still name the line after
the slice's last line; only a read that reached the end of the file names
nothing further. The tool passes whether the read was bounded by an end_line
separately from the joined-lines hint, because a start_line-only read also
runs to the end of the file.
A blank line and a line past the end both read back as an empty slice; the
tool now tells them apart with a two-line probe, so a continuation named by a
marker that lands on a blank line answers "(empty)" rather than
"(start_line exceeds file length)".
* test(sandbox): adapt upstream continuation checks after rebase
---------
Co-authored-by: Totoro-qaq <279883115+Totoro-qaq@users.noreply.github.com>
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
* fix(sandbox): report the line a read_file truncation lands on
The read_file tool head-truncates at a character offset and tells the
model to continue with start_line/end_line, but the marker reported only
character counts ("showing first N of M chars"), so the model had no way
to know which line the cut fell on — the cut almost always lands mid-line
and read_file output carries no line numbers (#5475).
The marker now also reports the 1-indexed line holding the first hidden
character and the file's total line count, and names the exact resume
point: "... [truncated: showing first N of M chars (cut lands in line L
of T). Use start_line=L — optionally with end_line — to continue without
a gap] ...". Resuming at the reported line is gap-free whether the cut
lands mid-line or exactly after a newline.
The marker length budget accounts for the new fields, so the
len(result) <= max_chars contract still holds.
* fix(sandbox): report absolute lines in ranged-read truncation markers
Review follow-up on #5478: read_file_tool runs the same truncation on
ranged reads (start_line/end_line), where the slice's line 1 is the
requested start_line, not the file's first line. The marker's reported
lines were slice-relative while the model reasons in absolute file lines,
so the resume hint could re-issue the identical start_line forever
(repro: resume at 831 -> "cut lands in line 831 of 5170" -> start_line=831).
_truncate_read_file_output gains a line_offset parameter (the 0-based
absolute line of the slice's first line) and reports absolute lines for
both the cut position and the range end; read_file_tool threads
effective_start - 1 through. Full reads pass the default offset 0 and are
byte-identical.
Regression tests pin the absolute coordinates and that the resume point
strictly advances past the slice start.
* fix(sandbox): report an exactly-full AIO glob result as complete
AioSandbox.glob's include_dirs branch returned as soon as it had
collected max_results matches, without looking at the rest of the
listing. A listing that held exactly that many matches and nothing more
was therefore reported as truncated, and the glob tool told the model
the result was incomplete — prompting a re-search or distrust of a
complete answer. The same line returned one match for max_results=0,
one past the caller's cap.
Look one match past the cap before deciding, which is what the
include_dirs=False branch in the same function already does and what
#5427 moved parse_remote_search_output to for BoxLite, Tenki, E2B and
OpenSandbox.
* review: filtered-tail cases, the glob contract docstring, and the cap wording
Addresses the three items from the review on #5449.
- Two regression cases over a tail of ignored / out-of-root / pattern-miss
entries: an exactly-full result stays complete when only filtered entries
follow, and a third eligible match after that tail still reports
truncation. Both fail against the previous return-on-the-max-th-match
behaviour.
- 'Sandbox.glob' promised the conservative flag ('``max_results`` was
reached') that this change deliberately stops producing on the AIO branch.
The contract now reads as 'may be incomplete' and records that providers
differ in how precisely they can decide it.
- The changelog no longer lumps 'parse_remote_search_output' in with the
filtered-match cap: its raw-output cap is a separate limit with its own
one-line-past accounting, and the other providers' filtered-match cap is
unchanged.
Also corrects the docstring on the existing test, which still described the
removed early return in the present tense.
---------
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
* fix(sandbox): force UTF-8 console for PowerShell so CJK output is not garbled
LocalSandbox captures PowerShell output through a UTF-8 pipe reader
(errors=replace), but Windows PowerShell 5.1 writes console output in
the legacy OEM codepage (GBK on zh-CN Windows) unless told otherwise,
so every CJK character in tool output arrives as mojibake and the
decode never raises. Prepend a UTF-8 preamble
([Console]::InputEncoding/[Console]::OutputEncoding/$OutputEncoding)
to the -Command payload so both directions of the console are UTF-8
before the user command runs.
* fix(sandbox): pair PowerShell UTF-8 capture and guard console setup
---------
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
* fix(sandbox): report truncated remote glob and grep results
BoxLite, Tenki, E2B, and OpenSandbox run find/grep in the sandbox, cap
the raw output with `| head`, and then filter those lines in Python:
ignored directories such as node_modules are dropped and grep's glob
scope is applied. They reported truncated only when max_results matches
survived the filter. When the capped lines were mostly filtered out, a
search with real matches past the cap came back short or empty with
truncated=False, and glob_tool/grep_tool rendered it as "No files
matched" / "No matches found". With the default max_results=200 and
1,200 files under node_modules, glob("**/*.py") reported no matches for
a workspace that has src/app.py.
remote_search_command now lets one line past its limit through, and
parse_remote_search_output(..., limit=) returns RemoteSearchOutput(text,
truncated): the first `limit` lines and whether the extra line arrived.
Exactly `limit` lines stays a complete result. Each provider passes the
cap it already computed to both calls and returns that truncated from
glob and grep when fewer than max_results results survive filtering.
The glob and grep tools now describe an empty truncated result as
incomplete instead of reporting no matches, which also covers AIO grep's
forwarded truncated flag. Sandbox.glob/grep document truncated as "the
matches may be incomplete".
* docs(changelog): reference #5427 in the remote search truncation entry
---------
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
* fix(sandbox): mask every host path in a colon-joined list
Host-to-virtual output masking matched a host root and then consumed the
path tail up to whitespace or shell punctuation, but not `:`. A
`:`-joined list such as $PATH or $PYTHONPATH was therefore swallowed into
the first match's tail, and scanning resumed after it, so every later
entry under the same root reached the model as a raw host path. The regex
matcher (process-stable skill roots) and the direct scanner (per-thread
roots, LocalSandbox) shared the gap. Each redundant masking pass --
separator variants, the realpath spelling, the /mnt/user-data root
mapping, LocalSandbox's own reverse resolution -- happened to recover one
entry, which hid the leak for short lists: bash output leaked from the
fourth entry, single-pass consumers from the third.
The shared tail in path_patterns.py now ends at `:` in both matchers.
`;`, the Windows list separator, already ended it. A `:` inside one path
(grep -n output, a file name) only shortens the match; the remaining text
is copied through verbatim.
Shortening the match exposed a second leak. LocalSandbox reverse
resolution realpaths the matched path and returned that realpath when no
mount contained it, so a symlink inside a mount whose target lies outside
every mount was shown as the target's host path. grep -n lines used to
hide this only because the whole line resolved as one nonexistent file;
whitespace-terminated output and LocalSandbox.glob results already leaked
it on main. Reverse resolution now falls back to the link's own spelling,
normalized so `mount/../x` does not pass, before giving up. A symlink into
another mount still reports that mount's path.
* docs(changelog): reference #5418 in the colon-joined path masking entry
* docs(changelog): split the #5418 and #5419 entries fused by the merge
Resolving the CHANGELOG conflict when main was merged in dropped the
opener of the #5419 entry, so the BoxLite grep fix continued inside this
PR's bullet in both CHANGELOG.md and CHANGELOG_zh.md. Restore it as its
own bullet; the #5419 entry is byte-identical to main again.
---------
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
* fix(sandbox): scope BoxLite grep globs to the search root
BoxliteBox.grep omits grep --include for busybox portability and applies
the glob in Python, but it kept only the glob's last path segment and
matched it against each file's basename. A scoped pattern therefore lost
its directory part: grep(glob="src/*.js") returned every .js file in the
tree, including vendor/ and nested src/ subdirectories that glob() with
the same pattern excludes.
The glob now goes through path_matches against the path relative to the
search root, with the file's basename when the root is a single file --
the same scope glob() uses and the one Tenki, E2B, OpenSandbox, AIO and
LocalSandbox already enforce. Like those providers, an empty glob is now
passed to path_matches instead of being treated as no filter.
* docs(changelog): reference #5419 in the BoxLite grep glob scope entry
* fix(sandbox): reverse-resolve forward-slash spellings of Windows host paths
Forward resolution deliberately spells resolved paths with forward
slashes in commands and file content (#3869: backslashes break bash
escapes), but the reverse scanner anchored its matches on the native
backslash base, so on Windows every forward-resolved path that came back
in command output or agent-written files leaked the raw host path
instead of mapping to its container path. Match separator-agnostically
in LocalSandbox like sandbox.tools already does, align the two
regex-cache tests with the documented spellings, and refresh the
path_patterns rationale comments that described the old asymmetry.
* test(sandbox): pin the reverse mask to separator-agnostic matching
The flag is the entire Windows fix but is invisible on POSIX CI, so
assert the routing kwargs in the direct-helper wiring test — the same
pin test_tools_mask_patterns_route_through_the_helper already applies to
the sandbox.tools copy. A revert to separator-exact matching now fails
on every platform instead of silently reintroducing the host-path
leak on Windows.
---------
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
* fix(sandbox): stop remote grep/glob from reporting failures as no matches
E2B, OpenSandbox, BoxLite and Tenki ran grep/find behind `2>/dev/null | head`, so a missing search root, a missing grep/find binary or an unreadable tree exited 0 with empty stdout and the tools reported "No matches found". Wrap the search in sandbox/remote_search.py, which checks the root first and records the search's own status after head, as remote_list_dir does for list_dir: a missing root raises FileNotFoundError, a failed search raises OSError, and a genuine no-match still returns []. glob's find gains -H for symlinked roots, OpenSandbox's BusyBox fallback keeps the primary grep status, and E2B no longer swallows client errors. Regression tests run each provider's real command in a local POSIX sh.
Fixes#5376
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(sandbox): fail remote grep/glob on partial traversal errors
grep 2 / find 1 after some results were printed (an unreadable file or
subdirectory) were returned as a complete search. Callers have no
partial-result channel, and #5376 asks for permission and command
failures to raise, so these statuses now raise OSError like any other
failure. Only grep 0/1/141 and find 0/141 pass.
The error for grep 2 / find 1 says that some files or directories could
not be read and asks for a narrower path, so the agent can retry instead
of giving up.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Totoro-qaq <279883115+Totoro-qaq@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix(sandbox): stop list_dir from reporting failures as empty
Remote providers swallowed find/client errors as [] and 2>/dev/null
missing paths as empty stdout. ls_tool then told the agent the
directory was (empty). Raise OSError/FileNotFoundError instead so
the tool returns Error.
* fix(sandbox): list_dir raises on missing local paths and uses find -H
Empty stdout is not a missing path when find's start point is a
symlink (E2B /mnt/acp-workspace). Dereference only the start point
with find -H. LocalSandbox now raises FileNotFoundError for a
non-directory root, matching remote providers. AIO maps a missing
result.data to OSError rather than FileNotFoundError.
* fix(sandbox): group AIO list_dir find type predicates
Without parentheses, find PATH -maxdepth N -type f -o -type d applies
-type d without maxdepth and can drop files from the listing.
* fix(sandbox): distinguish list_dir command failure from missing path
Tenki, Boxlite, and OpenSandbox treated any empty find stdout as
FileNotFoundError, so a missing find binary (exit 127) or SDK error
looked like a missing directory. Raise OSError when find status is
outside (0, 1); keep FileNotFoundError for the find-ran-but-empty case.
* fix(sandbox): apply list_dir exit-status contract to AIO and E2B
Same gap as Tenki/Boxlite/OpenSandbox: empty find stdout with exit 127
was FileNotFoundError. Raise OSError when the status is outside (0, 1).
* fix(sandbox): classify list_dir by find status not head status
find | head under sh -lc reports head's exit code, so a missing find
binary (127) became FileNotFoundError. Record find's own status after
the bounded listing, treat SIGPIPE 141 as truncation success, and add
a shell-level regression test.
* test(auth): include projects permissions in /me contract pins
#5265 added projects:read/write/delete to the registered route set.
The /auth/me tests still pinned the pre-projects list, so CI failed
after merging main.
* fix(sandbox): do not treat missing list_dir marker as success
The generated script ended on `rm -f`, so process status was 0/1 even
when find's marker never landed. Both codes are in _FIND_OK, and the
parser fallback then classified an empty listing as FileNotFoundError —
the 127 misclassification this helper was meant to close.
Exit with find's status (126 if unknown). A missing marker is now
OSError unless the process status is already a non-OK failure.
* test(sandbox): emit list_dir status marker in provider fixtures
Parser now requires __DF_FIND_STATUS__ and refuses marker-less stdout.
Update AIO/Boxlite/E2B stubs and OpenSandbox/Tenki find fakes so listings
carry :0 and missing paths carry :1 with matching exit codes.
* style(sandbox): format list dir test fixture
* style(sandbox): format remote list dir helper
* docs(sandbox): keep guidance within the tested size budget
---------
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
* fix(sandbox): mask output tails into POSIX-style virtual paths
The output maskers slice the matched path tail from the original
output. With separator-agnostic matching, a Windows-spelled nested
tail kept its backslashes and was spliced into the POSIX-style virtual
path, so glob results and masked read output showed mixed paths like
/mnt/user-data/workspace/pkg\util.py or
/mnt/skills/integrations/lark-cli\lark-doc\SKILL.md. Virtual paths are
always POSIX-style, so normalize nested tails to forward slashes the
same way depth-1 tails already end up. Depth-1 tails and the callable
replacer (LocalSandbox._reverse_resolve_path) were unaffected.
Pin the nested-tail contract in test_sandbox_path_patterns; the
previously failing glob-tool and skills-masking regressions now pass
on Windows hosts.
* refactor(sandbox): share the mask tail-splicing rule; guard it on Linux CI
Review follow-up for #5247:
- hoist the tail-splicing rule (slice off the base, strip leading
separators, normalize the rest to "/") into
path_patterns.normalize_mask_tail and import it at both call sites,
so the two maskers can only drift in their matching logic, not in
the splice;
- add test_mask_local_paths_normalizes_windows_spelled_skill_tails,
which spells the skills host root and the output with Windows-style
strings so the nested tail keeps backslashes on every platform.
Reverting the mask_local_paths_in_output-side normalization now goes
red on Linux CI too, not only on Windows hosts.
SSH_AUTH_SOCK points at the host's ssh-agent socket. A sandbox
subprocess that inherits it can sign and authenticate with every key
the agent holds (git push, ssh logins) without reading any key file --
the same credential-pointer leak class as the *_ASKPASS helpers the
env policy already scrubs deliberately. No wildcard pattern fits
(*AUTH* would strip benign names), so add an exact entry to
_BLOCKED_EXACT_NAMES.
A skill that genuinely needs the agent socket can still declare it via
required-secrets: injected values win over the blocklist by design.
Co-authored-by: zhouyujie <zhouyujie@keep.com>
* feat(harness): deterministic acceptance checklist for subagent delegations (RFC #4651, layer 2)
PR4 of RFC #4651: check lead-supplied acceptance_criteria in code when a
subagent completes, so objectively checkable requirements can never be
silently passed by a self-report.
- subagents/acceptance_checks.py: deterministic leaf families —
file:<path> exists|non-empty and file_written:<path> read through
read_current_file_content scoped to the shared thread workspace; the
read uses the sandbox-native virtual path form (the local read
validator and provider mount tables resolve /mnt/user-data/... paths,
not host paths); the scope decision canonicalizes with realpath on the
local sandbox so workspace symlinks cannot escape into uploads; a
remote provider's "Error: ..." return string is normalized to a
failed check (provider-typed via is_local_sandbox); a
UnicodeDecodeError marks a binary deliverable as existing and
non-empty; out-of-scope paths degrade to UNVERIFIED.
tests_passed:<command> anchors to a matching recorded bash execution
with status=success and a test-summary shape; matching is
shell-structure aware with control-flow attribution (span must end at
the last segment with provable execution), negating-option values are
ineligible evidence and a target negated anywhere in the command
degrades the match, extra flags must be selection-preserving, extra
positionals widen only after a path-scoped criterion, truncated
commands degrade via command_truncated, the summary shape is read
only from output attributable to the matched segment (preceding
segments provably silent by invocation form), and pass shapes require
a nonzero passed count. Criterion text is neutralized with
neutralize_untrusted_tags before storage/rendering. Anything else
renders UNVERIFIED, never silently passed.
- executor: accumulate bounded bash command/output evidence per streamed
chunk (merged by tool_call_id, newest-capped) so subagent
summarization compacting earlier messages cannot erase a recorded
execution; the recorded status is the actual shell exit status parsed
from the output's exit marker (signed codes included; the remote
Command exited with code N form is accepted only as the whole trimmed
output), falling back to deerflow_tool_meta only when no marker
exists.
- sandbox providers: e2b/opensandbox/tenki/boxlite append the
LocalSandbox-style "Exit Code: N" marker on nonzero exit even with
non-empty output; aio propagates the SDK's structured exit_code on
both exec paths the same way; local timeouts append Exit Code: 124;
and _truncate_bash_output always preserves a trailing exit marker
(signed included) inside its budget, with a 32-char floor raising any
smaller configured limit, so the actual shell outcome always survives
in the output text.
- task_tool: run the checklist offloaded (asyncio.to_thread) on the
completed branch, failure-isolated; stamp the verdict into result
metadata and render the per-criterion section into the model-visible
result text.
- status contract: additive subagent_acceptance_verdict transport with
read-side structural validation.
- delegation ledger: entry carries the verdict and renders a compact
acceptance segment; gateway strips caller-forged verdicts from both
ledger entries and message metadata, like the citation verdict.
- blocking-IO anchor pins the offload (teeth proven red->green); leaf
read errors catch only OSError/SandboxError so unexpected errors reach
the task-tool-level isolation instead of being mislabeled.
* fix(harness): close acceptance evidence gaps from review (RFC #4651 PR4)
- negating options: overlap with a matched criterion target is now
checked by path/nodeid prefix, not exact token equality — excluding a
sub-path of the criterion's selection (pytest tests --deselect
tests/unit/test_auth.py) degrades to UNVERIFIED instead of holds
- output attribution: any redirection token in the matched final segment
makes the recorded tail non-attributable (> / >> / 2> are word
characters to the parser, so redirection was invisible to the matcher)
- silent-source allowlist narrowed from any *activate suffix to the
*/bin/activate shape
- status_contract docstring: restore the shared-fixture sentence and
note subagent_acceptance_verdict is deliberately outside the fixture
- executor: update_bash_executions publishes [] (stream carried no
bash-family calls) instead of collapsing it into None, mirroring
update_tool_receipts
* fix(harness): close acceptance residual gaps from re-review (RFC #4651 PR4)
- tests_passed: add error outcomes to the fail shapes — "4 passed, 1 error"
and pytest's "ERROR <nodeid>" short summary no longer satisfy the pass
shape when the exit status is swallowed (|| true) or absent; zero-error
counts stay clean.
- file leaves: bound the deliverable read — a "wc -c" shell size probe
answers files above 50k bytes without loading ~2x their size, honoring
the host-bash kill switch and falling back to the full read on any
non-integer rendering, so verdicts never get less sound.
- executor: record the exit marker text as status_marker on harvested bash
evidence; the leaf detail now reports the marker actually seen instead of
asserting a failure indistinguishable from the command's own trailing text.
- extend the blocking-IO anchor to drive the probe branch inside the
offload; teeth re-verified red->green.
* fix(harness): close acceptance forgery and bound gaps from P2 re-review (RFC #4651 PR4)
- file leaves: never read unbounded — size is established first (os.stat on
the validated local host path, so the host-bash-disabled configuration
needs no shell; a guarded wc -c on remote providers that renders
missing/unreadable in its own words). Above the 50k cap the leaf answers
from the size alone, at/below it the full read runs, and an
unestablishable size degrades to UNVERIFIED instead of an unlimited
fallback read.
- output attribution: source/. prefixes are never provably silent — a
crafted */bin/activate path shape says nothing about what the script
prints, so sourced segments can no longer lend a passing summary.
- executable identity: an explicitly path-spelled criterion now requires
the same normalized executable path; the basename rule stays only for
deliberately bare criterion commands.
* fix(harness): run acceptance size probe outside subagent-controlled state (RFC #4651 PR4)
- remote probe no longer runs in the sandbox's persistent shell: a fresh
env -i /bin/sh with absolute-path stat/realpath (poisoned functions,
aliases, PATH, exported functions, IFS, locale cannot steer it), plus a
marker env routing AIO onto a fresh per-call bash.exec session.
- metadata-only: stat never opens content, so a FIFO deliverable cannot
block the parent for the provider's idle timeout; non-regular files
(fifo/dir/symlink) degrade to UNVERIFIED.
- containment canonicalized against the literal mount root: a
final-component symlink or a swapped parent directory (root included)
cannot redirect the check outside shared storage; unprovable layouts
degrade to UNVERIFIED.
* fix(harness): canonicalize probe containment against the canonical mount root (RFC #4651 PR4)
Literal-root equality made every remote file leaf permanently UNVERIFIED
on e2b and Tenki, which realize /mnt/user-data as a symlink to the home
dir by default (e2b bootstrap 'sudo ln -sfn', Tenki best-effort symlink).
Containment now compares the file's realpath against the mount root's
realpath — exactly what the provider's own read path resolves, so probe
and read-back stay consistent; final-component symlinks stay rejected by
the non-dereferencing stat, and an intermediate dir-link escape under a
sane root still lands ESCAPED. The inner script is a module constant and
the suite now executes the composed probe for real against on-disk
layouts (real dir, symlinked prefix, final symlink, fifo, missing,
dir-link escape), which the canned-output stub could not see.
* fix(harness): close bare-criterion negation and CDPATH summary channels (RFC #4651 PR4)
- matching: a criterion with no positional selection target (bare pytest,
make test) stands for the runner's default selection, so ANY negating
option (--ignore/--deselect/...) makes the recorded run a different
selection — unprovable. The overlap guard only sees consumed criterion
tokens, which a bare criterion does not have; scoped criteria keep the
unrelated-exclusion behavior.
- attribution: cd is no longer blanket-silent — CDPATH makes cd print the
resolved (subagent-chosen) destination and the pass shapes match as
substrings, so one mkdir 'all tests passed' plus an export minted a pass
for any quiet command. A cd argument or CDPATH= value (export or leading
assignment) carrying any summary shape makes the segment non-silent;
shape-free cd dir wrappers keep matching.
- docs: _truncate_bash_output states the effective 32-char floor (the
guarantee previously read as an unconditional max_chars bound).
* fix(harness): close env-assignment and expansion channels in acceptance matching (RFC #4651 PR4)
Self-audit in the shape of the last review rounds — channels the matcher
classified as accounted-for that can change what runs, narrow the
selection, or lend the summary text:
- env assignments are no longer blanket-stripped: only an allowlist of
inert display/CI knobs (CI, NO_COLOR, PY_COLORS, ...) may prefix a
matched span, and a non-allowlisted assignment in any preceding segment
(pure-assignment or export NAME=) is state pollution — PATH redirects
the executable, LD_PRELOAD/PYTHONPATH/NODE_OPTIONS inject code,
PYTEST_ADDOPTS/GOFLAGS/MAKEFILES inject selection-changing inputs,
BASH_ENV runs arbitrary shell startup. All degrade to unprovable.
- runtime expansions: any span token carrying /$( )/backticks, any
negating-option value carrying an expansion or glob (unknown excluded
set), and any extra executed token carrying glob metacharacters
(crafted option-looking filenames narrow invisibly) are unprovable.
Criterion-side globs stay self-consistent (literal match).
- cd: an argument carrying a runtime expansion or glob is non-silent
(unknown destination, unknown print); CDPATH= assignments are now
handled as state pollution at the match layer, subsuming the
value-shape special case.
* fix(harness): persistent-shell evidence, exact env sets, option-arity scoping (RFC #4651 PR4)
- tests_passed: on a persistent-shell provider (new
Sandbox.persistent_shell_sessions capability, set by AioSandbox) every
leaf degrades to UNVERIFIED — any earlier call in the shared session
could have mutated the state the clean-looking run executed in, and
only a fresh controlled session (RFC section 6 verifier) can prove
otherwise. The flag is read from the provider registry without
acquiring a sandbox.
- env assignments: the allowlist is gone — no variable is provably inert
across repositories (CI/DEBUG are routinely read by tests). The span's
assignment prefix must equal the criterion's exactly (values included,
order-insensitive); any assignment or export NAME= in a preceding
segment is state pollution.
- scoping: positional targets are now read by option arity, so a path
embedded in an option (--basetemp=/tmp/p, --junitxml=/tmp/r.xml) never
counts as a selection target and an extra positional after such a
criterion narrows the default selection it denotes.
* fix(harness): stamp shell provenance at harvest, close export/unset and arity gaps (RFC #4651 PR4)
* fix(harness): split physical newlines as shell separators in acceptance matching (RFC #4651 PR4)
* fix(harness): scope cd wrappers to thread data roots, pin accepted boundaries (RFC #4651 PR4)
* fix(harness): preserve criterion connectors, prove file_written readable, fail-closed shell capability (RFC #4651 PR4)
* fix(harness): compare only the connector prefix, tolerate trailing criterion semicolons (RFC #4651 PR4)
* fix(harness): preserve continuation-line operators, keep ./-spelled executable identity (RFC #4651 PR4)
* fix(harness): render criteria single-line so a multiline criterion cannot inject a forged checklist line (RFC #4651 PR4)
* fix(harness): reject parent-traversal executable tokens in acceptance matching (RFC #4651 PR4)
* fix(harness): reject parent-traversal negated values in acceptance matching (RFC #4651 PR4)
* fix(deps): depend on renamed tenki package instead of tenki-sandbox
tenki-sandbox has been removed from PyPI and republished as tenki. Its old wheel URL still resolves, so existing lockfiles keep installing and the breakage is invisible to anyone with a warm lock; any fresh resolution fails with 'tenki-sandbox was not found in the package registry'.
tenki 1.0.2 still ships the tenki_sandbox module, so the imports in community/tenki/provider.py and sandbox.py are unchanged.
Fixes#5081
* fix(tenki): point install guidance at the renamed distribution
The rename to `tenki` left the user-facing remediation still naming the
removed package. `_import_client` raised "pip install tenki-sandbox" on the
missing-extra path — the exact instruction this change proves now 404s on
PyPI, handed to the user at the exact moment they need it to work.
Update that message and the remaining `tenki-sandbox` references in the
provider, sandbox adapter, README, sandbox AGENTS.md and the test docstring.
The imported module stays `tenki_sandbox`, so the distribution and module
names now differ; each mention says so rather than just swapping the string.
No behavior change beyond the error text.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(tenki): migrate the provider to the 1.x workspace-only API
Renaming the dependency was not enough. tenki 1.0.2 keeps the tenki_sandbox
module name but not its contract: Client.create dropped project_id and has no
**kwargs to absorb it, and IdentityWorkspace no longer carries `projects`
(the attribute is gone from the package entirely). Both configuration paths
therefore failed before a sandbox could be created — explicit project scope
raised TypeError, and automatic scope raised AttributeError walking
workspace.projects.
Scope is now the workspace alone. _resolve_scope returns a single workspace id,
auto-selecting when the account has exactly one, and project_id is gone from
create_kwargs and from the documented config surface.
A stale project_id in config.yaml warns rather than fails. SandboxConfig is
extra="allow", so simply not reading the key would leave it scoping nothing
with no signal; it also used to short-circuit the identity lookup, so operators
with more than one workspace need to know they must now set workspace_id.
The suite passed against the broken provider because the fake client took
**kwargs and swallowed the project_id the real SDK rejects. The double now
mirrors 1.0.2 — keyword-only, no **kwargs — so an unexpected argument is a
TypeError in tests exactly as it is against the SDK. Reintroducing the old
create call fails 20 tests; before this change it failed none.
Verified against the exact locked wheels: every other kwarg the provider
passes (name, workspace_id, sticky, wait, max_duration, image, cpu_cores,
memory_mb, env) and every SDK surface it touches (who_am_i, Identity.workspaces,
wait_ready, exec, close, the fs API, the four terminal exception classes) is
unchanged in 1.0.2.
Reported by willem-bd in review.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* docs(config): drop sandbox.project_id from the Tenki example
The canonical example still documented project_id as a supported optional key
after the provider stopped honouring it, so an operator following it could set
the key, get no scope from it, and hit a workspace-resolution failure with
nothing in the example to explain why.
Replaced with a migration note rather than a silent deletion: someone upgrading
already has the key in their config.yaml and needs to know it is inert now and
that workspace_id is what scopes a sandbox on Tenki 1.x.
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Aniket Wagh <aniketwaghh@users.noreply.github.com>
* feat(sandbox): share sandbox identity derivation and acquire serialization (#4741)
Remote providers (AIO, E2B, BoxLite, Tenki, OpenSandbox) each inlined the
same sha256(user:thread)[:16] sandbox-id expression and kept per-scope lock
dicts that grew unboundedly until shutdown. This extracts both mechanisms
into shared components without changing provider lifecycle, ids, capacity
semantics, or public tool behavior:
- sandbox/identity.py: keyword-only derive_sandbox_scope_token (byte-pinned
compatibility contract) + is_sandbox_scope_token; per-provider golden
vectors pin current behavior including BoxLite's raw-None quirk and each
provider's private user_id resolution.
- sandbox/acquire_serialization.py: AcquireSerializer — per-key lock table
with holder/waiter refcount reclamation, bounded dedicated executor
(async waits off both the event loop and the default executor),
worker-owned cancellation cleanup (no event-loop callback dependency), idempotent close().
- Each provider adopts both components; AIO/E2B key by (user_id, thread_id)
with acquire and (E2B) release serialized; BoxLite/Tenki/OpenSandbox key
by derived sandbox id and offload the whole sync acquire to the
serializer's executor so a cancelled awaiter cannot overlap a retried
same-scope body (leaked-remote-VM regression caught in review).
- thread_id=None acquires stay unserialized; provider shutdown()/reset()
close the serializer; E2B capacity/ledger/reconciliation and AIO
ownership/flock machinery untouched.
- blocking-IO anchor proves contended OpenSandbox acquire_async stays off
the event loop (teeth verified red/green); AGENTS.md documents the
shared components.
* refactor(sandbox): address review on acquire serialization (#5089)
- Replace unreachable checkin branch with an assertion: run() returns
False only after abandon(), which the except handler always re-raises;
the old _checkin would have double-decremented the refcount.
- Document the task.cancelling() == 0 assumption in hold_async.
- Drop unused thread_id/user_id kwargs from BoxLite and Tenki
_acquire_scope_locked (OpenSandbox still forwards them).
* fix(sandbox): preserve request ContextVars in acquire executor bridge (#5089)
loop.run_in_executor() does not copy contextvars, unlike the inherited
SandboxProvider.acquire_async() which used asyncio.to_thread(). The
BoxLite/OpenSandbox/Tenki acquire_async bridges introduced in this PR
therefore dropped the request trace id (logged as trace_id=-).
Add AcquireSerializer.run_on_executor(), which copies the calling
context and runs the callable through ctx.run, and route all three
providers through it. Add regression tests binding request_trace_context
and verifying the worker thread observes it.
* Preserve Windows CLI compatibility for local sandbox commands
MSYS path conversion must remain disabled for DeerFlow virtual paths, but applying a blanket environment override to every POSIX command breaks host-native CLI shims on Windows. Limit MSYS argument-conversion exclusions to safe non-root virtual path prefixes, omit values that would broaden the exclusion pattern, and document the contract.
Constraint: Preserve the virtual-path protection introduced by #2765/#2766
Rejected: Disable MSYS conversion for every command | breaks Windows CLI shims
Rejected: Toggle blanket conversion only for commands containing virtual paths | host CLIs can receive virtual-path arguments and still need normal conversion for their own paths
Confidence: high
Scope-risk: moderate
Reversibility: clean
Directive: Keep regression coverage for virtual-path arguments, root mounts, and host-native CLI launchers
Tested: test_local_sandbox_encoding.py (12 passed); related sandbox suite (197 passed, 8 skipped, 7 failures matching origin/main); ruff check; ruff format --check; git diff --check; direct LocalSandbox CLI and virtual-path smoke tests
Not-tested: Full offline suite completion; stopped at 6% after unrelated Windows and optional-runtime failures
Related: #2765
Related: #2766
* Keep MSYS regression tests portable across CI operating systems
The Windows-shell environment tests patched os.name to nt while mounting Windows-specific paths. On Linux and macOS, pathlib then attempted to construct WindowsPath during command resolution or output masking, so the backend merge gate failed before exercising the environment contract. Stub the exclusion boundary in execute-command tests and retain mapping-specific filtering coverage in the helper test.
Constraint: Backend unit tests run on Linux, while the behavior under test is Windows-only
Rejected: Skip the tests outside Windows | would remove CI coverage of the environment contract
Rejected: Patch pathlib internals | couples tests to implementation details and hides the platform boundary
Confidence: high
Scope-risk: narrow
Reversibility: clean
Directive: Keep OS-specific subprocess assertions independent from host-path resolution
Tested: test_local_sandbox_encoding.py (12 passed); ruff check; ruff format --check; git diff --check
Not-tested: Linux runner execution locally because Docker Desktop is unavailable and WSL cannot access this linked worktree
Related: #5003
Related: https://github.com/bytedance/deer-flow/pullrequestreview-5013380238
Sandbox is an execution environment, not a named resource: multiple tools
(bash, read_file, write_file, glob, grep, ...) depend on it, all funneled
through ensure_sandbox_initialized / ensure_sandbox_initialized_async. Gate
the single acquisition entry point (single source of truth) instead of
maintaining a sandbox-tool-name set in middleware:
- authorize_sandbox_execution helper (authz/sandbox_authz.py) checks
authorize("sandbox", "execute", target="*") — a binary judgment
(can this role use the sandbox at all); RBAC allow:"*"/true permits,
allow:[]/false denies.
- lazy path: ensure_sandbox_initialized (+ async) calls the gate before
provider.acquire.
- eager path: SandboxMiddleware.before_agent / abefore_agent call the gate
before _acquire_sandbox.
- deny raises SandboxAuthorizationError (SandboxError subclass) which
propagates through tool execution as a friendly ToolMessage (RFC §9:
'not a crash').
- authorization.enabled: false is a no-op everywhere; provider errors
follow fail_closed (deny) / fail_open (allow).
12 tests in tests/test_sandbox_authorization.py cover disabled/allow/deny/
deny-via-bool/no-policy-unrestricted/provider-error-fail-closed/open/
internal-caller + ensure_sandbox_initialized deny (never acquires) and
allow (acquires) integration paths.
* docs: govern agent guidance size
* refactor: split agent guidance by code scope
* Clarify virtual path handling in AGENTS.md
Updated the translation section to clarify the role of `LocalSandboxProvider` and the handling of virtual paths in the tool layer.
---------
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
* fix(sandbox): project enabled skills into sandbox views
* fix(skills): keep projection mutations consistent
* fix(skills): fail closed on projection errors
* fix(skills): isolate per-scope failures during boot projection rebuild
rebuild_all_skill_projections() propagated any exception from the public
rebuild or from a single user's rebuild straight out of the gateway
lifespan startup, uncaught. A single broken user directory (bad
permissions, corrupted _skill_states.json, unreadable content) would
therefore abort gateway boot for every user, not just that one -
_rebuild_*_locked already fails closed internally (clears the view and
re-raises), so the boot loop only needed to stop treating that re-raise
as fatal.
Each scope's rebuild now fails closed independently and boot continues;
a scope left empty by a boot failure self-heals on the next sandbox
acquire via ensure_skill_projections().
Also patches deerflow.skills.projection.rebuild_all_skill_projections in
the memory-flush lifespan test fixture, matching the two sibling
fixtures in the same file — this call is now on the lifespan startup
path and the fixture's minimal SimpleNamespace config predates it.
* test(skills): update authz test for the projection-aware public toggle
_persist_shared_skill_state (introduced earlier in this branch) reads
the shared extensions_config.json fresh from disk under the projection
lock instead of through the cached get_extensions_config() singleton -
that's the whole point of the fix (stale worker caches must not clobber
another worker's concurrent update). The name no longer exists on the
skills router module, so the test's monkeypatch of it started raising
AttributeError instead of exercising the endpoint.
The mock storage in this test isn't a real LocalSkillStorage instance,
so _persist_shared_skill_state's projection-mutation branch is already
skipped (nullcontext) and it falls back to a fresh ExtensionsConfig()
for the nonexistent tmp config_path - no replacement monkeypatch needed.
* fix(sandbox): make skill projection ensure best-effort in acquire
acquire() called _ensure_skills_projection() directly, outside any
try/except, in both LocalSandboxProvider and AioSandboxProvider. Every
other skill-mount setup path in these providers has always caught
exceptions and logged a warning rather than failing sandbox acquire
outright (e.g. when config.yaml can't be resolved) - these two new call
sites broke that contract, so any projection failure (including simply
not having a config.yaml, as in CI's test environment) now failed
acquire() itself instead of just leaving skill mounts off.
_ensure_skills_projection now catches its own exceptions and returns
None; both providers' callers already tolerate that (a None projection
skips the skill-specific mounts, matching the existing degrade path)
after making _append_public_skill_mapping and the custom/legacy mount
block in LocalSandboxProvider explicitly None-safe.
Caught by running the full suite with config.yaml removed, matching
CI's environment - not caught locally because a real config.yaml was
present, masking the failure.
* fix(sandbox): make E2B skill projection mounts best-effort
_skill_projection_mounts called ensure_skill_projections with no guard,
unlike Local/AIO's _ensure_skills_projection. A raise propagated out of
_apply_mounts before the configured-mounts loop ran, so a skills
projection failure dropped the operator's own configured mounts too -
only caught by create()'s outer warning, with nothing applied at all.
Swallow here and return an empty mount list on failure, matching the
Local/AIO pattern: still fail-closed for skills, but no longer widens
the blast radius to unrelated configured mounts.
Review feedback from PR #4178.
* docs(skills): document projection trade-offs flagged in review
- _update_tree_digest: note the metadata-only (not content) hashing
trade-off and why runtime writes through this codebase are still
covered regardless (rebuild-under-lock + rename always changes inode).
- LocalSandboxProvider.acquire: note the acquire-time self-heal cost
(cheap on a fresh manifest, ~400ms rebuild under lock on stale/drift).
- skill_projection_mutation: drop the no-op except-Exception-then-raise;
a raise from the mutation already propagates past the yield with the
view left cleared, no explicit re-raise needed.
- provisioner README: spell out that hostPath skills volumes require
the gateway and K8s node to share DEER_FLOW_HOST_BASE_DIR (single-node
or shared storage), and that the custom/legacy volumes' hostPath type
Directory (not DirectoryOrCreate) makes a violation of that assumption
a visible Pod-creation failure instead of a silent empty mount.
Review feedback from PR #4178.
* fix(skills): lazily repair user projections
* fix(skills): close projection review gaps
* fix(skills): refresh user projection enable state
* fix(skills): close projection review follow-ups
* fix(skills): preserve state across projection writes
---------
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
* fix(sandbox): unwrap Overwrite-wrapped state in ensure_sandbox_initialized
The same fork-restored wrapper that crashed after_agent also reaches the sandbox init path, where sandbox_state.get() on the Overwrite object raises AttributeError. Share the unwrap helper from #4381's follow-up module deerflow/sandbox/overwrite.py and apply it at both init sites.
Signed-off-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com>
* fix(sandbox): note why discarding fork_restored at the reuse sites is safe
* fix(sandbox): unify the Overwrite unwrap helper and pin the fall-through
- middleware.py now imports unwrap_sandbox from overwrite.py instead of
keeping a second local copy whose docstring had already drifted; the
shared helper covers both crash forms (subscript TypeError and the
.get()-form AttributeError)
- test the acquire fall-through: when the fork-restored id is gone from
the provider, a fresh sandbox is acquired and the stale wrapped state
is replaced by the plain acquired dict
- the reuse-path test now also asserts runtime.state["sandbox"] stays
wrapped, pinning the don't-treat-as-owned contract after_agent relies on
Signed-off-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com>
* fix(sandbox): unwrap Overwrite state in the sibling sandbox readers
Signed-off-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com>
---------
Signed-off-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com>
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
* feat(lark): sidecar credential broker for sandbox lark-cli (Pattern B)
Removes the plaintext Lark credential mounts (appSecret + OAuth tokens)
from the sandbox container. A long-running broker sidecar owns lark-cli
and the per-user config/data dirs and serves the command surface over
Pod loopback; the sandbox gets only a forwarding shim on PATH, so the
raw credential files never exist in the sandbox filesystem.
- lark_broker.py: stdlib-only loopback broker (argv passthrough with
shell=False, server-injected credential env, bounded I/O) + shim
script constant + install-shim mode.
- docker/lark-cli-broker: init(install-shim) + serve image.
- provisioner: LARK_CLI_BROKER_IMAGE + provision_lark_cli_broker →
shim init container + lark-cli-broker sidecar (config/data mounted
sidecar-only); credentials dropped from the sandbox container;
/api/capabilities reports lark_cli_broker_image. Broker supersedes
the Pattern A init-container binary when both are configured.
- gateway: lark_cli_env_overlay(broker=True) omits config/data env;
sandbox_lark_broker_active() TTL-cached mode resolver; broker added
to sandbox_runtime_mode / readiness and the settings UI.
Opt-in and off by default (empty LARK_CLI_BROKER_IMAGE ⇒ no change).
Closes#4338
* fix(lark): address Pattern B broker review findings (#4501)
Follow-up to the sidecar credential broker addressing the PR #4501 review:
- shim: split the on-PATH lark-cli into a /bin/sh launcher + Python shim body
so broker mode fails loudly (exit 127, actionable message) instead of ENOEXEC
when the sandbox image ships no python3; interpreter pinnable via
DEERFLOW_LARK_BROKER_PYTHON. Launcher bakes in the shim's absolute path since
$0 is the bare command name when run off PATH.
- broker: drop the dead cwd payload field (broker can't see the sandbox FS) and
document the command-surface-only / no-file-IO limitation.
- broker: return a structured 500 JSON on unexpected exec errors so the shim
gets a meaningful message, not an opaque transport failure; set a handler
socket timeout to bound slow/stuck connections.
- broker: add an opt-in DEERFLOW_LARK_BROKER_DENY_SUBCOMMANDS denylist that
refuses secret-dumping subcommands before spawning the binary, forwarded from
the provisioner sidecar.
- gateway: tighten the per-bash-call broker probe timeout (1.5s) and cache
negatives longer (300s) so non-broker remote-provisioner users don't pay a
latency hit; guard the mode cache with a lock; drop the dead
_probe_provisioner_lark_cli_init_image wrapper.
- docs: remove the broken design-doc link from the broker README.
Adds tests for launcher python resolution, cwd omission, denylist enforcement,
500-on-error, hot-path probe timeout + negative caching, and provisioner
denylist-env wiring.
* fix(sandbox): unwrap Overwrite-wrapped sandbox state in after_agent
Fork-restored checkpoints can deliver the sandbox channel still wrapped
in langgraph.types.Overwrite: the rollback restore applies replace-style
writes through a state-mutation graph in delta checkpoint mode, and
after_agent/aafter_agent then crash subscripting the wrapper
("TypeError: 'Overwrite' object is not subscriptable") on the next
sandbox tool run in the forked conversation. Unwrap before reading the
sandbox id, and pin both hooks against an Overwrite-wrapped state.
Refs #4380 (bug 1 of 2; the history-loss half is a separate display path)
Signed-off-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com>
* fix(sandbox): don't release fork-restored parent sandboxes
The Overwrite-wrapped value replays the parent thread's sandbox state, so releasing it from the forked run would evict the parent's warm sandbox. _unwrap_sandbox now reports the wrapped form, and both after_agent hooks skip the release for it while keeping the normal path unchanged.
Signed-off-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com>
---------
Signed-off-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com>
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
* feat: add lark cli integration
* fix: polish lark integration actions
* feat: support lark incremental permissions
* fix: detect lark authorization completion
* fix: harden lark integration install
* feat: expand lark auth scopes and reuse host auth in sandbox
Default lark auth to least-privilege (recommend=false, base sign-in only)
and expose the full set of lark-cli --domain business domains as native
--domain grants instead of a 4-domain read-only mapping. Resolve the
skill pack from the latest larksuite/cli GitHub release at install time
with content-hash integrity, and surface version/runtime drift in status.
Share the per-user lark-cli config/data profile between the Gateway
Settings auth flow and agent conversations by mounting the integration
dirs into the AIO sandbox and injecting the matching env for lark-cli
commands, with an allowlisted extra_mounts path in the provisioner/K8s
backend and traversal guards on integration paths.
* style: fix lint issues from ruff and prettier
Sort imports in the provisioner PVC test and re-wrap two long i18n
description strings to satisfy backend ruff and frontend prettier CI.
* fix(lark): address managed integration review feedback
* fix(frontend): stabilize integrations settings e2e
* test(sandbox): isolate remote backend legacy visibility check
* test: fix backend unit failures after merge
* Harden Lark integration review fixes
* Format Lark integration E2E test
* fix(lark): harden sandbox credential exposure and status disclosure
Address willem_bd's security review on PR #3971:
- Mount the per-user lark-cli config dir (long-lived appSecret) read-only
into the AIO sandbox; only the refreshable-token data dir stays writable.
- Redact host filesystem paths (install_path, cli.path) from
GET /lark/status and the config/auth complete responses for non-admin
callers, fail-closed on any auth error.
- Document the npm postinstall trade-off (--ignore-scripts is not viable
because @larksuite/cli fetches its platform binary in postinstall).
- Document the sandbox credential trust boundary in AGENTS.md and README,
pointing at the sidecar-broker follow-up (#4338).
---------
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
str_replace guards the replacement with `if old_str not in content`, which
cannot reject an empty old_str -- `"" in content` is always true. So an
empty old_str reached `str.replace("", new_str)`, which inserts new_str at
every character boundary, and the tool rewrote the file while still
returning "OK":
old_str='', new_str='# H\n' -> OK, file silently prepended
old_str='', new_str='X', replace_all -> OK, 'XdXeXfX XmXaXiXnX(X)X:X\nX...'
The empty-file branch above it already handles this case (`if not content:
if not old_str: return "OK"`), and the existing test states the intent
directly: "An empty old_str is a no-op edit and remains a benign OK". That
contract just never held once the file had content.
The tool is registered by default (config.example.yaml) and its schema
declares old_str as a plain string with no minLength, so a model can emit
"" legitimately; read-before-write only compares a hash and lets it past.
Check old_str first so the no-op holds whatever the file contains. The
empty-file case folds into the same not-found branch, which keeps its
message and behaviour.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(sandbox): give the host→virtual output-mask regex a single owner
Two call sites rewrite host paths back to their virtual form in text that
reaches the model — LocalSandbox._reverse_output_patterns (bash output) and
sandbox.tools._compiled_mask_patterns (glob/grep/ls results) — and each built
the same `escape(base) + boundary + tail` rule from its own copy.
That duplication has already produced two bugs: #4035 added the segment
boundary to the reverse patterns and missed the masking patterns, and #4053
had to add the same boundary to the other copy. Extract the rule into
sandbox/path_patterns.py so a third copy cannot silently disagree.
The extraction is not a pure move: the two sites disagree on the base. tools.py
derives bases from _path_variants (which yields Windows spellings) and matches
them against output whose separators it does not control, so it relaxes the
separators inside the base; LocalSandbox resolves its bases from the running
platform and must not be widened. That difference is now an explicit
`separator_agnostic` parameter rather than an accident of two implementations.
The boundary and tail constants are private: build_output_mask_pattern is the
only supported spelling, so a third site cannot import the pieces and hand-roll
a variant.
Behavior is unchanged at both sites — pinned by tests that reproduce each
pre-extraction expression byte-for-byte.
* test(sandbox): pin the base the helper must not normalize
Review notes on #4108.
The committed snapshot compares the helper against hand-copied literals of the
pre-extraction expressions, so its red-ness rests on those literals, not on the
length of _BASES -- both sides compute the same expression, and 5k fuzzed bases
produce zero byte-differences. Mutating the helper one clause at a time (12
mutations over the boundary, the tail and the escape/replace) shows the seven
committed bases catch 11: the miss is a helper that normalizes its input by
rstripping a trailing separator. Only a trailing-slash base or a Windows drive
root catches that, and Path.resolve() / str(Path(...)) strip trailing slashes,
so neither call site can produce the former. C:\ survives resolve() with its
separator intact, so that is the one base worth adding.
Also point local_sandbox's comment at path_patterns, the owner, instead of
citing _content_pattern as the class reference, and drop the rationale it now
duplicates from the owner's docstring -- a second copy of the explanation drifts
the same way the second copy of the regex did. The site-specific half stays.
Comments and test data only; no behavior change.
* fix(sandbox): use os.sep in reverse-resolve containment check on Windows
Path.resolve() always renders with the native separator (backslash on
Windows), but _reverse_resolve_path's containment check hardcoded a
"/" suffix when testing whether a resolved path is nested under a
mapping's local root. Only the exact-root case (no separator needed)
ever matched; every nested path fell through to the "no mapping
found" branch and returned the raw host path -- leaking the real
username and full directory tree into list_dir/glob/grep results and
bash output masking instead of the virtual /mnt/user-data/... path.
_is_read_only_path already does the equivalent check correctly via
os.sep, so this aligns _reverse_resolve_path with that pattern: the
containment check now compares with os.sep, and the extracted
relative portion is normalized to forward slashes before being
spliced into the (always POSIX-style) container path.
Also fixes a same-file cosmetic bug in list_dir's virtual
sub-directory overlay: it compared a bare child name (e.g.
"workspace") against a set of full container paths, so the
already-listed guard never matched and a mount whose subdirectory the
underlying scan already found (the common case for
/mnt/user-data/workspace, uploads, outputs) was appended a second
time.
Continues the same separator-bug class already fixed in this file by
#3869 (forward-direction command resolution) and #4035 (reverse
regex-boundary matching); neither touched this containment check.
* test(sandbox): add host-OS-independent regression test for the os.sep containment fix
_reverse_resolve_path's os.sep containment check (and the paired
lstrip(os.sep).replace(os.sep, "/") extraction) has no test that would
fail if reverted: backend CI runs only on ubuntu-latest, where
os.sep == "/" makes the pre-fix hardcoded "/" and the current os.sep
form observationally identical, so a plain POSIX-path test can't
discriminate between them.
Add a test that forces the Windows code path independent of host OS by
monkeypatching os.sep to "\" and stubbing both the module's Path name
and the sandbox's cached _resolved_local_paths to return
backslash-joined strings, mirroring what real WindowsPath.resolve()
produces -- without touching the filesystem or requiring an actual
Windows host. Verified this fails with the raw host path leaking
through when the os.sep fix is reverted to the hardcoded "/" form, and
passes with the fix in place.
replace_virtual_paths_in_command matches the virtual root with no
segment-boundary lookahead:
re.compile(rf"{re.escape(VIRTUAL_PATH_PREFIX)}(/[^\s\"';&|<>()]*)?")
The trailing group needs a "/" to consume anything, so when the character after
/mnt/user-data is "-", ".", "_", a digit or a letter, the group matches empty
and the bare root still matches. The substitution then rewrites it to the
thread's host user-data directory and the rest of the sibling name rides along,
so a command naming a prefix sibling of the mount root is pointed at a real host
directory outside the mount contract:
cat /mnt/user-data-backup/secret.txt -> cat <host>/user-data-backup/secret.txt
which reads the host file. This is the same defect as #4035 (reverse patterns)
and #4053 (masking patterns), mirrored into the virtual->host direction; it is
the last unguarded member of that family.
The boundary class mirrors LocalSandbox._content_pattern's rather than
_command_pattern's: a virtual root can legitimately be followed by ":"
(PATH-style concatenation) or ",", which the shell-oriented class rejects, so
narrowing to it would stop translating paths that translate today. "$" covers a
command ending exactly at the root.
str_replace returned "OK" whenever the target file was empty, silently
reporting success even when the model asked to replace a non-empty string
that could not possibly be present. Only short-circuit to "OK" when old_str
is also empty; otherwise return the standard not-found error.
* fix(sandbox): stop glob/grep/ls from surfacing disabled skills' files
The disabled-skill gate checks the path a tool is given, but ls, glob and
grep all descend from it and return other paths, so a root above a disabled
skill still serves its files. glob and grep never called the gate at all;
ls called it only on its own argument and still leaked from a category root.
Add the entry gate to glob/grep, and filter what all three return through
the existing fail-closed _is_disabled_skill_path. The verdict is memoized per
skill because ExtensionsConfig.from_file() is uncached, so a per-match check
would turn a 100-match grep into 100 config reads.
* fix(sandbox): normalize trailing slashes in the disabled-skill path check
Review follow-ups on the disabled-skill gate:
- _extract_skill_name_from_skills_path returned "" instead of None for a
category directory carrying a trailing slash. LocalSandbox.list_dir appends
"/" to directories, so `ls /mnt/skills` yields "/mnt/skills/public/", giving
parts ["public", ""]. The empty name skipped the `skill_name is None`
short-circuit and fell through to a config read, landing on the right outcome
only because unknown skills default to enabled. Drop empty segments so a
trailing-slash category root takes the existing category-root branch.
- ls_tool resolved the runtime user id twice per call; hoist it into a local,
matching glob_tool/grep_tool.
- Cover the CUSTOM path: custom/legacy skills resolve their enabled state
through the per-user _skill_states.json, a different store from the public
skills' extensions_config.json, and no automated test exercised it.
* fix(sandbox): guard the output-masking regex with a segment boundary
`_compiled_mask_patterns` builds the same class of host→virtual matcher as
`LocalSandbox._reverse_output_patterns`, but without the segment-boundary
lookahead that one carries. The trailing group needs a separator to consume
anything, so when the character after a host base is `-`, `.`, `_`, a digit or
a letter, the group matches empty and the regex still matches the bare base.
`replace_match` then takes its `matched_path == base` branch and rewrites the
sibling: with `/mnt/skills` mounted at `.../skills`, output naming a sibling
`.../skills-extra/data.txt` is handed to the model as `/mnt/skills-extra/data.txt`
— a container path forward resolution explicitly refuses to map back, so
reading it raises FileNotFoundError.
This is the sibling site of #4035, which fixed the identical bug in
`local_sandbox.py`. That PR's scope argument enumerated the prefix matchers in
that file and missed this one; `mask_local_paths_in_output` runs on every
glob/grep match and on local bash output.
The boundary class mirrors `_content_pattern`'s, not `_command_pattern`'s: this
runs over arbitrary command output, where a base can legitimately be followed
by `,`, `:` or `\`, all of which the shell-oriented class rejects.
* test(sandbox): anchor the sibling boundary at the ACP source too
The sibling-rejection cases only fed the skills source. `_compiled_mask_patterns`
builds every source's matcher in one loop, so the ACP workspace carried the same
defect: nothing maps its parent, and `/mnt/acp-workspace-backup/hello.py` is
unresolvable in both directions.
User-data is the exception and is now pinned as such: `_thread_virtual_to_actual_mappings`
also maps the virtual root `/mnt/user-data` to the three dirs' common parent, so a
sibling of `outputs` is still inside a mount and has a real virtual path — the output
is byte-identical with and without the boundary. That test is green on main; it guards
the boundary from being narrowed into one that stops translating a mapped path.
Reverting only `boundary` turns the 5 skills + 4 ACP cases red and leaves the 3
user-data cases green.