* fix(sandbox): report an exactly-full search result as complete in the remote providers
`glob` and `grep` decide `truncated` twice: once for the raw output cap
(`parse_remote_search_output`, unchanged) and once for `max_results` after the
Python-side filters have run. The second decision returned as soon as
`max_results` matches had been collected, which cannot tell a search that held
exactly that many from one that held more — a tree holding exactly
`max_results` eligible matches came back flagged as cut off, and the tool then
told the model the result was incomplete.
These providers hold the whole listing (the raw stream is capped at
`max(max_results * 4, max_results + 50)` lines and reports its own cut-off), so
like AIO's `glob` branches they can look one match past the cap before
deciding: `AioSandbox.grep`, plus `glob`/`grep` in E2B, OpenSandbox, Tenki and
BoxLite now use the same `len(matches) > max_results` rule. This completes what
#5449 started for AIO's `glob`; the local provider's half is #5491.
Co-Authored-By: Claude Code <noreply@anthropic.com>
* fix(sandbox): let remote grep see one match past the per-file cap
E2B and OpenSandbox stopped each file's grep at max(max_results, 50)
matches, so a single file holding more than max_results hits — with a
raw stream far below its limit — ended the Python loop exactly at the
cap and reported the result as complete (#5534 review).
Retain one extra match per file so the one-match lookahead can observe
the overflow and report truncation. A single-file regression at
max_results=50 covers 50 matches (complete) vs 51 (truncated) for both
providers.
Co-Authored-By: Claude Code <noreply@anthropic.com>
---------
Co-authored-by: Claude Code <noreply@anthropic.com>
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
* fix(ragflow): batch validation for large document selections
* docs(ragflow): align documentation language with repository conventions
* docs(ragflow): preserve spacing before validation heading
---------
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
* feat(knowledge): add verifiable RAGFlow source citations
* docs(knowledge): scope RAGFlow guidance to its own directory
* fix(knowledge): preserve citations through rendering and budgets
* fix(sandbox): default to loopback bind on Docker Desktop for DooD sandboxes (#5445)
* fix(sandbox): memoize desktop detection and clarify bind host docstring (#5445)
* fix(sandbox): latch desktop detection on success only to permit retry on transient failure (#5445)
* fix(sandbox): restrict desktop loopback bind to local DooD hostnames (#5445)
* fix(sandbox): add Desktop legacy aliases and parametrize DooD host tests (#5445)
* fix(sandbox): report an exactly-full AIO glob result as complete
AioSandbox.glob's include_dirs branch returned as soon as it had
collected max_results matches, without looking at the rest of the
listing. A listing that held exactly that many matches and nothing more
was therefore reported as truncated, and the glob tool told the model
the result was incomplete — prompting a re-search or distrust of a
complete answer. The same line returned one match for max_results=0,
one past the caller's cap.
Look one match past the cap before deciding, which is what the
include_dirs=False branch in the same function already does and what
#5427 moved parse_remote_search_output to for BoxLite, Tenki, E2B and
OpenSandbox.
* review: filtered-tail cases, the glob contract docstring, and the cap wording
Addresses the three items from the review on #5449.
- Two regression cases over a tail of ignored / out-of-root / pattern-miss
entries: an exactly-full result stays complete when only filtered entries
follow, and a third eligible match after that tail still reports
truncation. Both fail against the previous return-on-the-max-th-match
behaviour.
- 'Sandbox.glob' promised the conservative flag ('``max_results`` was
reached') that this change deliberately stops producing on the AIO branch.
The contract now reads as 'may be incomplete' and records that providers
differ in how precisely they can decide it.
- The changelog no longer lumps 'parse_remote_search_output' in with the
filtered-match cap: its raw-output cap is a separate limit with its own
one-line-past accounting, and the other providers' filtered-match cap is
unchanged.
Also corrects the docstring on the existing test, which still described the
removed early return in the present tense.
---------
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
* fix(sandbox): report truncated remote glob and grep results
BoxLite, Tenki, E2B, and OpenSandbox run find/grep in the sandbox, cap
the raw output with `| head`, and then filter those lines in Python:
ignored directories such as node_modules are dropped and grep's glob
scope is applied. They reported truncated only when max_results matches
survived the filter. When the capped lines were mostly filtered out, a
search with real matches past the cap came back short or empty with
truncated=False, and glob_tool/grep_tool rendered it as "No files
matched" / "No matches found". With the default max_results=200 and
1,200 files under node_modules, glob("**/*.py") reported no matches for
a workspace that has src/app.py.
remote_search_command now lets one line past its limit through, and
parse_remote_search_output(..., limit=) returns RemoteSearchOutput(text,
truncated): the first `limit` lines and whether the extra line arrived.
Exactly `limit` lines stays a complete result. Each provider passes the
cap it already computed to both calls and returns that truncated from
glob and grep when fewer than max_results results survive filtering.
The glob and grep tools now describe an empty truncated result as
incomplete instead of reporting no matches, which also covers AIO grep's
forwarded truncated flag. Sandbox.glob/grep document truncated as "the
matches may be incomplete".
* docs(changelog): reference #5427 in the remote search truncation entry
---------
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
* fix(sandbox): scope BoxLite grep globs to the search root
BoxliteBox.grep omits grep --include for busybox portability and applies
the glob in Python, but it kept only the glob's last path segment and
matched it against each file's basename. A scoped pattern therefore lost
its directory part: grep(glob="src/*.js") returned every .js file in the
tree, including vendor/ and nested src/ subdirectories that glob() with
the same pattern excludes.
The glob now goes through path_matches against the path relative to the
search root, with the file's basename when the root is a single file --
the same scope glob() uses and the one Tenki, E2B, OpenSandbox, AIO and
LocalSandbox already enforce. Like those providers, an empty glob is now
passed to path_matches instead of being treated as no filter.
* docs(changelog): reference #5419 in the BoxLite grep glob scope entry
_get_firecrawl_client only read api_key from the tool config and ignored
base_url, so self-hosted Firecrawl deployments were unreachable: the SDK
defaulted to https://api.firecrawl.dev and raised 'No API key provided'.
Now base_url is forwarded as api_url; api_url is omitted when unset so
cloud behavior is unchanged.
Also document the optional self-hosted base_url on both Firecrawl entries
in config.example.yaml and reconcile their headers to the fastCRW house
style (Cloud requires FIRECRAWL_API_KEY; self-host may need no key.).
* fix(sandbox): stop remote grep/glob from reporting failures as no matches
E2B, OpenSandbox, BoxLite and Tenki ran grep/find behind `2>/dev/null | head`, so a missing search root, a missing grep/find binary or an unreadable tree exited 0 with empty stdout and the tools reported "No matches found". Wrap the search in sandbox/remote_search.py, which checks the root first and records the search's own status after head, as remote_list_dir does for list_dir: a missing root raises FileNotFoundError, a failed search raises OSError, and a genuine no-match still returns []. glob's find gains -H for symlinked roots, OpenSandbox's BusyBox fallback keeps the primary grep status, and E2B no longer swallows client errors. Regression tests run each provider's real command in a local POSIX sh.
Fixes#5376
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(sandbox): fail remote grep/glob on partial traversal errors
grep 2 / find 1 after some results were printed (an unreadable file or
subdirectory) were returned as a complete search. Callers have no
partial-result channel, and #5376 asks for permission and command
failures to raise, so these statuses now raise OSError like any other
failure. Only grep 0/1/141 and find 0/141 pass.
The error for grep 2 / find 1 says that some files or directories could
not be read and asks for a narrower path, so the agent can retry instead
of giving up.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Totoro-qaq <279883115+Totoro-qaq@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix(web-fetch): resolve relative URLs in extracted Markdown
Pass the request URL through Jina and Browserless extraction and resolve
link/image destinations before Readability removes document base tags.
Preserve the optional legacy API and fallback text behavior.
Fixes#5307
Signed-off-by: tiammomo <26957354+tiammomo@users.noreply.github.com>
* fix(web-fetch): address provider and base URL review feedback
Signed-off-by: tiammomo <26957354+tiammomo@users.noreply.github.com>
* fix(web-fetch): preserve HTML source when resolving destinations
Signed-off-by: tiammomo <26957354+tiammomo@users.noreply.github.com>
---------
Signed-off-by: tiammomo <26957354+tiammomo@users.noreply.github.com>
* fix(sandbox): stop list_dir from reporting failures as empty
Remote providers swallowed find/client errors as [] and 2>/dev/null
missing paths as empty stdout. ls_tool then told the agent the
directory was (empty). Raise OSError/FileNotFoundError instead so
the tool returns Error.
* fix(sandbox): list_dir raises on missing local paths and uses find -H
Empty stdout is not a missing path when find's start point is a
symlink (E2B /mnt/acp-workspace). Dereference only the start point
with find -H. LocalSandbox now raises FileNotFoundError for a
non-directory root, matching remote providers. AIO maps a missing
result.data to OSError rather than FileNotFoundError.
* fix(sandbox): group AIO list_dir find type predicates
Without parentheses, find PATH -maxdepth N -type f -o -type d applies
-type d without maxdepth and can drop files from the listing.
* fix(sandbox): distinguish list_dir command failure from missing path
Tenki, Boxlite, and OpenSandbox treated any empty find stdout as
FileNotFoundError, so a missing find binary (exit 127) or SDK error
looked like a missing directory. Raise OSError when find status is
outside (0, 1); keep FileNotFoundError for the find-ran-but-empty case.
* fix(sandbox): apply list_dir exit-status contract to AIO and E2B
Same gap as Tenki/Boxlite/OpenSandbox: empty find stdout with exit 127
was FileNotFoundError. Raise OSError when the status is outside (0, 1).
* fix(sandbox): classify list_dir by find status not head status
find | head under sh -lc reports head's exit code, so a missing find
binary (127) became FileNotFoundError. Record find's own status after
the bounded listing, treat SIGPIPE 141 as truncation success, and add
a shell-level regression test.
* test(auth): include projects permissions in /me contract pins
#5265 added projects:read/write/delete to the registered route set.
The /auth/me tests still pinned the pre-projects list, so CI failed
after merging main.
* fix(sandbox): do not treat missing list_dir marker as success
The generated script ended on `rm -f`, so process status was 0/1 even
when find's marker never landed. Both codes are in _FIND_OK, and the
parser fallback then classified an empty listing as FileNotFoundError —
the 127 misclassification this helper was meant to close.
Exit with find's status (126 if unknown). A missing marker is now
OSError unless the process status is already a non-OK failure.
* test(sandbox): emit list_dir status marker in provider fixtures
Parser now requires __DF_FIND_STATUS__ and refuses marker-less stdout.
Update AIO/Boxlite/E2B stubs and OpenSandbox/Tenki find fakes so listings
carry :0 and missing paths carry :1 with matching exit codes.
* style(sandbox): format list dir test fixture
* style(sandbox): format remote list dir helper
* docs(sandbox): keep guidance within the tested size budget
---------
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
* feat(community): add Sofya web search provider
Add a community provider backed by Sofya (https://sofya.co). Its search
endpoint returns the content of the result pages, not only their snippets,
and its fetch endpoint returns a page as markdown. Both are plain JSON over
HTTP, so this needs no extra Python package (uses httpx, already a
dependency).
Changes:
- backend/packages/harness/deerflow/community/sofya/__init__.py
- backend/packages/harness/deerflow/community/sofya/tools.py
Implements web_search_tool and web_fetch_tool using httpx.
API key is read from the config.yaml `api_key` field or the SOFYA_API_KEY
env var. Follows the same interface and output shape as the existing
ddg_search and serper providers, including the max_results parameter with
config override and the structured "No results found" error.
- backend/tests/test_sofya_tools.py
Unit tests covering API key resolution, config overrides, result mapping,
time range, HTTP errors, empty results, and fetch failures.
- config.example.yaml: add commented-out Sofya web_search and web_fetch
examples alongside the other providers
- .env.example: add SOFYA_API_KEY placeholder
- backend/docs/CONFIGURATION.md: list Sofya under web_search, web_fetch and
the environment variables
* fix(sofya): honor caller max_results, validate search_depth, join time_range contract test
- Caller-supplied max_results now wins; config is used only when the
argument is omitted, matching GroundRoute.
- search_depth is clamped to basic/snippets; an unsupported value logs a
warning and falls back to basic.
- Sofya added to the shared time_range schema contract test.
* fix(sofya): cap per-result content so a search stays inline
An unbounded search payload (up to 20 read pages) crossed the tool output
budget middleware's externalize_min_chars threshold, which replaces the
result list with a file reference. Cap each result's content at
contents_max_characters (default 2000, 0 disables), matching Exa's config
key. Five capped results stay under the 12000 char threshold.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016TZhyPNCX2GYyBPkvTgJV5
* fix(sofya): list Sofya in the recency contract, coerce non-string content
_clip subscripted its input, so a non-string content or description from
the API raised TypeError instead of degrading. Coerce to text first, the
way _sofya_post and _response_results guard the shapes around it. Also add
Sofya to the Web Search Recency section in backend/AGENTS.md.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016TZhyPNCX2GYyBPkvTgJV5
* fix(sofya): coerce web_fetch content, list sofya in the tools guide, add changelog
web_fetch sliced its content the same way web_search did before the last
push: a truthy non-string from the API passed the falsiness guard and then
raised TypeError. Reuse _clip, keeping the `or ""` so empty content still
reports "No content found".
Also add sofya to the community provider inventory in
packages/harness/deerflow/tools/AGENTS.md and an [Unreleased] changelog entry.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016TZhyPNCX2GYyBPkvTgJV5
* docs(zh): add the missing InfoQuest and Firecrawl web_fetch tabs
The ZH web_fetch tab list named five providers where EN names seven. Both
tabs mirror their EN counterparts, so the two locales list the same
web_fetch providers again.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016TZhyPNCX2GYyBPkvTgJV5
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
* fix(sandbox): stop E2B append from overwriting on read failure
E2B has no native append, so write_file(append=True) read-modify-writes.
Treat only FileNotFoundException/FileNotFoundError as an empty file; any
other pre-read error must abort so a timeout cannot replace the original
contents with just the new fragment.
* fix(sandbox): distinguish E2B append pre-read refusal in logs
A non-not-found pre-read error now logs as a refused overwrite instead
of a write failure. Tests pin the successful read-modify-write path,
including a bytes pre-image, so dropping `existing` cannot go green.
* feat(e2b-sandbox): make mount upload deadline configurable
Replace the hardcoded 120-second mount upload deadline with a
configurable `mount_upload_deadline_seconds` key read from
SandboxConfig (extra=allow). The value is validated: zero and
negative inputs are clamped to 1 second. Omitting the key
preserves the existing 120-second default.
This addresses the follow-up from PR #4842 review: operators
with large mounts or slow networks can now size the deadline to
their deployment without changing code.
* fix(e2b-sandbox): address review feedback on configurable deadline
- Remove import-time default capture from _mount_deadline_reason()
and _MountUploadBudget.deadline_seconds to prevent silent drift.
- Add warning log when mount_upload_deadline_seconds is clamped to 1
(was silent before).
- Update AGENTS.md E2B Mount Uploads section: deadline is now
configurable, not fixed 120.
- Add mount_upload_deadline_seconds to YAML examples in provider
docstring and __init__.py.
- Add config-path test that exercises SandboxConfig -> _load_config ->
_apply_mounts end-to-end.
* feat(e2b-sandbox): surface structured mount upload result on sandbox
Introduce MountUploadResult dataclass and attach it to
E2BSandbox.mount_upload_result after creation. This makes mount
truncation observable in code without re-parsing Gateway logs.
_apply_mounts() now returns MountUploadResult with truncated, reason,
and upload totals. _create_sandbox() captures the result, stores it on
the sandbox instance, and records it in a provider-level map so the
result survives warm-pool reclaim and reconnect.
MountUploadResult.truncated is True only when the upload pass was
stopped early by a resource limit (deadline, file count cap, or byte
budget). Individual mount failures (missing host path, SDK errors) are
logged but do NOT set truncated.
Tests cover: success totals, deadline truncation, file-count truncation,
byte-budget truncation, non-limit failure not reported as truncation,
missing host path not reported as truncation, create→sandbox wiring,
and create→release→warm-pool→acquire result preservation.
* fix(e2b-sandbox-provider): fix _mount_results lifecycle leak and review findings
- Add _forget_mount_result() helper and call it at all terminal sandbox
paths: _reuse_in_process_sandbox dead-evict, _reclaim_warm_pool_sandbox
reconnect/dead/bootstrap/ownership/shutdown failure branches,
_forget_local_sandbox, _kill_and_close. Prevents unbounded dict growth
over a long-running Gateway process.
- Make MountUploadResult @dataclass(frozen=True) to prevent silent mutation
of the shared reference between provider map and sandbox attribute.
- Move _mount_results insert under self._lock in _create_sandbox to match
the read discipline in _register_connected_sandbox.
- Guard _resolve_mount_upload_deadline against None (YAML explicit null)
to avoid int(None) TypeError.
- Add 5 regression tests covering each bypass path and the frozen invariant.
* fix(e2b-sandbox-provider): add _forget_mount_result to _evict_oldest_warm branches
Add _forget_mount_result() calls to all four terminal exit paths in the
E2B _evict_oldest_warm override (reconnect failure, already-gone, kill
failure, kill success). The peer-owned path already cleans up via
_forget_local_sandbox. Add test_evict_oldest_warm_cleans_mount_result to
pin the kill-success branch.
* docs: reduce agent guidance size
---------
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
BrowserSession scheduled three coroutines with a bare
asyncio.ensure_future(), so nothing held a reference to the resulting
tasks. The event loop only keeps weak references, so such a task can be
garbage collected before it finishes.
For the two live-frame schedulers the consequence is worse than losing
the task. Each sets a *_pending guard before scheduling and clears it in
a finally block:
self._settle_live_frames_pending = True
asyncio.ensure_future(self._settle_live_frames())
If the task is collected, the finally never runs, the guard stays True
forever, and every later _schedule_settle_live_frames()/
_schedule_input_live_frame() call returns early — silently stopping live
frame refresh for that session with no error.
Add _spawn_background(), which retains the task in a set and discards it
on completion, and route the three call sites through it. This matches
the pattern already used in task_tool, session_pool, notify and others.
Add regression tests asserting the task is retained across a gc.collect()
and released once it completes.
* fix(sandbox): stop stripping filenames when parsing find output in remote providers
The list_dir and glob parsers in the e2b, OpenSandbox, AIO, Tenki, and
BoxLite providers called .strip() on every line of find output. A
filename that legitimately ends (or begins) in whitespace was corrupted,
so the listed path never resolved on any follow-up file API call, and
the remote providers diverged from LocalSandbox, which preserves such
names via pathlib.
splitlines() already removes the line terminators, so filter empty lines
only and keep each entry verbatim. Same class of bug as the e2b
_sync_outputs_to_host fix (#4861), applied to the search parsers.
Adds a trailing-space regression test per provider at the seam each
suite already uses.
* fix(sandbox): split find output on \n only, and rename the tenki test
Review follow-ups from willem-bd:
- aio_sandbox.list_dir used str.splitlines(), which also breaks records on
\v, \f, \x1c-\x1e and \x85 - all legal inside a Linux filename, and all
contrary to this PR's own rule that the newline is the only delimiter.
find emits \n and nothing else, so split("\n") is the correct parse.
- Renamed test_search_preserves_trailing_space_in_filename to
test_list_dir_and_glob_preserve_trailing_space_in_filename, matching the
sibling tests in test_opensandbox_provider.py and test_boxlite_provider.py.
The body covers list_dir and glob; it never touches grep.
* feat(harness): deterministic acceptance checklist for subagent delegations (RFC #4651, layer 2)
PR4 of RFC #4651: check lead-supplied acceptance_criteria in code when a
subagent completes, so objectively checkable requirements can never be
silently passed by a self-report.
- subagents/acceptance_checks.py: deterministic leaf families —
file:<path> exists|non-empty and file_written:<path> read through
read_current_file_content scoped to the shared thread workspace; the
read uses the sandbox-native virtual path form (the local read
validator and provider mount tables resolve /mnt/user-data/... paths,
not host paths); the scope decision canonicalizes with realpath on the
local sandbox so workspace symlinks cannot escape into uploads; a
remote provider's "Error: ..." return string is normalized to a
failed check (provider-typed via is_local_sandbox); a
UnicodeDecodeError marks a binary deliverable as existing and
non-empty; out-of-scope paths degrade to UNVERIFIED.
tests_passed:<command> anchors to a matching recorded bash execution
with status=success and a test-summary shape; matching is
shell-structure aware with control-flow attribution (span must end at
the last segment with provable execution), negating-option values are
ineligible evidence and a target negated anywhere in the command
degrades the match, extra flags must be selection-preserving, extra
positionals widen only after a path-scoped criterion, truncated
commands degrade via command_truncated, the summary shape is read
only from output attributable to the matched segment (preceding
segments provably silent by invocation form), and pass shapes require
a nonzero passed count. Criterion text is neutralized with
neutralize_untrusted_tags before storage/rendering. Anything else
renders UNVERIFIED, never silently passed.
- executor: accumulate bounded bash command/output evidence per streamed
chunk (merged by tool_call_id, newest-capped) so subagent
summarization compacting earlier messages cannot erase a recorded
execution; the recorded status is the actual shell exit status parsed
from the output's exit marker (signed codes included; the remote
Command exited with code N form is accepted only as the whole trimmed
output), falling back to deerflow_tool_meta only when no marker
exists.
- sandbox providers: e2b/opensandbox/tenki/boxlite append the
LocalSandbox-style "Exit Code: N" marker on nonzero exit even with
non-empty output; aio propagates the SDK's structured exit_code on
both exec paths the same way; local timeouts append Exit Code: 124;
and _truncate_bash_output always preserves a trailing exit marker
(signed included) inside its budget, with a 32-char floor raising any
smaller configured limit, so the actual shell outcome always survives
in the output text.
- task_tool: run the checklist offloaded (asyncio.to_thread) on the
completed branch, failure-isolated; stamp the verdict into result
metadata and render the per-criterion section into the model-visible
result text.
- status contract: additive subagent_acceptance_verdict transport with
read-side structural validation.
- delegation ledger: entry carries the verdict and renders a compact
acceptance segment; gateway strips caller-forged verdicts from both
ledger entries and message metadata, like the citation verdict.
- blocking-IO anchor pins the offload (teeth proven red->green); leaf
read errors catch only OSError/SandboxError so unexpected errors reach
the task-tool-level isolation instead of being mislabeled.
* fix(harness): close acceptance evidence gaps from review (RFC #4651 PR4)
- negating options: overlap with a matched criterion target is now
checked by path/nodeid prefix, not exact token equality — excluding a
sub-path of the criterion's selection (pytest tests --deselect
tests/unit/test_auth.py) degrades to UNVERIFIED instead of holds
- output attribution: any redirection token in the matched final segment
makes the recorded tail non-attributable (> / >> / 2> are word
characters to the parser, so redirection was invisible to the matcher)
- silent-source allowlist narrowed from any *activate suffix to the
*/bin/activate shape
- status_contract docstring: restore the shared-fixture sentence and
note subagent_acceptance_verdict is deliberately outside the fixture
- executor: update_bash_executions publishes [] (stream carried no
bash-family calls) instead of collapsing it into None, mirroring
update_tool_receipts
* fix(harness): close acceptance residual gaps from re-review (RFC #4651 PR4)
- tests_passed: add error outcomes to the fail shapes — "4 passed, 1 error"
and pytest's "ERROR <nodeid>" short summary no longer satisfy the pass
shape when the exit status is swallowed (|| true) or absent; zero-error
counts stay clean.
- file leaves: bound the deliverable read — a "wc -c" shell size probe
answers files above 50k bytes without loading ~2x their size, honoring
the host-bash kill switch and falling back to the full read on any
non-integer rendering, so verdicts never get less sound.
- executor: record the exit marker text as status_marker on harvested bash
evidence; the leaf detail now reports the marker actually seen instead of
asserting a failure indistinguishable from the command's own trailing text.
- extend the blocking-IO anchor to drive the probe branch inside the
offload; teeth re-verified red->green.
* fix(harness): close acceptance forgery and bound gaps from P2 re-review (RFC #4651 PR4)
- file leaves: never read unbounded — size is established first (os.stat on
the validated local host path, so the host-bash-disabled configuration
needs no shell; a guarded wc -c on remote providers that renders
missing/unreadable in its own words). Above the 50k cap the leaf answers
from the size alone, at/below it the full read runs, and an
unestablishable size degrades to UNVERIFIED instead of an unlimited
fallback read.
- output attribution: source/. prefixes are never provably silent — a
crafted */bin/activate path shape says nothing about what the script
prints, so sourced segments can no longer lend a passing summary.
- executable identity: an explicitly path-spelled criterion now requires
the same normalized executable path; the basename rule stays only for
deliberately bare criterion commands.
* fix(harness): run acceptance size probe outside subagent-controlled state (RFC #4651 PR4)
- remote probe no longer runs in the sandbox's persistent shell: a fresh
env -i /bin/sh with absolute-path stat/realpath (poisoned functions,
aliases, PATH, exported functions, IFS, locale cannot steer it), plus a
marker env routing AIO onto a fresh per-call bash.exec session.
- metadata-only: stat never opens content, so a FIFO deliverable cannot
block the parent for the provider's idle timeout; non-regular files
(fifo/dir/symlink) degrade to UNVERIFIED.
- containment canonicalized against the literal mount root: a
final-component symlink or a swapped parent directory (root included)
cannot redirect the check outside shared storage; unprovable layouts
degrade to UNVERIFIED.
* fix(harness): canonicalize probe containment against the canonical mount root (RFC #4651 PR4)
Literal-root equality made every remote file leaf permanently UNVERIFIED
on e2b and Tenki, which realize /mnt/user-data as a symlink to the home
dir by default (e2b bootstrap 'sudo ln -sfn', Tenki best-effort symlink).
Containment now compares the file's realpath against the mount root's
realpath — exactly what the provider's own read path resolves, so probe
and read-back stay consistent; final-component symlinks stay rejected by
the non-dereferencing stat, and an intermediate dir-link escape under a
sane root still lands ESCAPED. The inner script is a module constant and
the suite now executes the composed probe for real against on-disk
layouts (real dir, symlinked prefix, final symlink, fifo, missing,
dir-link escape), which the canned-output stub could not see.
* fix(harness): close bare-criterion negation and CDPATH summary channels (RFC #4651 PR4)
- matching: a criterion with no positional selection target (bare pytest,
make test) stands for the runner's default selection, so ANY negating
option (--ignore/--deselect/...) makes the recorded run a different
selection — unprovable. The overlap guard only sees consumed criterion
tokens, which a bare criterion does not have; scoped criteria keep the
unrelated-exclusion behavior.
- attribution: cd is no longer blanket-silent — CDPATH makes cd print the
resolved (subagent-chosen) destination and the pass shapes match as
substrings, so one mkdir 'all tests passed' plus an export minted a pass
for any quiet command. A cd argument or CDPATH= value (export or leading
assignment) carrying any summary shape makes the segment non-silent;
shape-free cd dir wrappers keep matching.
- docs: _truncate_bash_output states the effective 32-char floor (the
guarantee previously read as an unconditional max_chars bound).
* fix(harness): close env-assignment and expansion channels in acceptance matching (RFC #4651 PR4)
Self-audit in the shape of the last review rounds — channels the matcher
classified as accounted-for that can change what runs, narrow the
selection, or lend the summary text:
- env assignments are no longer blanket-stripped: only an allowlist of
inert display/CI knobs (CI, NO_COLOR, PY_COLORS, ...) may prefix a
matched span, and a non-allowlisted assignment in any preceding segment
(pure-assignment or export NAME=) is state pollution — PATH redirects
the executable, LD_PRELOAD/PYTHONPATH/NODE_OPTIONS inject code,
PYTEST_ADDOPTS/GOFLAGS/MAKEFILES inject selection-changing inputs,
BASH_ENV runs arbitrary shell startup. All degrade to unprovable.
- runtime expansions: any span token carrying /$( )/backticks, any
negating-option value carrying an expansion or glob (unknown excluded
set), and any extra executed token carrying glob metacharacters
(crafted option-looking filenames narrow invisibly) are unprovable.
Criterion-side globs stay self-consistent (literal match).
- cd: an argument carrying a runtime expansion or glob is non-silent
(unknown destination, unknown print); CDPATH= assignments are now
handled as state pollution at the match layer, subsuming the
value-shape special case.
* fix(harness): persistent-shell evidence, exact env sets, option-arity scoping (RFC #4651 PR4)
- tests_passed: on a persistent-shell provider (new
Sandbox.persistent_shell_sessions capability, set by AioSandbox) every
leaf degrades to UNVERIFIED — any earlier call in the shared session
could have mutated the state the clean-looking run executed in, and
only a fresh controlled session (RFC section 6 verifier) can prove
otherwise. The flag is read from the provider registry without
acquiring a sandbox.
- env assignments: the allowlist is gone — no variable is provably inert
across repositories (CI/DEBUG are routinely read by tests). The span's
assignment prefix must equal the criterion's exactly (values included,
order-insensitive); any assignment or export NAME= in a preceding
segment is state pollution.
- scoping: positional targets are now read by option arity, so a path
embedded in an option (--basetemp=/tmp/p, --junitxml=/tmp/r.xml) never
counts as a selection target and an extra positional after such a
criterion narrows the default selection it denotes.
* fix(harness): stamp shell provenance at harvest, close export/unset and arity gaps (RFC #4651 PR4)
* fix(harness): split physical newlines as shell separators in acceptance matching (RFC #4651 PR4)
* fix(harness): scope cd wrappers to thread data roots, pin accepted boundaries (RFC #4651 PR4)
* fix(harness): preserve criterion connectors, prove file_written readable, fail-closed shell capability (RFC #4651 PR4)
* fix(harness): compare only the connector prefix, tolerate trailing criterion semicolons (RFC #4651 PR4)
* fix(harness): preserve continuation-line operators, keep ./-spelled executable identity (RFC #4651 PR4)
* fix(harness): render criteria single-line so a multiline criterion cannot inject a forged checklist line (RFC #4651 PR4)
* fix(harness): reject parent-traversal executable tokens in acceptance matching (RFC #4651 PR4)
* fix(harness): reject parent-traversal negated values in acceptance matching (RFC #4651 PR4)
* feat(e2b-sandbox): make mount upload deadline configurable
Replace the hardcoded 120-second mount upload deadline with a
configurable `mount_upload_deadline_seconds` key read from
SandboxConfig (extra=allow). The value is validated: zero and
negative inputs are clamped to 1 second. Omitting the key
preserves the existing 120-second default.
This addresses the follow-up from PR #4842 review: operators
with large mounts or slow networks can now size the deadline to
their deployment without changing code.
* fix(e2b-sandbox): address review feedback on configurable deadline
- Remove import-time default capture from _mount_deadline_reason()
and _MountUploadBudget.deadline_seconds to prevent silent drift.
- Add warning log when mount_upload_deadline_seconds is clamped to 1
(was silent before).
- Update AGENTS.md E2B Mount Uploads section: deadline is now
configurable, not fixed 120.
- Add mount_upload_deadline_seconds to YAML examples in provider
docstring and __init__.py.
- Add config-path test that exercises SandboxConfig -> _load_config ->
_apply_mounts end-to-end.
* fix(e2b-sandbox-provider): handle non-numeric mount_upload_deadline_seconds
Guard _resolve_mount_upload_deadline against None, non-numeric strings,
and other invalid values. None returns the default; non-numeric strings
like '120s' or 'abc' log a warning and fall back to the 120-second
default instead of crashing provider init with TypeError/ValueError.
Extend the parametrized clamp test with None, suffix, and alpha cases,
and add a warning assertion. Update CONFIGURATION.md with the new
mount_upload_deadline_seconds key and its behavior.
* fix(sandbox): handle infinite mount deadline
* fix(deps): depend on renamed tenki package instead of tenki-sandbox
tenki-sandbox has been removed from PyPI and republished as tenki. Its old wheel URL still resolves, so existing lockfiles keep installing and the breakage is invisible to anyone with a warm lock; any fresh resolution fails with 'tenki-sandbox was not found in the package registry'.
tenki 1.0.2 still ships the tenki_sandbox module, so the imports in community/tenki/provider.py and sandbox.py are unchanged.
Fixes#5081
* fix(tenki): point install guidance at the renamed distribution
The rename to `tenki` left the user-facing remediation still naming the
removed package. `_import_client` raised "pip install tenki-sandbox" on the
missing-extra path — the exact instruction this change proves now 404s on
PyPI, handed to the user at the exact moment they need it to work.
Update that message and the remaining `tenki-sandbox` references in the
provider, sandbox adapter, README, sandbox AGENTS.md and the test docstring.
The imported module stays `tenki_sandbox`, so the distribution and module
names now differ; each mention says so rather than just swapping the string.
No behavior change beyond the error text.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(tenki): migrate the provider to the 1.x workspace-only API
Renaming the dependency was not enough. tenki 1.0.2 keeps the tenki_sandbox
module name but not its contract: Client.create dropped project_id and has no
**kwargs to absorb it, and IdentityWorkspace no longer carries `projects`
(the attribute is gone from the package entirely). Both configuration paths
therefore failed before a sandbox could be created — explicit project scope
raised TypeError, and automatic scope raised AttributeError walking
workspace.projects.
Scope is now the workspace alone. _resolve_scope returns a single workspace id,
auto-selecting when the account has exactly one, and project_id is gone from
create_kwargs and from the documented config surface.
A stale project_id in config.yaml warns rather than fails. SandboxConfig is
extra="allow", so simply not reading the key would leave it scoping nothing
with no signal; it also used to short-circuit the identity lookup, so operators
with more than one workspace need to know they must now set workspace_id.
The suite passed against the broken provider because the fake client took
**kwargs and swallowed the project_id the real SDK rejects. The double now
mirrors 1.0.2 — keyword-only, no **kwargs — so an unexpected argument is a
TypeError in tests exactly as it is against the SDK. Reintroducing the old
create call fails 20 tests; before this change it failed none.
Verified against the exact locked wheels: every other kwarg the provider
passes (name, workspace_id, sticky, wait, max_duration, image, cpu_cores,
memory_mb, env) and every SDK surface it touches (who_am_i, Identity.workspaces,
wait_ready, exec, close, the fs API, the four terminal exception classes) is
unchanged in 1.0.2.
Reported by willem-bd in review.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* docs(config): drop sandbox.project_id from the Tenki example
The canonical example still documented project_id as a supported optional key
after the provider stopped honouring it, so an operator following it could set
the key, get no scope from it, and hit a workspace-resolution failure with
nothing in the example to explain why.
Replaced with a migration note rather than a silent deletion: someone upgrading
already has the key in their config.yaml and needs to know it is inert now and
that workspace_id is what scopes a sandbox on Tenki 1.x.
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Aniket Wagh <aniketwaghh@users.noreply.github.com>
* feat(sandbox): share sandbox identity derivation and acquire serialization (#4741)
Remote providers (AIO, E2B, BoxLite, Tenki, OpenSandbox) each inlined the
same sha256(user:thread)[:16] sandbox-id expression and kept per-scope lock
dicts that grew unboundedly until shutdown. This extracts both mechanisms
into shared components without changing provider lifecycle, ids, capacity
semantics, or public tool behavior:
- sandbox/identity.py: keyword-only derive_sandbox_scope_token (byte-pinned
compatibility contract) + is_sandbox_scope_token; per-provider golden
vectors pin current behavior including BoxLite's raw-None quirk and each
provider's private user_id resolution.
- sandbox/acquire_serialization.py: AcquireSerializer — per-key lock table
with holder/waiter refcount reclamation, bounded dedicated executor
(async waits off both the event loop and the default executor),
worker-owned cancellation cleanup (no event-loop callback dependency), idempotent close().
- Each provider adopts both components; AIO/E2B key by (user_id, thread_id)
with acquire and (E2B) release serialized; BoxLite/Tenki/OpenSandbox key
by derived sandbox id and offload the whole sync acquire to the
serializer's executor so a cancelled awaiter cannot overlap a retried
same-scope body (leaked-remote-VM regression caught in review).
- thread_id=None acquires stay unserialized; provider shutdown()/reset()
close the serializer; E2B capacity/ledger/reconciliation and AIO
ownership/flock machinery untouched.
- blocking-IO anchor proves contended OpenSandbox acquire_async stays off
the event loop (teeth verified red/green); AGENTS.md documents the
shared components.
* refactor(sandbox): address review on acquire serialization (#5089)
- Replace unreachable checkin branch with an assertion: run() returns
False only after abandon(), which the except handler always re-raises;
the old _checkin would have double-decremented the refcount.
- Document the task.cancelling() == 0 assumption in hold_async.
- Drop unused thread_id/user_id kwargs from BoxLite and Tenki
_acquire_scope_locked (OpenSandbox still forwards them).
* fix(sandbox): preserve request ContextVars in acquire executor bridge (#5089)
loop.run_in_executor() does not copy contextvars, unlike the inherited
SandboxProvider.acquire_async() which used asyncio.to_thread(). The
BoxLite/OpenSandbox/Tenki acquire_async bridges introduced in this PR
therefore dropped the request trace id (logged as trace_id=-).
Add AcquireSerializer.run_on_executor(), which copies the calling
context and runs the callable through ctx.run, and route all three
providers through it. Add regression tests binding request_trace_context
and verifying the worker thread observes it.
Add deerflow.community.serply.tools:web_search_tool, a Google SERP
provider for the web_search slot that also covers Google News and Google
Scholar through an optional `vertical` config option. Reads the key from
api_key in config.yaml or SERPLY_API_KEY, clamps max_results to Serply's
1-100 range, and returns the same structured JSON errors as the Serper
and Brave tools.
Register the provider in config.example.yaml, scripts/doctor.py,
scripts/wizard/providers.py, .env.example, backend/docs/CONFIGURATION.md,
the en/zh tools.mdx provider tabs, and tools/AGENTS.md. Tests mock httpx.
* fix(sandbox): harden local Docker sandbox containers and port binding
Root causes (security audit SBX-1/SBX-2) in the local container backend:
- _resolve_docker_bind_host published sandbox ports on 0.0.0.0 whenever
DEER_FLOW_SANDBOX_HOST was non-loopback (docker-compose defaults to
host.docker.internal), exposing the unauthenticated /v1/shell/* exec
API on every host interface.
- _start_container ran every sandbox with seccomp=unconfined and no
capability, privilege-escalation, or resource limits, so untrusted
model-authored code could exhaust the host, escalate privileges, and
reach internal networks / cloud metadata endpoints directly.
Hardening changes and defaults:
- Port binding: non-loopback sandbox hosts now bind the Docker default
bridge gateway instead of 0.0.0.0, discovered dynamically via
`docker network inspect bridge` with a static 172.17.0.1 fallback.
host.docker.internal resolves to that gateway through host-gateway,
so DooD gateways and the Docker host still reach the sandbox while
external interfaces no longer see the port.
DEER_FLOW_SANDBOX_BIND_HOST=0.0.0.0 restores the legacy broad bind.
- seccomp=unconfined is no longer unconditional: sandboxes run with
Docker's default seccomp profile; opt back in with
DEER_FLOW_SANDBOX_SECCOMP_UNCONFINED=1, only when the sandbox image
is verified to require syscalls the default profile blocks.
- Add --cap-drop=ALL and --security-opt no-new-privileges (Docker only;
the Apple Container CLI does not support these flags).
- Bounded resources with env overrides: --memory 2g
(DEER_FLOW_SANDBOX_MEMORY), --cpus 2 (DEER_FLOW_SANDBOX_CPUS),
--pids-limit 512 (DEER_FLOW_SANDBOX_PIDS_LIMIT); each also accepts
"0"/"none" to disable the limit.
- No --user is forced by default (the default AIO sandbox image's user
is upstream-controlled and unverified), but
DEER_FLOW_SANDBOX_CONTAINER_USER passes one through for deployments
that know their image.
- DEER_FLOW_SANDBOX_NETWORK passes --network so sandboxes can be
attached to a dedicated egress-controlled network; default networking
is unchanged.
backend/docs/CONFIGURATION.md documents the new bind behavior and every
override; tests cover each default and escape hatch.
* fix(sandbox): follow host-gateway mapping for binds; keep image-required seccomp default
Review follow-ups on the hardening change:
- Bind: resolve the sandbox host itself and bind that address, instead of
assuming the default bridge IPv4. host.docker.internal follows the
daemon host-gateway-ip mapping (customizable, possibly IPv6), so the
resolved address is exactly where the gateway connects — the published
port and advertised URL always match. IPv6 is bracketed for docker -p,
zone ids stripped, wildcard resolutions ignored; unresolved hosts fall
back to the bridge gateway with a warning pointing at
DEER_FLOW_SANDBOX_BIND_HOST.
- seccomp: the shipped AIO image needs seccomp=unconfined for its
Chromium browser (upstream quick-start always passes it; the upstream
FAQ documents the browser failing under Docker default profile), so
that option returns as the default. Tightening stays possible via
DEER_FLOW_SANDBOX_SECCOMP_PROFILE=<path to a restricted,
Chromium-compatible profile> or DEER_FLOW_SANDBOX_SECCOMP_UNCONFINED=0
for images verified to work with Docker's default profile.
- cap-drop/no-new-privileges and the resource limits are unchanged.
- Tests updated for both behaviors; 37 pass.
* fix(sandbox): bracket bare IPv6 bind overrides; state seccomp default accurately
DEER_FLOW_SANDBOX_BIND_HOST was returned verbatim, so a bare IPv6 literal
like fd00::1 produced an invalid publish spec (fd00::1:port:8080); Docker
requires the bracketed form. Normalize raw and already-bracketed IPv6
literals (IPv4/hostnames untouched), with resolver-level and argv-level
tests covering the explicit IPv6 override.
The CONFIGURATION.md overview claimed Docker's default seccomp profile
stays active, contradicting the seccomp=unconfined default the table (and
the code) actually ship for the Chromium-based image; spell out the relaxed
default and where to change it.
* style(sandbox): apply ruff format to local_backend
* fix(sandbox): reject host networking, force builtin seccomp opt-out, resolve hostname binds
Review follow-up on #4986 (willem-bd):
- P1: DEER_FLOW_SANDBOX_NETWORK=host (and container:<name>) now raise a
RuntimeError at start instead of silently voiding the hardened port
bind — Docker discards -p/--publish in host mode and shares the
network namespace for container:<name>, which would re-expose the
unauthenticated exec API on the host's interfaces. Two regression
tests cover both rejections.
- P2: the seccomp opt-out now passes seccomp=builtin explicitly instead
of omitting the option, so a daemon configured with an unconfined or
custom default cannot weaken the documented opt-out; the test asserts
the flag.
- P2: hostname values in DEER_FLOW_SANDBOX_BIND_HOST resolve to an
address before use (Docker publish specs require an IP literal as the
host part, so host.docker.internal previously produced an invalid
spec that prevented every sandbox from starting); unresolvable names
raise a clear configuration error. Tests cover resolution and
rejection; CONFIGURATION.md updated for all three behaviors.
43/43 pass in tests/test_aio_sandbox_local_backend.py; ruff check +
format clean.
* fix(sandbox): reject DEER_FLOW_SANDBOX_NETWORK=none (loopback-only, breaks published API port)
* fix(sandbox): validate the effective Docker network target; normalize IPv6 sandbox hosts once
name=host / name=none dodge raw-string checks but attach like the bare
words; strip name= prefixes and validate the effective target (network IDs
keep passing). Bracketed IPv6 sandbox hosts now resolve for the bind and
bare IPv6 hosts produce bracketed URL authorities — both input forms give
identical bind and URL addresses.
* fix(sandbox): parse the full Docker network long syntax before validating
Docker accepts comma-separated key=value fields in any order (name=, gw-priority=,
alias=, ...); a name=host field hides the host network behind surrounding fields.
Parse the CSV and validate the parsed name= target (last occurrence wins, fields
lowercased, mirroring opts/network.go); no-name values fall through like Docker's
own rejection.
* fix(sandbox): keep CHOWN/SETUID/SETGID through cap-drop=ALL for the default image
The shipped image's entrypoint starts as root, creates the gem user,
chowns /opt/jupyter and drops to that user via su; without those three
capabilities the set -e script dies before the readiness endpoint exists.
no-new-privileges stays (it blocks gaining privileges via exec, not using
the added caps). Adds a docker-gated real-image startup smoke test.
* fix(sandbox): let pre-initialized non-root images drop the startup capabilities
The CHOWN/SETUID/SETGID re-add only exists for the shipped image's root
entrypoint handoff. A custom image that never runs as root gets an explicit
opt-out (DEER_FLOW_SANDBOX_IMAGE_STARTUP_CAPS=0) so those capabilities are
not left available to sandboxed code (chown on bind mounts, UID/GID
impersonation).
* test(sandbox): gate the real-image smoke test behind the live marker
The default offline suite (make test = -m 'not live') must not depend on a
third-party registry: mark the smoke test live, probe the daemon inside the
test body (never at collection time), and allow pinning the image reference
via DEER_FLOW_SANDBOX_SMOKE_IMAGE for a dedicated integration job.
* test/docs: isolate DEER_FLOW_SANDBOX_IMAGE_STARTUP_CAPS in tests; add table row; split custom-image guidance
_clear_hardening_env now clears the new knob so a developer shell or .env
preset cannot flip the default-path tests. CONFIGURATION.md gains the table
row, and the custom-image guidance becomes its own paragraph with the
no-new-privileges scope stated correctly (it does not mitigate the retained
CAP_SETUID/SETGID risk).
* test(sandbox): make the live smoke test diagnosable
300s readiness budget (cold pull + cold start must not be conflated with
broken capabilities) and dump the container's last 40 log lines on failure
so the next live run tells us whether the capability set is incomplete
(chown/useradd/su errors) or the services are merely slow.
* test(ci): align the smoke test with the 60s provider deadline; add a dedicated live smoke workflow
Single-source the readiness deadline as SANDBOX_LOCAL_PROVIDER_READY_TIMEOUT
(used by both provider paths and the smoke test) so the validation cannot
drift from the production contract again. New sandbox-image-smoke.yml runs
the live test on a dedicated job, with the image reference pinnable via the
SANDBOX_SMOKE_IMAGE repository variable (digest resolved and recorded in the
job summary when falling back to :latest).
* test(sandbox): pull the failing program's own logs on smoke failure
supervisord only surfaces exit codes in docker logs; nginx's stderr lands in
files inside the container. Dump supervisor program logs, nginx -t, and the
nginx error log on failure so the next run names the exact broken line.
* ci(sandbox): export an immutable repo@digest reference for the smoke run
docker pull once on the runner platform, resolve RepoDigests[0], and pass
that immutable reference to the test via GITHUB_ENV — the recorded and
executed images can no longer diverge when the tag moves, and platform
selection is left to the daemon instead of jq over the manifest index.
* fix(sandbox): add DAC_OVERRIDE — the root nginx master writes gem-owned logs
The image's root nginx master opens /var/log/nginx/{access,error}.log,
which belong to the gem user, for the container's lifetime; without
CAP_DAC_OVERRIDE it dies with 'open() failed (13: Permission denied)' on
every start (FATAL under supervisord) and readiness never arrives. Four
capabilities now: CHOWN/SETUID/SETGID for the entrypoint handoff plus this
runtime log-write need.
* fix(e2b): preserve trailing whitespace in filenames and survive mtime overflow
_sync_outputs_to_host iterated the NUL-delimited find output with
entry.strip() on each record. NUL already guarantees record boundaries,
so the strip is redundant and harmful: a filename that legitimately ends
in whitespace (e.g. "report ") had its trailing space trimmed, pointing
host_path at the wrong file and recording a manifest key that can never
match — the file was re-downloaded on every release.
The same host-write block wrapped only os.utime in the outer
except OSError, but os.utime raises OverflowError (not an OSError) when
the ns value is out of range (a far-future remote mtime, e.g.
`touch -d '99999 years'`). That escaped the loop, skipping the manifest
write and forcing a full re-download next release. Wrap os.utime in its
own (OSError, OverflowError) so only the timestamp restoration is
dropped; the file is still written and the manifest still updated.
* test(e2b): rely on monkeypatch cleanup
---------
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
* fix(sandbox): bound aggregate E2B mount upload work
* fix(sandbox): preserve mount guards on upload failure
* fix(sandbox): cover mount preflight with deadline
* refactor(sandbox): clarify mount deadline checks
* refactor(sanbox): deduplicate mount deadline reason
* fix(sandbox): evaluate mount deadline reason lazily
Admission derived the live-entry count as `HLEN - 3`, where 3 was the number
of `meta:*` fields written 35 lines earlier in initialize(). Nothing tied the
two together, so adding a fourth meta field would shift the capacity ceiling
by one.
Name the offset `META_FIELD_COUNT` next to initialize(), and add a guard test
asserting a freshly initialized ledger holds exactly those three fields, plus
one pinning that a hard_limit of N admits exactly N reservations.
References #4575
Co-authored-by: icn5381 <255778606+icn5381@users.noreply.github.com>
* fix(sandbox): project enabled skills into sandbox views
* fix(skills): keep projection mutations consistent
* fix(skills): fail closed on projection errors
* fix(skills): isolate per-scope failures during boot projection rebuild
rebuild_all_skill_projections() propagated any exception from the public
rebuild or from a single user's rebuild straight out of the gateway
lifespan startup, uncaught. A single broken user directory (bad
permissions, corrupted _skill_states.json, unreadable content) would
therefore abort gateway boot for every user, not just that one -
_rebuild_*_locked already fails closed internally (clears the view and
re-raises), so the boot loop only needed to stop treating that re-raise
as fatal.
Each scope's rebuild now fails closed independently and boot continues;
a scope left empty by a boot failure self-heals on the next sandbox
acquire via ensure_skill_projections().
Also patches deerflow.skills.projection.rebuild_all_skill_projections in
the memory-flush lifespan test fixture, matching the two sibling
fixtures in the same file — this call is now on the lifespan startup
path and the fixture's minimal SimpleNamespace config predates it.
* test(skills): update authz test for the projection-aware public toggle
_persist_shared_skill_state (introduced earlier in this branch) reads
the shared extensions_config.json fresh from disk under the projection
lock instead of through the cached get_extensions_config() singleton -
that's the whole point of the fix (stale worker caches must not clobber
another worker's concurrent update). The name no longer exists on the
skills router module, so the test's monkeypatch of it started raising
AttributeError instead of exercising the endpoint.
The mock storage in this test isn't a real LocalSkillStorage instance,
so _persist_shared_skill_state's projection-mutation branch is already
skipped (nullcontext) and it falls back to a fresh ExtensionsConfig()
for the nonexistent tmp config_path - no replacement monkeypatch needed.
* fix(sandbox): make skill projection ensure best-effort in acquire
acquire() called _ensure_skills_projection() directly, outside any
try/except, in both LocalSandboxProvider and AioSandboxProvider. Every
other skill-mount setup path in these providers has always caught
exceptions and logged a warning rather than failing sandbox acquire
outright (e.g. when config.yaml can't be resolved) - these two new call
sites broke that contract, so any projection failure (including simply
not having a config.yaml, as in CI's test environment) now failed
acquire() itself instead of just leaving skill mounts off.
_ensure_skills_projection now catches its own exceptions and returns
None; both providers' callers already tolerate that (a None projection
skips the skill-specific mounts, matching the existing degrade path)
after making _append_public_skill_mapping and the custom/legacy mount
block in LocalSandboxProvider explicitly None-safe.
Caught by running the full suite with config.yaml removed, matching
CI's environment - not caught locally because a real config.yaml was
present, masking the failure.
* fix(sandbox): make E2B skill projection mounts best-effort
_skill_projection_mounts called ensure_skill_projections with no guard,
unlike Local/AIO's _ensure_skills_projection. A raise propagated out of
_apply_mounts before the configured-mounts loop ran, so a skills
projection failure dropped the operator's own configured mounts too -
only caught by create()'s outer warning, with nothing applied at all.
Swallow here and return an empty mount list on failure, matching the
Local/AIO pattern: still fail-closed for skills, but no longer widens
the blast radius to unrelated configured mounts.
Review feedback from PR #4178.
* docs(skills): document projection trade-offs flagged in review
- _update_tree_digest: note the metadata-only (not content) hashing
trade-off and why runtime writes through this codebase are still
covered regardless (rebuild-under-lock + rename always changes inode).
- LocalSandboxProvider.acquire: note the acquire-time self-heal cost
(cheap on a fresh manifest, ~400ms rebuild under lock on stale/drift).
- skill_projection_mutation: drop the no-op except-Exception-then-raise;
a raise from the mutation already propagates past the yield with the
view left cleared, no explicit re-raise needed.
- provisioner README: spell out that hostPath skills volumes require
the gateway and K8s node to share DEER_FLOW_HOST_BASE_DIR (single-node
or shared storage), and that the custom/legacy volumes' hostPath type
Directory (not DirectoryOrCreate) makes a violation of that assumption
a visible Pod-creation failure instead of a silent empty mount.
Review feedback from PR #4178.
* fix(skills): lazily repair user projections
* fix(skills): close projection review gaps
* fix(skills): refresh user projection enable state
* fix(skills): close projection review follow-ups
* fix(skills): preserve state across projection writes
---------
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>