* fix(sandbox): stop E2B reconciliation from reviving warm-pool sandboxes
Periodic reconciliation probed every discovered remote sandbox with
Sandbox.connect() before the locality check, and the check itself only
consulted _sandboxes, not _warm_pool. A sandbox parked by release() was
therefore adopted back to active on the first pass, and because the SDK
normalizes connect(timeout=None) to its 300s default and the control
plane extends a running sandbox's expiry when now+timeout is later,
each 60s pass kept pushing the expiry forward — idle warm sandboxes
never hit their configured idle_timeout.
Treat _sandboxes and _warm_pool ids as locally tracked up front: skip
probing them (no timeout-mutating connect), keep them canonical, and
route only genuinely remote candidates through the duplicate-reap path.
Extend the post-probe adoption recheck to _warm_pool so a release that
lands mid-probe cannot be promoted back to active either.
Fixes#5550
* fix(sandbox): keep active E2B VMs alive and sweep expired warm entries
Address review on #5562:
- Reconciliation now refreshes the remote TTL of locally active
sandboxes through their cached client (never connect()), restoring
the keepalive for turns that outlive idle_timeout without reviving
warm-pool VMs.
- Warm-pool entries parked longer than idle_timeout are dropped during
reconciliation — their VMs are expected to be reaped by the control
plane — releasing the ownership lease and the capacity slot they
would otherwise pin until reclaim, eviction, or shutdown.
- Remove the now-dead thread-local canonical sort; locally tracked ids
are skipped unconditionally, so the ordering hint had no effect.
* fix(sandbox): preserve active E2B keepalive and shared capacity
* fix(sandbox): serialize E2B reconciliation lifecycle transitions
* fix(sandbox): fence E2B ownership and timeout lifecycle writes
* fix(sandbox): isolate ownership heartbeats from E2B timeout IO
---------
Co-authored-by: Totoro-qaq <279883115+Totoro-qaq@users.noreply.github.com>
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
* feat(mcp): re-scope to MCP task claim lifecycle only
Keep PR #4966 a small, closed MCP lease/cancellation state-machine change and
move RunJournal and Run lifecycle work into dedicated follow-ups. This branch
contains only the MCP task claim lifecycle:
- mcp task release/snapshot fencing by owner + per-claim lease token
- phase-level single-flight poll/cancel/notification owners with retained handoff
- routine cancellation no longer persisted as a task failure diagnostic
- bounded ordinary release ownership retention past the drain deadline
- 0018_mcp_task_lease_tokens migration + migration/bootstrap head assertions
- wait_for_task_until helper (MCP uses it); worker-specific capture helper moved
to the run-finalization follow-up
RunJournal (journal.py + test_run_journal.py) and run lifecycle
(manager/worker/store/run sql + run tests) are preserved on
backup/cancellation-safety-full and will be raised as separate follow-ups.
* fix(mcp): unblock claims after ambiguous handoff resolves
A phase-level single-flight owner only guards an ambiguous claim outcome. Once
the claim resolves, the phase owner is released immediately; the handoff may
continue releasing returned rows as bounded, service-owned background work
(transferred to _compensation_tasks on timeout). Per-claim token fencing rejects
a late release against a newer claim generation, so a stuck release no longer
locks the whole phase until process restart.
- README: drop the stale progress-snapshot sentence from the bounded ordinary
release description.
- service: pop the identity-checked phase owner as soon as the claim outcome is
known, then release returned rows with the bounded path; carry the release in
_compensation_tasks if it exceeds the drain deadline.
- mcp/AGENTS.md: document that only an unresolved claim outcome (not the handoff)
blocks later phase scans, and that returned-row releases may continue in the
background once the owner is released.
- tests: pin that the phase owner is released before a stuck release finishes
while the release stays service strong-owned.
* refactor(mcp): remove unused single-record claim wrappers
_poll_one, _cancel_one, and _notify_one are unreachable in production: the
worker always processes claimed records through _run_claimed_batch, so these
wrappers preserved a second, dead single-record lifecycle (state is None)
whose only observable behavior was a wrapper-specific cancellation release.
Remove the three wrappers and migrate the regressions that guarded their
cancel/release invariants to exercise the production _run_claimed_batch path
(operation=_*_one_claimed, release=_release_*_after_cancellation). The single
wrapper-only "state is None" contract (test_poll_release_hang_without_batch)
is deleted; all 11 remaining invariants (CancelledError preservation, repeated
cancellation, poll-only token-fenced lease release, notification claimed vs
dispatched phase release, hung compensation -> service ownership, and
background compensation exactly-once observation) are now covered through the
real batch lifecycle.
* fix(mcp): fence claim-owned mutations against stale generations
The per-claim token check in the ORM release/apply paths was only in the
SELECT; the final write went out by primary key. On SQLite (where
with_for_update() is a no-op) a mutation from an older claim generation
could therefore clear a claim that a newer generation had reclaimed after lease
expiry — the exact distributed lease-fencing failure the per-claim token was
meant to prevent.
Make every claim-owned mutation a single atomic conditional UPDATE with the
owner and per-claim token in the WHERE clause (rowcount 0 => stale, return
False, no mutation):
- release_claim: atomic fence; record the poll-failure event after the fence
wins (same transaction, holding the write lock).
- apply_snapshot / apply_cancel_snapshot: atomic fence; record the event after.
- finish_notification_run: atomic fence; use a CASE on event_version >>
dispatch_version to keep a newer event pending for redelivery instead of
swallowing it as delivered.
Add one regression per path: a stale generation's release/apply/finish after a
same-worker reclaim is rejected and never clears the newer claim.
* test(mcp): pin the migration chain head to the lease-token revision
0026_mcp_task_lease_tokens becomes the alembic head, so the chain-head pin in the 0025 repair test had to move on. Follow the 0023 precedent there (single head plus expected predecessor) instead of pinning a literal head, and give the new revision its own migration test, which owns the pin and covers the nullable claim-token columns on upgrade and their removal on downgrade.
* refactor(mcp): close cancellation cleanup leftovers
* fix(mcp): retain cancelled release diagnostics
* test(mcp): remove obsolete settled compensation case
* test(mcp): cover interleaved lease reclaim races
---------
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
* feat(knowledge): add verifiable RAGFlow source citations
* docs(knowledge): scope RAGFlow guidance to its own directory
* fix(knowledge): preserve citations through rendering and budgets
* feat(settings): persist account preferences across browsers
* docs(settings): scope preference guidance to user persistence
* fix(settings): preserve SSR and fence custom-agent defaults
* test: include user persistence in scoped guidance inventory
* fix(settings): sync explicit edits and preserve local tab updates
* fix(tools): run tool assembly off-loop at async entry points
get_available_tools() may block on MCP cache initialization while it is
called on async agent-assembly paths (task_tool, durable batch execution),
stalling the calling event loop for the full discovery duration.
Dispatch the (unchanged, synchronous) assembly call to a worker thread via
asyncio.to_thread at the two async entry points so the loop keeps processing
requests, SSE frames, cancellations, and timers.
Fixes#5172
* fix(tools): offload lead-agent assembly off-loop and pin with blocking-io anchors
Review follow-up for #5224:
- run_agent now dispatches agent_factory(...) through asyncio.to_thread, so
lead-agent assembly (including both get_available_tools call sites in
_assemble_lead_agent) runs off the event loop — the Gateway headline
scenario from issue #5172.
- _ensure_sync_invocable_tool takes a double-checked threading.Lock, making
the in-place tool.func wrap on the shared tool singletons explicitly
single-shot now that assembly can run concurrently on worker threads.
- Add backend/tests/blocking_io/test_tool_assembly_offloop.py: blocking-probe
anchors for task_tool and SubagentBatchService._execute_item under the
strict Blockbuster gate, plus a meta-check proving the gate trips on the
exact syscall class (ExtensionsConfig.from_file on the loop). Verified the
anchor goes red when the offload is flattened back to a plain call.
* fix(gateway): build checkpoint state accessor off-loop; anchor run_agent offload
Review follow-up for #5224:
- Add abuild_checkpoint_state_accessor (asyncio.to_thread around the
unchanged sync builder) and switch every async call site to it: the
stateless_wait route, thread_runs, both threads call sites, and the
build_thread_checkpoint_state_accessor boundary. The agent-factory
assembly re-enters get_available_tools() and may block on MCP cache
initialization; repeat calls hit _state_accessor_graph_cache and only
pay the thread hop.
- Add a third blocking-io anchor driving the real run_agent with minimal
RunManager/bridge stubs; the factory performs a real production blocking
read (ExtensionsConfig.from_file()) and the test asserts assembly never
runs on the main thread. Verified the anchor goes red when the run_agent
offload is flattened back to a plain call.
- Adapt the test_threads_router checkpoint-builder patch sites to the new
async name.
* refactor(tools): carry assembly offloads on a dedicated bounded pool
Review follow-up for #5224:
- Add utils/assembly_io.py: a dedicated ThreadPoolExecutor (default 8
workers, DEER_FLOW_ASSEMBLY_WORKERS-overridable, mirroring
utils/file_io.py and tools/sync.py) with run_assembly(), which copies
contextvars explicitly. A hung stdio MCP server parks its worker for
the full MCP timeout; carrying assembly hops on the loop's default
executor would let a few parked assemblies queue every other
to_thread/run_in_executor(None, ...) caller behind them.
- Switch all four offloads (run_agent, task_tool, batch _execute_item,
abuild_checkpoint_state_accessor) to run_assembly().
- State the cold-path behavior in the accessor docstring: the graph
cache validates factory identity, so non-identity-stable factories may
duplicate lead-agent assembly across concurrent readers (MCP discovery
stays process-wide single-flight); the pool bounds the duplicates.
- Add a fourth blocking-io anchor driving build_thread_checkpoint_state_
accessor with a per-resolution fresh factory (always a cache miss) and
the real production blocking read; enumerate all four offloads in the
gate's module docstring. Verified the anchor goes red when
abuild_checkpoint_state_accessor is flattened back to a plain call.
* fix(subagents): revalidate batch item before launch; make assembly pool observable
Review follow-up for #5224:
- _execute_item() revalidates the durable state right after assembly and
before executor.execute_async(): renew_item_lease() returns valid=False
when cancel_batch() terminalized the item or the lease was lost while
assembly was parked, and the launch is skipped (the canceller already
finalized the item). Previously the launch was unconditional and the
poll loop's cancellation checks only started after execution began.
- Regression test driving the real SQLite repository: a blocking assembly
probe parks _execute_item, cancel_batch() lands, and the launch is
skipped with the item staying cancelled. Verified the test goes red
when the revalidation is removed.
- run_assembly() tracks pending assemblies and logs a throttled WARNING
once the pending count exceeds the worker count, so assembly starvation
(workers parked on a hung MCP server) is distinguishable from idle.
- The run_agent blocking-io anchor now binds a sentinel extension
snapshot via ctx.extensions and asserts the factory observed it through
get_agent_build_extensions(), pinning run_assembly()'s ContextVar
propagation. Verified red when ctx.run is dropped.
- Document the assembly pool in backend/AGENTS.md.
* fix(utils): decrement the assembly pending count on the pool thread
The pending-assembly counter behind the starvation warning decremented
from the asyncio future's done callback, which never fires once the
submitting loop is closed while its worker is still running: the count
ratcheted up permanently and eventually fired the starvation warning
with no starvation behind it (reproduced at 97dc9bec by review).
Decrement instead from the dispatched work item: run_assembly() wraps
func so a finally drops the count under the pending lock on the pool
thread, and the done callback is gone.
Pin the counter with tests/test_assembly_io.py: a healthy call returns
the count to zero, and an abandoned loop (stopped while the worker is
parked) does not wedge it — the abandoned case goes red against the old
done-callback decrement.
* docs(utils): fix the pending-counter comment after the decrement move
The comment still described the removed done-callback decrement,
contradicting _work()'s own comment; state the actual mechanism
(increment on the loop before dispatch, decrement from the dispatched
work item's finally on a pool thread).
* test(gateway): retarget checkpoint-accessor stubs to the services seam
thread_runs and runs now call abuild_checkpoint_state_accessor, so the
upstream wait-reader, regenerate-prepare, and idempotency tests must stub
the sync builder where abuild resolves it (app.gateway.services); stubbing
the removed router re-exports fails with AttributeError at setup. The
async seam semantics are unchanged: run_assembly invokes the stubbed
sync builder off-loop and propagates its return values and exceptions.
Move the agent/tool assembly off-load note from backend/AGENTS.md to
deerflow/utils/AGENTS.md (next to assembly_io.py) so the effective
instruction chain for agents/middlewares no longer grows past the AG002
hard limit.
* fix(runtime): serialize same-key accessor assembly and release queued-cancel slots
Address the three review follow-ups on the assembly off-load:
- assembly_io: a job cancelled while still queued never runs its work
item, so the dispatched finally never fired and _pending_assemblies
stayed elevated until a false starvation warning. Exactly-once cleanup
now rides the concurrent future's cancelled() state — cancel() only
succeeds before the executor starts the item, so cancelled() is true
precisely when the finally will never run — plus a submit-failure
release; the one-worker queued-cancellation case is pinned red/green.
- services: overlapping cold readers sharing one cache key could both
run full agent assembly. _state_accessor_graph now serializes per key
through a thread-side KeyedLockTable (pool threads, no running loop)
and re-validates factory/app-config identity under the lock, so the
factory runs exactly once while identity changes still rebuild. Cache
dict access is lock-guarded now that construction runs off-loop.
- guidance inventory: register deerflow/utils/AGENTS.md in
EXPECTED_GUIDANCE_PATHS so test_repository_has_the_approved_scoped_
guidance_shape matches the relocated assembly note (CI shard 4).
* test(keyed-lock): pin KeyedLockTable reclamation and waiter bypass directly
Thread-side counterparts of the async table's own tests: overlapping
hold() calls serialize (a late arrival joins the live entry instead of
creating a second lock that bypasses a queued waiter), the last check-in
pops the entry, and many unique keys leave the registry empty. Both
regressions verified red — popping unconditionally trips the late-arrival
test, never reclaiming trips the many-keys test.
---------
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
* docs: govern agent guidance size
* refactor: split agent guidance by code scope
* Clarify virtual path handling in AGENTS.md
Updated the translation section to clarify the role of `LocalSandboxProvider` and the handling of virtual paths in the tool layer.
---------
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>