* feat(subagents): add delegations ledger field + reducer to ThreadState
* feat(subagents): pure helpers to derive + format the delegation ledger
* feat(subagents): DelegationLedgerMiddleware records + injects the ledger
* feat(subagents): register DelegationLedgerMiddleware for lead when subagents enabled + docs
* add runtime log
* chore(subagents): make delegation-ledger injection log production-ready
* test(subagents): make delegation-ledger registration tests config-free; refresh ultra replay golden for delegations channel
* refactor(subagents): derive TERMINAL_STATUSES from SUBAGENT_STATUS_VALUES + pin it
Make thread_state's TERMINAL_STATUSES a frozenset over the status contract's
SUBAGENT_STATUS_VALUES instead of a hardcoded literal, so the terminal-status
set can never drift from the contract. Add a pinning test asserting the
derivation and that the non-terminal "in_progress" stays excluded.
Addresses PR #3877 review.
---------
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
* fix(sandbox): stop blocking bash commands from hanging the turn
Starting a server through the host bash tool (e.g. `python -m http.server`)
could hang the whole turn for the full 600s timeout. `LocalSandbox.execute_command`
used `subprocess.run(capture_output=True)`, whose captured pipes are inherited
by any process the command spawns — so a backgrounded long-lived process
(`server &`) keeps the read end open and blocks `communicate()` until the
timeout fires, even though the foreground command already returned. Commands
that read stdin blocked the same way, and on timeout only the direct child was
killed, leaving orphaned process groups.
Rework the POSIX path to capture stdout/stderr via temp files instead of pipes,
take stdin from /dev/null, and run the command in its own session/process group:
- Backgrounded long-lived processes (servers) now return immediately while the
process keeps running.
- A command reading stdin gets immediate EOF instead of blocking.
- A genuinely blocking foreground command is bounded by a configurable
wall-clock timeout; on timeout the whole process group is killed and the agent
gets an explanatory notice telling it to background long-lived processes.
The timeout is configurable via `sandbox.bash_command_timeout` (default 600).
The Windows path is unchanged. Adds focused regression tests and updates docs.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(sandbox): instruct the agent to background long-lived processes
The code fix bounds a foreground server with a timeout, but the turn still
waits the full timeout before the run continues. Add the prompt-side half:
the bash tool description now tells the model to ALWAYS start long-lived
processes (e.g. web servers) in the background with output redirected, so the
tool returns immediately. The timeout notice points at the same readable
workspace log path. Pins the guidance with a test.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(sandbox): make fallback-kill exception explicit and observable
Address automated review: the inner `except OSError: pass` in
_terminate_process_group silently swallowed the case where the direct-child
fallback kill found the process already gone. Make the intent explicit with a
comment and a debug log instead of a bare pass.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(sandbox): address bash timeout review feedback
* fix(sandbox): document fd cleanup races
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Split `_build_runtime_middlewares`'s flat list into three named declarative
sublists (outer_wrappers / thread_hooks / tail) and drop the
`middlewares.insert(2, UploadsMiddleware())` magic-index pattern. The
declarative structure makes the layering self-documenting and immune to
position drift when the head of the list changes.
Move UploadsMiddleware to run after ThreadDataMiddleware in the chain.
Under the previous order (a magic-index artifact introduced when #3662
prepended InputSanitizationMiddleware), UploadsMiddleware scanned the
uploads directory before ThreadDataMiddleware created it under
lazy_init=False, so historical files could be missed on the first run of
a thread. Narrow path — the upload endpoint normally pre-creates the
directory — but the order is the correct semantic and is now locked.
Documentation:
- backend/AGENTS.md middleware chain renumbered: ThreadDataMiddleware is
now #3, UploadsMiddleware #4 (was reversed).
Tests (backend/tests/test_tool_error_handling_middleware.py):
- test_build_lead_runtime_middlewares_orders_thread_data_before_uploads
— focused td_idx < um_idx assertion.
- test_build_lead_runtime_middlewares_chain_order_matches_agents_md
— full-chain order pin using real classes, so a swap between any pair
is caught (the existing FakeMiddleware-stubbed tests cannot detect
this).
- test_lead_runtime_chain_finds_historical_uploads_under_lazy_init_false
— integration anchor: under lazy_init=False, ThreadDataMiddleware
creates the dirs, then UploadsMiddleware surfaces a pre-existing
historical file in the injected <uploaded_files> context.
* docs: add root-level CLAUDE.md to orient the monorepo
Adds a thin top-level CLAUDE.md that maps the monorepo and delegates depth
to backend/CLAUDE.md and frontend/CLAUDE.md, per issue #3761.
Includes the project overview + service topology (Nginx 2026, Gateway 8001,
Frontend 3000, optional Provisioner 8002), a top-level repository map, root
`make` vs. per-module command sections, "where to go next" links to the module
guides and primary root docs, and the repo-wide cross-cutting conventions
(documentation-update policy, TDD expectation, format before pushing).
No code or behavior changes; root points down, modules own depth.
Closes#3761
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs: make AGENTS.md the source of truth, CLAUDE.md a thin @AGENTS.md importer
Adopt the AGENTS.md convention so the same agent guidance serves Claude Code,
Codex, and other tools. At each level (root, backend, frontend) the content
lives in AGENTS.md and CLAUDE.md just imports it via `@AGENTS.md`.
- root: move the monorepo orientation layer to AGENTS.md; CLAUDE.md -> @AGENTS.md.
Fix an incorrect "TUI" reference (not present on main) and repoint the module
links to the AGENTS.md files.
- backend: move the guide to AGENTS.md (was an AGENTS.md -> @CLAUDE.md pointer;
direction is now flipped). Refresh stale content: rebuild the full middleware
chain (~26 ordered steps incl. InputSanitization, ToolOutputBudget,
DynamicContext, TokenBudget, SafetyFinishReason) from the actual build
functions; drop the brittle "11 middleware components" count; expand the
community-tools list to the real set.
- frontend: merge the practical Next.js guide with the existing AGENTS.md's
unique sections (LangGraph diagram, tech-stack versions, interaction
ownership, resources) into one AGENTS.md (CLAUDE.md -> @AGENTS.md). Fix the
stale src/ layout (remove the no-longer-present server/ better-auth entry;
add the now-active auth/agents/blog/... modules and routes) and drop a bogus
interaction-ownership bullet referencing files that don't exist.
Docs only; no code or behavior changes.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
- Add AioSandboxProvider for Docker-based sandbox execution with
configurable container lifecycle, volume mounts, and port management
- Add TitleMiddleware to auto-generate thread titles after first
user-assistant exchange using LLM
- Add Claude Code documentation (CLAUDE.md, AGENTS.md)
- Extend SandboxConfig with Docker-specific options (image, port, mounts)
- Fix hardcoded mount path to use expanduser
- Add agent-sandbox and dotenv dependencies
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>