* fix(clarification): drop sibling tool calls before interrupt
- Rewrite the AIMessage in ClarificationMiddleware.after_model so a
parallel bash/write_file cannot run before the user answers
- langchain return_direct only inspects the last ToolMessage; siblings
both execute and can keep the agent loop alive
- Skip the rewrite when disable_clarification is set
- Prompt and tool docs: do not call other tools in the same turn
Fixes#4906
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(clarification): enhance sibling tool call handling in ClarificationMiddleware
- Update ClarificationMiddleware to ensure sibling tool calls are dropped when `ask_clarification` is invoked, preventing unintended execution before user input.
- Modify documentation to clarify that the `return_direct` router now inspects all client-side tool calls of the last AIMessage, ensuring proper routing behavior.
- Introduce a new integration test to validate that sibling tools do not execute when `ask_clarification` is present in the same turn.
This change addresses potential issues with tool execution order and improves the overall reliability of the middleware.
Fixes#4906
* fix(clarification): enhance tool call filtering in ClarificationMiddleware
- Update _filter_content_tool_use to handle Gemini-style function_call blocks by matching on name when no id is present, ensuring proper filtering of tool calls.
- Modify ClarificationMiddleware to maintain sibling tool call integrity by dropping unnecessary blocks, improving the clarity of the AIMessage content.
- Add a new test to validate the correct stripping of idless function call content blocks, ensuring that sibling tool calls do not execute prematurely.
This change improves the robustness of the middleware and addresses potential execution order issues.
Fixes#4906
* fix(clarification): drop siblings when ask_clarification is malformed
LangChain parks invalid args on invalid_tool_calls independently, so a
valid sibling would otherwise still execute before the user answers.
Co-authored-by: Cursor <cursoragent@cursor.com>
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
Adds scheduler.recursion_limit to config.yaml (default 1000, clamped by
max_recursion_limit) so scheduled background runs can use a different
recursion limit than the web UI. The value is read at dispatch time, so
a YAML edit applies to the next scheduled run without a Gateway restart.
Also logs a warning when the resolver falls back to the default or
clamps the configured value.
The cancel endpoint resolved McpTaskService whenever SQL persistence was
configured, even with mcp_tasks.enabled=false, and acknowledged the
request by recording cancel_requested_at. The background loop that owns
the remote cancel call only runs when enabled, so the fence was never
claimed and the remote task kept running indefinitely.
Gate the endpoint on app.state.mcp_tasks_available (set only after the
service is started) and return 503 before writing the fence, consistent
with the background-loop ownership contract. Read-only list/detail
endpoints remain available while the worker is stopped.
* feat(harness): add deterministic tool receipts with model-visible ledger
Stamp an immutable per-call fact record (tool name, status, args/output
hashes, byte count, timestamp) onto every tool result via a new
ToolReceiptMiddleware, and inject the derived receipt ledger (r1..rN)
into the model context so subagent reports can cite executed actions.
- tool_receipt.py: receipt core (make/extract/render), newest-first
budget eviction, ids derived from the append-only message stream
- ToolReceiptMiddleware: stamps ToolMessages directly or inside
Command-wrapped results; hidden ledger injection mirrors
DurableContextMiddleware; sits between ToolProgress and
ToolErrorHandling with a build-time ordering guard
- config: new verification section (receipts on, judge off), config
version 32 -> 33 with example/helm/docs updates
* feat(harness): split receipt rendering from stamping; address PR review
Review fixes (PR #4659):
- output_sha256 now uses sort_keys=True for structured content, matching
the order-invariant args fingerprint
- stamping failures log at warning (silent ledger gaps would corrupt
citations); tool execution remains never blocked
- _insert_after_leading_system_messages extracted to shared public
message_utils.insert_after_leading_system_messages; both middlewares
depend on it instead of a private cross-module helper
- code comments in English
RFC #4651 revision-2 alignment:
- receipts_render_mode config ('always' | 'delegation_only'): subagent
chains always render the ledger (citations are produced there); the
lead chain renders only while processing subagent results, removing
the always-on token tax from ordinary turns
- receipts gain bounded args_preview/output_preview (<=200 chars, tail
for output) so later typed claim bindings (tests_passed) can anchor
to a specific recorded execution
* docs(harness): state receipt freshness caveat and vocabulary layering in module docstring
* merge: upstream/main — resolve AGENTS.md split, bump config_version to 34, drop unused receipt previews
- backend/AGENTS.md: take upstream's slimmed root guidance (#4799); move the
ToolReceiptMiddleware chain entry into agents/middlewares/AGENTS.md and the
verification.* hot-reload mention into config/AGENTS.md
- config.example.yaml + helm values/README: config_version 33 -> 34 so existing
v33 configs get the outdated-config prompt (review: willem-bd)
- tool_receipt.py: drop args_preview/output_preview — no Layer 1 consumer reads
them; re-add with the Layer 2 claim-binding consumer (review: willem-bd)
* docs(harness): cover receipt id renumbering after compaction in module docstring
Positional display ids are stable only while history is append-only;
compaction drops ToolMessages and the survivors renumber, so Layer 2
citation verification must resolve [rN] against the ledger as of the
citing turn (review: willem-bd, doc-only).
* chore(config): bump config_version to 35
main reached 34 via #4780 without the verification section; publishing
the new schema at the same number would silently skip the outdated-config
prompt for configs synced from main in that window (review: willem-bd).
* fix(skills): restore errno import dropped upstream in #4830
upstream/main adf6c422 uses errno.ENOTDIR in the drift guard but removed
the import, so the PR merge ref fails lint-backend (F821).
* fix(harness): harden tool receipts against forgery and turn-scope delegation_only
Address willem-bd's pre-merge review on #4659:
1. Untrusted receipt metadata: the gateway now strips the server-owned
deerflow_tool_receipt key from external input messages; stamping always
overwrites any tool-supplied value instead of preserving it; and
extract_tool_receipts validates persisted receipt shapes (required typed
fields, unknown keys ignored) so malformed entries are skipped instead of
crashing render or passing as runtime-stamped evidence.
2. delegation_only no longer sticks on: _should_render now scopes the
subagent_status scan to the current turn (messages after the latest
genuine user message), so an old completed delegation stops rendering the
ledger on later ordinary turns. The genuine-user predicate moves to
message_utils.is_genuine_user_message, shared with input sanitization.
* fix(harness): stamp receipts outside short-circuiting tool middlewares
Address willem-bd's review on #4659: ToolReceiptMiddleware was registered
inside Guardrail/SandboxAudit/ReadBeforeWrite/ToolProgress, each of which
can return a ToolMessage without invoking its handler — blocked calls
(e.g. a read-before-write-denied write_file) never got a receipt, silently
gapping the ledger on a default-enabled path. SandboxAudit additionally
rebuilds medium-risk results, dropping an inner stamp.
ToolReceiptMiddleware is now the outermost wrap_tool_call layer in the
runtime tail. Normal results still carry deerflow_tool_meta (stamped by
ToolErrorHandling on the inner return path); short-circuit messages
self-stamp meta or fall back to message.status. The new invariant is
declared as ordering constraints in deerflow.extensions.ordering, with
composed-chain regression tests for a blocked write and a warn-rebuilt
bash result.
* feat(mcp): per-user credential injection for shared MCP servers
A single HTTP/SSE MCP server entry can now serve several users, each
authenticated to the remote service with their own credential. A server
opts in with a user_auth block mapping user ids to credential header
values ($ENV_VAR references supported):
"user_auth": {
"header": "Authorization",
"users": { "<user-id>": "$SERVICE_TOKEN_ALICE" }
}
The built-in user-scoped auth interceptor resolves the authenticated
runtime user on every tool call (request runtime -> ambient LangGraph
runtime -> auth config -> request-scoped user ContextVar) and rewrites
the configured header via request.override(), the same per-call
mechanism as the OAuth interceptor. It registers after OAuth in the
shared assembly so its per-user value wins the header when a server
declares both. The entry's static headers are used only for startup
tool discovery.
Fail-closed by default: an unmapped user - including the anonymous
default-user fallback - or a credential whose env reference resolved
empty gets an actionable ToolException instead of another user's
credential; on_missing: "passthrough" opts out per server. Combined
with the existing per-(user, thread) MCP session scoping this gives
credential isolation on shared servers.
Gateway API: user_auth.users values are masked in GET responses, and
PUT round-trips preserve stored credentials for masked values (same
contract as env/headers/oauth secrets); a masked value for a user id
not already stored is rejected.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(mcp): address review — preserve stored user_auth sub-fields on partial PUT, warn-and-skip stdio, allow extras
Review findings from #4868:
1. A partial user_auth payload (e.g. {"enabled": false}) merged to
users={} and default on_missing, irreversibly wiping stored
credentials on PUT. The merge is now sub-field-aware via
model_fields_set — omitted sub-fields carry the stored values, an
explicitly sent users map still replaces (so full-round-trip removal
works), masked values still swap back for stored credentials.
2. user_auth on a stdio server was a silent no-op: rewritten headers go
to call meta, never a transport header, while deny errors still fired.
The interceptor builder now warns and skips non-sse/http servers,
matching the tool_call_timeout transport-mismatch convention.
3. McpUserScopedAuthConfigResponse now allows extra keys like the
harness-side model, and extras survive masking and merge, matching
the server-level model_extra handling.
Adds four regression tests (partial-PUT preservation, explicit-map
replacement, extras round-trip, stdio warn-and-skip).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(mcp): reject blank user_auth.header at the gateway
A blank header passed the gateway response model, was persisted, then
failed the harness-side ExtensionsConfig validator on reload — the PUT
returned 500 after the write and every later config load/startup failed
until the file was hand-edited. Mirror the harness non-blank validator
on McpUserScopedAuthConfigResponse so the PUT fails with 422 before
anything is written.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* style: ruff format extensions_config.py
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(mcp): harden the trust chain and masking around user-scoped credential selection
Address review round 4:
- Scrub client-supplied user_id from run context/configurable for
external callers in inject_authenticated_user_context, before every
early return, and restamp only from request.state.user. Now that
user_id selects which user's credential user-scoped MCP auth injects,
a forged value must not survive any future path that skips the
restamp. Internal callers (IM channels, scheduler) keep supplying
end-user identity as before (PR #3294 contract). Regression tests pin
both the scrub and that a forged body.context.user_id can never
resolve as another user through merge + inject ordering.
- Include the caller's own resolved user id in the fail-closed deny
message so operators can copy the exact users key (it differs by
deployment path), and document the key formats in the mcp.mdx doc.
- Mask sensitive extra keys inside user_auth on GET like server-level
extras, and swap masked sentinels back for stored values on PUT via
_merge_extra_value_preserving_masked.
- Extract the interceptor wrap loop into compose_tool_interceptors and
pin the security property functionally: an OAuth interceptor that
actually sets Authorization loses the final header to the per-user
credential through the same composition the session-pool path uses.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* fix(docker): let the Gateway write extensions_config.json in production
AGENTS.md states extensions_config.json may be edited at runtime through the
Gateway API, and the Gateway implements that for the MCP enable switch,
PUT/PATCH /api/mcp/config and the skill update route. Two properties of the
production compose stack made every one of those writes fail:
- the file was mounted read-only, and
- Docker mounts it as its own mount point, so the temp-file-plus-rename in
atomic_write_extensions_config hit EBUSY. Linux refuses rename() over a
mount point whether or not the mount is writable, so making the mount
read-write alone is not enough.
Mount it read-write and fall back to an in-place overwrite on EBUSY only.
The fallback is deliberately non-atomic and says so in a warning; it is
reached only where the atomic route cannot work, and any other errno still
propagates. config.yaml stays read-only: no API writes it.
docker-compose-dev.yaml mounts the whole project directory, so the
destination is an ordinary file there and this never surfaced in development.
* test(docker): parse mount options instead of matching a :ro suffix
Docker's short-syntax options segment is comma-separated, so a read-only
mount can legally be spelled ":ro,z" or ":z,ro" — common with SELinux
relabelling. Matching the raw string for a ":ro" suffix reads those as
writable, which silently defeats the guard: the writability assertion would
pass on a read-only mount, and the config.yaml assertion would fail on a
correctly read-only one.
Parse the options segment and test membership instead, and cover the parser
with the spellings that broke the suffix check.
* fix(config): harden mutable extensions config
Move SQL-dependent connection identity and /start bind work onto the
Gateway main loop via _submit_threadsafe_coroutine, and send PTB replies
through _run_on_telegram_loop so the Telegram worker never blocks or
touches SQLAlchemy/HTTP across event loops.
When the main loop is not running (e.g. during gateway shutdown), the
bind path logs a warning and returns False instead of running SQLAlchemy
on the wrong loop, matching the Feishu bind pattern.
* feat(extensions): let an out-of-tree extension observe what the agent did
DeerFlow's extension system can contribute middleware, services and routes,
but an extension cannot answer basic questions about a run without reaching
into host internals. Several of the facts it would need are destroyed by the
operations that produce them:
* The middleware chain injects and rewrites a lot of context — date
reminders, recalled memory, compaction summaries, durable-context data,
image payloads, activated skill bodies. Downstream, none of it is
attributable: at the model-call boundary an injected HumanMessage is
indistinguishable from the user's own, and anything wanting to tell them
apart has to pattern-match prompt wording, which breaks on the next copy
edit.
* Two runs of "the same agent" are only comparable if the chain enforced the
same limits, prompts and thresholds. Recovering that from outside means
reading private attributes and guessing which of them change behaviour — a
guess that rots silently as middlewares gain fields.
* The lead-agent factory resolves a model after runtime overrides, renders a
prompt, filters tools through authorization and composes a stack, all
inside one synchronous call, and none of it survives: a middleware sees its
neighbours but not the prompt, the run worker sees a graph but not what
went into it.
* Summarization is destructive by design. N messages leave the context and
one summary enters it; afterwards only the summary exists, so "which
messages became this?" is not reconstructible.
This adds seven neutral facilities so those facts are recorded where they are
still true, and releases the contract package as 0.2.0.
Message provenance
Producers stamp `deerflow_content_kind` / `deerflow_producer_kind` onto the
messages they inject or rewrite. Stamping is unconditional — a fact whose
presence depends on whether an observer is installed is not a fact — and the
keys are server-owned, so provenance cannot be forged from a request.
Middleware self-description
Twelve middlewares declare their own behaviour-affecting parameters through
a duck-typed `release_policy_parameters()`. Long text is hashed rather than
embedded: a declaration is an identity, not a copy of the prompt.
Agent assembly descriptor
`assemble_lead_agent()` returns the graph plus a descriptor whose fingerprint
answers "did anything about this agent change between these two runs?".
`make_lead_agent()` keeps its graph-only signature — it is the LangGraph
Server ABI declared in langgraph.json. Tools and skills are sorted before
hashing because their assembly order is incidental; middlewares are not,
because stack order decides what wraps what. Host build identity is reported
but excluded from the fingerprint, so a redeploy does not invalidate every
agent's identity.
Context compaction observation
Summarization emits the content hashes of the messages it is about to remove
joined to the summary that replaced them. Content is the only identity
available at that seam: the summary does not become a message, and what later
projects it into a request renders it bounded and escaped rather than
verbatim.
Neutral policy, transform and MCP-source facts
Guardrail decisions are published to runtime context under a `__`-prefixed
key; result-rewriting middlewares append a declared, ordered transform trail;
MCP tools carry their credential-free logical origin.
Extension route identity
Contributed routes are session-authenticated and cannot opt out, but
"logged in" and "administrator" are different questions. Extensions get a
neutral projection of the caller rather than the host's auth context, and
`require_admin` fails closed when identity cannot be determined.
Extension-owned tables
An extension that persists data owns its own MetaData and migration chain, so
its tables are absent from Base.metadata and `alembic revision --autogenerate`
proposes dropping them. Extensions declare a table prefix, which is rejected
at registration if it would shadow a host table.
The contract package stays dependency-free and imports no host code; every new
Protocol method has a default so later additions remain additive. The loader's
pre-1.0 rule requires an exact major.minor match, so extensions written against
0.1 are now refused at startup with an actionable install hint rather than
loading into a host that implements a different surface.
uv.lock records the contract package's new version, so `uv sync --locked` still
resolves on a fresh checkout.
* fix(backend): sort gateway service imports
* fix(mcp): keep grant_type authoritative over extra_token_params
_fetch_token built the token request body as
{"grant_type": oauth.grant_type, **oauth.extra_token_params}, so an
operator-supplied extra_token_params that happened to contain
"grant_type" silently overwrote the value sent to the token endpoint
while the branch logic below still keyed off oauth.grant_type — the
sent grant_type and the chosen auth flow would disagree, and the
provider would almost certainly reject the request.
Spread extra_token_params first and set grant_type (and the other
reserved fields, which were already set after the spread) afterward, so
operator-supplied params can populate arbitrary extra fields but never
override the reserved ones the flow depends on.
* test(mcp): cover extra OAuth token parameters
---------
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
* fix(e2b): preserve trailing whitespace in filenames and survive mtime overflow
_sync_outputs_to_host iterated the NUL-delimited find output with
entry.strip() on each record. NUL already guarantees record boundaries,
so the strip is redundant and harmful: a filename that legitimately ends
in whitespace (e.g. "report ") had its trailing space trimmed, pointing
host_path at the wrong file and recording a manifest key that can never
match — the file was re-downloaded on every release.
The same host-write block wrapped only os.utime in the outer
except OSError, but os.utime raises OverflowError (not an OSError) when
the ns value is out of range (a far-future remote mtime, e.g.
`touch -d '99999 years'`). That escaped the loop, skipping the manifest
write and forcing a full re-download next release. Wrap os.utime in its
own (OSError, OverflowError) so only the timestamp restoration is
dropped; the file is still written and the manifest still updated.
* test(e2b): rely on monkeypatch cleanup
---------
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
* fix(buzz): drop replayed events across reconnects with a persistent seen-id store
The Buzz connector's resubscribe filter replays by design: 'since' is the
created_at of the last processed event and NIP-01 'since' is inclusive, so
every relay reconnect redelivers at least that event. The guard against
re-running the agent on it was the manager's inbound dedupe, whose default
store is in-process with a 10-minute TTL — so any reconnect more than ten
minutes after a channel's last message (or any gateway restart) re-answered
that message. Users saw the agent respond to an old question after every
relay restart.
Fix: persist the ids of fully processed events per channel
(BuzzSeenEventStore, JSON under {base_dir}/channels/, atomic writes) and
drop redelivered ids in _handle_chat_event before the /connect branch —
a replayed /connect would otherwise be re-answered with a spurious
'code invalid or expired'. Matching is by exact event id only, never
timestamp, so a genuinely new event (same-second or clock-skewed author)
can never be skipped, preserving the connector's fail-toward-replay
invariant. Only fully processed events are recorded, mirroring the
watermark rule: a gated drop or failed publish stays replayable.
Fail-open in both directions: an unreadable store loads empty (costs one
replayed reply, the previous behavior) and a failed write is logged and
retried on the next record. Id lists and the channel map are bounded like
the connector's other remote-fed maps. The persistent path is wired in
ChannelService (like channel_store); directly constructed channels get a
memory-only store so tests and tooling stay free of filesystem side
effects.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(buzz): coalesce seen-store writes, clean up temp files, harden docs and coverage
Address review on the seen-event store:
- record() now marks the store dirty and coalesces persistence to one
write per FLUSH_DELAY_SECONDS on the event loop, so a reconnect
backlog burst pays one O(store) file write instead of one per event;
sync callers (no running loop) keep immediate writes, and
BuzzChannel.stop() flushes so a clean shutdown loses nothing. A crash
inside the window only costs replay, never a skip.
- _save() unlinks its temp file on failure (ChannelStore parity), so a
persistently unwritable path no longer accumulates *.tmp litter.
- Module docstring now documents that restart protection is bounded to
the newest MAX_IDS_PER_CHANNEL ids per channel (and to raise it if a
relay ever serves a deeper default backlog), and pins the
single-event-loop assumption that makes the class safe without a lock.
- New tests: MAX_CHANNELS LRU eviction, coalescing behavior, flush
idempotence, temp-file cleanup, stop() flushing, and the
ChannelService wiring that injects seen_event_store_path (the line
that makes real deployments durable).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(buzz): reschedule the coalesced flush when the pending timer's loop is gone
A pending flush handle pinned to a since-closed event loop kept
_flush_handle non-None forever, so later record() calls on a new loop
never scheduled a timer and the store silently stopped persisting until
an explicit flush(). Track the scheduling loop (TimerHandle has no
public get_loop()) and reschedule when it differs from the running one.
Unreachable in production (one loop per process, stop() flushes), but
now hardened and tested.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* fix(mcp): exclude internal temp files from workspace changes
* fix(mcp): address review — shared tmp subdir constant, any-depth docs, nested test
- Export MCP_TMP_SUBDIR from constants.py and import it in both stdio
launch paths (tools.py, task_tool_caller.py) so the "/tmp" suffix is
composed once.
- Document that the .mcp exclusion matches by directory name at any
depth (like .git/node_modules) in README.md and mcp/AGENTS.md —
subagent work dirs below the workspace root get their own .mcp/tmp.
- Pin the any-depth semantic in test_workspace_changes.py with a nested
workspace/project/.mcp assertion.
* docs(mcp): correct any-depth exclusion rationale; re-home tmp pinning comment
The previous commit justified the any-depth `.mcp` exclusion with subagent
work dirs sitting below the workspace root — a mechanism that doesn't
exist: subagents share the parent's thread_id and both stdio launch paths
resolve sandbox_work_dir(thread_id), so `.mcp/tmp` is always pinned at the
workspace root. Reword mcp/AGENTS.md and the test comment to the real
justification (consistency with the other reserved dir names, robustness
against a server creating a relative `.mcp` from another cwd).
Also move the orphaned "pinning the process temp dir" rationale from
tools.py to constants.py next to MCP_TMP_SUBDIR, where both importers see it.
* fix(docker): keep runtime data out of the build context
backend/Dockerfile copies the backend tree wholesale, and .dockerignore did
not exclude the directories a running DeerFlow writes: DEER_FLOW_HOME
(backend/.deer-flow by default) and the local sandbox workspace root
(backend/sandbox).
Two consequences. Building on a host that has run DeerFlow bakes that state
into the image, including .jwt_secret and the sqlite user database. And once
the Gateway container has created directories owned by root, the build client
can no longer read them and the build fails outright:
target gateway: failed to solve: error from sender:
open .../.deer-flow/users/<uuid>/integrations/lark-cli: permission denied
Neither directory has tracked content, so excluding them costs the build
nothing. The new test pins both that the runtime paths are excluded and that
real build inputs still are not.
* fix(docker): exclude nested env files from builds
The Star History charts in the README files no longer render because the underlying chart service relies on GitHub stargazer data that is currently restricted. This switches the charts to a working alternative that needs no API token, updating the English, Simplified Chinese, Japanese, French, and Russian READMEs at the same time.
Co-authored-by: OctoBored <212877535+OctoBored@users.noreply.github.com>
* fix(sandbox): bound aggregate E2B mount upload work
* fix(sandbox): preserve mount guards on upload failure
* fix(sandbox): cover mount preflight with deadline
* refactor(sandbox): clarify mount deadline checks
* refactor(sanbox): deduplicate mount deadline reason
* fix(sandbox): evaluate mount deadline reason lazily
* feat: integrate MiniMax Code as an ACP agent
* Potential fix for pull request finding
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
---------
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
* fix(skills): fail closed on drifted projection namespace on all platforms
* test(skills): add regression test simulating swallowed unlink on drifted namespace
* fix(middleware): target the latest user message on first-turn fallback injection
When an earlier turn ends without any dynamic-context reminder — e.g.
the async abefore_agent degraded path times out and skips injection
(issue #3402's guard) — the next turn enters the first-injection branch
(last_date is None) on a history that already holds several turns.
That branch scanned from the start and attached the ID-swap to the
FIRST user message. The swap's {id}__user copy is appended by
add_messages, so the stale first prompt moved to the tail of history,
ahead of the current question — and the model answered the old prompt
as if it were the current turn.
Scan from the end instead (matching the midnight-crossing branch) so
the reminder attaches to the latest user message and history order is
preserved. Genuine first turns are unaffected: they have exactly one
message, which is both first and last. The pre-existing
test_injects_only_into_first_human_message_not_later_ones case encoded
the buggy target selection and is updated to the corrected contract.
* refactor: rename first_idx to target_idx after reversed scan
The branch now scans from the end, so the local holds the LAST user
injection target; first_idx read misleadingly. Match the
midnight-crossing branch's naming convention and clarify the log line
accordingly. No behavior change.
* feat(mcp): add durable task runtime foundation
* fix(chart): sync embedded config version
* fix(mcp): isolate task polls during shutdown
* feat(mcp): track consecutive poll errors on mcp_tasks
poll_attempt_count grows on every claim (successful polls included), so it
cannot drive a failure backoff without misjudging normal long tasks. Add
consecutive_poll_error_count: incremented when a claim is released after a
poll error, reset to zero by any applied snapshot. The backoff/terminal
policy that consumes it lands with the first concrete driver.
* fix(mcp): harden durable task lifecycle
* feat(mcp): add ordinary durable task driver
* test(mcp): address durable task review feedback
* fix(mcp): preserve submit tool descriptions
* fix(mcp): bound remote task calls
* fix(mcp): bound persisted task payloads
* fix(mcp): preserve task tool error details
* fix(mcp): enforce durable task boundaries
* test(mcp): cover task config snapshot lifecycle
---------
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
* Initial plan
* fix: enforce global concurrent-run budget for manual triggers
Manual triggers now check count_active_runs() before dispatching and
return a conflict result (409 at the router) when max_concurrent_runs
is already reached, preventing the global cap from being exceeded.
Co-authored-by: WillemJiang <219644+WillemJiang@users.noreply.github.com>
---------
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: WillemJiang <219644+WillemJiang@users.noreply.github.com>
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>