Zeren Wang a06a6fed7e
feat(harness): deterministic acceptance checklist for subagent delegations (RFC #4651, layer 2) (#5109)
* feat(harness): deterministic acceptance checklist for subagent delegations (RFC #4651, layer 2)

PR4 of RFC #4651: check lead-supplied acceptance_criteria in code when a
subagent completes, so objectively checkable requirements can never be
silently passed by a self-report.

- subagents/acceptance_checks.py: deterministic leaf families —
  file:<path> exists|non-empty and file_written:<path> read through
  read_current_file_content scoped to the shared thread workspace; the
  read uses the sandbox-native virtual path form (the local read
  validator and provider mount tables resolve /mnt/user-data/... paths,
  not host paths); the scope decision canonicalizes with realpath on the
  local sandbox so workspace symlinks cannot escape into uploads; a
  remote provider's "Error: ..." return string is normalized to a
  failed check (provider-typed via is_local_sandbox); a
  UnicodeDecodeError marks a binary deliverable as existing and
  non-empty; out-of-scope paths degrade to UNVERIFIED.
  tests_passed:<command> anchors to a matching recorded bash execution
  with status=success and a test-summary shape; matching is
  shell-structure aware with control-flow attribution (span must end at
  the last segment with provable execution), negating-option values are
  ineligible evidence and a target negated anywhere in the command
  degrades the match, extra flags must be selection-preserving, extra
  positionals widen only after a path-scoped criterion, truncated
  commands degrade via command_truncated, the summary shape is read
  only from output attributable to the matched segment (preceding
  segments provably silent by invocation form), and pass shapes require
  a nonzero passed count. Criterion text is neutralized with
  neutralize_untrusted_tags before storage/rendering. Anything else
  renders UNVERIFIED, never silently passed.
- executor: accumulate bounded bash command/output evidence per streamed
  chunk (merged by tool_call_id, newest-capped) so subagent
  summarization compacting earlier messages cannot erase a recorded
  execution; the recorded status is the actual shell exit status parsed
  from the output's exit marker (signed codes included; the remote
  Command exited with code N form is accepted only as the whole trimmed
  output), falling back to deerflow_tool_meta only when no marker
  exists.
- sandbox providers: e2b/opensandbox/tenki/boxlite append the
  LocalSandbox-style "Exit Code: N" marker on nonzero exit even with
  non-empty output; aio propagates the SDK's structured exit_code on
  both exec paths the same way; local timeouts append Exit Code: 124;
  and _truncate_bash_output always preserves a trailing exit marker
  (signed included) inside its budget, with a 32-char floor raising any
  smaller configured limit, so the actual shell outcome always survives
  in the output text.
- task_tool: run the checklist offloaded (asyncio.to_thread) on the
  completed branch, failure-isolated; stamp the verdict into result
  metadata and render the per-criterion section into the model-visible
  result text.
- status contract: additive subagent_acceptance_verdict transport with
  read-side structural validation.
- delegation ledger: entry carries the verdict and renders a compact
  acceptance segment; gateway strips caller-forged verdicts from both
  ledger entries and message metadata, like the citation verdict.
- blocking-IO anchor pins the offload (teeth proven red->green); leaf
  read errors catch only OSError/SandboxError so unexpected errors reach
  the task-tool-level isolation instead of being mislabeled.

* fix(harness): close acceptance evidence gaps from review (RFC #4651 PR4)

- negating options: overlap with a matched criterion target is now
  checked by path/nodeid prefix, not exact token equality — excluding a
  sub-path of the criterion's selection (pytest tests --deselect
  tests/unit/test_auth.py) degrades to UNVERIFIED instead of holds
- output attribution: any redirection token in the matched final segment
  makes the recorded tail non-attributable (> / >> / 2> are word
  characters to the parser, so redirection was invisible to the matcher)
- silent-source allowlist narrowed from any *activate suffix to the
  */bin/activate shape
- status_contract docstring: restore the shared-fixture sentence and
  note subagent_acceptance_verdict is deliberately outside the fixture
- executor: update_bash_executions publishes [] (stream carried no
  bash-family calls) instead of collapsing it into None, mirroring
  update_tool_receipts

* fix(harness): close acceptance residual gaps from re-review (RFC #4651 PR4)

- tests_passed: add error outcomes to the fail shapes — "4 passed, 1 error"
  and pytest's "ERROR <nodeid>" short summary no longer satisfy the pass
  shape when the exit status is swallowed (|| true) or absent; zero-error
  counts stay clean.
- file leaves: bound the deliverable read — a "wc -c" shell size probe
  answers files above 50k bytes without loading ~2x their size, honoring
  the host-bash kill switch and falling back to the full read on any
  non-integer rendering, so verdicts never get less sound.
- executor: record the exit marker text as status_marker on harvested bash
  evidence; the leaf detail now reports the marker actually seen instead of
  asserting a failure indistinguishable from the command's own trailing text.
- extend the blocking-IO anchor to drive the probe branch inside the
  offload; teeth re-verified red->green.

* fix(harness): close acceptance forgery and bound gaps from P2 re-review (RFC #4651 PR4)

- file leaves: never read unbounded — size is established first (os.stat on
  the validated local host path, so the host-bash-disabled configuration
  needs no shell; a guarded wc -c on remote providers that renders
  missing/unreadable in its own words). Above the 50k cap the leaf answers
  from the size alone, at/below it the full read runs, and an
  unestablishable size degrades to UNVERIFIED instead of an unlimited
  fallback read.
- output attribution: source/. prefixes are never provably silent — a
  crafted */bin/activate path shape says nothing about what the script
  prints, so sourced segments can no longer lend a passing summary.
- executable identity: an explicitly path-spelled criterion now requires
  the same normalized executable path; the basename rule stays only for
  deliberately bare criterion commands.

* fix(harness): run acceptance size probe outside subagent-controlled state (RFC #4651 PR4)

- remote probe no longer runs in the sandbox's persistent shell: a fresh
  env -i /bin/sh with absolute-path stat/realpath (poisoned functions,
  aliases, PATH, exported functions, IFS, locale cannot steer it), plus a
  marker env routing AIO onto a fresh per-call bash.exec session.
- metadata-only: stat never opens content, so a FIFO deliverable cannot
  block the parent for the provider's idle timeout; non-regular files
  (fifo/dir/symlink) degrade to UNVERIFIED.
- containment canonicalized against the literal mount root: a
  final-component symlink or a swapped parent directory (root included)
  cannot redirect the check outside shared storage; unprovable layouts
  degrade to UNVERIFIED.

* fix(harness): canonicalize probe containment against the canonical mount root (RFC #4651 PR4)

Literal-root equality made every remote file leaf permanently UNVERIFIED
on e2b and Tenki, which realize /mnt/user-data as a symlink to the home
dir by default (e2b bootstrap 'sudo ln -sfn', Tenki best-effort symlink).
Containment now compares the file's realpath against the mount root's
realpath — exactly what the provider's own read path resolves, so probe
and read-back stay consistent; final-component symlinks stay rejected by
the non-dereferencing stat, and an intermediate dir-link escape under a
sane root still lands ESCAPED. The inner script is a module constant and
the suite now executes the composed probe for real against on-disk
layouts (real dir, symlinked prefix, final symlink, fifo, missing,
dir-link escape), which the canned-output stub could not see.

* fix(harness): close bare-criterion negation and CDPATH summary channels (RFC #4651 PR4)

- matching: a criterion with no positional selection target (bare pytest,
  make test) stands for the runner's default selection, so ANY negating
  option (--ignore/--deselect/...) makes the recorded run a different
  selection — unprovable. The overlap guard only sees consumed criterion
  tokens, which a bare criterion does not have; scoped criteria keep the
  unrelated-exclusion behavior.
- attribution: cd is no longer blanket-silent — CDPATH makes cd print the
  resolved (subagent-chosen) destination and the pass shapes match as
  substrings, so one mkdir 'all tests passed' plus an export minted a pass
  for any quiet command. A cd argument or CDPATH= value (export or leading
  assignment) carrying any summary shape makes the segment non-silent;
  shape-free cd dir wrappers keep matching.
- docs: _truncate_bash_output states the effective 32-char floor (the
  guarantee previously read as an unconditional max_chars bound).

* fix(harness): close env-assignment and expansion channels in acceptance matching (RFC #4651 PR4)

Self-audit in the shape of the last review rounds — channels the matcher
classified as accounted-for that can change what runs, narrow the
selection, or lend the summary text:

- env assignments are no longer blanket-stripped: only an allowlist of
  inert display/CI knobs (CI, NO_COLOR, PY_COLORS, ...) may prefix a
  matched span, and a non-allowlisted assignment in any preceding segment
  (pure-assignment or export NAME=) is state pollution — PATH redirects
  the executable, LD_PRELOAD/PYTHONPATH/NODE_OPTIONS inject code,
  PYTEST_ADDOPTS/GOFLAGS/MAKEFILES inject selection-changing inputs,
  BASH_ENV runs arbitrary shell startup. All degrade to unprovable.
- runtime expansions: any span token carrying /$( )/backticks, any
  negating-option value carrying an expansion or glob (unknown excluded
  set), and any extra executed token carrying glob metacharacters
  (crafted option-looking filenames narrow invisibly) are unprovable.
  Criterion-side globs stay self-consistent (literal match).
- cd: an argument carrying a runtime expansion or glob is non-silent
  (unknown destination, unknown print); CDPATH= assignments are now
  handled as state pollution at the match layer, subsuming the
  value-shape special case.

* fix(harness): persistent-shell evidence, exact env sets, option-arity scoping (RFC #4651 PR4)

- tests_passed: on a persistent-shell provider (new
  Sandbox.persistent_shell_sessions capability, set by AioSandbox) every
  leaf degrades to UNVERIFIED — any earlier call in the shared session
  could have mutated the state the clean-looking run executed in, and
  only a fresh controlled session (RFC section 6 verifier) can prove
  otherwise. The flag is read from the provider registry without
  acquiring a sandbox.
- env assignments: the allowlist is gone — no variable is provably inert
  across repositories (CI/DEBUG are routinely read by tests). The span's
  assignment prefix must equal the criterion's exactly (values included,
  order-insensitive); any assignment or export NAME= in a preceding
  segment is state pollution.
- scoping: positional targets are now read by option arity, so a path
  embedded in an option (--basetemp=/tmp/p, --junitxml=/tmp/r.xml) never
  counts as a selection target and an extra positional after such a
  criterion narrows the default selection it denotes.

* fix(harness): stamp shell provenance at harvest, close export/unset and arity gaps (RFC #4651 PR4)

* fix(harness): split physical newlines as shell separators in acceptance matching (RFC #4651 PR4)

* fix(harness): scope cd wrappers to thread data roots, pin accepted boundaries (RFC #4651 PR4)

* fix(harness): preserve criterion connectors, prove file_written readable, fail-closed shell capability (RFC #4651 PR4)

* fix(harness): compare only the connector prefix, tolerate trailing criterion semicolons (RFC #4651 PR4)

* fix(harness): preserve continuation-line operators, keep ./-spelled executable identity (RFC #4651 PR4)

* fix(harness): render criteria single-line so a multiline criterion cannot inject a forged checklist line (RFC #4651 PR4)

* fix(harness): reject parent-traversal executable tokens in acceptance matching (RFC #4651 PR4)

* fix(harness): reject parent-traversal negated values in acceptance matching (RFC #4651 PR4)
2026-09-01 16:13:41 +08:00

216 lines
8.7 KiB
Python

"""Deterministic capture and rendering for task delegations."""
from __future__ import annotations
import hashlib
from datetime import UTC, datetime
from html import escape
from typing import Any
from langchain_core.messages import AIMessage, AnyMessage, ToolMessage
from deerflow.agents.middlewares.receipt_verification import render_citation_verdict, validate_receipt_verdict
from deerflow.agents.thread_state import DelegationEntry
from deerflow.subagents.acceptance_checks import render_acceptance_segment, validate_acceptance_verdict
from deerflow.subagents.status_contract import (
read_subagent_result_metadata,
)
_RESULT_BRIEF_CAP = 2000
_DESCRIPTION_CAP = 200
_LEDGER_RENDER_CHAR_BUDGET = 6000
_LEDGER_ENTRY_RESULT_RENDER_CAP = 120
_STATUS_ONLY_RESULT_BRIEFS = {
"failed": "Task failed.",
"cancelled": "Task cancelled by user.",
"timed_out": "Task timed out.",
"polling_timed_out": "Task polling timed out.",
}
def _utc_now_iso() -> str:
return datetime.now(UTC).isoformat().replace("+00:00", "Z")
def _bound_text(text: str, cap: int = _RESULT_BRIEF_CAP) -> str:
"""Deterministic head/tail truncation. This is not an LLM summary."""
if len(text) <= cap:
return text
if cap <= 0:
return ""
head = cap * 2 // 3
omitted_marker = "\n...\n"
if cap <= len(omitted_marker):
return text[:cap]
tail = cap - head - len(omitted_marker)
if tail <= 0:
return text[:cap]
return f"{text[:head]}{omitted_marker}{text[-tail:]}"
def _escape_context_text(value: object) -> str:
return escape(" ".join(str(value).split()), quote=False)
def _status_guidance(status: str, stop_reason: str | None = None) -> str:
if stop_reason:
# A guardrail cap ended this run early (#3875 Phase 2): the status is
# still completed/failed, and ``stop_reason`` carries *why* it stopped
# (token_capped / turn_capped / loop_capped). The old contract surfaced
# this as a separate ``max_turns_reached`` status; the additive
# ``stop_reason`` field replaced it so v1 consumers keep working.
if status == "completed":
return "hit a guardrail cap with a partial result; reuse the partial result, retry with a tighter scope, or raise the per-agent budget (max_turns / token_budget)"
return "hit a guardrail cap with no usable result; retry with a tighter scope or raise the per-agent budget (max_turns / token_budget)"
if status == "in_progress":
return "already delegated; do NOT delegate again; wait for or build on the result"
if status == "completed":
return "completed result; do NOT delegate again; reuse this result"
if status == "failed":
return "failed attempt; may retry with a changed plan"
if status == "cancelled":
return "cancelled attempt; may retry with a changed plan"
if status == "timed_out":
return "timed-out attempt; may retry with a changed plan"
if status == "polling_timed_out":
return "polling timed-out attempt; may retry with a changed plan"
return "prior attempt; inspect status before retrying"
def _tool_call_name(tool_call: dict[str, Any]) -> str:
name = tool_call.get("name")
if isinstance(name, str):
return name
function = tool_call.get("function")
if isinstance(function, dict) and isinstance(function.get("name"), str):
return function["name"]
return ""
def _tool_call_id(tool_call: dict[str, Any]) -> str | None:
tool_call_id = tool_call.get("id")
return str(tool_call_id) if tool_call_id else None
def _tool_call_args(tool_call: dict[str, Any]) -> dict[str, Any]:
args = tool_call.get("args")
return args if isinstance(args, dict) else {}
def extract_delegations(messages: list[AnyMessage]) -> list[DelegationEntry]:
"""Enumerate `task` delegations from AI tool calls and paired results."""
entries_by_id: dict[str, DelegationEntry] = {}
order: list[str] = []
now = _utc_now_iso()
for message in messages:
if not isinstance(message, AIMessage):
continue
for tool_call in message.tool_calls or []:
if _tool_call_name(tool_call) != "task":
continue
tool_call_id = _tool_call_id(tool_call)
if tool_call_id is None:
continue
args = _tool_call_args(tool_call)
description = str(args.get("description") or args.get("prompt") or "")[:_DESCRIPTION_CAP]
if tool_call_id not in entries_by_id:
order.append(tool_call_id)
entries_by_id[tool_call_id] = {
"id": tool_call_id,
"description": description,
"subagent_type": str(args.get("subagent_type") or ""),
"status": "in_progress",
"created_at": now,
}
for message in messages:
if not isinstance(message, ToolMessage):
continue
tool_call_id = str(message.tool_call_id) if message.tool_call_id else ""
entry = entries_by_id.get(tool_call_id)
if entry is None:
continue
structured = read_subagent_result_metadata(message.additional_kwargs)
if structured is None:
continue
entry["status"] = structured["status"]
stop_reason = structured.get("stop_reason")
if stop_reason:
entry["stop_reason"] = stop_reason
receipt_verdict = structured.get("receipt_verdict")
if receipt_verdict:
entry["receipt_verdict"] = receipt_verdict
acceptance_verdict = structured.get("acceptance_verdict")
if acceptance_verdict:
entry["acceptance_verdict"] = acceptance_verdict
result_text = structured.get("result_brief") or structured.get("error") or _STATUS_ONLY_RESULT_BRIEFS.get(structured["status"])
if result_text:
result_sha256 = structured.get("result_sha256") or hashlib.sha256(result_text.encode("utf-8")).hexdigest()
entry.update(
{
"result_brief": _bound_text(result_text),
"result_sha256": result_sha256,
"result_ref": str(message.id or tool_call_id),
}
)
return [entries_by_id[tool_call_id] for tool_call_id in order]
def _fits_budget(lines: list[str], candidate: str, max_chars: int) -> bool:
return len("\n".join([*lines, candidate])) <= max_chars
def _render_entry_line(entry: DelegationEntry) -> str:
status = _escape_context_text(entry["status"])
description = _escape_context_text(entry["description"])
subagent_type = _escape_context_text(entry["subagent_type"])
guidance = _status_guidance(entry["status"], entry.get("stop_reason"))
line = f"- [{status}] {description} (via {subagent_type}; {guidance})"
result_brief = entry.get("result_brief")
if result_brief:
line += f" -> {_escape_context_text(_bound_text(result_brief, _LEDGER_ENTRY_RESULT_RENDER_CAP))}"
receipt_verdict = validate_receipt_verdict(entry.get("receipt_verdict"))
if receipt_verdict is not None:
segment = render_citation_verdict(receipt_verdict)
if segment:
line += f" · {segment}"
acceptance_verdict = validate_acceptance_verdict(entry.get("acceptance_verdict"))
if acceptance_verdict is not None:
segment = render_acceptance_segment(acceptance_verdict)
if segment:
line += f" · {segment}"
return line
def render_delegation_ledger(entries: list[DelegationEntry], *, max_chars: int = _LEDGER_RENDER_CHAR_BUDGET) -> str:
"""Render the delegation ledger as model-visible system context."""
if not entries:
return ""
lines = [
"## Work already delegated",
"Newest entries are shown first. In-progress entries are already delegated. Completed entries are reusable results. Failed, cancelled, or timed-out entries are prior attempts.",
]
omitted = 0
for index, entry in enumerate(reversed(entries)):
line = _render_entry_line(entry)
if _fits_budget(lines, line, max_chars):
lines.append(line)
continue
omitted = len(entries) - index
break
if omitted:
omitted_line = f"- ... {omitted} older delegation entries omitted from this model view because of context budget"
while len(lines) > 1 and not _fits_budget(lines, omitted_line, max_chars):
lines.pop()
omitted += 1
omitted_line = f"- ... {omitted} older delegation entries omitted from this model view because of context budget"
if _fits_budget(lines, omitted_line, max_chars):
lines.append(omitted_line)
rendered = "\n".join(lines)
if len(rendered) <= max_chars:
return rendered
return rendered[: max(0, max_chars - 4)] + "\n..."