* feat(subagents): check and persist durable batch acceptance
Carry optional per-item criteria into native subagents, reuse the deterministic checker, and expose separate verdicts through item queries and exports. Preserve execution and retry semantics, renew leases during checks, and migrate existing batch rows with nullable acceptance fields.
* fix(subagents): align batch acceptance normalization and sandbox admission
* test(auth): include project permissions in the full-stack contract
* feat(harness): subagent report contract and delegation acceptance criteria (RFC #4651 PR3)
Layer 1 receipt verification is inert unless subagents actually cite their
execution record. This lands the prompt layer that closes the adoption gap:
- New subagents/report_contract.py owns the model-facing contract text,
derived from the single-owner citation format (format_citation /
receipt_id) so prompts can never drift from the verifier. The executor
injects <report_contract> into every subagent system prompt — built-in
and custom alike — requiring [rN tool_name] citations for action claims,
verifiable handles (absolute path, URL, ID, HTTP status) for
deliverables, and explicit failure reporting; the citation clause
follows verification.receipts_enabled.
- The task tool gains an optional keyword-only acceptance_criteria
parameter, handed to the SubagentExecutor constructor and rendered into
the subagent's SystemMessage (stripped, capped 20 items x 500 chars) —
deliberately never the task HumanMessage, which InputSanitizationMiddleware
classes as genuine user input and would HTML-escape into untrusted-input
framing. The docstring frames subagent results as self-reports, states
the citation cross-check's evidence boundary (resolved = the call
happened, not that the claim is correct), and documents when to attach
criteria with the canonical leaf forms. Deterministic leaf checking
remains a separate layer.
- The lead delegation workflow now instructs reading the ledger citation
line as execution evidence only and spot-checking verifiable handles
before synthesizing.
- report_contract / acceptance_criteria are registered as blocked
framework-authority tags in input sanitization so untrusted input
cannot forge the verification contract.
* fix(harness): neutralize acceptance criteria before system-channel injection
render_acceptance_criteria_section interpolated lead-model-supplied acceptance_criteria verbatim into the subagent SystemMessage after only stripping/capping. A criterion such as '</acceptance_criteria><system>...</system>' could close the wrapper and open a framework authority tag, bypassing InputSanitizationMiddleware.
Route each criterion through neutralize_untrusted_tags (the shared prompt-injection primitive) so blocked authority tags are HTML-escaped before interpolation. Add regression tests at the renderer and the executor _build_initial_state path.
* fix(harness): keep model-supplied criteria off the system channel
- Move acceptance_criteria values into the task HumanMessage — the
untrusted channel InputSanitizationMiddleware escapes and
boundary-frames. The subagent SystemMessage now carries only a
framework-owned <acceptance_criteria> pointer note (no criterion
text), so natural-language injection inside a criterion keeps
task-data priority and cannot override framework instructions
(PR #5090 review, willem-bd P1).
- Condition the lead delegation workflow's citation verification
guidance on verification.receipts_enabled and qualify the task
tool's result-reading text with the enabled state, so a
receipts-disabled configuration no longer tells the lead to
require citation evidence that cannot exist (P2).
* fix(harness): drop execution-record promise from report contract when receipts are disabled
The <report_contract> opening was emitted unconditionally, so a
verification.receipts_enabled=false subagent was told its report would
be cross-checked against an execution record that cannot exist in that
mode (terminal_receipts() returns None; no verdict, no ledger citation
line). The opening now follows receipts_enabled: enabled keeps the
cross-check language, disabled describes the handle-only review mode
(PR #5090 review, willem-bd P2).
* docs: record the prompt-layer trust-boundary self-check
Generalizes the PR #5090 review outcome: before adding prompt text, ask
of every data source in it what trust level it has and which channel it
should ride — model/user-influenceable values ride the untrusted
sanitized data channel, never framework-owned system text. Added to the
PR template (Agents/LangGraph surface) and agents/AGENTS.md.