Zeren Wang 4e35f0d1d4
feat(harness): deterministic tool receipts with model-visible ledger (RFC #4651, layer 1) (#4659)
* feat(harness): add deterministic tool receipts with model-visible ledger

Stamp an immutable per-call fact record (tool name, status, args/output
hashes, byte count, timestamp) onto every tool result via a new
ToolReceiptMiddleware, and inject the derived receipt ledger (r1..rN)
into the model context so subagent reports can cite executed actions.

- tool_receipt.py: receipt core (make/extract/render), newest-first
  budget eviction, ids derived from the append-only message stream
- ToolReceiptMiddleware: stamps ToolMessages directly or inside
  Command-wrapped results; hidden ledger injection mirrors
  DurableContextMiddleware; sits between ToolProgress and
  ToolErrorHandling with a build-time ordering guard
- config: new verification section (receipts on, judge off), config
  version 32 -> 33 with example/helm/docs updates

* feat(harness): split receipt rendering from stamping; address PR review

Review fixes (PR #4659):
- output_sha256 now uses sort_keys=True for structured content, matching
  the order-invariant args fingerprint
- stamping failures log at warning (silent ledger gaps would corrupt
  citations); tool execution remains never blocked
- _insert_after_leading_system_messages extracted to shared public
  message_utils.insert_after_leading_system_messages; both middlewares
  depend on it instead of a private cross-module helper
- code comments in English

RFC #4651 revision-2 alignment:
- receipts_render_mode config ('always' | 'delegation_only'): subagent
  chains always render the ledger (citations are produced there); the
  lead chain renders only while processing subagent results, removing
  the always-on token tax from ordinary turns
- receipts gain bounded args_preview/output_preview (<=200 chars, tail
  for output) so later typed claim bindings (tests_passed) can anchor
  to a specific recorded execution

* docs(harness): state receipt freshness caveat and vocabulary layering in module docstring

* merge: upstream/main — resolve AGENTS.md split, bump config_version to 34, drop unused receipt previews

- backend/AGENTS.md: take upstream's slimmed root guidance (#4799); move the
  ToolReceiptMiddleware chain entry into agents/middlewares/AGENTS.md and the
  verification.* hot-reload mention into config/AGENTS.md
- config.example.yaml + helm values/README: config_version 33 -> 34 so existing
  v33 configs get the outdated-config prompt (review: willem-bd)
- tool_receipt.py: drop args_preview/output_preview — no Layer 1 consumer reads
  them; re-add with the Layer 2 claim-binding consumer (review: willem-bd)

* docs(harness): cover receipt id renumbering after compaction in module docstring

Positional display ids are stable only while history is append-only;
compaction drops ToolMessages and the survivors renumber, so Layer 2
citation verification must resolve [rN] against the ledger as of the
citing turn (review: willem-bd, doc-only).

* chore(config): bump config_version to 35

main reached 34 via #4780 without the verification section; publishing
the new schema at the same number would silently skip the outdated-config
prompt for configs synced from main in that window (review: willem-bd).

* fix(skills): restore errno import dropped upstream in #4830

upstream/main adf6c422 uses errno.ENOTDIR in the drift guard but removed
the import, so the PR merge ref fails lint-backend (F821).

* fix(harness): harden tool receipts against forgery and turn-scope delegation_only

Address willem-bd's pre-merge review on #4659:

1. Untrusted receipt metadata: the gateway now strips the server-owned
   deerflow_tool_receipt key from external input messages; stamping always
   overwrites any tool-supplied value instead of preserving it; and
   extract_tool_receipts validates persisted receipt shapes (required typed
   fields, unknown keys ignored) so malformed entries are skipped instead of
   crashing render or passing as runtime-stamped evidence.

2. delegation_only no longer sticks on: _should_render now scopes the
   subagent_status scan to the current turn (messages after the latest
   genuine user message), so an old completed delegation stops rendering the
   ledger on later ordinary turns. The genuine-user predicate moves to
   message_utils.is_genuine_user_message, shared with input sanitization.

* fix(harness): stamp receipts outside short-circuiting tool middlewares

Address willem-bd's review on #4659: ToolReceiptMiddleware was registered
inside Guardrail/SandboxAudit/ReadBeforeWrite/ToolProgress, each of which
can return a ToolMessage without invoking its handler — blocked calls
(e.g. a read-before-write-denied write_file) never got a receipt, silently
gapping the ledger on a default-enabled path. SandboxAudit additionally
rebuilds medium-risk results, dropping an inner stamp.

ToolReceiptMiddleware is now the outermost wrap_tool_call layer in the
runtime tail. Normal results still carry deerflow_tool_meta (stamped by
ToolErrorHandling on the inner return path); short-circuit messages
self-stamp meta or fall back to message.status. The new invariant is
declared as ordering constraints in deerflow.extensions.ordering, with
composed-chain regression tests for a blocked write and a warn-rebuilt
bash result.
2026-08-23 15:43:37 +08:00

111 lines
5.2 KiB
Python

"""Declarative ordering invariants for the middleware stack.
Replaces hand-written index comparisons. Extension-contributed middlewares are
merged before validation runs, so a contribution cannot slip past an invariant,
and the failure names the extension responsible.
A broken invariant is the one hard failure in this system: unlike a missing
observation, it produces wrong behaviour without an error.
"""
from __future__ import annotations
from collections.abc import Mapping, Sequence
from dataclasses import dataclass
from functools import cache
from deerflow.extensions.isolation import IsolatedMiddleware
@dataclass(frozen=True)
class OrderingConstraint:
outer: type
inner: type
reason: str
def _indices_of(middlewares: Sequence[object], target: type) -> list[int]:
indices: list[int] = []
for index, middleware in enumerate(middlewares):
candidate = middleware.inner if isinstance(middleware, IsolatedMiddleware) else middleware
if isinstance(candidate, target):
indices.append(index)
return indices
def assert_ordering(
middlewares: Sequence[object],
provenance: Mapping[int, str],
constraints: Sequence[OrderingConstraint] | None = None,
) -> None:
"""Raise when a constraint is violated. No-op when both sides are absent."""
for constraint in constraints if constraints is not None else core_ordering_constraints():
outer_indices = _indices_of(middlewares, constraint.outer)
inner_indices = _indices_of(middlewares, constraint.inner)
if not outer_indices or not inner_indices:
continue
if max(outer_indices) < min(inner_indices):
continue
violating_indices = [index for index in outer_indices if index >= min(inner_indices)] + [index for index in inner_indices if index <= max(outer_indices)]
culprits = sorted({source for index in violating_indices if (source := provenance.get(index)) is not None})
blame = ", ".join(culprits) if culprits else "core middleware order"
raise RuntimeError(
f"Middleware ordering constraint violated: {constraint.outer.__name__} must be outer "
f"(lower index) of every {constraint.inner.__name__}, but found outer indices "
f"{outer_indices} vs inner indices {inner_indices}. Reason: {constraint.reason}. "
f"Contributed by: {blame}."
)
@cache
def core_ordering_constraints() -> tuple[OrderingConstraint, ...]:
"""The host's ordering invariants, resolved on first use.
Deferred deliberately, and the deferral is about dependency *direction*,
not just cycles: ``extensions/`` is the layer the middleware layer calls
into, so importing ``agents.middlewares`` at module scope here would point
the dependency backwards and close a cycle the moment any middleware
imports something under ``extensions/`` at module level. Resolution instead
happens at ``assert_ordering`` time, which already runs inside the
middleware builder — a forward reference within one layer.
Returns a plain tuple. The predecessor deferred by way of a ``tuple``
subclass overriding only ``__iter__``; because a tuple cannot populate its
own storage after construction, every operation reading that storage
(``len``, ``bool``, ``in``, indexing, slicing, ``reversed``, ``==``)
reported an empty sequence while iteration yielded the real constraints.
Deferring the call instead of faking the value keeps one answer.
"""
from deerflow.agents.middlewares.read_before_write_middleware import ReadBeforeWriteMiddleware
from deerflow.agents.middlewares.sandbox_audit_middleware import SandboxAuditMiddleware
from deerflow.agents.middlewares.tool_error_handling_middleware import ToolErrorHandlingMiddleware
from deerflow.agents.middlewares.tool_progress_middleware import ToolProgressMiddleware
from deerflow.agents.middlewares.tool_receipt_middleware import ToolReceiptMiddleware
from deerflow.guardrails.middleware import GuardrailMiddleware
return (
OrderingConstraint(
outer=ToolProgressMiddleware,
inner=ToolErrorHandlingMiddleware,
reason=("ToolProgressMiddleware reads deerflow_tool_meta in _update_state_from_result, so its wrap_tool_call chain must enclose the ToolErrorHandlingMiddleware step that stamps it"),
),
OrderingConstraint(
outer=ToolReceiptMiddleware,
inner=ToolErrorHandlingMiddleware,
reason=("ToolReceiptMiddleware reads the deerflow_tool_meta status stamped by ToolErrorHandlingMiddleware when building each receipt, so its wrap_tool_call chain must enclose the stamping step"),
),
*(
OrderingConstraint(
outer=ToolReceiptMiddleware,
inner=short_circuiter,
reason=(f"{short_circuiter.__name__} can return or rebuild a ToolMessage without invoking its handler; ToolReceiptMiddleware must wrap it or those results never get a receipt and the ledger silently gaps"),
)
for short_circuiter in (
GuardrailMiddleware,
SandboxAuditMiddleware,
ReadBeforeWriteMiddleware,
ToolProgressMiddleware,
)
),
)