deer-flow/backend/tests/test_tool_receipt.py
Zeren Wang 4e35f0d1d4
feat(harness): deterministic tool receipts with model-visible ledger (RFC #4651, layer 1) (#4659)
* feat(harness): add deterministic tool receipts with model-visible ledger

Stamp an immutable per-call fact record (tool name, status, args/output
hashes, byte count, timestamp) onto every tool result via a new
ToolReceiptMiddleware, and inject the derived receipt ledger (r1..rN)
into the model context so subagent reports can cite executed actions.

- tool_receipt.py: receipt core (make/extract/render), newest-first
  budget eviction, ids derived from the append-only message stream
- ToolReceiptMiddleware: stamps ToolMessages directly or inside
  Command-wrapped results; hidden ledger injection mirrors
  DurableContextMiddleware; sits between ToolProgress and
  ToolErrorHandling with a build-time ordering guard
- config: new verification section (receipts on, judge off), config
  version 32 -> 33 with example/helm/docs updates

* feat(harness): split receipt rendering from stamping; address PR review

Review fixes (PR #4659):
- output_sha256 now uses sort_keys=True for structured content, matching
  the order-invariant args fingerprint
- stamping failures log at warning (silent ledger gaps would corrupt
  citations); tool execution remains never blocked
- _insert_after_leading_system_messages extracted to shared public
  message_utils.insert_after_leading_system_messages; both middlewares
  depend on it instead of a private cross-module helper
- code comments in English

RFC #4651 revision-2 alignment:
- receipts_render_mode config ('always' | 'delegation_only'): subagent
  chains always render the ledger (citations are produced there); the
  lead chain renders only while processing subagent results, removing
  the always-on token tax from ordinary turns
- receipts gain bounded args_preview/output_preview (<=200 chars, tail
  for output) so later typed claim bindings (tests_passed) can anchor
  to a specific recorded execution

* docs(harness): state receipt freshness caveat and vocabulary layering in module docstring

* merge: upstream/main — resolve AGENTS.md split, bump config_version to 34, drop unused receipt previews

- backend/AGENTS.md: take upstream's slimmed root guidance (#4799); move the
  ToolReceiptMiddleware chain entry into agents/middlewares/AGENTS.md and the
  verification.* hot-reload mention into config/AGENTS.md
- config.example.yaml + helm values/README: config_version 33 -> 34 so existing
  v33 configs get the outdated-config prompt (review: willem-bd)
- tool_receipt.py: drop args_preview/output_preview — no Layer 1 consumer reads
  them; re-add with the Layer 2 claim-binding consumer (review: willem-bd)

* docs(harness): cover receipt id renumbering after compaction in module docstring

Positional display ids are stable only while history is append-only;
compaction drops ToolMessages and the survivors renumber, so Layer 2
citation verification must resolve [rN] against the ledger as of the
citing turn (review: willem-bd, doc-only).

* chore(config): bump config_version to 35

main reached 34 via #4780 without the verification section; publishing
the new schema at the same number would silently skip the outdated-config
prompt for configs synced from main in that window (review: willem-bd).

* fix(skills): restore errno import dropped upstream in #4830

upstream/main adf6c422 uses errno.ENOTDIR in the drift guard but removed
the import, so the PR merge ref fails lint-backend (F821).

* fix(harness): harden tool receipts against forgery and turn-scope delegation_only

Address willem-bd's pre-merge review on #4659:

1. Untrusted receipt metadata: the gateway now strips the server-owned
   deerflow_tool_receipt key from external input messages; stamping always
   overwrites any tool-supplied value instead of preserving it; and
   extract_tool_receipts validates persisted receipt shapes (required typed
   fields, unknown keys ignored) so malformed entries are skipped instead of
   crashing render or passing as runtime-stamped evidence.

2. delegation_only no longer sticks on: _should_render now scopes the
   subagent_status scan to the current turn (messages after the latest
   genuine user message), so an old completed delegation stops rendering the
   ledger on later ordinary turns. The genuine-user predicate moves to
   message_utils.is_genuine_user_message, shared with input sanitization.

* fix(harness): stamp receipts outside short-circuiting tool middlewares

Address willem-bd's review on #4659: ToolReceiptMiddleware was registered
inside Guardrail/SandboxAudit/ReadBeforeWrite/ToolProgress, each of which
can return a ToolMessage without invoking its handler — blocked calls
(e.g. a read-before-write-denied write_file) never got a receipt, silently
gapping the ledger on a default-enabled path. SandboxAudit additionally
rebuilds medium-risk results, dropping an inner stamp.

ToolReceiptMiddleware is now the outermost wrap_tool_call layer in the
runtime tail. Normal results still carry deerflow_tool_meta (stamped by
ToolErrorHandling on the inner return path); short-circuit messages
self-stamp meta or fall back to message.status. The new invariant is
declared as ordering constraints in deerflow.extensions.ordering, with
composed-chain regression tests for a blocked write and a warn-rebuilt
bash result.
2026-08-23 15:43:37 +08:00

121 lines
5.1 KiB
Python

"""Tests for tool receipt core (deterministic verification layer)."""
from __future__ import annotations
from langchain_core.messages import AIMessage, ToolMessage
from deerflow.agents.middlewares.tool_receipt import (
TOOL_RECEIPT_KEY,
extract_tool_receipts,
make_tool_receipt,
render_tool_receipts,
)
from deerflow.agents.middlewares.tool_result_meta import TOOL_META_KEY
def _msg(content: str, *, tool_call_id: str, name: str = "write_file", meta_status: str = "success") -> ToolMessage:
return ToolMessage(
content=content,
tool_call_id=tool_call_id,
name=name,
additional_kwargs={TOOL_META_KEY: {"status": meta_status}},
)
def _stamped_msg(content: str, *, tool_call_id: str, name: str, args: dict | None = None) -> ToolMessage:
message = _msg(content, tool_call_id=tool_call_id, name=name)
receipt = make_tool_receipt({"name": name, "id": tool_call_id, "args": args or {}}, message)
message.additional_kwargs[TOOL_RECEIPT_KEY] = receipt
return message
def test_make_tool_receipt_hashes_args_and_output():
receipt = make_tool_receipt(
{"name": "write_file", "id": "tc-1", "args": {"path": "/tmp/a.txt", "content": "hello"}},
_msg("ok", tool_call_id="tc-1"),
)
assert receipt["tool_call_id"] == "tc-1"
assert receipt["tool_name"] == "write_file"
assert receipt["status"] == "success"
assert len(receipt["args_sha256"]) == 16
assert len(receipt["output_sha256"]) == 16
assert receipt["output_bytes"] == 2
def test_make_tool_receipt_args_hash_is_key_order_invariant():
first = make_tool_receipt({"name": "t", "id": "x", "args": {"a": 1, "b": 2}}, _msg("r", tool_call_id="x", name="t"))
second = make_tool_receipt({"name": "t", "id": "x", "args": {"b": 2, "a": 1}}, _msg("r", tool_call_id="x", name="t"))
assert first["args_sha256"] == second["args_sha256"]
def test_make_tool_receipt_uses_meta_error_status():
receipt = make_tool_receipt(
{"name": "web_fetch", "id": "tc-2", "args": {"url": "https://x"}},
_msg("Error: 404", tool_call_id="tc-2", name="web_fetch", meta_status="error"),
)
assert receipt["status"] == "error"
def test_extract_assigns_sequential_ids_and_skips_unstamped():
messages = [
AIMessage(content="working", tool_calls=[{"name": "bash", "id": "tc-9", "args": {}}]),
_msg("unstamped", tool_call_id="tc-0", name="bash"),
_stamped_msg("first", tool_call_id="tc-1", name="write_file", args={"path": "/tmp/a"}),
_stamped_msg("second", tool_call_id="tc-2", name="bash"),
]
receipts = extract_tool_receipts(messages)
assert [r["id"] for r in receipts] == ["r1", "r2"]
assert receipts[0]["tool_name"] == "write_file"
assert receipts[1]["tool_name"] == "bash"
def test_render_empty_and_budget():
assert render_tool_receipts([]) == ""
receipts = extract_tool_receipts([_stamped_msg("ok", tool_call_id="tc-1", name="write_file", args={"path": "/tmp/a"})])
text = render_tool_receipts(receipts)
assert "r1" in text and "write_file" in text and "success" in text
# Anti-automation-bias (design rule 4): the ledger must always carry its evidence-boundary statement
assert "do not validate claim correctness" in text
assert len(render_tool_receipts(receipts, max_chars=10)) <= 14 # truncated + "\n..."
def test_render_budget_keeps_newest_receipts_with_original_ids():
receipts = extract_tool_receipts([_stamped_msg(f"result-{index}", tool_call_id=f"tc-{index}", name=f"tool-{index}") for index in range(1, 13)])
text = render_tool_receipts(receipts, max_chars=500)
assert len(text) <= 500
assert "[r12] tool-12" in text
assert "[r1] tool-1" not in text
assert "older receipts omitted" in text
def test_extract_skips_malformed_receipts():
"""Persisted/foreign receipt payloads must not crash or enter the ledger."""
good = _stamped_msg("ok", tool_call_id="tc-good", name="bash")
malformed = []
for payload in [
"not-a-dict",
{}, # missing every field
{"tool_call_id": "tc-1"}, # partial shape
{**make_tool_receipt({"name": "t", "id": "tc-2", "args": {}}, _msg("x", tool_call_id="tc-2", name="t")), "output_bytes": "2"}, # wrong type
]:
message = _msg("bad", tool_call_id="tc-bad", name="bash")
message.additional_kwargs[TOOL_RECEIPT_KEY] = payload
malformed.append(message)
# A future-schema receipt (extra keys, valid core shape) must not crash
# extraction either — its known fields are picked, unknown keys ignored.
forward_compat = _msg("newer", tool_call_id="tc-newer", name="bash")
forward_compat.additional_kwargs[TOOL_RECEIPT_KEY] = {
**make_tool_receipt({"name": "bash", "id": "tc-newer", "args": {}}, forward_compat),
"layer2_field": {"nested": True},
}
receipts = extract_tool_receipts([*malformed, good, forward_compat])
assert [r["tool_call_id"] for r in receipts] == ["tc-good", "tc-newer"]
assert [r["id"] for r in receipts] == ["r1", "r2"]
# And the render path never sees a shape it can KeyError on.
rendered = render_tool_receipts(receipts)
assert "[r1] bash" in rendered and "[r2] bash" in rendered