deer-flow/backend/tests/test_config_version.py
Zeren Wang 4e35f0d1d4
feat(harness): deterministic tool receipts with model-visible ledger (RFC #4651, layer 1) (#4659)
* feat(harness): add deterministic tool receipts with model-visible ledger

Stamp an immutable per-call fact record (tool name, status, args/output
hashes, byte count, timestamp) onto every tool result via a new
ToolReceiptMiddleware, and inject the derived receipt ledger (r1..rN)
into the model context so subagent reports can cite executed actions.

- tool_receipt.py: receipt core (make/extract/render), newest-first
  budget eviction, ids derived from the append-only message stream
- ToolReceiptMiddleware: stamps ToolMessages directly or inside
  Command-wrapped results; hidden ledger injection mirrors
  DurableContextMiddleware; sits between ToolProgress and
  ToolErrorHandling with a build-time ordering guard
- config: new verification section (receipts on, judge off), config
  version 32 -> 33 with example/helm/docs updates

* feat(harness): split receipt rendering from stamping; address PR review

Review fixes (PR #4659):
- output_sha256 now uses sort_keys=True for structured content, matching
  the order-invariant args fingerprint
- stamping failures log at warning (silent ledger gaps would corrupt
  citations); tool execution remains never blocked
- _insert_after_leading_system_messages extracted to shared public
  message_utils.insert_after_leading_system_messages; both middlewares
  depend on it instead of a private cross-module helper
- code comments in English

RFC #4651 revision-2 alignment:
- receipts_render_mode config ('always' | 'delegation_only'): subagent
  chains always render the ledger (citations are produced there); the
  lead chain renders only while processing subagent results, removing
  the always-on token tax from ordinary turns
- receipts gain bounded args_preview/output_preview (<=200 chars, tail
  for output) so later typed claim bindings (tests_passed) can anchor
  to a specific recorded execution

* docs(harness): state receipt freshness caveat and vocabulary layering in module docstring

* merge: upstream/main — resolve AGENTS.md split, bump config_version to 34, drop unused receipt previews

- backend/AGENTS.md: take upstream's slimmed root guidance (#4799); move the
  ToolReceiptMiddleware chain entry into agents/middlewares/AGENTS.md and the
  verification.* hot-reload mention into config/AGENTS.md
- config.example.yaml + helm values/README: config_version 33 -> 34 so existing
  v33 configs get the outdated-config prompt (review: willem-bd)
- tool_receipt.py: drop args_preview/output_preview — no Layer 1 consumer reads
  them; re-add with the Layer 2 claim-binding consumer (review: willem-bd)

* docs(harness): cover receipt id renumbering after compaction in module docstring

Positional display ids are stable only while history is append-only;
compaction drops ToolMessages and the survivors renumber, so Layer 2
citation verification must resolve [rN] against the ledger as of the
citing turn (review: willem-bd, doc-only).

* chore(config): bump config_version to 35

main reached 34 via #4780 without the verification section; publishing
the new schema at the same number would silently skip the outdated-config
prompt for configs synced from main in that window (review: willem-bd).

* fix(skills): restore errno import dropped upstream in #4830

upstream/main adf6c422 uses errno.ENOTDIR in the drift guard but removed
the import, so the PR merge ref fails lint-backend (F821).

* fix(harness): harden tool receipts against forgery and turn-scope delegation_only

Address willem-bd's pre-merge review on #4659:

1. Untrusted receipt metadata: the gateway now strips the server-owned
   deerflow_tool_receipt key from external input messages; stamping always
   overwrites any tool-supplied value instead of preserving it; and
   extract_tool_receipts validates persisted receipt shapes (required typed
   fields, unknown keys ignored) so malformed entries are skipped instead of
   crashing render or passing as runtime-stamped evidence.

2. delegation_only no longer sticks on: _should_render now scopes the
   subagent_status scan to the current turn (messages after the latest
   genuine user message), so an old completed delegation stops rendering the
   ledger on later ordinary turns. The genuine-user predicate moves to
   message_utils.is_genuine_user_message, shared with input sanitization.

* fix(harness): stamp receipts outside short-circuiting tool middlewares

Address willem-bd's review on #4659: ToolReceiptMiddleware was registered
inside Guardrail/SandboxAudit/ReadBeforeWrite/ToolProgress, each of which
can return a ToolMessage without invoking its handler — blocked calls
(e.g. a read-before-write-denied write_file) never got a receipt, silently
gapping the ledger on a default-enabled path. SandboxAudit additionally
rebuilds medium-risk results, dropping an inner stamp.

ToolReceiptMiddleware is now the outermost wrap_tool_call layer in the
runtime tail. Normal results still carry deerflow_tool_meta (stamped by
ToolErrorHandling on the inner return path); short-circuit messages
self-stamp meta or fall back to message.status. The new invariant is
declared as ordering constraints in deerflow.extensions.ordering, with
composed-chain regression tests for a blocked write and a warn-rebuilt
bash result.
2026-08-23 15:43:37 +08:00

243 lines
9.6 KiB
Python

"""Tests for config version check and upgrade logic."""
from __future__ import annotations
import logging
import os
import tempfile
from pathlib import Path
import yaml
from deerflow.config.app_config import AppConfig
def _make_config_files(tmpdir: Path, user_config: dict, example_config: dict) -> Path:
"""Write user config.yaml and config.example.yaml to a temp dir, return config path."""
config_path = tmpdir / "config.yaml"
example_path = tmpdir / "config.example.yaml"
# Minimal valid config needs sandbox
defaults = {
"sandbox": {"use": "deerflow.sandbox.local:LocalSandboxProvider"},
}
for cfg in (user_config, example_config):
for k, v in defaults.items():
cfg.setdefault(k, v)
with open(config_path, "w", encoding="utf-8") as f:
yaml.dump(user_config, f)
with open(example_path, "w", encoding="utf-8") as f:
yaml.dump(example_config, f)
return config_path
def test_missing_version_treated_as_zero(caplog):
"""Config without config_version should be treated as version 0."""
with tempfile.TemporaryDirectory() as tmpdir:
config_path = _make_config_files(
Path(tmpdir),
user_config={}, # no config_version
example_config={"config_version": 1},
)
with caplog.at_level(logging.WARNING, logger="deerflow.config.app_config"):
AppConfig._check_config_version(
{"sandbox": {"use": "deerflow.sandbox.local:LocalSandboxProvider"}},
config_path,
)
assert "outdated" in caplog.text
assert "version 0" in caplog.text
assert "version is 1" in caplog.text
def test_matching_version_no_warning(caplog):
"""Config with matching version should not emit a warning."""
with tempfile.TemporaryDirectory() as tmpdir:
config_path = _make_config_files(
Path(tmpdir),
user_config={"config_version": 1},
example_config={"config_version": 1},
)
with caplog.at_level(logging.WARNING, logger="deerflow.config.app_config"):
AppConfig._check_config_version(
{"config_version": 1},
config_path,
)
assert "outdated" not in caplog.text
def test_outdated_version_emits_warning(caplog):
"""Config with lower version should emit a warning."""
with tempfile.TemporaryDirectory() as tmpdir:
config_path = _make_config_files(
Path(tmpdir),
user_config={"config_version": 1},
example_config={"config_version": 2},
)
with caplog.at_level(logging.WARNING, logger="deerflow.config.app_config"):
AppConfig._check_config_version(
{"config_version": 1},
config_path,
)
assert "outdated" in caplog.text
assert "version 1" in caplog.text
assert "version is 2" in caplog.text
def test_no_example_file_no_warning(caplog):
"""If config.example.yaml doesn't exist, no warning should be emitted."""
with tempfile.TemporaryDirectory() as tmpdir:
config_path = Path(tmpdir) / "config.yaml"
with open(config_path, "w", encoding="utf-8") as f:
yaml.dump({"sandbox": {"use": "test"}}, f)
# No config.example.yaml created
with caplog.at_level(logging.WARNING, logger="deerflow.config.app_config"):
AppConfig._check_config_version({}, config_path)
assert "outdated" not in caplog.text
def test_string_config_version_does_not_raise_type_error(caplog):
"""config_version stored as a YAML string should not raise TypeError on comparison."""
with tempfile.TemporaryDirectory() as tmpdir:
config_path = _make_config_files(
Path(tmpdir),
user_config={"config_version": "1"}, # string, as YAML can produce
example_config={"config_version": 2},
)
# Must not raise TypeError: '<' not supported between instances of 'str' and 'int'
AppConfig._check_config_version({"config_version": "1"}, config_path)
def test_newer_user_version_no_warning(caplog):
"""If user has a newer version than example (edge case), no warning."""
with tempfile.TemporaryDirectory() as tmpdir:
config_path = _make_config_files(
Path(tmpdir),
user_config={"config_version": 3},
example_config={"config_version": 2},
)
with caplog.at_level(logging.WARNING, logger="deerflow.config.app_config"):
AppConfig._check_config_version(
{"config_version": 3},
config_path,
)
assert "outdated" not in caplog.text
def test_version_26_config_upgrades_to_checkpoint_channel_mode(tmp_path, caplog):
"""A v26 user config must be flagged outdated and merge the new persisted field.
`database.checkpoint_channel_mode` shipped with config_version 27; the
upgrade path must add it with the safe default (``full``) without touching
the user's existing database backend settings. Uses the repository's real
config.example.yaml and the real config-upgrade script.
"""
import subprocess
repo_root = Path(__file__).resolve().parents[2]
example_src = repo_root / "config.example.yaml"
example_data = yaml.safe_load(example_src.read_text(encoding="utf-8"))
expected_version = example_data["config_version"]
assert expected_version > 26, "config.example.yaml must be bumped past 26 for checkpoint_channel_mode"
config_path = tmp_path / "config.yaml"
(tmp_path / "config.example.yaml").write_text(example_src.read_text(encoding="utf-8"), encoding="utf-8")
user_config = {
"config_version": 26,
"sandbox": {"use": "deerflow.sandbox.local:LocalSandboxProvider"},
"database": {"backend": "sqlite", "sqlite_dir": "custom-data"},
}
config_path.write_text(yaml.dump(user_config), encoding="utf-8")
with caplog.at_level(logging.WARNING, logger="deerflow.config.app_config"):
AppConfig._check_config_version(dict(user_config), config_path)
assert "outdated" in caplog.text
assert "(version 26)" in caplog.text
env = {**os.environ, "DEER_FLOW_CONFIG_PATH": str(config_path)}
result = subprocess.run(
["bash", str(repo_root / "scripts" / "config-upgrade.sh")],
env=env,
capture_output=True,
text=True,
timeout=120,
)
assert result.returncode == 0, result.stderr
upgraded = yaml.safe_load(config_path.read_text(encoding="utf-8"))
assert upgraded["config_version"] == expected_version
assert upgraded["database"]["checkpoint_channel_mode"] == "full"
assert upgraded["database"]["backend"] == "sqlite"
assert upgraded["database"]["sqlite_dir"] == "custom-data"
assert upgraded["verification"]["receipts_enabled"] is True
assert upgraded["verification"]["receipts_render_mode"] == "delegation_only"
assert upgraded["verification"]["judge_enabled"] is False
assert upgraded["verification"]["judge_model_name"] is None
def _load_repo_example() -> dict:
"""Load the real repo config.example.yaml (first-run template)."""
example_path = Path(__file__).resolve().parents[2] / "config.example.yaml"
with open(example_path, encoding="utf-8") as f:
return yaml.safe_load(f) or {}
def _merge_missing(target: dict, source: dict) -> None:
"""Add-missing-keys-only recursive merge mirroring scripts/config-upgrade.sh."""
for key, value in source.items():
if key not in target:
import copy
target[key] = copy.deepcopy(value)
elif isinstance(value, dict) and isinstance(target[key], dict):
_merge_missing(target[key], value)
def test_security_fail_closed_bumped_config_version():
"""The example must ship security_fail_closed under a version > 26 so v26 configs upgrade."""
example = _load_repo_example()
assert example.get("config_version", 0) >= 27
assert example["skill_evolution"]["security_fail_closed"] is True
def test_version_26_config_reported_outdated_against_example(caplog):
"""A version-26 user config is flagged outdated against the real example version."""
example = _load_repo_example()
example_version = example["config_version"]
with tempfile.TemporaryDirectory() as tmpdir:
config_path = _make_config_files(
Path(tmpdir),
user_config={"config_version": 26},
example_config=example,
)
with caplog.at_level(logging.WARNING, logger="deerflow.config.app_config"):
AppConfig._check_config_version({"config_version": 26}, config_path)
assert "outdated" in caplog.text
assert "version 26" in caplog.text
assert f"version is {example_version}" in caplog.text
def test_config_upgrade_adds_security_fail_closed_preserving_user_values():
"""config-upgrade merges security_fail_closed: true without touching existing skill_evolution values."""
example = _load_repo_example()
# A version-26 user who customized skill_evolution but predates the new field.
user = {
"config_version": 26,
"skill_evolution": {
"enabled": True,
"moderation_model_name": "custom-moderation-model",
},
}
_merge_missing(user, example)
user["config_version"] = example["config_version"]
# New persisted field is merged in with the example's fail-closed default.
assert user["skill_evolution"]["security_fail_closed"] is True
# The user's existing skill_evolution values are preserved unchanged.
assert user["skill_evolution"]["enabled"] is True
assert user["skill_evolution"]["moderation_model_name"] == "custom-moderation-model"
assert user["config_version"] == example["config_version"]