mirror of
https://github.com/bytedance/deer-flow.git
synced 2026-09-22 04:26:18 +00:00
7 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
c56c7293f8
|
test(checkpoint): measure postgres storage growth (#5051) | ||
|
|
2eba65449f
|
eval(memory): add a reproducible hybrid eviction evaluation (#4810)
* eval(memory): scaffold reproducible eviction evaluation * refactor(eval): align with benchmark layout * eval(memory): add deterministic QA grading Implement the disclosed deterministic-overlap-v1 grader as a pure offline module. Grading is blind by construction: grade_answer() accepts only the prediction and reference strings, never a policy identity. The undisclosed stopword list is committed as a fixed part of this grader version; yes/no/not are deliberately excluded because negation can be the entire answer. Before freezing, the grader locally reproduced all 90 historical (prediction, grade) pairs disclosed in #4789 with zero mismatches and no post-hoc tuning. validate-contracts now rejects a config whose qa.grader_version does not match the committed grader. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * eval(memory): add environment-configured QA runner Add the exact answer-prompt renderer (retained facts sorted by ID, CURRENT DATE line omitted when absent), an OpenAI-compatible provider adapter configured only through the environment variable names pinned in the config, and a resumable run-qa command that calls both policies with identical versioned settings. Each row persists as its own file on success, so a partial paid run resumes without repeating completed calls; qa_run.json binds an output directory to one config identity. Row files and errors carry predictions and non-secret metadata only -- never questions, references, memory text, credentials, or response headers. All tests are offline via mocked transports; run-qa fails fast before touching the dataset when the provider environment is missing. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * eval(memory): add blind QA grading report and paired statistics Add grade-qa: it recomputes the deterministic selector output, rejects any answer row whose kept facts, capacity, or policy disagree with it, grades every prediction through the policy-blind grade_answer(prediction, reference) call, and only then joins grades back through stable row IDs. Published artifacts are qa.rows.jsonl (graded rows with non-secret metadata), qa.summary.json (accuracy by source/scenario/policy; official and synthetic suites never folded together), and qa.stats.json (exact paired McNemar and seeded paired bootstrap difference for the official, synthetic, and overall suites using the pinned statistics parameters). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * eval(memory): pin the official DeepSeek model ID The historical protocol recorded the answer model with an aggregator-style namespace (deepseek/deepseek-v4-flash). The live run calls the same underlying model (DeepSeek-V4-Flash-0731, released before the historical run) directly through DeepSeek's official OpenAI-compatible API, whose canonical ID is deepseek-v4-flash. The served model is recorded from the provider response in every answer row. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * eval(memory): publish paired eviction QA results Publish the equal-budget live QA artifacts for pr4789-reproduction-v1: provenance, 90 graded rows, per-scenario summary, and paired statistics. At capacity 7 with identical settings, confidence answers 24/45 and hybrid-v1 40/45 (official 23/40 vs 35/40, exact McNemar p=0.0042; overall p=0.0004). The noisy-signal control is the one scenario where hybrid-v1 scored below the baseline (10/10 vs 8/10) and is reported separately. The offline suite now verifies the published statistics are recomputable from the published rows and that the artifacts carry no dataset text or credentials. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(eval): decompose the noisy-signal QA cell Both policies retained the support fact in all ten noisy-signal cases, so the two rows hybrid-v1 lost are grader phrasing boundaries (verbose numeric answers rejected by the numeric-conflict rule), not eviction failures. Documented from the published rows; the grader stays frozen. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * eval(memory): address review hardening findings - ignore the responses/ directory the runner actually writes instead of the stale provider-responses/ entry - cover the official selection-rule recomputation with direct synthetic tests: matching manifests pass, rule-breaking IDs and missing eligible rows fail, and every published exclusion is load-bearing - align the report docstring and README with the statistics contract: the summary never folds sources; the explicitly labeled overall suite is reported alongside the separate official and synthetic suites - recompute the published bootstrap intervals (not only McNemar) in the published-results test - wire required_policy_version to the production EVICTION_POLICY_HYBRID_V1 constant so validate-contracts rejects policy drift Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * eval(memory): restore historical evidence rendering and harden resume identity Address both blocking findings from the #4789 artifact cross-check. The evidence renderer now emits the historical SESSION {id} AT {date} line instead of the divergent bracket format. The byte representation is protocol-critical: the witness record 35a27287 renders at 697 characters again, stays inside the 700-character distractor-bank bound, and 60d45044 leaves the bank, restoring row-level pool reproduction. Deterministic capacity-7 retention is unchanged at 27/45 vs 45/45. qa_run.json now binds a run directory to the SHA-256 of all five protocol inputs (config, both manifests, answer prompt, dataset) and names the changed artifact when it refuses to resume. Stored rows are reused only when row identity, policy, capacity, kept facts, and the request fingerprint recomputed from the current task all match; the disclosed probe (changed message under the same config) is now a regression test. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * eval(memory): republish QA results under the historical protocol Replace the published artifacts with the fresh equal-budget run executed at 497ff3d0 under the restored historical evidence rendering; the earlier run under the divergent rendering is discarded entirely. At capacity 7 with identical settings, confidence answers 24/45 and hybrid-v1 38/45 (official 23/40 vs 33/40, exact McNemar p=0.0129; overall p=0.0013). The confidence control is the one scenario below baseline for hybrid-v1 (8/10 vs 6/10); both lost rows retained the support fact and are grader phrasing/abstention boundaries, documented from the published rows. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * eval(memory): adopt the historical fact IDs and prompt serialization Pool facts now carry the historical protocol IDs (gold_{case} for the support fact, d_{case}_{index}_{source} for distractors in bank-draw order), and the rendered STORED MEMORY joins fact blocks with a blank line. Sorting by these IDs reproduces the historical selection tie-break: witness case 41698283 at capacity 7 again keeps the 58bf7951 distractor and evicts 001be529 under both policies. Deterministic capacity-7 retention is unchanged at 27/45 vs 45/45. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * eval(memory): republish QA results under the historical serialization Replace the published artifacts with the fresh equal-budget run executed at 01f99d61 under the historical fact IDs and prompt serialization; earlier runs under divergent serializations are discarded entirely. At capacity 7 with identical settings, confidence answers 24/45 and hybrid-v1 40/45 (official 23/40 vs 35/40, exact McNemar p=0.0018; overall p=0.0001). The single row below baseline (1cea1afa, confidence-control) retained its support fact; the model abstained. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * eval(memory): self-certify the publish path and pin offset coverage grade-qa now verifies (read-only) that the run marker's five protocol artifact hashes match the current inputs, rebuilds every answer task, and rejects any stored row whose request fingerprint does not match the task recomputed from the current protocol — the staleness class that previously required an out-of-band cross-check to detect. Verified end-to-end against the published run: all 90 rows pass and regrade to byte-identical artifacts, while a tampered fingerprint is refused by row ID. The distractor offset derivation and wraparound selection are now pinned by unit tests with hardcoded indices, including a wrapping offset, so a digest-slice or modulus regression can no longer stay green offline. Closes both non-blocking suggestions from the re-review. * fix(bench): bind persisted answer rows to their expected case identity Grading derived the reference case from the stored row's embedded case_id, so reassigning a valid row to another valid case passed every integrity check while silently changing the published grade. The resume path had the same gap: _row_matches_task() never compared case_id, source, or scenario. The recomputed task is now authoritative in both paths: grade_answer_rows() resolves the reference case from the expected PolicyResult and rejects any mismatch in the persisted row_id/case_id/source/scenario, and _row_matches_task() checks the same identity fields so a reassigned row is re-run instead of reused. Regressions tamper each field individually and exercise both paths. * docs(bench): document where to download the pinned LongMemEval file The README named the dataset but never said it lives on Hugging Face or how to fetch the pinned revision, so a reviewer could not run the offline commands. Add the direct download URL, the expected SHA-256, and the mirror and huggingface-cli alternatives; the CLI still never downloads anything. --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
c8cf1bf2fb
|
feat(checkpoint): checkpoint history cache (#4638)
* feat(checkpoint-cache): delta-mode checkpoint history cache with recursive compose
Read-only, invalidation-free cache for LangGraph delta-channel history
({writes, seed}) at the get_delta_channel_history choke point:
- database.checkpoint_cache config (memory|redis; max_entries 0=disabled;
redis bounded by TTL, Gateway/async only)
- memory LRU backend (copy-on-read, zero-serde hit path) and redis backend
(lazy import, degrades to all-miss on outage)
- CachedHistorySaver: recursive composition from the nearest warm ancestor
(depth budget 8), caching each level; depth-0 cold chains delegate one
inner fast-path walk. Entries keyed by immutable
(db, thread, ns, checkpoint_id, channel) — no invalidation, coherent
across workers
- provider wiring: wraps in delta mode only (async + sync), full mode
untouched; sync path is memory-only
- bench opt-in: DEERFLOW_CHECKPOINT_BENCH_HISTORY_CACHE=1
sqlite bench (500 updates, payload 2KB): write phase 2.28x at f=250,
1.32x at f=10; one delegated walk per thread cold start.
* chore(config): bump config_version to 32 for database.checkpoint_cache
The checkpoint history cache feature added the database.checkpoint_cache
section to config.example.yaml; bump the schema version so existing
deployments get the outdated-config warning and can run make config-upgrade.
* chore(helm): bump config_version to 32 in chart values and README
* fix(checkpoint-cache): purge thread history entries on delete paths
Addresses review on #4638: delete_thread/prune removed source-of-truth
checkpoints but left the thread's materialized history payloads in the
cache (memory: until LRU eviction; redis: until TTL, default 1 day) — a
data-lifecycle gap for tenant offboarding / GDPR-style erasure.
- Cache contract gains thread-scoped adelete_thread/delete_thread
(lifecycle purge, not invalidation; entries remain immutable)
- Memory backend: stem scan over the LRU map; redis: SCAN MATCH + UNLINK,
outage degrades to TTL-bounded retention without raising
- CachedHistorySaver purges on delete_thread/adelete_thread and
prune/aprune (prune rewrites chains, so pre-prune histories must go);
delete_for_runs stays delegation-only (run->thread mapping unavailable,
no in-tree callers), documented in code
- ttl_seconds description documents the residual-retention window
- Tests: thread-scoped purge on both backends, saver-level delete/prune
purge, prefix-safety (t1 vs t10), redis outage degradation, and the
pinned no-purge behavior of delete_for_runs
* fix(checkpoint-cache): stable db identity, prefix-aware sync singleton, explicit zero TTL
Addresses Copilot review on #4638:
- checkpoint_cache_db_hash now hashes the credential-free postgres
identity (host:port/database + schema): credential rotation no longer
changes the cache namespace (cold cache + orphaned keys until TTL).
Unparseable URLs fall back to the raw string.
- The sync-path memory cache singleton is also keyed by its key_prefix:
a namespace change (db identity change or operator override) recreates
the cache instead of leaving stale-prefix entries unreachable and
unpurgeable.
- ttl_seconds=0 is now an explicit, documented opt-out of redis expiry
(SET without EX; redis maxmemory policy only) instead of a silent
'ttl_seconds or None' coercion.
Tests: credential-rotation hash stability, unparseable-URL fallback,
prefix-change singleton recreation, same-prefix singleton reuse, and
zero-TTL wire behavior (ex=None).
---------
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
|
||
|
|
c48de5e70b
|
feat(checkpoint): make delta snapshot_frequency configurable (#4516)
* feat(checkpoint): make delta snapshot_frequency configurable * fix(config): carry legacy checkpoint_delta_snapshot_frequency with warning Addresses review on #4516: the rename from the flat database.checkpoint_delta_snapshot_frequency key to nested database.checkpoint_delta.snapshot_frequency silently dropped the old value (pydantic extra="ignore"). Add a before-validator that maps the legacy key onto the nested one with a deprecation warning (nested key wins when both are set), plus a CHANGELOG breaking-change note covering the rename and the 1000 -> 10 default change. * fix(checkpoint): validate frozen snapshot frequency |
||
|
|
e01173d8b2
|
bench(checkpoint): production-shaped full/delta benchmark with configurable snapshot frequency (#4467)
* feat(checkpoint): production-shaped full/delta benchmark with configurable snapshot frequency - Group benchmark scripts into per-family folders (checkpoint/, sandbox/) - Extract shared benchmark infrastructure into checkpoint_bench_common.py - Add checkpoint_delta_snapshot_frequency config (default 1000, process-frozen); freeze it in make_lead_agent and DeerFlowClient; key the state-schema adaptation cache by resolved frequency - New bench_production.py: per-case child processes run N ainvoke turns through the real lead-agent graph (scripted deterministic model, real AsyncSqliteSaver), then measure GET /state + POST /history through the real Gateway route stack in one event loop (httpx ASGITransport), cold/warm accessor-cache split, cross-mode digest gates - New summarize_production.py: delta/full ratios plus decision metrics (snapshot_write_spike, cache_effect_ms, checkpoint_write_share, auto-discovered history per-limit ratios) * fix(checkpoint): address production benchmark review |
||
|
|
62dd8d2b67
|
bench(checkpoint): add channel mode benchmark (#4395)
* bench(checkpoint): add channel mode benchmark * bench(checkpoint): harden benchmark reporting |
||
|
|
dc38d2d003
|
perf(boxlite): add benchmark-driven warm-pool reclaim tuning (#4001)
* bench: add provider-agnostic sandbox benchmark with BoxLite warm pool results - scripts/bench_sandbox_provider.py: CLI for measuring acquire/run/release across providers, scenarios, workloads, and concurrency levels - scripts/summarize_bench.py: JSONL aggregation with p50/p95/p99 tables - bench_results.jsonl: 110 turns across 7 scenarios on real BoxLite 0.9.7 Key findings: cold acquire: ~860ms warm reclaim: ~14ms (60x speedup) release: ~0ms warm_hit_rate: 95% (warm_same_thread) * perf(boxlite): skip health check for recently-released warm pool boxes Boxes released within health_check_skip_seconds (default 5.0s) are promoted directly without the ~14ms echo-ok round-trip. A VM alive seconds ago is overwhelmingly likely to still be alive. Add sandbox.health_check_skip_seconds config option. Set to 0 to always health-check (old behaviour). Benchmark (warm_same_thread, noop, 20 iters): acquire p50: 14.9ms → 0.0ms total p50: 29.9ms → 14.0ms * chore: move benchmark scripts into backend/scripts/benchmark/ * fix: address BoxLite benchmark review findings * fix(boxlite): only skip warm reclaim checks for released boxes * fix(benchmark): keep BoxLite shim workaround off the event loop * fix(boxlite): invalidate dead boxes from command path * test(boxlite): cover skip window and invalidation edge cases * fix(boxlite): treat sandbox-has-been-closed as terminal in _exec * fix(boxlite): harden warm-pool reclaim and benchmark accounting * fix(boxlite): validate warm-pool reclaims by default * fix(config): expose boxlite health-check skip setting * fix(boxlite): tighten failure classification and benchmark workaround * Update config_version to 21 in values.yaml --------- Co-authored-by: Willem Jiang <willem.jiang@gmail.com> |