* feat(memory): signals-based update pipeline + always-on watermark/trivial filter
Refactor the DeerMem memory update pipeline (message_processing -> queue ->
updater) around a signals frozenset seam, replacing the
(filtered, correction_detected, reinforcement_detected) 3-tuple with
(filtered, signals: frozenset[str]) end to end.
message_processing:
- Externalize signal-detection patterns to YAML (message_patterns/*.yaml).
- Extend signals from correction/reinforcement to a 6-class set
(correction/reinforcement/preference/identity/goal/decision); detect_signals
returns a frozenset aligned with the fact category enum.
- Pure-acknowledgment turns ("ok"/"好的"/...) are always filtered out before
enqueue (whole-message fullmatch), saving an extraction LLM call.
queue (core/queue.py):
- In-memory list + debounce timer, with flush_sync (graceful-shutdown drain
that joins an in-flight worker under a hard timeout) and queue_max_depth
backpressure (signal-bearing updates always admitted; QueueFull otherwise).
- Same-key updates coalesce with a signal union; per-batch success/fail summary.
updater (core/updater.py):
- head500+tail500 message truncation (replaces the 1000-char head chop).
- Always-on per-thread watermark: feed only messages added since the last
extraction. The watermark is in-memory and is not advanced on failure, so a
failed/lost update is re-fed on the next conversation turn.
- [MANUAL] prompt marker for user-authored facts (source.type="manual").
- Post-invoke extraction_callback (host-injected) emitting facts_extracted /
facts_accepted / rejected_low_confidence; the host default logs metrics and
flags >60% rejection.
Confidence filtering remains in _apply_updates (the existing
fact_confidence_threshold check); there is no separate write gate.
Consolidation stays opt-in (lossy). The ABC add/add_nowait signature is
unchanged, so the summarization flush hook and host are unaffected.
Tests: add test_message_processing_signals, test_updater_truncation,
test_updater_watermark; update queue/updater/consolidation/staleness/pluggable
tests for the signals seam.
Co-Authored-By: Claude <noreply@anthropic.com>
* fix(memory): harden update pipeline per PR review
- Catch QueueFull in DeerMem.add/add_nowait so backpressure degrades to
'update skipped' instead of propagating into after_agent /
summarization_hook and breaking the agent run (peer middlewares
self-guard; MemoryMiddleware was the lone exception). Emergency
(add_nowait) always admits under backpressure -- its data cannot be
re-fed next turn.
- Rewrite the watermark from index-based to content/identity-based
(_message_identity + _feed_after_watermark) so it stays correct when
summarization removes the conversation front -- an index watermark
pointed at the wrong message and silently skipped un-extracted tail
turns. The emergency flush bypasses the watermark (bypass_watermark on
ConversationContext, threaded through update_memory) and coexists with
(does not replace) a pending normal update, so a flush cannot drop a
pending update's un-extracted tail.
- Populate facts_accepted / rejected_low_confidence inside _apply_updates
at the real confidence-filter site (passed_threshold) instead of
re-deriving the threshold in _finalize_update -- eliminates metric drift.
- Emit extraction metrics in a finally with an 'attempted' flag so
exception failures (parse error, apply_changes raise after retry) are
observable, not only the happy path.
- Re-detect signals on the post-watermark feed for the extraction hint so
it no longer references turns the LLM cannot see; admission-time signals
still drive backpressure.
- Move the post-batch reschedule inside the queue lock to close a
non-atomic self._timer race with a concurrent add.
Co-Authored-By: Claude <noreply@anthropic.com>
* fix(memory): address follow-up review nits (LRU, metric name, docstring)
- Bound the in-memory watermark cache with a configurable LRU
(watermark_max_keys, default 4096, 0=unbounded). A dropped key re-extracts
one batch on that thread's next turn (the documented restart behavior), so
eviction is safe and preserves the content-identity watermark's
front-removal guarantee. Adds _watermark_get/_watermark_set helpers and a
bounded-LRU regression test.
- Rename the extraction metric facts_accepted -> facts_passed_confidence so
the name matches what the >60% rejection-rate warning assumes (a
confidence-gate signal, not a persisted-fact count); drop the stale
"historical semantics" justification. Brand-new callback, one consumer.
- Fix the stale test_message_processing_signals module docstring: the signals
seam is already swapped to frozenset, and a stale stage-numbering prefix is
removed.
Co-Authored-By: Claude <noreply@anthropic.com>
---------
Co-authored-by: Claude <noreply@anthropic.com>
Memory Backends
Each subfolder under agents/memory/backends/ is a pluggable memory backend. Swap the active one by changing one line in config.yaml - no deer-flow core changes required.
deermem/- the default backend (deer-flow's own: structured facts + JSON storage).noop/- an empty backend and the template to copy when adding a new one.
This guide tells you which files to touch when you change, swap, or add a memory system. Paths are relative to backend/ unless noted.
Table of Contents
- Add a New Backend
- Switch the Active Backend
- Backend Contract
- Do Not Modify
- Common Pitfalls
- Reference
Add a New Backend
Copy noop/ to backends/<yourname>/ and edit three files in this folder plus two outside it.
| File | What to change |
|---|---|
backends/<yourname>/config.py |
Declare your config fields + from_backend_config (parse backend_config; read storage_path from it - do not import deer-flow path helpers) |
backends/<yourname>/<yourname>_manager.py |
Rename the class; parse config in model_post_init; implement from_config + the tier-1 abstracts (add/get_context); override tier-2/3 methods as needed (see Backend Contract) |
backends/<yourname>/__init__.py |
MANAGER_CLASS = YourManager (relative import) |
config.yaml (repo root, parent of backend/) |
memory.manager_class: <yourname> + your knobs under memory.backend_config |
packages/harness/pyproject.toml |
Only if the backend needs external libs: declare the dependency; add [tool.uv.sources] for vendored source. Otherwise uv sync purges it (see Common Pitfalls) |
See the docstring at the top of noop/noop_manager.py for the full 6-step walkthrough.
Switch the Active Backend
Edit config.yaml (repo root) only:
memory:
manager_class: <name> # deermem / noop / <yourname>
backend_config: { ... } # that backend's private config
Then restart deer-flow - the memory manager is a process-level singleton; a running process does not hot-reload config or backend code.
Backend Contract
1. The three-tier contract
MemoryManager is a pydantic BaseModel (not a bare ABC). Methods are tiered:
- Tier 1 (abstract) --
add+get_context: every backend MUST implement (write + read-inject are the backend's fundamental duties; missing one is caught at instantiation). - Tier 2 (management, with defaults) --
add_nowait(delegates toadd),search/get_memory/clear_memory/import_memory/export_memory/delete_memory(defaultraise NotImplementedError),shutdown_flush(defaultTrue). Override the ones your backend supports. - Tier 3 (optional hooks, with defaults) --
warm(defaultTrue),reload_memory/create_fact/delete_fact/update_fact(default raise),on_pre_compress/on_turn_start(default no-op).
A new backend implements from_config + add + get_context and overrides only what it supports; the rest inherits defaults. Signatures must match (parameter names, keyword-only args). noop is the minimal reference.
2. Return shape (critical, easy to get wrong)
get_memory / export_memory / clear_memory / import_memory return a dict that the gateway casts to the DeerMem shape (MemoryResponse: version / lastUpdated / user / history / facts[]). Your backend must return a dict this shape accepts, or:
- the data is silently dropped (pydantic ignores unknown fields);
- the frontend gets empty defaults and
lastUpdated=""crashes the date formatter.
A non-DeerMem backend maps its native records (e.g. {"results": [...]}) into this shape via a small adapter helper.
3. Tier-3 hooks (contracted, no hasattr probing)
create_fact / delete_fact / update_fact / reload_memory / warm are tier-3 hooks ON the base contract (with defaults). Callers (gateway / client / tools) invoke them directly and catch NotImplementedError for unsupported backends -- no more hasattr probing.
create_fact/delete_fact/update_fact- the frontend's add/delete/edit-fact buttons. Default raises (caller returns 501).reload_memory- the frontend's reload button (caller falls back toget_memoryonNotImplementedError).warm- one-time warm-up at gateway startup (defaultTrue= nothing to warm).
Implement the ones your backend supports; the rest inherit the default raise.
4. Portability (the golden rule)
Important
A backend talks to the host through exactly two channels: (1) the ABC method arguments (
manager.py), and (2) thebackend_configdict. The onlyfrom deerflowimport allowed anywhere in your backend folder is the ABC contract line in<name>_manager.py:
from deerflow.agents.memory.manager import MemoryManager
Change that one line (and only that line) to port the backend to another agent. Do not import deer-flow path helpers, config singletons, or models - get storage_path and everything else from backend_config.
5. What the host provides
The factory (manager.py::get_memory_manager) resolves the backend class, injects storage_path into backend_config, then calls cls.from_config(backend_config, mode=cfg.mode, **host_hooks). The host hooks (passed as from_config kwargs, NOT in backend_config):
backend_config["storage_path"](str) - a writable state dir (the host'sruntime_homeby default, or whateverconfig.yamlsets). Use this as your storage root.callbacks(MemoryCallbacks| None) - observability;on_memory_llm_callmerges trace metadata before your LLM call (langfuse). Pass it to your LLM path; ignore if you don't trace.should_keep_hidden_message/trace_context_manager/host_llm_factory- other host hooks; consume infrom_configif relevant.- Plus whatever the user puts under
config.yaml::memory.backend_config(your backend's own knobs).
Each backend's from_config consumes the hooks it needs (DeerMem does; noop ignores them).
Do Not Modify
These are backend-agnostic. Don't touch them when swapping backends (unless you're changing the shared contract, which affects every backend):
| File | Role |
|---|---|
packages/harness/deerflow/agents/memory/manager.py |
ABC + factory + scanner |
packages/harness/deerflow/agents/middlewares/memory_middleware.py |
after_agent -> manager.add |
packages/harness/deerflow/agents/memory/summarization_hook.py |
summarization -> manager.add_nowait |
packages/harness/deerflow/agents/lead_agent/prompt.py |
_get_memory_context -> manager.get_context |
app/gateway/routers/memory.py |
HTTP endpoints -> manager.* (direct call + try/except NotImplementedError) |
packages/harness/deerflow/config/memory_config.py |
shared 4 fields (enabled / injection_enabled / manager_class / backend_config) |
frontend/src/components/workspace/settings/memory-settings-page.tsx |
frontend memory page (assumes DeerMem shape) |
Note
The gateway and frontend are currently hard-coded to the DeerMem shape - that's why backends must return DeerMem-shape data (contract #2). Making them fully backend-agnostic is a larger refactor.
Common Pitfalls
Lessons from integrating external backends:
- External deps must be declared in
pyproject.toml. A bareuv pip installis purged on the nextuv sync/langgraph dev. Declare the dep (and[tool.uv.sources]for vendored source). - Return the DeerMem shape. Otherwise the frontend crashes with
Invalid time valueand your data is silently dropped. Build a small adapter helper to map your native records into it. - Fact CRUD returns 501 if not implemented. The frontend's delete-fact button reports
Operation 'delete fact' not supported. Implementdelete_fact(and friends) to fix it. - Don't import
runtime_home. Readstorage_pathfrombackend_config. (Thenooptemplate shows the correct pattern; importing deer-flow path helpers breaks portability - contract #4.) - Restart deer-flow after changes. The manager is a process-level singleton; a running process does not hot-reload config or backend code.
- Cap
get_contextlength yourself. The host applies no token budget; the backend must truncate (DeerMem hasmax_injection_tokens; noop does not).
Reference
- Template:
noop/- minimal implementation with full docstrings; copy and go. - Contract + factory:
packages/harness/deerflow/agents/memory/manager.py(MemoryManagerbase,MemoryCallbacks,get_memory_managerfactory).