* feat(memory): add memory consolidation to synthesize fragmented facts
When a fact category accumulates many individual entries, the LLM
reviews them during the normal memory-update call (same invocation,
no extra API cost) and decides whether groups of related facts can be
synthesized into a single richer fact. This completes the memory
lifecycle: extraction → guaranteed injection → staleness review →
consolidation.
- Select fragmented categories by min-facts threshold, surface the most
fragmented groups first; prompt-layer caps aligned with apply-layer
guardrails so the LLM never sees groups it cannot act on
- Cap consolidated confidence at source maximum to prevent inflation;
reject results below fact_confidence_threshold
- Double-consume protection prevents a fact from being merged into
multiple consolidation targets
- Feature-gated at both prompt and apply time with per-cycle safety caps
- Add 26 tests covering candidate selection, normalization, apply
guardrails, and prompt integration
* fix(memory): address consolidation correctness issues from PR review
Six fixes based on maintainer review of #3996:
1. Deduplicate sourceIds in normalization — ["f1","f1"] previously
bypassed the ≥2-distinct-sources check; dict.fromkeys collapses it
to ["f1"] which is correctly rejected.
2. Run consolidation after max_facts trim — previously, sources were
deleted then the merged fact could be evicted by the trim, leaving
no record of either. Moving consolidation last ensures source facts
exist in the post-trim index before removal.
3. Fix count= attribute in consolidation prompt — advertised the full
category size but listed only max_sources IDs; now uses
min(len(group), max_sources) to match what the LLM can act on.
4. Exempt staleness_protected_categories from consolidation candidates
— mirrors the existing staleness-review contract so correction facts
are never surfaced for merging.
5. Strip and default category in consolidation normalization — " " or
" preference " are now normalised, matching _normalize_memory_update_fact.
6. Propagate sourceError from source facts into consolidated fact —
correction context is no longer silently lost on merge.
* fix(memory): add apply-time guardrails and tests for consolidation
P1: mirror the staleness-pass defense-in-depth pattern — build
allowed_source_ids from _select_consolidation_candidates at apply time
so a protected-category or below-threshold fact proposed by the LLM is
rejected regardless of model behavior.
P2a: test that LLM-returned confidence is capped at max source confidence
and that a capped result below fact_confidence_threshold is rejected.
P2b: test that factsToConsolidate with consolidation_enabled=False is a
no-op at apply time (35 tests, all pass).
* fix(memory): address three correctness issues from second review round
1. Default consolidation_enabled=False — consolidation is lossy (source
content is permanently replaced, only consolidatedFrom IDs preserved);
new lossy features default to off. config.example.yaml updated to match.
2. Unify confidence coercion between prompt and apply — _build_consolidation_section
now calls _coerce_source_confidence(fact) instead of an inline 0.0-default
coercion, so a null-confidence fact renders with 0.50 in the LLM prompt and
is capped at 0.50 at apply time (same value, same function).
3. Preserve staleness clock on merge — consolidated fact now carries the
newest source's createdAt (not now) so aged information does not gain a
fresh staleness-review window just by being consolidated; consolidatedAt
is added as an explicit audit field.
Three regression tests added (default=false, null-confidence consistency,
createdAt policy); all guardrail tests now set consolidation_enabled=True
explicitly so they test the guardrail, not the feature flag. 38 tests pass.
* fix(memory): harden createdAt comparison and confidence handling
1. createdAt max via _parse_fact_datetime — replaces string max() which
crashes on non-string createdAt (numeric unix timestamps) and sorts
Z/+00:00 mixed formats incorrectly. Mirrors how staleness computes age.
2. Remove dead min(..., 1.0) — _coerce_source_confidence already clamps
each source confidence to [0, 1], so max(source_confidences) ≤ 1.0
by contract; the outer min could never bind.
3. Clamp raw_llm_conf to [0, 1] before applying the source cap — out-of-
range values like 1.5 are safe today (pinned by the cap) but defensively
clamped first so the invariant holds even if the cap is ever loosened.
4. Doc: expand the apply-time guardrails comment to call out the protected-
category exclusion via allowed_source_ids — this is the central safety
property ("explicit user feedback is never silently merged away").
5. Test: add test_confidence_fallback_to_max_source_when_llm_omits_field
covering the else-branch (LLM omits confidence → uses max_source_conf).
6. Fix lint: reorder imports in test file (stdlib before third-party).
39 tests, all pass.
---------
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>