mirror of
https://github.com/bytedance/deer-flow.git
synced 2026-09-14 16:08:41 +00:00
* fix(summarization): resolve fraction triggers from declared context_window, degrade instead of crashing the agent build A fraction trigger/keep clause requires profile["max_input_tokens"], which any third-party OpenAI-compatible model lacks, so SummarizationMiddleware construction raised ValueError out of create_summarization_middleware and failed the whole agent build (#3103). - factory: translate a declared model context_window into the langchain profile (metadata-only, never reaches the provider payload); explicit caller/override profiles win - summarization factory: drop unusable fraction trigger clauses (absolute clauses survive), fall a fraction keep back to the messages default, and disable compaction with an actionable warning only when no usable trigger clause remains — the agent build never dies from summarization config - docs: config.example.yaml, ModelConfig.context_window, summarization.md * refactor(summarization): share the default keep constant with the fraction fallback The fraction-keep degradation fallback hardcoded ("messages", 20), duplicating SummarizationConfig.keep's default_factory literal. Move the value to a shared DEFAULT_KEEP constant so the two cannot drift apart. * fix(summarization): keep trigger-null + fraction-keep constructing after degradation A trigger of None with a fraction keep hit the all-clauses-dropped branch (has_usable_trigger=False) and disabled compaction, and the accompanying warning claimed configured triggers were all fraction-based when none were configured. Only report nothing-usable when trigger clauses actually existed; trigger:null keeps constructing the never-firing middleware with the degraded keep, matching its behavior outside the degradation path. * fix(summarization): address review — keep manual compaction, validate ContextSize, pin wiring Review follow-ups on #4901: - When every configured trigger is a dropped fraction clause, keep constructing the never-firing middleware (trigger=None) instead of returning None: manual /compact runs with force=True and never consults trigger clauses, so it must keep working for a profile-less model rather than reporting 'compaction is disabled'. The warning now says auto-compaction will not fire while manual compaction remains. - ContextSize gains a config-load validator: fraction values must be in (0,1] (a percent-style 80 instead of 0.8 previously produced a threshold the context could never reach — a silently inert trigger), absolute values must be positive. - New un-monkeypatched integration test pins the shipped wiring (context_window declared -> real factory attaches profile -> fraction clause survives -> middleware constructs), which the stubbed middleware-side tests and kwarg-capturing factory-side tests each stopped short of. - Docs (summarization.md + config.example.yaml) clarify that the fraction resolves against the summary/anchor model's context_window (summarization.model_name when set, else the run model), including the mismatch caveat for a larger-window summary model. * fix(summarization): reject non-finite ContextSize values at config load YAML .nan / .inf pass pydantic's float parsing, and nan <= 0 is False, so the positivity check alone let them through as dead thresholds (count >= nan is always False) — the same silent-inert-trigger class the range validator was added to close. Guard with math.isfinite first, consistent with the existing non-finite guards on mem0 timeout_seconds and poll_after_seconds. * fix(summarization): merge context_window into inferred profile, require whole message counts - construct the model first, then merge max_input_tokens into the provider-inferred langchain profile: passing profile= to the constructor replaced the whole inferred metadata (tool_calling, structured_output, io capabilities, output limits) with the single key. An explicitly configured profile is still never clobbered. - reject non-integral ContextSize values for type=messages at config load: langchain slices the message list with them, so a float index raised TypeError mid-compaction. --------- Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
393 lines
15 KiB
Markdown
393 lines
15 KiB
Markdown
# Conversation Summarization
|
|
|
|
DeerFlow includes automatic conversation summarization to handle long conversations that approach model token limits. When enabled, the system automatically condenses older messages while preserving recent context.
|
|
|
|
New checkpoints no longer use raw task-result or skill-read transcript content to derive durable context. The capture path consumes bounded structured metadata stamped on the corresponding `ToolMessage.additional_kwargs`; transcript text remains display/model content, not the state-capture protocol.
|
|
|
|
## Overview
|
|
|
|
The summarization feature uses LangChain's `SummarizationMiddleware` to monitor conversation history and trigger summarization based on configurable thresholds. When activated, it:
|
|
|
|
1. Monitors message token counts in real-time
|
|
2. Triggers summarization when thresholds are met
|
|
3. Keeps recent messages intact while summarizing older exchanges
|
|
4. Maintains AI/Tool message pairs together for context continuity
|
|
5. Stores the summary in `ThreadState.summary_text` and projects it ephemerally through durable context data
|
|
|
|
## Configuration
|
|
|
|
Summarization is configured in `config.yaml` under the `summarization` key:
|
|
|
|
```yaml
|
|
summarization:
|
|
enabled: true
|
|
model_name: null # null = summarize with the run's own model (see below); or name a lightweight model
|
|
|
|
# Trigger conditions (OR logic - any condition triggers summarization)
|
|
trigger:
|
|
- type: tokens
|
|
value: 4000
|
|
# Additional triggers (optional)
|
|
# - type: messages
|
|
# value: 50
|
|
# - type: fraction
|
|
# value: 0.8 # 80% of model's max input tokens
|
|
|
|
# Context retention policy
|
|
keep:
|
|
type: messages
|
|
value: 20
|
|
|
|
# Token trimming for summarization call
|
|
trim_tokens_to_summarize: 4000
|
|
|
|
# Custom summary prompt (optional)
|
|
summary_prompt: null
|
|
|
|
# Tool names treated as skill file reads for the durable skill_context channel
|
|
skill_file_read_tool_names:
|
|
- read_file
|
|
- read
|
|
- view
|
|
- cat
|
|
```
|
|
|
|
### Configuration Options
|
|
|
|
#### `enabled`
|
|
- **Type**: Boolean
|
|
- **Default**: `false`
|
|
- **Description**: Enable or disable automatic summarization
|
|
|
|
#### `model_name`
|
|
- **Type**: String or null
|
|
- **Default**: `null`
|
|
- **Description**: Model to use for generating summaries.
|
|
- **`null` (model ownership)**: summarize with the model the run actually executes with — the lead run's resolved model, a subagent's own model, or a thread's custom-agent model — **not** `config.models[0]`. This keeps compaction working on a run whose model is healthy even when `models[0]`'s provider is broken (expired key, quota, outage).
|
|
- **Set to a model name**: that model generates summaries. If its provider fails, compaction **falls back to the run's own model** so a broken summary provider cannot disable compaction while a working model is available. Recommended to use a lightweight, cost-effective model like `gpt-4o-mini` or equivalent.
|
|
- Ownership applies to all three paths — automatic lead compaction, subagent compaction, and manual `/compact`. Manual `/compact` resolves the run model with the same precedence as a normal run: the model selected for the request (`POST /api/threads/{id}/compact` body `model_name`, sent by the frontend from the composer's current model) → the thread's custom-agent model → the default. A whitespace-only summary response is treated as a generation failure (it is never committed as a valid empty summary).
|
|
|
|
#### `trigger`
|
|
- **Type**: Single `ContextSize` or list of `ContextSize` objects
|
|
- **Required**: At least one trigger must be specified when enabled
|
|
- **Description**: Thresholds that trigger summarization. Uses OR logic - summarization runs when ANY threshold is met.
|
|
|
|
**ContextSize Types:**
|
|
|
|
1. **Token-based trigger**: Activates when token count reaches the specified value
|
|
```yaml
|
|
trigger:
|
|
type: tokens
|
|
value: 4000
|
|
```
|
|
|
|
2. **Message-based trigger**: Activates when message count reaches the specified value
|
|
```yaml
|
|
trigger:
|
|
type: messages
|
|
value: 50
|
|
```
|
|
|
|
3. **Fraction-based trigger**: Activates when token usage reaches a percentage of the model's maximum input tokens
|
|
```yaml
|
|
trigger:
|
|
type: fraction
|
|
value: 0.8 # 80% of max input tokens
|
|
```
|
|
|
|
The percentage resolves from the **summary model's** declared `context_window`
|
|
— the anchor that generates summaries: `summarization.model_name` when set,
|
|
otherwise the run's own model. Declare `context_window` on that models entry
|
|
in `config.yaml`. Third-party OpenAI-compatible models carry no built-in
|
|
capacity profile, so without a declared `context_window` the fraction clause
|
|
is dropped with a warning at agent build — any remaining absolute clauses
|
|
(`tokens` / `messages`) keep working. Caveat: when a separate summary model
|
|
is configured, its window sizes the threshold — a 64k run model paired with
|
|
a 128k-window summary model resolves `fraction: 0.8` to ~102k tokens and
|
|
auto-summarization cannot fire before the run model overflows; in that setup
|
|
prefer absolute `tokens` thresholds sized for the run model.
|
|
|
|
**Multiple Triggers:**
|
|
```yaml
|
|
trigger:
|
|
- type: tokens
|
|
value: 4000
|
|
- type: messages
|
|
value: 50
|
|
```
|
|
|
|
#### `keep`
|
|
- **Type**: `ContextSize` object
|
|
- **Default**: `{type: messages, value: 20}`
|
|
- **Description**: Specifies how much recent conversation history to preserve after summarization.
|
|
|
|
**Examples:**
|
|
```yaml
|
|
# Keep most recent 20 messages
|
|
keep:
|
|
type: messages
|
|
value: 20
|
|
|
|
# Keep most recent 3000 tokens
|
|
keep:
|
|
type: tokens
|
|
value: 3000
|
|
|
|
# Keep most recent 30% of model's max input tokens
|
|
keep:
|
|
type: fraction
|
|
value: 0.3
|
|
```
|
|
|
|
#### `trim_tokens_to_summarize`
|
|
- **Type**: Integer or null
|
|
- **Default**: `4000`
|
|
- **Description**: Maximum tokens to include when preparing messages for the summarization call itself. Set to `null` to skip trimming (not recommended for very long conversations).
|
|
|
|
#### `summary_prompt`
|
|
- **Type**: String or null
|
|
- **Default**: `null` (uses LangChain's default prompt)
|
|
- **Description**: Custom prompt template for generating summaries. The prompt should guide the model to extract the most important context.
|
|
|
|
#### `skill_file_read_tool_names`
|
|
- **Type**: List of strings
|
|
- **Default**: `["read_file", "read", "view", "cat"]`
|
|
- **Description**: Tool names treated as skill file reads when `DurableContextMiddleware` captures loaded skills into the checkpointed `skill_context` channel. A tool call is captured only when its name appears in this list and its target path is under `skills.container_path`. Set this list to `[]` to disable durable skill-reference capture.
|
|
|
|
Legacy `preserve_recent_skill_*` settings are no longer used. Loaded skill retention is handled by the durable `skill_context` reference channel instead of by preserving raw skill-read messages in the summarization window.
|
|
|
|
**Default Prompt Behavior:**
|
|
The default LangChain prompt instructs the model to:
|
|
- Extract highest quality/most relevant context
|
|
- Focus on information critical to the overall goal
|
|
- Avoid repeating completed actions
|
|
- Return only the extracted context
|
|
|
|
## How It Works
|
|
|
|
### Summarization Flow
|
|
|
|
1. **Monitoring**: Before each model call, the middleware counts tokens in the message history plus the existing `summary_text`, because both are projected into the next model request
|
|
2. **Trigger Check**: If any configured threshold is met, summarization is triggered
|
|
3. **Message Partitioning**: Messages are split into:
|
|
- Messages to summarize (older messages beyond the `keep` threshold)
|
|
- Messages to preserve (recent messages within the `keep` threshold)
|
|
4. **Summary Generation**: The model generates a concise summary of the older messages
|
|
5. **Context Replacement**: The message history is updated:
|
|
- All old messages are removed
|
|
- Recent messages are preserved
|
|
- The generated prose summary is stored in `summary_text`
|
|
6. **AI/Tool Pair Protection**: The system ensures AI messages and their corresponding tool messages stay together
|
|
7. **Skill context channel**: Skill files read during the conversation (tool calls whose name is in `skill_file_read_tool_names` and whose path is under `skills.container_path`, narrowed to `.../SKILL.md`) are stamped with `skill_context_entry` metadata at the read-tool boundary, then captured by `DurableContextMiddleware` into the checkpointed `skill_context` channel as references: `name`, `path`, a one-line `description` parsed in-memory from the file's frontmatter, and `loaded_at`, deduped by path. On every model call they are rendered into a hidden durable-context data message as a compact "active skills" reminder that points at each `SKILL.md` for on-demand re-read, so which skills are active survives summarization without persisting or re-injecting the verbatim body. The channel keeps the most recently read skills (cap `_SKILL_CONTEXT_MAX_ENTRIES`; re-reading an existing skill refreshes its recency); sessions typically load only 1-3.
|
|
|
|
### Token Counting
|
|
|
|
- Uses approximate token counting based on character count
|
|
- For Anthropic models: ~3.3 characters per token
|
|
- For other models: Uses LangChain's default estimation
|
|
- Can be customized with a custom `token_counter` function
|
|
|
|
### Message Preservation
|
|
|
|
The middleware intelligently preserves message context:
|
|
|
|
- **Recent Messages**: Always kept intact based on `keep` configuration
|
|
- **AI/Tool Pairs**: Never split - if a cutoff point falls within tool messages, the system adjusts to keep the entire AI + Tool message sequence together
|
|
- **Summary Format**: Summary prose is stored in `summary_text` and rendered into an ephemeral hidden durable-context data message. Static handling rules live in a separate system message; summary text and other user/tool/model-derived values stay in the lower-authority data message.
|
|
```
|
|
<durable_context_data>
|
|
## Conversation summary so far
|
|
[Generated summary text]
|
|
</durable_context_data>
|
|
```
|
|
|
|
## Best Practices
|
|
|
|
### Choosing Trigger Thresholds
|
|
|
|
1. **Token-based triggers**: Recommended for most use cases
|
|
- Set to 60-80% of your model's context window
|
|
- Example: For 8K context, use 4000-6000 tokens
|
|
|
|
2. **Message-based triggers**: Useful for controlling conversation length
|
|
- Good for applications with many short messages
|
|
- Example: 50-100 messages depending on average message length
|
|
|
|
3. **Fraction-based triggers**: Ideal when using multiple models
|
|
- Automatically adapts to each model's capacity
|
|
- Example: 0.8 (80% of model's max input tokens)
|
|
|
|
### Choosing Retention Policy (`keep`)
|
|
|
|
1. **Message-based retention**: Best for most scenarios
|
|
- Preserves natural conversation flow
|
|
- Recommended: 15-25 messages
|
|
|
|
2. **Token-based retention**: Use when precise control is needed
|
|
- Good for managing exact token budgets
|
|
- Recommended: 2000-4000 tokens
|
|
|
|
3. **Fraction-based retention**: For multi-model setups
|
|
- Automatically scales with model capacity
|
|
- Recommended: 0.2-0.4 (20-40% of max input)
|
|
|
|
### Model Selection
|
|
|
|
- **Recommended**: Use a lightweight, cost-effective model for summaries
|
|
- Examples: `gpt-4o-mini`, `claude-haiku`, or equivalent
|
|
- Summaries don't require the most powerful models
|
|
- Significant cost savings on high-volume applications
|
|
|
|
- **Default**: If `model_name` is `null`, summarizes with the run's own model (not `models[0]`)
|
|
- Keeps compaction working when `models[0]`'s provider is broken but the run's model is healthy
|
|
- Good for simple setups; no separate summary provider to keep credentialed
|
|
|
|
### Optimization Tips
|
|
|
|
1. **Balance triggers**: Combine token and message triggers for robust handling
|
|
```yaml
|
|
trigger:
|
|
- type: tokens
|
|
value: 4000
|
|
- type: messages
|
|
value: 50
|
|
```
|
|
|
|
2. **Conservative retention**: Keep more messages initially, adjust based on performance
|
|
```yaml
|
|
keep:
|
|
type: messages
|
|
value: 25 # Start higher, reduce if needed
|
|
```
|
|
|
|
3. **Trim strategically**: Limit tokens sent to summarization model
|
|
```yaml
|
|
trim_tokens_to_summarize: 4000 # Prevents expensive summarization calls
|
|
```
|
|
|
|
4. **Monitor and iterate**: Track summary quality and adjust configuration
|
|
|
|
## Troubleshooting
|
|
|
|
### Summary Quality Issues
|
|
|
|
**Problem**: Summaries losing important context
|
|
|
|
**Solutions**:
|
|
1. Increase `keep` value to preserve more messages
|
|
2. Decrease trigger thresholds to summarize earlier
|
|
3. Customize `summary_prompt` to emphasize key information
|
|
4. Use a more capable model for summarization
|
|
|
|
### Performance Issues
|
|
|
|
**Problem**: Summarization calls taking too long
|
|
|
|
**Solutions**:
|
|
1. Use a faster model for summaries (e.g., `gpt-4o-mini`)
|
|
2. Reduce `trim_tokens_to_summarize` to send less context
|
|
3. Increase trigger thresholds to summarize less frequently
|
|
|
|
### Token Limit Errors
|
|
|
|
**Problem**: Still hitting token limits despite summarization
|
|
|
|
**Solutions**:
|
|
1. Lower trigger thresholds to summarize earlier
|
|
2. Reduce `keep` value to preserve fewer messages
|
|
3. Check if individual messages are very large
|
|
4. Consider using fraction-based triggers
|
|
|
|
## Implementation Details
|
|
|
|
### Code Structure
|
|
|
|
- **Configuration**: `packages/harness/deerflow/config/summarization_config.py`
|
|
- **Integration**: `packages/harness/deerflow/agents/lead_agent/agent.py`
|
|
- **Middleware**: Uses `langchain.agents.middleware.SummarizationMiddleware`
|
|
|
|
### Middleware Order
|
|
|
|
Durable context capture runs before summarization so task delegations and
|
|
loaded skill references are recorded before their raw tool messages can be
|
|
compacted. It records in-progress dispatches as well as terminal result
|
|
summaries. Summarization then reduces message history before downstream
|
|
middlewares such as title generation, memory queuing, and clarification:
|
|
|
|
1. Runtime middlewares, including ThreadData and Sandbox initialization
|
|
2. DynamicContextMiddleware
|
|
3. SkillActivationMiddleware
|
|
4. DurableContextMiddleware
|
|
5. **SummarizationMiddleware** ← Runs here
|
|
6. Downstream lead middlewares such as Title, Memory, and Clarification
|
|
|
|
### State Management
|
|
|
|
- Summarization configuration is loaded from `config.yaml`
|
|
- Generated summaries are stored in `ThreadState.summary_text`, not as regular `messages`
|
|
- The message reducer removes compacted raw messages while the checkpointer persists `summary_text`
|
|
- DurableContextMiddleware projects `summary_text` back into later model calls as hidden durable context data
|
|
|
|
## Example Configurations
|
|
|
|
### Minimal Configuration
|
|
```yaml
|
|
summarization:
|
|
enabled: true
|
|
trigger:
|
|
type: tokens
|
|
value: 4000
|
|
keep:
|
|
type: messages
|
|
value: 20
|
|
```
|
|
|
|
### Production Configuration
|
|
```yaml
|
|
summarization:
|
|
enabled: true
|
|
model_name: gpt-4o-mini # Lightweight model for cost efficiency
|
|
trigger:
|
|
- type: tokens
|
|
value: 6000
|
|
- type: messages
|
|
value: 75
|
|
keep:
|
|
type: messages
|
|
value: 25
|
|
trim_tokens_to_summarize: 5000
|
|
```
|
|
|
|
### Multi-Model Configuration
|
|
```yaml
|
|
summarization:
|
|
enabled: true
|
|
model_name: gpt-4o-mini
|
|
trigger:
|
|
type: fraction
|
|
value: 0.7 # 70% of model's max input
|
|
keep:
|
|
type: fraction
|
|
value: 0.3 # Keep 30% of max input
|
|
trim_tokens_to_summarize: 4000
|
|
```
|
|
|
|
### Conservative Configuration (High Quality)
|
|
```yaml
|
|
summarization:
|
|
enabled: true
|
|
model_name: gpt-4 # Use full model for high-quality summaries
|
|
trigger:
|
|
type: tokens
|
|
value: 8000
|
|
keep:
|
|
type: messages
|
|
value: 40 # Keep more context
|
|
trim_tokens_to_summarize: null # No trimming
|
|
```
|
|
|
|
## References
|
|
|
|
- [LangChain Summarization Middleware Documentation](https://docs.langchain.com/oss/python/langchain/middleware/built-in#summarization)
|
|
- [LangChain Source Code](https://github.com/langchain-ai/langchain)
|