deer-flow/backend/tests/test_gateway_checkpoint_mode.py
Amorend 85c3909c2e
feat: show real-time context window usage (#3125) (#3183)
* feat: show real-time context window usage in chat UI (#3125)

Adds a `context_usage` block to `GET /api/threads/{id}/token-usage`
(token count from the live checkpoint, the thread model's
`context_window`, and a percentage), introduces a new
`ModelConfig.context_window` distinct from the per-call `max_tokens`
output cap, and surfaces the percentage in the chat header — inside
`TokenUsageIndicator` when token-usage tracking is on, or as a
standalone badge when it's off so context capacity stays visible
independent of cost tracking.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* feat: per-category breakdown for context window usage

Replace the single-number context_usage payload with a Claude-Code-style
breakdown — messages, system prompt, skills, system/MCP tools (active +
deferred), custom agents, memory injection, autocompact buffer, and free
space — and surface it in the chat UI with a segmented progress bar and
per-row table.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* docs(config): document context_window across model examples

Add `context_window` to every example model in config.example.yaml so the
new chat-UI "% context used" indicator works out of the box for whichever
example a user adopts. Each value is the published default at the time of
writing; users are pointed at the official model spec to verify. Bumps
config_version to 11 so `make config-upgrade` flags outdated user configs.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* style: ruff format (line-length 240)

No behavior change — collapses two multi-line expressions that fit on
one line under the project's 240-char limit. Picked up by `make format`.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* review: address Copilot bot comments on #3183

- token-usage-indicator: switch `{contextPercentage && (...)}` to an
  explicit `!= null` check. (The string `"0"` is actually truthy in JS so
  the original code wasn't buggy, but the explicit check is clearer.)
- context-usage-breakdown: drop the `useMemo` around segments/totals — the
  computation is O(n) over a handful of rows and the previous memo deps
  omitted `t.contextUsage.categories`, so the bar's tooltips/aria-labels
  could stay in the old language after a locale switch.
- context_usage._split_tools: snapshot MCP names from
  `get_cached_mcp_tools()` directly instead of re-reading
  `extensions_config.json` after `get_available_tools()` already loaded
  it. Removes redundant file I/O on every `/token-usage` poll.
  (`get_available_tools()` still emits its own INFO logs — silencing
  those is out of scope here.)

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* style(frontend): prettier --write context-usage-breakdown

CI's `pnpm format` (prettier --check) caught two lines previously
formatted by hand. Collapses one comma to fit on one line; no behavior
change.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(gateway): correct context-usage breakdown + add exact token counting

The context-usage indicator shipped two bugs that silently zeroed whole
breakdown rows (both caught by try/except, so the feature looked alive but
produced wrong numbers):

1. _count_system_prompt passed app_config= to get_deferred_tools_prompt_section,
   which only accepts deferred_names -> TypeError swallowed -> system_prompt
   row always 0, and used_tokens/percentage undercounted by the full prompt.
   Also subtracted the deferred section twice (the rendered prompt already
   excluded it). Fix: derive deferred names deterministically and pass them to
   apply_prompt_template; drop the redundant subtraction.

2. _split_tools imported a non-existent get_deferred_registry -> ImportError
   swallowed -> all four tool-category rows always 0. Fix: classify via the
   public is_mcp_tool predicate + tool_search.enabled (mirrors
   build_deferred_tool_setup); the MCP tag is set by get_available_tools.

Added token_usage.counting (approximate|exact). 'exact' routes text/schema/
message counting through the model tokenizer (tiktoken cl100k_base) via the
existing memory-module machinery (lazy load + cache + cooldown + CJK-aware
fallback), so CJK-heavy threads stop being undercounted by chars//4.

Regression + e2e tests added; 6621 backend tests pass.

* fix(gateway): harden context usage accounting

* fix(gateway): count promoted MCP tools as active in context usage

Promoted tools (deferred MCP tools the thread has fetched via tool_search)
have their full schema bound on every subsequent turn by
DeferredToolFilterMiddleware, so they consume context like any active tool.
The breakdown previously left them in the reserved *_deferred rows, under-
counting the thread's used_tokens.

Classification now treats a tool as deferred only when tool_search is enabled,
it is MCP-sourced, AND it has not been promoted. The promoted set is read from
the checkpoint's channel_values and scoped by catalog hash — matching the
runtime middleware, so a stale promotion from MCP-config drift cannot inflate
the active count.

The static system prompt still lists all deferred tool names (promotions only
affect schema binding, not the prompt), so _count_system_prompt's deferred
rendering is intentionally left unchanged.

8 new tests cover classification, catalog-hash scoping (match / drift /
compute-failure / malformed), and checkpoint extraction.

* fix(context): address review feedback

* fix(context): count structured message payloads

* fix(context): harden usage accounting

* fix(config): bump schema for context usage fields

* refactor: narrow context usage to core indicator

---------

Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
2026-07-31 21:57:22 +08:00

322 lines
14 KiB
Python

"""Dual-mode (full/delta) parity for the gateway thread-state endpoints.
Drives ``GET /api/threads/{id}``, ``GET /api/threads/{id}/state``,
``POST /api/threads/{id}/history``, and the context-usage checkpoint reader
through the real materialization stack
(``build_thread_checkpoint_state_accessor`` -> factory-built graph ->
``CheckpointStateAccessor``) against a real ``InMemorySaver``, once per
checkpoint channel mode, and asserts the wire responses are identical apart
from checkpoint ids/timestamps. The delta storage layout must be invisible
to API consumers.
"""
from __future__ import annotations
import asyncio
from types import SimpleNamespace
from typing import Any
from unittest.mock import AsyncMock
import pytest
from _router_auth_helpers import make_authed_test_app
from fastapi.testclient import TestClient
from langchain_core.messages import AIMessage, HumanMessage
from langgraph.checkpoint.memory import InMemorySaver
from langgraph.graph import StateGraph
from langgraph.store.memory import InMemoryStore
from app.gateway import context_usage
from app.gateway import services as gateway_services
from app.gateway.routers import threads
from deerflow.agents.thread_state import get_thread_state_schema
from deerflow.config.app_config import AppConfig, reset_app_config, set_app_config
from deerflow.persistence.thread_meta.memory import MemoryThreadMetaStore
from deerflow.runtime import RunManager
from deerflow.runtime.checkpoint_mode import checkpoint_metadata_uses_delta, inject_checkpoint_mode
from deerflow.runtime.runs.store.memory import MemoryRunStore
_THREAD_ID = "thread-gateway-parity"
@pytest.fixture
def _stub_app_config():
set_app_config(AppConfig.model_validate({"sandbox": {"use": "deerflow.sandbox.local:LocalSandboxProvider"}}))
yield
reset_app_config()
def _build_reply_graph(mode: str, checkpointer: Any):
async def _reply(state: dict[str, Any]) -> dict[str, Any]:
n = len(state.get("messages") or [])
return {"messages": [AIMessage(content=f"answer-{n}", id=f"a{n}")]}
builder = StateGraph(get_thread_state_schema(mode))
builder.add_node("reply", _reply)
builder.set_entry_point("reply")
builder.set_finish_point("reply")
return builder.compile(checkpointer=checkpointer)
def _message_wire_shape(messages: list[dict[str, Any]]) -> list[tuple[str, str, str]]:
return [(message.get("type"), message.get("content"), message.get("id")) for message in messages]
def _run_gateway_flow(mode: str, monkeypatch: pytest.MonkeyPatch) -> dict[str, Any]:
app = make_authed_test_app()
store = InMemoryStore()
checkpointer = InMemorySaver()
app.state.store = store
app.state.checkpointer = checkpointer
app.state.thread_store = MemoryThreadMetaStore(store)
app.state.checkpoint_channel_mode = mode
app.state.run_event_store = SimpleNamespace()
app.include_router(threads.router)
graph = _build_reply_graph(mode, checkpointer)
monkeypatch.setattr(
gateway_services,
"resolve_agent_factory",
lambda assistant_id=None: lambda config: graph,
)
config: dict[str, Any] = {"configurable": {"thread_id": _THREAD_ID}}
inject_checkpoint_mode(config, mode)
for i in range(2):
asyncio.run(graph.ainvoke({"messages": [HumanMessage(content=f"question-{i}", id=f"h{i}")]}, config))
with TestClient(app) as client:
thread_response = client.get(f"/api/threads/{_THREAD_ID}")
state_response = client.get(f"/api/threads/{_THREAD_ID}/state")
history_response = client.post(f"/api/threads/{_THREAD_ID}/history", json={"limit": 10})
assert thread_response.status_code == 200, thread_response.text
assert state_response.status_code == 200, state_response.text
assert history_response.status_code == 200, history_response.text
thread_payload = thread_response.json()
state_payload = state_response.json()
history_payload = history_response.json()
return {
"thread_status": thread_payload["status"],
"thread_messages": _message_wire_shape(thread_payload["values"]["messages"]),
"state_messages": _message_wire_shape(state_payload["values"]["messages"]),
"history_messages": [_message_wire_shape(snapshot["values"].get("messages", [])) for snapshot in history_payload],
}
def test_thread_state_endpoints_are_mode_invariant(_stub_app_config, monkeypatch: pytest.MonkeyPatch) -> None:
full = _run_gateway_flow("full", monkeypatch)
monkeypatch.undo()
delta = _run_gateway_flow("delta", monkeypatch)
assert full == delta
# Guard against a vacuous pass: the flow must have observed real messages.
assert full["thread_messages"], "expected seeded messages in the thread response"
assert any(full["history_messages"]), "expected history snapshots with messages"
@pytest.mark.parametrize("mode", ["full", "delta"])
def test_context_usage_reads_materialized_messages_in_both_modes(
mode: str,
_stub_app_config,
monkeypatch: pytest.MonkeyPatch,
) -> None:
"""Context usage must not read raw delta ``channel_values``."""
app = make_authed_test_app()
store = InMemoryStore()
checkpointer = InMemorySaver()
app.state.store = store
app.state.checkpointer = checkpointer
app.state.thread_store = MemoryThreadMetaStore(store)
app.state.checkpoint_channel_mode = mode
app.state.run_event_store = SimpleNamespace()
graph = _build_reply_graph(mode, checkpointer)
monkeypatch.setattr(
gateway_services,
"resolve_agent_factory",
lambda assistant_id=None: lambda config: graph,
)
config: dict[str, Any] = {"configurable": {"thread_id": _THREAD_ID}}
inject_checkpoint_mode(config, mode)
for i in range(2):
asyncio.run(graph.ainvoke({"messages": [HumanMessage(content=f"question-{i}", id=f"h{i}")]}, config))
request = SimpleNamespace(app=app)
accessor, read_config = asyncio.run(gateway_services.build_thread_checkpoint_state_accessor(request, thread_id=_THREAD_ID))
messages = asyncio.run(context_usage._load_checkpoint_messages(accessor, read_config))
assert [(message.type, message.content, message.id) for message in messages] == [
("human", "question-0", "h0"),
("ai", "answer-1", "a1"),
("human", "question-1", "h1"),
("ai", "answer-3", "a3"),
]
def test_full_mode_gateway_rejects_delta_thread_with_409(_stub_app_config, monkeypatch: pytest.MonkeyPatch) -> None:
"""Fail-closed gate at the HTTP boundary, against a real checkpointer.
A full-mode process opening a delta thread must get a precise 409 naming
the cause — not a generic 500 that forces operators to grep logs. Seeds a
real delta checkpoint through the delta graph (marker + LangGraph delta
counters land in checkpoint metadata), then exercises every state surface
of the threads router in full mode.
"""
app = make_authed_test_app()
store = InMemoryStore()
checkpointer = InMemorySaver()
app.state.store = store
app.state.checkpointer = checkpointer
app.state.thread_store.get = AsyncMock(return_value=None)
app.state.checkpoint_channel_mode = "full"
app.state.run_event_store = SimpleNamespace()
app.state.run_manager = RunManager(store=MemoryRunStore())
app.include_router(threads.router)
full_graph = _build_reply_graph("full", checkpointer)
monkeypatch.setattr(
gateway_services,
"resolve_agent_factory",
lambda assistant_id=None: lambda config: full_graph,
)
# Seed through the delta graph so the checkpoint carries the delta marker.
delta_graph = _build_reply_graph("delta", checkpointer)
config: dict[str, Any] = {"configurable": {"thread_id": _THREAD_ID}}
inject_checkpoint_mode(config, "delta")
asyncio.run(delta_graph.ainvoke({"messages": [HumanMessage(content="question", id="h0")]}, config))
latest = asyncio.run(checkpointer.aget_tuple({"configurable": {"thread_id": _THREAD_ID, "checkpoint_ns": ""}}))
assert checkpoint_metadata_uses_delta(latest.metadata), "seed did not produce a delta checkpoint"
with TestClient(app) as client:
state_response = client.get(f"/api/threads/{_THREAD_ID}/state")
assert state_response.status_code == 409, state_response.text
assert "requires delta mode" in state_response.json()["detail"]
assert _THREAD_ID in state_response.json()["detail"]
update_response = client.post(f"/api/threads/{_THREAD_ID}/state", json={"values": {"title": "x"}})
assert update_response.status_code == 409, update_response.text
assert "requires delta mode" in update_response.json()["detail"]
history_response = client.post(f"/api/threads/{_THREAD_ID}/history", json={"limit": 10})
assert history_response.status_code == 409, history_response.text
assert "requires delta mode" in history_response.json()["detail"]
thread_response = client.get(f"/api/threads/{_THREAD_ID}")
assert thread_response.status_code == 409, thread_response.text
assert "requires delta mode" in thread_response.json()["detail"]
def test_full_mode_state_reads_degrade_to_raw_checkpointer_when_factory_fails(_stub_app_config, monkeypatch: pytest.MonkeyPatch) -> None:
"""Full-mode read endpoints survive a broken agent factory.
Full-mode checkpoints persist complete channel_values, so when the agent
factory cannot build the graph (bad model config, MCP server down), state
reads degrade to raw checkpointer reads instead of 500ing. The fail-closed
delta gate must still apply on the degraded path.
"""
app = make_authed_test_app()
store = InMemoryStore()
checkpointer = InMemorySaver()
app.state.store = store
app.state.checkpointer = checkpointer
app.state.thread_store.get = AsyncMock(return_value=None)
app.state.checkpoint_channel_mode = "full"
app.state.run_event_store = SimpleNamespace()
app.include_router(threads.router)
full_graph = _build_reply_graph("full", checkpointer)
config: dict[str, Any] = {"configurable": {"thread_id": _THREAD_ID}}
inject_checkpoint_mode(config, "full")
for i in range(2):
asyncio.run(full_graph.ainvoke({"messages": [HumanMessage(content=f"question-{i}", id=f"h{i}")]}, config))
latest = asyncio.run(checkpointer.aget_tuple(config))
assert latest is not None
latest_created_at = latest.checkpoint["ts"]
def _broken_factory(assistant_id=None):
def _factory(config):
raise RuntimeError("model config broken")
return _factory
monkeypatch.setattr(gateway_services, "resolve_agent_factory", _broken_factory)
with TestClient(app) as client:
state_response = client.get(f"/api/threads/{_THREAD_ID}/state")
assert state_response.status_code == 200, state_response.text
values = state_response.json()["values"]
assert state_response.json()["created_at"] == latest_created_at
assert state_response.json()["checkpoint"]["ts"] == latest_created_at
assert _message_wire_shape(values["messages"]) == [
("human", "question-0", "h0"),
("ai", "answer-1", "a1"),
("human", "question-1", "h1"),
("ai", "answer-3", "a3"),
]
# next/tasks are not derivable without the compiled graph.
assert state_response.json()["next"] == []
history_response = client.post(f"/api/threads/{_THREAD_ID}/history", json={"limit": 10})
assert history_response.status_code == 200, history_response.text
entries = history_response.json()
assert len(entries) >= 2
assert all(entry["created_at"] for entry in entries)
# History pagination: config.checkpoint_id is the *inclusive* anchor
# (pregel semantics), so the degraded path must include it too.
anchor_id = entries[1]["checkpoint_id"]
paged_response = client.post(f"/api/threads/{_THREAD_ID}/history", json={"limit": 10, "before": anchor_id})
assert paged_response.status_code == 200, paged_response.text
assert paged_response.json()[0]["checkpoint_id"] == anchor_id
app.state.thread_store.get = AsyncMock(
return_value={
"thread_id": _THREAD_ID,
"assistant_id": None,
"status": "interrupted",
"created_at": latest_created_at,
"updated_at": latest_created_at,
"metadata": {},
}
)
thread_response = client.get(f"/api/threads/{_THREAD_ID}")
assert thread_response.status_code == 200, thread_response.text
assert thread_response.json()["status"] == "interrupted"
assert _message_wire_shape(thread_response.json()["values"]["messages"]) == [
("human", "question-0", "h0"),
("ai", "answer-1", "a1"),
("human", "question-1", "h1"),
("ai", "answer-3", "a3"),
]
# The fail-closed gate still applies on the degraded path: a delta
# checkpoint is a precise 409, never silently served as partial state.
delta_graph = _build_reply_graph("delta", checkpointer)
delta_config: dict[str, Any] = {"configurable": {"thread_id": "thread-degraded-delta"}}
inject_checkpoint_mode(delta_config, "delta")
asyncio.run(delta_graph.ainvoke({"messages": [HumanMessage(content="q", id="h0")]}, delta_config))
delta_response = client.get("/api/threads/thread-degraded-delta/state")
assert delta_response.status_code == 409, delta_response.text
assert "requires delta mode" in delta_response.json()["detail"]
def test_mutation_accessor_fails_closed_when_thread_metadata_lookup_fails(monkeypatch: pytest.MonkeyPatch) -> None:
app = make_authed_test_app()
app.state.thread_store.get = AsyncMock(side_effect=RuntimeError("metadata store unavailable"))
resolve_factory = AsyncMock()
monkeypatch.setattr(gateway_services, "resolve_agent_factory", resolve_factory)
with pytest.raises(RuntimeError, match="metadata store unavailable"):
asyncio.run(
gateway_services.build_thread_checkpoint_state_mutation_accessor(
SimpleNamespace(app=app),
thread_id="custom-assistant-thread",
as_node="manual_state_update",
)
)
resolve_factory.assert_not_called()