mirror of
https://github.com/bytedance/deer-flow.git
synced 2026-09-21 20:16:18 +00:00
* feat: show real-time context window usage in chat UI (#3125) Adds a `context_usage` block to `GET /api/threads/{id}/token-usage` (token count from the live checkpoint, the thread model's `context_window`, and a percentage), introduces a new `ModelConfig.context_window` distinct from the per-call `max_tokens` output cap, and surfaces the percentage in the chat header — inside `TokenUsageIndicator` when token-usage tracking is on, or as a standalone badge when it's off so context capacity stays visible independent of cost tracking. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * feat: per-category breakdown for context window usage Replace the single-number context_usage payload with a Claude-Code-style breakdown — messages, system prompt, skills, system/MCP tools (active + deferred), custom agents, memory injection, autocompact buffer, and free space — and surface it in the chat UI with a segmented progress bar and per-row table. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * docs(config): document context_window across model examples Add `context_window` to every example model in config.example.yaml so the new chat-UI "% context used" indicator works out of the box for whichever example a user adopts. Each value is the published default at the time of writing; users are pointed at the official model spec to verify. Bumps config_version to 11 so `make config-upgrade` flags outdated user configs. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * style: ruff format (line-length 240) No behavior change — collapses two multi-line expressions that fit on one line under the project's 240-char limit. Picked up by `make format`. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * review: address Copilot bot comments on #3183 - token-usage-indicator: switch `{contextPercentage && (...)}` to an explicit `!= null` check. (The string `"0"` is actually truthy in JS so the original code wasn't buggy, but the explicit check is clearer.) - context-usage-breakdown: drop the `useMemo` around segments/totals — the computation is O(n) over a handful of rows and the previous memo deps omitted `t.contextUsage.categories`, so the bar's tooltips/aria-labels could stay in the old language after a locale switch. - context_usage._split_tools: snapshot MCP names from `get_cached_mcp_tools()` directly instead of re-reading `extensions_config.json` after `get_available_tools()` already loaded it. Removes redundant file I/O on every `/token-usage` poll. (`get_available_tools()` still emits its own INFO logs — silencing those is out of scope here.) Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * style(frontend): prettier --write context-usage-breakdown CI's `pnpm format` (prettier --check) caught two lines previously formatted by hand. Collapses one comma to fit on one line; no behavior change. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * fix(gateway): correct context-usage breakdown + add exact token counting The context-usage indicator shipped two bugs that silently zeroed whole breakdown rows (both caught by try/except, so the feature looked alive but produced wrong numbers): 1. _count_system_prompt passed app_config= to get_deferred_tools_prompt_section, which only accepts deferred_names -> TypeError swallowed -> system_prompt row always 0, and used_tokens/percentage undercounted by the full prompt. Also subtracted the deferred section twice (the rendered prompt already excluded it). Fix: derive deferred names deterministically and pass them to apply_prompt_template; drop the redundant subtraction. 2. _split_tools imported a non-existent get_deferred_registry -> ImportError swallowed -> all four tool-category rows always 0. Fix: classify via the public is_mcp_tool predicate + tool_search.enabled (mirrors build_deferred_tool_setup); the MCP tag is set by get_available_tools. Added token_usage.counting (approximate|exact). 'exact' routes text/schema/ message counting through the model tokenizer (tiktoken cl100k_base) via the existing memory-module machinery (lazy load + cache + cooldown + CJK-aware fallback), so CJK-heavy threads stop being undercounted by chars//4. Regression + e2e tests added; 6621 backend tests pass. * fix(gateway): harden context usage accounting * fix(gateway): count promoted MCP tools as active in context usage Promoted tools (deferred MCP tools the thread has fetched via tool_search) have their full schema bound on every subsequent turn by DeferredToolFilterMiddleware, so they consume context like any active tool. The breakdown previously left them in the reserved *_deferred rows, under- counting the thread's used_tokens. Classification now treats a tool as deferred only when tool_search is enabled, it is MCP-sourced, AND it has not been promoted. The promoted set is read from the checkpoint's channel_values and scoped by catalog hash — matching the runtime middleware, so a stale promotion from MCP-config drift cannot inflate the active count. The static system prompt still lists all deferred tool names (promotions only affect schema binding, not the prompt), so _count_system_prompt's deferred rendering is intentionally left unchanged. 8 new tests cover classification, catalog-hash scoping (match / drift / compute-failure / malformed), and checkpoint extraction. * fix(context): address review feedback * fix(context): count structured message payloads * fix(context): harden usage accounting * fix(config): bump schema for context usage fields * refactor: narrow context usage to core indicator --------- Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com> Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
183 lines
5.6 KiB
TypeScript
183 lines
5.6 KiB
TypeScript
"use client";
|
|
|
|
import type { Message } from "@langchain/langgraph-sdk";
|
|
import { ChevronDownIcon, CoinsIcon } from "lucide-react";
|
|
import { useMemo } from "react";
|
|
|
|
import { Button } from "@/components/ui/button";
|
|
import {
|
|
DropdownMenu,
|
|
DropdownMenuContent,
|
|
DropdownMenuLabel,
|
|
DropdownMenuRadioGroup,
|
|
DropdownMenuRadioItem,
|
|
DropdownMenuSeparator,
|
|
DropdownMenuTrigger,
|
|
} from "@/components/ui/dropdown-menu";
|
|
import { useI18n } from "@/core/i18n/hooks";
|
|
import {
|
|
formatTokenCount,
|
|
selectHeaderTokenUsage,
|
|
type TokenUsage,
|
|
} from "@/core/messages/usage";
|
|
import {
|
|
getTokenUsageViewPreset,
|
|
tokenUsagePreferencesFromPreset,
|
|
type TokenUsagePreferences,
|
|
type TokenUsageViewPreset,
|
|
} from "@/core/messages/usage-model";
|
|
import type { ContextUsage } from "@/core/threads/token-usage";
|
|
import { cn } from "@/lib/utils";
|
|
|
|
import { formatContextUsagePercentage } from "./context-usage-format";
|
|
|
|
interface TokenUsageIndicatorProps {
|
|
threadId?: string;
|
|
messages: Message[];
|
|
pendingMessages?: Message[];
|
|
backendUsage?: TokenUsage | null;
|
|
contextUsage?: ContextUsage | null;
|
|
enabled?: boolean;
|
|
preferences: TokenUsagePreferences;
|
|
onPreferencesChange: (preferences: TokenUsagePreferences) => void;
|
|
className?: string;
|
|
}
|
|
|
|
export function TokenUsageIndicator({
|
|
threadId,
|
|
messages,
|
|
pendingMessages,
|
|
backendUsage,
|
|
contextUsage,
|
|
enabled = false,
|
|
preferences,
|
|
onPreferencesChange,
|
|
className,
|
|
}: TokenUsageIndicatorProps) {
|
|
const { t } = useI18n();
|
|
|
|
const usage = useMemo(
|
|
() =>
|
|
selectHeaderTokenUsage({
|
|
backendUsage: threadId ? backendUsage : null,
|
|
messages,
|
|
pendingMessages,
|
|
}),
|
|
[backendUsage, messages, pendingMessages, threadId],
|
|
);
|
|
const preset = getTokenUsageViewPreset(preferences);
|
|
const contextPercentage = formatContextUsagePercentage(
|
|
contextUsage?.percentage,
|
|
);
|
|
|
|
if (!enabled) {
|
|
return null;
|
|
}
|
|
|
|
return (
|
|
<DropdownMenu>
|
|
<DropdownMenuTrigger asChild>
|
|
<Button
|
|
type="button"
|
|
variant="ghost"
|
|
className={cn(
|
|
"text-muted-foreground bg-background/70 hover:bg-background/90 flex h-auto items-center gap-1.5 rounded-full border px-2 py-1 text-xs font-normal",
|
|
className,
|
|
)}
|
|
>
|
|
<CoinsIcon size={14} />
|
|
<span>{t.tokenUsage.label}</span>
|
|
<span className="font-mono">
|
|
{preferences.headerTotal
|
|
? usage
|
|
? formatTokenCount(usage.totalTokens)
|
|
: "-"
|
|
: t.tokenUsage.presets[presetKeyToTranslationKey(preset)]}
|
|
</span>
|
|
{contextPercentage != null && (
|
|
<span
|
|
className="text-muted-foreground/80 border-l pl-1.5 font-mono"
|
|
aria-label={t.contextUsage.badgeAriaLabel(contextPercentage)}
|
|
>
|
|
{contextPercentage}%
|
|
</span>
|
|
)}
|
|
<ChevronDownIcon className="size-3" />
|
|
</Button>
|
|
</DropdownMenuTrigger>
|
|
<DropdownMenuContent side="bottom" align="end" className="w-80">
|
|
<DropdownMenuLabel>{t.tokenUsage.title}</DropdownMenuLabel>
|
|
<div className="px-2 py-1 text-xs">
|
|
{usage ? (
|
|
<div className="space-y-1">
|
|
<div className="flex justify-between gap-4">
|
|
<span>{t.tokenUsage.input}</span>
|
|
<span className="font-mono">
|
|
{formatTokenCount(usage.inputTokens)}
|
|
</span>
|
|
</div>
|
|
<div className="flex justify-between gap-4">
|
|
<span>{t.tokenUsage.output}</span>
|
|
<span className="font-mono">
|
|
{formatTokenCount(usage.outputTokens)}
|
|
</span>
|
|
</div>
|
|
<div className="border-t pt-1">
|
|
<div className="flex justify-between gap-4">
|
|
<span>{t.tokenUsage.total}</span>
|
|
<span className="font-mono font-medium">
|
|
{formatTokenCount(usage.totalTokens)}
|
|
</span>
|
|
</div>
|
|
</div>
|
|
</div>
|
|
) : (
|
|
<div className="text-muted-foreground">
|
|
{t.tokenUsage.unavailable}
|
|
</div>
|
|
)}
|
|
</div>
|
|
<DropdownMenuSeparator />
|
|
<DropdownMenuLabel>{t.tokenUsage.view}</DropdownMenuLabel>
|
|
<DropdownMenuRadioGroup
|
|
value={preset}
|
|
onValueChange={(value) =>
|
|
onPreferencesChange(
|
|
tokenUsagePreferencesFromPreset(value as TokenUsageViewPreset),
|
|
)
|
|
}
|
|
>
|
|
{(
|
|
["off", "summary", "per_turn", "debug"] as TokenUsageViewPreset[]
|
|
).map((value) => {
|
|
const translationKey = presetKeyToTranslationKey(value);
|
|
return (
|
|
<DropdownMenuRadioItem key={value} value={value}>
|
|
<div className="grid gap-0.5">
|
|
<span>{t.tokenUsage.presets[translationKey]}</span>
|
|
<span className="text-muted-foreground text-xs">
|
|
{t.tokenUsage.presetDescriptions[translationKey]}
|
|
</span>
|
|
</div>
|
|
</DropdownMenuRadioItem>
|
|
);
|
|
})}
|
|
</DropdownMenuRadioGroup>
|
|
<DropdownMenuSeparator />
|
|
<div className="text-muted-foreground px-2 py-2 text-xs leading-relaxed">
|
|
{t.tokenUsage.note}
|
|
</div>
|
|
</DropdownMenuContent>
|
|
</DropdownMenu>
|
|
);
|
|
}
|
|
|
|
function presetKeyToTranslationKey(preset: TokenUsageViewPreset) {
|
|
switch (preset) {
|
|
case "per_turn":
|
|
return "perTurn" as const;
|
|
default:
|
|
return preset;
|
|
}
|
|
}
|