* fix(frontend): preserve message order during long runs * test(frontend): fix history pagination regression mock * fix(frontend): validate thread history sequences
23 KiB
AGENTS.md
This file provides guidance to AI coding agents (Claude Code, Codex, and others) when working with the DeerFlow frontend. It is the source of truth; the sibling CLAUDE.md imports it via @AGENTS.md.
Project Overview
DeerFlow Frontend is a Next.js 16 web interface for an AI agent system. It communicates with a LangGraph-based backend to provide thread-based AI conversations with streaming responses, artifacts, and a skills/tools system.
Stack: Next.js 16, React 19, TypeScript 5.8, Tailwind CSS 4, pnpm 10.26.2. Requires Node.js 22+ and pnpm 10.26.2+.
Core dependencies
- LangGraph SDK (
@langchain/langgraph-sdk^1.5.3) — Agent orchestration and streaming - LangChain Core (
@langchain/core^1.1.15) — Fundamental AI building blocks - TanStack Query (
@tanstack/react-query^5.90.17) — Server state management - UI: Shadcn UI, MagicUI, React Bits, and Vercel AI SDK elements (generated from registries — see Code Style)
Commands
| Command | Purpose |
|---|---|
pnpm dev |
Dev server with Turbopack (http://localhost:3000) |
pnpm build |
Production build |
pnpm check |
Lint + type check (run before committing) |
pnpm lint |
ESLint only |
pnpm lint:fix |
ESLint with auto-fix |
pnpm format |
Prettier check (pnpm format:write to apply) |
pnpm test |
Run unit tests with Rstest |
pnpm test:e2e |
Run E2E tests with Playwright (Chromium) |
pnpm typecheck |
TypeScript type check (tsc --noEmit) |
pnpm start |
Start production server |
Unit tests live under tests/unit/ and mirror the src/ layout (e.g., tests/unit/core/api/stream-mode.test.ts tests src/core/api/stream-mode.ts). Powered by Rstest; import source modules via the @/ path alias.
Rstest runs them as two projects (rstest.config.ts). *.test.ts / *.test.tsx run in a plain node environment — that is nearly the whole suite, and it is the default for anything that is pure logic. *.dom.test.ts / *.dom.test.tsx run in happy-dom, for tests that need a document: hooks driven through renderHook from @testing-library/react, and components. Keep the split — a DOM environment costs roughly 3x the runtime of the node suite, so tests that do not render should not opt into it. A hook whose behavior only exists under real React (effect ordering, cleanup on unmount, re-render on store change) belongs in a .dom.test.* file rather than a node test that mocks react itself.
E2E tests live under tests/e2e/ and use Playwright with Chromium. They mock all backend APIs via page.route() network interception and test real page interactions (navigation, chat input, streaming responses). Config: playwright.config.ts.
Architecture
Frontend (Next.js) ──▶ LangGraph SDK ──▶ LangGraph Backend (lead_agent)
├── Sub-Agents
└── Tools & Skills
The frontend is a stateful chat application. Users create threads (conversations), send messages, set thread-scoped /goal completion conditions, and receive streamed AI responses. The backend orchestrates agents that can produce artifacts (files/code), todos, and goal state updates.
Source Layout (src/)
app/— Next.js App Router. Routes include/(landing),/workspace/chats/[thread_id](chat),/workspace/agents/[agent_name]and/workspace/agents/new(custom agents),/blog/…, the(auth)/{login,setup,auth/callback}flow,/[lang]/docs/…, and/api/…route handlers (e.g./api/memory).components/— React components:ui/— Shadcn UI primitives (auto-generated, ESLint-ignored)ai-elements/— Vercel AI SDK elements (auto-generated, ESLint-ignored)workspace/— Chat page components (messages, artifacts, settings)landing/— Landing page sectionsdocs/— Docs / MDX rendering components
core/— Business logic, the heart of the app. Domains includethreads/(creation, streaming, state),api/(LangGraph client singleton),agents/(custom agents),auth/(authentication),artifacts/,channels/(IM connections),integrations/(managed third-party integration status/install clients such as Lark CLI),i18n/(en-US, zh-CN),settings/,memory/,skills/,messages/,mcp/,models/,input-polish/(pre-send draft rewrite API),voice-input/(browser speech-recognition helpers),suggestions/,tasks/,todos/,tools/,workspace-changes/(run-scoped changed-file summaries and diff fetching),config/,notification/,blog/, plus rendering helpers (rehype/,streamdown/) andutils/.hooks/— Shared React hookslib/— Utilities (cn()from clsx + tailwind-merge)content/— MDX content (blog posts, docs) rendered by the appstyles/— Global CSS with Tailwind v4@importsyntax and CSS variables for themingtypings/— Ambient TypeScript declarations- Root files:
env.js(env validation),mdx-components.ts(MDX component map)
Data Flow
- Optional composer helpers such as
core/input-polishcan rewrite the local draft before submission, andcore/voice-inputcan transcribe browser microphone input into that same local draft; confirmed user input then flows to thread hooks (core/threads/hooks.ts) → LangGraph SDK streaming - Stream events update thread state (messages, artifacts, todos, goal)
File-tool artifact auto-open work must run in an effect with timer cleanup; never schedule timers while rendering streamed
write_fileorstr_replaceupdates. useThreadHistoryloads persisted conversation pages fromGET /api/threads/{id}/messages/page, preserving the backend's thread-global eventseq; rendering overlays checkpoint/live copies at their matching canonical identities (a summarized checkpoint may contain a protected early input plus a recent tail). Context-compaction rescue diffs every retained visible identity rather than slicing at the first anchor, and keeps a run-scoped ledger of committed visible messages so replacement updates and repeated rolling checkpoint windows cannot erase an already displayed step. The resolver suppresses checkpoint/transient prefixes whose canonical position is still behind an unloaded cursor page instead of collapsing that unknown gap before a recent anchor, then adds optimistic messages without timestamp re-sorting. History invalidation preserves already-loaded pages so their established ordering positions are not discarded.- Stop actions call the LangGraph SDK stream stop path;
core/threads/hooks.tsinvalidates current-thread, thread-history, token-usage, and sidebar/search caches immediately and schedules one follow-up refetch because SDK stop may finish via abort + fire-and-forget cancel before backend title finalization commits - TanStack Query manages server state; localStorage stores user settings
- Components subscribe to thread state and render updates
Run duration is run-scoped UI metadata even though the compatibility field additional_kwargs.turn_duration is repeated on historical AI messages. core/messages/run-duration.ts folds those copies into one display anchored after the run's last visible message group. MessageList owns the temporary client-side duration for a just-completed live turn until authoritative history arrives. The duration is total run wall-clock time, not per-message reasoning time; reasoning disclosure and run activity/duration are rendered separately.
Composer drafts are tab-scoped browser state. core/threads/composer-draft.ts stores only text plus the selected slash-skill name in sessionStorage, keyed by user, agent, and logical conversation scope. New-chat pages pass the stable scope "new" because their runtime threadId is a fresh UUID on every reload; established conversations use their real thread ID. InputBox waits for enabled skills before restoring a skill chip, degrades a missing/disabled skill back to editable slash text, and clears the stored draft through SendMessageOptions.onSent only after the send passes the in-flight guard. Attachments, sidecar quotes, voice state, and polish undo state are not persisted.
Auth UI note: the login page's "keep me signed in" option submits only remember_me to the Gateway and may persist only the email address through core/auth/remember-login.ts. Passwords and tokens must never be stored in frontend storage; the HttpOnly access_token and readable csrf_token cookies remain Gateway-owned.
/goal and /compact are built-in composer commands, not skill activations. src/components/workspace/input-box.tsx intercepts /goal, /goal clear, and /goal <condition> before normal chat submission, calling Gateway GET/PUT/DELETE /api/threads/{thread_id}/goal. Setting /goal <condition> also submits the condition text as the next user task so the agent starts running immediately; status and clear do not start a run. Goal and compact requests are tied to the current threadId with an AbortController, so switching threads or unmounting the composer aborts in-flight requests and stale responses cannot update the new thread's composer state. The chat pages render GoalStatus above the composer from AgentThreadState.goal, with local optimistic state until the next stream values update arrives. /compact calls POST /api/threads/{thread_id}/compact to summarize older active context while leaving the full visible chat history intact; it is skipped on new/empty threads and blocked server-side while a run is in flight. Thread rename uses the same serialized state-write route; the rename dialog stays open and surfaces the server error when an active run returns 409.
Human input requests are a structured message protocol layered on normal chat history. The backend writes request payloads to ToolMessage.artifact.human_input, src/core/messages/human-input.ts owns the runtime validators/types, and src/components/workspace/messages/human-input-card.tsx renders the reusable card. The protocol is versioned on the request side only: v1 covers free_text / choice_with_other, and v2 adds form (typed fields — text/textarea/number/select/multi_select/checkbox/date — with required-field validation in the card). Replies deliberately stay on the v1 response protocol: the form card submits a response_kind: "text" reply whose value is the human-readable summary plus one JSON block keyed by stable field names (buildHumanInputFormSubmissionValue — the readable part alone is ambiguous because labels/values may contain the separators), so the model can reconstruct the submitted mapping without a structured response kind. The validators reject unknown versions/modes (and field names colliding with JS Object.prototype members) so future protocol bumps degrade to the plain-text ToolMessage fallback rather than rendering a broken card. Form values are read through own-property access only (readHumanInputFormValue); select fields stay controlled from their empty-string placeholder state through selection; checkbox fields are native <input type="checkbox"> controls seeded to an explicit false (buildInitialHumanInputFormValues) so an untouched checkbox submits as "no" while a required checkbox keeps must-agree semantics (no HTML required attribute — native constraint validation would intercept the custom submit path), and form controls carry label/htmlFor, aria-required plus a visually-hidden localized "required" marker, and aria-invalid/error associations whose error node stays mounted while any field is still invalid. Legacy-fallback closure: deriveHumanInputThreadState treats a visible plain human message as answering the latest unanswered request opened before it (only the latest — nothing guarantees a single outstanding request across runs, and closing all would silently swallow older decisions; an older request left open simply becomes the active card again) — an old v1-only frontend degrades a v2 request to text and the user replies through the normal composer without response metadata, and without this rule an upgraded frontend would see that request as still open and lock the composer. MessageList owns answered/latest/pending state for visible cards, but derives answered responses from raw thread.messages because replies are hidden; pending cards clear when the hidden reply appears, when dispatch is dropped, or when a new thread.error reports an async stream failure. Page-level submit callbacks must send a normal human message and put hide_from_ui: true plus the response payload in the fourth sendMessage(..., options) argument as options.additionalKwargs; the third argument remains run context such as { agent_name }. Composer entry points should disable normal bottom input while hasOpenHumanInputRequest(...) is true so users answer through the card and preserve response metadata.
Tool-calling AI messages can contain user-visible text as well as tool_calls. core/messages/utils.ts keeps these turns in an assistant:processing group, and components/workspace/messages/message-group.tsx must render the visible text as a processing step instead of treating the message as only tool metadata. This preserves provider text such as error explanations or "trying another approach" notes during tool-heavy runs.
Edit-and-rerun is deliberately latest-turn-only. core/messages/utils.ts::getLatestEditableTurn() exposes a human turn only when the transcript is idle and the most recent visible turn ends in a terminal assistant message. core/threads/hooks.ts::editAndRegenerateMessage() calls POST /api/threads/{id}/runs/edit-regenerate/prepare, submits the returned replacement message/checkpoint/metadata through the same LangGraph stream path as regenerate, optimistically hides the superseded message ids, and clears the optimistic replacement once the persisted replacement arrives.
MessageGroup builds its tool-result and browser-preview lookups once per processing group before converting messages to steps. The lookup preserves the first non-empty result and first screenshot-bearing browser view for each tool-call ID, matching the streamed-message display semantics without repeatedly scanning the full group for every tool call.
Key Patterns
- Server Components by default,
"use client"only for interactive components - Thread hooks (
useThreadStream,useSubmitThread,useThreads) are the primary API interface - Thread routes — construct Web UI chat paths through
core/threads/utils.ts::pathOfThread(), which percent-encodes both custom agent names and thread IDs before inserting them into route segments - LangGraph client is a singleton obtained via
getAPIClient()incore/api/ - Run stream options are sanitized by
core/api/stream-mode.ts: the Gateway-supported set isvalues,messages-tuple,updates,debug,tasks,checkpoints, andcustom; any request containing an unsupported mode throws before HTTP instead of being partially forwarded or silently defaulting tovalues.streamResumableis retained by thread hooks only for SDK-side reconnect bookkeeping but stripped before the HTTP request because the Gateway does not accept that request option; actual replay uses the SSELast-Event-IDcursor. Keep this boundary aligned with the backend request schema;messagesandeventsare not supported and must not be forwarded. - SSE replay gaps are handled in
core/api/api-client.ts, which wraps both initial and joined run streams because the upstream SDK ignores unknown event names. An id-less backendgapcontrol frame clears stale reconnect metadata, emits an internalstream_replay_gapcustom event, reloads durable thread values, and rejoins after the server-provided retained tail, with up to five recovery rejoins after the original stream (six total stream calls on an all-gap exhaustion path). The wrapper remains a lazy async iterable because the SDK consumes it withfor await.core/threads/hooks.tsclears optimistic/transient/subtask state, invalidates durable history caches, and shows the localized recovery warning; never let a gap fall through as a normal stream finish or cancel the still-running backend run. - Streaming Markdown rendering is owned by
core/streamdown: Streamdown'sanimated/isAnimatingAPI handles incremental word animation, while the sharedstreamdownRenderingPluginsconfig registers the named code-highlighting and Mermaid plugins required by Streamdown 2.5. Keep wrappers and derived configs wired to that shared object; do not reintroduce a rehype plugin that wraps every word, because reparsing a growing block remounts old words and replays their animation. - Environment validation uses
@t3-oss/env-nextjswith Zod schemas (src/env.js). Skip withSKIP_ENV_VALIDATION=1 - Subtask step history and runtime metadata (
core/tasks/) — the subtask card shows a subagent's full step timeline (#3779): its assistant reasoning turns interleaved with the tools it ran.Subtask.steps[]is accumulated live fromtask_runningevents (appended viamergeSteps, not overwritten) and backfilled on expand for historical runs byfetchSubtaskSteps, which pages the events endpoint scoped to one task (GET/runs/{runId}/events?event_types=subagent.step&task_id=…&after_seq=…) until a short page, so the run-wide limit can't truncate the timeline.task_startedcarries the effectivemodel_name;task_runningcarries a cumulative usage snapshot after each completed LLM call.core/tasks/lifecycle.tsnormalizes these additive events, andcomputeNextSubtaskkeeps the largest cumulative total so replayed or late SSE frames cannot double-count or roll the folded card backward. Terminal ToolMessage metadata (subagent_model_name/subagent_token_usage) restores the same values from normal history after reload; no per-card event fetch is needed.core/tasks/steps.tsis the pure step model:messageToStep(live),eventsToSteps(reload),mergeSteps(dedup bymessage_index), andstepsForDisplay(what the card renders — keeps tool steps + AI steps with text, drops the trailing final-answer AI step when completed since it's shown asresult).core/tasks/context.tsx'suseUpdateSubtaskapplies updates against atasksRefmirroring the latest state (not a closure snapshot), so a late-resolvingfetchSubtaskStepsbackfill merges into current state instead of clobbering SSE steps or sibling subtasks that arrived meanwhile. The owningrun_idis carried onto history content messages inbuildVisibleHistoryMessagesso the card can resolve the events endpoint.
Interaction Ownership
src/app/workspace/chats/[thread_id]/page.tsxowns composer busy-state wiring.src/app/workspace/chats/[thread_id]/page.tsxowns branch-from-turn submission and navigation; sidecarMessageListinstances do not receive the branch action.src/app/workspace/chats/[thread_id]/page.tsxandsrc/app/workspace/agents/[agent_name]/chats/[thread_id]/page.tsxown edit-and-rerun submission wiring because the page must preserve normal/custom-agent run context;MessageListonly detects the latest editable user turn and renders the inline editor.src/app/workspace/chats/[thread_id]/page.tsxgates the Workspace Browser trigger and browser right panel on/api/features -> browser_control.enabled; default/failed feature discovery hides the browser control so optional backend installs do not show a dead Live socket.src/app/workspace/chats/[thread_id]/page.tsxandsrc/app/workspace/agents/[agent_name]/chats/[thread_id]/page.tsxown active-goal display state for their composer overlays.src/components/workspace/messages/message-list.tsxowns human-input card answered/latest/pending gating; entry pages only translate a submitted card response intosendMessagecalls.src/components/workspace/browser-view/browser-view-panel.tsxforwards each physical pointer click as oneclickinput; do not also emitdown/upfor the same gesture because the remote Playwright click would run twice.src/core/threads/hooks.tsowns pre-submit upload state and thread submission.src/components/workspace/chats/chat-box.tsxowns the desktop right-panel layout, and all three right panels (artifacts, sidecar, browser) share oneResizablePanelGroup— do not fork a non-resizable branch per panel kind, which is how the artifacts divider silently lost its drag handle (#4465). Open/close iscollapse()/resize()on the side panel's imperative handle, not conditional rendering, so the width can animate. Three constraints hold that together: the size transition is applied from the group as[&>[data-panel]]:transition-[flex-grow]because the sized flex item is the library's own[data-panel]element rather than the childclassNamelands on; it is applied only while an open/close is in flight, so a drag is not interpolated frame by frame; and during the animation the panel content is held at its final width incqwand clipped, because a reflowing message list re-runs its scroll-to-bottom (pinned bytests/e2e/sidecar-chat.spec.ts's no-animated-scroll test) and a re-wrapping composer changes which responsive labels it shows.
Code Style
- Imports: Enforced ordering (builtin → external → internal → parent → sibling), alphabetized, newlines between groups. Use inline type imports:
import { type Foo }. - Unused variables: Prefix with
_. - Class names: Use
cn()from@/lib/utilsfor conditional Tailwind classes. - Path alias:
@/*maps tosrc/*. - Components:
ui/andai-elements/are generated from registries (Shadcn, MagicUI, React Bits, Vercel AI SDK) — don't manually edit these.
Environment
Backend API URLs are optional; an nginx proxy is used by default:
NEXT_PUBLIC_BACKEND_BASE_URL=http://localhost:8001
NEXT_PUBLIC_LANGGRAPH_BASE_URL=http://localhost:8001/api
Leave these unset for the standard make dev / Docker flow, where nginx serves the public /api/langgraph/* prefix and rewrites it to Gateway's native /api/* routes.
To reach a dev server on anything other than localhost — a LAN address, or a proxied hostname — list the host in DEER_FLOW_DEV_ALLOWED_ORIGINS (comma-separated; a full URL is reduced to its host). It feeds Next's allowedDevOrigins, which gates /_next/*, fonts, and HMR. Without it those requests get a 403 and the page renders server-side but never hydrates, so nothing on it — including the login form — responds. Development only; production builds ignore it.
Resources
Contributing
When adding features:
- Follow the established
src/structure - Add TypeScript types and proper error handling
- Write unit tests under
tests/unit/(pnpm test) and E2E tests undertests/e2e/(pnpm test:e2e) - Run
pnpm checkbefore committing - Update this
AGENTS.mdwhen architecture, commands, or conventions change