deer-flow/docs/database-forward-revision-recovery.md
Zeren Wang 5951c89b5b
feat(projects): project workspaces with scoped chats and thread membership (#5265)
* feat(projects): project workspaces with scoped chats and thread membership

Backend:
- projects table model and migration; fail-closed ProjectRepository with
  ownership checks, CRUD/archive/restore/delete router, and atomic thread
  move between projects
- threads_meta.project_id column exposed as reserved deerflow_project_id
  metadata; project-aware thread create/search with pagination bounds and
  membership echoed in create responses
- first-run admission assigns the project only at genuine first run, seeded
  at write time and dropped when invalid; serialized against project
  deletion and thread assignment
- branch creation inherits the source thread's project membership (an
  archived/deleted project degrades the branch to unassigned instead of
  failing the request)

Frontend:
- projects data layer, thread move API, and sidebar projects section with
  flat/grouped modes, archived-project threads, and stable virtual-list
  offsets
- project detail page with project-scoped new chat
  (/workspace/chats/new?project=) and paginated thread list
- move-to-project thread menu, new-project dialog, archived-project gates
- project-scoped new chats pre-create the thread with membership before the
  first submit or /goal set, so runs never proceed outside the project
- goal-set preparation is fenced against conversation switches: a stale
  continuation is dropped instead of saving the goal or launching the
  abandoned submission on the newly opened conversation
- project thread lists join thread lifecycle invalidations (stop, pin) so
  an open project page never keeps stale titles, recency, or pagination

* fix(chats): keep archive undo toast when the sidebar row unmounts

The archive success toast was fired from per-mutate callbacks passed to
mutation.mutate. React Query drops those handlers when the observer
component unmounts before the mutation settles; archiving the open chat
removes its sidebar row mid-flight, so the undo toast never appeared and
the e2e archive-undo test timed out waiting for it.

Move the success/error handlers to the mutation level (useArchiveThread
options, same pattern as useMoveThreadToProject) where callbacks are
delivered even after the originating row unmounts.

* fix(projects): pin project thread listing contract and exclude archived chats

GET /api/projects/{id}/threads returned the thread store row verbatim
(list[dict], no response_model): user_id/assistant_id leaked, any future
ThreadMetaRow column would auto-leak, and the OpenAPI schema was empty.
Return a narrow ProjectThreadResponse (the exact fields ProjectThread
declares) with the same metadata secret redaction the surrounding thread
endpoints get from _MetadataRedactingResponse.

The listing also ran search() without the archived filter, so a retired
chat rendered as a normal row on the project page while the sidebar hid
it. Search archived=False to mirror the sidebar's archived:false lists;
restore stays on the global Archived tab.

Both regressions pinned by new router tests: wire-shape allowlist and
archived-member exclusion.

* docs(migrations): record the 0019/0020 chain against the bootstrap reservation

The tree now chains 0018 -> 0019_projects -> 0020_threads_meta_project_id,
so migrations/AGENTS.md was stale twice over: the revision index stopped at
0018 and the rolling-forward section still claimed the tree 'deliberately
remains at 0018'.

Document the new head and record the intentional numeric-prefix reuse of
0019: 0019_projects is in-chain while 0019_thread_incarnations stays the
reserved, allowlisted out-of-tree rollout id. The owning rollout revision
must re-parent onto this tree's head when it merges so alembic never sees
two heads off 0018; bootstrap.py now cross-references that note next to
_FORWARD_COMPATIBLE_REVISION.

* fix(chats): invalidate project thread lists on archive/restore

useArchiveThread refreshed the infinite sidebar cache, threads/search and
the per-thread metadata cache but not the project-scoped list
([...PROJECTS_QUERY_KEY, 'threads', id]) this PR adds — the one thread
mutation not wired to that key, after usePinThread, useRenameThread,
useDeleteThread, useMoveThreadToProject and invalidateStoppedThreadCaches.

An archive from a sidebar row while a project page is open therefore left
the archived chat rendered as a normal row until remount (and undo left it
missing). Invalidate the prefix in the mutation-level success handler.

Regression test asserts the project-list prefix is invalidated on success.

* fix(projects): fetch project discovery only in grouped sidebar mode

RecentChatList mounted two useProjects queries per sidebar render, but
knownProjectIds is consumed only by the grouped-mode exclusion filter; in
the default flat mode every page load paid two GET /api/projects?status=
round trips for data nothing read. Gate both queries on grouped mode —
GroupedProjectList fetches the same keys when the toggle is on and
TanStack dedupes the observers.

Also set retry: false on useProject: a deleted or foreign project 404s
deterministically, and the page renders a dedicated not-found state for
it, so the default 1s/2s/4s retry backoff kept deep links in 'loading'
for ~7s before that state appeared. Matches useThreadMetadata /
useThreadTokenUsage.

* fix(threads): fail closed on project-scoped create in memory mode

MemoryThreadMetaStore.create accepted project_id and silently ignored it,
making memory mode the one membership path that fails open: POST
/api/threads with a project id returned 200 and the run started
unassigned, violating the invariant that a run never proceeds outside the
selected project (the SQL store raises ProjectNotAssignableError inside
the insert transaction for the same request).

Raise ProjectNotAssignableError whenever project_id is present so the
router's existing 404 mapping applies, the frontend keeps the composer
text for a retry, and memory mode behaves exactly like SQL mode.
set_project already reports rejection; create now matches it.

Store-level test (raises, nothing persisted, project filter stays empty,
unscoped creates still work) plus a router-level test asserting the 404
and that no row is left behind.

* fix(projects): window the project page thread list

ProjectThreadsSection rendered every loaded page as a plain Link row, so a
long-lived project accumulated unbounded DOM on the page's scroll surface:
each load-more appended another 100 rows and every formatTimeAgo tick
re-rendered the whole list.

Reuse VirtualThreadList (now generic over any row shape with a
thread_id), pointing its scroll parent at this page's ScrollArea viewport
via the shared [data-slot="scroll-area-viewport"] selector used by
/workspace/chats; under the 60-row threshold it falls back to the plain
render, so small projects are unchanged.

* fix(projects): restore row dividers and pin them with a render test

The row class template literal concatenated transition-colors directly
with the conditional border-b token, so non-final rows rendered the
invalid class 'transition-colorsborder-b' and lost both the divider and
the transition. Compose the row classes with cn() and a boolean guard
instead.

The section moved out of page.tsx into a testable component so the row
markup finally has coverage: a DOM test asserts every row except the
final data row carries border-b (index-based, not last: — correct under
virtualization where the last mounted row is not the last data row), and
the untitled fallback plus load-more button render for a partial page.

* fix(projects): validate forward schemas and fence membership reads
2026-09-08 17:00:26 +08:00

3.6 KiB

Recovering the original thread-incarnation database revision

The Projects build requires projects and threads_meta.project_id. An older deployment may have stamped 0019_thread_incarnations on a database containing only 0018_oauth_identity_pg_partial plus two nullable VARCHAR(32) columns: threads_meta.incarnation and mcp_tasks.thread_incarnation. That shape cannot serve this build's repositories. Startup now rejects it without changing the schema or revision, and reports the missing tables/columns.

Normal databases on this tree's known migration chain upgrade automatically. The procedure below is only for the exact original incarnation rollout shape. An incarnation-stamped database that already has all current ORM tables and columns can still use the audited compatibility exception without re-stamping.

Offline migration

  1. Stop every Gateway, scheduler, and other process writing to the database. Take a restorable database backup and rehearse these steps on a copy.

  2. Verify there is exactly one alembic_version row, containing 0019_thread_incarnations. Inspect the owning deployment's migration and actual database schema: it must be the local 0018 schema plus only the two nullable columns above, with no added defaults, constraints, tables, indexes, or data backfills. Neither projects nor threads_meta.project_id may already exist for this recovery path. If the shape differs, use a migration reviewed for that deployment; do not use the commands below.

  3. From this checkout's backend/, set DEERFLOW_RECOVERY_DATABASE_URL to the target async SQLAlchemy URL (sqlite+aiosqlite:////absolute/path/database.db or postgresql+asyncpg://…). For Postgres, also set DEERFLOW_RECOVERY_POSTGRES_SCHEMA to the configured application schema, if one is used. Keep credentials out of shell history.

  4. Rebase the version marker to the verified common parent and run the normal migrations. purge=True is necessary because this tree does not contain the out-of-tree revision; it replaces the version row, not application data.

    uv run python - <<'PY'
    import asyncio
    import os
    
    from alembic import command
    from sqlalchemy.ext.asyncio import create_async_engine
    
    from deerflow.persistence.bootstrap import _get_alembic_config
    
    engine = create_async_engine(os.environ["DEERFLOW_RECOVERY_DATABASE_URL"])
    cfg = _get_alembic_config(
        engine,
        postgres_schema=os.environ.get("DEERFLOW_RECOVERY_POSTGRES_SCHEMA", ""),
    )
    command.stamp(cfg, "0018_oauth_identity_pg_partial", purge=True)
    command.upgrade(cfg, "head")
    asyncio.run(engine.dispose())
    PY
    

    This applies 0019_projects and 0020_threads_meta_project_id, preserving the two incarnation columns and their existing values. Do not stamp directly to head: that would skip the DDL and reproduce the missing-column failure.

  5. Confirm the version is 0020_threads_meta_project_id, the project table and membership column/index exist, and existing incarnation values are retained. Start this build, verify existing conversations load and a new conversation can be created, then resume service. Do not restart older binaries that cannot read this tree's head revision.

Bootstrap never performs this re-stamp itself. The regression in backend/tests/test_persistence_forward_revision_compat.py constructs the original schema, verifies startup rejection, and exercises the recovery while checking thread reads/inserts and preservation of incarnation data. The future incarnation migration must chain from the current local head and handle these already-present nullable columns idempotently.