deer-flow/docs/database-forward-revision-recovery.md
Zeren Wang 5951c89b5b
feat(projects): project workspaces with scoped chats and thread membership (#5265)
* feat(projects): project workspaces with scoped chats and thread membership

Backend:
- projects table model and migration; fail-closed ProjectRepository with
  ownership checks, CRUD/archive/restore/delete router, and atomic thread
  move between projects
- threads_meta.project_id column exposed as reserved deerflow_project_id
  metadata; project-aware thread create/search with pagination bounds and
  membership echoed in create responses
- first-run admission assigns the project only at genuine first run, seeded
  at write time and dropped when invalid; serialized against project
  deletion and thread assignment
- branch creation inherits the source thread's project membership (an
  archived/deleted project degrades the branch to unassigned instead of
  failing the request)

Frontend:
- projects data layer, thread move API, and sidebar projects section with
  flat/grouped modes, archived-project threads, and stable virtual-list
  offsets
- project detail page with project-scoped new chat
  (/workspace/chats/new?project=) and paginated thread list
- move-to-project thread menu, new-project dialog, archived-project gates
- project-scoped new chats pre-create the thread with membership before the
  first submit or /goal set, so runs never proceed outside the project
- goal-set preparation is fenced against conversation switches: a stale
  continuation is dropped instead of saving the goal or launching the
  abandoned submission on the newly opened conversation
- project thread lists join thread lifecycle invalidations (stop, pin) so
  an open project page never keeps stale titles, recency, or pagination

* fix(chats): keep archive undo toast when the sidebar row unmounts

The archive success toast was fired from per-mutate callbacks passed to
mutation.mutate. React Query drops those handlers when the observer
component unmounts before the mutation settles; archiving the open chat
removes its sidebar row mid-flight, so the undo toast never appeared and
the e2e archive-undo test timed out waiting for it.

Move the success/error handlers to the mutation level (useArchiveThread
options, same pattern as useMoveThreadToProject) where callbacks are
delivered even after the originating row unmounts.

* fix(projects): pin project thread listing contract and exclude archived chats

GET /api/projects/{id}/threads returned the thread store row verbatim
(list[dict], no response_model): user_id/assistant_id leaked, any future
ThreadMetaRow column would auto-leak, and the OpenAPI schema was empty.
Return a narrow ProjectThreadResponse (the exact fields ProjectThread
declares) with the same metadata secret redaction the surrounding thread
endpoints get from _MetadataRedactingResponse.

The listing also ran search() without the archived filter, so a retired
chat rendered as a normal row on the project page while the sidebar hid
it. Search archived=False to mirror the sidebar's archived:false lists;
restore stays on the global Archived tab.

Both regressions pinned by new router tests: wire-shape allowlist and
archived-member exclusion.

* docs(migrations): record the 0019/0020 chain against the bootstrap reservation

The tree now chains 0018 -> 0019_projects -> 0020_threads_meta_project_id,
so migrations/AGENTS.md was stale twice over: the revision index stopped at
0018 and the rolling-forward section still claimed the tree 'deliberately
remains at 0018'.

Document the new head and record the intentional numeric-prefix reuse of
0019: 0019_projects is in-chain while 0019_thread_incarnations stays the
reserved, allowlisted out-of-tree rollout id. The owning rollout revision
must re-parent onto this tree's head when it merges so alembic never sees
two heads off 0018; bootstrap.py now cross-references that note next to
_FORWARD_COMPATIBLE_REVISION.

* fix(chats): invalidate project thread lists on archive/restore

useArchiveThread refreshed the infinite sidebar cache, threads/search and
the per-thread metadata cache but not the project-scoped list
([...PROJECTS_QUERY_KEY, 'threads', id]) this PR adds — the one thread
mutation not wired to that key, after usePinThread, useRenameThread,
useDeleteThread, useMoveThreadToProject and invalidateStoppedThreadCaches.

An archive from a sidebar row while a project page is open therefore left
the archived chat rendered as a normal row until remount (and undo left it
missing). Invalidate the prefix in the mutation-level success handler.

Regression test asserts the project-list prefix is invalidated on success.

* fix(projects): fetch project discovery only in grouped sidebar mode

RecentChatList mounted two useProjects queries per sidebar render, but
knownProjectIds is consumed only by the grouped-mode exclusion filter; in
the default flat mode every page load paid two GET /api/projects?status=
round trips for data nothing read. Gate both queries on grouped mode —
GroupedProjectList fetches the same keys when the toggle is on and
TanStack dedupes the observers.

Also set retry: false on useProject: a deleted or foreign project 404s
deterministically, and the page renders a dedicated not-found state for
it, so the default 1s/2s/4s retry backoff kept deep links in 'loading'
for ~7s before that state appeared. Matches useThreadMetadata /
useThreadTokenUsage.

* fix(threads): fail closed on project-scoped create in memory mode

MemoryThreadMetaStore.create accepted project_id and silently ignored it,
making memory mode the one membership path that fails open: POST
/api/threads with a project id returned 200 and the run started
unassigned, violating the invariant that a run never proceeds outside the
selected project (the SQL store raises ProjectNotAssignableError inside
the insert transaction for the same request).

Raise ProjectNotAssignableError whenever project_id is present so the
router's existing 404 mapping applies, the frontend keeps the composer
text for a retry, and memory mode behaves exactly like SQL mode.
set_project already reports rejection; create now matches it.

Store-level test (raises, nothing persisted, project filter stays empty,
unscoped creates still work) plus a router-level test asserting the 404
and that no row is left behind.

* fix(projects): window the project page thread list

ProjectThreadsSection rendered every loaded page as a plain Link row, so a
long-lived project accumulated unbounded DOM on the page's scroll surface:
each load-more appended another 100 rows and every formatTimeAgo tick
re-rendered the whole list.

Reuse VirtualThreadList (now generic over any row shape with a
thread_id), pointing its scroll parent at this page's ScrollArea viewport
via the shared [data-slot="scroll-area-viewport"] selector used by
/workspace/chats; under the 60-row threshold it falls back to the plain
render, so small projects are unchanged.

* fix(projects): restore row dividers and pin them with a render test

The row class template literal concatenated transition-colors directly
with the conditional border-b token, so non-final rows rendered the
invalid class 'transition-colorsborder-b' and lost both the divider and
the transition. Compose the row classes with cn() and a boolean guard
instead.

The section moved out of page.tsx into a testable component so the row
markup finally has coverage: a DOM test asserts every row except the
final data row carries border-b (index-based, not last: — correct under
virtualization where the last mounted row is not the last data row), and
the untitled fallback plus load-more button render for a partial page.

* fix(projects): validate forward schemas and fence membership reads
2026-09-08 17:00:26 +08:00

71 lines
3.6 KiB
Markdown

# Recovering the original thread-incarnation database revision
The Projects build requires `projects` and `threads_meta.project_id`. An older
deployment may have stamped `0019_thread_incarnations` on a database containing
only `0018_oauth_identity_pg_partial` plus two nullable `VARCHAR(32)` columns:
`threads_meta.incarnation` and `mcp_tasks.thread_incarnation`. That shape cannot
serve this build's repositories. Startup now rejects it without changing the
schema or revision, and reports the missing tables/columns.
Normal databases on this tree's known migration chain upgrade automatically.
The procedure below is only for the exact original incarnation rollout shape.
An incarnation-stamped database that already has all current ORM tables and
columns can still use the audited compatibility exception without re-stamping.
## Offline migration
1. Stop every Gateway, scheduler, and other process writing to the database.
Take a restorable database backup and rehearse these steps on a copy.
2. Verify there is exactly one `alembic_version` row, containing
`0019_thread_incarnations`. Inspect the owning deployment's migration and
actual database schema: it must be the local 0018 schema plus only the two
nullable columns above, with no added defaults, constraints, tables, indexes,
or data backfills. Neither `projects` nor `threads_meta.project_id` may already
exist for this recovery path. If the shape differs, use a migration reviewed
for that deployment; do not use the commands below.
3. From this checkout's `backend/`, set `DEERFLOW_RECOVERY_DATABASE_URL` to the
target async SQLAlchemy URL (`sqlite+aiosqlite:////absolute/path/database.db`
or `postgresql+asyncpg://…`). For Postgres, also set
`DEERFLOW_RECOVERY_POSTGRES_SCHEMA` to the configured application schema, if
one is used. Keep credentials out of shell history.
4. Rebase the version marker to the verified common parent and run the normal
migrations. `purge=True` is necessary because this tree does not contain the
out-of-tree revision; it replaces the version row, not application data.
```bash
uv run python - <<'PY'
import asyncio
import os
from alembic import command
from sqlalchemy.ext.asyncio import create_async_engine
from deerflow.persistence.bootstrap import _get_alembic_config
engine = create_async_engine(os.environ["DEERFLOW_RECOVERY_DATABASE_URL"])
cfg = _get_alembic_config(
engine,
postgres_schema=os.environ.get("DEERFLOW_RECOVERY_POSTGRES_SCHEMA", ""),
)
command.stamp(cfg, "0018_oauth_identity_pg_partial", purge=True)
command.upgrade(cfg, "head")
asyncio.run(engine.dispose())
PY
```
This applies `0019_projects` and `0020_threads_meta_project_id`, preserving
the two incarnation columns and their existing values. Do not stamp directly
to head: that would skip the DDL and reproduce the missing-column failure.
5. Confirm the version is `0020_threads_meta_project_id`, the project table and
membership column/index exist, and existing incarnation values are retained.
Start this build, verify existing conversations load and a new conversation
can be created, then resume service. Do not restart older binaries that
cannot read this tree's head revision.
Bootstrap never performs this re-stamp itself. The regression in
`backend/tests/test_persistence_forward_revision_compat.py` constructs the
original schema, verifies startup rejection, and exercises the recovery while
checking thread reads/inserts and preservation of incarnation data. The future
incarnation migration must chain from the current local head and handle these
already-present nullable columns idempotently.