deer-flow/backend/tests/test_multi_worker_postgres_gate.py
heart-scalpel 3bc3af2530
fix(runs): close multi-worker ownership gaps in run atomicity (#3948) (#4003)
* feat(runs): cross-process run ownership with lease + reconciliation (#3948)

Implements work items 2 and 3 of the multi-worker P0 plan
(docs/multi_worker.md). Work item 1 (Postgres startup gate, #3960)
already landed; this PR makes run creation race-safe across worker
processes and lets Postgres deployments recover orphaned inflight runs
from crashed workers without mis-marking live runs as orphans.

Work item 2 — cross-process atomic create_or_reject

- Alembic revision 0004_run_ownership adds runs.owner_worker_id,
  runs.lease_expires_at, idx_runs_lease, and a partial unique index
  uq_runs_thread_active (one pending/running run per thread). The
  index is declared on RunRow.__table_args__ with sqlite_where +
  postgresql_where (mirroring uq_channel_connection_active_identity)
  so the empty-DB bootstrap path — which runs Base.metadata.create_all
  + alembic stamp head without executing any revision's upgrade() —
  also lands it on fresh deployments. Migration 0004 additionally
  creates it idempotently for legacy/versioned upgrades.
- RunRepository.create_run_atomic is the new atomic primitive:
  - reject: INSERT directly; the partial unique index catches
    duplicate active runs; the manager surfaces the result as
    ConflictError.
  - interrupt/rollback: SELECT FOR UPDATE the conflicting rows,
    skip rows whose lease is still valid AND owned by another live
    worker (raise ConflictError — the INSERT would have failed on
    the index anyway, and a retry loop cannot make progress),
    cancel the rest in the same transaction, then INSERT the new
    row. Rows owned by this worker are interruptible regardless of
    lease state.
- RunManager.create_or_reject dispatches to the store under the
  existing local lock; same-worker in-memory cancellation runs after
  the store commit succeeds. MemoryRunStore mirrors the same
  semantics for tests and database.backend=memory.

Work item 3 — lease heartbeat + Postgres reconciliation

- RunOwnershipConfig (lease_seconds=30, grace_seconds=10,
  heartbeat_enabled=false by default), registered as startup-only in
  reload_boundary.STARTUP_ONLY_FIELDS because the heartbeat background
  task is created once in langgraph_runtime() and is not rebuilt on
  config.yaml edits.
- When heartbeat_enabled, each worker renewes leases on its own
  active runs with interval = lease_seconds / 3. The loop is bounded
  and stop-event-cancellable so shutdown is prompt.
- reconcile_orphaned_inflight_runs now runs on every backend — the
  sqlite-only gate in app/gateway/deps.py is dropped in the same
  commit so there is no window where Postgres would mis-mark live
  Worker A runs as orphans. Reconciliation errors only runs whose
  lease is NULL (legacy pre-ownership rows) or older than
  grace_seconds. In single-worker mode (heartbeat off, NULL leases)
  all inflight rows reclaim immediately, preserving the pre-ownership
  recovery latency.
- Heartbeat starts AFTER startup reconciliation and stops BEFORE the
  in-flight run drain on shutdown so the two cannot race.

GATEWAY_WORKERS=1 with heartbeat_enabled=false keeps current behavior.

Verified: 170 related tests + full backend suite (minus Docker-gated
live tests) green; ruff check + ruff format clean.

* fix(runs): tighten unique-violation handling and document clock-sync budget

Three follow-up fixes to the cross-process run ownership work in #3948,
surfacing during review.

1. _is_unique_violation: detect by driver-native signal, not message text

   The previous substring heuristic ("unique" + "violat", or "duplicate")
   missed SQLite's actual phrasing "UNIQUE constraint failed: <table>.<index>"
   — SQLite says "failed", not "violates", and never "duplicate". On SQLite
   the detector returned False, the reject path re-raised the raw
   IntegrityError, and clients saw HTTP 500 instead of ConflictError 409.
   The conversion is the load-bearing piece of the "store is source of
   truth" design but was untested — every atomic test used MemoryRunStore,
   which raises ConflictError directly and never reached this branch.

   Now prefers driver-native signals: psycopg pgcode/sqlcode "23505" and
   sqlite3 sqlite_errorcode SQLITE_CONSTRAINT_UNIQUE (reachable through
   SQLAlchemy IntegrityError.orig). Message matching stays as a fallback
   with SQLite's exact "unique constraint failed" phrase added.

2. interrupt/rollback: convert exhausted-retry IntegrityError to ConflictError

   The reject branch converts unique violations to ConflictError. The
   interrupt/rollback retry loop did not — on the 3rd attempt it re-raised
   the raw IntegrityError, leaking HTTP 500 for the same race condition
   that reject surfaces as 409. Symmetric conversion added after the loop;
   callers now see a consistent ConflictError regardless of strategy.

3. Document clock-sync requirement for multi-worker lease reconciliation

   reconcile_orphaned_inflight_runs compares another worker's UTC
   lease_expires_at against this worker's datetime.now(UTC). The only skew
   budget is grace_seconds (default 10s) — worst case, with the owning
   worker's heartbeat just about to fire, a peer whose clock is more than
   ~grace_seconds ahead can mis-reclaim a still-live run as an orphan.

   Documented in RunOwnershipConfig's docstring (with the math) and in
   config.example.yaml (with operational guidance), so operators in
   NTP-poor environments know to raise grace_seconds. Default unchanged:
   10s is reasonable for NTP-synced K8s/cloud, and bumping it would slow
   recovery of genuinely dead workers (lease_seconds + grace_seconds from
   last heartbeat to reclaim).

Tests:
- test_create_run_atomic_reject_propagates_conflict_on_unique_violation:
  end-to-end against a real SQLite-backed RunRepository, pre-inserts an
  active run, asserts reject-strategy create surfaces as ConflictError
  rather than raw IntegrityError.
- test_is_unique_violation_detects_real_sqlite_integrity_error: unit test
  for the detector against a real SQLite-raised IntegrityError; asserts
  driver-level sqlite_errorcode is SQLITE_CONSTRAINT_UNIQUE.
- test_interrupt_exhausted_retries_surface_as_conflict_error: pins the
  symmetric 409 behavior after the retry loop exhausts.

Verified: ruff check + ruff format clean; multi-worker + run_repository
+ owner_isolation + reload_boundary suites green.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(runs): close multi-worker ownership gaps in lease heartbeat and unique-violation detection

Five code-review fixes from docs/multi_worker.md:

1. Drop unused ``claim_inflight_runs`` primitive — no caller anywhere.
   ``create_run_atomic`` does its own inline claim (SELECT FOR UPDATE +
   cancel) inside the INSERT transaction; a separate claim primitive
   would split that into two transactions and open a claim→INSERT race.
   Removes ~40 lines across base.py / memory.py / sql.py plus the
   unused ``now_iso`` parameter, freeing future RunStore implementations
   from providing it.

2. Broaden ``_renew_leases`` filter to renew pending/running runs owned
   by this worker even when ``record.task is None``. The previous
   ``task is not None`` requirement skipped the brief window between
   ``create_run_atomic`` inserting the row and the worker spawning the
   agent task; under event-loop load that window can approach
   ``lease_seconds``, after which peer reconciliation marks the run
   ``error`` (visible) or a peer's ``create_or_reject("interrupt")``
   silently kills the queued run. Filter now:
   ``task is None or not task.done()``.

3. Document the unsynchronised ``record.lease_expires_at = new_expiry``
   write. ``lease_expires_at`` is the only field on an existing record
   this path mutates; ``set_status`` / ``_persist_status`` touch other
   fields, so there is no concurrent writer to race against. Re-acquiring
   ``self._lock`` would serialise unrelated run mutations for no gain.

4. Gate ``_is_unique_violation`` message fallbacks on
   ``isinstance(current, (SAIntegrityError, sqlite3.IntegrityError))``.
   The driver-code path (pgcode/sqlite_errorcode) remains load-bearing;
   substring fallbacks are now belt-and-suspenders only for cases where
   the driver attribute isn't reachable through the cause chain. Without
   the gate, any application exception whose ``str()`` happens to contain
   "duplicate key" / "unique" + "violat" (CHECK constraint, validation
   error) would silently surface as HTTP 409 instead of 500.

5. Route ``update_lease`` through ``_call_store_with_retry`` for
   consistency with every other store call, and wrap
   ``await self._renew_leases()`` in ``_heartbeat_loop`` with
   ``except Exception: logger.warning(...)``. Previously a transient
   error from the snapshot path or an unexpected exception would kill
   the heartbeat task silently — after which no lease is ever renewed
   again and every active run eventually looks orphaned.
   ``except Exception`` lets ``CancelledError`` (BaseException since
   3.8) propagate so shutdown cancellation still works.

Regression tests:
- ``test_heartbeat_renews_pending_run_before_task_is_spawned``
- ``test_is_unique_violation_does_not_misclassify_application_exception``

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(runs): harden multi-worker migration, memory atomicity, and tz-naive lease comparison

Three follow-up fixes to the multi-worker run ownership work:

- migration 0004 dedupe pass: cancel superseded duplicate active rows per
  thread before creating the partial UNIQUE index ``uq_runs_thread_active``
  so dirty DBs (Postgres deployments that had reconciliation skipped by the
  old sqlite-only gate, or any env that ran GATEWAY_WORKERS>1 before this PR)
  do not abort the alembic upgrade and block gateway startup. Keeps the
  newest active row per thread, marks the rest as error with an explanatory
  message.

- MemoryRunStore.create_run_atomic interrupt/rollback path: split the single-
  pass loop into two passes (collect candidates, validate, then mutate) so a
  ConflictError raised on a later candidate does not leave earlier candidates
  half-interrupted. Mirrors the SQL store's transactional rollback semantics;
  the entire test_multi_worker_run_ownership.py suite runs against memory so
  this divergence was giving false confidence.

- RunRepository.create_run_atomic interrupt path: coerce tz-naive
  ``row.lease_expires_at`` to UTC before comparing against the aware
  ``cutoff``. SQLite drops tzinfo on read despite ``DateTime(timezone=True)``
  (this file's own comment acknowledges it), so the Python-side comparison
  raised ``TypeError: can't compare offset-naive and offset-aware datetimes``
  whenever heartbeat was enabled on SQLite and a lease was non-NULL. Defaults
  (heartbeat off -> leases always NULL) masked it, but there was no guard
  against the combination. Follows the existing "naive is UTC" convention
  from ``coerce_iso``.

Each fix ships with a regression test pinning the behavior.

Co-Authored-By: heart-scalpel <heart-scalpel@users.noreply.github.com>
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(runs): enforce heartbeat for multi-worker, fix memory-store datetime comparison, lazy-import ConflictError in store layer

Three fixes from code review:

1. Extend the startup gate (GATEWAY_WORKERS>1) to also require
   run_ownership.heartbeat_enabled=true. Without heartbeat every run has
   a NULL lease, so reconciliation treats all inflight rows as orphans
   and Worker B would kill Worker A's live runs on every rolling update
   or scale-up.

2. Fix MemoryRunStore.list_inflight_with_expired_lease to parse
   created_at as datetime instead of ISO string lexical comparison,
   and handle tz-naive lease values uniformly with the SQL store.

3. Store layer (sql.py, memory.py) now lazy-imports ConflictError
   inside create_run_atomic instead of importing from the higher
   RunManager layer at module level.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(runs): add owner check to update_lease, document create() assumption, restore deleted comment

- update_lease (SQL + memory) now requires owner_worker_id match in WHERE
  clause so the primitive is safe by construction against misuse
- create() docstring notes it bypasses atomic create_run_atomic and
  assumes no active run exists for the thread
- restore explanatory comment in MemoryRunStore.aggregate_tokens_by_thread
  that was dropped in an earlier commit

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(runs): add psycopg3 sqlstate detection and periodic orphan reconciliation

- _is_unique_violation now checks sqlstate attribute (psycopg3 uses this
  instead of pgcode). On Postgres, the only supported multi-worker backend,
  detection was falling through to the message-substring fallback.
- _heartbeat_loop now runs reconcile_orphaned_inflight_runs every 3rd
  cycle (every lease_seconds) to catch orphans whose lease expires between
  pod restarts. Single-worker deployments are unaffected (heartbeat off).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
Co-authored-by: heart-scalpel <heart-scalpel@users.noreply.github.com>
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
2026-07-11 16:05:30 +08:00

181 lines
7.4 KiB
Python

"""Tests for the multi-worker Postgres startup gate.
Pins the contract documented in ``docs/multi_worker.md`` work item 1
(issue #3948): when ``GATEWAY_WORKERS > 1`` and the configured
database backend is not Postgres, the Gateway must refuse to start.
The gate runs inside :func:`langgraph_runtime` *before* any
persistence engine is initialised so operators see a clear error
instead of intermittent SQLite ``database is locked`` failures in
production.
"""
from __future__ import annotations
from contextlib import asynccontextmanager
from types import SimpleNamespace
from unittest.mock import AsyncMock, MagicMock, patch
import pytest
from fastapi import FastAPI
from app.gateway.deps import _enforce_postgres_for_multi_worker, langgraph_runtime
from deerflow.config.database_config import DatabaseConfig
from deerflow.config.run_ownership_config import RunOwnershipConfig
def _config_with_backend(backend: str, *, heartbeat_enabled: bool | None = None) -> SimpleNamespace:
run_ownership = RunOwnershipConfig(heartbeat_enabled=heartbeat_enabled) if heartbeat_enabled is not None else None
return SimpleNamespace(database=DatabaseConfig(backend=backend), run_ownership=run_ownership)
# ---------------------------------------------------------------------------
# Unit tests of the gate function itself
# ---------------------------------------------------------------------------
def test_gate_noop_when_gateway_workers_unset(monkeypatch):
"""With GATEWAY_WORKERS unset, every backend must be accepted."""
monkeypatch.delenv("GATEWAY_WORKERS", raising=False)
for backend in ("sqlite", "memory", "postgres"):
_enforce_postgres_for_multi_worker(_config_with_backend(backend))
def test_gate_noop_for_single_worker(monkeypatch):
"""GATEWAY_WORKERS=1 preserves the historical single-worker behavior."""
monkeypatch.setenv("GATEWAY_WORKERS", "1")
for backend in ("sqlite", "memory", "postgres"):
_enforce_postgres_for_multi_worker(_config_with_backend(backend))
def test_gate_allows_multi_worker_with_postgres_and_heartbeat(monkeypatch):
monkeypatch.setenv("GATEWAY_WORKERS", "2")
_enforce_postgres_for_multi_worker(_config_with_backend("postgres", heartbeat_enabled=True))
def test_gate_rejects_multi_worker_with_sqlite(monkeypatch):
monkeypatch.setenv("GATEWAY_WORKERS", "2")
with pytest.raises(SystemExit) as exc_info:
_enforce_postgres_for_multi_worker(_config_with_backend("sqlite"))
msg = str(exc_info.value)
assert "GATEWAY_WORKERS=2" in msg
assert "postgres" in msg.lower()
assert "sqlite" in msg.lower()
def test_gate_rejects_multi_worker_with_memory(monkeypatch):
"""The gate is not sqlite-specific: memory is also unsafe across processes."""
monkeypatch.setenv("GATEWAY_WORKERS", "2")
with pytest.raises(SystemExit):
_enforce_postgres_for_multi_worker(_config_with_backend("memory"))
def test_gate_rejects_high_worker_counts(monkeypatch):
"""The threshold is >1, not ==2; prod-scale counts must also be gated."""
monkeypatch.setenv("GATEWAY_WORKERS", "4")
with pytest.raises(SystemExit) as exc_info:
_enforce_postgres_for_multi_worker(_config_with_backend("sqlite"))
assert "GATEWAY_WORKERS=4" in str(exc_info.value)
def test_gate_treats_invalid_env_as_single_worker(monkeypatch):
"""Non-integer GATEWAY_WORKERS values must not crash startup.
Uvicorn itself rejects these later; the gate should not preempt
that with its own crash. Falling back to 1 keeps the gate inert.
"""
for invalid in ("", "auto", "1.5", "abc", "0x4"):
monkeypatch.setenv("GATEWAY_WORKERS", invalid)
_enforce_postgres_for_multi_worker(_config_with_backend("sqlite"))
def test_gate_treats_zero_and_negatives_as_single_worker(monkeypatch):
"""GATEWAY_WORKERS <= 1 (including 0 and negatives) skips the gate."""
for value in ("0", "-1", "-999"):
monkeypatch.setenv("GATEWAY_WORKERS", value)
_enforce_postgres_for_multi_worker(_config_with_backend("sqlite"))
def test_gate_error_message_lists_both_remediations(monkeypatch):
"""Operators must see both fix options without reading docs."""
monkeypatch.setenv("GATEWAY_WORKERS", "2")
with pytest.raises(SystemExit) as exc_info:
_enforce_postgres_for_multi_worker(_config_with_backend("sqlite"))
msg = str(exc_info.value)
assert "GATEWAY_WORKERS=1" in msg, "must mention the rollback knob"
assert "Postgres" in msg, "must mention the alternative backend"
# ---------------------------------------------------------------------------
# Heartbeat enforcement: multi-worker requires heartbeat_enabled=true
# ---------------------------------------------------------------------------
def test_gate_rejects_multi_worker_without_heartbeat(monkeypatch):
monkeypatch.setenv("GATEWAY_WORKERS", "2")
with pytest.raises(SystemExit) as exc_info:
_enforce_postgres_for_multi_worker(_config_with_backend("postgres", heartbeat_enabled=False))
msg = str(exc_info.value)
assert "heartbeat_enabled=true" in msg
def test_gate_rejects_multi_worker_without_run_ownership_config(monkeypatch):
monkeypatch.setenv("GATEWAY_WORKERS", "2")
with pytest.raises(SystemExit) as exc_info:
_enforce_postgres_for_multi_worker(_config_with_backend("postgres", heartbeat_enabled=None))
msg = str(exc_info.value)
assert "heartbeat_enabled=true" in msg
def test_gate_heartbeat_check_not_triggered_for_single_worker(monkeypatch):
"""GATEWAY_WORKERS=1 skips the heartbeat check entirely."""
monkeypatch.setenv("GATEWAY_WORKERS", "1")
_enforce_postgres_for_multi_worker(_config_with_backend("postgres", heartbeat_enabled=False))
def test_gate_heartbeat_check_not_triggered_for_sqlite(monkeypatch):
"""The gate exits on Postgres check before reaching heartbeat check."""
monkeypatch.setenv("GATEWAY_WORKERS", "2")
with pytest.raises(SystemExit) as exc_info:
_enforce_postgres_for_multi_worker(_config_with_backend("sqlite", heartbeat_enabled=True))
msg = str(exc_info.value)
assert "postgres" in msg.lower()
# ---------------------------------------------------------------------------
# Integration: the gate is wired into langgraph_runtime before init_engine
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_langgraph_runtime_invokes_gate_before_persistence_setup(monkeypatch):
"""When the gate trips, no persistence / stream-bridge setup may run.
Guards against regressions that reorder the gate behind
``init_engine_from_config`` (or any other expensive startup step).
"""
monkeypatch.setenv("GATEWAY_WORKERS", "2")
init_engine_from_config = AsyncMock(name="init_engine_from_config")
@asynccontextmanager
async def _noop_stream_bridge(_config):
yield MagicMock()
with (
patch(
"deerflow.persistence.engine.init_engine_from_config",
init_engine_from_config,
),
patch("deerflow.runtime.make_stream_bridge", side_effect=_noop_stream_bridge) as make_stream_bridge,
patch("deerflow.runtime.make_store", side_effect=_noop_stream_bridge) as make_store,
):
app = FastAPI()
startup_config = _config_with_backend("sqlite")
with pytest.raises(SystemExit):
async with langgraph_runtime(app, startup_config):
pass
init_engine_from_config.assert_not_called()
make_stream_bridge.assert_not_called()
make_store.assert_not_called()