rayhpeng 7852421c68 feat(schedule): complete the outer ring with the launch adapters
Adds the three remaining adapters plus the poller. Nothing is wired yet --
the composition root is the next commit -- so this is additive and the
legacy `app/scheduler/service.py` still serves production.

`run_launcher.py` is the pivot of the whole slice. The Gateway signals a
busy thread two ways -- `ConflictError` from the run manager, or an
`HTTPException(409)` from the route-level path -- which is why the legacy
scheduler service imported fastapi to tell them apart. Both are one domain
fact, and saying so here is what lets that import disappear without the
busy/failed distinction disappearing with it. Everything else becomes
`LaunchFailedError`, because the port promises the domain that nothing but
its two errors escapes. `CancelledError` is deliberately not caught:
shutdown is control flow, not a launch outcome.

`thread_lookup.py` narrows `ThreadMetaStore` to the one question this
context asks. `require_existing=True` is load-bearing -- the store's
default treats an absent row as accessible, which is right for a thread
not yet written and wrong for binding a task to it.

Both inherit their port explicitly, matching every other adapter in the
codebase including feedback's own anti-corruption layer, and both carry
the TODO naming the published contract that would replace them once the
upstream context has been through a slice of its own.

`run_outcome_mapping.py` implements no port: it is the inbound translation
the composition root will install on the completion hook, and it owns the
filtering the legacy hook did inline. Returning None means "this run is
none of the schedule context's business", so the service is simply never
called and needs no guard clauses.

`poller.py` keeps the two behaviours the legacy loop got right: a failing
poll must not end the loop (one transient "database is locked" used to
stop scheduling for the rest of the process life), and reconciliation must
not block startup.

One deliberate behaviour change: the legacy `start()` swept stale runs and
stuck once-tasks under separate try/excepts, so the first failing did not
stop the second. `reconcile_on_startup` is one call that lets failures
propagate -- the domain's position is that fatality is the caller's policy
-- so the poller's single except means a failed first sweep now skips the
second. Both end up logged and non-fatal, as before.

Tests: 50 new cases across the four modules, each port method called and
asserted on its return value. That is not decoration: inheriting a
Protocol means a misspelled method silently inherits its `...` body and
returns None, so the suite was verified by mutation -- renaming `launch`
and `exists_for_user` turns 16 and 6 cases red respectively.
2026-07-28 18:28:46 +08:00

90 lines
3.9 KiB
Python

"""Secondary adapter (anti-corruption layer) -- starting a run through Gateway.
Implements ``RunLauncher`` from ``deerflow.domain.schedule.ports``. This context
owns no part of the run lifecycle: it asks the Gateway to start one and
translates whatever comes back into the two outcomes the domain distinguishes.
That translation is the reason this file exists. The Gateway signals a busy
thread two different ways -- ``ConflictError`` from the run manager, or an
``HTTPException(409)`` from the route-level path -- and the legacy scheduler
service therefore imported ``fastapi`` to tell them apart. Both are the same
domain fact, and saying so here is what keeps the web framework and the run
runtime out of the inner ring.
TODO(hexagonal): this depends on ``launch_scheduled_thread_run``, a Gateway
service function returning an untyped dict, rather than on a contract published
by the run context -- that context has not been through a hexagonal slice yet.
When it publishes one (a DTO, not its aggregate and not its repository),
replace the body of this class. The ``RunLauncher`` port does not move.
"""
from __future__ import annotations
from collections.abc import Awaitable, Callable, Mapping
from typing import Any
from fastapi import HTTPException
from deerflow.domain.schedule.model import LaunchFailedError, ThreadBusyError
from deerflow.domain.schedule.ports import LaunchedRun, RunLauncher
from deerflow.runtime import ConflictError
LaunchRun = Callable[..., Awaitable[Mapping[str, Any]]]
class GatewayRunLauncher(RunLauncher):
"""Adapts the Gateway's scheduled-run launch path to the ``RunLauncher`` port.
Takes the launch callable rather than importing it, because the production
one is bound to the FastAPI app (``launch_scheduled_thread_run(app=app,
...)``) and that binding belongs to the composition root.
Explicit inheritance is a readability aid only: a misspelled method would
still instantiate fine and silently inherit the Protocol's ``...`` body,
so the contract tests must call every port method and assert on what it
returns.
"""
def __init__(self, launch_run: LaunchRun) -> None:
self._launch_run = launch_run
async def launch(
self,
*,
thread_id: str,
assistant_id: str | None,
prompt: str,
owner_user_id: str | None,
metadata: dict[str, str],
) -> LaunchedRun:
try:
result = await self._launch_run(
thread_id=thread_id,
assistant_id=assistant_id,
prompt=prompt,
owner_user_id=owner_user_id,
metadata=metadata,
)
except ConflictError as exc:
raise ThreadBusyError(str(exc)) from exc
except HTTPException as exc:
if exc.status_code == 409:
raise ThreadBusyError(str(exc.detail)) from exc
raise LaunchFailedError(str(exc.detail)) from exc
except Exception as exc:
# Deliberately broad: the port promises the domain that nothing but
# its two errors escapes, so an unclassifiable failure has to become
# the "genuine failure" branch rather than unwinding the poll loop.
# `CancelledError` derives from BaseException and is not caught --
# shutdown is control flow, not a launch outcome.
raise LaunchFailedError(str(exc)) from exc
run_id = result.get("run_id")
launched_thread_id = result.get("thread_id")
if not isinstance(run_id, str) or not isinstance(launched_thread_id, str):
# The run path broke its own contract. Reporting it as a failure
# keeps the task's bookkeeping honest instead of recording a launch
# whose run can never be traced.
raise LaunchFailedError(f"run launch returned no usable identity: {result!r}")
return LaunchedRun(run_id=run_id, thread_id=launched_thread_id)