rayhpeng 06dda5045a fix(schedule): stop a corrupt stored row surfacing as a client error
SqlScheduledTaskRepository._to_domain raised InvalidScheduleError for a
row whose stored schedule no longer parses -- the same error the
aggregate raises for a client-submitted schedule, which the router maps
to 422. A stored fault therefore told the client its perfectly fine
request was wrong, and made the row unrepairable over HTTP: PATCH reads
the task before writing, so the fix path 422'd too. Enum rebuild
failures were worse -- a raw ValueError crossed the boundary
untranslated.

The rebuild is now translated to a dedicated CorruptStoredScheduleError,
raised only by the persistence adapter and deliberately absent from the
router's status table, so it falls through to the unclassified-500
branch: a server-side fault reported as one. List reads keep skipping
and logging the bad row. Pinned by tests/test_schedule_corrupt_rows.py
against a real sqlite database.
2026-07-29 19:46:44 +08:00

68 lines
2.2 KiB
Python

"""The known errors of the schedule context.
One family under one base class, so the primary adapter can map the whole
family onto protocol codes in a single table. Class names keep the PEP 8
``Error`` suffix; the module is named ``exceptions`` after the AWS
hexagonal guidance's domain folder of the same name.
"""
from __future__ import annotations
class ScheduleError(Exception):
"""Base error for the schedule domain."""
class InvalidScheduleError(ScheduleError):
"""Timezone, cron expression, or run_at is not usable."""
class InvalidContextModeError(ScheduleError):
"""context_mode is unknown, or reuse_thread is missing its thread_id."""
class TaskNotFoundError(ScheduleError):
"""The task does not exist or does not belong to the user."""
class TaskNotMutableError(ScheduleError):
"""The task is currently running and cannot be edited."""
class ThreadNotFoundError(ScheduleError):
"""reuse_thread points at a thread the user cannot access."""
class CorruptStoredScheduleError(ScheduleError):
"""A stored task row can no longer be rebuilt into a valid aggregate.
Raised by the persistence adapter, never by the aggregate: it means the
*storage* is damaged, not that a client submitted something invalid --
which is why it is deliberately absent from the router's status table and
falls through to the unclassified-500 branch instead of riding
``InvalidScheduleError``'s 422.
"""
class ActiveRunConflictError(ScheduleError):
"""The task already holds its single active run slot.
Raised by the run repository when the partial unique index
``uq_scheduled_task_run_active`` rejects a second active row. Moved here
from ``persistence/scheduled_task_runs/sql.py`` (was
``ActiveScheduledRunConflict``) so the domain owns its own vocabulary.
"""
class ThreadBusyError(ScheduleError):
"""The execution thread already has an in-flight run.
Translated by the RunLauncher adapter from ConflictError / HTTP 409.
This is what removes `from fastapi import HTTPException` from the
orchestration layer.
"""
class LaunchFailedError(ScheduleError):
"""The run could not be launched for any non-conflict reason."""