mirror of
https://github.com/bytedance/deer-flow.git
synced 2026-08-01 19:06:01 +00:00
2 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
152e82e25b |
test(schedule): make the domain suite fail when the domain breaks
All 102 domain cases were green, and four of them would have stayed green through the exact regression they were named after. Assertions that could not fail ------------------------------ `test_a_claimed_task_is_marked_running_before_dispatch` asserted the lease was released *after* dispatch -- the opposite of the ordering its name and docstring describe. The claim is what makes a task uneditable while it is being dispatched, so the only place that ordering is observable is inside the launch; the launcher double now reads the repository from there. `test_active_statuses_are_exactly_queued_and_running` restated the constant it was checking, so editing the constant edits the assertion with it. Replaced by `is_active` over all six statuses, which also covers `RUNNING` and the two terminal statuses that had none. `test_reuse_thread_with_an_empty_thread_falls_back_to_a_fresh_one` compared its result against `task.thread_id`, which is `None` on the default task -- it asserted "not None" against a method whose body is `str(uuid.uuid4())`. Now asserts the fresh-thread semantics it is named for: a real uuid, distinct per call. `test_a_task_deleted_mid_flight_is_not_an_error` had no assert at all. That path does have observable behaviour: the hook writes the run record before it reads the task, so a task deleted mid-flight must still leave a finalized record and a freed active slot. Contracts stated in a docstring and nowhere else ------------------------------------------------ - a cron overlap must not leave `last_error` behind (service.py:455 branches on it; only the `once` half was covered, so dropping the branch was free) - a failed launch replaces the launch bookkeeping instead of carrying it over the way a skip does -- which is what `last_run_id=None` in `_fail` means for a task that had already run successfully - `SchedulePolicy`'s defaults are the permissive ones, so a deployment that configured no policy cannot have a business constraint invented for it - transitions leave `updated_at` to the repository, rather than becoming a second source of truth for the same column ACTIVE_RUN_STATUSES' promised assertion --------------------------------------- Its docstring says the check that it stays in lockstep with the partial unique index's predicate "lives in a separate test module rather than the domain tests". It did not exist, in that module or any other. `test_scheduled_task_ models.py` now reads both dialect predicates off `__table_args__` and compares the values it extracts against the constant. The domain suite cannot do this -- it is deliberately dependency-free and cannot import an ORM model -- so the new domain case names where the other half of the rule lives. Removed ------- Three duplicates: an `INTERRUPTED -> CANCELLED` case the parametrize directly above it already made, a trailing-Z case identical in path to the aware-run_at case beside it (the `from_primitives` one is the real one, because it parses a string), and an `ensure_launchable == next_after` case whose value another case already asserts outright. Their reasoning moved into comments where it still applies. The second copy of the tautology, in `test_schedule_fakes.py`, goes with it. Verification ------------ A green run is not evidence for this kind of change, so every new or rewritten assertion was checked by mutation -- break the production rule, confirm the guarding test fails. All nine caught, re-run after `ruff format` to confirm the reformat did not soften any of them. Net +207/-31 across four test files; 425 passed, 3 skipped for the schedule and composition suites. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
4fc08b4f15
|
feat: add scheduled tasks MVP (#3898)
* feat: add scheduled tasks MVP
* fix: harden scheduled task execution semantics
* feat(scheduled-tasks): preset-driven schedule form with timezone and live preview
Replace the raw cron input with a preset Select (hourly/daily/weekly/monthly/custom)
plus structured inputs (time picker, weekday toggles, day-of-month), datetime-local
for one-time tasks, a timezone selector defaulting to the browser timezone, and a
live human-readable preview. Reuses one ScheduledTaskScheduleInput for create and
edit; backend contract unchanged; zero new deps (pure Intl + DST-safe offset helpers).
* feat(scheduled-tasks): full-page i18n + recipe templates + E2E locale pin
Localize the rest of the scheduled-tasks page (filters, detail pane, actions,
edit form, run list, enum values) via t.scheduledTasks.* in en/zh. Add four
built-in recipe templates (GitHub Trending, news digest, issue triage, weekly
report) exposed as a chip row that pre-fills title + prompt + schedule. Pin
Playwright locale to en-US so E2E selectors stay stable against i18n. No backend
change, no new deps.
* fix(scheduled-tasks): idempotent 0003 migration, update head constants, future-date once test
Merge with main surfaced three CI failures:
- 0003_scheduled_tasks create_table collided with legacy test seeds that
build from full metadata; guard with inspector.has_table so the revision
no-ops when the table already exists (0004/0005 are already idempotent via
_helpers.py).
- persistence bootstrap concurrency/regression tests pinned HEAD to main's
0002_runs_token_usage; bump to the new head 0005_scheduled_task_thread_nullable.
- once-task router test used a fixed past run_at and tripped the
must-be-in-the-future validation; use a future date.
* address review: ok-check, 502 for trigger failure, mock fields, migration filename, doc fences
- fetchThreadScheduledTasks now checks response.ok like the other fetchers.
- trigger endpoint returns 502 (not 409) when dispatch fails outright, so
clients can distinguish a real conflict from a server-side failure.
- E2E mock normalizes scheduled-task objects with context_mode/last_thread_id
and nullable thread_id, matching the backend contract the UI renders against.
- Rename 0002_scheduled_tasks.py -> 0003_scheduled_tasks.py to match its
revision id (file was renamed in spirit already; filename now follows).
- CONFIGURATION.md: close the Tool Groups yaml fence and drop the stray fence
after the Scheduler notes so the sections render correctly.
* fix(scheduled-tasks): harden lease, poller, config, and frontend UX after review
* fix(scheduled-tasks): harden run lifecycle, overlap skip, non_interactive gating, and DST conversion after review
- defer a once task's terminal status to the run-completion hook; the task
stays running until the real outcome, and a startup sweep cancels once
tasks orphaned by a crash (launch-time 'completed' could stick forever)
- record interrupted runs as a distinct 'interrupted' run status with a
readable message; an interrupted once task ends 'cancelled', not 'failed'
- enforce overlap_policy=skip for fresh_thread_per_run via an active-run
pre-check (same-thread ConflictError can never fire across fresh threads)
- protect terminal run statuses from the late launch-path 'running' write
- honor context.non_interactive only for internally-authenticated callers;
arbitrary clients can no longer strip ask_clarification
- fix DST-stale timezone offset in zonedLocalToUtcIso by re-deriving the
offset at the resolved instant (once tasks fired an hour late around
spring-forward and the create->edit round-trip diverged)
- drop dead ScheduledTaskRunRepository.update_by_run_id; share one Gateway
API error helper between channels and scheduled-tasks frontends
* fix(scheduled-tasks): close review round-3 gaps in guards, concurrency, and API ergonomics
- scrub internal-only context keys (non_interactive) from the assembled run
config for non-internal callers: gating body.context alone left the same
key smuggle-able through the free-form body.config copied verbatim by
build_run_config
- guard update_after_launch with protect_terminal so the launch bookkeeping
write cannot clobber a once task already finalized by a fast-failing run's
completion hook (parent-row sibling of the run-row guard)
- reject a manual trigger while the task has an active run (409) instead of
launching a duplicate concurrent run on fresh_thread_per_run
- re-arm a terminal once task to enabled when PATCH pushes run_at into the
future; previously the endpoint returned 200 with a next_run_at that could
never be claimed
- make max_concurrent_runs a real global cap: each poll claims only into the
remaining budget of active (queued/running) scheduled runs
- paginate GET /scheduled-tasks/{id}/runs (limit<=200, offset) and push the
thread filter of /threads/{id}/scheduled-tasks into SQL
- stamp context.user_id on scheduler-launched runs, matching IM channels, so
user-scoped guardrail providers see the owning user
---------
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
|