now-ing 2a96341889
sandbox: claim ownership before readiness-timeout destroy (#4248) (#4505)
* fix(sandbox): claim ownership before readiness-timeout destroy (#4248)

When a freshly-created sandbox fails to become ready within the 60s timeout, _create_sandbox / _create_sandbox_async destroyed the container with a bare self._backend.destroy(info) call. Ownership is published by _register_created_sandbox only after the readiness gate, so for the whole timeout window the container ran unowned — exactly the state a peer gateway startup reconciliation is built to adopt across. With the claim skipped, a peer could adopt the not-yet-ready Pod and this instance subsequent stop landed on the turn the peer had just handed it: a cross-instance kill of an active turn.

Route the destroy through a new _destroy_unready_sandbox helper that first claims the teardown lease (claim(..., for_destroy=True) writes the del: marker) and holds it via _held_teardown_lease for the duration of the stop, matching the ownership guard every other reap path already uses (_destroy_warm_entry, _drop_unhealthy_reserved). Fail closed if a peer already holds the lease or the ownership store cannot answer: leave the container for the peer own reconciliation rather than stopping it from underneath an active turn.

Fixes #4248.

* fix(sandbox): reserve local teardown around readiness-timeout destroy

Review feedback (AnnaSuSu) on #4505: the ownership claim is only the
cross-instance half of the guard — claim() succeeds against our own
lease by design, so in the window between the readiness timeout and the
claim a same-process _reconcile_orphans (idle checker, every 60s) can
adopt the unready container into _warm_pool; the subsequent claim still
succeeds and the stop lands on an entry this instance has just adopted,
leaving a dead warm entry for the next reclaim to hand out.

Wrap the still-untracked check, claim, and stop in
_reserve_local_teardown / _finish_local_teardown, with the predicate
checking the id is absent from the active and warm maps — the same
pairing _destroy_warm_entry already uses.

Add an interleaving regression test mirroring
test_reconcile_does_not_adopt_a_container_this_instance_is_tearing_down:
reconcile runs while the destroy thread is parked after reserving but
before its `del:` claim lands, and must not adopt. A mirror test
asserts a genuinely unowned container is still adopted when no teardown
is in flight, so the reservation cannot over-block reconciliation.

* style(sandbox): apply ruff format to readiness-timeout destroy changes

---------

Co-authored-by: now-ing <24534365+now-ing@users.noreply.github.com>
2026-07-30 07:12:07 +08:00
..